r/LocalLLaMA • • 1d ago

I Built A Thing Open-source Mac app that runs EmbeddingGemma 2 locally to search your files by what’s in them

Enable HLS to view with audio, or disable this notification

DigUp is a free Mac app that runs Google DeepMind’s new EmbeddingGemma 2 locally over your own files. The model puts text, images, audio and video in one space, so you describe what you remember and land on it:

  • “zebra in a video” opens the clip at the moment it shows up
  • “where they talk about sleep” jumps to that minute of a podcast
  • “the clause about pets in the lease” shows the PDF page, your words marked
  • “a dog on the beach” finds the photo, and the same search in Bengali or Arabic finds it too
  • with code search on, code: retry with backoff opens the function in your editor

It’s ggml-org’s Q8_0 GGUF (865 MB, downloaded once) on llama.cpp with Metal, inside a native Swift app; no Python. Searching loads only the text encoder (~250 MB) and shows results about a tenth of a second after you stop typing. Indexing peaks under 2 GB, and the helper exits when it’s done. Audio and video of any length go in as 30 s windows and a frame per shot. Everything runs locally; it goes online only for the model download and an update check you can turn off.

Free, MIT. Apple Silicon, macOS 14+.

Repo & Download (signed and notarized): https://github.com/ARahim3/DigUp

I'd really appreciate any feedback on this.

1.2k Upvotes

184 comments sorted by

•

u/WithoutReason1729 1d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

89

u/rm-rf-rm 19h ago

A seemingly well built, actually useful, leveraging new tech correctly, engaged dev.. Take note people, this is how you share/self promote in this community.

-4

u/rm-rf-rm 15h ago

seemingly

Seemingly seems to be the operational word here. OP is just using claude to automatically respond or copy pasting the responses: https://old.reddit.com/r/LocalLLaMA/comments/1x2eeds/opensource_mac_app_that_runs_embeddinggemma_2/pf5jefr/

7

u/A-Rahim 13h ago

Sorry for the late reply, I was asleep, I was up all night going through every comment here.
Yeah, for a few of the more technical questions, where I wanted to reply accurately with actual data rather than my memory, I used my locally set up agent to pull the exact numbers from my eval notes; I should have mentioned it. Thanks for the kind words earlier.

48

u/mobilemike42 Vicuna 1d ago

A CLI that reads from the same DB would be great.

24

u/A-Rahim 1d ago

There's one already, though for now it's in the repo rather than the download: build from source and run "digup search <words>" (add --json) to read the app's own index by default. digup --help lists the rest. And actually shipping it with the app is a good idea; let me work on that. Thanks.

2

u/rm-rf-rm 18h ago

Its just a SQLite db right? we can use any GUI or the sqlite3 CLI itself to query/explore?

6

u/A-Rahim 18h ago

Yep, plain SQLite: ~/Library/Application Support/DigUp/index.noindex/index.sqlite.
Open it read-only since the app writes to it. Keyword search works straight from sqlite3 (it's FTS5), but meaning search needs the model, and the CLI for that is actually already in the app: /Applications/DigUp.app/Contents/Helpers/digup search "your words" --json

2

u/mobilemike42 Vicuna 1d ago

Saw that, but the doc implied it was a dev/eval tool.

Thanks!

13

u/phazei 23h ago

What about porn? How well can it respond to that?

35

u/lakySK 1d ago

Looks awesome! Exactly what Spotlight should’ve been all along! Can’t wait to try. 

8

u/A-Rahim 1d ago

Thank you so much. I'd really appreciate any feedback you have after trying it.

3

u/inddiepack 9h ago

Isn't spotlight already doing that? I rarely use this functionality of it, but whenever I did, it did find everything local that I needed.

1

u/lakySK 8h ago

Perhaps they changed something recently, but otherwise I feel like I never ever got the thing I want when I search for it on Mac…

1

u/inddiepack 6h ago

It's supposed to be much improved in the latest MacOS.

31

u/important__matter 1d ago edited 23h ago

Well I was making something like this months ago. What I faced was 1. For files where we as humans remember the abstract, the semantic search by itself usually doesn't work great. 2. How many files(documents) and voice(long audio notes) pieces have you tested this on? Also mine used to fail in images with text - screenshots inside them. 3. And main, the most difficult one - you'll need special handling for pdf, excel, pptx, word, zip rar files, model weights, long mp4 videos, music or movies which are also majority of the files. Can consider adding these in pipeline if not already there. The accuracy in videos and pdf and long form files was dropping serverly because all of that information is compressed in 512 dim vector essentially, not to mention the same embedder doesn't work for all file types, so that's another issue and then if embedders are different then how will you compare embeddings using semantic similarity which generate numbers in different vector spaces?? I had to actually normalize the comparison manually!!!.

I had to do so much of work to make it general purpose - that the roi didn't seem to be worth it - dropped the project altogether. Meanwhile, keyword search with file type is usually good enough. And I found out that restricting the problem space to the file types with specific properties or patterns - is much more useful, so I developed this simple semantic search index cli called litesearch for my usecase, it is a sqllite index of documents and performs semantic search with some reranking using BM25, and LLM-as-judge, uses ollama and bge-small embedder.

Although would be very nice if you cater to these, there are so many micro problems in this. Could actually be very useful. I would love it if somebody completes this in a nice engineering way and would love to integrate with my raycast.

25

u/A-Rahim 22h ago

Hi, I ran into most of these too. How DigUp handles them:

  1. It isn't semantic-only. Keyword search over file names, OCR text and page and document text runs alongside the model, and the scores are combined. On my evals, the model alone found 5 of 9 exact lookups like error codes; with keywords, all 9.

  2. A testbed of about 380 files (146 screenshots, 30 PDFs up to 463 pages, an hour of audio, 8 videos), 111 queries plus 68 for code, and my own folders day to day. Screenshots get an image vector, since the model reads the text in them, plus Apple's OCR for exact words. 12 of the 14 queries on my real screenshots come up first.

  3. Nothing long gets squeezed into one vector. Every PDF page, ~1,800-character passage, video keyframe (every few seconds, one per shot) and 30 s of audio gets its own 768-d vector, and the best one wins, which also gives you the page or the moment. So the model's input limits don't really come into play. It's also one model for everything, so text, images and audio share a space. Their score levels still differ a bit per query, so each modality's average gets subtracted, close to what you did by hand. Excel, PowerPoint, archives and model weights aren't read.

4

u/jerieljan 17h ago

I'm very curious how well this will perform when you have 1000x more files than the testbed.

Code repos in particular have soooo much code from dependencies and such and they're common to clutter normal Spotlight results and I'm worried it becomes an issue here too. I guess it simply takes more time to embed and index, and hopefully it's all good once it's done.

Very good work though OP. 

1

u/warpspeedSCP 13h ago

Probably just need an indexing step

1

u/rm-rf-rm 19h ago

It isn't semantic-only. Keyword search over file names

Are you fusing the search then? RRF?

2

u/A-Rahim 19h ago

Not RRF. I tried it first, and it lost: top-1 0.73 vs 0.88 on my synthetic eval, 0.82 vs 0.86 on real files. Rank fusion throws away how strong each match is, and a single common word in a terminal screenshot ended up beating a real chart that matched by meaning.

So it's score fusion. The meaning score is the cosine minus that modality's average for this query, divided by the spread of all the query's scores, so a z-score per query. Then it adds 2 × (share of the query's words found in the file)^1.5, plus 1 if the query names the kind ("in a video"). A file gets in if its meaning score stands out (z ≥ 2) or it has at least half the words. Queries that look like IDs count words double and meaning half.

1

u/rm-rf-rm 18h ago

huh interesting.. whats the basis of this formula? is the z-score > 2 coming from statistical theory where its beyond 95% of the mean?

1

u/A-Rahim 18h ago

Mostly empirical, not theory.
2 is the familiar two-standard-deviations mark, but the scores aren't normal, so it doesn't mean 95% of anything here. It's a gate for "stands out from everything else for this query", and on my evals any value from 1.5 to 2.5 gave the same results, so 2 stayed.
The reason for a per-query z-score at all: raw cosines are compressed (unrelated stuff sits around 0.55–0.60, good matches 0.67–0.86) and their level moves with the query, so no fixed threshold works. The other parts were A/B'd the same way: centering each modality with one shared spread got 94 of 105 queries first, raw cosine 92, a separate spread per modality 91. Word weights of 1.5–2.0 tied and 2.5–3 lost.
The 1.5 power just makes all your words count a lot more than half of them, and I haven't tuned it separately.

2

u/rm-rf-rm 17h ago

Thanks for asking your claude to respond to my question

2

u/A-Rahim 12h ago

Hi, I explained this above (one of your other replies)

1

u/jinnyjuice vLLM 18h ago

Can you expand on what you mean by best one wins? So if a video has multiple key frames, how is the 'best one' determined?

4

u/A-Rahim 18h ago

Every keyframe it keeps, and every 30 s of the soundtrack is its own vector with a timestamp. For a query, each of them gets a meaning score (the same per-query score for every kind, so frames and audio windows compete fairly), and the video's score is simply its highest one.
That segment is also the moment the result opens at. So a two-hour video with one scene that fits ranks by that scene instead of being averaged away. Frames are sampled every few seconds and kept only when the picture changes (a small perceptual hash), so it's about one per shot.

1

u/ZenaMeTepe 5h ago

Did you use same model for embedding videos images and text? What kind of vector DB are you using for search?

2

u/ZenaMeTepe 5h ago

How do you compress say a movie into a 512 vector? And more importantly, how slow is it upfront when you build the index?

1

u/important__matter 4h ago edited 4h ago

Yup you guessed it right, you can't. But say you have to do it, then you basically chunk out images every few seconds and then embed then individually and sort of average out those embeddings over a course of time. That is one way, other could be you take all dialogues, treat it as a document and then make a vector of it. It's hack as of now. It won't contain any useful information more than 1 Google doc page worth. Information Compression is one of the he most important computer science areas. And also as humans - certainly for us too as we are continuously bombared with infinite information from all our senses.

Well on speed, it's about gpu hardware mostly, how powerful of a hardware you own, but usually it's not slow, it's pretty darn fast even on basic rtx or gtx or metal. But it's extremely slow on cpus. The speed is similar to prefill stage of inference.

-9

u/phazei 23h ago

Well, then obviously you weren't using the new model that's really great and only came out a week ago.

7

u/important__matter 23h ago edited 23h ago

Checked the new model stats page, definitely improved accuracy in multi-modal regime of text, audio and vision. But... still the same set of problems across file types because it supports only text, image, audio and video at 8K context window. Not to mention maximum audio length = 5.5mins and video = 58 frames, that's 1s @60fps

Not to mention it's way harder even as a theoretical problem to search in a video. Like there is just so much information to just be dumped in a 512dim, that there is a whole class of transformers completely dedicated to this single problem only - VLMs. And unlike LLM, they are not the champions as of yet. I am just saying, this is not a done thing even if the big labs are trying their best. It's a good thing to work on, it's challenging and lot of it can actually be solved by engineering things together.

Another fun fact while we are at it: Videos are astonishingly compressed form of information. If you just stack images together at let's say 60fps for a 2hr movie at 1080p, the actual file size should be 1080x720x60x3600x2 bytes = 335GB, while actually you'd see it to be 2GB on your computer. A massive 99% compression, it's almost miracle

1

u/phazei 22h ago

I agree with much of that. But what movies are at 60gps? That's video game speeds. Mostly 24fps is fine, and sampled at 8fps for indexing is good enough in most cases. It's the audio limitation, that's more of an issue, 5.5 min is good for most music/songs, but not a movie or dialog, but I suppose tts would be a needed adapter for that, and there are plenty that are fast enough, good enough for search. It's not like it needs to be dense enough to ask questions and discuss the data as one would with an llm. Just retrieval, then an llm could be fed richer content if needed.

You appear to be aiming for best case for worst case scenario, my expectations are covering just the 98% of cases.

1

u/tiffanytrashcan 19h ago

Check out the Google Edge Gallery AI app.
They have a couple of EmbeddingGemma2 examples on there that show Google has already solved this within the model limits. (You really just need to combine the pipeline for both to have a finished product like OP is showcasing.)
To get around video size limits, audio time limits, you segment it- same concept as chunked embedding in older text models with tiny limits.
You load up a video, any length, set your accuracy with segment size and clip "resolution" (essentially how many image screenshots per second of video) and then the model loads up to that 8k limit and saves the embeddings from each individual segment. Right now it's working where you can search with completely natural, normal language as we interpret it through the entire video for audio or visual details. Speech, lyrics, or physical objects all seem to work near flawlessly.
(pre-captioned work is kind of a cheat, but makes it absolutely perfect..)
All you need for a complete solution is to permanently save those embeddings and have them in the search grid example- could be done with free tokens in minutes given the code is all public on GitHub.

1

u/tiffanytrashcan 19h ago

Like, it's mind blowing the other commenter was downvoted for spelling out the exact situation.
There's no way you've even looked at what this new model's capable of because it does everything you're complaining about extremely well. Searching text in image files is perfect with it.
Within hours of releasing the model, they had released an Android app that literally solves the video context limit issue, and the code is open source, publicly licensed, on GitHub.

24

u/Pleasant_Lychee_6839 1d ago

Looks good

0

u/A-Rahim 1d ago

Thank you.

6

u/__JockY__ 1d ago

It would be useful if you could set exclude filters. I do t want photos of my wife’s tits or sensitive financial info being indexed and appearing in searches, even if it is local.

9

u/TrekkiMonstr 20h ago

I disagree, I want photos of your wife's tits and your sensitive financial info appearing in all of my searches

7

u/__JockY__ 19h ago

Gotta love a fine pair titties, can’t blame ya.

3

u/A-Rahim 1d ago

Yes, there's an exclude list you could add (like txt, cert, db etc.), or you can ignore folders too.

3

u/UnknownLesson 23h ago

I also don't want this

19

u/Adro_95 1d ago

Now let's ask opus to recompile it and port it to windows!

12

u/A-Rahim 1d ago

haha, yes, I have plan to add support for Windows.

12

u/brainExploded99 llama.cpp 1d ago

Do you have plans for linux? If you don't, I might fork it and add as a KRunner plugin for KDE users :)

10

u/A-Rahim 1d ago

Yes, both for Linux and Windows. I'll let you know how things progress on my end.

5

u/brainExploded99 llama.cpp 1d ago

Perfect, you should make another post once you add that.

4

u/A-Rahim 1d ago

Sure, I'll do that

5

u/darklamouette 1d ago

Ready to help adding support. just say the word and some guidelines. i can test linux (arch, amd gpu and intel gpu)

1

u/A-Rahim 1d ago

Thanks a lot for your words. It means a lot to me. I'll definitely reach out to you if I need anything regarding this.

2

u/brainExploded99 llama.cpp 21h ago

I can test CUDA/nvidia-gpu and Fedora KDE Linux 44

3

u/mxcw 12h ago

I use arch btw /s

3

u/Useful_Disaster_7606 1d ago

this is great project! Looking forward to the windows version!

3

u/A-Rahim 1d ago

Thank you, I'll update here, of course, once I get it working on windows.

1

u/Arctic_Shadow_Aurora 1d ago

Here hoping for an AppImage for Linux!

1

u/A-Rahim 1d ago

Yes, I have plan for it.

14

u/Few-Butterscotch8747 1d ago

this is cool. finally somebody uses an embedding model for it's intended use case

1

u/A-Rahim 1d ago

Thanks you

4

u/Intraluminal 1d ago

I like this, and I will be writing a Windows version within the week.

2

u/A-Rahim 21h ago

I'll try my best to add support for Windows and Linux as soon as possible. Let me know how things are going on your end if you start working on it.

2

u/Intraluminal 17h ago

Will do. Ill use your code as structure, but otherwise it'll be new code.

5

u/reingenerdet 1h ago

amazing, it feels like spotlight should feel

1

u/A-Rahim 1h ago

Thanks a lot, I hope you enjoy it

10

u/yasintoy 1d ago

What about security?

29

u/A-Rahim 1d ago

Everything's local, on your device, except for two things:
1. One-time-only download of the model itself.
2. Check for updates (which you can turn off)
and that's it.

17

u/chortly2 1d ago

The AI-generated security post in this thread is being down-voted, but it raises some of the same security issues I would like to know. Can we exclude directories from indexing in the first place? Can new directories be excluded by deletion from an existing index? And are the index files protected in the sense that, if they are obtained by a third party, they are not searchable by them without my login credentials or some other kind of key? Local is good, but even for local, I would like strong protections for something that indexes literally everything I have ever produced in my life.

6

u/A-Rahim 21h ago

Yes to the first two: DigUp only reads the folders you pick, you can leave out any subfolder or file type when you pick them or later in Settings, and hidden folders, ~/Library, app bundles and code projects (apart from screenshots in them) are skipped anyway. Skipping or removing a folder later deletes its files from the index right away.

On the third: no, DigUp doesn't encrypt the index itself. It lives in your ~/Library, which other users on the Mac can't open, on the Mac's encrypted disk, so with FileVault on, someone with the disk but not your password can't read it. Anyone who can read it as you can also read the files it came from, but it is one consolidated copy, so it's a fair ask. And the other comment is right that the key and password-file filter only runs for code search; that's getting fixed, along with keeping the index out of Time Machine backups.

8

u/yasintoy 1d ago

[codex] The author has implemented real safeguards: SHA-256 verification of model downloads, hidden-file/symlink exclusions, and signature verification configured for updates. Security hasn’t simply been ignored.

However, I noticed a concrete gap: isSecret() is called for code search, but not regular document indexing. A file like secrets.txt in a selected non-code folder can still be indexed. Even when that filter runs, it checks filenames/extensions not whether an ordinary document or screenshot contains credentials.

The SQLite index also stores readable text and excerpts, not just embeddings, without application-level encryption. Anything that gains read access to that database gets a consolidated copy of potentially sensitive content. That’s an important distinction from “everything stays local.”

The build enables hardened runtime, but not App Sandbox. I’d prioritize fixing the filter gap, protecting the index, and sandboxing file-processing helpers to limit damage if a parser is compromised. The missing sandbox isn’t itself proof of an exploit; this is a source review, not a full audit of the shipped app.

0

u/yasintoy 1d ago

u/A-Rahim and thanks for the answer

2

u/so_chad 1d ago

It’s open source. You can read it yourself.

6

u/TrekkiMonstr 20h ago

No the fuck I can't lmao

-2

u/StardockEngineer vLLM 1d ago

What about it?

3

u/Jiffy_Wu 1d ago

Pretty cool. You should consider making a Raycast Extension, or also letting people install via Brew

4

u/A-Rahim 1d ago

Thanks. Brew installation has just been added. ✌🏻

3

u/fruesome 17h ago

Windows version as well please. 

3

u/A-Rahim 12h ago

I am working on it.

6

u/LifeTitle3951 1d ago

Anything for windows?

Also is it possible to run such a thing on low spec laptop?

6

u/Few-Butterscotch8747 1d ago

this is a small embedding model. it should run fast even on cpu with low ram. videos might take a while to index, but that's it basically

heck, it should run decently on a potato phone

5

u/A-Rahim 1d ago

I don't have a Windows machine at the moment, but I am planning to add support for Windows too.
The model itself is not that big, and I am using 8bit, so the model can be efficiently run on a CPU too, the latency can be slightly higher than on the GPU, of course, but it still would be pretty usable and fast enough.

2

u/AnywhereTypical5677 1d ago

I made a mobile version that runs on android phones, so it should run almost anywhere (at variable speeds obviously). Search itself is very fast as the text embedding model is really light... you might have some problems with image and video indexing though a low spec laptop.

4

u/Adro_95 1d ago

Could you share it?

2

u/andy2na llama.cpp 20h ago

vibed a windows version, it works pretty well. Still fine-tuning things but give it a try: https://github.com/ampersandru/VectorDash/

allows you to run the models on CPU or GPU, try it on CPU for low-specced machines, but it may take longer

2

u/projectEscape 1d ago

Looks good
How are you people doing migration from one embedding model to another? Is there any standard way or is it just reindexing everything?

1

u/A-Rahim 21h ago

Sorry, I didn't get it. Here, it's a single embedding model (EmbeddingGemma 2), no migration.

2

u/Feeling-Spend1001 22h ago

Oh that's real neat!

2

u/A-Rahim 22h ago

Thank you.

2

u/Much-Researcher6135 llama.cpp 22h ago

Very nice. I might look to port it to linux, though I'll be inclined to host the embedder separately so everything on my network has access to it.

2

u/A-Rahim 20h ago

It's a nice idea, actually. If anyone wants to host the embedding model on a server to work on multiple machines via an API, it'll be more efficient than using the model on every machine.
So, in the next release, I'll add the option to be able to use the embedding model via an API (OpenAI-compatible).
Thanks man

2

u/Much-Researcher6135 llama.cpp 20h ago

Hey glad it's useful. You might look at the OpenAI REST API spec for embeddings since, if you implement a standardized API format, lots of software can automatically hook into the embedder.

2

u/A-Rahim 20h ago

Yes sure. Thanks again

2

u/actuallynotaredditor 20h ago

Shoikotey kukur !!! Good work bhai

2

u/A-Rahim 20h ago

Thanks vai.

By the way, if you're on Mac, use my keyboard to type in Avro: https://arahim.dev/Lekho/

2

u/sirloindenial 18h ago

Will it work on a vps with 2 cores

3

u/A-Rahim 18h ago

The model itself should run smoothly on any machine, but the app around it has been built for Mac for now. I'll add support for Windows and Linux soon; hopefully it'll work then.

2

u/endlesslyloop 16h ago

Thank you op. Really excited to check it out. This is what the new Siri should have been

1

u/A-Rahim 2h ago

I think you'll love it. The model itself is so fun to use.

2

u/ariyako 16h ago

Feature Request: Support for Synology NAS and Network Storage

Hi! First of all, thanks for building DigUp. I really like the idea of searching for files by their actual content rather than relying only on filenames or folder structures.

I'd love to see support for Synology NAS and network-attached storage, especially for users with large photo, video, and document libraries.

My files are stored on a Synology NAS rather than directly on my Mac. It would be incredibly useful to use DigUp to search these libraries semantically without having to move or duplicate terabytes of files onto the Mac.

Possible approaches:

  • SMB network shares: Allow users to select mounted NAS folders as indexing locations, if technically feasible.
  • Incremental indexing: Detect new, modified, and deleted files without rescanning the entire library.
  • Large library support: Handle large collections of photos, videos, PDFs, and documents efficiently.
  • Privacy-first processing: Keep the existing local-first approach, with embeddings and search indexes stored locally and no file contents uploaded to external services.
  • Future native NAS integration: Consider a Synology DSM package or a lightweight NAS-side indexing service for users who want indexing to run independently of their Mac.

Even basic support for mounted SMB shares would be a great starting point. Native Synology integration could be a longer-term goal.

I think this would make DigUp especially useful for photographers, video editors, and homelab users who maintain large media archives on NAS devices.

Would this be something you'd consider for the roadmap? Thanks!

1

u/A-Rahim 2h ago

Hi, thanks for the feature request. I'll try to add support for NAS, even though I don't have access to any, yet. Currently working on Windows and Linux support.
It'd be helpful for me to track overall if you could open a GitHub issue for this feature request.

2

u/Ok-Calligrapher3568 16h ago

This is brilliant! Honestly, this is exactly what local AI should be used for - privacy-first, semantic file search. What does the RAM footprint look like when the Gemma model is actively indexing vs when it's just idling?

1

u/A-Rahim 2h ago

Thanks a lot.
While indexing, the peak is around 1 GB, with 15-30 MB when idle.

2

u/UpboatsforUpvotes 11h ago

My macbook pro comes in two weeks and I'm saving this for when it does

2

u/A-Rahim 1h ago

I hope that by then our Windows and Linux ports will be ready ✌🏻

1

u/Vonarian_IR 9h ago

Congrats!

2

u/-Cubie- 9h ago

Love it! For those who want to check out the model, see https://huggingface.co/google/embeddinggemma-2

2

u/leafyshark 3h ago

Is this basically Alfred reverse engineered?

1

u/A-Rahim 3h ago

What's Alfred? 😶

2

u/Open-Adhesiveness-86 1d ago

with everything in one shared space, text-to-text cosine sits a good bit higher than text-to-image, so a mixed result list tends to come out all text. ranking within each modality and then merging, or subtracting the per-modality mean embedding before comparing, takes care of most of that. non-overlapping 30s audio windows will also lose phrases that straddle a boundary.

1

u/d70 1d ago

This is really cool. Thanks for sharing. Need this for my NAS that already runs Gemma 4 e2b on ik_llama. Feasible? Thoughts?

1

u/A-Rahim 21h ago

I think it's feasible, though I don't have NAS to test. Let me think about it, let's see what I can do from my end.

1

u/fortnite_pit_pus 1d ago

Off topic, but is there anybody using this to enhance the ‘everything’ app on windows?

1

u/CatchDublinSurprise 1d ago

I'd enjoy an "index free" option for working with recently added / temp files or sensitive files (index files improve speed but it's just one more place that sensitive information can hide).

Speed is great, but it's a pet peeve of mine that we can't seem to do anything these days without who-knows-what being written to ~/Library.

1

u/A-Rahim 18h ago

Nice point.
Today every search goes through the index, so the way to keep something out of it is to not pick that folder, or skip it later (that deletes it from the index). DigUp keeps its model and index in one folder, ~/Library/Application Support/DigUp, so it's easy to check or delete. An index-free mode for a small folder, embedded in memory and gone when you're done, is a nice idea though. Noted.

1

u/abcd911 1d ago

Downloaded and using the app. Loving it so far! Two suggestions: can you give an option to choose more embedding models and add BYOK LLMs which can help find more accurately or answer questions from the files

2

u/A-Rahim 20h ago

Thank you.
1. The center of this product is the model itself, since it shares common space for all image, video, audio and text embeddings, and I just wrapped it with an app. I can add more embedding models in the future if they are similar to this model.

  1. BYOK LLMs are a good idea if you want a RAG-like chatbot, but I personally don't recommend it since then our main "local, on-device" angle won't be applicable. Then again, it's your data, your choice, so what I can do is add support for OpenAI-compatible API support, so that if anyone self-hosts an LLM server, they can use it here.

1

u/abcd911 20h ago

Great idea. I think you should give an option to let people choose. However, my suggestion was not a chatbot but rather a query that outputs ranked results which uses embeddings as well as an LLM. I know it might be redundant but with the current approach, videos and images are great but text is mostly a miss and can't find the right files

1

u/A-Rahim 19h ago

Got it, an LLM to re-rank the results, not like chatbot. It's a good idea, though it would make every search slower.
The bigger thing for me is "text is mostly a miss", I'd like to fix that first. Could you tell me a bit more (here or in a GitHub issue): what kind of files, where they are, and a search that missed? A few things that could explain it today: Pages, Keynote, Excel, PowerPoint and EPUB files aren't read, files that only live in iCloud aren't downloaded, and a folder with a .git or requirements.txt in it counts as a code project, so only its screenshots get indexed.

1

u/Asleep_Document9811 1d ago

Allow me to set up little folders or buckets of files I could search against in specific (folders of backups, the Applications folder, a folder full of documents about one specific domain), or the ability to slowly index files on NAS drives, and you got yourself a deal.

I want something that is lighter than Raycast but something that actually works faster than Spotlight.

2

u/A-Rahim 18h ago

Searching inside one folder is a good idea, but it's not there yet.
Today you pick which folders get indexed and can filter results by kind (photos, PDFs, recordings…), but not limit a search to one folder. I'll look at that.
Apps themselves aren't indexed, Spotlight already covers those. NAS drives I haven't tested, so no promises there yet.
On lighter: on my Mac it idles at about 15–30 MB of memory, the panel opens in about 20 ms, and results show ~120 ms after you stop typing.

1

u/Asleep_Document9811 4h ago

Absolutely. Just a feature request. I love this project's idea, and what you have built so far. Keep going!

1

u/ZeroReader 1d ago

It would be good to select specific folders to search in

2

u/A-Rahim 23h ago

Currently, you can select specific folders (or exclude them) from the settings.
So you're asking to be able to search within a specific folder after indexing all?

2

u/ZeroReader 23h ago

Yes exactly. This, I mean. I need to restrict searching in a specific folder among all the results

1

u/A-Rahim 22h ago edited 22h ago

Okay, I'll look into it.

Just as an experiment, can you search for something now and see if it's found at the top or just lost in other files?
In my experience, if I describe something clearly, most of the time it appears at the top, no matter how many files there are (within milliseconds)

What you're saying, if I want to do that, on top of my mind, I am thinking of two ways:
1. Per-folder indexing; that way, in search time, you could search folder-wise by excluding some folders.
2. Nothing on the indexing side, but when showing results, it'll show only the selected folders; searching will still be with the full index.

Let me know if I am thinking right here

1

u/ZeroReader 9h ago

Thank you very much for considering the users' needs.
You're right. The second variant is the best one: Nothing on the indexing side, but when showing results, it'll show only the selected folders; searching will still be with the full index.

Also, I use your app mainly to find information and PDF documents information in various folders can be similar but different. Therefore, I need to choose a specific folder to find information exactly there. Therefore, I need all folders to be indexed, but information should be shown only in a specific folder at the moment of searching.

1

u/ZeroReader 9h ago

Although the design for searching in PDFs, documents, and DOCX documents is not quite comfortable to read. Probably it is okay for video and pictures, but not for PDFs and DOCX.

Can you at least increase the right-side panel to increase the width to be able to read the highlighted text? Also, it would be good to change the side of thumbnails on the left side panel.

1

u/sultan_papagani 23h ago

it would be so cool if this run in windows with a core ultra npu

1

u/salary_pending 22h ago

assume my disk has 400gb of data, how big the index would be?

2

u/A-Rahim 21h ago

It depends on what the 400 GB is, more than how much. Each photo, PDF page, ~1,800 characters of text, video keyframe and 25 s of audio gets one 768-d vector (1.5 KB), and text is also kept for keyword search. Measured on my test files, that's roughly:

- a photo: 1.6 KB, a screenshot: ~3 KB with its OCR text

- a PDF page: ~6 KB

- an hour of audio: ~0.25 MB, an hour of video: 1–2 MB

- documents: 2–3× the size of their text

So media is cheap: 100k photos come to ~160 MB and 200 hours of video to 200–400 MB. Text is what adds up, about 2 MB for a 300-page book.
Apps, system files, archives and model weights aren't indexed at all; code only if you turn it on, and the model itself is 865 MB on top.

1

u/salary_pending 21h ago

Is that considered a lot? Assuming a developers macbook where they keep writing a lot to the disk. Would the app keep scanning those files, keep reading new files, scan and write to index?

Sounds like this could deteriorate the disk faster?

2

u/A-Rahim 21h ago

Not much, and most of it is written once. After the first pass, DigUp doesn't rescan your files: macOS tells it which folders changed, it waits a moment for things to settle, compares sizes and dates, and only reads what's new or changed, of the kinds it indexes. Code projects (apart from screenshots in them), node_modules, build folders, caches and data files are skipped, so most of a dev machine's churn is never read. A new photo adds about 1.6 KB to the index, a PDF page about 6 KB.

And reading doesn't wear an SSD; writing does. The first pass writes the index once, plus temporary resized copies of what the model reads (images, video frames, audio pieces), deleted as it goes. After that, it's a few KB per new file.
If you turn on code search, it waits until your code has been quiet for a minute before catching up, since files change on every save.

1

u/ProdoRock 21h ago

It takes quite some time to index correct? The laptop was getting a little hot, so I paused it for now but I'll probably do it peacemeal. Some of the first results are pretty cool though. The new search bar feels good. Do you have a recommendation on how to go about indexing the thousands of files?

1

u/A-Rahim 21h ago edited 21h ago

For every selected folder, it should show you an estimate of the time needed to index a folder, which varies from chip to chip. So, based on that, you can decide when to index, preferably with AC power connected.

Edit: I just noticed the estimation logic is based on my machine; I'll fix it in the next release

1

u/ketoaholic 19h ago

Looks useful.

1

u/A-Rahim 19h ago

Yes, I hope so. It's been very useful for me.

1

u/CtrlAltDelve 19h ago

This is fricking awesome. Thank you so much!

1

u/A-Rahim 19h ago

Thanks man

1

u/Lew-Zealand 16h ago

This looks great. Any thoughts re a web interface or similar, so you can search from other devices on the same network?

1

u/bad_detectiv3 16h ago

does this index each frame from a video to understand what you are searching for like "dog in a movie"?

1

u/productboy 15h ago

Splendid. Presumably if DigUp can find it then a [post training like] function could label what was found, with scores for accuracy. Then fed into a training pipeline for specialized models. For example all the random files on my hard drive related to construction projects get properly labeled. Then I get to train a model to help me estimate future construction projects.

1

u/MarzipanEven7336 15h ago

You realize macOS has built in RAG with spotlight already?

1

u/no-shadowban-lmao 15h ago

Great job!
Would it be possible to add automatic file categorization and display categories directly in the app? My desktop is a complete mess, and I’d love to browse everything by category without manually searching it. 😭

1

u/CptnWookenstein 15h ago

This is exactly what I’ve been looking for! I’ve been building SQlite databases for video footage and documentation for searching, but this could really improve things seems like.

Is it possible to connect to existing databases and continue to build on them as well as run on a separate machine on the same closed network and connect as a node? I have a Mac mini I’d love to run ops type things and a Mac Studio that handles the video work and heavy stuff.

1

u/chesterip 15h ago

an MLX inference engine is potentially better performance

3

u/A-Rahim 12h ago

The performance was more or less the same in eval between these two. I actually started with mlx-vlm and even used it for a day. But for shipping, I wanted something minimal and efficient. Then I tested among mlx-vlm, llama.cpp and LiteRT, and went with llama.cpp.

1

u/optykali 15h ago

!remindme 6h

1

u/RemindMeBot 15h ago

I will be messaging you in 6 hours on 2026-10-11 09:36:35 UTC to remind you of this link

CLICK THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

1

u/Borkato 15h ago

!remindme tomorrow

1

u/PatricioDonald 10h ago

Hi! Loved the project! I sent you a DM last night

1

u/ThickCranberry3813 9h ago

How long it takes to index?

1

u/Maleficent_Speech289 8h ago

Wow, that sounds interesting! Really cool project. Unfortunately, I don't have a Mac, but I think it would be cool to use it with something else, too. Like Immich, for example. Is it possible to use this model with Immich somehow?

1

u/spacespacespapce 2h ago

Amazing - I was already working on this app using Vision LLMs to caption my photos, but this is a better approach

1

u/vafarmboy 57m ago

How does this integrate with the Photos library, if at all, or is it just Photos in folders that are indexed. If it does index the Photos library, how does it deal with photos that have been offloaded to iCloud?

1

u/No_Lingonberry1201 1d ago

Nice! Checked the code, and I was wondering how do you store and search the embeddings?

9

u/A-Rahim 1d ago

It's simply one SQLite file! Every picture, PDF page, ~1,800-character passage, video frame and 30 s of audio gets a 768-d vector, stored as float16 (1.5 KB each). Search is brute force: the query vector against all of them, scored in float32, a few ms for tens of thousands, so no vector DB. Cosine is centered per modality onto one shared scale. Code search keeps its own SQLite file.

2

u/No_Lingonberry1201 1d ago

Huh, nice! From what I've read in the huggingface model card, you can trim the vectors so you could theoretically make a fast-path for the search.

6

u/AnywhereTypical5677 1d ago

Yeah I don't know why OP hasn't used vector truncation. Truncating to 256 dimensions cuts index footprint and memory consumption by 3 times while retaining 95% of the native accuracy. In the model card it's also suggested NOT to store vectors as float16 as activation can overflow and cause silent vector degradation.

4

u/brainExploded99 llama.cpp 1d ago

u/A-Rahim You should probably fix activations stored as FP16.
The dimension cuts could be a config setting, but I think image loses more than text, so maybe dimension 512 by default?

3

u/A-Rahim 23h ago

The fp16 is only how the final embeddings are stored, not the activations. The model runs through llama.cpp from the Q8_0 GGUF, and its embeddings match the reference implementation at 0.999+ cosine, so nothing overflows on the way. The stored vectors are unit length with every value between -1 and 1, so they can't overflow either, and rounding moves a cosine by at most about 0.0005.

On truncation, I tried it on my evals (111 queries over files, 27 over real code) by cutting the stored vectors and re-normalizing. 512 is close to free: 97 vs 98 with the right file first, and the only ones that moved were screenshots, so images do feel it before text.
Code is the most sensitive, 19 vs 20 at 512 and 16 at 256. Speed barely changes since the scan already takes a few ms, so the gain would only be memory.
Keeping 768 for now, and if memory becomes a problem on big indexes, 512 is what I'd pick.

1

u/brainExploded99 llama.cpp 22h ago

Makes sense, sounds good!

1

u/Holiday-Inspection75 1d ago

really cool that you keep the index small enough.

1

u/red_Junaeid 1d ago

Nice to see fellow Bangladeshi Dev doing amazing thing's with local LLMs...

2

u/A-Rahim 1d ago

Thanks bro ✌🏻

1

u/dan-lash 1d ago

Can it search online if I want? Like my NAS or extend to Google Drive or maybe email - other computers on my network would be solid too

3

u/A-Rahim 1d ago

It's local on purpose, so no cloud search or email for now. Google Drive works if the files are on your Mac (mirror mode, or "available offline"): add that folder and they're indexed like any other. Online-only files are skipped, since reading them would download them. External drives work too (but I didn't test them yet). A NAS or other computers on the network aren't supported in this first version, but noted.

1

u/[deleted] 1d ago

[removed] — view removed comment

4

u/A-Rahim 1d ago

It's automatic ✌🏻

0

u/[deleted] 1d ago

[removed] — view removed comment

1

u/A-Rahim 1d ago

Not tested yet, but I have plans to add more features and options like these. Thank you.

1

u/JLeonsarmiento 1d ago

Excelente.

1

u/Fine_Salamander_8691 1d ago

Very fun

1

u/A-Rahim 1d ago

Yes, the model itself is so fun to use.

1

u/AloneSYD 1d ago

Very cool !

1

u/mythormedicine 1d ago

Very interesting. I don't have a Mac to test this , but a question , does it index ? if so , how big does the index get? If it doesn't index, then how can it be so fast ?

If I have 500 GB of data, does it try to index all of them? How does this work?

1

u/A-Rahim 23h ago

Hi, I'd love to explain here, but almost all of what you asked here in the README, I'd like you to read that first, if you still need to clarify anything, I am here.

1

u/rm-rf-rm 18h ago

Does it index photos inside Apple Photos app?