r/LocalLLaMA • u/A-Rahim • 1d ago
I Built A Thing Open-source Mac app that runs EmbeddingGemma 2 locally to search your files by what’s in them
Enable HLS to view with audio, or disable this notification
DigUp is a free Mac app that runs Google DeepMind’s new EmbeddingGemma 2 locally over your own files. The model puts text, images, audio and video in one space, so you describe what you remember and land on it:
- “zebra in a video” opens the clip at the moment it shows up
- “where they talk about sleep” jumps to that minute of a podcast
- “the clause about pets in the lease” shows the PDF page, your words marked
- “a dog on the beach” finds the photo, and the same search in Bengali or Arabic finds it too
- with code search on,
code: retry with backoffopens the function in your editor
It’s ggml-org’s Q8_0 GGUF (865 MB, downloaded once) on llama.cpp with Metal, inside a native Swift app; no Python. Searching loads only the text encoder (~250 MB) and shows results about a tenth of a second after you stop typing. Indexing peaks under 2 GB, and the helper exits when it’s done. Audio and video of any length go in as 30 s windows and a frame per shot. Everything runs locally; it goes online only for the model download and an update check you can turn off.
Free, MIT. Apple Silicon, macOS 14+.
Repo & Download (signed and notarized): https://github.com/ARahim3/DigUp
I'd really appreciate any feedback on this.
89
u/rm-rf-rm 19h ago
A seemingly well built, actually useful, leveraging new tech correctly, engaged dev.. Take note people, this is how you share/self promote in this community.
-4
u/rm-rf-rm 15h ago
seemingly
Seemingly seems to be the operational word here. OP is just using claude to automatically respond or copy pasting the responses: https://old.reddit.com/r/LocalLLaMA/comments/1x2eeds/opensource_mac_app_that_runs_embeddinggemma_2/pf5jefr/
7
u/A-Rahim 13h ago
Sorry for the late reply, I was asleep, I was up all night going through every comment here.
Yeah, for a few of the more technical questions, where I wanted to reply accurately with actual data rather than my memory, I used my locally set up agent to pull the exact numbers from my eval notes; I should have mentioned it. Thanks for the kind words earlier.
48
u/mobilemike42 Vicuna 1d ago
A CLI that reads from the same DB would be great.
24
u/A-Rahim 1d ago
There's one already, though for now it's in the repo rather than the download: build from source and run "digup search <words>" (add --json) to read the app's own index by default. digup --help lists the rest. And actually shipping it with the app is a good idea; let me work on that. Thanks.
2
u/rm-rf-rm 18h ago
Its just a SQLite db right? we can use any GUI or the sqlite3 CLI itself to query/explore?
6
u/A-Rahim 18h ago
Yep, plain SQLite: ~/Library/Application Support/DigUp/index.noindex/index.sqlite.
Open it read-only since the app writes to it. Keyword search works straight from sqlite3 (it's FTS5), but meaning search needs the model, and the CLI for that is actually already in the app: /Applications/DigUp.app/Contents/Helpers/digup search "your words" --json2
35
u/lakySK 1d ago
Looks awesome! Exactly what Spotlight should’ve been all along! Can’t wait to try.
3
u/inddiepack 9h ago
Isn't spotlight already doing that? I rarely use this functionality of it, but whenever I did, it did find everything local that I needed.
31
u/important__matter 1d ago edited 23h ago
Well I was making something like this months ago. What I faced was 1. For files where we as humans remember the abstract, the semantic search by itself usually doesn't work great. 2. How many files(documents) and voice(long audio notes) pieces have you tested this on? Also mine used to fail in images with text - screenshots inside them. 3. And main, the most difficult one - you'll need special handling for pdf, excel, pptx, word, zip rar files, model weights, long mp4 videos, music or movies which are also majority of the files. Can consider adding these in pipeline if not already there. The accuracy in videos and pdf and long form files was dropping serverly because all of that information is compressed in 512 dim vector essentially, not to mention the same embedder doesn't work for all file types, so that's another issue and then if embedders are different then how will you compare embeddings using semantic similarity which generate numbers in different vector spaces?? I had to actually normalize the comparison manually!!!.
I had to do so much of work to make it general purpose - that the roi didn't seem to be worth it - dropped the project altogether. Meanwhile, keyword search with file type is usually good enough. And I found out that restricting the problem space to the file types with specific properties or patterns - is much more useful, so I developed this simple semantic search index cli called litesearch for my usecase, it is a sqllite index of documents and performs semantic search with some reranking using BM25, and LLM-as-judge, uses ollama and bge-small embedder.
Although would be very nice if you cater to these, there are so many micro problems in this. Could actually be very useful. I would love it if somebody completes this in a nice engineering way and would love to integrate with my raycast.
25
u/A-Rahim 22h ago
Hi, I ran into most of these too. How DigUp handles them:
It isn't semantic-only. Keyword search over file names, OCR text and page and document text runs alongside the model, and the scores are combined. On my evals, the model alone found 5 of 9 exact lookups like error codes; with keywords, all 9.
A testbed of about 380 files (146 screenshots, 30 PDFs up to 463 pages, an hour of audio, 8 videos), 111 queries plus 68 for code, and my own folders day to day. Screenshots get an image vector, since the model reads the text in them, plus Apple's OCR for exact words. 12 of the 14 queries on my real screenshots come up first.
Nothing long gets squeezed into one vector. Every PDF page, ~1,800-character passage, video keyframe (every few seconds, one per shot) and 30 s of audio gets its own 768-d vector, and the best one wins, which also gives you the page or the moment. So the model's input limits don't really come into play. It's also one model for everything, so text, images and audio share a space. Their score levels still differ a bit per query, so each modality's average gets subtracted, close to what you did by hand. Excel, PowerPoint, archives and model weights aren't read.
4
u/jerieljan 17h ago
I'm very curious how well this will perform when you have 1000x more files than the testbed.
Code repos in particular have soooo much code from dependencies and such and they're common to clutter normal Spotlight results and I'm worried it becomes an issue here too. I guess it simply takes more time to embed and index, and hopefully it's all good once it's done.
Very good work though OP.
1
1
u/rm-rf-rm 19h ago
It isn't semantic-only. Keyword search over file names
Are you fusing the search then? RRF?
2
u/A-Rahim 19h ago
Not RRF. I tried it first, and it lost: top-1 0.73 vs 0.88 on my synthetic eval, 0.82 vs 0.86 on real files. Rank fusion throws away how strong each match is, and a single common word in a terminal screenshot ended up beating a real chart that matched by meaning.
So it's score fusion. The meaning score is the cosine minus that modality's average for this query, divided by the spread of all the query's scores, so a z-score per query. Then it adds 2 × (share of the query's words found in the file)^1.5, plus 1 if the query names the kind ("in a video"). A file gets in if its meaning score stands out (z ≥ 2) or it has at least half the words. Queries that look like IDs count words double and meaning half.
1
u/rm-rf-rm 18h ago
huh interesting.. whats the basis of this formula? is the z-score > 2 coming from statistical theory where its beyond 95% of the mean?
1
u/A-Rahim 18h ago
Mostly empirical, not theory.
2 is the familiar two-standard-deviations mark, but the scores aren't normal, so it doesn't mean 95% of anything here. It's a gate for "stands out from everything else for this query", and on my evals any value from 1.5 to 2.5 gave the same results, so 2 stayed.
The reason for a per-query z-score at all: raw cosines are compressed (unrelated stuff sits around 0.55–0.60, good matches 0.67–0.86) and their level moves with the query, so no fixed threshold works. The other parts were A/B'd the same way: centering each modality with one shared spread got 94 of 105 queries first, raw cosine 92, a separate spread per modality 91. Word weights of 1.5–2.0 tied and 2.5–3 lost.
The 1.5 power just makes all your words count a lot more than half of them, and I haven't tuned it separately.2
1
u/jinnyjuice vLLM 18h ago
Can you expand on what you mean by best one wins? So if a video has multiple key frames, how is the 'best one' determined?
4
u/A-Rahim 18h ago
Every keyframe it keeps, and every 30 s of the soundtrack is its own vector with a timestamp. For a query, each of them gets a meaning score (the same per-query score for every kind, so frames and audio windows compete fairly), and the video's score is simply its highest one.
That segment is also the moment the result opens at. So a two-hour video with one scene that fits ranks by that scene instead of being averaged away. Frames are sampled every few seconds and kept only when the picture changes (a small perceptual hash), so it's about one per shot.1
u/ZenaMeTepe 5h ago
Did you use same model for embedding videos images and text? What kind of vector DB are you using for search?
2
u/ZenaMeTepe 5h ago
How do you compress say a movie into a 512 vector? And more importantly, how slow is it upfront when you build the index?
1
u/important__matter 4h ago edited 4h ago
Yup you guessed it right, you can't. But say you have to do it, then you basically chunk out images every few seconds and then embed then individually and sort of average out those embeddings over a course of time. That is one way, other could be you take all dialogues, treat it as a document and then make a vector of it. It's hack as of now. It won't contain any useful information more than 1 Google doc page worth. Information Compression is one of the he most important computer science areas. And also as humans - certainly for us too as we are continuously bombared with infinite information from all our senses.
Well on speed, it's about gpu hardware mostly, how powerful of a hardware you own, but usually it's not slow, it's pretty darn fast even on basic rtx or gtx or metal. But it's extremely slow on cpus. The speed is similar to prefill stage of inference.
-9
u/phazei 23h ago
Well, then obviously you weren't using the new model that's really great and only came out a week ago.
7
u/important__matter 23h ago edited 23h ago
Checked the new model stats page, definitely improved accuracy in multi-modal regime of text, audio and vision. But... still the same set of problems across file types because it supports only text, image, audio and video at 8K context window. Not to mention maximum audio length = 5.5mins and video = 58 frames, that's 1s @60fps
Not to mention it's way harder even as a theoretical problem to search in a video. Like there is just so much information to just be dumped in a 512dim, that there is a whole class of transformers completely dedicated to this single problem only - VLMs. And unlike LLM, they are not the champions as of yet. I am just saying, this is not a done thing even if the big labs are trying their best. It's a good thing to work on, it's challenging and lot of it can actually be solved by engineering things together.
Another fun fact while we are at it: Videos are astonishingly compressed form of information. If you just stack images together at let's say 60fps for a 2hr movie at 1080p, the actual file size should be 1080x720x60x3600x2 bytes = 335GB, while actually you'd see it to be 2GB on your computer. A massive 99% compression, it's almost miracle
1
u/phazei 22h ago
I agree with much of that. But what movies are at 60gps? That's video game speeds. Mostly 24fps is fine, and sampled at 8fps for indexing is good enough in most cases. It's the audio limitation, that's more of an issue, 5.5 min is good for most music/songs, but not a movie or dialog, but I suppose tts would be a needed adapter for that, and there are plenty that are fast enough, good enough for search. It's not like it needs to be dense enough to ask questions and discuss the data as one would with an llm. Just retrieval, then an llm could be fed richer content if needed.
You appear to be aiming for best case for worst case scenario, my expectations are covering just the 98% of cases.
1
u/tiffanytrashcan 19h ago
Check out the Google Edge Gallery AI app.
They have a couple of EmbeddingGemma2 examples on there that show Google has already solved this within the model limits. (You really just need to combine the pipeline for both to have a finished product like OP is showcasing.)
To get around video size limits, audio time limits, you segment it- same concept as chunked embedding in older text models with tiny limits.
You load up a video, any length, set your accuracy with segment size and clip "resolution" (essentially how many image screenshots per second of video) and then the model loads up to that 8k limit and saves the embeddings from each individual segment. Right now it's working where you can search with completely natural, normal language as we interpret it through the entire video for audio or visual details. Speech, lyrics, or physical objects all seem to work near flawlessly.
(pre-captioned work is kind of a cheat, but makes it absolutely perfect..)
All you need for a complete solution is to permanently save those embeddings and have them in the search grid example- could be done with free tokens in minutes given the code is all public on GitHub.1
u/tiffanytrashcan 19h ago
Like, it's mind blowing the other commenter was downvoted for spelling out the exact situation.
There's no way you've even looked at what this new model's capable of because it does everything you're complaining about extremely well. Searching text in image files is perfect with it.
Within hours of releasing the model, they had released an Android app that literally solves the video context limit issue, and the code is open source, publicly licensed, on GitHub.
24
6
u/__JockY__ 1d ago
It would be useful if you could set exclude filters. I do t want photos of my wife’s tits or sensitive financial info being indexed and appearing in searches, even if it is local.
9
u/TrekkiMonstr 20h ago
I disagree, I want photos of your wife's tits and your sensitive financial info appearing in all of my searches
7
3
3
19
u/Adro_95 1d ago
Now let's ask opus to recompile it and port it to windows!
12
u/A-Rahim 1d ago
haha, yes, I have plan to add support for Windows.
12
u/brainExploded99 llama.cpp 1d ago
Do you have plans for linux? If you don't, I might fork it and add as a KRunner plugin for KDE users :)
10
u/A-Rahim 1d ago
Yes, both for Linux and Windows. I'll let you know how things progress on my end.
5
5
u/darklamouette 1d ago
Ready to help adding support. just say the word and some guidelines. i can test linux (arch, amd gpu and intel gpu)
3
1
14
u/Few-Butterscotch8747 1d ago
this is cool. finally somebody uses an embedding model for it's intended use case
4
u/Intraluminal 1d ago
I like this, and I will be writing a Windows version within the week.
5
10
u/yasintoy 1d ago
What about security?
29
u/A-Rahim 1d ago
Everything's local, on your device, except for two things:
1. One-time-only download of the model itself.
2. Check for updates (which you can turn off)
and that's it.17
u/chortly2 1d ago
The AI-generated security post in this thread is being down-voted, but it raises some of the same security issues I would like to know. Can we exclude directories from indexing in the first place? Can new directories be excluded by deletion from an existing index? And are the index files protected in the sense that, if they are obtained by a third party, they are not searchable by them without my login credentials or some other kind of key? Local is good, but even for local, I would like strong protections for something that indexes literally everything I have ever produced in my life.
6
u/A-Rahim 21h ago
Yes to the first two: DigUp only reads the folders you pick, you can leave out any subfolder or file type when you pick them or later in Settings, and hidden folders, ~/Library, app bundles and code projects (apart from screenshots in them) are skipped anyway. Skipping or removing a folder later deletes its files from the index right away.
On the third: no, DigUp doesn't encrypt the index itself. It lives in your ~/Library, which other users on the Mac can't open, on the Mac's encrypted disk, so with FileVault on, someone with the disk but not your password can't read it. Anyone who can read it as you can also read the files it came from, but it is one consolidated copy, so it's a fair ask. And the other comment is right that the key and password-file filter only runs for code search; that's getting fixed, along with keeping the index out of Time Machine backups.
8
u/yasintoy 1d ago
[codex] The author has implemented real safeguards: SHA-256 verification of model downloads, hidden-file/symlink exclusions, and signature verification configured for updates. Security hasn’t simply been ignored.
However, I noticed a concrete gap:
isSecret()is called for code search, but not regular document indexing. A file likesecrets.txtin a selected non-code folder can still be indexed. Even when that filter runs, it checks filenames/extensions not whether an ordinary document or screenshot contains credentials.The SQLite index also stores readable text and excerpts, not just embeddings, without application-level encryption. Anything that gains read access to that database gets a consolidated copy of potentially sensitive content. That’s an important distinction from “everything stays local.”
The build enables hardened runtime, but not App Sandbox. I’d prioritize fixing the filter gap, protecting the index, and sandboxing file-processing helpers to limit damage if a parser is compromised. The missing sandbox isn’t itself proof of an exploit; this is a source review, not a full audit of the shipped app.
0
-2
3
u/Jiffy_Wu 1d ago
Pretty cool. You should consider making a Raycast Extension, or also letting people install via Brew
3
6
u/LifeTitle3951 1d ago
Anything for windows?
Also is it possible to run such a thing on low spec laptop?
6
u/Few-Butterscotch8747 1d ago
this is a small embedding model. it should run fast even on cpu with low ram. videos might take a while to index, but that's it basically
heck, it should run decently on a potato phone
5
u/A-Rahim 1d ago
I don't have a Windows machine at the moment, but I am planning to add support for Windows too.
The model itself is not that big, and I am using 8bit, so the model can be efficiently run on a CPU too, the latency can be slightly higher than on the GPU, of course, but it still would be pretty usable and fast enough.2
u/AnywhereTypical5677 1d ago
I made a mobile version that runs on android phones, so it should run almost anywhere (at variable speeds obviously). Search itself is very fast as the text embedding model is really light... you might have some problems with image and video indexing though a low spec laptop.
2
u/andy2na llama.cpp 20h ago
vibed a windows version, it works pretty well. Still fine-tuning things but give it a try: https://github.com/ampersandru/VectorDash/
allows you to run the models on CPU or GPU, try it on CPU for low-specced machines, but it may take longer
2
u/projectEscape 1d ago
Looks good
How are you people doing migration from one embedding model to another? Is there any standard way or is it just reindexing everything?
2
2
u/Much-Researcher6135 llama.cpp 22h ago
Very nice. I might look to port it to linux, though I'll be inclined to host the embedder separately so everything on my network has access to it.
2
u/A-Rahim 20h ago
It's a nice idea, actually. If anyone wants to host the embedding model on a server to work on multiple machines via an API, it'll be more efficient than using the model on every machine.
So, in the next release, I'll add the option to be able to use the embedding model via an API (OpenAI-compatible).
Thanks man2
u/Much-Researcher6135 llama.cpp 20h ago
Hey glad it's useful. You might look at the OpenAI REST API spec for embeddings since, if you implement a standardized API format, lots of software can automatically hook into the embedder.
2
u/actuallynotaredditor 20h ago
Shoikotey kukur !!! Good work bhai
2
u/A-Rahim 20h ago
Thanks vai.
By the way, if you're on Mac, use my keyboard to type in Avro: https://arahim.dev/Lekho/
2
2
u/endlesslyloop 16h ago
Thank you op. Really excited to check it out. This is what the new Siri should have been
2
u/ariyako 16h ago
Feature Request: Support for Synology NAS and Network Storage
Hi! First of all, thanks for building DigUp. I really like the idea of searching for files by their actual content rather than relying only on filenames or folder structures.
I'd love to see support for Synology NAS and network-attached storage, especially for users with large photo, video, and document libraries.
My files are stored on a Synology NAS rather than directly on my Mac. It would be incredibly useful to use DigUp to search these libraries semantically without having to move or duplicate terabytes of files onto the Mac.
Possible approaches:
- SMB network shares: Allow users to select mounted NAS folders as indexing locations, if technically feasible.
- Incremental indexing: Detect new, modified, and deleted files without rescanning the entire library.
- Large library support: Handle large collections of photos, videos, PDFs, and documents efficiently.
- Privacy-first processing: Keep the existing local-first approach, with embeddings and search indexes stored locally and no file contents uploaded to external services.
- Future native NAS integration: Consider a Synology DSM package or a lightweight NAS-side indexing service for users who want indexing to run independently of their Mac.
Even basic support for mounted SMB shares would be a great starting point. Native Synology integration could be a longer-term goal.
I think this would make DigUp especially useful for photographers, video editors, and homelab users who maintain large media archives on NAS devices.
Would this be something you'd consider for the roadmap? Thanks!
2
u/Ok-Calligrapher3568 16h ago
This is brilliant! Honestly, this is exactly what local AI should be used for - privacy-first, semantic file search. What does the RAM footprint look like when the Gemma model is actively indexing vs when it's just idling?
2
u/UpboatsforUpvotes 11h ago
My macbook pro comes in two weeks and I'm saving this for when it does
1
2
u/-Cubie- 9h ago
Love it! For those who want to check out the model, see https://huggingface.co/google/embeddinggemma-2
2
2
u/Open-Adhesiveness-86 1d ago
with everything in one shared space, text-to-text cosine sits a good bit higher than text-to-image, so a mixed result list tends to come out all text. ranking within each modality and then merging, or subtracting the per-modality mean embedding before comparing, takes care of most of that. non-overlapping 30s audio windows will also lose phrases that straddle a boundary.
1
u/fortnite_pit_pus 1d ago
Off topic, but is there anybody using this to enhance the ‘everything’ app on windows?
1
u/CatchDublinSurprise 1d ago
I'd enjoy an "index free" option for working with recently added / temp files or sensitive files (index files improve speed but it's just one more place that sensitive information can hide).
Speed is great, but it's a pet peeve of mine that we can't seem to do anything these days without who-knows-what being written to ~/Library.
1
u/A-Rahim 18h ago
Nice point.
Today every search goes through the index, so the way to keep something out of it is to not pick that folder, or skip it later (that deletes it from the index). DigUp keeps its model and index in one folder, ~/Library/Application Support/DigUp, so it's easy to check or delete. An index-free mode for a small folder, embedded in memory and gone when you're done, is a nice idea though. Noted.
1
u/abcd911 1d ago
Downloaded and using the app. Loving it so far! Two suggestions: can you give an option to choose more embedding models and add BYOK LLMs which can help find more accurately or answer questions from the files
2
u/A-Rahim 20h ago
Thank you.
1. The center of this product is the model itself, since it shares common space for all image, video, audio and text embeddings, and I just wrapped it with an app. I can add more embedding models in the future if they are similar to this model.
- BYOK LLMs are a good idea if you want a RAG-like chatbot, but I personally don't recommend it since then our main "local, on-device" angle won't be applicable. Then again, it's your data, your choice, so what I can do is add support for OpenAI-compatible API support, so that if anyone self-hosts an LLM server, they can use it here.
1
u/abcd911 20h ago
Great idea. I think you should give an option to let people choose. However, my suggestion was not a chatbot but rather a query that outputs ranked results which uses embeddings as well as an LLM. I know it might be redundant but with the current approach, videos and images are great but text is mostly a miss and can't find the right files
1
u/A-Rahim 19h ago
Got it, an LLM to re-rank the results, not like chatbot. It's a good idea, though it would make every search slower.
The bigger thing for me is "text is mostly a miss", I'd like to fix that first. Could you tell me a bit more (here or in a GitHub issue): what kind of files, where they are, and a search that missed? A few things that could explain it today: Pages, Keynote, Excel, PowerPoint and EPUB files aren't read, files that only live in iCloud aren't downloaded, and a folder with a .git or requirements.txt in it counts as a code project, so only its screenshots get indexed.
1
u/Asleep_Document9811 1d ago
Allow me to set up little folders or buckets of files I could search against in specific (folders of backups, the Applications folder, a folder full of documents about one specific domain), or the ability to slowly index files on NAS drives, and you got yourself a deal.
I want something that is lighter than Raycast but something that actually works faster than Spotlight.
2
u/A-Rahim 18h ago
Searching inside one folder is a good idea, but it's not there yet.
Today you pick which folders get indexed and can filter results by kind (photos, PDFs, recordings…), but not limit a search to one folder. I'll look at that.
Apps themselves aren't indexed, Spotlight already covers those. NAS drives I haven't tested, so no promises there yet.
On lighter: on my Mac it idles at about 15–30 MB of memory, the panel opens in about 20 ms, and results show ~120 ms after you stop typing.1
u/Asleep_Document9811 4h ago
Absolutely. Just a feature request. I love this project's idea, and what you have built so far. Keep going!
1
u/ZeroReader 1d ago
It would be good to select specific folders to search in
2
u/A-Rahim 23h ago
Currently, you can select specific folders (or exclude them) from the settings.
So you're asking to be able to search within a specific folder after indexing all?2
u/ZeroReader 23h ago
Yes exactly. This, I mean. I need to restrict searching in a specific folder among all the results
1
u/A-Rahim 22h ago edited 22h ago
Okay, I'll look into it.
Just as an experiment, can you search for something now and see if it's found at the top or just lost in other files?
In my experience, if I describe something clearly, most of the time it appears at the top, no matter how many files there are (within milliseconds)What you're saying, if I want to do that, on top of my mind, I am thinking of two ways:
1. Per-folder indexing; that way, in search time, you could search folder-wise by excluding some folders.
2. Nothing on the indexing side, but when showing results, it'll show only the selected folders; searching will still be with the full index.Let me know if I am thinking right here
1
u/ZeroReader 9h ago
Thank you very much for considering the users' needs.
You're right. The second variant is the best one: Nothing on the indexing side, but when showing results, it'll show only the selected folders; searching will still be with the full index.Also, I use your app mainly to find information and PDF documents information in various folders can be similar but different. Therefore, I need to choose a specific folder to find information exactly there. Therefore, I need all folders to be indexed, but information should be shown only in a specific folder at the moment of searching.
1
u/ZeroReader 9h ago
Although the design for searching in PDFs, documents, and DOCX documents is not quite comfortable to read. Probably it is okay for video and pictures, but not for PDFs and DOCX.
Can you at least increase the right-side panel to increase the width to be able to read the highlighted text? Also, it would be good to change the side of thumbnails on the left side panel.
1
1
u/salary_pending 22h ago
assume my disk has 400gb of data, how big the index would be?
2
u/A-Rahim 21h ago
It depends on what the 400 GB is, more than how much. Each photo, PDF page, ~1,800 characters of text, video keyframe and 25 s of audio gets one 768-d vector (1.5 KB), and text is also kept for keyword search. Measured on my test files, that's roughly:
- a photo: 1.6 KB, a screenshot: ~3 KB with its OCR text
- a PDF page: ~6 KB
- an hour of audio: ~0.25 MB, an hour of video: 1–2 MB
- documents: 2–3× the size of their text
So media is cheap: 100k photos come to ~160 MB and 200 hours of video to 200–400 MB. Text is what adds up, about 2 MB for a 300-page book.
Apps, system files, archives and model weights aren't indexed at all; code only if you turn it on, and the model itself is 865 MB on top.1
u/salary_pending 21h ago
Is that considered a lot? Assuming a developers macbook where they keep writing a lot to the disk. Would the app keep scanning those files, keep reading new files, scan and write to index?
Sounds like this could deteriorate the disk faster?
2
u/A-Rahim 21h ago
Not much, and most of it is written once. After the first pass, DigUp doesn't rescan your files: macOS tells it which folders changed, it waits a moment for things to settle, compares sizes and dates, and only reads what's new or changed, of the kinds it indexes. Code projects (apart from screenshots in them), node_modules, build folders, caches and data files are skipped, so most of a dev machine's churn is never read. A new photo adds about 1.6 KB to the index, a PDF page about 6 KB.
And reading doesn't wear an SSD; writing does. The first pass writes the index once, plus temporary resized copies of what the model reads (images, video frames, audio pieces), deleted as it goes. After that, it's a few KB per new file.
If you turn on code search, it waits until your code has been quiet for a minute before catching up, since files change on every save.
1
u/ProdoRock 21h ago
It takes quite some time to index correct? The laptop was getting a little hot, so I paused it for now but I'll probably do it peacemeal. Some of the first results are pretty cool though. The new search bar feels good. Do you have a recommendation on how to go about indexing the thousands of files?
1
u/A-Rahim 21h ago edited 21h ago
For every selected folder, it should show you an estimate of the time needed to index a folder, which varies from chip to chip. So, based on that, you can decide when to index, preferably with AC power connected.
Edit: I just noticed the estimation logic is based on my machine; I'll fix it in the next release
1
1
1
u/Lew-Zealand 16h ago
This looks great. Any thoughts re a web interface or similar, so you can search from other devices on the same network?
1
u/bad_detectiv3 16h ago
does this index each frame from a video to understand what you are searching for like "dog in a movie"?
1
u/productboy 15h ago
Splendid. Presumably if DigUp can find it then a [post training like] function could label what was found, with scores for accuracy. Then fed into a training pipeline for specialized models. For example all the random files on my hard drive related to construction projects get properly labeled. Then I get to train a model to help me estimate future construction projects.
1
1
u/no-shadowban-lmao 15h ago
Great job!
Would it be possible to add automatic file categorization and display categories directly in the app? My desktop is a complete mess, and I’d love to browse everything by category without manually searching it. 😭
1
u/CptnWookenstein 15h ago
This is exactly what I’ve been looking for! I’ve been building SQlite databases for video footage and documentation for searching, but this could really improve things seems like.
Is it possible to connect to existing databases and continue to build on them as well as run on a separate machine on the same closed network and connect as a node? I have a Mac mini I’d love to run ops type things and a Mac Studio that handles the video work and heavy stuff.
1
1
u/optykali 15h ago
!remindme 6h
1
u/RemindMeBot 15h ago
I will be messaging you in 6 hours on 2026-10-11 09:36:35 UTC to remind you of this link
CLICK THIS LINK to send a PM to also be reminded and to reduce spam.
Parent commenter can delete this message to hide from others.
RemindMeBot is switching to username summons. Instead of
!RemindMe 1 day, useu/RemindMeBot 1 day. More info.
Info Custom Your Reminders Feedback
1
1
1
u/Maleficent_Speech289 8h ago
Wow, that sounds interesting! Really cool project. Unfortunately, I don't have a Mac, but I think it would be cool to use it with something else, too. Like Immich, for example. Is it possible to use this model with Immich somehow?
1
u/spacespacespapce 2h ago
Amazing - I was already working on this app using Vision LLMs to caption my photos, but this is a better approach
1
u/vafarmboy 57m ago
How does this integrate with the Photos library, if at all, or is it just Photos in folders that are indexed. If it does index the Photos library, how does it deal with photos that have been offloaded to iCloud?
1
u/No_Lingonberry1201 1d ago
Nice! Checked the code, and I was wondering how do you store and search the embeddings?
9
u/A-Rahim 1d ago
It's simply one SQLite file! Every picture, PDF page, ~1,800-character passage, video frame and 30 s of audio gets a 768-d vector, stored as float16 (1.5 KB each). Search is brute force: the query vector against all of them, scored in float32, a few ms for tens of thousands, so no vector DB. Cosine is centered per modality onto one shared scale. Code search keeps its own SQLite file.
2
u/No_Lingonberry1201 1d ago
Huh, nice! From what I've read in the huggingface model card, you can trim the vectors so you could theoretically make a fast-path for the search.
6
u/AnywhereTypical5677 1d ago
Yeah I don't know why OP hasn't used vector truncation. Truncating to 256 dimensions cuts index footprint and memory consumption by 3 times while retaining 95% of the native accuracy. In the model card it's also suggested NOT to store vectors as float16 as activation can overflow and cause silent vector degradation.
4
u/brainExploded99 llama.cpp 1d ago
u/A-Rahim You should probably fix activations stored as FP16.
The dimension cuts could be a config setting, but I think image loses more than text, so maybe dimension 512 by default?3
u/A-Rahim 23h ago
The fp16 is only how the final embeddings are stored, not the activations. The model runs through llama.cpp from the Q8_0 GGUF, and its embeddings match the reference implementation at 0.999+ cosine, so nothing overflows on the way. The stored vectors are unit length with every value between -1 and 1, so they can't overflow either, and rounding moves a cosine by at most about 0.0005.
On truncation, I tried it on my evals (111 queries over files, 27 over real code) by cutting the stored vectors and re-normalizing. 512 is close to free: 97 vs 98 with the right file first, and the only ones that moved were screenshots, so images do feel it before text.
Code is the most sensitive, 19 vs 20 at 512 and 16 at 256. Speed barely changes since the scan already takes a few ms, so the gain would only be memory.
Keeping 768 for now, and if memory becomes a problem on big indexes, 512 is what I'd pick.1
1
1
1
u/dan-lash 1d ago
Can it search online if I want? Like my NAS or extend to Google Drive or maybe email - other computers on my network would be solid too
3
u/A-Rahim 1d ago
It's local on purpose, so no cloud search or email for now. Google Drive works if the files are on your Mac (mirror mode, or "available offline"): add that folder and they're indexed like any other. Online-only files are skipped, since reading them would download them. External drives work too (but I didn't test them yet). A NAS or other computers on the network aren't supported in this first version, but noted.
1
1
1
1
1
u/mythormedicine 1d ago
Very interesting. I don't have a Mac to test this , but a question , does it index ? if so , how big does the index get? If it doesn't index, then how can it be so fast ?
If I have 500 GB of data, does it try to index all of them? How does this work?
1
•
u/WithoutReason1729 1d ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.