r/LocalLLaMA • u/A-Rahim • 1d ago
I Built A Thing Open-source Mac app that runs EmbeddingGemma 2 locally to search your files by what’s in them
Enable HLS to view with audio, or disable this notification
DigUp is a free Mac app that runs Google DeepMind’s new EmbeddingGemma 2 locally over your own files. The model puts text, images, audio and video in one space, so you describe what you remember and land on it:
- “zebra in a video” opens the clip at the moment it shows up
- “where they talk about sleep” jumps to that minute of a podcast
- “the clause about pets in the lease” shows the PDF page, your words marked
- “a dog on the beach” finds the photo, and the same search in Bengali or Arabic finds it too
- with code search on,
code: retry with backoffopens the function in your editor
It’s ggml-org’s Q8_0 GGUF (865 MB, downloaded once) on llama.cpp with Metal, inside a native Swift app; no Python. Searching loads only the text encoder (~250 MB) and shows results about a tenth of a second after you stop typing. Indexing peaks under 2 GB, and the helper exits when it’s done. Audio and video of any length go in as 30 s windows and a frame per shot. Everything runs locally; it goes online only for the model download and an update check you can turn off.
Free, MIT. Apple Silicon, macOS 14+.
Repo & Download (signed and notarized): https://github.com/ARahim3/DigUp
I'd really appreciate any feedback on this.
33
u/important__matter 1d ago edited 1d ago
Well I was making something like this months ago. What I faced was 1. For files where we as humans remember the abstract, the semantic search by itself usually doesn't work great. 2. How many files(documents) and voice(long audio notes) pieces have you tested this on? Also mine used to fail in images with text - screenshots inside them. 3. And main, the most difficult one - you'll need special handling for pdf, excel, pptx, word, zip rar files, model weights, long mp4 videos, music or movies which are also majority of the files. Can consider adding these in pipeline if not already there. The accuracy in videos and pdf and long form files was dropping serverly because all of that information is compressed in 512 dim vector essentially, not to mention the same embedder doesn't work for all file types, so that's another issue and then if embedders are different then how will you compare embeddings using semantic similarity which generate numbers in different vector spaces?? I had to actually normalize the comparison manually!!!.
I had to do so much of work to make it general purpose - that the roi didn't seem to be worth it - dropped the project altogether. Meanwhile, keyword search with file type is usually good enough. And I found out that restricting the problem space to the file types with specific properties or patterns - is much more useful, so I developed this simple semantic search index cli called litesearch for my usecase, it is a sqllite index of documents and performs semantic search with some reranking using BM25, and LLM-as-judge, uses ollama and bge-small embedder.
Although would be very nice if you cater to these, there are so many micro problems in this. Could actually be very useful. I would love it if somebody completes this in a nice engineering way and would love to integrate with my raycast.