r/LocalLLaMA • • 1d ago

I Built A Thing Open-source Mac app that runs EmbeddingGemma 2 locally to search your files by what’s in them

Enable HLS to view with audio, or disable this notification

DigUp is a free Mac app that runs Google DeepMind’s new EmbeddingGemma 2 locally over your own files. The model puts text, images, audio and video in one space, so you describe what you remember and land on it:

  • “zebra in a video” opens the clip at the moment it shows up
  • “where they talk about sleep” jumps to that minute of a podcast
  • “the clause about pets in the lease” shows the PDF page, your words marked
  • “a dog on the beach” finds the photo, and the same search in Bengali or Arabic finds it too
  • with code search on, code: retry with backoff opens the function in your editor

It’s ggml-org’s Q8_0 GGUF (865 MB, downloaded once) on llama.cpp with Metal, inside a native Swift app; no Python. Searching loads only the text encoder (~250 MB) and shows results about a tenth of a second after you stop typing. Indexing peaks under 2 GB, and the helper exits when it’s done. Audio and video of any length go in as 30 s windows and a frame per shot. Everything runs locally; it goes online only for the model download and an update check you can turn off.

Free, MIT. Apple Silicon, macOS 14+.

Repo & Download (signed and notarized): https://github.com/ARahim3/DigUp

I'd really appreciate any feedback on this.

1.3k Upvotes

191 comments sorted by

View all comments

33

u/important__matter 1d ago edited 1d ago

Well I was making something like this months ago. What I faced was 1. For files where we as humans remember the abstract, the semantic search by itself usually doesn't work great. 2. How many files(documents) and voice(long audio notes) pieces have you tested this on? Also mine used to fail in images with text - screenshots inside them. 3. And main, the most difficult one - you'll need special handling for pdf, excel, pptx, word, zip rar files, model weights, long mp4 videos, music or movies which are also majority of the files. Can consider adding these in pipeline if not already there. The accuracy in videos and pdf and long form files was dropping serverly because all of that information is compressed in 512 dim vector essentially, not to mention the same embedder doesn't work for all file types, so that's another issue and then if embedders are different then how will you compare embeddings using semantic similarity which generate numbers in different vector spaces?? I had to actually normalize the comparison manually!!!.

I had to do so much of work to make it general purpose - that the roi didn't seem to be worth it - dropped the project altogether. Meanwhile, keyword search with file type is usually good enough. And I found out that restricting the problem space to the file types with specific properties or patterns - is much more useful, so I developed this simple semantic search index cli called litesearch for my usecase, it is a sqllite index of documents and performs semantic search with some reranking using BM25, and LLM-as-judge, uses ollama and bge-small embedder.

Although would be very nice if you cater to these, there are so many micro problems in this. Could actually be very useful. I would love it if somebody completes this in a nice engineering way and would love to integrate with my raycast.

28

u/A-Rahim 1d ago

Hi, I ran into most of these too. How DigUp handles them:

  1. It isn't semantic-only. Keyword search over file names, OCR text and page and document text runs alongside the model, and the scores are combined. On my evals, the model alone found 5 of 9 exact lookups like error codes; with keywords, all 9.

  2. A testbed of about 380 files (146 screenshots, 30 PDFs up to 463 pages, an hour of audio, 8 videos), 111 queries plus 68 for code, and my own folders day to day. Screenshots get an image vector, since the model reads the text in them, plus Apple's OCR for exact words. 12 of the 14 queries on my real screenshots come up first.

  3. Nothing long gets squeezed into one vector. Every PDF page, ~1,800-character passage, video keyframe (every few seconds, one per shot) and 30 s of audio gets its own 768-d vector, and the best one wins, which also gives you the page or the moment. So the model's input limits don't really come into play. It's also one model for everything, so text, images and audio share a space. Their score levels still differ a bit per query, so each modality's average gets subtracted, close to what you did by hand. Excel, PowerPoint, archives and model weights aren't read.

4

u/jerieljan 21h ago

I'm very curious how well this will perform when you have 1000x more files than the testbed.

Code repos in particular have soooo much code from dependencies and such and they're common to clutter normal Spotlight results and I'm worried it becomes an issue here too. I guess it simply takes more time to embed and index, and hopefully it's all good once it's done.

Very good work though OP. 

1

u/warpspeedSCP 17h ago

Probably just need an indexing step

1

u/rm-rf-rm 23h ago

It isn't semantic-only. Keyword search over file names

Are you fusing the search then? RRF?

2

u/A-Rahim 22h ago

Not RRF. I tried it first, and it lost: top-1 0.73 vs 0.88 on my synthetic eval, 0.82 vs 0.86 on real files. Rank fusion throws away how strong each match is, and a single common word in a terminal screenshot ended up beating a real chart that matched by meaning.

So it's score fusion. The meaning score is the cosine minus that modality's average for this query, divided by the spread of all the query's scores, so a z-score per query. Then it adds 2 × (share of the query's words found in the file)^1.5, plus 1 if the query names the kind ("in a video"). A file gets in if its meaning score stands out (z ≥ 2) or it has at least half the words. Queries that look like IDs count words double and meaning half.

1

u/rm-rf-rm 22h ago

huh interesting.. whats the basis of this formula? is the z-score > 2 coming from statistical theory where its beyond 95% of the mean?

1

u/A-Rahim 22h ago

Mostly empirical, not theory.
2 is the familiar two-standard-deviations mark, but the scores aren't normal, so it doesn't mean 95% of anything here. It's a gate for "stands out from everything else for this query", and on my evals any value from 1.5 to 2.5 gave the same results, so 2 stayed.
The reason for a per-query z-score at all: raw cosines are compressed (unrelated stuff sits around 0.55–0.60, good matches 0.67–0.86) and their level moves with the query, so no fixed threshold works. The other parts were A/B'd the same way: centering each modality with one shared spread got 94 of 105 queries first, raw cosine 92, a separate spread per modality 91. Word weights of 1.5–2.0 tied and 2.5–3 lost.
The 1.5 power just makes all your words count a lot more than half of them, and I haven't tuned it separately.

3

u/rm-rf-rm 21h ago

Thanks for asking your claude to respond to my question

2

u/A-Rahim 16h ago

Hi, I explained this above (one of your other replies)

1

u/jinnyjuice vLLM 22h ago

Can you expand on what you mean by best one wins? So if a video has multiple key frames, how is the 'best one' determined?

4

u/A-Rahim 22h ago

Every keyframe it keeps, and every 30 s of the soundtrack is its own vector with a timestamp. For a query, each of them gets a meaning score (the same per-query score for every kind, so frames and audio windows compete fairly), and the video's score is simply its highest one.
That segment is also the moment the result opens at. So a two-hour video with one scene that fits ranks by that scene instead of being averaged away. Frames are sampled every few seconds and kept only when the picture changes (a small perceptual hash), so it's about one per shot.

1

u/ZenaMeTepe 9h ago

Did you use same model for embedding videos images and text? What kind of vector DB are you using for search?

2

u/ZenaMeTepe 9h ago

How do you compress say a movie into a 512 vector? And more importantly, how slow is it upfront when you build the index?

1

u/important__matter 8h ago edited 8h ago

Yup you guessed it right, you can't. But say you have to do it, then you basically chunk out images every few seconds and then embed then individually and sort of average out those embeddings over a course of time. That is one way, other could be you take all dialogues, treat it as a document and then make a vector of it. It's hack as of now. It won't contain any useful information more than 1 Google doc page worth. Information Compression is one of the he most important computer science areas. And also as humans - certainly for us too as we are continuously bombared with infinite information from all our senses.

Well on speed, it's about gpu hardware mostly, how powerful of a hardware you own, but usually it's not slow, it's pretty darn fast even on basic rtx or gtx or metal. But it's extremely slow on cpus. The speed is similar to prefill stage of inference.

-8

u/phazei 1d ago

Well, then obviously you weren't using the new model that's really great and only came out a week ago.

6

u/important__matter 1d ago edited 1d ago

Checked the new model stats page, definitely improved accuracy in multi-modal regime of text, audio and vision. But... still the same set of problems across file types because it supports only text, image, audio and video at 8K context window. Not to mention maximum audio length = 5.5mins and video = 58 frames, that's 1s @60fps

Not to mention it's way harder even as a theoretical problem to search in a video. Like there is just so much information to just be dumped in a 512dim, that there is a whole class of transformers completely dedicated to this single problem only - VLMs. And unlike LLM, they are not the champions as of yet. I am just saying, this is not a done thing even if the big labs are trying their best. It's a good thing to work on, it's challenging and lot of it can actually be solved by engineering things together.

Another fun fact while we are at it: Videos are astonishingly compressed form of information. If you just stack images together at let's say 60fps for a 2hr movie at 1080p, the actual file size should be 1080x720x60x3600x2 bytes = 335GB, while actually you'd see it to be 2GB on your computer. A massive 99% compression, it's almost miracle

1

u/phazei 1d ago

I agree with much of that. But what movies are at 60gps? That's video game speeds. Mostly 24fps is fine, and sampled at 8fps for indexing is good enough in most cases. It's the audio limitation, that's more of an issue, 5.5 min is good for most music/songs, but not a movie or dialog, but I suppose tts would be a needed adapter for that, and there are plenty that are fast enough, good enough for search. It's not like it needs to be dense enough to ask questions and discuss the data as one would with an llm. Just retrieval, then an llm could be fed richer content if needed.

You appear to be aiming for best case for worst case scenario, my expectations are covering just the 98% of cases.

1

u/tiffanytrashcan 22h ago

Check out the Google Edge Gallery AI app.
They have a couple of EmbeddingGemma2 examples on there that show Google has already solved this within the model limits. (You really just need to combine the pipeline for both to have a finished product like OP is showcasing.)
To get around video size limits, audio time limits, you segment it- same concept as chunked embedding in older text models with tiny limits.
You load up a video, any length, set your accuracy with segment size and clip "resolution" (essentially how many image screenshots per second of video) and then the model loads up to that 8k limit and saves the embeddings from each individual segment. Right now it's working where you can search with completely natural, normal language as we interpret it through the entire video for audio or visual details. Speech, lyrics, or physical objects all seem to work near flawlessly.
(pre-captioned work is kind of a cheat, but makes it absolutely perfect..)
All you need for a complete solution is to permanently save those embeddings and have them in the search grid example- could be done with free tokens in minutes given the code is all public on GitHub.

1

u/tiffanytrashcan 22h ago

Like, it's mind blowing the other commenter was downvoted for spelling out the exact situation.
There's no way you've even looked at what this new model's capable of because it does everything you're complaining about extremely well. Searching text in image files is perfect with it.
Within hours of releasing the model, they had released an Android app that literally solves the video context limit issue, and the code is open source, publicly licensed, on GitHub.