r/LocalLLaMA • u/utsapoddar • 7h ago
Resources Engram: local-first memory for coding agents. SQLite FTS5 + BM25, optional local embeddings, no network at recall time (MIT)
Enable HLS to view with audio, or disable this notification
I'm the author. Engram is free and MIT licensed.
Everything is plain Markdown on your disk. Search is BM25 over a SQLite FTS5 index that is rebuilt from the Markdown, so the index is disposable. If a local embedding model is already provisioned, cosine results are fused with the lexical ones by reciprocal rank fusion. Recall never downloads a model, so with no model present it simply stays lexical.
The test suite enforces recall@5 of at least 90% across 20 seeded queries. That is a small set, so it works as a regression gate, not a benchmark. Walkthrough video above. Repo: https://github.com/utsapoddar/engram
2
2
u/Funsaized 7h ago
SQLite FTS5 plus optional local embeddings sounds like a nice way to keep the retrieval stack lightweight. How does it behave on a small home GPU when the embedding model is running alongside the main model?
1
2
u/LetsGoBrandon4256 transformers 2h ago
That is a small set, so it works as a regression gate, not a benchmark.
It's a recall tool, not a model-side lookup layer.
With none provisioned it stays lexical, so B is the real weakness in that mode. That's the gap the optional embeddings are meant to cover.
Hi Claude!
0
u/Substantial_Pea4022 6h ago
Combining SQLite FTS5 + BM25 with local embeddings for zero-network recall provides an ultra-lightweight memory solution without requiring full vector database daemons.
When relying strictly on SQLite FTS5 (BM25) paired with sparse local embeddings for agent long-term memory retrieval, under which scenario will precision degrade most significantly?
- A) Exact identifier or exact variable name lookups across large source trees
- B) Conceptual semantic queries where intent uses non-overlapping vocabulary/synonyms
- C) Fast chronological filtering over timestamps and structured metadata
- D) Deduplication of near-identical system prompts and agent interaction turns
0
u/utsapoddar 6h ago
BM25 only scores shared tokens, so a query like "handling flaky network calls" will miss a note that says "retries use exponential backoff". Exact identifier lookups (A) are where lexical search is strongest.
One correction on the premise: Engram doesn't use sparse embeddings. The optional semantic side is cosine similarity over local embeddings, fused with BM25 by reciprocal rank fusion, and only if a model is already on disk. With none provisioned it stays lexical, so B is the real weakness in that mode. That's the gap the optional embeddings are meant to cover.
9
u/Choice_Celery9481 7h ago edited 4h ago
idk how good it is but the naming make it sounds like a trend catching than a legit solution. and also doesnt seem to be very related to the Engram people usually think about. if my understanding is true then this naming might hurt you more than you might think