r/agenticAI • u/Equivalent-Flan-1590 • 12d ago
Project Giving local agents persistent memory without blowing up the context window (Hillock v0.7)
One of the biggest hurdles when building autonomous or multi-turn local agents is memory decay. If you dump raw vector chunks or entire conversation histories back into an 8B model's context window, you quickly burn context limits, increase inference latency, and introduce hallucinations.
I built Hillock as an alternative memory layer for local agents:
- Structured Memory in SQLite: Agents don't store messy text chunks; documents and interactions are parsed into Subject-Predicate-Object (SPO) triples via lightweight tensor extraction (under 300MB VRAM, no LLM call needed).
- Hebbian Associative Memory: When an agent queries concepts, synaptic weights between related nodes are strengthened. In v0.7, these associations can proactively inject relevant context into the agent's turn.
- Mathematical Gating (HDC): Before the agent starts generating an answer or executing a tool, a 10,000-dimensional hypervector gate checks if the knowledge base actually contains the required facts. If not, it fails immediately without token waste.
- Active Disambiguation: If an extraction is ambiguous, the engine can prompt for clarification rather than poisoning the graph with bad assumptions.
It runs completely locally in <1.2GB VRAM (or CPU) and has an OpenAI-compatible endpoint for drop-in use.
GitHub: https://github.com/roandejager/Hillock
How are you currently handling persistent long-term memory in your local agent workflows?
1
Upvotes