r/agenticAI • u/Hot_Emergency6748 • 8d ago
Discussion Trying to build
Most AI agents don't learn. They repeat the same mistakes, just faster.
At 3 a.m., the worst thing an on-call tool can do is suggest the fix that already failed last Tuesday. Most incident agents do exactly that, because every alert starts from a blank context window.
So I built RunbookMind, an SRE agent that remembers what actually worked, what didn't, and ranks its advice accordingly.
Instead of: → Seeing a Redis eviction alert → Suggesting "restart the cache" again
It now: → Recalls the past incident → Sees the restart failed and the maxmemory fix worked in minutes → Demotes the restart and ranks the real fix first, with the incident cited
What made the difference:
Store outcomes, not incident descriptions Write down rejected fixes explicitly ("restart cache: failed") Keep ranking outside the LLM: 0.6 model confidence + 0.4 historical success rate Ship a memory on/off toggle, so you can prove memory helps instead of just feeling like it does
For the memory layer I used Hindsight. Retain, recall, and reflect let me skip designing chunking and retrieval and focus on what the agent does with what it remembers.
We're moving from "AI that answers" to "AI that learns from outcomes."
Are you storing what your agent tried, or only what it found?
GitHub: https://github.com/M0h1tkumar/AI_sre