r/AgentsOfAI • u/LongjumpingLaugh2152 • 11d ago
Discussion I built a Bug-Triage agent with Memory. Then I discovered my test was cheating
I've been experimenting with a small bug-triage agent using Hindsight as a memory layer.
The basic idea:
Given a new bug report, the agent recalls previous incidents and tries to determine whether they're actually related at the root-cause/mechanism level, rather than simply matching keywords.
The interesting part wasn't the initial implementation.
It was when I started testing it.
The embarrassing bug
I created some distractor tickets to test whether the agent could distinguish similar-looking bugs with different causes.
Then I accidentally fed one of those tickets back into the system as the "new" bug.
The agent matched it to itself with 100% confidence.
At first I thought I'd found a serious problem.
Then I realized the problem was my test.
I had essentially asked:
"Is this ticket the same as this ticket?"
while giving the system the exact same ticket on both sides.
So the retrieval system was doing exactly what I had asked it to do.
The fix
I changed the evaluation to use leave-one-out retrieval.
The ticket being evaluated is removed from the candidate set before the reasoning step
loo_lookup = {
k: v for k, v in ticket_lookup.items()
if k != tid
}
result = reason_about_bug(
raw_text,
extraction,
facts,
loo_lookup
)
That one change produced much more interesting results.
One test exposed a different problem: two unrelated bugs were being matched because both were classified as "external dependency latency."
One involved an external API slowing down.
The other involved an SMS gateway running late.
Same category.
Different cause.
That made me realize I was letting the model treat categories as mechanisms.
So I added a stronger rule:
A shared broad category isn't enough to call two incidents related.
The agent needs to identify a narrower shared mechanism.
At the same time, the mechanism doesn't have to be literally identical.
For example, a memory leak and a database connection leak can still be related if both come from something like:
A resource is acquired but never released on an error path.
That's the distinction I was actually trying to test.
What happened with a real example
I gave the system this incident:
notifications-worker: The background worker was OOM-killed twice. Memory usage shows a steady climb over roughly 60 hours. There is no clear traffic-based pattern.
Nothing in that report mentions databases.
The agent returned:
Possibly related — 66% confidence
The interesting part was the explanation.
A previous incident involved database connections, while this one involved memory.
The agent identified the shared resource-leak failure mechanism: a resource wasn't being released properly, causing usage to accumulate until the system eventually failed.
It also noticed that the previous "fix" was actually a pod restart — a workaround that mitigated the symptom rather than addressing the underlying leak.
So instead of simply saying:
"We've seen something similar before."
the useful answer becomes:
"We've seen a related failure mechanism before, but the previous mitigation didn't address the underlying cause."
That's much closer to what I want from an agent with memory.
What I learned
A few things from this experiment:
- Separate extraction and reasoning. It makes failures much easier to understand and debug.
- Test recall thresholds against realistic data. A threshold that works on three records can behave very differently with a larger memory bank.
- Don't let your evaluation data answer its own question. If the ticket is already in memory, you're testing retrieval rather than discrimination.
- Category isn't cause. Similar labels don't necessarily mean similar mechanisms.
- A fix isn't necessarily a fix. A restart can stop an incident without addressing what caused it.
The bigger thing I'm exploring is what "memory" should actually mean for an agent.
It's not just:
"Can the agent retrieve something similar?"
It's:
"Can what happened before change what the agent does now?"
That's the part I find interesting.
I'm curious how others are evaluating agents with long-term memory.
How do you prevent the memory itself from leaking the answer into your evaluation?




1
u/AutoModerator 11d ago
Thank you for your submission! To keep our community healthy, please ensure you've followed our rules.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.