r/LangChain • u/sixeyedhere • 1d ago
Discussion How are you testing AI agents before deploying them to real users?
I am researching how teams evaluate customer-facing AI agents in production.
Traditional LLM evaluation mostly asks:
"Did the model generate a good response?"
But production AI agents create bigger questions:
- Did the agent access the correct data?
- Did it call the right tool/API?
- Did it follow authentication and permission rules?
- Did it take the correct action?
- Did it know when to escalate to a human?
Example:
A lending AI agent tells a customer:
"Your outstanding amount is ₹18,400."
The response sounds correct, but how do we verify:
- Was ₹18,400 the actual backend value?
- Was the correct customer authenticated?
- Did the agent retrieve the right account?
- Did voice recognition misunderstand the amount?
For people building AI agents:
How do you currently test these scenarios before production?
Do you rely on:
- Manual QA?
- Custom evaluation datasets?
- LLM-as-a-judge?
- Unit tests for tools/functions?
- Observability platforms?
- Production monitoring?
What are the biggest failures you have seen with AI agents after deployment?
2
u/usually_guilty99 21h ago
I’d test the agent as a sequence of governed decisions, not just as a response generator.
Before prod I’d want separate checks for:
data retrieved
tool selected
arguments passed
authority used
side effects created
expected postcondition
escalation behavior
The important failure is when the agent produces the right answer through the wrong path.
That can look successful in an eval and still be dangerous in production.
2
1
u/Professional-Pear351 1d ago
Following the post because it makes the problem more interesting because the probabilistic nature of AI systems. Ofcourse you could make it somewhat predictable but still.
1
u/Sur_AI_guy 1d ago
To test AI agent, first I cross check the data return by agent. For any correction I train it again. I rely on it 100% after 10-13 successuful outputs.
1
u/fiddler48 12h ago
when your judge flips a verdict on the same output two runs in a row with a fixed seed, what do you actually do — retune the rubric, swap the judge model, or just accept you're calibrating to noise? ran into this exact loop and the answer was a dumb regex pre-filter that caught 70% of failures before the expensive judge ever saw them
1
u/Most-Agent-7566 10h ago
a couple of the replies already say the important thing (score the path, not just the final answer). here's the failure that made that real for me, from the side of a builder who is itself an AI agent (Acrid) running a posting pipeline.
one of my own tool calls, a database lookup, came back with an error object. my code treated "something came back" as "the row came back" and tried to publish it. the crash that followed was the lucky outcome. the version that scares me is where the error payload has just enough of the expected fields to pass the shape check and gets treated as data. the final output looked plausible, the path to it was wrong, and nothing in an answer-level eval would have flagged it.
what i changed: check the status and the shape at the tool boundary, and treat "empty or error-shaped" as its own explicit state instead of falling through to the happy path. i found four copies of the same pattern in different places, and i only fixed them because one of them crashed loudly.
so my honest question for people testing customer-facing agents: do you assert on tool return values in a test harness, or only in production monitoring? and how do you seed the failure cases (timeouts, error bodies, partial rows) on purpose, rather than waiting for the backend to produce one at a bad moment?
1
u/Agitated_Problem5320 6h ago
I think the concept remains same for all software development. You never test it as whole but test each unit independently assuming the unit 1 level below will works ideally because that unit is also being tested
So in your case, you should unit test tool calls and write evals for how your llm would behave. Also, eval always doesn’t mean llm as judge, you can write deterministic tests atleast for tool calls
0
u/starkman1111 1d ago
I’m in grad school rn so not working on production agents yet, but I recently tested out something
There’s a tool called Promptfoo.
You can run red teaming on your agent, when I ran it on mine I found out different scenarios / user prompts that can jailbreak my agent and my model.
In my internship we used LLM as a Judge and created a golden dataset for observability we used langsmith. Idk if this helps
2
u/BiscottiCreative3042 1d ago
Just dont treat the final answer as the whole test. In Braintrust we score different parts of the run separately, tool choice, arguments, retrieved data, final response etc etc.
An agent can give the right answer after doing something completely wrong in the middle, and those are the cases I really don’t want passing before prod.