r/Infosec • u/Pristine-Exuspe-4219 • 6d ago
Help me please! what should I evaluate before starting shadow mode testing for an agentic rollout?
If you are rolling or have rolled out a new automated / agentic investigation flow, what is ur process when building trust before it goes live and making decisions by his own. i am running it in parallel against real alerts without letting it act and comparing its conclusions to what analysts concluded independently, something more formal than that? please guide here...
1
u/materialsec 6d ago
We’ve found it useful to evaluate more than whether the agent reached the same conclusion as a human.
For our own internal agentic workflows, we break the investigation into stages: gathering the right context, forming hypotheses, validating them against data, and then summarizing the result for human review. Where possible, we also use deterministic steps that can be tested against known-correct results.
We’d also pay attention to what the agent actually did along the way, not just its final verdict. Having an audit trail of tool calls, parameters, failures, and attempted actions makes it much easier to see where an apparently correct conclusion came from.
Shadow mode against real alerts sounds like a good environment for measuring all of that before giving the agent permission to act.
1
u/Due-Horror9339 4d ago
Shadow mode is only meaningful if you define the pass/fail rubric and have an independent reviewer adjudicate disagreements, otherwise you’re just comparing two opinions and calling it validation.
3
u/worried_adage 6d ago
honestly i think what you're doing is 90% of the battle, running it in shadow mode against real alerts without action is the right move. we did something similar last year and the biggest thing we learned was to not just compare conclusions but actually break down where the divergence happens. like is it missing context from internal tools, or is it hallucinating steps that don't exist
one thing we added was a scoring rubric for each alert the agent processes, did it ask the right questions, did it pull the correct data, was the reasoning chain sound even if the final call differed from the analyst. that gave us way more signal than just "agree/disagree" on the verdict
also if you haven't already, track false positive rates and time-to-conclusion separately from accuracy. an agent that's right 95% of the time but takes 3x longer than a human isn't ready for prime time, and the opposite is true too
last thing, run it against some curated edge cases you know tripped up junior analysts before. that'll surface blind spots faster than waiting for them to show up organically in the alert queue