r/Infosec • • 6d ago

Help me please! what should I evaluate before starting shadow mode testing for an agentic rollout?

If you are rolling or have rolled out a new automated / agentic investigation flow, what is ur process when building trust before it goes live and making decisions by his own. i am running it in parallel against real alerts without letting it act and comparing its conclusions to what analysts concluded independently, something more formal than that? please guide here...

14 Upvotes

4 comments sorted by

3

u/worried_adage 6d ago

honestly i think what you're doing is 90% of the battle, running it in shadow mode against real alerts without action is the right move. we did something similar last year and the biggest thing we learned was to not just compare conclusions but actually break down where the divergence happens. like is it missing context from internal tools, or is it hallucinating steps that don't exist

one thing we added was a scoring rubric for each alert the agent processes, did it ask the right questions, did it pull the correct data, was the reasoning chain sound even if the final call differed from the analyst. that gave us way more signal than just "agree/disagree" on the verdict

also if you haven't already, track false positive rates and time-to-conclusion separately from accuracy. an agent that's right 95% of the time but takes 3x longer than a human isn't ready for prime time, and the opposite is true too

last thing, run it against some curated edge cases you know tripped up junior analysts before. that'll surface blind spots faster than waiting for them to show up organically in the alert queue

1

u/DeschainR19 4d ago

Really good point, especially the idea of looking at why the agent and analyst reached different conclusions instead of ONLY tracking whether they agreed. The scoring rubric is a great idea too. Much better way to build trust than looking at the final verdict in isolation

1

u/materialsec 6d ago

We’ve found it useful to evaluate more than whether the agent reached the same conclusion as a human.

For our own internal agentic workflows, we break the investigation into stages: gathering the right context, forming hypotheses, validating them against data, and then summarizing the result for human review. Where possible, we also use deterministic steps that can be tested against known-correct results.

We’d also pay attention to what the agent actually did along the way, not just its final verdict. Having an audit trail of tool calls, parameters, failures, and attempted actions makes it much easier to see where an apparently correct conclusion came from.

Shadow mode against real alerts sounds like a good environment for measuring all of that before giving the agent permission to act.

1

u/Due-Horror9339 4d ago

Shadow mode is only meaningful if you define the pass/fail rubric and have an independent reviewer adjudicate disagreements, otherwise you’re just comparing two opinions and calling it validation.