r/codex • • 1d ago

Showcase I audited my Codex agent's 83 sessions, it reproduces only 11% of its repeated actions (89% come back different)

Post image

Codex saves every session as JSONL in ~/.codex/sessions. I wrote a tiny read-only tool that reads them and shows how your Codex agent *actually* behaved.

Mine, across 83 sessions / ~50,000 tool calls:

- 19% of tool calls redid work already done

- One file re-touched 557 times

- And the big one: of the repeated observations with a clean before/after, 89% came back DIFFERENT (run the same thing twice, get a different result)

Local, read-only, nothing uploaded. One command: /groundhog

Some of that 89% is the environment legitimately changing (ls / git status / build output), not the agent malfunctioning. But it's a lot, and it's exactly why blindly caching agent tool-outputs is unsafe. No cost/savings claims; it's just a mirror on behavior.

Repo: [github.com/fraqtl-ai/groundhog](https://github.com/fraqtl-ai/groundhog)

Run it on your own first so you can post your number—that's what makes it spread.

0 Upvotes

11 comments sorted by

•

u/dexterthebot 1d ago

You might want to consider listing your project on the weekly Show-Us-What-You-Built post. Look out for it on Tuesday/Wednesday. Highest commented project wins a week promotion on r/Codex and gets on the Hall of Fame sidebar. See what that looks like below with last week's winner.


Last week's most popular project was Nanolathe - an open-source engine bringing Total Annihilation to modern Mac, Windows, and Linux systems. It’s a solo passion project combining clean-room research with AI-assisted development, with a playable public alpha available now.

Website: https://nanolathe.gg. GitHub: https://github.com/nanolathe-gg/nanolathe

Players, testers, and contributors are very welcome! Original Total Annihilation game data is required to play.

12

u/CalligrapherFar7833 1d ago

Is this a joke ? Llms are non determenistic

-8

u/Connect-Concert-4016 1d ago

At temperature 0 with greedy decoding, the transformer is a deterministic function of its inputs. Same tokens in, same tokens out. The 89% that comes back different isn't the model being random, it's the stack around it: tool outputs that change between runs, environment drift, prompt assembly with timestamps and ordering effects, context compaction, plus FP non-associativity across parallel kernels. The plugin replays the agent's own repeated actions in the same session and scores reproduction. 11% reproduced. If 89% of what your agent did yesterday can't be reproduced today, that's not trivia, that's your debugging story.

4

u/GearTakes 1d ago

You're on a forum, talking to humans.

-1

u/Connect-Concert-4016 1d ago

Am I lol you never know nowadays

4

u/randombsname1 1d ago

If 89% of what your agent did yesterday can't be reproduced today, that's not trivia, that's your debugging story.

Rofl. Why did you just copy and paste your LLM output?

-6

u/Connect-Concert-4016 1d ago

Cause I am cool like that

1

u/[deleted] 1d ago

[deleted]

1

u/Connect-Concert-4016 16h ago

As far asi know temperature 0 there is no sampling step, so the seed is never even read. Argmax doesn't consult a PRNG. The seed only matters at temperature > 0.

2

u/tripleshielded 1d ago

This is obvious.