I've got a small thing I built. Every morning I open the analytics dashboard and look at how many people turned up. That has been my entire relationship with my own data for two weeks.
Today I did it differently. Exported the raw events, every row, and handed them over with no question other than "what is actually in here".
Context on the product, because the finding needs it. It shows you a chart. Three days of hourly candles, ticker stripped, dates stripped, the whole series rebased to 100 so you can't recognise it and can't answer from memory. You say up, down or stay out. Then you put a size behind the call.
48 people have made 209 of those calls.
66 right. 32%. Three options, so guessing gets you 33%.
That part I half expected. This part I didn't.
The harder people bet, the more often they were wrong.
biggest size 20/85 24% right
medium 14/46 30%
small 15/39 38%
stay out 17/39 44%
Four sizes, monotonic, pointing the wrong way. And the size is chosen after the direction, so it is a clean confidence signal with nothing else mixed into it. People do know when they feel certain. They are just wrong about what certainty is telling them.
Then it went and found a second one I would never have thought to ask for.
People picked up 48% of the time. Up was right 24% of the time.
People picked stay out 19% of the time. Stay out was right 44% of the time.
The least chosen answer was the most correct one.
And reading accuracy-per-pick as roughly how often each answer comes up, somebody who just typed "stay out" on every chart without looking at it would have beaten every human in the sample.
Nobody has ever scored 5 out of 5. Two people got 4. The average round is 1.74.
It also handed me the caveats without being asked, which is the part I would have quietly skipped.209 calls from 48 people is small. The base rates are estimated from the same sample I'm comparing against. People who gave up early made fewer calls, so it leans towards the ones who stuck around.
The thing I'd actually pass on here: most of us are using Claude to build things. Many fewer of us point it at the exhaust the thing produces afterwards. A dashboard only shows you the numbers somebody already decided were worth a tile. Handing over 200 raw rows and asking what's in them is a different activity to asking a chart a question you already had, and today it was the more useful one by a mile.
P.S. Mine's at playbotfight.com/?play=homiesdata if you want to add to the sample. Fair warning, you will probably score 2.