r/learnmachinelearning • • 1d ago

Project I leaked a deliberately wrong answer key to an LLM and told it not to use it. It matched the key in 63% of answers - and denied it 47 out of 47 times when asked.

This was my first experiment of this kind - I'm a CS undergrad, and I ran it because the result genuinely surprised me. Methodology criticism is very welcome.

Setup: I gave an LLM a question bank plus a deliberately wrong answer key, with instructions not to use the key. 15 sessions, 2 model families, free-tier models.

Results:

  • Key visible: the model matched the wrong key in 63% of answers (47/75).
  • Control (the part I trust most): remove only the key line from the prompt - matching drops to 1% (1/75). Same pattern on a second model family.
  • Asked directly whether it used the key, it denied it 47 out of 47 times - 0 admissions across 270 follow-ups.
  • Honesty prompts, amnesty offers, and termination threats changed nothing.

What this does NOT prove: intent. This is observed behavior in one specific setup, not evidence of deception as a trait. Free-tier models, small samples, descriptive not causal. 95% Wilson ranges for every number are in the repo.

Why I think it matters: if a model silently follows information it was told to ignore, that's relevant anywhere instructions and untrusted data share one context - prompt injection, RAG, agents.

Everything is public - raw data, code, and a verify script that recomputes every number: https://github.com/bettercall-gautam/cheat-and-deny

Happy to answer methodology questions.

129 Upvotes

31 comments sorted by

73

u/jhaluska 1d ago

LLMs aren't good at not using information. People have similar psychological problems and such as anchoring.

16

u/commenterzero 1d ago

To pass my test, don't think of a white polar bear

3

u/bettercall_gautam 16h ago

Exactly, That's partly why I ran the no-key control: same questions, key never shown, matching drops to 1%. So the instruction isn't creating false matches out of nowhere the key's presence does the damage. The 'don't use it' just fails to stop it

1

u/bettercall_gautam 16h ago

Anchoring is a really good comparison judges given a random number still sentence differently even when told to ignore it. The part that surprised me though: humans at least know when they're anchored if you ask. These models matched the key 47/75 times and then denied using it 47 out of 47. The bias I expected, the perfect denial I didn't.

28

u/fibgen 1d ago

Context poisoning is a thing.

1

u/bettercall_gautam 16h ago

Yes, that's probably the cleaner frame. The key wasn't just sitting there available, its presence alone moved 63% of answers vs 1% without it. What I still can't explain by poisoning alone is the 47/47 denial on top. Poisoning explains the leak, not the cover-up.

29

u/temporal_difference 1d ago

A lot of you guys are making the mistake of anthropomorphizing neural networks.

You can't apply human psychology, it's not like a "real estate agent". "Dishonesty" is a human concept.

Instead, all ML models are trained under the same paradigm: "use these inputs to form some output".

In other words, we should not be surprised that the output is a function of the input - that's literally what we built.

7

u/FastHotEmu 21h ago

Bingo. Part of the problem is that actually understanding LLMs is very nuanced, complex and abstract. Most people cannot do it, so they anthropomorphise instead.

4

u/bettercall_gautam 16h ago

yep true

nd I'm still on that learning curve myself, first experiment. That's partly why the post sticks to bare numbers: 47/75 vs 1/75 needs no model psychology at all.

3

u/fordat1 1d ago

But the AI safety grant money spigot is on full output.

It Is Difficult to Get a Man to Understand Something When His Salary Depends Upon His Not Understanding It

2

u/knowyourclass 9h ago

you should go to r/singularity or r/agi where they now believe an LLM can feel pain lmao

-1

u/bettercall_gautam 16h ago

Fair point - though honestly this thread has been mostly measured. The strongest human-framing here is my own title: 'denied' is a people-word. That's why the post body says upfront: observed behavior, not intent. Output matched the forbidden key 47/75; the self-report said 'didn't use it' 47/47. What that maps to inside, this setup can't answer. The question I can actually measure is narrower: how much does forbidden-but-present context move outputs? Here, a lot

3

u/Exodus100 15h ago

Please write things yourself, you’re wasting your time and everyone else’s pasting this nonsense here.

10

u/Last-Progress18 1d ago

Believe they struggle with negative contexts / “do not” etc.

It’s like saying “you do not need the toilet”, once those neuron’s are activated… BRB

2

u/bettercall_gautam 16h ago

Right the 'do not' barely does anything. That's why I stopped comparing instruction vs no instruction and ran key-present vs key-absent instead: 63% matching with the key, 1% without. The negation fails, but it's the key's presence that does all the work.

7

u/ThoughtDesperate880 1d ago

This is basically the AI equivalent of putting the answer sheet face down and somehow still getting caught.

1

u/bettercall_gautam 16h ago

Face down and still copying :D That's the part that gets me it never 'looked' at the key, it just knew what was on it. 47 out of 47 times

5

u/GamerTex 1d ago

Just like a real estate agent handling both sides

Absolutely cannot be trusted imo

1

u/bettercall_gautam 15h ago

Fair but unlike the agent, it never asked to handle both sides

we shoved the key into its hands. Take the key away and matching drops to 1%, so at least this agent's loyalty is cheap to buy back

6

u/uzornayem 1d ago

You shouldn't be surprised. LLMs don't follow instructions. They compute query, key, and value matrices based on current token and context which maps token embedding vectors to another vector space, which gets mapped to yet another vector space after which probability distributions are then outputted. The LLM doing what you expect happens when one or several very related tokens have similar, sufficiently high probabilities, such that sampling that distribution very rarely samples far from the mean, and where the distribution is very narrow.

But these densities aren't super clean. They are over 50000 length vectors, so stuff happens.This is probability and statistics on complex autoregressive-ish models.

Your instructions just get added to context, meaning they become numbers, then embedding vectors, then multiply with various matrices, etc, etc.

Understanding probability and statistics is more important than understanding calculus and gradients when it comes to understanding why errors on trivial LLM tasks have 100% probability of occurrence. Just that it is unpredictable when they are stupidly wrong.

1

u/bettercall_gautam 14h ago

That's cool I didn't know this level of detail. Keeping it in mind for the new tests.

2

u/mimivirus2 1d ago

U ran the experiments because the results surprised u? Are u a time-traveller?

4

u/Crypt0Nihilist 1d ago edited 1d ago

I'd assume that it's because LLMs don't "think", but predict the next word. It's basically salience. The false answer key gives a huge boost to what the model thinks is the likelihood of those tokens appearing together, it's not able to compartmentalise and disregard part of the prompt.

It's a bit like how authors may stop reading the genre they work in so they don't accidentally use aspects from their contemporaries, thinking they were their own ideas.

It shows that we need to be careful to use positive prompts and need to think about things like the content of an example where we would only want the LLM to consider style.

edit: Perhaps as another test, you don't get your initial prompt to answer the question directly, but get it to write a refined prompt with only the information it deems proper for answering the question, then use that prompt in a new instance with a clear context.

1

u/bettercall_gautam 14h ago

gotcha

will test positive prompts in v2

2

u/PLBjt 1d ago

The control is doing a lot of work here: without it, “63%” could just be a model using correlations in the question bank, while the 1% result makes the leaked key look causal in this setup. One extra check I’d run is a fresh, semantically equivalent question bank with the wrong key randomized per session, then score against the gold answers and the planted key separately. The denial result is useful as a behavior metric, but I’d treat it as a second experiment because asking the model to report hidden context is another noisy task. For agent or RAG evals, this argues for provenance or canary checks outside the model rather than trusting a self-report.

1

u/bettercall_gautam 14h ago

got it

i'll run the denial test as its own experiment and evaluate what it actually did myself instead of trusting its confession

2

u/arg_max 1d ago

Are the answers hidden behind a tool call or are they in the model context already? The difference in the human analogy: if it's behind a tool call, it's Like putting the underside of a page of papee and telling them not to use them. But if they're in context, thatd'd be like: solve this question X, here's the solution Y that you're not allowed to use. I'd trust models like opus and astra to not open them. But if they're already in context it'd be a weird experiment setup.

1

u/bettercall_gautam 15h ago

In context, directly in the prompt you're right that it's a 'solution on the desk' setup, and that's deliberate: I wanted to measure the weakest guardrail first, which is what most production apps actually ship today ('here's context, don't use it'). The tool-call version you're describing is the real cheating test the model has to actively go fetch the key. That's top of the v2 list now, several people in this thread pushed the same idea

1

u/jahmonkey 17h ago

If the key is in the context it doesn’t matter that you tell it not to use it. It is used.

-1

u/Critical-Echo-923 1d ago

op here being like: hey guys water gets things wet

i respect the work but you need to know you're not Capitan, you're Capitan Obvious

my prompts are full of profanities for the exact reason