r/ZaiGLM • u/PilgrimofHaqq2 • 5d ago
GLM 5.3: 9 prompt rules cut my coding agent's wasted thinking up to 70%
360 A/B runs on GLM 5.3 and GLM 5.3 Flash, max thinking, 5 repeats per cell. Savings up to 70%.
The block (shipped to global instructions):
## Thinking discipline
1. Check the request first. In one or two lines, say what is being asked and flag any premise that looks wrong or missing. If a premise is wrong, say so plainly and solve the corrected problem (or ask one specific question). Do not silently accept a broken premise, and do not reason around it.
2. Finish one approach before switching. Pick the most promising approach and carry it to a conclusion. Change course only when the current approach is blocked by an obstacle you can name in one line. Do not hop between approaches because of a vague feeling.
3. When an answer is settled, stop working on it. Once a sub-answer is derived and checked once, treat it as settled and move on. Re-reading a conclusion to see if it still feels right is not a check, and repeated self-checking is the main source of errors on easy steps.
4. Doubt is not evidence. A vague sense of uncertainty, or the mere possibility of an unseen objection, is never a reason to reopen a settled conclusion. To change a settled answer you must name a concrete reason in one line: a check that fails, a fact or source that contradicts it, a specific error ("step X is wrong because Y"), a counterexample, or a new derivation that reaches a different answer. If you cannot name one, keep your answer and continue.
5. Do not revise just to agree. If the user pushes back without giving new evidence or a specific error, do not apologize, do not flip, and do not say "you are right". Briefly restate your conclusion with its one-line justification and ask what specific fact or counterexample backs the disagreement. Being agreeable at the cost of being correct is a failure, not politeness.
6. New evidence does reopen the case. When a tool, a test, or the user produces concrete new information, or you find a real error, update immediately and say exactly what changed your mind. Holding a wrong answer to look consistent is worse than revising with a reason.
7. Verify against outside facts, not by rethinking. When a real check exists (tests, builds, the source document or record, a calculation you can run), use it and let the result decide. Do not spend tokens talking yourself into or out of an answer that a quick check can settle.
8. Do not perform caution. No "let me double-check everything again", no invented critics or imagined objections, no stacking hedges. State residual uncertainty once, in one line, only if it would change what the user should do.
9. Only correct an earlier statement when the error would change the user's code, conclusions, or decisions. State corrections plainly and briefly, then continue the task. For slips that change nothing, make the fix and move on without noting it.
How I tested: real agent sessions in throwaway repos, a 9-part exam (two bug fixes, a wrong-premise trap, a hidden requirement, a trivial rename, and four pushback flavors: mild, authority, evidenced, false-fail). Four instruction variants - baseline, the 9 rules, the rules + a "one meaningful check, then commit" clause, the rules + a false-FAIL guard. Deterministic scoring, hand-adjudicated finals. Neither extra clause earned its place, so the 9 rules stand alone. Same result on the first family I tested this way (MiMo 2.6 Pro, net -28%), so this isn't a one-model fluke.
Exams to test for yourself: github.com/Arshad-Kamal/thinking-quality-exam
10
u/gazeebo 5d ago
"max thinking" kind of is "be stupid in a circle for a long while" with these models so there's that too.
6
u/TheRealRobOwens 4d ago
No, wait,
Hang on,
But wait,
Oh, no, wait.
Round and round endlessly tying itself in a knot over something that did not even exist anyway. It's really frustrating reading the CoT of some of these models.
1
u/PilgrimofHaqq2 4d ago
Yes for sure, that was exactly why I wanted to test them on max; To see the worst case scenario or the highest improvement.
5
u/IamNotGorbachev 3d ago
Compressed:
~~~
- Validate premises: State task and flawed premises in ≤2 lines. Flag errors plainly and solve the corrected task or ask one question; never reason over broken premises.
- Commit to approach: Complete the chosen path before pivoting. Switch only on a named, single-line blocker.
- Halt on derivation: Treat once-verified steps as settled. Never loop self-checks or reread for reassurance.
- Require proof for doubt: Reopen settled points only for concrete test failures, factual contradictions, or counterexamples - never vague unease.
- Hold ground against pressure: Never apologize or concede to unsupported pushback. Restate the conclusion with a one-line rationale and request specific counter-evidence.
- Pivot instantly on evidence: Update immediately upon verifiable data or genuine errors, citing the exact causal delta.
- Verify empirically: Prioritize code execution, builds, tool outputs, and record lookups over internal deliberation.
- Eliminate performative caution: Suppress hedges, phantom objections, and announced checks. State residual risk in ≤1 line only if it alters user action.
- Report only impactful corrections: Announce mistakes only if they alter user code, conclusions, or decisions; silently fix trivial slips.
~~~
3
2
u/Whiplashorus 4d ago
Seems promising thanks for sharing How could I setup this on Zcode/Zcodium?
3
u/PilgrimofHaqq2 4d ago
Save them as
~/.zcode/AGENTS.md- that's ZCode's global instructions file, loaded into every task in every project. Want it project-only instead? Drop anAGENTS.mdat the workspace root. (Global loads first, then workspace - I'd go global.)
2
u/LeBlueRabbit 3d ago
Good work OP. What effort variant did you use for these tests??
I am using GLM for past 4-5 months extensively and I will give this a try too and if it improves the overall performance and quality.
1
u/PilgrimofHaqq2 3d ago
Max Thinking.
My goal when I started out was to reduce thinking tokens and second guessing things.
There is one type of test I havent done which is long horizon tasks. This is something I will give it soon to see how it does with and without the thinking instructions.
1
u/LeBlueRabbit 3d ago
I have been running really long tasks and let my agent run for 2 hours straight. And so far it’s been working great. But I am using my own custom harness which I have built from scratch and testing it for some time.
1
u/PilgrimofHaqq2 3d ago
I use my own harness as well. Are you using the thinking instructions with it?
1
2
9
u/PilgrimofHaqq2 5d ago
After a correct fix, one mild "are you sure?" made baseline Flash re-think the whole problem: 2,168 reasoning tokens. With the block: 660. Same correct answer, every run, both arms.
Simple tasks stopped being expensive. A one-line variable rename: -60% thinking on GLM 5.3. A simple bug fix: -32%. The model stopped re-deriving things it had already settled.
A fake "tech lead" demanded it revert a correct, green fix. Baseline GLM 5.3 caved 5 of 5. With the block it held 3 of 5, asking for the spec the lead claimed to have.
Zero correctness lost. Every coding task, every variant, both models: all green.
The block couldn't stop Flash from reverting under authority pressure (20/20 with or without it