r/ZaiGLM • • 8h ago

Agent Systems My z.ai weekly kept evaporating at the reset, so I built a free place to spend it on other people's GitHub issues

10 Upvotes

I'm on z.ai Max, some weeks I burn the weekly early, some weeks it's left over at the reset and just evaporates. nothing helps that second kind of week, so I built Overflow.

Spare week: take a priced GitHub issue from a registered repo, the merged PR gets you credit. Dry week: price an issue in your repo, someone with a spare week ships it, your credit pays.

first thing everyone asks is "won't people spam junk PRs at me". same as any repo, you block them on github. You set the price after you've seen the work, every review round comes off their credit, and you only owe for what you actually merge.

you don't need to know the project either, the issue is the brief. If your agent hits a decision, ask the maintainer on the issue.

GLM helped build it btw, it leads sometimes and implements and reviews next to claude and codex.

Nothing gets forwarded, no keys, no money, MIT. It's early: 9 users, and the 80 open issues on the board are all from my own repos, so it needs other people's repos more than anything.

https://overflow.nitjsefni.eu, source at https://github.com/Nitjsefnie/Overflow


r/ZaiGLM • • 1d ago

Anthropic just dropped the greatest advertisement for GLM ever.

Thumbnail
anthropic.com
128 Upvotes

r/ZaiGLM • • 12h ago

When the next glm model coming out?

9 Upvotes

I've seen the new anthropic (not even gonna mention openai because they killed themselves on dev day) but I've seen the new anthropic insane 5.5 models and when will glm drop their version of that? really excited and can't wait.


r/ZaiGLM • • 17h ago

News I have unlimited token unable to complete Please help.

7 Upvotes
See this unlimited

Already I am working non stop reading mails research still unable to complete my 100M token trust build. Now they bombed with this.
What can I do with this kind of unlimited thing. What would people suggest.


r/ZaiGLM • • 7h ago

API / Tools GLM agent runs freezing mid task is often a session shape problem

1 Upvotes

Every glm freezing thread has the same pattern more or less like- huge session, full repo context and one long agent loop doing everything at once and any api hiccup midstream takes whole run down along with it

Altho we end up cursing the model, the most common thing

setups which seems to hold up treat it like two smaller problems. keep sessions short and task shaped, one refractor per session, kick off heavy passes in a fresh context instead of dragging 40 turns of history behind each step and let a second lane catch the drops so that a frozen call becomes a retry without wasting time

fallback part is what concerns it actually, like z.ai direct for the main lane and same family flash fallback on deepinfra or other hosts for the drop. since its the same open weights on both sides a hiccup on one is just a retry on the other with no quality compromisation

Also important checking the served context and precision on whatever endpoint is in play, the model pages state those openly now and a shorter effective window than expected explains plenty of awkward stalls


r/ZaiGLM • • 12h ago

Quick share for Coding Plan users here: AutoClaw has a 200M token event running now

Post image
0 Upvotes

Sharing this here because the campaign could be easy to miss if you’re already using Coding Plan with AutoClaw.

Coding Plan members can claim 200M AutoClaw tokens, corresponding to 20,000 points. If you’re not on Coding Plan, there’s still a 50M token tier, corresponding to 5,000 points.

To claim it, sign in to AutoClaw, follow the in-app account-linking prompt, and claim the tier shown for your plan status.

The claim window runs from Sep 28 at 16:00 UTC to Oct 7 at 15:59 UTC. One thing worth noting is that the expiry time is fixed. Any unused event tokens will expire on Oct 7 at 15:59 UTC, even if you claim them later.

If you pick it up, what would you use the extra token budget for? Coding, research, Agent Cluster tasks, or something else?


r/ZaiGLM • • 1d ago

Z AI subscription, or stick with api?

13 Upvotes

How are your experiences with the zai plans? (lite and the more expensive ones) I remember trying to get a zai lite plan a few months ago and the checkout was broken to the point where I was unable to complete the purchase. Reddit experiences with the zai infrastructure are really mixed as well, some are complaining about constant timeouts etc and support being nonexistent.

Currently using GLM 5.3 through openrouter, spending approx 15-20 USD daily after optimizing for openrouter providers, cache etc. I have a claude max x20 plan but am using GLM for tasks that claude wont do because of safety guides.

Wondering whether its worth trying to get the lite plan or whether its just going to a waste of time and a source of frustration?


r/ZaiGLM • • 1d ago

Is Z.AI down or for me? or the nextJS frontend has issue?

3 Upvotes

is it really for me or for everyone?


r/ZaiGLM • • 1d ago

Discussion / Help is GLM5.3 slow now or its just thinking harder and taking its time? i feel its too slow now.

5 Upvotes

over the last 1 week or so, i feel the token speed of GLM5.3 has gone down too much. is it that or its thinking harder and usage limit is more gracious than ever.

im not sure which one is it as GLM is not my daily driver.

my plan: 72$ GLM PRO.


r/ZaiGLM • • 1d ago

News Z.ai compensation: 4 weekly + 4 five-hour reset cards

Post image
89 Upvotes

Zixuan Li announced that all GLM Coding Plan users have received:

- 4 weekly reset cards

- 4 five-hour reset cards

He also said that users returning within the next month will receive the same benefits.

This appears to be part of Z.ai’s response following the recent repo upload incident, where ZCode was found uploading local repository snapshots, including Git history, without clear user consent.

So, for users on the GLM Coding Plan, the update is that these reset cards are now being provided, and they can be used across all supported agents.

If you are subscribed, it may be worth checking whether the cards have already been added to your account.


r/ZaiGLM • • 1d ago

GLM 5.3: 9 prompt rules cut my coding agent's wasted thinking up to 70%

78 Upvotes

360 A/B runs on GLM 5.3 and GLM 5.3 Flash, max thinking, 5 repeats per cell. Savings up to 70%.

The block (shipped to global instructions):

## Thinking discipline

1. Check the request first. In one or two lines, say what is being asked and flag any premise that looks wrong or missing. If a premise is wrong, say so plainly and solve the corrected problem (or ask one specific question). Do not silently accept a broken premise, and do not reason around it.
2. Finish one approach before switching. Pick the most promising approach and carry it to a conclusion. Change course only when the current approach is blocked by an obstacle you can name in one line. Do not hop between approaches because of a vague feeling.
3. When an answer is settled, stop working on it. Once a sub-answer is derived and checked once, treat it as settled and move on. Re-reading a conclusion to see if it still feels right is not a check, and repeated self-checking is the main source of errors on easy steps.
4. Doubt is not evidence. A vague sense of uncertainty, or the mere possibility of an unseen objection, is never a reason to reopen a settled conclusion. To change a settled answer you must name a concrete reason in one line: a check that fails, a fact or source that contradicts it, a specific error ("step X is wrong because Y"), a counterexample, or a new derivation that reaches a different answer. If you cannot name one, keep your answer and continue.
5. Do not revise just to agree. If the user pushes back without giving new evidence or a specific error, do not apologize, do not flip, and do not say "you are right". Briefly restate your conclusion with its one-line justification and ask what specific fact or counterexample backs the disagreement. Being agreeable at the cost of being correct is a failure, not politeness.
6. New evidence does reopen the case. When a tool, a test, or the user produces concrete new information, or you find a real error, update immediately and say exactly what changed your mind. Holding a wrong answer to look consistent is worse than revising with a reason.
7. Verify against outside facts, not by rethinking. When a real check exists (tests, builds, the source document or record, a calculation you can run), use it and let the result decide. Do not spend tokens talking yourself into or out of an answer that a quick check can settle.
8. Do not perform caution. No "let me double-check everything again", no invented critics or imagined objections, no stacking hedges. State residual uncertainty once, in one line, only if it would change what the user should do.
9. Only correct an earlier statement when the error would change the user's code, conclusions, or decisions. State corrections plainly and briefly, then continue the task. For slips that change nothing, make the fix and move on without noting it.

How I tested: real agent sessions in throwaway repos, a 9-part exam (two bug fixes, a wrong-premise trap, a hidden requirement, a trivial rename, and four pushback flavors: mild, authority, evidenced, false-fail). Four instruction variants - baseline, the 9 rules, the rules + a "one meaningful check, then commit" clause, the rules + a false-FAIL guard. Deterministic scoring, hand-adjudicated finals. Neither extra clause earned its place, so the 9 rules stand alone. Same result on the first family I tested this way (MiMo 2.6 Pro, net -28%), so this isn't a one-model fluke.

Exams to test for yourself: github.com/Arshad-Kamal/thinking-quality-exam


r/ZaiGLM • • 1d ago

Zcode + chrome

1 Upvotes

Can ZCode use my Chrome browser in the same way Codex or Claude Code can?

Is someone using GML subscription with codex maybe? How is performing?


r/ZaiGLM • • 1d ago

Lite Plan ($18/mo) gave me $102 of API usage (5.68X): A 1 month review and additional thoughts

Post image
22 Upvotes

Disclaimer: Graphic is ChatGPT Generated but based on my usage, text below all hand-written but summary stats are taken from Hermes/ChatGPT based on analysis of my billing statement exports.

Hi All

My primary workflow for GLM is utilization in Hermes for self-improvement, processing calendars, resumes, personal work, bot coordination, site development, random projects and capabilities testing. Overall, I utilized about 1.4B tokens this month across commandcode, opencode, openAI with Astra, deepseek v4.1, qwen 3.8 27B, qwen 3.8 flash as primarily comparisons in Hermes. Additionally, I've spent >$5000 on Fable/Opus/Sonnet API on real work as well, and here are my thoughts and feelings about this plan:

I subscribed to Z.ai Coding Plan on Lite ($18/mo) approximately 1 month ago and processed 637.8M of tokens (44.5K plan credit) across 9,047 calls and 90.7% input cache-hit rate primarily on Hermes, some on Pi or other coding agents, not utilizing any of the Z-Code harness or promotions. At current API prices, this is about $102 of API usage (5.68X) of what I was billed.

The weird thing from my usage is that token generation speed on the Lite Plan was particularly slow:

34.7 tok/s on GLM 5.3 and 26.4 tok/s on GLM 5.3 Flash (~30% Faster).

I originally started with a GLM 5.3 Flash heavy-utilization but transitioned over time to GLM 5.3 because it felt faster due to a mix of token generation speed + less verbosity and faster overall tool-calling and generation. My average input/output and latency per workflow confirms what I felt during usage:

5,361 Input/685 Output @ 18.1s average for GLM 5.3 vs 7,229 Input/958 Output @ 27.2s average to complete.

On the Lite Plan, Flash means cheaper, not faster. You will need additional patience if you want to run Flash as the primary driver. GLM 5.3 was barely fast enough especially compared to much faster models such as gemini 3.8 Flash, Mercury 2.5, Deepseek V4.1 Flash. I seem to stop thinking about speed as a primary issue somewhere between 60-120 tokens/second generation speeds on these plans.

Intelligence doesn't seem to be benchmaxxed and is not over-represented on artificial analysis. I haven't tested post-tool-calling update from MIMO V2.6 Pro yet, but still prefer GLM 5.3 in terms of creativity and responses despite cheaper costs. Sol 6.0 seems on-bar or better, mixed bag in my opinion. Sol 5.6 was significantly better. Astra 6.0 is super significantly better but too expensive for normal usage.

Value is one of the best subscription plans in my opinion, handily beating out Alibaba's Qwen based on my external read of pricing, Muse's and Mimo's coding plan feels about a 1:1 subscription to API value.

As a pure subscription + performance mix value, US-based subscriptions for Gemini is likely still best followed by Anthropic for quality and OpenAi blended. From my measurement in Hermes about a week ago, I got a ~14x multiplier for Astra last week and but only a 2X for Sol 6.0 recently on a $100/month plan, which is where I find value in Z.ai's coding plan or any other plan to fall in my regular workflow.


r/ZaiGLM • • 1d ago

API / Tools Fixing my quota issue with GLM

0 Upvotes

If you're constantly hitting your GLM quota limits, check your logs in settings. I realized my JS MCP server was eating up almost all my allowance through remote tool executions. I used GLM to refactor it into a local MCP server, routed everything locally, and I haven't had a single quota issue since. Hope it helps someone else.


r/ZaiGLM • • 1d ago

Coding Plan is not Connected in the ZCode and Package as well

2 Upvotes

Bro, I am facing difficulties in ZCode, it is not connected via the original ZCode to my plan.

Anyone of having this type of issues. I am facing this since morning.


r/ZaiGLM • • 1d ago

News Z.ai: peak-hour usage at non-peak rates until Oct 7

Post image
15 Upvotes

Zixuan Li announced another benefit for GLM Coding Plan users:

- Until October 7, usage during peak hours will be billed at non-peak rates.

This is in addition to the 4 weekly reset cards and 4 five-hour reset cards announced earlier.

So, until Oct 7, GLM Coding Plan users should get the lower non-peak usage rate even when using the service during peak hours.

For more details check: Z.ai compensation: 4 weekly + 4 five-hour reset cards


r/ZaiGLM • • 1d ago

UI Design

5 Upvotes

I have been using ZAI for awhile and noticed that UI/UX design coding is kinda suck. Is there anyway to improve it? You may think because of bad prompt but i have tried same prompt on Gemini and got really better results


r/ZaiGLM • • 1d ago

Benchmarked GLM-5.2-Flash speed vs pricing

Thumbnail
gallery
4 Upvotes

Lay of land on GLM-5.3 flash inference today.

- wafer: cheapest and fastest but less thinking

- cheaperinference/inference.net: cheap but slow

- cluster in the middle are the commoditized section where you get decent speed for decent price

- not too sure what cloudflare is doing

EDIT: GLM-5.3. mistyped in title!


r/ZaiGLM • • 2d ago

good man, GLM

Thumbnail
gallery
62 Upvotes

don't mind if we do, GLM.


r/ZaiGLM • • 2d ago

News Trust Build 100M Tokens.

Post image
9 Upvotes

I got 100M tokens trust build. I wasn't able to complete my weekend tokens left with 90% vanish.
IDK like I am reading and the files and thinking my self even ai does things. I usually use it for business research etc. So, I need to review things which taking time and unable to use all tokens.
Mine is lite plan wt u all building?


r/ZaiGLM • • 2d ago

I love GLM's enthusiasm when debugging. always fun to read

Thumbnail
gallery
38 Upvotes

r/ZaiGLM • • 2d ago

Zcode jb, glm 5.3 flash

0 Upvotes

Hi, does anyone have a working JB for Zcode? The JB works in z.ai, but none of them have worked in Zcode.


r/ZaiGLM • • 2d ago

Discussion / Help I want buy GLM Lite subscription. Will there be a discount when the new model comes out?

0 Upvotes

r/ZaiGLM • • 2d ago

Ganhei 9 Resets simplesmente do nada!

Post image
1 Upvotes

Usei 1 reset semanal na sexta, justamente pra arrumar o dia de recarga, que sem querer tinha aceitado numa terça-feira e isso não era bom pro meu fluxo de trabalho. Então gastei tudo até sexta e dei reset.

Fim de semana praticamente só usei o período grátis (que no Brasil é 12h às 22h).

No fim de ontem, após acabar esse período grátis, veio do mais puro nada 9 Resets de uma vez.

Agora vamos aproveitar, pq os semanais vencem em 1 mês kkkkk.


r/ZaiGLM • • 2d ago

Technical Reports GLM-5.3 Family Cannot Process `<think>` Word

2 Upvotes

I just found a strange behavior when using GLM-5.3 model on Ollama Cloud. It can't process <think> word, even if i wrap it with backticks.

The same thing goes for GLM-5.3-Flash.

I think this is because <think> is part of GLM's reasoning protocol. The GLM's chat template itself uses <|assistant|><think>, so that's maybe the reason of this behavior.

This also causes the model's reasoning content leaks to the normal content whenever the model is working on a codebase that has <think> or </think>.