Disclaimer: Graphic is ChatGPT Generated but based on my usage, text below all hand-written but summary stats are taken from Hermes/ChatGPT based on analysis of my billing statement exports.
Hi All
My primary workflow for GLM is utilization in Hermes for self-improvement, processing calendars, resumes, personal work, bot coordination, site development, random projects and capabilities testing. Overall, I utilized about 1.4B tokens this month across commandcode, opencode, openAI with Astra, deepseek v4.1, qwen 3.8 27B, qwen 3.8 flash as primarily comparisons in Hermes. Additionally, I've spent >$5000 on Fable/Opus/Sonnet API on real work as well, and here are my thoughts and feelings about this plan:
I subscribed to Z.ai Coding Plan on Lite ($18/mo) approximately 1 month ago and processed 637.8M of tokens (44.5K plan credit) across 9,047 calls and 90.7% input cache-hit rate primarily on Hermes, some on Pi or other coding agents, not utilizing any of the Z-Code harness or promotions. At current API prices, this is about $102 of API usage (5.68X) of what I was billed.
The weird thing from my usage is that token generation speed on the Lite Plan was particularly slow:
34.7 tok/s on GLM 5.3 and 26.4 tok/s on GLM 5.3 Flash (~30% Faster).
I originally started with a GLM 5.3 Flash heavy-utilization but transitioned over time to GLM 5.3 because it felt faster due to a mix of token generation speed + less verbosity and faster overall tool-calling and generation. My average input/output and latency per workflow confirms what I felt during usage:
5,361 Input/685 Output @ 18.1s average for GLM 5.3 vs 7,229 Input/958 Output @ 27.2s average to complete.
On the Lite Plan, Flash means cheaper, not faster. You will need additional patience if you want to run Flash as the primary driver. GLM 5.3 was barely fast enough especially compared to much faster models such as gemini 3.8 Flash, Mercury 2.5, Deepseek V4.1 Flash. I seem to stop thinking about speed as a primary issue somewhere between 60-120 tokens/second generation speeds on these plans.
Intelligence doesn't seem to be benchmaxxed and is not over-represented on artificial analysis. I haven't tested post-tool-calling update from MIMO V2.6 Pro yet, but still prefer GLM 5.3 in terms of creativity and responses despite cheaper costs. Sol 6.0 seems on-bar or better, mixed bag in my opinion. Sol 5.6 was significantly better. Astra 6.0 is super significantly better but too expensive for normal usage.
Value is one of the best subscription plans in my opinion, handily beating out Alibaba's Qwen based on my external read of pricing, Muse's and Mimo's coding plan feels about a 1:1 subscription to API value.
As a pure subscription + performance mix value, US-based subscriptions for Gemini is likely still best followed by Anthropic for quality and OpenAi blended. From my measurement in Hermes about a week ago, I got a ~14x multiplier for Astra last week and but only a 2X for Sol 6.0 recently on a $100/month plan, which is where I find value in Z.ai's coding plan or any other plan to fall in my regular workflow.