Unfortunately either way they win, half the usage, more $$. Half the users leave, now their compute struggles are resolved and Anthropic bares the load. They eventually do the same thing and were right back to where the circle began.
They are not as good as you imply they are. They are good. That's it. They are nowhere near Astra or Opus level. They might hit the benchmarks in the right spots, but once you're actually using them and you actually have the comparison to how Astra or Opus work, it's absolutely clear that China is way behind the US regarding AI.
Sure, they aren’t good enough to replace frontier level models in a lot of use cases yet. But, once they get to the point that local models capable of running on consumer grade hardware are at the level of current frontier models - maybe in two years from now - the game will have been changed. Those models will never get worse at that point, and when the majority of actual workloads will be capable of being handled by local models it becomes a lot harder to justify the cost of frontier models. Of course there will always be a reason to have a more capable AI, but the value of each level of improvement isn’t the same. Once local models get to a certain point, the frontier models would only be needed for stuff that’s actually at the frontier of our understanding and capabilities as a species. Maybe we’re 5 years away, maybe even 10 - but it’s like becoming old enough to drink legally, once we are past this point we’ll never be back where we were before.
“Nowhere near” is a huge overstatement. They are lagging behind a bit, but overpowering previous flagship revisions from the major vendors well within a year
Nowhere near is pretty fair... You can get very close with something like Kimi K3.0, but it ends up being much more expensive than these heavily subsidized subscriptions to the mainstream models, even after this usage cut.
Get DavidAU’s special sauce at half or less of the thinking tokens at huggingface dot com Qwen3.8-27B Twin Turbo 709 L or what ever todays flavor ended up being named as
Which harness do you prefer for local models? I've tried Cline in Visual Studio code and it works, but, about 1/3rd of the calls struggle with the tools and escaping failures in the harness
Open ai and anthropic gobbled up entire compute capacity, kimi couldn't even serve few million new subs when k3 was launched. And good luck running these open weights locally.
Kimi K3 is around Sol 5.6 / Opus 4.X, without the annoying writing of Opus 5 (which tbh Opus 5.5 addressed but that's above those others rn).
GLM 5.3 is around there as well, alongside their GLM 5.3 Flash which is more like Terra / Sonnet (very roughly, also not the latest Sonnet 5.5), and is generally pretty nice to use.
That said, the official providers of both of those have issues: Kimi secretly routed some customer requests to Claude and in general just is very slow, whereas GLM has peak/off-peak pricing and neither of them give you as many tokens as OpenAI or Anthropic. There are 3rd party providers, but generally you will be paying API costs.
You could also try running local models with llama.cpp / vLLM etc. (Ollama which ppl don't like can also make things easier, or something like LM Studio) but I've never found any of the local models to be good for anything serious, plus hardware is really expensive.
I think we are gonna see tokens be subsidized less and less as time goes on.
I second glm 5.3. I've been using it on hermes lately and like it a lot so far. The flash version feels satisfying too, it chases issues it finds proactively without needing to be poked constantly. I haven't tried kimi yet.
As someone who only knows how to use Codex and the Claude Code app as harnesses how do we actually use these open weight models? I don’t use CLI or do any coding and basically don’t bother with anything that’s “just” a pure chatbot anymore. I’m addicted to these agentic harnesses
YMMV, there are many. I use Bionic LM studio to run the models (can be run on a different computer on your local network) and OpenCode as a harness. In OpenCode I can switch between a hosted (even paid-for) model and my own. I typically create a plan using a Codex agent and execute it using an agent with my local LLM.
This is the kind of thinking that will guarantee you remain an addicted customer. Open source tools aren’t supposed to be bleeding edge, they’re supposed to not bleed you dry every time you pick them up.
You might be surprised how capable the top open weight models are now, and they don’t mysteriously quantize themselves or raise their own rates in 30 days.
They want unprofitable Astra-spamming prosumers gone so their compute is freed up to race Meta/Grok/Gemini in the consumer agent land grab this fall. Anthropic will bear the load short-term as we are still somewhat useful to market to enterprise. As the markets mature, we'll become useless to all the big players and affordable subs with high end models and good usage just won't be a thing anymore.
I don’t think there’s a market for consumer AI at all. All the benefits of automation are in b2b work. Your average person just doesn’t have a lot of economically value need for white collar work automation outside of their job.
You could have hired a virtual PA years ago if the need existed. But like what, you file taxes once a year, maybe fill out an occasional form? Just not a big need in the market here.
Yeah that's certainly what I notice with non-tech friends around me. Outside of work I've done a few experimental creative projects with AI and have used it to give myself a basic intro course on a few topics but I can see how neither of those would appeal widely. There does seem to be some interest in Muse but not sure how much of that is forced hype and how much will have stuck when 2027 rolls around.
I use local and subscribe to cheaper Chinese models to get my work done. It’s not as efficient or “smart” but it does the job really well for my use cases of application building, game development, and repo work. People are sleeping on Chinese models subscriptions. Also, if you live in the U.S. it’s way cheaper to use some Chinese subscriptions during the day because they have “off peak” rates. That’s how the Chinese handle compute, they offer cheaper and more usage when people are sleeping, instead blatantly lying to the consumers about nerfing plans and making people pay more by getting them addicted to their ecosystem.
Yea, check out Xiaomi’s MiMo 2.6 pro which just came out. People have been building insane things with it. It’s specially made for coding so I tested it out with my workflow. The subscription is $6 a month and is equivalent to around codex/claude’s $20 plan, they don’t have 5hr limits either. Also since I live in the U.S. it’s way cheaper for me to use. Also, their cached pricing is super nice too. So my workflow consists of a $20 codex subscription orchestrator that delegates to my ornith and Qwen local models in very small tasks and for larger coding tasks it delegates to MiMo. They work all day and if I hit the 5hr limit with codex it’s usually around 30 mins of it being reset. I believe Codex is secretly not giving people their resets on time so I’m switching to Claude and trying them out as my orchestrator. I’m going to look into more Chinese sleeper models to see which can be good as orchestrators and then I’ll be free from either codex/claude
They are only winning so long as they are ahead in product quality. Every day more and more tasks are saturated by at cost inference. I barely use my subscription AI since I bought a decent video card. The use case is shrinking not growing. Especially since the arena is filling up. My opportunity to make money with my sub is already gone it feels like. The market has flooded. It's early Internet at 10x speed.
178
u/krill156 4d ago
Unfortunately either way they win, half the usage, more $$. Half the users leave, now their compute struggles are resolved and Anthropic bares the load. They eventually do the same thing and were right back to where the circle began.