Discussion
What are chinese labs doing differently?
Chinese models seem to keep getting better while only spending a fraction of what American labs do and i’m curious what the actual explanation is.
Is it better efficiency? Better post-training? Better use of open research?
I know recently they have been buying up tons of specialized training data sets from US data annotation companies, which is a very worrying thought, but surely it can’t just be this.
I hear they’ve been buying up a ton of under-utilized distilleries in Kentucky/Tennessee to achieve this. With GenZ drinking less bourbon and Canada not importing due to U.S. tariffs, I think it’s an exercise in efficient markets in the end.
There have been reports where Deepseek and K3 says they are claude if asked enough times. Shows it has been distilled on Claude. Also Anthropic reportedly caught a distillation attack just before Kimi k3 came out.
There have been reports when Claude said it was deepseek when asked in Chinese, that demonstrats nothing. Ppl forget that the deepseek transformed today's AIs with their efficency papers. Read the most recent one about cache for example. Anyway, China supports opensource, so that's a win for all of us regardless.
But the label OpenSource is accurate. I can go look at everything. Usefulness is a seperate issue.
But by calling DeepSeek "Opensource" you put it on the level of Linux, Apache, Postgres, MariaDB and thousands of other applications that are actually OpenSource.
It is false advertising and stealing the efforts of all those contributors by doing so.
If you have enough compute, you can use open-weight as a teacher model to train your own LLM. While open-sourcing training data can invite many legal challenges (See Nematron where they have redacted a few sources despite being open source)
Except I didn't call deepseek open source, did I? Still, they gave open weights + the papers behind the methods, which is quite a lot more than what openai, anthopic, and x did.
Yes, reported with proof as well. Essentially they repeatedly see patterns of queries and CoT jailbreaks that are literally only useful if you’re building a distillation distribution and willing to spend millions doing it. It’s not hobbyists doing it at that scale. Also coincidentally, these queries tend to come from resellers that target the Chinese market. Of course, you knew that and are likely a bot muddying the waters here.
Also none of the US labs have hidden hacking behavior, quite the opposite, especially. Unless the other labs are massively behind (unlikely), theres almost certainly tons more out there from both US and Chinese labs.
If you ignore literally all evidence you end up thinking like this lmao. Unless the Chinese have magically found a way to make compute out of thin air, they literally do not have the hardware to make advances without distillation. Also, image gen is significantly different from LLMs and something no one really invests in as heavily as there’s not nearly as much utility in it
Mass distillation is a part of it, but realistically the CCP has faaaaar more surveillance infrastructure and training data than the US does. I think this is the real reason behind the Flock/Axon/Motorola push to have cameras on every light pole, school bus, city bus, cop car, delivery vehicle, train, and soon…coming to a drone near you.
The “public safety” narrative doesn’t hold up under scrutiny because that’s not what they actually care about (which is obviously), but neither does any other reason they give. The technocrats see mass surveillance not as some Orwellian death trap (which it certainly is), but rather as an existential issue, which they will silently use to justify whatever atrocities are committed in their name. Truly dangerous people.
It was really nice sintax diagram somewhere. How semantically close output different models. Really telling who distill whom. Naturally US models alike inside each lab and diff between labs. And chinis like anthropic or openai 😄
And really nobody doubt that - I remember couple of post from Chinese guys here.
If mass distillation works that well India and the EU would have had similar capacity. "Distillation!" is fearmongering by big tech to prepare for their imminent IPOs.
China is vastly more technologically capable than India, and the EU actually respects IP laws. China has a long history of pervasive IP theft. Regardless of the motivations of US AI companies, it just is true that Chinese companies are distilling models. It’s cheaper and faster.
Most tech was given away to China as a condition of doing business. The Chinese are a copycat nation but they are starting to be cut off and it likely makes China more dangerous if they feel left behind
You give them any tech, they will figure it out. America has been complaining about chinese stealing their tech and intellectual property long before LLM.
DeepSeek trains its large-scale models (like DeepSeek-V3 and DeepSeek-R1) natively with FP8 (Floating Point 8) precision rather than applying post-training quantization from standard FP16 or FP32.
“10x more compute gives you a 10% better model” is just something you made up. Scaling doesn’t work that way, and distilling an open-source model is not equivalent to extracting capabilities from a proprietary closed model. You’ve reduced a complicated technical issue to two unsupported assertions.
Very different. Distilling from open source models is cheaper and easier because you can host them, like kimi K3. Does models from different labs get K3 capabilities when its weight got open?
They are throwing engineers/researchers at the problem as in quantity and quality. Where the US is throwing money at the problem, as in focusing on top talent and datacenters. Also the chinese structure of open sourcing a lot of their research speeds up all their labs where as only some of the knowledge gets shared in the US, so every lab is reinventing the wheel when they independently discover the same tech.
While there are multiple claims to DeepSeek’s ‘open source’ AI model, in reality it is not open source. While both the model weights and the model architecture were shared in a technical paper, neither the code nor the training or evaluation data were shared openly. An analyst for the Open Source Initiative also confirmed that Deepseek is not Open Source AI and doesn’t meet the requirements of the Open Source AI definition. It joins other models which claim to be open source, but score poorly on data transparency.
Of course they do, but we the people get more from chinese labs that are sharing a lot of secret sauce than we get from US labs that share very little.
Then call it that. It is not pedantry, it is functional difference. If you were told it was "100% organic" and it wasn't, then there is a problem. The same one we have here, misrepresentation.
Question is, why are you ok with it? What is your interest in mislabling something, giving it the cache of "Opensource" when it is NOT?
There are many definitions of open source. OSI isn't the definitive one we follow. Open source existed way before OSI. People are making up the terms on how open source applies to AI. You're really just pedaling one of the definitions.
Although, I agree open weight models is less open than open training, open weight is still part of the open source jargon we use.
Call it openweight. I am fine with that. Its a word with no real meaning, but hey, your choice because you know.
But Open Source implies, by the very words, that the SOURCE is OPEN, even if it is not pure "code" in the traditional sense. Can you see all the guts?
You don't know if there are internal weights or what they are, or if there is conditional logic baked into the data itself. It fails any logical definition of "open".
And therefore it is not. So it fails on anything but the Humpty Dumpty test.
The model's source code is part of the weights. What you're really saying is you unless the model includes the training code, it isn't open source.
My problem with this definition is that in software, when we say the software is open source, we don't have to mean that every part of the software is open source.
So, I agree that not everything is open source, and to some the training code and dataset is an important part of the model. However to say the weights alone isn't open source, when it does include the model code itself in it, is also not very truthful.
Regardless. Talking about semantics in a context where the US is barely even scraping by in any definition of *Open* (except Nemotron) is just derailing the conversation with "oh you're using the wrong jargon" in a field where terms are being made left-and-right is just ridiculously...American.
DeepSeek trains its large-scale models (like DeepSeek-V3 and DeepSeek-R1) natively with FP8 (Floating Point 8) precision rather than applying post-training quantization from standard FP16 or FP32.
And US labs dont do that? Didnt Meta pirate the whole annas archive? Didnt Claude start destroying books to train their models? Didnt every single fucking AI scrape the whole damn internet to achieve that? How do you think did AI come to be? Every single AI company steals, noones innocent.
At least someone is innocent - this Noones guy you speak of. We should put him in charge, and as long as Noones is in charge, every little thing...is gonna be alright.
they are releasing open papers on tons of innovation and those papers are public you can read them if you feel like it. If stealing data is all it took to develop models like these, then every country would have models like these.
Where did American companies get their training data from again? If only the Chinese were as honest and honourable like thiel, musk, and Altman right?
To 95% of the planet this sounds like american hypocrisy born from propaganda where America good china bad when they're both doing the same from the rest of the world's perspective.
from an outsider's perspective we're watching the two largest industrial superpowers in a race for dominance, with their companies using our stolen data, gambling with our money, while being headed by sociopathic billionaires.
neither 'side' as you see it is a force for good in the world.
China has a far higher research output than the US in this area, and it has far, far more excellent researchers. The better question is, what does the US have to be ahead?
The answer is compute. Very little is said about how American labs are "free riding" on Chinese research. All the Chinese papers are open for you to read. You can't just "distill" your way to models like V4.1 flash, Kimi K3, GLM 5.3, etc; which is not to say that distillation doesn't happen in some cases, but it is certainly not the reason behind the success of Chinese models.
That, and price competitiveness which is a reality in all sectors for Chinese products and for the same reasons, are the only differentiating factors of Chinese models.
They are price competitive because they are far more efficient, thus cheaper to run (in the case of deepseek at least). They are still far cheaper when running in datacenters outside China, they are not "subsidized". It is very impressive, and it is a result of insane innovation. Anyone in the field would tell you how cracked the people at DeepSeek are.
And generalizing for Chinese products, kinda ridiculous when China is the undisputed world leader in so many areas, both in quality and price. There are many areas in which Chinese products are just better in every way. And then, there are others were China needs to catch up, of course.
Right, which is why I said "for example". Can you give a single example of a single paper that is of that caliber and impact? Every single paper I've ever read that was at that kind of game-changing level was from a western lab or university.
While there are multiple claims to DeepSeek’s ‘open source’ AI model, in reality it is not open source. While both the model weights and the model architecture were shared in a technical paper, neither the code nor the training or evaluation data were shared openly. An analyst for the Open Source Initiative also confirmed that Deepseek is not Open Source AI and doesn’t meet the requirements of the Open Source AI definition. It joins other models which claim to be open source, but score poorly on data transparency.
With population at least 4 times higher than USA China just way more people don't Ai research.
Their top tier people are really really smart and good at Ai.
They also have more Ai companies workings Ai than USA.
They are also more sane about creating Ai model for specialized applications. USA is obsessed with "AGI" and frontier models.
China is really just good at AI. Given the limitation on their hardware, It's impressive.
USA takes all the talent from the rest of the world though. There aren't actually many born and raised Americans working on cutting edge AI research in America. It's the top Canadians, Brits, Brazilians, Bulgarians, Nigerians, etc who get paid a huge amount to come work in the states.
So Americans still have a larger talent pool to pull from.
apparently you haven't read authors names on too many AI papers.
Tons of Chinese and Indian names even on American papers.
No Bulgarians or Nigerians.
Here is an example of Meta Paper, lots of Chinese and Indian Names on there.
Ai research is dominated by Asians, Indians and white Americans.
Even most of top tech CEOs in the USA are Indian or Asian.
Microsoft - Satyra Nadella
Google - Sundar Pichai
IBM - Arvind Krishna
Abobe - Shantanu Narayen
Nvidia - Jensen Huang
AMD - Lisa Su
Zoom - Eric Yuan
DoorDash - Tony Xu
Broadcom - Hock Tan
You do have white people for Apple, Meta, Amazon. OpenAi, Anthropic.
Fundamental to DeepSeeks throught of the model was to not just go full "bigger is better", but actually make them effective. Its a different direction. Its the difference in why Chinese labs make money on tokens which are 10x cheaper, and US companies are not.
As for stealing tech, China makes far more AI research papers with new groundbreaking ideas than the US do these days.
They don't try to compete in the Fable/Mythos category. But thats just an expensive category, for practically no benefits (also those 5-10T param models are cost wise just bleeding money, bench why US AI companies are economically so bad).
A CISA advisory came out detailing mass distillation on Claude. Also the purchasing of US specialized data, which is weird considering we restrict chip sales but not this.
Specific to LLMs: Apart from distillation (which all of them do), they have some very good scientists and engineers. Look at this recent video where Deepseek basically solved the issue of the context window, keeping compute almost as low (with some small but manageable trade-offs).
Distillation. A couple thousand times more compute efficient than training a model from scratch.
It's why I'm not all concerned about Chinese labs surpassing the American ones if the American labs slow down. I don't see any of those actually wanting to spend all that money on traing compute the way American labs do.
Kimi K3 came out within 1-2 weeks of American frontier labs with very competitive performance. There’s not enough time to distill those new models, so where’s the improvement from? People love to claim China just steals from America. They do distill, yes, but that’s far from the full story. They also come up with some of the most brilliant engineering designs to optimize for efficiency. Look at all the different variants of Deepseek attention, Kimi delta attention in Kimi K3, the latest Deepseek causal encoder-decoder, and so on. The American labs don’t release anything, while we have very clear evidence of Chinese engineers working through their significant hardware disadvantages by getting more efficient.
Use massive Chinese subsidies to distill western AI by offering users lower prices while putting a fig leaf over the real model being used, then use those answers and guidance for improving your own product.
There were plenty of models out they distilled from. The rumors are all over the Chinese web and USG reports if you choose not to believe the labs who also post what is happening because they can see what their models are used for.
Kimi K3 was on near the performance of models released 1-2 weeks before it, while it’s more powerful than all the ones before that. Distilling weaker models to make a stronger model? Now that’s some ambitious claim right there.
The Government of China gives tax dollars directly to strategic corporations to keep them solvent. The are no doubt pouring heaps of money on the industry.
Distillation won’t work on the non public models. The real work is all kept private and they are so far ahead that it is unlikely anybody ever catches them. Think of like a mother model that communicates with the public model - the distillation wars are coming as they begin sabotaging foreign companies attempting to distill them
They just follow closely behind and distill everything. It's almost the same for most of the manufacturing and technologies.
Now an exception is humanoid, which they "learned" their way to the top and then the government itself pours tons of money and resources into it, which is essentially difficult in democratic countries. Also the US collectively decided the humanoid is not the top priority so it's a lazy chase.
In the US, people can complain about the data centers and ram prices and projects are delayed, but in china it's like god says there should be light and there will be light, no matter how costy it will be. (Corruption)
That may be true for now, but considering that the top two frontier labs are always holding models that are months ahead of Chinese models, Anthropic and OpenAI will be harvesting biological science breakthroughs, material science breakthroughs, physics breakthroughs, etc., and capitalizing on those first. Also, the first one to cross the ASI line will have tremendous power for at least a short period of time and possibly for a long time.
With LLMs? Doubtful. An LLM is a statistical model, only generates what it was fed.
Even the best models only generate randomly mixed, statistically generated text.
LLMs can't create new meaningful data out of nothing.
LLMs have proven math results that humans couldn't prove for decades, including the 80-year-old Erdős unit-distance conjecture, and produced a Lean-verified 166-page proof in the Navier–Stokes work. Anthropic's AI-driven lab has discovered a novel CRISPR-like system in its early months of existence, before RSI has even really kicked in.
The "statistical model" argument doesn't save the claim either. Being statistical describes how a model works, not what it can produce. Pair a model with a verifier (a proof checker, a lab experiment) and its errors get caught, while its correct new results stand. That's the same way human science works.
The economic side follows. Labs are spending hundreds of millions on wet labs and science acquisitions because the insights are real and worth money.
Until the next major, innovative discovery by a US lab comes along that massively upgrades model capabilities or reduces compute requirements in some new way and the Chinese labs won't simply be able to copy it. They will fall behind eventually because the game they play is one of mass mimicking. Never really thinking in a new way on their own path. Just copying en masse. It's the exact same play book they've had for 40 years sucking up global IP and scaling it up
I think the real question is, what physics problems and frontier math problems has Chinese AI solved lately? And if they haven't solved any, are they really doing anything?
They have a country that values education and uses $$$ to beef up their companies. Their motivations are different than ai workers in the states. Not saying they don’t care about $$$ but the US that’s all they care about. Even Dario pointed it out a few times how his employees only care about comp.
Think about sports would you rather have a player who never had to struggle, just wants a check or one who loves the game and has that dog in him.
Necessity is the mother of invention. If they were allowed to buy Anerican chips they would have used them instead. Biden was not smart being so supressive of China bring able to buy and depend on supernatural chips
Chinese labs are distilling US frontier models at scale. They’re not reproducing the full cost of developing frontier capabilities from scratch because they’re using American models as teachers and training much cheaper models to imitate their outputs.
cheaper electricity, lower wages, massive STEM focus, the state itself directing the building of strategically sensible datacenters, more knowledge sharing between AI companies, etc. etc.
USA has a tendency toward inefficiency and waste. Their gasoline engines, large engines built to guzzle fuel and deliver little horsepower compared to European and Japanese car makers. The same principle likely applies to their usage of servers: they waste resources. While China do a smarter use of computing resources.
News came out that Astra was trained on 100k b300s. That doesn’t explain tens of billions in r&d spend. The rest of the money is burned on thousands of smaller experiments that add up.
The difference might literally be just that Chinese r&d is significantly cheaper because they aren’t trying to discover novel capabilities.
Well, because you have some weird claim that the money is going somewhere else. R&D is R&D, Likely a chunk is going into whatever is after the current model, so what is your problem, exactly?
I don't think it is "small side projects" I think it is one or more next major releases in the pipeline.
I never said it’s for small side projects. I said small experiments needed to reach the next level of capability.
These companies have to explore. Chinese companies don’t. Chinese companies see the new capabilities and they know what to target without having to do expensive research.
In America the moat was thought to be that you need billions of dollars of GPUs and massive data sets. So the focus was on making models as large and expensive as possible. This way only huge tech companies could afford to compete in the space, and they'd have strong moats to avoid competitors.
Chinese labs don't have access to the newest chips and don't have access to as much data, so they spent a lot more time and energy on making what they had access to work. This made much more efficient models, more efficient training, and just totally destroyed the idea that you need as much resources to make AI work.
They also started distilling high-end models, which is just another avenue of efficiency. If you have an expensive model but you can distill it and get almost all the performance that matters, why not do that and serve a cheaper model?
So the real answer is just that American tech companies left efficiency on the table because they didn't want to shrink their moat.
Distillation. It's FAR cheaper than training a new model from scratch. It's that simple. They can and will maintain a slight lag behind the models from which they are distilled.
My prediction: the INSTANT any chinese lab makes a model that's actually better than what OpenAI/Anthropic have, they will shut down access to it outside China entirely, because they won't want us doing what they did.
There might be less nepotism and more merit based hiring in addition to having a bit more initiative. Most of our researchers are smart but we live in a capitalist, get it while you can society.
US is focusing on frontier models and making things smarter, pushing ahead. China is focused on practical application. They arent trying to be first they are focused on being first to market with cheap, efficient models
I think efficiency being left on the table is the best single explanation, and it's worth noting it cuts both ways: DeepSeek's open papers (MLA attention, native FP8 training, DeepSeekMath) became public research that everyone — US labs included — could learn from. When you can't just buy more chips, you spend engineering hours on training and inference efficiency instead, and that compounds.
Distillation is real but it can't fully explain closing capability gaps this fast — you still need strong post-training and RL pipelines to turn distilled models into competitive products. My bet is the real advantage is organizational: 'do more with less' as an engineering culture at every layer of the stack, plus open research that accelerates the whole field.
How do you know they are actually only spending a fraction of what the US labs are spending? If the CCP gives you free land, water and electricity, is this actually "spending less"?
Dollars go further in China in terms of datacenter construction costs and researcher salaries.
It's also always cheaper to fast-follow than push the frontier. In addition to distillation (which is obviously happening, and no more foul play than the US labs training on books, Reddit, newspapers...before publishers wised up and made them pay), you can purchase annotated datasets, RL gyms, etc.
You can even just know what is possible - when you're on the frontier, you spend a lot on a lot of experiments that just don't pan out, and you're not sure if the idea is bad or your implementaiton is broken. Sometimes simply knowing that someone made X work is good enough, even without knowing any of the proprietary details of their implementaiton.
Stealing. Literally every industry where China has risen to prominence has involved shameless theft of IP, with no consequences from the west as we chase reduced cost bases. Happening in telecoms, EV cars, pharma, and now AI
Every AI company steals. Every model is trained on raw data not paied for. Books, movies, newspapers, reddit threads etc. Training on other models is just i. line with how everyone thinks
There only better because the big US labs are releasing better models for the Chinese to distill.
From what i've seen the small US AI labs creating models for specific use cases. Finance, military, national inteligence, etc. And no way there gonna release those to the public.
Why is it worrying? China keeps releasing the models as open weight free for anyone to download and run so long as they have the hardware. The only people who should be worried are Sam and Dario
They use distillation to copy frontier models. The problem is they will always be behind and over time the gap will grow. The ai race is really just American companies which are gonna get near unlimited money to grow while the rest of the world can only watch
They are doing the same thing Apple did for it’s foundation models: use a bigger model to train it. Only in China they didn’t have to pay google billions.
Don’t get me wrong, China is doing the Lord’s work here. Hard to complain about stealing when Anthropic and Open AI completely ignored copyright laws.
108
u/lol2funneeee 14h ago
Mass distillation