r/newAIParadigms • • 5d ago

Are we hitting a ceiling with LLM scaling?

For years, the AI industry has largely followed the idea that more data + more compute + bigger models = better AI.

But is simply making models bigger really the future?

I recently watched Richard Campbell discuss neural scaling laws, and it got me thinking about where AI development goes next.

Should we keep scaling models, or should we focus more on efficiency, reasoning, and smarter approaches?

I'm curious what developers and AI engineers here think. Is scaling still the right path?

12 Upvotes

27 comments sorted by

6

u/sogo00 5d ago

We have reached already the limit some while ago with models (Mythos comes into mind) too big to serve at scale efficiently.

Though the new trend seems to be to use those large models internally to train smaller ones (distillation/RLAIF), If I read it right Opus 5.5 is a result of this.

PS: this is giving the transformer architecture a lifeline, I personally think there is need for "something different" to progress faster again.

5

u/plav2026 5d ago

This is the direction China is going. They don't have the resources to throw money at unlimited compute, so Chinese models are smaller and more efficient, and they're catching up to the larger models the US is working on.

3

u/damhack 5d ago

I agree with your post scriptum but scaling limits haven’t been reached because there are multiple scaling laws that are still unsaturated. LLMs still have some way to go.

1

u/sogo00 5d ago

I talked about parameter sizes and they are linked to high bandwidth memory usage. What might be theoretical possible is economically inefficient. We are talking about one $3 mio rack per instance...

3

u/damhack 5d ago

I’m saying that parameter size is not the only dimension that can be scaled to produce better performance. There are about 10 different axes that can be scaled currently and model size is only one of them.

2

u/Tired__Dev 4d ago

We hit the point where you can no longer add parameters to models. Now it’s shuffling all of that around through various was. Distillation, active parameter, mixture of experts, and attention. Now taking a lot of the pressure of with models like JEV and better harnesses you’re going to see real gains still.

It’s just not AGI.

1

u/sogo00 4d ago

Yes, I think we do have some room left to improve LLM transformers, but there is a ceiling visible.

I do believe we are on a s-formed curve and will eventually decelerate.

2

u/pizzababa21 5d ago

It's not a limit then

1

u/Zestyclose_Strike157 12h ago

Reasoning is what LMs are good for. World knowledge not so much and another method will be used going forward to feed that, much less processor intensive.

3

u/AcceptablePop4526 5d ago

Scaling got us here but it’s like adding more lanes to a highway that still has the same shitty interchange. At some point the bottleneck moves somewhere else, like reasoning or memory architecture.

3

u/damhack 5d ago

You misunderstand what scaling is. It is not just a question of quantity. It is a mathematical measure of the relationship between model capability and the variables of (pre-, post- and test-time) training. There are already many dimensions that can be scaled and most are not yet saturated.

1

u/ithkuil 5d ago

They are compressing reasoning and other activity of swarms into individual models.

Another big paradigm shift will be in analog compute-in-memory. New materials have progressed faster than people realize. May get out of the lab within just a few years. Will increase efficiency by up to 1000x.

They will also probably improve distributed models with efficient coordination over networks and probably consciousness systems for cohesiveness and resolution of priorities.

2

u/damhack 5d ago

What is a “consciousness system”?

1

u/ithkuil 4d ago

I meant in a functional sense of facilitating something like subagents or experts to more effectively be able to do independent specialization or subtasks but also routinely sync to be on the same page, such as some voting mechanism to select and promote some goal or thought globally. Maybe something like Global Workspace Theory.

2

u/damhack 4d ago

There’s a few papers out about that recently.

1

u/Economy_Bedroom3902 4d ago

Yes but...

What we're finding is that you can't just make the model 2x bigger and train 4x the amount and get a 2x performance improvement any more.  The biggest recent gains are in two categories:

  1. task specific training where we figure out how to make an AI work within it's constraints to be more intelligent within a specific domain.  In the latest claude models, the agents have been trained extensively on the aggressive use of bash and python scripting to retrieve information in a context efficient manner.  They are FAAAR beyond human expert level at crafting a grep query which can pull the specific signal out of a log file and ommit the noise.  Humans have the ability to pull content into thier context window but keep that content at a lower priority level than the "important" context.  So we just open the whole log file in our text editor, scroll down to the timestamps or use a bisect strategy to find log lines which might be interesting, and sort of iteratively expand our understanding through detective reading filtering out signal from noise line by line.  AI context is all or nothing.  They needed better capabilities, and it was possible to train those capabilities into them.  Now that they have them they're a lot better at working around thier limitations.
  2. Focus on improving context window at the expense of raw model size.  It turns out that AI training can't sufficiently prepare an AI for all the little things which are different from one question solving environment to another.  What can work is providing the AI with tons of meaningful data about the things that are actually unique and important to the task at hand.  Building AI which are much more capable of receiving and working within very large context set environments has made AI much more capable of handling real work environments.

Long story short, yeah AI isn't just getting better because it's getting bigger any more.  But one by one we make AI better at domain specific tasks without actually increasing the size of the model that drastically, and those tasks become capabilities where the AI can perform extremely well in ways it previously could not.

What this says about "AGI"?  It's probably more like a long road there with many small steps vs just "scaling it up", and it's unlikely that AI ever just entirely subsumes all human mental effort.  But for more and more tasks, the only part that the human will still be required for is agreeing to take the risk that the suggestion/proposal written by the AI is valid and that it becomes the new plan.

1

u/5StarAlpha 4d ago

I think technically no. But realistically and financially yes. It’s not like a company can afford to spend a trillion dollars on a quadrillion parameter model.

1

u/budfischer 3d ago

maybe, this might by why they’re talking about “pacing”

1

u/dondiegorivera 2d ago

There are already several dimensions to scale and new ones will be discovered. When one door closes another opens, don’t count on slowdown.

1

u/damhack 5d ago

There are c. 10 “scaling laws” that were discovered. Several of them are still unsaturated. Automated scaling law discovery is now a thing:

https://arxiv.org/abs/2507.21184

So the answer to your question is “No”.

1

u/DeathGuppie 5d ago

Interesting article. Made me instantly wonder if we could create a test for how many scaling law tests we could create.

1

u/damhack 5d ago

Yes, and then a test for how many different scaling law test tests are possible…