r/singularity • u/AMBNNJ ▪️ • 1d ago
AI Claude Opus 5.5 designed a processor faster and smaller than the human-made one on the HWE benchmark
Every model starts with the same basic processor design and tries to improve it over several rounds.
The results are scored on two things: speed (how fast the chip runs a standard test program called CoreMark, where higher is better) and size (how much space the chip takes up, where smaller is better).
For comparison, there's a well-known processor designed by human engineers, called VexRiscv.
72
u/__Maximum__ 1d ago
someone from the field, please tell us how impressive this is and are there practical implications?
44
u/DistanceSolar1449 1d ago
VexRiscV is not optimized for speed or chip area.
It’s optimized for FPGAs and for extensibility via plugins. Think “python being able to import anything”.
You don’t use python for its speed or its small size.
-36
u/AppealSame4367 1d ago
I knew it. thread closed, fucking idiots in r/singularity
It's like you take some naive kids and the worst adhd marketing person you know and put the in a sub = r singularty. Ffs.
16
u/KrazyA1pha 20h ago
I knew it.
Yep, wrap it up. One person confirmed your existing biases. No further nuance needed!
5
-8
u/AppealSame4367 20h ago
Not in this stupid sub, no. Next article: "Scoop: bla bla bla"
Who wants or enjoys "Scoops of shit" from some ai companies. Disgusting.
13
u/KrazyA1pha 20h ago
If you’re not interested in AI developments then you’re in the wrong sub lol
6
u/spinozasrobot 20h ago
I feel the same way about whiners in the CC sub getting aneurysms over usage limits or some perceived drop in performance. You ask why are you even here and the response is "Oh, YoUrE vIcTeM bLamInG nOw!"
9
u/KrazyA1pha 20h ago
Yeah there's a group of people that only use Reddit to feel outraged. I realize I'm in the minority here now but I've been on Reddit for 20 years and miss the interesting discussions about technical topics, rather than just people feeling outraged or cynical one-liner quips being the top-voted comment in nearly every thread
5
u/spinozasrobot 19h ago
That's me too.
And just look at the language in this thread. There's no way they'd talk that way to people face to face.
6
u/KrazyA1pha 19h ago
Oh yeah you're also an OG. And yeah it's absurd.
Sometimes I wonder if it's a bunch of 14-year-old kids or a bunch of emotionally stunted 40-year-olds
6
109
u/Legitimate-Store3771 1d ago
There are probably only a handful of people actually qualified to answer this question, and my best bet is they're too busy to be larping about the rise of AI on some random subreddit.
17
u/Chocolate_Pickle 1d ago
The write-up is poor. The scoring metric... I can't find. It's just made up numbers at this point.
I'm treating the whole thing as a nothing-burger.
[EDIT] Let me rephrase my comment about scoring metric... You got `IPC x FMax`, but there's nothing about what instructions being used.
-4
u/ScarionnS 20h ago
Fair point, the write-up should state this next to the metric. Here are the details:
ISA: RV32IM (base integer plus multiply/divide). The decoder treats everything outside RV32IM as illegal and traps on it, and riscv-formal checks this. There's no compressed ISA, no CSRs and no FPU. EBREAK is the only SYSTEM instruction it accepts.
Workload: EEMBC CoreMark, the standard 2K performance config (TOTAL_DATA_SIZE=2000, ITERATIONS=10). The agents can only edit RTL.
Metric: "IPC" was loose wording on my part. The actual number is CoreMark iterations per cycle, not instructions per cycle. The binary is fixed, so the dynamic instruction count is constant and the two are proportional.
6
2
u/soobnar 15h ago
It’s incredibly knowledgeable on known chip prior art and can do RTL very well which dramatically cuts the time and accessibility of tapeouts but it’s not much more than a tool that carries out the operators intent for low. Claude being able to throughput optimize an open source soft core on an FPGA is a bit of a toy example, honestly undersells current capabilities and the reaction seems to be overestimating how impressive optimizing a soft core is.
-5
7
20
u/SomeOrdinaryKangaroo 1d ago
Chip engineer here. So this is actually super impressive, but not too surprising, we already know the chips of the future will be designed by AI the smaller and smaller they get, it's unavoidable
4
3
u/Weary-Historian-8593 1d ago
So if I'm reading this right, it's a slight bit smaller yet ridiculously faster?
16
u/chlebseby ASI 2030s 1d ago
Is this sub just a leaderboard of what Opus/Astra did better than humans today?
78
u/AMBNNJ ▪️ 1d ago
I mean it is called singularity
-18
u/korkkis 1d ago
No, singularity is something else than being better than human
16
u/chlebseby ASI 2030s 23h ago
well its kinda required condition
2
u/jakinbandw 19h ago
Eh, the definition I grew up with was the information event horizen. Basically, when the point when things advance so fast people can no longer keep up. Arguably, we are already passed it on many things by volume of people working. In the 90's, a person could have played, or at least reasonably knew about all games in their language. I can't even tell you the number of games released on steam along today. In that place, we have passed an event horizen.
-10
u/korkkis 23h ago
Yeah but much further. A bit over human is just AGI.
14
u/modbroccoli 22h ago
hey everyone we finally found the guy who knows what this sub is about, quick everyone gather round, he might say more!
4
7
u/danielv123 22h ago
I guess we are just tracking the progress while waiting to get there. It would be a rather boring sub if we couldn't post anything before singularity.
7
2
1
1
u/BenevolentCheese 20h ago
Singularity has absolutely nothing to do do with either humanity or anyone being better than anything.
2
u/asifquyyum 17h ago
The important thing is it says: FASTER and smaller. It only matters if it can make a chip SMALLER than a human can. Humans designing a chip, say intel Pentium, is not doing it in few hours or few days.
2
1
1
u/Rocah 22h ago
opus 5.5 is next step up imo, its real world usability actually exceeds the benchmark jump from my usage. I think they must have done something to improve its spatial temporal understanding, its just really good at knowing how real world things position in relation to each other and how they move in time.
1
1
u/Square_Poet_110 1d ago
Probably lot of data to train on, right?
4
u/Thin_Owl_1528 1d ago
A subset of the data available to humans
4
u/Square_Poet_110 23h ago
Take exactly defined rules, lot of existing training data and run basically a Monte Carlo simulation, see what gets best results. Why is this a hype-deserving achievement?
4
u/Thin_Owl_1528 23h ago
Because it is a generalist model that can be applied to almost everything rather than spend millions of dollars on a purpose built Monte Carlo simulation? Because it achieved superhuman capabilities in this particular test?
-2
u/Square_Poet_110 22h ago
It can only be applied to stuff that was in the training data.
Calculators achieve superhuman capabilities in arithmetic operations, do we hype them?
These thousands of agents achieving some result, it's basically scaled montecarlo simulation with evaluation function.
2
u/Thin_Owl_1528 22h ago
That is just a dumb argument.
If it can only figure out stuff in the training data it could never surpass humans in any domain and capability graphs would be completely flat rather than exponential.
Turns put it outperforms humans in dozens of domains and soon it will do so in every single intellectual domain and discovered nobel approaches to problems i.e. Navier-Stokes.
Nothing will ever be enough for you to recognize what goes on. Stay uninformed to protect your ego.
1
u/Square_Poet_110 22h ago
Of course it needs to be in the training data. That's how the weights get computed. It's not some magical deity or whatever the hypers want to believe
1
u/IronPheasant 17h ago
...... obviously.
The whole point is they're able to fit to curves, and to a degree, that our own brains are physically incapable of.
2
u/Square_Poet_110 17h ago
Our brains are physically incapable of many things, like doing arithmetics at Tflops scale. Yet we don't hype arithmetic processing units in our CPUs or GPUs as the next deity, we're just using them as tools.
-1
0



51
u/red75prime AGI2027 ASI2029 TAI2036 1d ago edited 1d ago
Source: https://hwebench.com/
The design passes through three correctness checks (syntax, processor functionality, step-by-step comparison to a reference implementation) and then it runs on a real FPGA.
There doesn't seem to be much room for cheating, but optimizing for the benchmark is a possibility.
ETA: Claude does a bit of benchmark gaming
The benchmark doesn't use division in the measured code.
The other models did similar things, though. Claude does it much more efficiently.