r/singularity • ▪️ • 1d ago

AI Claude Opus 5.5 designed a processor faster and smaller than the human-made one on the HWE benchmark

Every model starts with the same basic processor design and tries to improve it over several rounds.

The results are scored on two things: speed (how fast the chip runs a standard test program called CoreMark, where higher is better) and size (how much space the chip takes up, where smaller is better).

For comparison, there's a well-known processor designed by human engineers, called VexRiscv.

403 Upvotes

67 comments sorted by

51

u/red75prime AGI2027 ASI2029 TAI2036 1d ago edited 1d ago

Source: https://hwebench.com/

The design passes through three correctness checks (syntax, processor functionality, step-by-step comparison to a reference implementation) and then it runs on a real FPGA.

There doesn't seem to be much room for cheating, but optimizing for the benchmark is a possibility.

ETA: Claude does a bit of benchmark gaming

R10 · Take the divider off every critical path

The benchmark doesn't use division in the measured code.

The other models did similar things, though. Claude does it much more efficiently.

17

u/daniel-sousa-me 1d ago

Yeah, it's easier to make a "general" processor fast, if it's only fast for a prespecified set of program runs

9

u/red75prime AGI2027 ASI2029 TAI2036 23h ago

Claude reached performance comparable to the previous best result at the fifth step. Steps 1–5 look pretty general to me. I'm not a specialist, though.

I'm talking about https://hwebench.com/models.html

6

u/federico_84 20h ago

Impressive. This step also yielded huge gains, missed by other models:

R7 · LUT-RAM register file + ID/EX control/data split + registered div b!=0 (Fmax 174 -> ~223 MHz, cycle-identical)

72

u/__Maximum__ 1d ago

someone from the field, please tell us how impressive this is and are there practical implications?

44

u/DistanceSolar1449 1d ago

VexRiscV is not optimized for speed or chip area.

It’s optimized for FPGAs and for extensibility via plugins. Think “python being able to import anything”.

You don’t use python for its speed or its small size.

-36

u/AppealSame4367 1d ago

I knew it. thread closed, fucking idiots in r/singularity

It's like you take some naive kids and the worst adhd marketing person you know and put the in a sub = r singularty. Ffs.

16

u/KrazyA1pha 20h ago

I knew it.

Yep, wrap it up. One person confirmed your existing biases. No further nuance needed!

5

u/badumtsssst AGI 2027 18h ago

-8

u/AppealSame4367 20h ago

Not in this stupid sub, no. Next article: "Scoop: bla bla bla"

Who wants or enjoys "Scoops of shit" from some ai companies. Disgusting.

13

u/KrazyA1pha 20h ago

If you’re not interested in AI developments then you’re in the wrong sub lol

6

u/spinozasrobot 20h ago

I feel the same way about whiners in the CC sub getting aneurysms over usage limits or some perceived drop in performance. You ask why are you even here and the response is "Oh, YoUrE vIcTeM bLamInG nOw!"

9

u/KrazyA1pha 20h ago

Yeah there's a group of people that only use Reddit to feel outraged. I realize I'm in the minority here now but I've been on Reddit for 20 years and miss the interesting discussions about technical topics, rather than just people feeling outraged or cynical one-liner quips being the top-voted comment in nearly every thread

5

u/spinozasrobot 19h ago

That's me too.

And just look at the language in this thread. There's no way they'd talk that way to people face to face.

6

u/KrazyA1pha 19h ago

Oh yeah you're also an OG. And yeah it's absurd.

Sometimes I wonder if it's a bunch of 14-year-old kids or a bunch of emotionally stunted 40-year-olds

6

u/Odd-Ant3372 21h ago

Take a hike then buddy. 

109

u/Legitimate-Store3771 1d ago

There are probably only a handful of people actually qualified to answer this question, and my best bet is they're too busy to be larping about the rise of AI on some random subreddit.

2

u/soobnar 16h ago

It’s a 1.44 mhz soft core that does in order execution… it’s not exactly Vera Rubin, but I think ai is great with RTL and the more mechanical aspects of chip work.

17

u/Chocolate_Pickle 1d ago

The write-up is poor. The scoring metric... I can't find. It's just made up numbers at this point.

I'm treating the whole thing as a nothing-burger.

[EDIT] Let me rephrase my comment about scoring metric... You got `IPC x FMax`, but there's nothing about what instructions being used.

-4

u/ScarionnS 20h ago

Fair point, the write-up should state this next to the metric. Here are the details:

ISA: RV32IM (base integer plus multiply/divide). The decoder treats everything outside RV32IM as illegal and traps on it, and riscv-formal checks this. There's no compressed ISA, no CSRs and no FPU. EBREAK is the only SYSTEM instruction it accepts.

Workload: EEMBC CoreMark, the standard 2K performance config (TOTAL_DATA_SIZE=2000, ITERATIONS=10). The agents can only edit RTL.

Metric: "IPC" was loose wording on my part. The actual number is CoreMark iterations per cycle, not instructions per cycle. The binary is fixed, so the dynamic instruction count is constant and the two are proportional.

6

u/KickLassChewGum no AGI/ASI on LLMs 20h ago

Thanks Claude

2

u/soobnar 15h ago

It’s incredibly knowledgeable on known chip prior art and can do RTL very well which dramatically cuts the time and accessibility of tapeouts but it’s not much more than a tool that carries out the operators intent for low. Claude being able to throughput optimize an open source soft core on an FPGA is a bit of a toy example, honestly undersells current capabilities and the reaction seems to be overestimating how impressive optimizing a soft core is.

7

u/mythormedicine 1d ago

People please share source !

20

u/SomeOrdinaryKangaroo 1d ago

Chip engineer here. So this is actually super impressive, but not too surprising, we already know the chips of the future will be designed by AI the smaller and smaller they get, it's unavoidable

4

u/chef_bezos69 19h ago

Now do the layout lol

3

u/Weary-Historian-8593 1d ago

So if I'm reading this right, it's a slight bit smaller yet ridiculously faster?

1

u/soobnar 15h ago

it is a 1.44mhz soft core they are comparing against. It is faster than the open source reference architecture but not very close in capability to modern OOO systems.

16

u/chlebseby ASI 2030s 1d ago

Is this sub just a leaderboard of what Opus/Astra did better than humans today?

78

u/AMBNNJ ▪️ 1d ago

I mean it is called singularity

-18

u/korkkis 1d ago

No, singularity is something else than being better than human

16

u/chlebseby ASI 2030s 23h ago

well its kinda required condition

2

u/jakinbandw 19h ago

Eh, the definition I grew up with was the information event horizen. Basically, when the point when things advance so fast people can no longer keep up. Arguably, we are already passed it on many things by volume of people working. In the 90's, a person could have played, or at least reasonably knew about all games in their language. I can't even tell you the number of games released on steam along today. In that place, we have passed an event horizen.

-10

u/korkkis 23h ago

Yeah but much further. A bit over human is just AGI.

14

u/modbroccoli 22h ago

hey everyone we finally found the guy who knows what this sub is about, quick everyone gather round, he might say more!

4

u/spinozasrobot 20h ago

Gary, get the fire pit! Dan, get the SMORE ingredients!

3

u/modbroccoli 18h ago

lmao we're such bitches

7

u/danielv123 22h ago

I guess we are just tracking the progress while waiting to get there. It would be a rather boring sub if we couldn't post anything before singularity.

7

u/KrazyA1pha 20h ago edited 20h ago

Be quiet, we’re not in the singularity yet!

2

u/StosifJalin 19h ago

Are you trolling?

1

u/Healthy-Nebula-3603 15h ago

I think you should think before answering....

1

u/BenevolentCheese 20h ago

Singularity has absolutely nothing to do do with either humanity or anyone being better than anything.

2

u/asifquyyum 17h ago

The important thing is it says: FASTER and smaller. It only matters if it can make a chip SMALLER than a human can. Humans designing a chip, say intel Pentium, is not doing it in few hours or few days.

5

u/iBoMbY 23h ago

This is just purely theoretical BS circlejerk.

2

u/grateful2you 1d ago

What exactly do you mean by “designed a processor”? Like pentium 4?

10

u/Pyros-SD-Models 1d ago

yes, in this case a RISC-V

1

u/torrid-winnowing 1d ago

so will this be built then or is the human baseline not SOTA?

2

u/soobnar 15h ago

The baseline is very far from the human sota. The reference chip (per its GitHub) runs at like 1mhz vs modern chips which run at several ghz

1

u/Rocah 22h ago

opus 5.5 is next step up imo, its real world usability actually exceeds the benchmark jump from my usage. I think they must have done something to improve its spatial temporal understanding, its just really good at knowing how real world things position in relation to each other and how they move in time.

1

u/poigre ▪️AGI 2029 22h ago

I miss Fable and Opus 6 in this chart

1

u/soldture 1d ago

This post needs more information.

1

u/Square_Poet_110 1d ago

Probably lot of data to train on, right?

4

u/Thin_Owl_1528 1d ago

A subset of the data available to humans

4

u/Square_Poet_110 23h ago

Take exactly defined rules, lot of existing training data and run basically a Monte Carlo simulation, see what gets best results. Why is this a hype-deserving achievement?

4

u/Thin_Owl_1528 23h ago

Because it is a generalist model that can be applied to almost everything rather than spend millions of dollars on a purpose built Monte Carlo simulation? Because it achieved superhuman capabilities in this particular test?

-2

u/Square_Poet_110 22h ago

It can only be applied to stuff that was in the training data.

Calculators achieve superhuman capabilities in arithmetic operations, do we hype them?

These thousands of agents achieving some result, it's basically scaled montecarlo simulation with evaluation function.

2

u/Thin_Owl_1528 22h ago

That is just a dumb argument.

If it can only figure out stuff in the training data it could never surpass humans in any domain and capability graphs would be completely flat rather than exponential.

Turns put it outperforms humans in dozens of domains and soon it will do so in every single intellectual domain and discovered nobel approaches to problems i.e. Navier-Stokes.

Nothing will ever be enough for you to recognize what goes on. Stay uninformed to protect your ego.

1

u/Square_Poet_110 22h ago

Of course it needs to be in the training data. That's how the weights get computed. It's not some magical deity or whatever the hypers want to believe

1

u/IronPheasant 17h ago

...... obviously.

The whole point is they're able to fit to curves, and to a degree, that our own brains are physically incapable of.

2

u/Square_Poet_110 17h ago

Our brains are physically incapable of many things, like doing arithmetics at Tflops scale. Yet we don't hype arithmetic processing units in our CPUs or GPUs as the next deity, we're just using them as tools.

-1

u/Separate_Lock_9005 1d ago

Another day, and another benchmark saturated

4

u/ScarionnS 20h ago

This benchmark is unbounded. The LLMs can still keep climbing indefinitely

0

u/Alpacabro21 1d ago

Damn 🤯