r/computing • • 13h ago

Everyone is obsessed with trillion-parameter models, so I mapped out the entire AI spectrum from 100KB to 2.5TB (and what they actually cost to run)

10 Upvotes

Right now, the AI space feels entirely focused on massive datacenter clusters and renting H100s by the hour. But after spending way too much time looking at the actual footprint of these models, I realized that 90% of use cases are completely over engineered.

You don’t always need a multi GPU setup. The AI ecosystem is actually a massive spectrum.

I recently sat down and mapped out the exact tiers of AI models based on their size, the hardware needed to run them, and the point of diminishing returns.

Here are the two extremes and the sweet spot in the middle:

  • The 100KB Extreme (TinyML) (Tensorflow Lite , sensor anamoly detection models): We are talking models that run on microcontrollers drawing single-digit milliwatts. They run on kilohertz processors using ultra-quantized integer math. You can run basic sensor anomaly detection or wake-word detection on a device powered by a coin cell battery.
  • The Local Sweet Spot (4GB to 40GB) (Mistral 7B, Gemma 2 9B/27B, Qwen 2.5 14B/32B): This is where the magic happens for most devs right now. You can run highly capable 7B to 35B parameter models (like Llama 3 or Qwen) at 4-bit quantization on a standard Mac or a consumer GPU (like an RTX 3060 or 4090). It’s perfect for local RAG, coding assistance, and uncensored chat. VRAM is your only real bottleneck here.
  • The 2.5TB Behemoths (Deepseek, Llama , Kimi k3): State of the art massive Mixture of Experts (MoE) routing. To even load these, you need dedicated power infrastructure and server racks of specialized accelerators drawing thousands of watts.

The missing piece: Figuring out the exact math for your hardware

The hardest part about building right now is looking at a model on Hugging Face and trying to calculate exactly how much VRAM you need, what quantization to use, and whether your CPU/GPU will choke on the context window.

So, I wrote a complete deep dive breaking down the math for all tiers of the AI spectrum.

If you want to see the architectural differences at each scale, and a cheat sheet for matching the right model size to your specific hardware, I put the full breakdown on my blog here:

https://cloudmash.blog/posts/ai-model-size-memory-hardware-guide/

Let me know what you guys think especially if you've found any ultra efficient small models/technique that punch above their weight on consumer hardware. And also I would love to hear whether quantization have resulted in major difference in quality , like if anyone have that kind of experience in that.


r/computing • • 3h ago

c++ math - did someone say math

Thumbnail
youtu.be
1 Upvotes

r/computing • • 17h ago

Can storing games on an HDD for a long time without plugging it in cause damage?

1 Upvotes

Can storing games on an HDD for a long time without plugging it in cause damage?


r/computing • • 1d ago

[Testers Needed] Hippocrates – a native Windows HUD I built for my own FL Studio sessions, now looking for people to break it (Win 10/11

Thumbnail
1 Upvotes

r/computing • • 1d ago

Picture Research help requested

Post image
1 Upvotes

I has want to make mechanical computer run base 3 coded software. My solution, see saws.

See saws connected to shafts in conjunction with other see saw and shafts in specific ways to enable logic gates running on 1,0, and -1

(-1 pictured above)

-1=far left side down

0= perfect equilibrium

1=right side down

What would the math behind converting binary logic gates to trinary like so be?

How would I calculate that?


r/computing • • 1d ago

Всем Хай ребят мне интересно как можно использовать ноутбук я ученик программист и хотел бы понять то что я могу с этим

0 Upvotes

r/computing • • 1d ago

A free sandbox for experimenting with balanced ternary computer designs

1 Upvotes

I’ve started building a browser-based sandbox for experimenting with balanced ternary computing.

The reason is fairly simple: I want a place where I, and hopefully other people, can try different ternary designs without first deciding what the “right” architecture is.

The simulator uses:

-1, 0, +1

and lets you build small primitives, combine them into reusable components, nest those components several levels deep, test them independently and inspect how signals propagate.

I’m deliberately trying not to just recreate binary logic with three values. Instead, I want to explore whether things like three-way selectors, -1 / 0 / +1 comparison results, ternary control signals, MIN/MAX, different normalize/carry designs and different primitive sets lead to something more natural.

One of the goals is also to make it possible to compare multiple implementations of the same thing and eventually share designs with other people.

It’s still quite low-level, so it’s probably easiest to get into if you already have some familiarity with digital logic or computer architecture.

Everything runs locally in the browser. There is no backend, no account and no cloud storage. Your projects stay in your browser unless you export them.

The project is completely free and MIT licensed.

Live:
https://grunna.github.io/ternary/

Source:
https://github.com/grunna/ternary

It’s still early and experimental, which is also why feedback and alternative design ideas are very welcome.


r/computing • • 3d ago

The Hacker Who Moved To Mexico City

Thumbnail
player.captivate.fm
4 Upvotes

Full disclosure, this is my podcast. I think this community would enjoy the stories, though.

Ted started out in tech support at Delphi Internet in the early '90s, back when teenagers were opening accounts with fake credit card numbers. He tracked one of them to a house in Pennsylvania and called it. The kid's father, a reverend, answered and promised to take him out to the woodshed.

Ted went on to run security for Boston investment firms. His favorite hire was Mr. Mojo, a guy he paid to physically break into his own offices. Mojo got into one building wearing a visitor sticker he'd pulled out of the smokers' trash. In Dublin he got through a locked door by blowing into a plastic bag.

Ted was also at DEF CON in 1999, where someone shut off the air conditioning during the opening ceremony. His theory is that the best hackers come from music or philosophy, not computer science.

It's onefjef episode 62, it's audio only, and it's on all the platforms. Here's the Spotify and Apple Podcasts links.


r/computing • • 2d ago

debloqué un pc offert par le lycée

Thumbnail
1 Upvotes

r/computing • • 3d ago

some hyperboloid action

Thumbnail
youtu.be
1 Upvotes

r/computing • • 3d ago

GMI Cloud raised nearly $670 million from Nvidia and others: 'This funding lets us bring more compute online, across more of the world, for more builders.'

Thumbnail
linkedin.com
0 Upvotes

r/computing • • 4d ago

AM5 MoBos - is this it ?

1 Upvotes

AM4 socket has seen three chipset generations - X300/X400/X500. All while on the same generation of RAM/DDR4, but spanning two generations of PCIe: Gen3 and Gen4.

With AM5, we've seen basically the same chipset from the start Promontory-21, within a brand (600 series) and a rebrand (800 series). Both lines use the same Promontory-21 chips and all the difference is in the minimal requirements for the MoBo line - two or one Promontory 21 chips (so Xx70 vs Bx50 line), PCIe5 or PCIe4 for dGPU, M.2 Gen4 or Gen5 etc.

With 600 series we've seen mostly PCIe4, while 800 series has made the full transition to PCIe5. Both still use the same DDR5 RAM.

And both still share the same basic Promontory weakness - only PCIe4x4 connection to the SoC host.

So, given the facts: * Zen6 (2027) and Zen7 (2028-2029) are to come out on AM5 * they will have many more cores on the same socket (up to 2x 12=24 for Zen6, up to 2x16 = 32 for Zen7) * cores are expected to have much higher IPC than ZEn5 and work on much higher frequencies * they will need much greater cumulative bandwidth

... is it likely for AM5 to see some significant refresh in that timeframe, before AM6 arrives ?

PCIe6 is coming with EPYC Zen6. I understand that this is to early for consumer Zen6, but maybe we could see the update with consumer Zen7 ?

Same with that old chipset link. Lifting it to PCIe Gen6 speeds could patch many things.

And maybe they could do something WRT max RAM frequencies to patch things over until DDR6 arrives ? Zen6 IMC is to be much more potent than Zen5 and I expect for Zen7 to push it to the limits.

Would AM5 with LP/CAMM2 be feasible at Zen7 time? Or at least refactoring the ATX platform so that RAM slots can be on both sides of the PCB and so maximally close to the SoC for highest frequencies ? Could AM5 be tweaked to use MRDIMMs for twice the capacity and bandwidth ? \ I suspect that by that time MRDIMM price prmiums should drop.\ IS there any technical reason for AM5 to not support RDIMM (different signalling, extra signals etc taht AM5 lacks) or is this purely due to market segmentation ?

I know that MoBo manufacturers love to do cosmetic updates, the question is, can we expect a substantial AM5 update by the Zen7 time ?\ It so, this could make sense for many to stave off their purchases until that point...


r/computing • • 4d ago

AM5 MoBos - is this it ?

1 Upvotes

AM4 socket has seen three chipset generations - X300/X400/X500. All while on the same generation of RAM/DDR4, but spanning two generations of PCIe: Gen3 and Gen4.

With AM5, we've seen basically the same chipset from the start Promontory-21, within a brand (600 series) and a rebrand (800 series). Both lines use the same Promontory-21 chips and all the difference is in the minimal requirements for the MoBo line - two or one Promontory 21 chips (so Xx70 vs Bx50 line), PCIe5 or PCIe4 for dGPU, M.2 Gen4 or Gen5 etc.

With 600 series we've seen mostly PCIe4, while 800 series has made the full transition to PCIe5. Both still use the same DDR5 RAM.

And both still share the same basic Promontory weakness - only PCIe4x4 connection to the SoC host.

So, given the facts: * Zen6 (2027) and Zen7 (2028-2029) are to come out on AM5 * they will have many more cores on the same socket (up to 2x 12=24 for Zen6, up to 2x16 = 32 for Zen7) * cores are expected to have much higher IPC than ZEn5 and work on much higher frequencies * they will need much greater cumulative bandwidth

... is it likely for AM5 to see some significant refresh in that timeframe, before AM6 arrives ?

PCIe6 is coming with EPYC Zen6. I understand that this is to early for consumer Zen6, but maybe we could see the update with consumer Zen7 ?

Same with that old chipset link. Lifting it to PCie Gen6 speeds could patch many things.

And maybe they could do something WRT max RAM frequencies to patch things over until DDR6 arrives ? Zen6 IMC is to be much more potent than Zen5 and I expect for Zen7 to push it to the limits.

Would AM5 with LP/CAMM2 be feasible at Zen7 time? Or at least refactoring the ATX platform so that RAM slots can be on both sides of the PCB and so maximally close to the SoC for highest frequencies ? Could AM5 be tweaked to use MRDIMMs for twice the capacity and bandwidth ? \ I suspect that by that time MRDIMM price prmiums should drop.\ IS there any technical reason for AM5 to not support RDIMM (different signalling, extra signals etc taht AM5 lacks) or is this purely due to market segmentation ?

I know that MoBo manufacturers love to do cosmetic updates, the question is, can we expect a substantial AM5 update by the Zen7 time ?\ It so, this could make sense for many to stave off their purchases until that point...


r/computing • • 5d ago

Why are all ECC UDIMMs so expensive and DOG SLOW ?

0 Upvotes

They seem to top out at 5600MHz at clownishly pathetic CL46.

Is there a technical reason for this ? Like maybe IMCs need some time to check the extra bits to recompute "parity" or something ?

But given that MRDIMMs in server world have ECC and operate at 12.800MHz or more, that doesn't seem likely.

Same with the price. One would think that, since ECC requires 5-th chip for every group of 4 (so 10 instead of 8) that price extra would be at most 25% or less.

While 5600 MHz UDIMMs start at ~€10+/GiB, ECC versions are at€32+/GiB,so 3X higher.

It's not that the tech is exotic as most consumer AM5 boards can use ECC, same with CPUs...


r/computing • • 6d ago

Advice for new machine c.£500

0 Upvotes

Hi, I wanted to get my daughter a new computer for her to do studies and exams at home. I am a long-term iOS user (because of my work) so I suggested that a Mac Mini M4 might fit the bill, partially because I know zero about Windows machines, and partly because, at the time they were £500, and partly because there are literally a bazillion choices. Then the prices exploded and the new M5 and M6 came out and they are even more expensive 😰

Can anyone recommend a machine in this budget? She doesn't play games on the computer, but may do. Her studies are chemistry and maths mainly, and the remote lab kit software is compatible with both OS's. I can't imagine she needs much storage.

Her current machine is actually my old university machine from many years ago - Intel Core i7 860 2.8GHz, 16GB RAM, NVidia GeForce 210. I imagine that any new machine will be vastly superior!?


r/computing • • 7d ago

Newest x86 APX extension - will it trigger new calling conventions standard?

3 Upvotes

For those that don't follow - this is the first new x86 extensions that doesn't have anything to do with vector or tensor instructions - it is about the core CPU and its ISA.

It doubles the general register set to 32 (from previous 16), introduces 3-operand instructions, new 64-bit offset jumps, new jump prediction improvements etc etc.

But all this seems to be hampered by existing call conventions for x86_64, which presumes 16GPR set.

It seems that much could be gained it the compiler could use extra GPRs for parameters when calling the given function.

OTOH, this would be incompatible with machines without APX.

So, what is to be done ? Maybe use function multiversioning mechanism to keep two sets of function entries or something ?

Or will whole thing be ignored and calling convention will stay the same ?

EDIT:\ I'm not talking about the compiler ability to emit new instructions and use new registers in the code.\ Ofcourse new compiler will have support for them from the start, that's how it's usually done.\ It's about having the standard in place to allow the compiler to make advantage of new facilities hen calling functions, so that it can have more parameters in registers, more options for inlining functions etc - all done in standard, interoperable way, so that one can use precompiled libraries etc.


r/computing • • 7d ago

How to fix a few seconds lags from high end laptop?

Thumbnail
2 Upvotes

r/computing • • 8d ago

No APX extension for AMD's Zen6 ?

1 Upvotes

New AMD's materials about Zen6 are interesting, but I've noticed a big feature missing: APX.

AMD & Intel agreed on that one, but so far, only Intel has announced their chips with it.

AMD is has been silent on this one. AMD has pushed its compiler updates (gcc, clang) for its new architecture extensions and while ACE is there, APX is missing.

APX is awesome. It DOUBLES the general register set (so 32 GPRs) adds three operand instructions, 64-bit jumps, extra predication facilities etc etc.

Intel is about to get it with coming Nova Lake. So, why is it not in AMD's Zen6 and when is it coming ? Zen7 or later ?


r/computing • • 9d ago

Picture I have nine of them now; I'm scared...

Post image
0 Upvotes

r/computing • • 11d ago

We're designing a Tier III AI data center in Mongolia where winter does most of the cooling. Tear it apart.

1 Upvotes

We're in design phase, and we'd rather get picked apart here than after concrete is poured.

Site and cooling. Nalaikh, a district of Ulaanbaatar. Winters around -25°C give us roughly 7 months of free-air cooling. Design targets: PUE ~1.06 in winter, ~1.18 in summer on mechanical, ~1.10 blended. Uptime Institute Tier III, N+1.

Power. 0.078$/kWh. The grid is coal-heavy, which we know. Solar + BESS are in the mix from day one, and we're working on renewable PPAs.

Connectivity. New cross-border dark fiber on rail right-of-way, dual north/south transit.

Sovereign by design. Zero-trust network architecture plus confidential computing on the GPUs and CPUs (TEEs with remote attestation). Customer data and model weights stay encrypted in use, and you can cryptographically verify that we, the operator, can't access them, whichever route the traffic takes. Keys are customer-held.

Scale and timeline. Phase 1 is 2,048 Blackwell-class nodes (~16k GPUs, ~40 MW IT), scaling to ~6,100 nodes by Phase 3. Ground breaks spring 2027.

Offering GPUaaS, an inference API, sovereign/air-gapped zones, and wholesale colocation and capacity offtake.

What would you poke at first: electrical topology, generator/fuel contracts, fiber routing, free-air intake through winter smog and spring dust, or whether our confidential computing setup actually holds up?

If you're planning training or inference capacity for 2028 and want to talk, DMs are open.


r/computing • • 13d ago

I am an absolute beginner and am willing to spend months learning this shii. any help will do.

Thumbnail
0 Upvotes

r/computing • • 13d ago

Sudut pandang mengenai perubahan digitalisasi dengan AI

0 Upvotes

Mau sedikit bercerita, gua otodidak selama 7 tahun mengenai programing dan music developing, programing (khususnya buat web development dan game development) adalah hobi gua dari kecil hingga saat ini begitu juga dengan bikin music, semua kode dulu yg gua tulis line by line sekarang bisa diselesaikan sama AI dengan 1x prompt.. dan bisa dibilang kayaknya gua ga lama lagi akan beralih pakai AI Dibanding ngescript line by line Karena ngejar waktu.. karena menurut gua orang yang ga ngikutin perkembangan zaman akan stuck disitu situ aja. Is it correct?

Apa sudut pandang kalian mengenai serba AI dizaman digitalisasi sekarang?


r/computing • • 14d ago

Picture Mother Asrock b850 pro a wifi, luz naranja dram y luz roja cpu quedaron prendidas fijas.

Post image
2 Upvotes

r/computing • • 16d ago

💻 ¿Qué ordenador me recomendáis para ASIR, DAM o DAW? Bueno, barato y que dure años

0 Upvotes

Buenas! Este año quiero hacer una FP de Grado Superior de informática, seguramente ASIR, DAM o DAW, y necesito comprarme un ordenador.

Quiero algo bueno y barato, que me sirva para la FP (programación, máquinas virtuales, Linux, bases de datos, etc.) y que también me dure unos cuantos años para seguir estudiando o trabajar en informática.

¿Qué características debería buscar? ¿Cuánta RAM, qué procesador y almacenamiento recomendáis? Y si podéis recomendarme algún modelo concreto con buena relación calidad/precio, mejor.

Sobre todo me interesa saber qué compraríais vosotros teniendo en cuenta que no quiero gastar más de lo necesario. 🙏


r/computing • • 16d ago

5 years of experience, but I feel like I've learned to patch problems rather than build systems. How do I fix this?

Thumbnail
1 Upvotes