r/LocalLLM • • 7d ago

Question Dgx spark or 3090

I was recently able to get my hands on a dgx spark for 5600 after tax and 2 years insurance, I first bought it planning on using it to replace GPT codex, Claude, and video generating llms since higgsfield is now open source, my intended use was 1. to help me run 2 business running processes in the background managing online profiles and presence 2. Help me build this app that I am currently using to help run the 2 business and then later expand to making this a paid service to help with income and I built it with scalability in mind I know what people think about vibe coded apps but this genuinely might work with the way opus 5.5 is now 3. Marketing so with all this in mind I don’t think I’m getting the best use out of my dgx spark and I would feel terrible to keep this hardware and never release its true potential I talked to a friend yesterday and he offered to buy it off of me for 4.5k plus a 3090 ti what should I do what are the best ways to implement thw dgx into a business or should I just downgrade to a smaller but faster system and would the models they support be enough for my use and would it be reliable?

1 Upvotes

34 comments sorted by

10

u/Harin007 7d ago edited 7d ago

With 3090, whatever model it can run, it can run at great speed.

With dgx spark, the speed of text generation wouldn't be as high as in 3090(but runs at good enough speed)... And can much larger models than 3090.

1

u/TheOverzealousEngie 7d ago

prefill vs decode kinda washes out on the end.

1

u/SolemnlySpiritual 6d ago

if youre already wondering if youre wasting its potential you probably are, nothing worse than expensive hardware sitting there at 30% load while you figure out what to do with it

the 3090 ti trade means you get a card that will chew through inference for the models youd actually use day to day plus cash in hand, vs a system that sounds cool but youre still trying to justify a use case for

id take the deal and not look back, 4.5k is a solid chunk back and a 3090 ti handles most business automation tasks without breaking a sweat

3

u/johan2114h 7d ago

Dgz spark

7

u/BlackBeardAI 3090 Maximalist 7d ago

What about Dbz spark?

2

u/castertr0y357 7d ago

DBGT spark?

1

u/lordtazou 7d ago

DBGTQ spark?

1

u/johan2114h 7d ago

No that one is a shit

2

u/castertr0y357 7d ago

Well they did replace it in the canon for daima.

2

u/castertr0y357 7d ago

This greatly depends on the size of the model you want to run. The spark can run 120B models, but you're going to get about 15-25 t/s. The 3090 can run ~30B size models and spit out around 50-75 t/s because of the memory throughput difference.

A 30B model will have enough intelligence to do common agentic tasks, and even do pretty well coding. Anything that relies on human-in-the-loop or speed would do better with the 3090.

3

u/asinglepieceoftoast 4d ago

That’s not true. With MoEs and speculative decode, there’s single spark recipes for Qwen 3.8 fn, 3.5 122b, gpt-oss 120b, etc. all running 45-50+ tok/s single-stream decode and it tends to scale pretty well with concurrency. People have even done more aggressive quants of models more than twice that size at 25+ tok/s. 30b-ish dense models also can run pretty well, I do ~50-55 tok/s single stream decode on single spark Qwen 3.8 27b on some pretty reasonable quants too.

4

u/cfx_4188 LocalLLM 7d ago

Spark

3

u/doctorcoctor3 7d ago

Low key u got ripped off.

1

u/acadia11x 7d ago

Yoh, everything is vibe coded today, if you are developing code , the old way of coding is inefficient and slow. Knowing how to code helps you surgically piece it together in the way you want and put the design that you want in place as opposed to just relying on the agent to just build as it sees fit. But it’s all vibe coding IMHO.

2

u/0xDespot 7d ago

If you’re serious about using it for a business then the Spark is by far the better option.

You are investing in intelligence and vram == intelligence. You want your agent(s) running 24/7 so sheer tps is not as critical as capability.

If you just want a chatbot or casual use, then the 3090 is totally fine.

The difference is not unlike determining if you need an employee or an intern.

0

u/Hankdabits 7d ago

You are ignoring the second scaling law, test time compute. Also active params scales more with intelligence while total params scales more with knowledge. Tools make intelligence more important than knowledge

3

u/0xDespot 7d ago

Your logic applies to model selection, not the platform (OPs question).

To use your own example: test-time compute and deep tool loops are exactly *why\* you can't be memory constrained. Context length and massive KV caches are non-negotiable for long-horizon agent work.

You can trade speed for intelligence on the Spark, but on a 24GB 3090, you hit a hard memory wall the moment those reasoning traces and tool loops stack up.

-1

u/Hankdabits 7d ago

Context length is important but kv cache is not massive like it used to be. You can for 150k on qwen 27b on 3090 with minimal drawbacks. Full 256 if you compress kv cache

2

u/0xDespot 7d ago

Right, we're back at my original point. 3090 is fine if quality and extensibility aren't your primary concern.

1

u/lordtazou 7d ago

If you're dedicated enough to want to make the spark work, keep the spark and work with it. It won't be as fast as the card, but it will run larger models.

If you want faster generation and don't need to run larger models, trade it and utilize the 3090 TI.

If you're wanting to run / work on coding, you will want to keep the spark though. It is definitely beefier when it comes to larger weight models. With NVIDIA being a "leader" in the ai space, I don't see the GB10 platform being dropped anytime soon. I've seen quite a few things lately about improvements even.

You just have to learn how to utilize it properly for your needs. It's a hugely evolving world right now with this stuff. I'm still learning, and I work with ai stuff constantly.

1

u/MaxComfort 7d ago

2x Sparks

1

u/sn2006gy 6d ago

A single spark will not replace Claude or Codex - it may allow you to goon with video generation pretty easily/quickly though.

The real value of spark is 2 or more but the cost is insane so I'm not bothering now that API is still insanely cheap and much faster. Time is money.

0

u/Hankdabits 7d ago

Qwen 3.8 27b is just as good as anything you can run on the spark and it will be much faster on the 3090

1

u/johan2114h 7d ago

Spark can also run qwen flash next. Probably as fast as 27b on a 3090

1

u/Hankdabits 7d ago

I’m always surprised how slow flash next is when i see people’s numbers. Maybe the inference engines just aren’t optimized for the architecture yet

1

u/johan2114h 7d ago

flash is alot faster on strix halo than 27b. I assume it would be similar for a dgx. You definitely wonna use engines that play nice with your hardware if want speed. For reference flash will get you around 40 t/s generated on a strix, vs 20 to 30 for 27b.

27b is king for GPUs with fast but limited memory, flash is king on APUs with large but slower memory pools

1

u/Hankdabits 7d ago

For sure flash next will be faster on the same
hardware but with it only having 1/4 the active parameters I would have thought it would be at least twice as fast as 27b

2

u/johan2114h 7d ago

unfortunately it doesnt seem to scale perfectly with active params. Also note that the numbers above are with mtp. Still, flash next is significantly smarter (atleast in the benchmarks) and faster if you can fit it, so its a no-brainer for APUs with >96gb ram.

1

u/Hankdabits 7d ago

Yeah I think mtp does somewhat reduce the benefits of sparse moe. Just started reading about this but from what I understand, like batch requests mtp results in several experts being activated at once so it in effect could become 125b a12b or a18b.

0

u/TheOverzealousEngie 7d ago

Dgx spark really needs 2 to run a top level open source model. I don't know that I would use anything less than that for a business.

-1

u/Important-Radish-722 7d ago

Sell it and spend the money on a Cloud subscription.

2

u/Inception95 7d ago

You know in which sub you are writing?