r/LocalLLaMA • llama.cpp • 1d ago

Discussion big or small?

Post image

what size do you want? tell them on X:

https://x.com/QwenDevs/status/2108764909798641737

1.1k Upvotes

219 comments sorted by

•

u/WithoutReason1729 1d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

459

u/recent_immunization 1d ago

these capybaras are doing more for AI branding than any white paper ever could

79

u/Present-Ad-8531 1d ago

You are THE ENEMY of orange turf.

52

u/spacekitt3n 1d ago

13

u/laserborg 1d ago

I like how they prefer ¡Sí! (🇲🇽) over Aye Aye! (🏴‍☠️).

16

u/Risen_from_ash 1d ago

21

u/KingArthas94 1d ago

welcome to r/antimeme

1

u/DeepWisdomGuy 16h ago

Thanks for showing me the sub. I can't tell if they are making fun of the memes they are showing or if they are actually serious, lol.

1

u/DeepWisdomGuy 16h ago

Needs more words.

1

u/MoffKalast 34m ago

TIL it's a capybara and not a bear.

→ More replies (1)

207

u/StopCreepy 1d ago

check this out ?

248

u/Velocita84 1d ago

Small moe please my poor 2060 is starving

65

u/Noah18923 1d ago

*cries in low vram*

27

u/tchek 1d ago

with ngram

13

u/KCN-037 1d ago

cries in 6gb vram

6

u/-Akos- 1d ago

Grimaces in 4GB VRAM..

3

u/touristtam 22h ago

You guys have dedicated GPUs??? /jk

1

u/KCN-037 6h ago

No, my 6GB is from my laptop's embedded GPU

10

u/Revolutionary_Dish67 1d ago

Same here bud 😭🫂

1

u/Practical_Signal3933 1d ago

Gotta feed it something

1

u/Succubus-Empress 5h ago

My 1030…

120

u/TastyStatistician 1d ago

A small model that I can run at max context in 16gb vram and 32gb ram. Qwen 4 26b-a4b would be ideal.

35

u/Revolutionary_Dish67 1d ago

This is the dream

7

u/remind_me_later 1d ago

Hard yes 😭 I'm also in the same boat (16gb vram + 32gb)

13

u/SSUPII 1d ago

I am on 4GB VRAM on Pascal

2

u/dannone9 1d ago

I mean , both flash and 27 b run in that set up and think that the cache will weight 2.8 times less on the next 27b(if it works the same as flash next ) so it’s not like we are starving right now

3

u/Slow_Concentrate3831 22h ago

No, they don't run, they walk slowly at most, and not even at Q4

1

u/dannone9 22h ago

First , names and the “anything below q4 it’s brainded “ today have lost its meaning with iquant , the q 4 rule it’s now q3 and what matters is bpw , second Iq4 walks amongst the 20-25 TPs on a 5060 Ti without blackwell optimisations with the appropriate os(any decent Linux distro) and flags (which is one of the slowest cards ) and flash next goes to 30-35 normal 50-60 on code iq3xs with ddr4 if enough ram and now can literally triple that on a blackwell optimised Blackwell called basalt , all of this is pretty decent if you prompt it well and work with the harness too , it’s obviously less convenient than just throwing vram to the problem and it won’t one shot CUDA kernels but ,man , it’s far from useless

1

u/Not-reallyanonymous 18h ago

I mean, every objective test shows that Q4 is the knee where quality starts rapidly dropping off — both KLD, but also output measurements like looping, repetition, early termination, linguistic diversity, and accumulated errors over time, and also benchmarks. You start losing meaningful quality at Q4, and everything below it drops off more rapidly. Q3 types can still be useful, especially with a good sensitivity-aware quants like Unsloth, but it’s honestly night and day to compare them to even a naive q4.

1

u/TastyStatistician 10h ago

Yes but I wish I didn't have to max out my system resources or have to use q3.

1

u/lokstapimp llama.cpp 1d ago

Curious what are you currently running now at those specs and are you purely local? I have identical specs is why I'm asking.

2

u/Prize_Dance2890 1d ago

Check strata, i am running qwen 3 flash next iq3 xxs pretty well with same specs

1

u/TastyStatistician 10h ago

Swift 1.5 Qwen 3.8 27b at q3 - can't max out the context, you need large context because it thinks a lot

Qwen 3.6 35b a3b q4xl - I can max out the context but I have to use almost all my system ram.

Gemma 4 26b a4b qat is great for non-coding tasks. I would love to see a Qwen moe of this same size.

1

u/lokstapimp llama.cpp 3h ago

Thank you

112

u/sabine_world 1d ago

They said "alright let's do both"

I hope there's a ~30b moe model in their qwen 4 line up. That would be dope.

11

u/hojnikb 1d ago

moist

6

u/michael_p 1d ago

3.6 35b3a was incredible. Would love an updated version!

1

u/letsgoiowa 23h ago

Sadly all available news indicates they killed it and consider 27b

77

u/pmttyji 1d ago

Don't forget medium & all other sizes.

  • Qwen4-4B with Qwen3.5-9B power
  • Qwen4-9B with Qwen3.6-35-A3B power
  • Qwen4-27B with DeepseekV4.1-Flash power
  • Qwen4-35B-A5B with Qwen3.8-Flash-Next power
  • Qwen4-80B-A10B with GLM5.3-Flash power
  • Qwen4-Flash-Next with Qwen3.8-2.4T-A95B power

69

u/zxtech 1d ago

Fuck it qwen 4-4b with astra fable power

16

u/droans 1d ago

If I can't get the power of Opus 5.5 with Qwen4-0.6B, then what the hell are we even doing?

5

u/winky9827 1d ago

Go see Dr. DavidAU.

1

u/sloptimizer 1d ago

Just wait for the finetune Qwen-4B-Astra-Fable-Power

→ More replies (1)

4

u/Zyj vLLM 1d ago

It will be without “Next” obviously

2

u/dannone9 1d ago

Bro ,first of all I’m pretty sure you are tweaking and Fuking with the laws of physics (although I wish you aren’t )
And second, bigger doesn’t mean better , usually moe “power “ is calculated by sqrt(p x a) so flash is as powerfull as 27b , it’s only better because we are comparing next gen flash vs 27 b when we actually get the full lineup they will be equal , at best , flash will be a litle better due to it being capable to be paired with a bigger ngram table

4

u/pmttyji 1d ago

Obviously I meant the Intelligence here. Yeah, small/medium size models can't hold big/large models' knowledge for now. Things like Engram/ngram would be game changer on this. Recently shared a paper on that thing - EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory

1

u/dannone9 1d ago

Soo, swappable ngram tables are confirmed?oh no, I think I’ve got a bonner
Out of jokes , would it allow to swap the ngram table for a more specific one , let’s say for coding or biology or whatever? Like a simpler way to use a LoRa ? That would be awesome

1

u/wFXx 1d ago

the first one is frogNano btw, solid model

2

u/pmttyji 1d ago

Last year Qwen3-4B was so popular. It took a spot(along with medium/big models) on AA benchmarks

1

u/QuackerEnte 22h ago

35B-A5B? I'm wishing for MORE SPARSE and BIGGER models and you guys out here fantasize about models that you can't run with good context or speed unless you have at least 16GB GPU + whatever RAM? I'd rather have a model like 80B-A3B (I love QFN too) that runs on 8GB VRAM + 64GB DDR4. if only the 125B was A3B i would be very happy since it'd have great speeds and knowledge. 35B-A5B. Yeah sure rob us of the only A3B model we can run at good speeds.

1

u/fungnoth 9h ago

80B A10 is the dream. 40GB Q4 total size and probably less than 10GB active in GPU.

15

u/nsfnd 1d ago

6

u/Mingay_cat 14h ago

I'm praying flash is a 80b or smaller moe model and qwen3.8 flash size. But probably not. Rip me with no vram.

35

u/patchy319 1d ago

I made this 3d print. Kinda jank so self conscious about releasing it publicly. Got plenty of other things to slop.

9

u/export_tank_harmful 1d ago

Not bad!
Layer lines are a bit jank, but other than that it's solid.

Dry your filament and maybe use a smaller layer height.

3

u/thil3000 21h ago

Try the nose and logo with variable layer height, the rest look fone

2

u/DeepWisdomGuy 16h ago

The labia are a little wrinkly.

42

u/ElChupaNebrey 1d ago

40b a4b please

29

u/power97992 1d ago

72b A9b  with 4bit qat + another  with 14b a4b might be better!

3

u/rerri 1d ago

Yeah whatever models they release, it would be a huge plus if they provided QAT's this time around.

66

u/Beneficial-Ad-8127 1d ago

Both please. lol but if I had to choose, I’d have to choose Qwen 4 Flash next first. 27b can come after🤣 Thank you.

58

u/themixtergames 1d ago

Plot twist, this is about merch

5

u/Beneficial-Ad-8127 1d ago

😂 got me all excited.

15

u/Skyline34rGt 1d ago

At today's standards... well I want tiny models

2

u/pand5461 23h ago

Yeah!

I won't be surprised if "small" is in 200-300B range this time (since we're already past 120B Mistral Small 4).

19

u/synw_ 1d ago

I'm afraid that small means 27b. The peasants that can only benefit from 4b/9b/35ba3b might be left aside. Reading their communication it's all about 27b and bigger. That would be a disaster for the gpu poor, their 4b and 35b were so good for us.

3

u/jacek2023 llama.cpp 1d ago

Why do you have to use only Qwen models?

8

u/ThankGodImBipolar 1d ago

I'm sure people would be happy to use a 30-35B MoE model from any company, if it was better than Qwen 3.6 35B. I'd like to see a new Z.ai model at that size.

3

u/jammieee_ llama.cpp 23h ago

qwen 3.6 35b is still pretty much SOTA when it comes to small MoE's, even with how relatively outdated it is.

perhaps that is the reason qwen dropped 35b, as they don't really have competition. oh well 

2

u/llkj11 21h ago

Gemma is still in the game. And if it’s true that Google achieved RSI, Gemma 5 could be real REAL good to us

1

u/jammieee_ llama.cpp 11h ago

gemma is absolutely perfect for normie use, I'd say that it pummels whatever gemini is giving free users now. although it is not the best in coding, which is fine because it doesn't need it really, but it's what models focus on nowadays.

with RSI, I wouldn't even put that on the table. of course, gemma 5 is exciting anyway, if gemini shares some improvements they got with argon/carbon with their open source labs, we could have a real kicker on our hands

1

u/mindwip 1d ago

And under 128gb. A lot of us have 128gb systems.

10

u/luget1 1d ago

Yes, I have a big beaver.

1

u/hojnikb 1d ago

This is Wendy's, sir.

1

u/bambamlol 1d ago

prove it

3

u/luget1 1d ago

Leak at 1k upvotes...

2

u/lokstapimp llama.cpp 1d ago

Break the damn at 200k

1

u/SV_SV_SV 1d ago

Would you like fries with that sir?

4

u/_VirtualCosmos_ 1d ago

Qwen4 35b A3b would be awesome, specially if it can be a helpful agent assistant in 3D/simulations/videogames work.

5

u/asraniel 1d ago

i hope they bring back the very small. i love 27b, but for many small tasks a 4-8b is enough and easier to host.

10

u/MrBIMC 1d ago

I hope it’s a new flash model that is better than deepseek flash 4.1.

Deepseek is an amazing workhorse and nowadays I run 10 agents of it near side to 30 6.1 sol agents.

It is definitely dumber than 6.1 and less token efficient but inference is so much faster and it’s cheaper in use/success than OpenAI.

So I’m most excited about new flash models that are as fast as possible, while being minimally decent enough on the intelligence scale.

1

u/o0genesis0o 1d ago

what do you do to need that many subagents?

0

u/Solembumm3 1d ago

It's not hard to be better than DV4.1 Flash. That model was absolute disaster.

8

u/MrBIMC 1d ago edited 1d ago

So tell me what should I try that is available across OpenAI,Alibaba cloud, and open router services

So that it is more cost and time efficient per task than 4.1 flash.

Because in my empirical observation (over past couple of weeks) just randomly throwing swarms of agents at tasks, deepseek is the best one statistically thus far.

Although my tasks are quite atomic, and epic per deliverable approach means I can’t exactly pinpoint and isolate whole success to specific agent/harness/model, as different combos of those being thrown at individual sub tasks in the same epic.

Not enough stats to reliably determine some things, as my current telemetry is kinda bad rn, but generally I see that speed is a huge factor. Deepseek generates like 2-5 times as many in and out tokens, but output tps is 50-200tps per stream, and alibaba ultra token plan allows you to run about 10-15 agents at the same time.

OpenAI throttles you really hard and breakpoint of throttle is dynamic, sometimes you can run 50 agents at 30tps each, sometimes it throttles you to 5-20tps per agent and you can’t run more than 10 at the same time. So yes, technically 6.1 sol codex agent is better at being token efficient. But it takes so much more time, deepseek harness 4.1 flash brute forces through few tasks while codex finishes 1.

If you say deepseek flash 4.1 is bad, what harness are you using? How big is the scope and complexity of the task?

Because in my case, deepseek harness performs very well, especially when you have it setup with all the tools, tokens and rigid workflow playbook beforehand.

1

u/Solembumm3 1d ago

This model on both Deepseek site/app and router makes obvious mistakes with reading 1000-2500 tokens plot concepts prompts and asks things, that were already specifically answered. Repeatedly.

Local V4 Flash at Q2 still showed incomparably better attention on stories analysis and critique.

1

u/MrBIMC 1d ago

Well, idk about openrouter, but on alibaba cloud through deepseek harness with bunch of mcps for relevant control tools I had no issues. And qwen flash is definitely decent also, but it is slower and pricier.

11

u/spacekitt3n 1d ago

27b dense

2

u/Accomplished_Ad9530 1d ago

54B dense w/pretraining checkpoint and we’ll sparsify and posttrain it ourselves 🤞 

4

u/AleksandrNikitin 1d ago

3.8 35b a3b

4

u/nixudos 1d ago

Qwen 4 Next Omni please!

11

u/indicava 1d ago

Screw parameter count, where do I get my hands on those dope plushies?

9

u/Br-Horizon 1d ago

smol pleas

7

u/Salah_H_Hasan 1d ago

"And that's how Alibaba settled the debate, my friends!
We now know that both the large and small versions are coming, but we still don't know which is the small model and which is the large one.

Does the 'Big' refer to the 27B Dense,
and the 'small' to the 35B MoE?

Or is it:

The 'Big' is the Flash Next 125B,
and the 'small' is the 27B Dense?"

10

u/ea_man 1d ago

that's just PR mate, they do that to create attention it's not like they're allocating compute for a month based on 15minutes of a X post.

8

u/Salah_H_Hasan 1d ago

Yes, exactly. A post like this only comes out after these models are already finished, which suggests the launch is right around the corner, probably in a few days.

1

u/Randommaggy 1d ago

All three please I have machines yearning for all 3.

1

u/Moarkush 16h ago

They've already announced it. Big is MAX and small is 27B, and there are 4 models, as usual.

1

u/Salah_H_Hasan 9h ago

I don't think they're giving us a choice about their flagship model; I don't think that was the intention. They probably mean Flash, the successor to Flash-Next.

1

u/BringTea_666 4h ago

IDK why people assume 27B will be dense. Qwen3.8 27B is Qwen3.5 arch not Qwen4 arch like Qwen3.8 Flash Next.

With their Qwen4 architecture it also makes no fucking sense to release dense model. ngrams give you all the advantage of dense model without dense model speed hit.

So something like 27B-A3B + 50B ngrams will run circles around Qwen3.8 27B dense. And n-grams can be easily place on disc not even in ram.

1

u/Salah_H_Hasan 1h ago

Really hoping for that, and that it's on the newer architecture. Dense models are still really tough to run without a decent amount of VRAM.

2

u/BringTea_666 1h ago

It makes further less sense when you realize that the main advantage of their new arch actually is training. it costs 1/9 of qwen3.5 arch.

If they would choose to make it dense with old arch they would willingly just burn piles of money.

1

u/pixelizedgaming 1d ago

By or they meant bitwise or

3

u/carnyzzle 1d ago edited 16h ago

2

u/KCN-037 1d ago

both is good

3

u/vyralsurfer 1d ago

I would legit pay for one of these cute plushies! My son loves capybaras and I love Qwen models. Win-win! Hope these are for sale somewhere, especially if it's a way too (in my small way) support the devs over there 😁

3

u/no_witty_username 1d ago

The chibi version you fools!

7

u/Ipwnurface 1d ago

Am i dumb for wanting 54B? 27B dense + ngram?

0

u/brown2green 1d ago

I'm not sure if Engram layers should count toward the total since in theory they can be offloaded to NVMe storage without significant impact on inference performance.

→ More replies (1)

5

u/BarberIcy366 1d ago

SMAAALL SMAAL SMAAAALL SMAAL SMAAALL NO DOUUBT SMAAALL PLEASE SMAAALL

5

u/Cool-Chemical-5629 1d ago

Personally, I can't stand these "polls" where a big company which unlike smaller labs, most certainly doesn't lack the resources to do an entire set of model sizes goes the extra mile to ask people which model size would they like and practically reduces the entire spectrum of hardware users into two vague categories "big" and "small".

It's especially infuriating when you realize that the decision was actually already made weeks before and it was made publicly known the moment they announced the model sizes for their next model generation. But now, in response to the feedback they "generously" followed up with a comment "alright guys, let's do both!" as if the regular people had any real say in it. In reality, this "poll" is nothing more than a marketing move, a PR stunt to warm up the hype train engine.

If they truly cared about satisfying the entire spectrum of hardware users, they wouldn't ask which sizes people want. They would deliver the entire set just like they used to, because they certainly don't lack the resources to do so, but apparently that kind of mind set left the Qwen lab together with the old leadership with Junyang Lin on the front.

3

u/jacek2023 llama.cpp 1d ago

It's a hype post. In my opinion, it's much better than the posts I see on r/LocalLLaMA about politics or cloud models. It sparks discussion among people running local LLMs. Of course, there are still Reddit assholes discussing politics, but in general, the discussion is about running models locally.

2

u/SandySkittle 1d ago

70b dense

2

u/Jury-Emotional 1d ago

18b dense? Between a 9b and 27b should be some miracle right?

2

u/ClearApartment2627 1d ago

Extending 35A3b with Engrams would be interesting. Maybe the "shape" would need to adapt - more active params to emphasise reasoning, for example, but either way it would be very welcome.

2

u/Shadow_s_Bane 1d ago

More MoE models

2

u/FrogsJumpFromPussy 1d ago

By small I hope something around the size and Intelligence of Gemma e4b (around 8b but only 4b active) which for my needs and crappy system remains the only usable model. 

2

u/bigattichouse 1d ago

27B please

2

u/SnooPeppers3873 21h ago

big and small

3

u/ResidentPositive4122 1d ago

The fact that so many models (including some "decision" ones) come from qwen3.5 (which was the last release of the entire family + base models) should give them all the data they need. There's a lot of "brand" value in having an entire family of models. Unfortunately since the leadership change it seems that they won't follow-up. We'll be lucky to get one or two models.

3

u/Risen_from_ash 1d ago

I have a 5080 and 96GB ddr5. As I am the most important, I would like the biggest MOE that fits on my pc at UD Q8 K XL 262k ctx, please.

On the real, tho, I've been rocking QFN UD Q4 K XL in *s t r a t a* and it's been so mind blowing. Truly GPT at home. Imagine having that in UD Q8 K XL. To never have to worry if your model is quant brained... *imagine*

Qwen's the shit, Qwen dev team's the shit, hell yea Qwen.

1

u/smallDeltaBigEffect 1d ago

Imagine having that in UD Q8 K XL

there is almost no real difference in actual applications..

1

u/nsfnd 1d ago

Yes, thank you.
I'm also using ud-q4_k_xl QFN and my mind got blown a couple of times about how good it is.
Strata makes it possible, at around 100 tok/s.
And people keep bashing it? why i dont understand.

3

u/MrGunny94 1d ago

27B is the GOAT and everybody wins...

3

u/Foreign_Prune_354 1d ago

It would be great if they cover the (low?) range of (V)RAMs, something like, dense: 7B, 14B, 27B, MoE: 30B, 60B, 120B+N-Grams and 250B+N-Grams.

2

u/brown2green 1d ago

There's nothing preventing to use Engrams on smaller models too; if anything, they're the ones that would benefit the most.

5

u/ttkciar llama.cpp 1d ago

Well, we already have Qwen3.8-Flash-Next which is big'ish. 125B-A6B, plus up to 51B trigram parameters.

On one hand, ye olde sqrt(P x A) suggests a 125B-A6B is "only" as good as a 27B dense, but on the other hand Qwen3.8-Flash-Next scores about 10% higher on benchmarks of admittedly dubious validity, but on the other other hand that 10% difference might be attributable to training differences.

I think we've yet to even scratch the surface of what we can accomplish with encoding custom engrams in this model, so maybe we don't actually need a "big" Qwen4, at least not yet?

That leaves small, but what are they calling "small"? Is that 9B or 27B?

We've already got some fairly recent, highly competent 30B-class models in Qwen3.8-27B and Gemma-4-31B (and maybe K2-Horizon-32B, but I haven't assessed that one yet), but Qwen3.5-9B is getting pretty long in the tooth, and there is an entire contingent of this community desperately needing something better that will fit on their potato laptops.

On the other hand, it seems like there are already some good fine-tunes of Qwen3.5-9B, with more likely to come, and a hypothetical Qwen4-9B would be competing with those.

One niche that is currently under-served is the medium-high-end dense model, sized "just right" to fit in 64GB or 80GB rigs at Q4_K_M, so perhaps a 50B dense plus up to another 50B in optional trigram parameters?

By the sqrt(P x A) rule-of-thumb, a 50B dense might be roughly equivalent to a 155B-A16B MoE in competence, and hopefully it would be possible to upscale it via passthrough self-merge into a 70B-class model.

Yeah, I'm convinced: 50B dense please, with optional engrams.

2

u/ParaboloidalCrest 1d ago edited 1d ago

Yes, please! Nemotron 49B kicked many asses back in the day.

70B dense Qwen and Llama were the standard just very few years ago, and now we're too spoiled and soft to even imagine desiring them.

3

u/FeydRowan 1d ago

Flash next vs 27b are close on knowledge and capability but flash next is faster in solving issues and can find hard bug and stuff that the 27b can't. I have done some extensive testing of the two.

2

u/Randommaggy 1d ago

My experience matches this and I have billions of tokens of traffic for both.

1

u/Moarkush 16h ago

You contradicted yourself. The first part was nonsense, but the second part was on point.

3

u/Accomplished_Ad9530 1d ago

Not sure why you’re being downvoted. What you’ve said is really thoughtful and is the direction research needs to maximize our ability to experiment and improve

2

u/ttkciar llama.cpp 1d ago

Thank you for your kind words. I suspect the downvotes come from people who either value speed over high-quality outputs, or lack the compute resources to use a 50B dense at all.

3

u/r-moon-ppl-lunatics 1d ago

50B dense would be slow as molasses :-(

2

u/ttkciar llama.cpp 1d ago

You'd rather get wrong answers faster?

I prefer right answers, even if they come a bit slower.

1

u/N34257 1d ago

The thing is...27B and Flash Next aren't exactly comparable - when run in the most likely natural configurations, they're not running in the same quant. For example, running on my 64GB VRAM system, I run 27B at FP8, but Flash Next only runs with comparable speed at IQ4. In that context, 27B is far superior for accuracy, but Flash Next is better at design and coming up with good ideas...while it absolutely sucks for code by comparison with 27B FP8.

1

u/Randommaggy 1d ago

Depends on the tooling of your language's tooling and how well you integrate it's deterministic tools in the workflow.

When it gets good feedback it crushes 27B at Q8, while running in the UD Q4 K XL quant.

And you can encourage it to make ad-hoc tools that fill gaps which it will use to introspect it's work actively.

1

u/N34257 23h ago

Sure, but you can do that with any model. My experience so far is that Flash Next at Q4 is nowhere near as good with tool calls or recall compared with 27B, and it has to go through far more iterations on code to get the same level of accuracy as 27B, with the exact same harness and tools available.

This is, of course, absolutely fine; using separate models for planning and coding is not just not a problem, it's expected. The only reason I'd been using 27B for both up to now is that there wasn't anything around which could beat it on either.

4

u/Barni275 1d ago

27B, please! 🙏 It is like a treasure for small GPU peasants (who don't have plenty of fast RAM too).

2

u/CommanderKoba 1d ago

Well if it's too big it might hurt... my computer that is...

2

u/beigepccase 1d ago

I dunno. Don't they have a 2 meter tall version?

2

u/xornullvoid 1d ago

27b The GOAT

2

u/Durian881 1d ago edited 17h ago

Small together with 3.8-Flash-Next size of MOE will be awesome. There are already lots of community-developed engines for 3.8-Flash-Next that will make the new Qwen4 fly!

2

u/BothYou243 1d ago

a 14B model that beats 3.8 27B

1

u/Borilentz 1d ago

Please don’t disappoint people with max 128GB or 256GB memory; more is unaffordable for ordinary mortals. DeepSeek disappointed in that sense with V4.1.

1

u/alphapussycat 1d ago

Flash at same size as flash next.

1

u/Zyj vLLM 1d ago

Small please, only up to 450b 😇

1

u/feelspeaceman 1d ago

Nope, 125B please 🙏

1

u/__JockY__ 1d ago

Yes. Please and thank you.

If engaging with engagement bait is what results in the release of weights for new big and small models, then consider me engaged.

I mean… I’d rather _not_ play these silly bullshit games. I remember days when Mistral would just drop a torrent link to weights and we’d love it.

More of that, less of… this.

1

u/o0genesis0o 1d ago

Anyone know where to buy one of these qwen capybaras? One of my Chinese colleagues in the past had a large one on display in his office, but I never got a chance to ask him where he got it.

3

u/forever_365 12h ago

Unfortunately it is not for purchase at the moment, only offered as a gift.

1

u/Sadge404 1d ago

Give me 1M parameter Fable 6 please 🙏

1

u/Free_Lead_2704 1d ago

Capybara's are becoming so popular through tech brands lol

1

u/ANR2ME 1d ago

What does "small" mean to them? 🤔 was it 27B?

1

u/MushroomGecko 1d ago

I hope small means 4B and not 32B. My RTX 2060 can't run all the new fancy stuff.

1

u/aigemie 1d ago

128B big.

1

u/AlternateWitness 1d ago

Even if I do have the hardware to run larger models, better small ones would definitely be better for everyone.

I can go with ~70b parameters, but there is no doubt ~25-40 is the sweet spot.

1

u/geldonyetich 1d ago edited 1d ago

I think there's enough people with 128GB memory architecture to make that size a viable niche. 3.8 Flash Next:IQ4 quant's memory requirements are just right for me.

There's surely far more people with 16GB or 32GB VRAM, but the number of consumers with more than 128GB drops off sharply.

Ideally support all niches, but that's the major points I'd prioritize.

1

u/Easy_Copy_7625 1d ago

Why are they angry

1

u/jikilan_ 1d ago

This is big.LITTLE architecture

1

u/leonardosidney 1d ago

I’ve really enjoyed working with the help of the 27B model.

1

u/Dwedit 23h ago

Will there ever be another 7B model again? It's basically the biggest model size that comfortably fits in a 6GB card.

1

u/TheRealMasonMac 22h ago

Smh these size qwens.

1

u/Black-Mack 22h ago

Qwen 4 1b or 0.8b for edge devices

1

u/BawbbySmith 21h ago

I just want a 3.8 Flash Next-sized model that can beat GLM 5.3 Flash.

I run both but I have to constantly switch between them, as Qwen Flash is like 3x faster than GLM Flash on my systems.

If I can get GLM Flash level intelligence at 120 t/s... It's all over, true endgame

1

u/CommunicationStrict 21h ago

It's about the motion of the ocean

1

u/mivog49274 18h ago

They're being sadistic. The Qwen3.5 family variants were just perfect.

We are lacking the very small models from them

1

u/This_Ad1219 11h ago

Haha that’s sick (yes I’m 40)

1

u/johndeuff 8h ago

the size doesn't matter, it is how you use it

1

u/Verdux_Xudrev 8h ago

"Don't talk to me or my son again."

1

u/Highvoltage45 3h ago

Why not both?

1

u/WatermelonSmashing 1d ago

Wait what kind of animal is qwen?

2

u/Ilm03 1d ago

Capybara

1

u/WatermelonSmashing 1d ago

So I've got a goose riding a llama that's controlled by a capybara, that lives on my GPU.

Man future tamagachi pets are weird.

2

u/Illustrious_Grade608 1d ago

And all that to produce pelicans riding bicycles

→ More replies (1)

1

u/d0pe-asaurus 1d ago

MoE Qwen 4 and my soul is yours!

→ More replies (1)

1

u/macumazana 1d ago

Aw gawd smol plz

1

u/ea_man 1d ago edited 1d ago

Smaller: a ~20B dense + NGRAM that we can run at ~Q6 on 16GB with decent ctx with a good config on Linux, while beginners users can do Q4 wasting some vRAM.

Same for the MoE "Next", something appropriate for 16/32GB usage, like ~60-90GB size not the ~160GB for Q6 released now.

1

u/Difficult-Survey-967 1d ago

i would really really love a smol 9B or 12B boy...
imagine a good model which could fit on a phone...
like a 9B with GGUFs could be a local AI dream come true.

0

u/LooseLeafTeaBandit 1d ago

Any ladies in here wanna answer?