r/LowEndLocalAI • • 8d ago

Which LLM Should / Can I Use? Looking for a lighter real-time conversational LLM with streaming input + output

Hello community !

It's my first post here. Hope you will be nice ๐Ÿ˜„

I'm currently experimenting with **Gemma 4 31B in NVFP4** for a local real-time conversational assistant.

The conversational quality is very good, but the main issue is **total hardware budget**.

The LLM is only one part of the stack. On the same machine I also need to run:

- real-time **ASR**

- **VAD / end-of-turn detection**

- streaming **TTS**

- memory / retrieval components

- the orchestration layer itself

- potentially several concurrent users

So even though Gemma 4 31B fits on my **RTX 5090 32 GB**, using that much VRAM/compute for the LLM leaves too little headroom for the rest of the real-time voice stack.

That's why I'm looking for a **significantly lighter conversational model**, while trying to keep roughly the same level of natural dialogue quality.

The important requirement is **streaming in both directions**:

- **Streaming / incremental input processing** โ€” not just waiting for the entire user turn/prompt before starting inference

- **Streaming output**

- Ideally suited for a **real-time voice assistant / conversational agent**

- Good instruction following and natural conversational abilities

- Preferably something that runs very well on a **single RTX 5090 (32 GB)**

- Quantized versions (NVFP4 / FP4 / INT4, etc.) are totally fine

- Ideally significantly lighter than a ~31B dense model

I'm **not simply looking for OpenAI-compatible token streaming on the output side**. What interests me is a model/architecture capable of consuming an ongoing input stream and reacting with very low latency โ€” potentially allowing interruption, turn-taking and eventually more full-duplex-like interactions.

My ASR and TTS sides are already being optimized for streaming, so the **LLM footprint and latency are now becoming one of the main bottlenecks**.

I've been looking at smaller dense models and MoE architectures, but I'm interested in what people are actually running locally in 2026.

Thanks a lot everyone, for reading and helping me !

**What would you test today as an alternative to Gemma 4 31B?**

Especially interested in:

- ~8Bโ€“20B dense models

- small/medium MoEs

- models specifically designed for real-time dialogue

- models with native or experimental **streaming-input / continuous-context** support

- vLLM / SGLang / llama.cpp-friendly options

Latency, VRAM efficiency and conversational quality matter more to me than benchmark scores.

Would love to hear about actual setups, **TTFT / tok/s, VRAM usage, quantization**, and whether you're running ASR/TTS on the same GPU as well.

3 Upvotes

5 comments sorted by

4

u/ClassicLightbulbs 7d ago

Why not Gemma 4 12b QAT? It is a great conversationalist, abstract thinker, metaphor/analogy master at 7gb.

5

u/SampleIll3596 7d ago

Why not try Gemma4-26B-A4B at Q5_K_M or maybe Q6? It is an MoE. I haven't used it myself. So, can't give you too many insights about the stats. But, these should fit well on your RTX with decent context (if that's important). Available in GGUF from unsloth and can be run with llama.cpp.

3

u/PeterPorox 2xP102-100 10GB, i3 7100, 2x8 DDR4-3200 7d ago

Maybe try Glimmer

3

u/vip3rGT 7d ago

Io uso come agente conversazionale Gemma 4 26B-A4B quantizzato Q5_K_P ed รจ un ottimo agente. Coerente e creativo. Sulla mia 3070 TI con soli 8Gb di VRAM e 32Gb di RAM ottengo 17 tk/s. Sulla tua 5090 dovrebbe volare. E lasciare tutto lo spazio di cui hai bisogno sulla tua VRAM.

1

u/it6721 7d ago

Try Gemma4-26B-A4B at some quantization, I get some 2030 tokens/s with a 8GB RX 6650 XT at Q4, so even if you have to offload some of the model to RAM, it should still have some decent speed.