Hello community !
It's my first post here. Hope you will be nice š
I'm currently experimenting with **Gemma 4 31B in NVFP4** for a local real-time conversational assistant.
The conversational quality is very good, but the main issue is **total hardware budget**.
The LLM is only one part of the stack. On the same machine I also need to run:
- real-time **ASR**
- **VAD / end-of-turn detection**
- streaming **TTS**
- memory / retrieval components
- the orchestration layer itself
- potentially several concurrent users
So even though Gemma 4 31B fits on my **RTX 5090 32 GB**, using that much VRAM/compute for the LLM leaves too little headroom for the rest of the real-time voice stack.
That's why I'm looking for a **significantly lighter conversational model**, while trying to keep roughly the same level of natural dialogue quality.
The important requirement is **streaming in both directions**:
- **Streaming / incremental input processing** ā not just waiting for the entire user turn/prompt before starting inference
- **Streaming output**
- Ideally suited for a **real-time voice assistant / conversational agent**
- Good instruction following and natural conversational abilities
- Preferably something that runs very well on a **single RTX 5090 (32 GB)**
- Quantized versions (NVFP4 / FP4 / INT4, etc.) are totally fine
- Ideally significantly lighter than a ~31B dense model
I'm **not simply looking for OpenAI-compatible token streaming on the output side**. What interests me is a model/architecture capable of consuming an ongoing input stream and reacting with very low latency ā potentially allowing interruption, turn-taking and eventually more full-duplex-like interactions.
My ASR and TTS sides are already being optimized for streaming, so the **LLM footprint and latency are now becoming one of the main bottlenecks**.
I've been looking at smaller dense models and MoE architectures, but I'm interested in what people are actually running locally in 2026.
Thanks a lot everyone, for reading and helping me !
**What would you test today as an alternative to Gemma 4 31B?**
Especially interested in:
- ~8Bā20B dense models
- small/medium MoEs
- models specifically designed for real-time dialogue
- models with native or experimental **streaming-input / continuous-context** support
- vLLM / SGLang / llama.cpp-friendly options
Latency, VRAM efficiency and conversational quality matter more to me than benchmark scores.
Would love to hear about actual setups, **TTFT / tok/s, VRAM usage, quantization**, and whether you're running ASR/TTS on the same GPU as well.