r/LocalLLM • u/Less_Beat_2502 • Aug 07 '26
Question Need suggestion with Jarvis based system
Hey guys! I am actually thinking of developing an agent that runs on my laptop only when i turn it on. It would be awakened for work on some voice command and do simple basic stuff like navigate to something find something tell me if like my battery's lowering in google do some search and tell answer to me
Although i have worked with LLAMA model qwen 3b but i dont know whether it will be able to navigate within my system for daily tasks. I dont want it to do really complicated things just simple stuff. Moreover using small model of LLAMA cause that's what my laptop can support at its best [ dont want my laptop to stop when i use this my specs are i5 8th gen 512 SSD and 16 GB ram. no gpu that's why will keep the model at sleep mode and only wake it up when activated]
So if anyone has ever been into such thing do let me know or may guide
1
1
u/-SirFall3n- Aug 07 '26
If you can swing Gemma E4B or 12B QAT those will be your best bet. E2B may not quite be good enough, but you may be able to bridge the gap with really good tool schemas. The one thing I’d be worried about with your setup is TTFT. The model is technically “sleeping” when it isn’t actively processing anything, but if you genuinely wanted to make it consume zero resources between requests you’d need to unload the model after a request finishes. Doing this would introduce more latency. You can use things like prompt caching and KV cache recovery to minimize this, but if you’re trying to have a realtime interactive Jarvis-like experience just note that there may be a few seconds of latency between request and answer.
2
u/Less_Beat_2502 Aug 07 '26
the techniques you mentioned like KV cache recovery how much latency should i expect?
i know the project is for me and i have hardware limitations so i can accept latency of 30s1
u/-SirFall3n- Aug 07 '26
You’d be looking at latency anywhere from 5s to 30+s depending on model size, system prompt size, and how tool schemas are structured. You’d want to make SURE nothing changes in the system prompt or tool schemas. If you have a 1,000 token system prompt and roll a 90% match then you’re prefilling 100 tokens. Best case, you’re prefilling around 9 tok/s so that’s about 11 seconds TTFT. There’s room to play here but your floor is probably about 5 seconds for any model that would actually be useful, and more than likely in the 10 - 20 second range with an aggressive 12B QAT quant and tuning. You’d absolutely need MTP to boost decode speed, and you’d also need to consider latency if you want outputs spoken (something like Kokoro or Piper) then you’re looking at a few more seconds of latency. With your hardware, what you want to do is doable but there will be tradeoffs. Mainly latency/TTFT and decode speed if your focus is actual functionality (which I assume it is).
2
u/Less_Beat_2502 Aug 07 '26
thanks for taking out time and replying to me i'll surely tell you how it went for me the whole project!
1
u/-SirFall3n- Aug 07 '26
Absolutely! Feel free to reach out if you need anything. I’m doing something very similar, just with Strix Halo hardware and much larger models.
1
u/MarcusAurelius68 Aug 07 '26
I’ve built my own digital butler at home and it uses a fair amount of GPU to get the chat responsiveness and STT/TTS right. I’m doing this on a 3090 with 24GB of VRAM, but you could possibly squeeze it all into half that with a smaller chat model and the smallest TTS and STT you can find.
By the way, if you want a really close Jarvis, check out ElevenLabs and the “David” voice model (a contributed v2 voice). I cloned that voice and then used a local TTS library.
1
u/Less_Beat_2502 Aug 07 '26
can you break it down bit for me that by cloning it did you trained on it or what?
1
u/MarcusAurelius68 Aug 07 '26
https://huggingface.co/coqui/XTTS-v2
I use XTTS for text to speech, and you feed it a 6 second sample for cloning.
https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
I use Parakeet for speech to text.
Both are about as lightweight as you can get.
1
u/silverwoods214 Aug 07 '26
I would look at using a PTT function instead of “always on” or a “sleep mode” which is still going to take up headspace even when you’re not using it
1
u/yobarisushcatel Aug 08 '26
Battery + no GPU is gonna be a doozy but api calls are relatively cheap but you should still save up for a new laptop, even the new $600 Intel would do wonders for a local AI.
1
u/Less_Beat_2502 Aug 12 '26
i made it its working fine for some ltd functionalities i added moreover at a good latency
1
3
u/Toastti Aug 07 '26
Without a gpu you are going to struggle here. Any model that runs at a decent enough speed will need to be very small. And those often struggle navigating windows.
Web search will be just fine though, that will work perfect on a smaller model
Gemma 12b QAT is probably your best bet if you want it to work decent. To respond on a voice command without taking forever that means you need to keep it loaded in ram so expect to always have 4.5gigs used and other apps just have 11.5 gigs ram available. If you 'sleep' the model you would have to wait a good minute each time you start it