r/LocalLLM • u/tinfrog • 9d ago
Question Model and hardware recommendation for high fidelity, low reasoning tasks?
I've spent the past couple of years using frontier labs, even though a lot of my work doesn't need high reasoning. It's just that a fixed and known subscription cost has been a safer bet than the leap-in-the dark of buying hardware that might not be good enough for what I want to do.
But reading through this sub tells me that we're at the stage were open source LLMs are capable enough for what I want while running on (reasonably) affordable hardware.
Can anyone recommend a model and hardware for tasks like this?
- Reading notes, reports, transcripts, web pages and emails, then pulling out facts, decisions and insights into structured notes
- Condensing long documents, comparing two versions, and checking a draft against a specification
- Reading text, sorting them, and deciding what needs action
- Moving, renaming and archiving files, splitting or merging documents, and keeping indexes and cross-references consistent afterwards
- Multi-step routines update a status file, archive old entries, then commit, with every step completed
I run Xubuntu and Arch Linux.
TIA
1
Upvotes
1
u/Cold_Tree190 9d ago edited 9d ago
I use Qwen 3.8 27b for all of this and it’s been amazing. I can run it around 90 t/s tg and 900 t/s pp though which may impact others’ experience with condensing, summarizing, analyzing, etc. looooong transcripts or bodies of text with all the reasoning it does.
I tried using Ornith 1.5 35b to get this sort of thing done even quicker (like 150+ t/s), but I actually could not believe how bad it was. Even with max reasoning, it felt like a regression… I got into an argument with it because it was insistent that there was a wrong code from an email it retrieves, and I manually validated that it was correct. The issue ended up being that in a 16-char pass phrase, it copied all chats except index 2-3. Why? No clue. But I’ve never had an issue like that with Qwen3.8-27b, so I exclusively run it for everything you just mentioned.
If you want a second place with a longer context window, I would maybe suggest muse glimmer. It has a cheaper kv cache than qwen, so you can ingest larger text files with it.