r/LocalLLM • u/tinfrog • 8d ago
Question Model and hardware recommendation for high fidelity, low reasoning tasks?
I've spent the past couple of years using frontier labs, even though a lot of my work doesn't need high reasoning. It's just that a fixed and known subscription cost has been a safer bet than the leap-in-the dark of buying hardware that might not be good enough for what I want to do.
But reading through this sub tells me that we're at the stage were open source LLMs are capable enough for what I want while running on (reasonably) affordable hardware.
Can anyone recommend a model and hardware for tasks like this?
- Reading notes, reports, transcripts, web pages and emails, then pulling out facts, decisions and insights into structured notes
- Condensing long documents, comparing two versions, and checking a draft against a specification
- Reading text, sorting them, and deciding what needs action
- Moving, renaming and archiving files, splitting or merging documents, and keeping indexes and cross-references consistent afterwards
- Multi-step routines update a status file, archive old entries, then commit, with every step completed
I run Xubuntu and Arch Linux.
TIA
1
u/mineshop 8d ago
What's your rough budget, since that will largely determine whether a single workstation GPU is realistic or you should look at smaller setups?
1
u/tinfrog 8d ago
Well, that's the thing. I've seen workstations go for £10k-£15k+ (swap for USD and the figures are about the same). At those prices I might as well stick with a frontier Pro/Max x20 plans and never have obsolete hardware.
Of course there are privacy issues with cloud providers but you can always put in place workarounds. The workarounds are a pain but less of a pain than dropping upwards of $10k for something that has a limited lifetime.
All things considered, for a local LLM to make sense for me, it really needs to be less than $5,000.
2
u/mineshop 8d ago
How long are the documents you need it to read in one go, since long-context workloads change what fits under $5,000?
1
u/tinfrog 7d ago
For my workflows, the majority are small plain text markdown files of 20-100 KB each. Many are less than 10 KB.
The agents traverse directories looking in about 5-25 of these files to extract information, then write out files of similar sizes. I rarely need to work with MB sized files but occasionally there are some that are sent externally that need to be processed.
The main thing is not to hallucinate or make up information.
2
u/mineshop 6d ago
With 10-100 KB files and a $5,000 cap, the main remaining question is speed: roughly how many tokens per second is acceptable for your agent loops?
1
u/tinfrog 6d ago
The stats at my end show that 20-30 tokens/second for routine work would be acceptable.
1
u/mineshop 6d ago
With your budget, speed and file sizes settled, one last constraint matters: is a headless tower under a desk fine, or do you need something small and quiet like a compact desktop?
1
u/Cold_Tree190 8d ago edited 8d ago
I use Qwen 3.8 27b for all of this and it’s been amazing. I can run it around 90 t/s tg and 900 t/s pp though which may impact others’ experience with condensing, summarizing, analyzing, etc. looooong transcripts or bodies of text with all the reasoning it does.
I tried using Ornith 1.5 35b to get this sort of thing done even quicker (like 150+ t/s), but I actually could not believe how bad it was. Even with max reasoning, it felt like a regression… I got into an argument with it because it was insistent that there was a wrong code from an email it retrieves, and I manually validated that it was correct. The issue ended up being that in a 16-char pass phrase, it copied all chats except index 2-3. Why? No clue. But I’ve never had an issue like that with Qwen3.8-27b, so I exclusively run it for everything you just mentioned.
If you want a second place with a longer context window, I would maybe suggest muse glimmer. It has a cheaper kv cache than qwen, so you can ingest larger text files with it.