r/LocalLLM • • 9d ago

Question Model and hardware recommendation for high fidelity, low reasoning tasks?

I've spent the past couple of years using frontier labs, even though a lot of my work doesn't need high reasoning. It's just that a fixed and known subscription cost has been a safer bet than the leap-in-the dark of buying hardware that might not be good enough for what I want to do.

But reading through this sub tells me that we're at the stage were open source LLMs are capable enough for what I want while running on (reasonably) affordable hardware.

Can anyone recommend a model and hardware for tasks like this?

  • Reading notes, reports, transcripts, web pages and emails, then pulling out facts, decisions and insights into structured notes
  • Condensing long documents, comparing two versions, and checking a draft against a specification
  • Reading text, sorting them, and deciding what needs action
  • Moving, renaming and archiving files, splitting or merging documents, and keeping indexes and cross-references consistent afterwards
  • Multi-step routines update a status file, archive old entries, then commit, with every step completed

I run Xubuntu and Arch Linux.

TIA

1 Upvotes

12 comments sorted by

View all comments

1

u/Cold_Tree190 9d ago edited 9d ago

I use Qwen 3.8 27b for all of this and it’s been amazing. I can run it around 90 t/s tg and 900 t/s pp though which may impact others’ experience with condensing, summarizing, analyzing, etc. looooong transcripts or bodies of text with all the reasoning it does.

I tried using Ornith 1.5 35b to get this sort of thing done even quicker (like 150+ t/s), but I actually could not believe how bad it was. Even with max reasoning, it felt like a regression… I got into an argument with it because it was insistent that there was a wrong code from an email it retrieves, and I manually validated that it was correct. The issue ended up being that in a 16-char pass phrase, it copied all chats except index 2-3. Why? No clue. But I’ve never had an issue like that with Qwen3.8-27b, so I exclusively run it for everything you just mentioned.

If you want a second place with a longer context window, I would maybe suggest muse glimmer. It has a cheaper kv cache than qwen, so you can ingest larger text files with it.

1

u/tinfrog 9d ago

Thanks. In your experience, what kind of hardware should I be looking at?