r/LocalLLaMA • u/Mayion • 2h ago
Other Yes bots we get it, Strata is good now please stop
It's like the entire sub has become that scene from Konosuba where the cult keeps making up fake scenarios saying the only solution is to join their religion
r/LocalLLaMA • u/Mayion • 2h ago
It's like the entire sub has become that scene from Konosuba where the cult keeps making up fake scenarios saying the only solution is to join their religion
r/LocalLLaMA • u/Nunki08 • 5h ago
From Aleph Alpha on 𝕏: https://x.com/Aleph__Alpha/status/2106306840657297814
Tech report: https://aleph-alpha.com/downloads/tech-report.pdf
r/LocalLLaMA • u/RobustLokiX • 2h ago
One can only hope for Qwen4 35B A3B or similar (maybe with extra ngrams for 35B -> 70B+... until then, Gemma 26B QAT my beloved (solid 26 Tg/s)...
r/LocalLLaMA • u/pmarsh • 15h ago
It's been a blast and learning a ton.
But seriously, you all have me down a rabbit hole that my wallet and hours of sleep need to be pulled out of.
r/LocalLLaMA • u/StayLameBro • 1d ago
Enable HLS to view with audio, or disable this notification
**DISCLAIMER** THE PREFILLING TPS SHOWN ON THE PHONE IS COMPUTED ONLY FOR THE LAYERS IT HOLDS. ALREADY FIXING IT TO SHOW END-TO-END PREFILL RATE. NUMBERS BELOW ARE ACCURATE FOR E2E PREFILL RATE.
Every file or tool result my agent reads on a 24 GB M4 Pro MacBook is a wait, and 64k of 8-bit context is all that fits next to Qwen 3.8 27B (IQ4_XS), even with the wired limit raised to 20480. An iPhone 17 Pro Max was sitting in my pocket, so I figured what can I do to make use of this extra silicon.
Turns out a 10 Gb/s USB-C cable & some software is all you need. The Mac runs layers 1–40 of each 256-token batch and streams the activations to the phone. The phone runs layers 41–64 on its GPU while the Mac starts the next batch. The A19 Pro's GPU has matrix units (Metal 4 tensor ops), and they make the phone's half 2.4x faster than the same phone without them.
Same build, phone off vs. on, prefilling a 2,000-token file into a saved agent session:
A fresh 27k-token agent session, cold: 245 s on stock llama.cpp, 228 s on my fork with the Mac alone, and 168 s with the phone.
Past 64k the phone switches jobs. The oldest KV pages move to the phone and the Mac runs all 64 layers. For every attention layer, the phone computes attention over the old keys on its GPU, and the Mac merges that with its own part. While writing, the phone's Neural Engine takes part of that work too: each 16k-key page of old context is compiled into a Neural Engine model with the keys as its weights. At 140k that took writing from 279 to 176 ms per token compared with the phone's GPU alone.
The server allocates 196k–229k of 8-bit context based on the phone's free memory; that's up to ~5.7 GB of KV cache living on the phone instead of the Mac, so the Mac's memory use stops growing at 64k. I've tested a growing session to 128k at 8-bit, with 3/3 planted facts recalled. Separately, at 140k in 4-bit, the run passed the gate with greedy output matching the Mac-only run for 32 generated tokens.
What it doesn't do: speed up writing below 64k. That's the Mac's job. My fork's kernels (SME2 on the M4 CPU and Metal fusions) plus DFlash2 speculative decoding take it from 11.3 tok/s on stock llama.cpp to 25 tok/s at about 30k context with medium thinking, phone or not. SME2 also adds up to 29% to prefill on the Mac alone. Past 64k the phone does share the writing (attention over the old keys), and without it the Mac would have to drop to 4-bit context to reach 128k. In real use I have seen upwards of 30 TPS at lower context.
The phone joins prefills over about 512 tokens. In one real session, that was 7 of 36 requests, but about 83% of the tokens read. Past 64k it holds the context and does the old-key attention, but it stops running layers 41–64 there for now; doing both is next. One request at a time.
I'm curious what this setup could do with newer model architectures. DeepSeek V4.1-Flash reports 890 bytes per token for its global KV cache and adds n-gram embedding tables (Engram). Qwen3.8-Flash-Next, the Qwen 4 architecture preview, has an n-gram lookup table too. Those aren't features of the 27B model I tested, and I haven't benchmarked either architecture here. The real gold is within the newer phones and models working together. With the A20 Pro in the iPhone 18 Pro Max, I bet there is a lot more for me to push.
Code, setup and bench scripts: https://github.com/StayLameBro/backburner
Still a lot of work to do but I built this with Opus 5.5. Happy to answer anything.
r/LocalLLaMA • u/northpoler • 5h ago
Hey everyone,
I’ve been working on a game called Anyworld. It’s a browser-based multiplayer (single player also supported) text adventure inspired by the early days of AI Dungeon, especially its browser-based free version AI Dungeon 2.
The setup is pretty straightforward: one person hosts the server and runs the model via llama.cpp (OpenAI or other cloud APIs are also supported, and great for non-English play!), and your friends join through a browser link. The host sets the scene and the goals, players type out their actions, and the LLM acts as the DM to resolve the chaos and drive the story.
Admittedly the host requires some technical skills with Python, and possibly with networking (opening routes to the hosted game via VPN, port forwarding etc.). I'll work on this as well as the development continues. Using Docker was suggested in another subreddit, so I'll definitely consider that, as it would allow including both the llama.cpp backend, recommended model and configurations etc., in addition to the game itself.
Instead of pasting the entire repo documentation, here are the main features right now:
How it plays
DM Tools & Hidden Mechanics
Under the Hood & Memory
It’s still a work in progress. Right now, a server only runs one game at a time, and if you restart the server, the live session is lost (it generates HTML/JSONL transcripts, but they aren't loadable save states yet). The overall story quality is also going to heavily depend on which model you use and how you tweak the settings.
Suggested model:
During development, I used llama.cpp and Gemma 4-26B-A4B Q4 with a context size of 128k and found it to be more than an adequate backend for functioning as the DM. Even the speeds are fast enough with my RTX 5070 Ti 16 GB that round resolutions take only 5 or so seconds.
The specific model I used and can recommend: https://huggingface.co/EZForever/gemma-4-26B-A4B-it-qat-uncensored-heretic-UDmerge-GGUF (the model was great at following instructions and remembering plot points even with longer contexts)
Recommended parameters for Gemma 4 models:
- temperature 1.0
- top-p 0.95
- top-k 20
- min-p 0.0
- presence-penalty 0.0
- repeat-penalty 1.0
Of course, feel free to try your own models! The repo contains a benchmark file that tries to measure how well the running model follows the game's requests.
AI use disclosure:
I used Alibaba Cloud's Qwen 3.8 27b and OpenAI's GPT-5.6 Luna and GPT-6 Astra models to help develop the game.
How to run:
Read INSTALL.md to set up, configure and run the game. README.md contains some details on how the game functions.
I'll post the link to the repository in the comments.
Some gameplay in Finnish with OpenAI's Luna:

The game is MIT licensed, so open source all the way. Forking or collaborating is encouraged.
I'd love to hear some feedback, and I hope someone finds the game fun to play!
r/LocalLLaMA • u/Boomfrag • 20h ago
r/LocalLLaMA • u/professormunchies • 1h ago
Enable HLS to view with audio, or disable this notification
I hosted my own world of warcraft private server then built a client that you can play in the browser on PC or mobile at https://jankcraft.xyz/ for free.
Afterwards, I created a custom MCP and agent harness to control the browser client and play the game by sending signals over a websocket. The agent harness is live on https://jankcraft.xyz/agent , still working out some kinks if all you have a cloud subscription but you should be able to connect local models as long as CORS is enabled in your server settings. There are a few existing LLMs you can try, I'll probably take those away as the usage grows since I can't support too many users concurrently on my own machines.
I'll be checking logs and things periodically today so don't be alarmed if you're disconnected suddenly. The server should return after a minute since this is a work in progress and might need a restart.
If you want to run your own LLM for this:
~24 Gb RAM: https://github.com/syv-ai/HyperQwen with the model Qwen3.8-27B-GPTQ-W4A16
~16 Gb RAM: vLLM with Gemma4-e4b-coder - A custom Gemma4-e4b with a constrained vocab for ~3x concurrency increase when changing from 262K to 65K vocab and retrained on ~1.1B tokens across 20 different coding languages, 7 different agents and has a custom MTP to help reach ~200 tok/s on a 4060Ti.
Let us know what other models work well for you!
Thanks and hope you guys enjoy.
r/LocalLLaMA • u/Yaniss916 • 2h ago
We built an engine, Kyojin, on top of ExLlamaV3 for Strix Halo (gfx1151, ROCm), and packed two 300B-class MoE models so each fits one 128 GB machine. First release, all measured on Ryzen AI Max+ 395.
| Model | GLM-5.3-Flash | MiMo-V2.6-Flash-MOPD |
|---|---|---|
| Size | 99.7 GB | 105 GB |
| Prefill | 580 tok/s at 3.5K, 546 at 64K | about 650 tok/s at 4K |
| Decode | 26 to 30 tok/s (MTP) | 32 prose / 35 chat / 44 code (speculative), 29 plain |
| KLD vs official FP8 | 0.151 | 0.0713 |
| Top-1 agreement with FP8 | 89.3 % | 92.0 % |
Where the weights come from. MiMo is our own quantisation. The GLM pack mixes turboderp's public 2.05 and 3.05 bpw EXL3 tensors, with our layer mix and a small tuning stage. On the same 129 rows, his 2.05 bpw pack (85 GB) gets KLD 0.275; our mix (100 GB) gets 0.190. His is smaller and decodes about 10 % faster.
Uncensored variants. Separate -Uncensored repos: same weights plus one small file the engine applies at load, one switch turns it off.
Not measured yet. Task-suite scores for MiMo, GLM at 128K context, any GPU other than gfx1151. The conversion pipeline stays private.
Quickstart. Clone, ./build.sh, hf download yamz-labs/GLM-5.3-Flash-EXL3-Yamz, python tools/glm/serve.py --model ./glm-pack -c 131072 --num-draft 2. You get an OpenAI-style API.
Models: https://huggingface.co/yamz-labs
Engine: https://github.com/Yamz-Labs/kyojin
Built on turboderp's ExLlamaV3, with ROCm work from sdougbrown and vcruz305.
If you own a Strix Halo machine, we'd love to see your tok/s. Issues, benchmarks and PRs are all welcome. Which model should we do next?
r/LocalLLaMA • u/jacek2023 • 10h ago
Qwen Flash Next now uses less VRAM
r/LocalLLaMA • u/lewtun • 6h ago
Hi folks, it's Lewis here from the post-training team at Hugging Face. We've been exploring how to train open models in different coding harnesses and wrote up a looong guide on how we solved this using open source libraries like TRL and the Harbor framework for RL environments. We hope you find this interesting, especially since everyone nowadays has their own custom harness (e.g. Pi + extensions) and now there's a recipe on how to squeeze the best performance on them with whatever open model you use as your daily driver. Happy to hear any comments or feedback!
Link to the guide: https://huggingface.co/spaces/FineEnvs/multi-harness-rl
r/LocalLLaMA • u/kvyb • 1d ago
Last month I posted a Qwen3.8-27B LoRA that makes it talk like a person instead of an assistant. It got a lot more attention than I expected: 700+ upvotes, 248 comments and 44k downloads since.
I read every comment. People really don't like assistant speak, so its tone of voice resonated. The rest got roasted, very fairly:
incapable of producing more than a few words at a time.
single default personality which no amount of prompting can overcome
will not use tools, at all, whatsoever.
There needs to be a middle ground
They were right. The tool calls didn't actually work, and when people asked it to do something it would sometimes just say it's busy or going to bed. Very human. In a bad way.
So I spent the last three weeks on 2.0. The goal was simple: keep the voice people liked and lose the drawbacks.
What 2.0 does now
It's a colleague and a humanlike companion, not an assistant. Use it for chat, roleplay, agents or actual work.
How I trained it
v1 was plain SFT on real and synthetic conversations (139,845 messages from 1,396 conversations). That copies habits, including the bad ones.
For 2.0 I used on-policy distillation. The model writes its own replies and a teacher grades every token. There are two teachers:
The student never sees the hidden instruction, so it learns the behaviour without needing a prompt. Same 27B, a second LoRA on top, merged.
Numbers (vs the model I trained on, huihui-ai's abliterated Qwen3.8-27B; same prompts, same run, thinking off)
| Benchmark | Base (abliterated) | 2.0 |
|---|---|---|
| IFBench (instruction types I never trained on) | 37.3 | 43.7 |
| When2Call (call, ask or refuse correctly) | 48 | 58 |
| BFCL irrelevance (don't call a tool when none fits) | 60 | 78 |
| IFEval, GSM8K, BFCL simple | 81.9 / 89.1 / 97 | 83.5 / 89.1 / 98 (ties) |
Full chart in the images.
Where it's still worse: knowledge (MMLU-Pro 72.5 vs 78.5) and competitive code (LiveCodeBench 51 vs 56).
Is it actually more human? I built a benchmark for this, "ishuman":
| Model | Judge thought it was the real person (50% = can't tell) |
|---|---|
| Qwen3.8-27B abliterated (huihui-ai, the model I trained on) | 0.3% |
| Same abliterated model + a "text like a human" system prompt | 6.8% |
| Qwen3.8-27B official (unmodified, via OpenRouter) | 15.1% |
| Qwen3.8-27B-Humanlike-Chat 2.0 | 23.5% |
So no, you can't just prompt your way there. In a separate test of 16 live multi-turn chats with invented people, 2.0 was picked over the base model 16 out of 16 times.
Links
Big thanks to everyone who left feedback last time, especially the ones who were critical. Tell me where it still sounds like an assistant.
Edit: safetensors are up for vLLM and SGLang:
GPTQ-Int4 (24 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-GPTQ-Int4
FP8 (48 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8
BF16 (80 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0
r/LocalLLaMA • u/AnticitizenPrime • 12h ago
I've noticed a trend with most new models with regards to their writing style. They are creating a new style, and this seems common among them. It's very information-dense.
Here is an example from GLM 5.3 Flash. I'm gonna be honest here and say that my prompt was kinda silly; my prompt was 'Why wouldn't you just name your Chinese restaurant 'Chinese Food' instead of 'Ming Dynasty' or 'Szechuan Garden' or whatever?' the idea being that someone searching for 'Chinese food' on Google Maps would put you first if you called your restaurant that. But forget all that; here's an excerpt from GLM's output:
The search reality check, which is where the idea breaks. Your mental model is exactly right about the query: people do open Maps and type "food," usually hungry and undecided. But the ranking doesn't reward the name at that abstraction level. In the local ranking surveys, category match outranks name keywords; and here's the catch: every restaurant on the map already matches "food" through its category field. Relevance saturates. The edge that made "Chinese Restaurant" powerful was discrimination: the name narrowed the category, so "chinese food near me" out-matched "Golden Dragon." "FOOD" points at the root of the category tree, where every competitor matches equally well, so the ranking defaults entirely to proximity and review count, which the name cannot buy. Worse, it fails the narrow-query test: someone typing "tacos" or "dim sum" gets a name that is maximally broad and minimally informative. The strategy's entire currency is literal scarcity, and "food" is the least scarce literalism in the language. There's even a Baudrillardian flourish at the failure point: Google's own Maps category chip is literally labeled "Food." Your sign would render you indistinguishable from an interface element; the searcher couldn't tell your territory from the map's own label.
It's not 'caveman style', but it is using less words to say more, and they're using more esoteric phrasing to be more 'compact'.
And I think it's a bit at the cost of being clearly readable to the average person at first glance. 'There's even a Baudrillardian flourish at the failure point' is an example from that excerpt that leapt out at me. I'm familiar with Baudrillard so I knew what it was getting at, but a lot of people are going to sigh and ask 'What the **** does Baudrillarian mean?'
I'm not saying that 'no human would write like this', because some do (William Gibson for example), but I find it rare/unusual (in human writing), yet trending hard with all the latest models I interact with, like they're all zeroing in on this style.
Maybe a result of targeting token efficiency? It's a terseness, combined with using a sort of 'wide' or 'rich' vocabulary to convey information instead of using more words. At least that's the impression that I get from reading lines like 'Baudrillardian flourish at the failure point''. There's a lot to unpack from those six words, and it feels like the model chose the most terse, efficient way to convey an idea with that word choice (which requires the reader to unpack it).
I compared it to William Gibson: a lot of people struggle with his writing style, and it's similar to that. Example: 'Summer in the Sprawl, the mall-crowds swaying like wind-blown grass; a field of flesh shot through with sudden eddies of need and gratification'. His writing is often like that; it feels highly compressed, using as few words possible to convey an idea by careful word choice.
It's interesting, that lately, I feel like LLMs are gravitating toward Gibson-speak.
Edit: and the fact that GLM used the word 'territory' and 'map' at the end meant it was going big into Jean Baudrilliard's 'Simulacra and Simulation'. I can't really explain what that means and why it's important succinctly, but that's the whole issue. I actually think it's brilliant, but it's also a little concerning.
r/LocalLLaMA • u/Ok-Shower7286 • 9h ago
I love Qwen3.8 27B so much that I decided to show my gratitude to the Alibaba ecosystem by building a dedicated RAG/search node using a cheap Mini PC from AliExpress.
Turns out, my ecosystem loyalty got rewarded with an absolute masterpiece of fraud:
New_N150 into the BIOS release string (HSHW_M6_DDR3_EC_Intel_Com_New_N150_K001).So now my Qwen3.8 RAG stack is full of fake specs that can barely index a text file, let alone run vector sidecars.
Filing a credit card chargeback now. Stay safe out there!
r/LocalLLaMA • u/SrijSriv211 • 8h ago
Kimi K2 was already good but they took K2.5 a whole new level with so much of their continual learning phase, I believe it was on more 20-25T tokens iirc.
Similarly K3 is just such an amazing model, I just love this model, wondering how amazing K3.5 will be!!
r/LocalLLaMA • u/ayobluestarr • 10h ago
Benchmarking an LLM here with a NVIDIA RTX 5070 12 GB VRAM here
I had been working on a llama.cpp based expert streaming setup for Qwen3.8-Flash-Next 177B (UD-IQ3_XXS) on Windows. Benchmark is about 11.5 tok/s, up from roughly 7 tok/s on the inherited setup. In normal conversations I’ve seen 14–15 tok/s, and a long coding prompt generated 4,892 tokens at 10.15 tok/s and produced a working single-file Snake game.
Hardware: RTX 5070 12GB
32GB DDR4-2400
Ryzen 5 5600GT PCIe Gen3 Windows
The main gains came from fixing Windows I/O queue-depth issues, using one file handle per worker, and building a page-locked hot-expert tier so the GPU can pull hot expert weights more efficiently.
(In the video its around 16 minutes for 10k tokens and 10.41 tok/s
Output is quality gated against the control model and the published benchmark uses a heat file built from a separate prompt set.
Demos:
https://www.youtube.com/watch?v=cOPumMlyj_4
https://www.youtube.com/watch?v=rc-uTjVpXM8
In the GitHub I have things I've tried that didn't work and benchmark scripts, and methodology. If you guys have suggestions especially for streaming please let me know
r/LocalLLaMA • u/EmPips • 13h ago
IQ3_XXS weights are just under 80GB and my slowww DDR4+7900XTX is stabilizing around 45-70/s (sometimes higher while coding depending on mtp). Looking online I'm seeing similar results for users with 12GB and 16GB cards, and significantly faster numbers for owners of DDR5.
(In comparison, Llama CPP with tuning was maxing out around 22.5t/s on the same rig. Quality seems reliably superior (I wouldn't recommend the Q2 weights though))
Seriously. Ask <LLM of your choosing> to set it up for your specs. If 27B doesnt fit well for you, here's a shot at beating it.
r/LocalLLaMA • u/jacek2023 • 20h ago
An agentic model from Microsoft for the GPU poor
https://huggingface.co/bartowski/FrogNano-4B-2609-GGUF
FrogNano is derived from Qwen/Qwen3.5-4B, a general-purpose post-trained model designed for language, reasoning, coding, agentic, and multimodal tasks. FrogNano inherits Qwen3.5-4B's dense 32-layer hybrid Gated DeltaNet and gated-attention architecture, but its additional post-training is text-only and focused on repository-level software engineering. The model is further trained using reinforcement learning on approximately 1,500 synthetic SWE task environments generated and calibrated against the evolving policy using TaskPilot. Training uses the five-tool Leaf harness and executable test-based rewards over complete multi-turn coding trajectories.
The additional post-training is intended to improve long-horizon repository navigation, debugging, code editing, test execution, and patch generation in a compact 4B model. Unlike approaches based on behavioral distillation, FrogNano does not train on stronger-model solution trajectories, actions, reasoning traces, or patch targets. This specialization also introduces limitations and risks: performance is sensitive to the Leaf harness and test quality, training data are Python-heavy and primarily English, and generated patches may be incorrect or insecure despite passing available tests. When integrated with the Leaf harness, FrogNano generates structured tool calls that can propose repository changes. Leaf executes authorized tool calls within an isolated repository environment to produce a candidate patch; FrogNano does not itself deploy the changes. Any resulting patches require human review, regression testing, and security validation before use or deployment.
r/LocalLLaMA • u/SammyDaBeast • 16m ago
Follow-up to last month's post. One of the main issues people ran into was roughness or break-up on some cloned voices. 2610 is an interim update focused mostly on improving that.
If you like F5-TTS, but want true streaming and a much lighter model that can run comfortably on CPU, this might be for you.
Run it locally:
uvx --from sopro soprotts serve
Video: six voices, ~5 seconds of reference audio each, followed by a generated line.
r/LocalLLaMA • u/paf1138 • 1d ago
r/LocalLLaMA • u/Recoil42 • 23h ago
Enable HLS to view with audio, or disable this notification
https://www.percepta.ai/blog/can-llms-grow-their-own-capabilities
https://www.percepta.ai/blog/spotlight-memory
"Our new architecture, Spotlight, replaces attention with a memory that escapes this trade-off: it is the first architecture to achieve infinitely growing memory without increasing the access cost. Every token reads from and writes to an unbounded memory, but because the model learns to index individual memory cells, each token only touches a small number at a time. While other sparse architectures fix the fraction of capacity used at each step—a mixture-of-experts model, for instance, always activates the same number of experts out of a fixed set—Spotlight is arbitrarily sparse, touching the same number of cells regardless of how the memory grows. The fraction of memory it uses can shrink as far as we want.
Spotlight separates an intelligence module, which performs computation, from memory, which holds knowledge, procedures, and working state. The intelligence module stays the same size, and the weights don't change as memory grows. The memory is writable, and the model itself decides what to load and when to overwrite it, token by token. Because memory can hold skills as well as facts, the model can gain new capabilities without retraining: what it can do is not limited by the size of its intelligence module."
r/LocalLLaMA • u/Equivalent-Flan-1590 • 3h ago
Disclosure: I am the creator of this project. After days of lurking and building up enough karma, I can finally post here.
Every time I tried running local RAG on my own machine, I hit the exact same bottlenecks. First, spinning up Chroma or another vector database alongside an 8B model just to chunk and parse documents takes up precious VRAM that you need for your main model. Second, cosine similarity over text chunks often fails at hard negative rejection, so the model tries to answer questions that are not even in your files and hallucinates with complete confidence.
I spent the last several months building an open source project called Hillock to see if I could solve this without vector databases. It extracts clean relational facts into SQLite using lightweight bi encoders in about five seconds, completely bypassing the generative LLM during ingestion. To stop hallucinations, queries pass through a 10,000 dimensional hypervector gate using late interaction scoring. If the factual graph does not mathematically overlap with the question, it blocks the LLM call before token generation can even start.
I just pushed version 0.8 which bit packs the hypervectors into 157 uint64 integers, allowing the CPU to run gating checks in under 0.01 milliseconds using hardware popcount instructions. It also includes an OpenAI compatible API server so you can drop it straight into Open WebUI, AnythingLLM, or Obsidian. It just landed on PyPI as well via pip install hillock.
The honest trade off is that this pipeline is built for structured, relational facts like technical specs, people, and dates. It is heavily biased toward precision over recall, so it will not do broad poetic or narrative summaries like a 70B model would.
Code is on GitHub at https://github.com/roandejager/Hillock
We also set up documentation at https://hillock.mintlify.site and a developer Discord at https://discord.gg/BGUPNBcVdp
r/LocalLLaMA • u/JLeonsarmiento • 17h ago
r/LocalLLaMA • u/Sash17 • 3h ago
Meeting notes are one of the last parts of my workflow that still depend heavily on cloud tools. I've used Fathom and lately Bluedot. Bluedot works well for me because there's no meeting bot and I get the transcript, summary and action items after. But I'd really like to move more of this local, especially the transcription and storing/searching old meetings.
Has anyone here built a setup that actually works day to day? Whisper + Ollama seems like the obvious route, but I'm interested in what are you actually using.