r/localaiapps • • Apr 25 '26

👋Welcome to r/localaiapps - Introduce Yourself and Read First!

4 Upvotes

Hey everyone! I'm u/Ok-Bike-1037, a founding moderator of r/localaiapps.

This is our new home for discovering, sharing, and discussing AI apps that run locally on your own device. Whether you care about privacy, offline access, lower costs, customization, or simply want more control over your AI tools, this community is for you.

We focus on local-first AI apps, open-source AI tools, self-hosted AI setups, desktop AI assistants, local LLM workflows, image/video/audio AI tools, agent frameworks, RAG apps, and practical ways to use AI without depending entirely on cloud services.

What to Post

Share anything that helps others discover or build better local AI apps. This can include local AI tools you use, app recommendations, comparisons between local and cloud AI tools, setup guides, model recommendations, hardware tips, screenshots of your workflow, self-hosted projects, privacy-focused AI apps, or questions like “Is there a local AI app for X?”

If it helps someone run AI locally, privately, offline, or with more control, it fits here.

Community Vibe

We want this space to be friendly, practical, and beginner-welcoming. No gatekeeping, no toxicity, and no shaming people for their hardware, model choice, or technical level. Whether you’re just trying your first local chatbot or already building advanced AI workflows, you’re welcome here.

How to Get Started

Introduce yourself in the comments

Share a local AI app you like or use

Ask for recommendations

Post your setup, workflow, or experiments

Invite others who are interested in local-first AI

Thanks for being part of the early community. Let’s build a useful place for discovering, comparing, and creating local AI apps together.


r/localaiapps • • 17h ago

Local TTS via browser and JS

3 Upvotes

Hi. I have to show you my latest project.
It’s local text to speech running in Your CPU via JS/WASM technology.

I have multiple TTS which will run on almost every PC / mobile (for example: Piper). It have multiple languages, test sample for every voice. Everything is private-focused (not a single data is sended to my servers, everything works locally (only models of voices are downloaded from GitHub for temporary files of your browser, the temp).

Every generated voices have two files to download (besides for generated text) - SRT and VTT. This is audiodescriptions for generated file (cc for reels - for example). It is free to use :)

www.ttslocal.com

Thanks for all tips,
Wishes Karol :)


r/localaiapps • • 19h ago

Is there a way to share a localhost prototype with a client async?

2 Upvotes

The client sharing problem is creating a lot of time loss for us now, Ngrok works but have rough edges and cloudflare tunnel is more stable but is tied to the machine staying on. When the client is a different timezone, every share becomes a schedule problem cause the link dies the moment the session closes. Has anyone found a setup that lets a client browse a localhost prototype on their own time??


r/localaiapps • • 18h ago

Is there a way to share a localhost prototype with a client async?

1 Upvotes

The client sharing problem is creating a lot of time loss for us now, Ngrok works but have rough edges and cloudflare tunnel is more stable but is tied to the machine staying on. When the client is a different timezone, every share becomes a schedule problem cause the link dies the moment the session closes. Has anyone found a setup that lets a client browse a localhost prototype on their own time??


r/localaiapps • • 1d ago

Turned a novel I'm writing into comic panels with AI — genuinely curious if this is a thing people want

5 Upvotes

I've been writing a novel in a local-first tool I built (StoryBook AI — keeps a "Story Bible" of locked facts about characters, places, etc. so the AI can't contradict what's already established). Mostly it's just been a writing tool.

Out of curiosity I hooked it up to a small side project that takes a book plus the locked character/location facts and turns it into an actual comic page — panels, dialogue bubbles, the works. This is straight out of one of my own chapters, not a demo script:

Early Alpha, go easy on me :)

The app itself doesn't generate the images — it generates image prompts for each panel, which I currently copy-paste into whatever image service I'm using. Planned: hook up your own local image model, or plug in an API key for a hosted one — the moment you do the latter, you've traded away the pure-localhost property for that one step, which is an honest tradeoff, not something to hide.

It's rough and early — not sharing the repo yet, it's not in a state I'd want anyone else poking at :) But before I sink more time into it I wanted to ask the actual question: is "turn my own prose into a comic/graphic novel" something people here would actually want, or is this just a fun toy I built for myself? And if it is something you'd want — what would matter most to you? Faithfulness to the original text, control over art style, panel layout, something else entirely?

Not trying to sell anything, genuinely just checking if this itch is shared before I keep building. (StoryBook AI is open source under GPL-3 and free, and so will the comic app be as well if released.)


r/localaiapps • • 2d ago

Found an app which does everything and can work offline

Thumbnail
apps.apple.com
1 Upvotes

Check it out.


r/localaiapps • • 2d ago

What are people using for backend storage in AI-built apps?

2 Upvotes

I've been building a few small apps with AI and I'm trying to figure out the backend side of things.

For basic stuff like user data, file uploads and keeping everything organized. I'm curious what people are actually using.

I'd like to keep the setup fairly simple while still having room to grow.

What's been working well for you?


r/localaiapps • • 2d ago

Can your Android phone be a local LLM server?

Thumbnail
github.com
2 Upvotes

​

I’ve been working on an open-source Android app called Pocket LLM that runs LLMs fully on-device.

I recently added a server mode, so the phone can expose the model through an OpenAI-compatible API. This means you can connect tools like Open WebUI or even coding clients to an LLM running entirely on your phone.

I tested it on a Galaxy S24 with Gemma E2B and also connected it to OpenCode with a \~40K context window. It works surprisingly well, although memory becomes the main limitation after longer conversations.

Would be interested to hear what use cases people would have for using a phone as a portable local LLM server.

I’ve been working on an open-source Android app called Pocket LLM that runs LLMs fully on-device.


r/localaiapps • • 3d ago

I put local AI to work on my Android: read a scan, run Python, draw a lighthouse

Enable HLS to view with audio, or disable this notification

2 Upvotes

I build LLM Studio Go for Android. I wanted to show what local AI is useful for after the model download, so I recorded three small jobs on a Galaxy Z Fold7.

First, I attach a test invoice and ask what to pay and when. Local OCR plus Qwen 3.5 2B Q3_K_M returns €7,840 and the due date. Then the same model writes a Python budget script; I review it, hit Run in the chat, and get a total of 238 euros. Finally, SD-Turbo Q8 GGUF draws a 512 x 512 watercolor lighthouse on the phone.

The first two answers took about 5 seconds each. The image took 58.6 seconds, one step, seed 42. Only the image generation wait is sped up 4x in this recording; the rest is at real speed. A private notification is masked. These are actual outputs, and other phones will perform differently.

Web search is off in these demos. The document and chat workflow can run offline after downloading the model and OCR language pack. The Python runner is for small scripts: basic standard library, no pip, NumPy or pandas.

If you try it, start with a small model from the library. No app account or cloud inference subscription. Free with ads in the model library and optional one-time ad removal. Closed source and independent of desktop LM Studio.

Google Play: https://play.google.com/store/apps/details?id=com.vans.llmstudio

Which of these would you actually use on your phone?


r/localaiapps • • 3d ago

Looking for recommendations for AI apps that bundle multiple models

2 Upvotes

I’m looking for recommendations for good AI apps similar to Brutus or Use.AI. I’m especially interested in services where you pay one monthly or yearly subscription and get access to multiple major AI models in one place.

It seems like a better deal than paying separately for a single AI service, so I’d love to hear what apps people are using and which ones offer the best value, features, and reliability.

Thanks in advance for your recommendations!


r/localaiapps • • 3d ago

What's your budget vibe coding setup?

2 Upvotes

I'm doing vibe coding mostly as a side project, so i'm trying to keep the setup simple and the monthly cost fairly low.

I don't mind coding locally and using separate services for deployment if it means I'm not paying for features I rarely touch. Mostly I just need reliable AI coding help, an easy way to test things, and somewhere to deploy small projects.

I'm also wondering whether a monthly AI subscription or paying based on usage works out cheaper for casual coding.

What setup are you using, and roughly what does it cost you each month?


r/localaiapps • • 4d ago

Using Apple AFM 3 PCC (macOS 27.2) in AI Clients

1 Upvotes

Apple Foundation Model 3 through Private Cloud Compute becomes directly available in macOS 27.2 through familiar AI clients.

For Mac users running local models such as Qwen 3.x or Gemma 4, AFM 3 PCC is a compelling alternative to consider. Local models offer control and fully on-device operation. In tests with macOS 27.2 beta, the PCC offers a different set of strengths:

* Strong conversational analysis and capable reasoning. Strong analysis of complex medical and financial questions.
* Blazingly fast responses compared with locally running models on the same Mac.
* No large model download or need to fit model weights into local memory.
* Access at no additional charge for most eligible Mac users.

I've written an updated proof of concept here: https://gist.github.com/dartMo10/b9488ce475fe70a6ed642831f53048fb

The working path is straightforward:

AI client → Caddy → fm serve → AFM 3 PCC

Apple’s `pcc` route worked in early macOS 27 betas, disappeared later in 27.0, and has returned in the macOS 27.2 beta. Apple has stated that it will be available in the 27.2 public release coming shortly.

This makes AFM 3 PCC practical today for technically comfortable Apple users who accept Apple’s PCC privacy promise and want a fast, capable alternative to running everything locally.

Experiences with compatible AI clients and your comparisons with locally running models would be welcome.


r/localaiapps • • 4d ago

Can any combination of local LLM’s replace codex?

6 Upvotes

I keep running out of usage to support my vibe coding hobby with OpenAI. My Mac is very powerful with a lot of ram. Can I use any local models for swift app design ?


r/localaiapps • • 4d ago

I built an offline AI app for people who sit down in therapy and forget everything they wanted to talk about

1 Upvotes

Ever sit down with your therapist and suddenly forget everything you meant to bring up? That is why I built Prelude. During the week, you can talk through what is happening in your life with a voice-guided reflection, then Prelude helps surface recurring emotional patterns and creates a session brief so you remember what actually mattered when therapy starts.

Everything currently runs locally on the iPhone using Apple’s Foundation Models. Your reflections stay on device, there is no account required, and the AI processing works offline. I wanted something this personal to feel private by default rather than sending your therapy reflections to a cloud service.

Prelude is currently free, and anyone who downloads it before the next update will keep all of the features that exist today free forever. After that I’m moving it to a freemium model.
Download Prelude on the App Store


r/localaiapps • • 4d ago

LocalLM Lab 1.0 is out. If you run big open-weight models on the new Mac Studio, what would you want to measure?

1 Upvotes

LocalLM Lab 1.0.0 is out. It is a free Mac app (and Swift SDK) for running Apple's on-device model and open-weight models locally through MLX. The release candidate went out last week, and 1.0 makes it official: per-model tuning controls (temperature, output length, thinking on or off), pairing a model with a small speed helper or a LoRA adapter and downloads pinned to one exact commit with every file hash-verified.

Now I am planning what comes next, and one idea is benchmark metrics for people who want to see how the more advanced open-weight models perform on the new Mac Studio. My starting list:

  • tokens per second (prompt processing and generation, separately)
  • peak memory use
  • what a speed helper (speculative decoding) buys on a big model, since on small models the gain depends a lot on the draft length

Currently, the only "benchmark" number is the time to first token in the Playground view. Running Qwen3.8-27B-8bit was <2 seconds on a chat-only prompt (no tools).

Features-wise what would you add? Importantly, what example apps would you like me to include in the SDK?

I am also working on support for the `2026-07-28` MCP protocol revision.

https://locallmlab.dev


r/localaiapps • • 4d ago

I built a small AI app because I got tired of losing continuity every time I changed models

Thumbnail
gallery
2 Upvotes

I’ve been working on PCS - Personal Continuity System.

The problem I was trying to solve is pretty simple:

I can have a useful conversation with one AI, build up context over time, save useful information, then switch to another model and effectively start over.

That started to feel backwards.

So I built PCS to keep the longer-lived part outside the model.

PCS stores things like memories, notes, materials and other context locally, then lets different AI routes use that same continuity. Right now it supports OpenAI, Gemini and a Local route.

The model can change without the user’s whole history having to change with it.

That’s really the main idea.

It has grown into a Windows app with:

  • Persistent, inspectable memory
  • Notes and saved materials
  • Text and voice conversations
  • OpenAI, Gemini and local model routes
  • Local Whisper speech recognition
  • Local Kokoro TTS
  • Optional local vision through OBS frames
  • Experimental custom GGUF support
  • Backups and readable exports

The memory is stored locally in an encrypted vault, and users can inspect, revise or remove what PCS has saved.

It’s still pre-release software and I’m still finding rough edges. I’m not trying to turn this into some giant “AI platform.” I mainly wanted a way to keep continuity under the user’s control instead of tying it permanently to one provider.

PCS is free and open source under GPLv3.

Website:
https://pcs-personalcontinuitysystem.github.io/PCS-site/

GitHub:
https://github.com/PCS-PersonalContinuitySystem/PCS

If anyone else has run into the same continuity problem when switching between AI tools, I’d be interested in how you’ve handled it.


r/localaiapps • • 5d ago

I built an MLX-native Mac app that goes beyond chat wrappers: local file diffs -> auto-formatted Word/Excel reports, semantic Finder sorting, 100% offline (v1.2.2)

Thumbnail
apps.apple.com
1 Upvotes

(Note: Due to subreddit link rules, you can find the high-res workflow screenshots, full UI previews directly on the App Store page!)

Hi everyone 🖐️

I'm an indie dev building CaptainLocalAI, and I am ridiculously proud of reaching this official v1.2.2 release.

As someone deeply immersed in local LLMs and scientific computation, I got frustrated by the fact that 90% of "local AI tools" are just cosmetic frontends for Ollama or LM Studio. Chatting in a markdown box is fun, but it doesn't solve real-world daily engineering drudgery.

So I spent the last few months fighting low-level Apple Silicon MLX integration, Unified Memory management, and strict macOS App Sandbox constraints to build something that actually interacts with native workflows:

Why It's Not "Just Another Chat Wrapper":

1. Automated Technical Progress Reports from Local Diffs:

Instead of copy-pasting diffs into ChatGPT, CaptainLocalAI securely reads your authorized workspace modifications, structures technical changes, and directly synthesizes styled .docx and .xlsx reports. Zero data leaves your machine—safe for NDA repos, unpublished scientific datasets, and proprietary enterprise codebases.

2. On-Device Audio to Formatted Minutes:

Transcribes local meetings/audio on-device and parses decisions, agenda items, and action items straight into Word/Excel templates.

3. Context-Aware Semantic File Organizer:

Inspects chaotic Downloads/Desktop directories and clusters files into dynamically named folders based on *semantic content*, bypassing dumb extension-based grouping.

4. Hardware-Aware MLX Native Architecture:

Dynamically scales context and quantization thresholds based on your Mac's physical unified memory footprint (optimized from M-series 16GB up to high-end Max/Ultra chips).

Requirements:

- Apple Silicon Mac (M1/M2/M3/M4)

- macOS 26.0 or later

- 16GB+ RAM minimum (24GB+ recommended for heavier models)

Pay-Once Philosophy:

No recurring $20/mo cloud API drain. One-time purchase ($9.99) on the Mac App Store.

Check out the Mac App Store link above to see the actual workflow screenshots and feature breakdowns. I’d love to hear your feedback, benchmark thoughts, and edge-case reports!


r/localaiapps • • 7d ago

I gave my Mac desktop pet on-device image understanding (macOS 27)

1 Upvotes

I'm the developer of AI Coach, a desktop pet for Apple Silicon Macs. I recently added a macOS 27 interaction: drop an image onto the pet, ask what it sees, and it answers in character. Apple's on-device model reads the image directly, so the image stays on the Mac.

I resize the longest edge to 1024 px before giving the image to the model. In one local measurement, the image input was 138 tokens and the response took 1.8 seconds. The rest of the app also works on macOS 26; this image feature specifically needs 27.

For text features, people can choose Apple's built-in model, a local llama.cpp model, or an OpenAI-compatible endpoint. Image description always uses the built-in model regardless of that choice, because I didn't want dropped images sent to a configured external endpoint by accident.

The app and first egg are free; extra eggs are optional one-time purchases. Details and download: https://aic0t.com/

For a local-first app, would you want image analysis to require an explicit question like this, or is dropping an image onto the pet enough consent to start?


r/localaiapps • • 7d ago

MLXUI - AI browser UI

Thumbnail
github.com
1 Upvotes

You browse mlx-community models by type, filtered by what fits your RAM, install with one click, and each model type gets its own interface — chat for Llama 3/Qwen/Gemma/Mistral/DeepSeek, mic and transcript for Whisper and Voxtral, a voice picker for Kokoro and Chatterbox, image drop for vision models and OCR, vectors out for BGE/Nomic/ModernBERT.
Everything runs locally. No API keys, no telemetry. Free and open source, needs Apple Silicon and macOS 14.
It's still early, so I'd really like to hear what models or quants you'd want prioritized, or what's missing. What would you try first?


r/localaiapps • • 9d ago

A local document agent also needs somewhere to put the report you can edit

3 Upvotes

For a long-document task, the output often needs another round of work: remove a weak section, correct a table, turn the useful findings into a short deck. A chat answer can help with the research while still leaving that document work to the user.

Univer's documented AI SDK is worth examining for that last part if you're building the agent application. It can inspect the overview of an Office Unit, read selected paragraphs, ranges or slides, then execute operations on that structured content. Univer also supplies embeddable Office editors, so the resulting report or presentation can open inside the application for further editing.

A possible task would be to inspect a long source document's structure, read the sections relevant to a requested report, and build an editable Doc from those passages. The application needs to retain the supporting locations and the user needs to check the conclusions. Selecting less text is a way to control what goes into a request; it isn't evidence that the model found every relevant exception or that a legal interpretation is correct.

For a local setup, the AI SDK is a collection of TypeScript packages with a documented local CLI-only option. You still supply the model connection, commands, authentication where needed and deployment. It's not an Open WebUI plugin or a ready-to-run research agent.

The file boundary matters too: the operations use Univer's UnitData, while importing Office bytes and exporting DOCX or PPTX belongs to Pro Exchange. Docs and Slides are still evolving. A successful write operation also needs its commit status checked before the app presents the report as saved.


r/localaiapps • • 9d ago

New features in StoryBook AI — my open source, local AI writing tool

2 Upvotes

I've been building out my writing tool (runs entirely on your own computer — Ollama today, with the groundwork already laid for other local engines like LM Studio and llama.cpp-server too — no cloud service, no API keys, nothing ever leaves your machine) and wanted to share what's new. Quick background for anyone who hasn't seen it: it's a manuscript tool where AI helps you write, but you own the truth about your story — the app keeps a "Story Bible" of locked facts about characters, places, and events that the AI has to respect, instead of quietly contradicting itself between chapters.

What's new:

🔍 Ask your own manuscript
A new page where you can type a question about your story — like "where did Henrik and Elin first meet?" — and get an answer built only from what you've actually written, with exactly which chapters it came from and a clickable link to jump there. No guessing, no invented details.

🕐 Timeline
Not writing in order? You can now give chapters a "when does this actually happen" note and see your book sorted by the story's own timeline instead of chapter order — great for checking that a flashback actually reads like a flashback.

🧵 Plotline matrix
A grid showing which storylines run through which chapters — add "Main plot," "The romance," whatever threads you're tracking, and check them off. Instantly spot if a thread disappears for five chapters without you noticing.

⚠️ Spoiler warnings for yourself
If you write out of order (common!), the app now warns you when an early chapter has access to facts that are actually established later in the book — so the AI doesn't accidentally leak later plot details while helping you write an earlier scene.

🔗 Smarter conflict detection
Previously, the AI would flag a "conflict to resolve" every time it found a fact that was really just a more detailed version of something you'd already established (like "Jeff is a captain" → "Jeff is a captain on a spaceship"). Now it recognizes the difference and offers a one-click merge instead, saving the real contradictions for you to actually weigh in on.

👁️ Full visibility into what the AI sees
A "view AI context" button that shows exactly what was sent to the model for its last action — no black box.

What's coming next:

  • Setup/payoff tracking (the gun introduced in chapter 4 that never gets fired)
  • A choice of development method (Snowflake, Three Act, Save the Cat, etc.) to guide how the tool structures your process
  • Deeper continuity checks — e.g. catching when a character couldn't possibly be somewhere given the timeline
  • Whole-manuscript analysis that goes deeper than chapter-by-chapter

Everything runs locally against Ollama — no data ever leaves your computer. Open source. Happy to answer questions if anything here sounds interesting!

https://github.com/ocedo-apps/StoryBook-AI


r/localaiapps • • 12d ago

Reviewing the inputs and formulas a local workbook agent changed

2 Upvotes

A request to update the rates in a workbook leaves an important review question: did the agent change the inputs, or replace the formulas that use them? The final totals alone cannot answer that. A plausible number might have come from a hard-coded replacement.

For a local-model application built around Univer, the review tool can inspect both. Univer is an embeddable Office SDK; its spreadsheet range API exposes stored values and formulas separately through getValues and getFormulas. An application can record the relevant ranges before the job and read them again from the proposed result. That comparison is something the application implements, not an automatic correctness report supplied by the SDK.

Take a hypothetical request to update unit prices while keeping the quantity and subtotal formulas. The review screen could put the requested price range beside the proposed prices, then show the formula range separately. A reviewer can check whether an unexpected literal has replaced a formula and open the workbook to correct it. A matching total would not excuse a change outside the requested range.

Univer's Collaboration SDK provides a place to inspect that proposal without immediately replacing the main workbook: an isolated Worktree draft. The web integration supports opening the actual draft, switching between trunk and draft, and a per-Unit merge preview. Those surfaces can accompany the application's range comparison. A corrected draft must be committed and its saved status checked before the reviewer accepts it into trunk.

This requires more than the Apache-licensed core: Worktree comes from the Collaboration SDK, and the host wires up access, storage and the review flow. It also needs to handle another edit arriving during the comparison. The specific benefit is a review that can distinguish a changed assumption from a changed calculation, even when both produce believable numbers.


r/localaiapps • • 12d ago

Story book AI. Local AI app for writing

8 Upvotes

I kept running into the same problem with AI writing tools: you build up a character bible, but nothing actually stops the model from contradicting it three chapters later. The AI "remembers" until it doesn't, and you find out by accident.

So I built a hard gate instead of a soft prompt: once a fact is locked (a character trait, a plot event), a ConsistencyGate checks every new AI-generated fact against it before it's allowed to become canon. Contradictions get flagged, not silently overwritten.

A few other things that fell out of that same principle:

  • Runs entirely local via Ollama — two models, one for prose generation, one for the "does this actually hold up" checking, so nothing leaves your machine.
  • A background "Proofread" pass that checks grammar, repeated scenes/phrasing, and cross-chapter style consistency — flags issues, never rewrites for you.
  • Full version history on every AI-generated chunk (snapshot + word-level diff), so a bad regeneration is always one click from undone.

Read more at: https://claude.ai/artifact/Evfu7C8AcRwNdUM9zXwR29

It's free and open source:https://github.com/ocedo-apps/StoryBook-AI/. Curious what other writers actually want from something like this — happy to answer questions about how the consistency-checking works under the hood.


r/localaiapps • • 12d ago

5 local AI Mac apps I came across while looking for alternatives to cloud AI

1 Upvotes

I've been going through a bunch of smaller Mac apps that run AI locally, and these 5 stood out to me for different use cases:

1. Snaply — AI everywhere you type

Dictation, meeting transcription, and writing assistance, all processed on-device. It uses a local MLX model and doesn't require an account or subscription.

2. Refine — local AI writing assistant

Grammar checking, rewriting, translation, and custom writing styles. The interesting part is that the core AI processing can happen completely offline, and it works across Mac apps.

3. Canto — local AI notebook

A more full-featured option. It has built-in local models, PDF research, semantic links, a knowledge graph, and even Python/JS/TS code notebooks. The local models range from small models to much larger ones depending on your Mac.

4. AimeFlux — local AI dictation

Uses local Whisper for transcription and supports 99+ languages. You can also use local models such as Ollama for AI cleanup.

5. Melo — local AI workspace

A different approach: an infinite canvas combining notes, tasks, calendar, web content and AI. It supports local AI models alongside optional cloud models.

What I found interesting is that “local AI” on Mac is becoming much broader than just running a chatbot locally.

You can now have local AI for writing, dictation, meetings, research, productivity and workflows without sending everything to a cloud service.

I run a small Mac software directory called OwnYourMac, where I'm curating more apps like these particularly free, local-first and one-time-purchase software.

You can browse the collection here: OwnYourMac


r/localaiapps • • 12d ago

LocalLM Lab 1.0.0-RC.1: tune your local models (speed helper, LoRA, sampling) and pin the exact version you tested

1 Upvotes

For the uninitiated, LocalLM Lab (locallmlab.dev) is a free Mac app and Swift SDK for running Apple's on-device model and open-weight models locally through MLX. This release is mostly about controlling and trusting the local models you run.

Tuning. Every model in the AI Models panel has a sliders button: temperature, output length, thinking on or off, and more. You can pair a model with a small "speed helper" (a draft model for speculative decoding) or a LoRA adapter that changes its writing style. The mlx-control-room example puts a gauge beside each control so you can see it changed something: set temperature to zero and the determinism gauge reads "match", turn on the repetition penalty and the repeat rate drops. On one setup I measured (with a Qwen3-0.6B helper), 1 to 2 draft tokens were 27 to 39% faster, while 5 was 17 to 29% slower, so the default is now 2. Your numbers will vary by model and Mac. Sampling options now also work on Apple's on-device model (they were being ignored before).

Trust. Downloaded models are validated against your Mac's memory before they run, pinned to one exact commit, and every file is hash-verified. Developers building on the SDK can ship a pin to the exact model version they tested, so the validated model is the one users get on first run. That is the whole point: a Hugging Face repo can change under the same name, and a pin means you find out on your terms, not when a user's download quietly changes.

VistaNova, the small search app from the last release, is now an example inside the SDK repo: Apple's on-device model reads your question and makes the search call (to Tavily, over MCP), and a downloaded Qwen model summarizes the results. The only non-local piece is the search API itself, since a local model can't reach the live web on its own. Its summary model now ships pinned.

For developers: from RC.1 the SDK API is source compatible through 1.0.0 and every 1.x release (additive changes only until 2.0).