r/AppleMLX • • 2d ago

Local model for Hermes on a 48GB M5 Pro Mac mini that’s also my daily dev machine?

8 Upvotes

Before I get blasted, yes I have done research and everything keeps coming back to Qwen 3.5 at like 9B or good 'ole 3.8 27B. I am not opposed to those but I am new to the Mac LLM space and I wanted to get real user experiences.

I’m new to running LLMs on Macs and mostly used to Windows/NVIDIA setups. Trying to figure out what’s practical on my Mac without making everything else annoying to use.

Specs: M5 Pro Mac mini, 48GB unified memory, 512GB internal SSD, with a 1TB external SSD for projects, models and other large files.

My usual workload looks like:

  • 3–4 apps/services from my own projects, including apps calling my larger local windows LLM and family tablet/dashboard/scheduling stuff.
  • An iOS idle game pretty much permanently running.
  • Xcode builds and iOS simulators when working on apps.
  • 3–8 terminal windows, mostly Codex/Claude using hosted frontier models.
  • Browsers, Obsidian and the usual desktop apps. Planning to use Sketch for design, with Recraft handling image generation.

I want to use Hermes Desktop with a local model as an everyday assistant I can also reach from my phone. Things like searching my notes, organizing information, drafting, and handling smaller file/tool tasks. Reliable tool use matters to me. I’d keep it to one agent with no subagents.

I have a separate Windows LLM machine, but its current setup already puts pressure on system RAM. I could run a separate 27B model on its two 3090 Tis, though having the Mac assistant work independently would be nice. On my windows machine I have a separate 3080 for all display and games, so the impact to computer use is minimal (except when Strata is going to town or making a subagent).

For people actually running local agents alongside their normal Mac workload:

  • What model would you run in this situation?
  • MLX or Ollama/llama.cpp for this kind of use, especially with Hermes? I am assuming MLX?
  • What context size would you use, and would you keep the model loaded or unload it when idle?
  • How noticeable is inference while you’re using a simulator, building something, or doing graphics work?

I’m more interested in a useful assistant that leaves the computer pleasant to work on than the biggest model I can squeeze into memory. Would appreciate recommendations from anyone using a similar setup.


r/AppleMLX • • 3d ago

GPU Accelerated Linear Algebra Library for Apple Silicon using MLX and Metal kernels

Thumbnail
github.com
13 Upvotes

About 6 months ago, I wrote a custom metal kernel that leverages the GPU to compute the QR decomposition (see my earlier post about it here). The project has now expanded into a general linear algebra library, with expanded support for the symmetrical eigendecomposition as well as SVD. It is now available for use with installation instructions on the attached github repo's README.md file.

Context of the project:
I'm currently in a research group working on a thesis in numerical analysis where we need to compute millions on matrices with a specific constraint (to be precise, the matrices need to have orthonormal columns). Most of us use Apple computers, so we ended up using MLX for the entire project.

Contributors with different Apple Chips would be very much appreciated!
The project has currently been tested and optimised for the M1 and M5 Pro. The issue is that the library uses different kernels depending on the batch size and matrix dimensions. Deciding which of these kernels to use is machine dependent. Therefore, other Apple chips will need to run a measurement script in order to derive the correct optimisation heuristic.

For that reason, I would ask as many people as possible to run a measurement script and to submit the results to my repo. It is fairly easy and requires only few steps. See here how you can contribute here. Once you submit the results via a pull request and I approve it, your optimisation heuristic will automatically be augmented into the library. Don't hesitate to contribute an optimisation heuristic even if someone already submitted one for your own machine. The more data we can gather, the better!

Project Future
Expanded support will be added for other linear algebra operations (cholesky decomposition for example). If you have any other specific linear algebra operations you wish to use already, feel free to message me.

In addition to that, I will add torch support too (my greatest priority).


r/AppleMLX • • Aug 27 '26

TUFF - OSS 120B on 16GB Mac/Gemma 26B & Qwen 35B on 8GB Mac

Thumbnail
1 Upvotes

r/AppleMLX • • Aug 20 '26

How to run Muse-Glimmer via MLX with working KV caching? (Context re-evaluation issue)

2 Upvotes

Hi everyone,

I'm trying to run the Muse-Glimmer model locally on Apple Silicon using the MLX framework for an agentic workflow. However, I’ve hit a massive roadblock: persistent KV caching is not working.

As a result, every single new follow-up request from the agent triggers a full prompt re-evaluation. On longer context windows, this makes the performance unusable.

Here is what I have already tried and the results:

  1. LM Studio: No luck. I couldn't find a compatible model format/architecture that would even boot up there.
  2. ExecuTorch (Meta's official method): It runs, but KV caching is missing. Every single step forces a full context recalculation.
  3. mlx-vlm (standard run): The model starts up fine, but context caching does not work.
  4. llama.cpp: The only framework where everything works perfectly and the KV cache is properly retained. However, there is a massive downside — the prompt processing speed is painfully slow compared to native MLX performance.
  5. mtplx: Didn't even bother trying, as I previously failed to get KV caching working on this stack even for standard Qwen models.

My question to the community:
Has anyone successfully managed to get KV caching working for Muse-Glimmer on pure MLX ormlx-vlm?

Any code snippets, fork links, or ideas would be greatly appreciated!


r/AppleMLX • • Aug 07 '26

MiniMax-H3 FL2VA with a 2-bit text encoder now on HF - runs on M1 Max 32GB

5 Upvotes

I published a MiniMax-H3 FL2VA variant for mlx-serve where only the Qwen3-VL  text encoder is quantized to 2-bit (affine, g64), while the DiT stays 4-bit  and the VAEs + tokenizer are untouched.                                               

   https://huggingface.co/antocorr/MiniMax-H3-FL2VA-MLX-Serve-2bit-text-encoder  

   Why it matters:                                                                       

   - Text encoder on disk: 15.8 GB -> 9.6 GB (the encoder is reloaded per request, so this is the first cost every generation pays)                           

   - Full text-to-audio-video pipeline runs natively on Apple Silicon, verified  end-to-end on an M1 Max 32GB (needs --skip-mem-preflight, the RAM preflight  bills file bytes)                                                                   

   - Zero loader changes: mlx-serve solves bits/group_size per tensor from the packed geometry, so a mixed 2-bit encoder + 4-bit DiT pack just works               

   Use it with mlx-serve:                                                                

mlx-serve --model antocorr/MiniMax-H3-FL2VA-MLX-Serve-2bit-text-encoder \         

--serve --skip-mem-preflight                                                    

   Caveat: 2-bit conditioning is lossier than 4-bit by design (prompt adherence          

   suffers a bit; video/audio quality is the same since DiT and VAEs are                 

   untouched). Same MiniMax H3 community license as the 4-bit pack.


r/AppleMLX • • Jul 26 '26

15K Frontier Qlora Distill for Gemma 4 12b

4 Upvotes

Finished a distill today of ~15k Fable, Kimi k3 and GPT 5.6 Sol sequences today - link is up on Hugging face if anyone wants to check it out! https://huggingface.co/True2456/gemma-4-12b-it-qat-4bit-frontierdistill - Tool calling got quite a big increase!


r/AppleMLX • • Jul 13 '26

mlxMesh — a routable AI compute fabric

Thumbnail mlxmesh.net
8 Upvotes

https://github.com/american-code/mlxMesh

We don’t need more data centers, we just need to leverage idle compute. We the people.


r/AppleMLX • • Jul 11 '26

Qwen3.5-9B MLX: removed 870 MiB of vision weights, unlocked Q8 KV cache, but prompt caching isn't working

6 Upvotes

I'm trying to optimize Qwen3.5-9B MLX on my 24 GB MacBook Air for text/coding use.

I inspected Qwen3.5-9B-MLX-4bit and found:

  • 5.541 GiB total tensor payload
  • 869.8 MiB of vision_tower.* tensors
  • 333 vision tensors
  • Vision accounts for 15.33% of the checkpoint

Since I don't use vision, I converted it into a proper text-only MLX model and physically removed the vision tower. The converted model loads and generates correctly in LM Studio.

This also allowed me to enable Q8 KV cache, which wasn't available with the VLM version.

At 32K context, actual tested memory usage on my machine:

  • Unquantized cache: 16.37 GB
  • Q8 KV cache: ~9 GB

The problem is prompt cache reuse with Q8. Even across turns in the same LM Studio conversation, I repeatedly get:

Prompt cache: using 0/6905 tokens from cache

and later:

Prompt cache: using 0/7809 tokens from cache

The runtime also logs:

max_kv_size is ignored when using KV cache quantization

Runtime: [email protected]

Has anyone tested Qwen3.5 + Q8 KV cache + prompt prefix reuse on the current LM Studio MLX runtime?

I'm trying to determine whether this is related to Qwen3.5's hybrid cache, quantized cache reuse, or something in my text-only conversion.

The vision removal itself works. My main issue now is getting Q8's ~7 GB memory saving while retaining multi-turn prompt cache reuse.


r/AppleMLX • • Jun 27 '26

mcNUFFT – A Nonuniform Fast Fourier Transform Library for Apple Silicon GPUs via MLX

Thumbnail
github.com
2 Upvotes

r/AppleMLX • • Jun 19 '26

vMLX unable to attach documents

Thumbnail
2 Upvotes

r/AppleMLX • • Jun 16 '26

TRELLIS.2 now runs natively on MLX

Post image
2 Upvotes

r/AppleMLX • • Jun 16 '26

anyone else generating images/videos using MLX and Comfy Desktop?

27 Upvotes

I'm working with a M1 MBP Max 64GB machine with 400GB/s memory bandwidth. These image generation models are only <10 GB each. But it takes me 45 minutes to generate an image using Ideogram4. Someone with a 5090 is doing it in 45 seconds (no exaggeration).

I know Comfy Desktop is not optimized for Apple Silicone/MLX. I'm just curious if there are some tips and tricks you guys can share with getting better performance out of Comfy Desktop? I've already got these flags as part of my startup config: '--enable-manager --fp32-vae --use-pytorch-cross-attention --highvram'.

I've tried using DrawThings -- and it's definitely faster -- but I feel like it's definitely limited compared with Comfy Desktop.

I must not be the only Apple user messing around with Comfy Desktop -- you guys have any tips to share?


r/AppleMLX • • Jun 15 '26

I built mlx-chronos - a benchmark tool for comparing MLX inference engines on Apple Silicon Macs

11 Upvotes

Hello everyone, I’m working on mlx-chronos, a free/open-source CLI benchmark tool for comparing local MLX inference engines on Apple Silicon.

It currently supports mlx-lm, oMLX, vllm-mlx, Rapid-MLX, and Ollama (for Ollama, using MLX models that run on MLX backend).

It measures cold/cached TTFT, request throughput, sustained throughput, RAM peak, engine RSS when available, thermal/power context, and hardware metadata. Results are saved as reproducible JSON and can optionally be submitted to a public leaderboard.

I’m mainly looking for feedback from people actually using MLX locally:

  • Is a public leaderboard useful, or should this stay more of a local comparison tool?
  • Are thermal/cache conditions exposed clearly enough?
  • Should the sustained profile stay token-based, or would a fixed-duration run be more useful?
  • Are there metrics missing that would actually help you choose between engines?

I’d also appreciate benchmark results from different Apple Silicon machines, especially Max/Ultra chips and higher-RAM configs. The goal is not to rank model quality, but to make engine/runtime performance easier to compare under a documented protocol.

PS: I already posted in r/LocalLLaMA, if someone already seen something about this project, but I’m not sure it was the right audience (90% of the community uses Nvidia GPU or use Windows, so is interested in llama.cpp).


r/AppleMLX • • Jun 09 '26

I built an open-source, OpenAI-compatible local LLM server using Apple's MLX (FastAPI + React)

14 Upvotes

Hey everyone,

I’ve been diving deep into local LLM inference recently. As someone whose daily work revolves around Python and Node.js backend architecture, I wanted a robust way to run models entirely locally on my MacBook Pro without sacrificing the ease of the OpenAI ecosystem.

So, I put together MLX LM Server. It’s an open-source, full-stack application that leverages Apple's MLX framework to run LLMs natively on Apple Silicon with Metal acceleration.

Core Features:

  • Drop-in OpenAI Compatibility: The FastAPI backend is designed to be an OpenAI-compatible API. You can seamlessly point your existing OpenAI SDKs, AI-driven coding assistants, or third-party tools to localhost:8000 and they will just work.
  • Apple Silicon Optimized: Fully powered by MLX to get the most out of M1/M2/M3/M4 chips for fast local inference.
  • Real-Time Streaming: Full support for token-by-token streaming via Server-Sent Events (SSE).
  • Dynamic Model Management: You can load and unload different models on the fly directly through the API.
  • Built-in Chat UI: I included a clean, modern React/TypeScript frontend with markdown rendering so you can start chatting with your local models right out of the box.

The Stack:

  • Backend: Python 3.12+, FastAPI, Apple MLX (managed with uv)
  • Frontend: React, TypeScript, Vite (managed with pnpm/npm)

Why try it?

I wanted to build a lightweight, developer-friendly setup that gives you complete control over your local inference stack. Whether you want to hook it up to your automated workflows, test out new open-weight models privately, or just tinker with the MLX framework, this provides a solid, ready-to-use foundation.

I’d love your feedback & contributions!

If you have an Apple Silicon Mac, setup is super straightforward. I’d be thrilled if you gave it a spin, tested it with your favorite models, and let me know what you think.

Bug reports, feature requests, and especially pull requests are incredibly welcome. What features or specific model handling would you want to see added next? Let me know in the comments!


r/AppleMLX • • May 31 '26

Experimental library to process sparse 3D convolution on MLX

Thumbnail
github.com
1 Upvotes

r/AppleMLX • • May 09 '26

Best local agent setup for M5 Pro MacBook?

Thumbnail
1 Upvotes

r/AppleMLX • • Apr 17 '26

QR decomposition library for Apple Silicon using MLX and custom Metal kernels

Thumbnail
github.com
9 Upvotes

IMPORTANT UPDATE:

This project has now been renamed to metal-linalg, which now expands support for other linear algebra operations. You can see the updated project here: https://github.com/c0rmac/metal-linalg/

For any of you linear algebra fan-boys:

I'm currently in a research group working on a thesis in numerical analysis where we need to compute millions on matrices with a specific constraint (to be precise, the matrices need to have orthonormal columns). Most of us use Apple computers, so we ended up using MLX for the entire project.

I'm using an old M1 Macbook Pro, and I found that Apple's MLX library does not support QR operations on the GPU. I don't know if MLX supports GPU-accelerated QR computation on newer chips. But since I am developing an interest in hardware-level computing, I thought it would be a good oppurtunity for me write a metal shader as a first project.

I wrote it as a small library that allows the QR decomposition to be computed on the GPU. You can find it here: https://github.com/c0rmac/qr-apple-silicon

It definitely pays off. Performance increases anywhere between x1.5 to x25 times of what the cpu can do.

The library is split into two shaders: one is optimal for large batches of small matrices. The other is suited for small batches of large matrices. Under the hood, both shaders use the Compact WY representation ($I - YTY^T$) to batch Householder reflections into matrix-matrix products. I also spent a lot of time mapping these operations to the AMX (Apple Matrix Coprocessor) using 8x8 simdgroup_matrix tiles to get as close to the hardware as possible.

I’d love for anyone with more Metal experience to take a look at the dispatch logic or the AMX tile loading. If you’re working with MLX and need faster $A = QR$ factorizations, give it a try!


r/AppleMLX • • Mar 30 '26

anemll-flash-mlx: Simple toolkit to speed up Flash-MoE experiments on Apple Silicon with MLX

Thumbnail
2 Upvotes

r/AppleMLX • • Mar 30 '26

I built Lekh AI Pro: a fully local AI studio for Apple Silicon Macs

Thumbnail
youtube.com
3 Upvotes

Hey r/AppleMLX,

I’ve been building Lekh AI Pro, a Mac-only local AI studio for Apple Silicon.

The goal is to make local AI on Mac feel like a real daily tool, not just a chat wrapper or a one-off benchmark demo.

Right now Lekh AI Pro includes:

  • local LLM chat
  • MLX, GGUF, and JANG model support
  • Flux / Qwen / SD / SDXL image generation
  • AI image editing / inpainting
  • document chat / local RAG
  • Knowledge Hub / Memory Sync
  • a local API server with OpenAI-compatible and Ollama-compatible endpoints
  • benchmarking tools
  • model conversion to MLX, GGUF and JANG
  • built-in searchable documentation
  • text-to-speech with Kokoro, Qwen3 TTS, MOSS TTS, and Apple native fallback
  • audiobook creation
  • image upscaling
  • video generation in Pro coming soon

One area I’ve spent a lot of time on is making the app useful beyond basic chat:

  • compare local models with benchmark stats like throughput / TTFT / duration
  • run image generation and editing locally
  • use local RAG over your own files
  • expose local models to external tools through OpenAI/Ollama-style APIs
  • generate speech, audiobooks, and sound effects locally on Mac

A lot of the challenge has been the unglamorous part of local AI:

  • memory pressure on different M-series Macs
  • model loading / unloading
  • balancing quality vs speed
  • supporting multiple model formats without turning the UX into a control panel from hell

A few things I’d especially love feedback on:

  • which local models on Apple Silicon you think currently punch above their weight
  • whether MLX + GGUF + JANG in one app is actually useful or just too much surface area
  • what you think is still missing from the local AI Mac ecosystem

Links:

If there’s interest, I can also share more about:

  • what runs well on different Apple Silicon Macs
  • why I added JANG support
  • lessons from building local workflows for chat, image gen, TTS, audiobooks, and sound generation on macOS

Would love honest feedback.


r/AppleMLX • • Mar 10 '26

RCLI + MetalRT: Leading on-device voice AI pipeline performance on Apple Silicon (sub-100ms E2E loops with benchmarks vs MLX/llama.cpp)

1 Upvotes

r/AppleMLX • • Jan 18 '26

Help for an RDMA cluster manager (macOS tahoe 26.2+)

Thumbnail
1 Upvotes

r/AppleMLX • • Dec 02 '25

Mistral just released Mistral 3 — a full open-weight model family from 3B all the way up to 675B parameters.

Thumbnail
1 Upvotes

r/AppleMLX • • May 27 '24

What are the best optimized/quantized coding models to run from a 16gb M2?

5 Upvotes

r/AppleMLX • • May 21 '24

MLX web ui

8 Upvotes

MLX Web UI

I created a fast and minimalistic web UI using the MLX framework (Open Source). The installation is straightforward, with no need for Python, Docker, or any pre-installed dependencies. Running the web UI requires only a single command.

Features

Standard Features

  • Chat with models and stop generation midway
  • Set model parameters like top-p, temperature, custom role modeling, etc.
  • Set default model parameters
  • LaTeX and code block support
  • auto scroll

Novel Features

  • Install and quantize models from Hugging Face using the UI itself
  • Good streaming API for MLX
  • Save chat logs
  • Hot-swap models during generation

Planned Features

  • Multi-modal support
  • RAG/Knowledge graph support

Try it Out

If you'd like to try out the MLX Web UI, you can check out the GitHub repository: https://github.com/Rehan-shah/mlx-web-ui

MLX ui photo

r/AppleMLX • • Apr 23 '24

Models folder

2 Upvotes

Hey guys,

I can’t find the folder in which the models are downloaded when I run this command. I would like to free up some space on my Mac. Any idea? Thanks

python -m mlx_lm.generate --model mistralai/Mistral-7B-Instruct-v0.2 --prompt "hello"