read it

my writeup on running models on iPhones: how MLX uses Apple Silicon's unified memory, quantization, which iPhones can handle it, and what it costs in battery and heat.

overview

I spent a couple of weeks in April 2026 trying to run my own AI instead of renting someone else's. It went in two directions. On my iPhone, a small model running entirely on the phone, no internet needed. On my MacBook, something much bigger: a fully local assistant I could text from my phone that could search the web, built from open-source parts and costing nothing per message.

The phone half is the one I kept. The Mac half worked end to end, and then I measured it honestly and shelved it. This page is both, and why.

on the phone

Google's Gemma 4 E4B, instruction-tuned and quantized to 4 bits, running on my iPhone 15 Pro. At 4-bit it's about 2.2 GB, which fits comfortably in the phone's 8 GB of unified memory, where the CPU and GPU share one pool so nothing gets copied between them. I started with it in Google's AI Edge Gallery, and for quick questions I ended up in Locally AI, an MLX-based app. Both work in airplane mode.

E4B over E2B: Gemma 4's edge models come in two sizes. The 15 Pro has the memory for the larger one, so that's the one I run.
the use case: A quick search engine that lives on my phone, for the dozens of small lookups a day that don't need a frontier model or my data leaving the device.
the costs: Local inference is hard on a phone. There's no fan, so sustained use throttles, and it drains the battery far faster than sending the same question to a server. It's the right tool for short questions, not long sessions.

on the mac

The goal: text my own local model from my iPhone, from anywhere, with no cloud APIs and no subscriptions, on a base M2 Pro MacBook with 16 GB of RAM. I got there. A message went from Discord on my phone, over Tailscale, into an OpenClaw gateway on the Mac, through a local model, out to a self-hosted search engine and back. Getting there meant replacing most of the parts I started with.

runtime: I started on Ollama, then moved to oMLX. MLX runs roughly 1.5 to 2 times faster than llama.cpp (what Ollama uses) on Apple Silicon, because it's built around unified memory instead of retrofitted to it. Model sizes don't translate between formats either: the same Gemma 4 E4B was 2.2 GB as an MLX bundle and 9.6 GB as an Ollama GGUF.
model: Llama 3.2 3B was too small. Asked for 23 × 47, it emitted a call to a "math" tool that didn't exist. I moved to Qwen 2.5 7B at 4-bit, about as large as 16 GB allows.
tools: Osaurus gave me around 55 MCP tools (mail, calendar, files, web, browser control), but its chat endpoint returned two JSON chunks per reply, which broke OpenClaw's parser. Search ended up as SearXNG in Docker, the cleanest part of the whole build: one container, one config file.
chat: I started with Matrix for end-to-end encryption. A headless bot can't verify its own device, so it couldn't read encrypted rooms, and its access tokens kept getting invalidated. I switched to a Discord bot, which just works.

why i shelved it

It worked, and it was too slow and not smart enough to be worth using. Two numbers made the decision for me.

three minutes per reply: OpenClaw rebuilds its full context on every message: system prompt, workspace files, tool schemas, history. That came to 25,094 tokens, even for "what time is it". At about 250 tokens a second, the model spent roughly 100 seconds reading before it wrote a word. Add generation and Discord on both ends and a reply took two to three minutes.
a capability floor: Asked to fetch a page it didn't have a URL for, the 7B model invented a plausible-looking one instead of searching first. Given three weather sources that disagreed, it listed all three instead of picking one. No prompt fixed either, because they weren't prompt problems. They were what a 7B model at 4-bit can and can't do.
the RAM ceiling: macOS, the model, its cache and the stack's own services left almost nothing on 16 GB. Once memory pressure went yellow, the cache paged to disk and replies stalled. I couldn't run the stack and a browser at the same time.

Meanwhile the phone answered quick questions faster, with no network in the way, and a frontier model was smarter for hard ones. The Mac stack sat in a middle that nothing in my week actually needed. So now it's a hybrid: the local model on my phone for fast, private lookups, and a frontier model for real work.

what i kept

tool calling is four problems: the framework has to serialize the tools, the runtime has to pass them through, the model has to emit a valid call, and the model has to pick the right tool. My model was fine at the first three. I spent a session debugging plumbing before realizing the failure was the fourth.
a socat proxy: When a runtime won't log what it receives, put a logging proxy between the two sides and read every byte. I proved OpenClaw was sending tool schemas this way, and I've reached for the trick since.
verify the README: OpenClaw had CLI commands for adding MCP servers, and they showed up as configured, but the connector that would actually start them wasn't shipped. If nobody has an end-to-end example of a feature working, I now assume it doesn't yet.
ask about the ceiling first: Every session I thought the next fix would make it good enough. The weakest part, the model, set the ceiling, and nothing else could raise it. Now I test the model on the hardest prompt I care about before building anything around it.

construction

Gemma 4 E4B MLX Locally AI AI Edge Gallery oMLX Ollama Qwen 2.5 7B OpenClaw Osaurus MCP SearXNG Docker Discord Tailscale launchd

architecture

iPhone (Discord) → Tailscale → OpenClaw gateway → oMLX + Qwen 2.5 7B → SearXNG → reply

Everything ran on the MacBook, with the gateway kept alive as a LaunchAgent. On the phone there's no architecture at all: the model and the app are the whole thing.

Source? The writeup is on my GitHub.