read it
my writeup on running models on iPhones: how MLX uses Apple Silicon's unified memory,
quantization, which iPhones can handle it, and what it costs in battery and heat.
overview
I spent a couple of weeks in April 2026 trying to run my own AI instead of renting someone else's.
It went in two directions. On my iPhone, a small model running entirely on the phone, no internet
needed. On my MacBook, something much bigger: a fully local assistant I could text from my phone
that could search the web, built from open-source parts and costing nothing per message.
The phone half is the one I kept. The Mac half worked end to end, and then I measured it honestly
and shelved it. This page is both, and why.
on the phone
Google's Gemma 4 E4B, instruction-tuned and quantized to 4 bits, running on my
iPhone 15 Pro. At 4-bit it's about 2.2 GB, which fits comfortably in the phone's 8 GB of unified
memory, where the CPU and GPU share one pool so nothing gets copied between them. I started with it
in Google's AI Edge Gallery, and for quick questions I ended up in Locally AI, an MLX-based app. Both
work in airplane mode.
E4B over E2B: Gemma 4's edge models come in two sizes. The 15 Pro has the memory
for the larger one, so that's the one I run.
the use case: A quick search engine that lives on my phone, for the dozens of
small lookups a day that don't need a frontier model or my data leaving the device.
the costs: Local inference is hard on a phone. There's no fan, so sustained use
throttles, and it drains the battery far faster than sending the same question to a server. It's
the right tool for short questions, not long sessions.
on the mac
The goal: text my own local model from my iPhone, from anywhere, with no cloud APIs and no
subscriptions, on a base M2 Pro MacBook with 16 GB of RAM. I got there. A message went from
Discord on my phone, over Tailscale, into an OpenClaw gateway on the Mac, through a local model, out
to a self-hosted search engine and back. Getting there meant replacing most of the parts I started
with.
runtime: I started on Ollama, then moved to oMLX. MLX runs
roughly 1.5 to 2 times faster than llama.cpp (what Ollama uses) on Apple Silicon, because it's
built around unified memory instead of retrofitted to it. Model sizes don't translate between
formats either: the same Gemma 4 E4B was 2.2 GB as an MLX bundle and 9.6 GB as an Ollama GGUF.
model: Llama 3.2 3B was too small. Asked for 23 × 47, it emitted a call to a
"math" tool that didn't exist. I moved to Qwen 2.5 7B at 4-bit, about as large as 16 GB allows.
tools: Osaurus gave me around 55 MCP tools (mail, calendar, files, web,
browser control), but its chat endpoint returned two JSON chunks per reply, which broke
OpenClaw's parser. Search ended up as SearXNG in Docker, the cleanest part of the whole build:
one container, one config file.
chat: I started with Matrix for end-to-end encryption. A headless bot can't
verify its own device, so it couldn't read encrypted rooms, and its access tokens kept getting
invalidated. I switched to a Discord bot, which just works.
why i shelved it
It worked, and it was too slow and not smart enough to be worth using. Two numbers made the
decision for me.
three minutes per reply: OpenClaw rebuilds its full context on every message:
system prompt, workspace files, tool schemas, history. That came to 25,094 tokens, even for "what
time is it". At about 250 tokens a second, the model spent roughly 100 seconds reading before it
wrote a word. Add generation and Discord on both ends and a reply took two to three minutes.
a capability floor: Asked to fetch a page it didn't have a URL for, the 7B model
invented a plausible-looking one instead of searching first. Given three weather sources that
disagreed, it listed all three instead of picking one. No prompt fixed either, because they
weren't prompt problems. They were what a 7B model at 4-bit can and can't do.
the RAM ceiling: macOS, the model, its cache and the stack's own services left
almost nothing on 16 GB. Once memory pressure went yellow, the cache paged to disk and replies
stalled. I couldn't run the stack and a browser at the same time.
Meanwhile the phone answered quick questions faster, with no network in the way, and a frontier model
was smarter for hard ones.
The Mac stack sat in a middle that nothing in my week actually needed. So now it's a hybrid: the
local model on my phone for fast, private lookups, and a frontier model for real work.
what i kept
tool calling is four problems: the framework has to serialize the tools, the
runtime has to pass them through, the model has to emit a valid call, and the model has to pick
the right tool. My model was fine at the first three. I spent a session debugging plumbing before
realizing the failure was the fourth.
a socat proxy: When a runtime won't log what it receives, put a logging proxy
between the two sides and read every byte. I proved OpenClaw was sending tool schemas this way,
and I've reached for the trick since.
verify the README: OpenClaw had CLI commands for adding MCP servers, and they
showed up as configured, but the connector that would actually start them wasn't shipped. If
nobody has an end-to-end example of a feature working, I now assume it doesn't yet.
ask about the ceiling first: Every session I thought the next fix would make it
good enough. The weakest part, the model, set the ceiling, and nothing else could raise it. Now I
test the model on the hardest prompt I care about before building anything around it.
construction
Gemma 4 E4B
MLX
Locally AI
AI Edge Gallery
oMLX
Ollama
Qwen 2.5 7B
OpenClaw
Osaurus
MCP
SearXNG
Docker
Discord
Tailscale
launchd
architecture
iPhone (Discord)
→
Tailscale
→
OpenClaw gateway
→
oMLX + Qwen 2.5 7B
→
SearXNG
→
reply
Everything ran on the MacBook, with the gateway kept alive as a LaunchAgent. On the phone there's no
architecture at all: the model and the app are the whole thing.
Source? The writeup is on my GitHub.