What I Run on a Local Model — and What I Still Send to Claude
I ran evals on real past inputs to decide which of my AI automations should move to a local model, and one of the two failed.
Ingredients
- A desktop computer with plenty of memory — memory is what lets a machine hold a 30-billion-parameter model
- Ollama — runs open models locally (free)
- Open models — Gemma, Qwen and Llama, from 3B to 70B (free downloads)
- Claude Code — for writing the eval harnesses and the plumbing (already had it)
The Rule: Run an Eval Before You Switch
Several of my scheduled jobs called Claude: a morning market email, a street camera that describes what passes my window, a headline classifier, and search. The tempting move was to port all of it in a weekend.
I didn’t, because a worse model doesn’t crash. It just writes a slightly less true sentence, and nobody files a bug about a sentence. So every job had to pass an eval before it moved: a side-by-side test of the current model and the challenger, run on the job’s own real inputs and graded on the one thing the job has to get right.
Test One: The Morning Market Email (Local Won)
The hardest part of the market email is one paragraph: why markets moved, written from that morning’s headlines and prices. Read twenty headlines, pick the two that mattered, tie them to the numbers.
There was no eval set, because the job had never saved what it read. So I rebuilt twelve past mornings from the emails it had already sent, and ran four local models on them, graded on one question: is the explanation true, and does it match the numbers it was given?
What settled it: on two of the twelve days, the winning local model tied an oil move to the actual crude price in the data, while the Claude model then writing the email reached for a round “$100 a barrel” that wasn’t there. That’s exactly the quiet failure I was worried about, and this time Claude was the one making it.
When Gemma 4 came out a few days later, I re-ran the same test. The 26B version was fastest at about four seconds; the 31B took about seventeen. I picked the 31B. This is the one email I read every morning, and a job that runs at 7am can spare thirteen seconds.
🔧 Developer section: turn “thinking” off
With thinking mode on, one model produced six to seven thousand tokens of internal reasoning to write a single sentence. For a short, grounded paragraph, turn it off. The job now also saves its inputs every morning, so the next eval won’t need a rebuilt eval set.
Test Two: The Street Camera (Local Lost)
The camera job has two AI steps: spot the thing in the frame, then describe the event in a sentence. I tested a 32B vision model on both, using clips I’d already checked by hand for the Bike, or Not Bike? post.
Rejected. The camera keeps its small trained detector and keeps sending descriptions to Claude, and the project notes now say “don’t re-run this” so a future me with a shiny new model doesn’t lose a weekend rediscovering it.
The same model was careful when asked to draw a box around an object and careless when asked to describe a scene in free text. A local model with a tight output (a box, a label, a yes or no) is a very different product from one writing prose.
The Part Nobody Warns You About: Sharing One Model Server
Once several jobs share one model server, they need to coordinate. Without it, two big jobs arriving at once will each push the other’s model out of memory to load their own, and both slow to a crawl. The fix is deliberately boring: take turns.
🔧 Developer section: the setting that un-pins your model
Ollama keeps a model loaded forever if you send keep_alive: -1, but every request re-applies its own value to the loaded model. One job that sends "30m" quietly downgrades your forever-pin to thirty minutes. Callers of a pinned model should send nothing, and a small checker re-pins anything that slips.
Five Weeks Later
Where Every Job Landed
Local models are excellent at narrow work: classify this, embed this, write one grounded paragraph from this exact data. They’re much less trustworthy at open-ended description, where nothing stops them from inventing a scooter.
One policy changed as a result. New automations now start on the local model by default and move to Claude only when a test says so. Migrating one job at a time left a long tail of half-moved work; starting local is simpler.
What went fast
- Getting models running. The first 27B model was a one-command download.
- Re-running the eval. Once the harness existed, a new model family took an afternoon.
- The fallback. “If local doesn’t answer, use Claude” was a few lines, and it’s why I could switch a daily email without worrying.
What needed patience
- Building an eval set that didn’t exist. Rebuilding past inputs was the unglamorous part that made the comparison honest.
- Accepting the loss. I wanted the camera test to pass. Writing “rejected, don’t re-run” took more discipline than running it.
- Waiting to write this. A week of good test numbers is a demo; five weeks of mornings is a track record.
If you’re a PM wondering whether a local model belongs in your stack, skip the leaderboards and build a small eval. Pick one job, pull its real past inputs, run the current model and the challenger side by side, and grade the one thing that job exists to get right. The biggest model may lose, and the answer will be different for every job.