← All Writing
August 28, 20266 min read

What I Run on a Local Model — and What I Still Send to Claude

I ran evals on real past inputs to decide which of my AI automations should move to a local model, and one of the two failed.

YieldA simple rule for moving an AI job onto a local model (or not), two real tests with opposite results, and one model server shared by several jobs
DifficultyIntermediate (running open models, designing evals from real production inputs)
Total Cook TimeA week of tests in late July, then five weeks of letting the winner run

Ingredients

The Rule: Run an Eval Before You Switch

Several of my scheduled jobs called Claude: a morning market email, a street camera that describes what passes my window, a headline classifier, and search. The tempting move was to port all of it in a weekend.

I didn’t, because a worse model doesn’t crash. It just writes a slightly less true sentence, and nobody files a bug about a sentence. So every job had to pass an eval before it moved: a side-by-side test of the current model and the challenger, run on the job’s own real inputs and graded on the one thing the job has to get right.

STEP 1Pick one jobthat runs on ClaudeSTEP 2Replay real pastinputs through bothSTEP 3Grade the one thingit must get rightLocal winsor ties?yesnoSwitch to localClaude stays as fallback✓ the market emailStay on ClaudeWrite down: don’t re-run✗ the street camera
The eval set is what the job actually saw on actual days, not a public benchmark. A loss gets written down so I don’t re-run it in a month out of optimism.

Test One: The Morning Market Email (Local Won)

The hardest part of the market email is one paragraph: why markets moved, written from that morning’s headlines and prices. Read twenty headlines, pick the two that mattered, tie them to the numbers.

There was no eval set, because the job had never saved what it read. So I rebuilt twelve past mornings from the emails it had already sent, and ran four local models on them, graded on one question: is the explanation true, and does it match the numbers it was given?

Finishing order, with average seconds per morning (12 recovered mornings)0s20s40s60s1stgemma3 27Bmost accurate causal reads37.8s2ndqwen3.5 9Bsolid, surprisingly for its size26.3s3rdqwen3 32Bfast, sometimes vague18.0sLastllama3.3 70Bterse, one hallucination58.8s
The biggest model came last. Speed wasn’t the gate; a correct, grounded explanation was.

What settled it: on two of the twelve days, the winning local model tied an oil move to the actual crude price in the data, while the Claude model then writing the email reached for a round “$100 a barrel” that wasn’t there. That’s exactly the quiet failure I was worried about, and this time Claude was the one making it.

When Gemma 4 came out a few days later, I re-ran the same test. The 26B version was fastest at about four seconds; the 31B took about seventeen. I picked the 31B. This is the one email I read every morning, and a job that runs at 7am can spare thirteen seconds.

🔧 Developer section: turn “thinking” off

With thinking mode on, one model produced six to seven thousand tokens of internal reasoning to write a single sentence. For a short, grounded paragraph, turn it off. The job now also saves its inputs every morning, so the next eval won’t need a rebuilt eval set.

Test Two: The Street Camera (Local Lost)

The camera job has two AI steps: spot the thing in the frame, then describe the event in a sentence. I tested a 32B vision model on both, using clips I’d already checked by hand for the Bike, or Not Bike? post.

Detection: seconds per framelog scale · both caught 20 of 20 two-wheelers10ms100ms1s10s100sfine-tuned detector~40 ms32B vision model~18 s~450× slowerA 15-second clip would take about 35 minutes.Descriptions: the same 15 eventseach dot is one event32B vision modelClaude7 of 15: invented a scooterjoggers, dog-walkers, people walking a bikeClaude got those 7 rightinvented scootercorrectother events
Tied on spotting bikes, 450 times slower at it, and wrong about scooters on nearly half the events. Any one of those was enough to keep this job on Claude.

Rejected. The camera keeps its small trained detector and keeps sending descriptions to Claude, and the project notes now say “don’t re-run this” so a future me with a shiny new model doesn’t lose a weekend rediscovering it.

What the failure taught me

The same model was careful when asked to draw a box around an object and careless when asked to describe a scene in free text. A local model with a tight output (a box, a label, a yes or no) is a very different product from one writing prose.

The Part Nobody Warns You About: Sharing One Model Server

Once several jobs share one model server, they need to coordinate. Without it, two big jobs arriving at once will each push the other’s model out of memory to load their own, and both slow to a crawl. The fix is deliberately boring: take turns.

Big-model job ABig-model job BBig-model job Casks for a turnTake turnsone heavy job at a timea crashed job loses its turnnext in lineHeadline classifiersmall model, must be instantskips the lineShared model serverone machine, one Ollamabig model, loadedfor whoever has the turnclassifier modelpinned in memory
Big jobs wait their turn; the one job that must be instant skips the line. If a job crashes while it has the turn, the turn expires on its own.

🔧 Developer section: the setting that un-pins your model

Ollama keeps a model loaded forever if you send keep_alive: -1, but every request re-applies its own value to the loaded model. One job that sends "30m" quietly downgrades your forever-pin to thirty minutes. Callers of a pinned model should send nothing, and a small checker re-pins anything that slips.

Five Weeks Later

MonTueWedThuFriJul 20Jul 27Aug 3Aug 10Aug 17Aug 2426 / 26trading morningswritten locally · 0 fallbackslocal modelbefore the switch
Every trading morning since the switch was written by the local model. The Claude fallback is wired in and hasn’t had to write a single one.
26 / 26trading mornings written locallyzero fallbacks to Claude5,325 / 5,325health checks passedone every ten minutes106 / 106heavy jobs finished and handed offnone got stuck holding the model
The machine underneath held up too.

Where Every Job Landed

Runs locallyStays on ClaudeMarket-email thesis 31BWon the bake-off on grounded accuracyHeadline classifier 9B, pinnedA narrow yes/no call that has to be instantSearch embeddingsTurning text into numbers; no judgmentPrivate chat windowConversations I'd rather keep at homeCamera event descriptionsLocal invented a scooter on 7 of 15 eventsAsk Goose, this site's chatbotAnswers visitors on the public siteWriting the code itself Claude CodeNothing local is close yet
Narrow jobs with a checkable answer went local. Open-ended description and writing code stayed on Claude.

Local models are excellent at narrow work: classify this, embed this, write one grounded paragraph from this exact data. They’re much less trustworthy at open-ended description, where nothing stops them from inventing a scooter.

One policy changed as a result. New automations now start on the local model by default and move to Claude only when a test says so. Migrating one job at a time left a long tail of half-moved work; starting local is simpler.

What went fast

What needed patience

If you’re a PM wondering whether a local model belongs in your stack, skip the leaderboards and build a small eval. Pick one job, pull its real past inputs, run the current model and the challenger side by side, and grade the one thing that job exists to get right. The biggest model may lose, and the answer will be different for every job.

← Back to all writing