Author: Ansel Robateau 10/1/2026
I still remember the first time I beat Mike Tyson's Punch-Out!!. Not with reflexes. With pattern recognition. Every fighter telegraphed their punches; you just had to watch, learn the tells, and respond. The NES didn't have the hardware to be impressive, so the games had to be smart about the hardware they had.
I thought about that a lot over the last few weeks, while trying to answer a deceptively simple question: can a two-billion-parameter language model do real work?
Not chat. Work. Write files. Run commands. Fix bugs. Chain steps together without wandering off into the weeds.
The model in question is gemma4:e2b, roughly two billion parameters, running locally via Ollama on my Mac. No data center, no API bill, no vast resources. And that last part matters to me. My heart is still for Belize, a country without America's vast resources but full of people with enormous promise. I've always believed technology should reach people like that, not just the ones with a cluster. If useful AI can run on a laptop, it can run anywhere.
So the little model is my underdog, and I wanted to see how far it could go.
If Ollama and Gemma are new names to you, here is the short version.
Ollama is a free, open source tool that runs large language models on your own machine. Install it, pull a model with one command, and you are chatting with an AI that never sends your data to anyone's cloud. It has become the standard way to run open models locally. Learn more at ollama.com (https://ollama.com).
Gemma is Google's family of open models, built from the same research and technology as their flagship Gemini models, and released openly for anyone to download and run. Gemma 4, the current generation, comes in sizes from E2B (effective 2 billion parameters, small enough for a laptop) up to 31B, all under the Apache 2.0 license. The two models in this post, gemma4:e2b and gemma4:e4b, are the two smallest of the family. Full documentation at ai.google.dev/gemma/docs (https://ai.google.dev/gemma/docs).
I started the way everyone starts: I gave the model a task, a set of tools, and told it to loop until done. Like handing a teenager the car keys and saying "drive to Belize." Technically an instruction. Practically a disaster.
It scored 1 out of 5 on my little eval: five small, checkable jobs. The failures were almost funny. The model would carefully type out what a tool call looked like instead of actually calling it, like a kid making vroom-vroom noises instead of turning the ignition. It looped forever. It declared victory immediately after failing.
Here's the thing I kept coming back to, the way you come back to a cut that isn't square: the problem wasn't the saw. It was the absence of a jig.
Anyone who's done woodworking knows the jig, the humble fixture that guides your tool so the cut comes out straight every time. A jig doesn't make the saw smarter. It makes it harder for the saw to go wrong. You don't ask a chisel to be a hammer; you build the setup around what the tool actually is.
That became the thesis of the whole project: strategy and executive function belong in the harness, not in the model. The small model should only ever receive work that has been decomposed, checked, and shaped to fit it, the way NES developers wrote tight code for a 1.79 MHz CPU instead of wishing for a PlayStation.
So every failure went into the harness. The model mimicked tool calls? Native tool calling, with a text fallback. It looped forever? Labeled results after every call, with a nudge: that worked, so call the next tool or say you're done. It hallucinated filenames it had never seen? A dedicated search tool, so it never touches a filename it hasn't been shown. It declared victory over its own errors? The harness now pushes back: you just failed, so engage with that before you tell me you're finished.
The biggest jump came when the checkers stopped being graders and started being teachers. Now, when an acceptance check fails, the harness retries: same workspace, the failure message as feedback, instructions to fix rather than redo. Test-driven development, but the process plays the developer. The oldest rule in the shop, automated: measure twice, cut once.
The full eval went 5/5.
Then came the part I care about most. The planner now writes its work up as Gherkin, a plain-language feature file, one scenario per step, and I review it before anything runs. Like a blueprint the client signs before the first cut:
Scenario: Capture the script's output
Given src/main.py prints its own filename
When I run python3 src/main.py and save the output to out.txt
Then out.txt exists
Then out.txt contains the name of the script that ran
Notice what the Then doesn't say: nothing about implementation. The contract states the observable outcome; the agent is free on the how. Gherkin turns out to be a fantastic bridge language between humans and agents, a universal translator between what I mean and what the machine does, because it's a work order both sides can genuinely agree on.
One honest caveat: the harness verifies the Thens it can phrase as mechanical checks (file exists, file contains text) and marks the rest as taken on trust, right in the plan review. The contract is always readable by both sides; the machine just checks whatever it can.
The default workflow: the agent proposes the plan, I review it, the harness runs it. Auto mode exists, but review-first is the point.
My favorite mechanism is a small act of paranoia: the code-smell gate. The multi-step task kept producing scripts with the filename hardcoded; print("main.py") passes if you squint. So the checker runs a renamed copy in a temp directory and demands the output change. Behavior check plus intent check. You can't fake understanding anymore.
And the punchline of the whole experiment: the small model's most stubborn failure was never a reasoning failure. It just didn't know the idiom, os.path.basename(__file__). An apprentice who was never shown the trick. So I built a suggestion library: trigger phrases mapped to copy-pasteable idioms, slipped into the planner's prompt and the retry feedback. One entry took the model from 4/5 to 5/5 on the first attempt. Knowledge gaps, it turns out, are fillable in the harness, not the model.
The twist: I ran the same eval against the bigger sibling, gemma4:e4b. It also scored 5/5, in twice the time and failing differently: sloppier plans, literal placeholders. A V8 in a go-kart doesn't win the race if the steering is loose. Bigger didn't mean better; it meant different failure modes, all absorbed by the same harness.
The code is public at github.com/robaone-silas/agent-harness. Zero dependencies, standard library only. But here's why I actually built it, and why I'm writing about it:
Every capability I moved into the harness is capability that doesn't require a bigger model. And smaller models run on smaller machines: laptops, old desktops, the kind of hardware that's actually available in places like Belize. The Death Star approach to AI, ever-larger models behind ever-pricier APIs, concentrates power where the resources already are. The R2-D2 approach, a small, competent droid with the right socket and clear instructions, puts it in more hands.
That's always been my motivation: use technology to help people. All people, including the ones the vast resources haven't reached yet. If a two-billion-parameter model, properly jigged, can do real work on a laptop, then the future doesn't belong only to whoever can afford the biggest cluster.
It belongs to anyone with a good jig and the patience to measure twice.
This is the first iteration. Right now the harness does coding tasks, because coding tasks are easy to check. But the pattern underneath is general: propose a plan in a language both sides understand, let the human review it, execute it stepwise, verify each step, and retry with feedback when something fails. There is nothing about that loop that only works for code.
So the roadmap is to widen the aperture. Richer verifiers that can check command output, not just files. A suggestion library that grows with every gap the evals reveal. And task types beyond code: research, analysis, writing, the everyday knowledge work that eats afternoons.
The goal hasn't changed since the first line of this post. I want this harness to do real work for real people. Not demos, not benchmarks. Work that matters to someone, on hardware they already own, wherever they happen to live.
The jig is built. Now let's see what it can cut.
APLS Code: H-C-H (What is this?)
Genesis: Human (Core thesis, experiments, and article concept)
Elaboration: Collaborative (Iterative drafting with an AI assistant, redirected and rewritten at the author's direction)
Validation: Human (Reviewed, fact-checked, and edited by the author before publishing)