A code-editing agent in Go · agent no. 3 by Aurimas Nausėdas
An LLM, a loop, and three tools.
That is all a code-editing agent is. This one is built from a 317-line tutorial file, then made safe enough to put on the internet: sandboxed, rate-limited, traced and tested. Try it, or build it yourself with a guide that shows where every line goes.
- 39/39
- deterministic evals pass on every commit
- 6
- tutorial checkpoints, each one compiled in CI
- 0
- lines of model-written code run on the server
#eli51/4You ask a question. It goes on the pile of notes and the whole pile is sent to Claude.
The whole idea
Claude can't touch anything. Your program can.
Claude is a clever friend on the phone. It can't see your desk or pick anything up. So you agree on three favours it may ask for: “read me that page”, “tell me what's on the desk”, and “swap these words for those”.
Your program listens. When Claude asks for a favour, the program does it and reads the result back. Then Claude might ask for another, or say “done, here's your answer”. That back-and-forth is the loop, and the loop is the agent.
The colours on this site mean something, and they come from the tutorial's own terminal output: blue is you, yellow is Claude, green is a tool.
From tutorial to production
What it takes before strangers can use it
The loop itself did not change. What changed is everything around it. Each item links to the evidence.
A sandbox, not a promise
The tutorial agent can read and write any file its user can. Here every path is checked in one place and resolved inside a workspace the agent cannot leave. That is enforced by Go code, so no prompt can talk its way around it.
See the sandbox evals →Edits that can't silently wreck a file
Two inputs made the tutorial's edit_file damage files without an error. Both are now refused with a message the model can act on.
See the edit-rule evals →Limits on everything that costs money
Rounds per message, messages per session, messages per visitor, and a daily budget in dollars that survives restarts. When the budget is gone the model is not called.
See the limits →Every step visible and stored
Each model call, tool call, result and guardrail is an event. The browser shows them live, the database keeps them as a trace, and the tests assert on them.
See the X-ray panel →Web research with sources
A second mode adds web search. Answers come with the pages they rely on, and are kept private until a person reviews and publishes them.
See the research mode →Evals, not vibes
Hostile inputs against each tool, a scripted model driving the real loop over the real wire format, and a live-model suite that checks the files the agent leaves behind.
See the eval report →
Who made this
Aurimas Nausėdas
I build AI products and teach Python and AI to working professionals in Lithuania and Latvia. This is the third agent I have built from first principles, each one to understand a different part of the stack by making it work end to end.
The agent's core follows Thorsten Ball's tutorial. The refactor, the sandbox, the limits, the evals and this site were built with Claude as a pair programmer; the design decisions and the reasons for them are written down in the repository.
The other two agents
- Calculator AgentPython, from scratch
- Claude Agent From ScratchFastAPI and Next.js
- Code-Editing AgentGo, this site