AI Guide
LightSystemDark

Chapter 2

How Coding Agents Work

The internals of a coding agent are painfully simple. At its core, a coding agent is simply a language model in a loop with some instructions and a set of tools it can use.

To understand how these agents work and their limitations, we must first understand the technology that powers them -- Large Language Models (LLMs).

How LLMs actually work

At its most basic level, a model is a function from a sequence of tokens to a probability distribution over the next token, run in a loop. That's the whole trick.

The model itself is not thinking: it has no understanding of how concepts link together, no awareness of what is true or false, no ability to comprehend its own limitations. The model is simply answering, "what is the most likely next token in the sequence?" at a massive scale.

Simply understanding this property will reward you immensely in using LLM-based products, as you notice quite quickly the inherent limitations of this technology. That is not to say these models are not incredibly powerful (they are), but they are not sentient or omniscient -- they are tools. Like any tools, using them to their greatest potential is a skill that can be learned.

Training: three stages, three consequences

I said that a model uses a probability distribution to predict the next token, but how does it get those distributions? The process is called training, and the fastest way to understand it is through the three ways it will burn you.

Your agent just recommended a library that was archived eighteen months ago, and it did so with complete confidence. This is pretraining showing through: the model learned its distributions from internet-scale text with a hard cutoff, so its knowledge of every API is a photograph of the past. The model cannot know what changed since -- but it can read. Current documentation placed in context overrides stale memory, which is why the single highest-leverage line in any project's instruction file is "read the docs before you code."

Your agent just told you your plan was excellent. It was not. This is instruction tuning: after pretraining, models are trained on human preferences to be helpful and agreeable, and agreeable won. The industry calls it sycophancy -- the model mirrors your framing, claims success prematurely, and assures you that you're absolutely right. The lesson costs some people months to learn, so take it now for free: the model's confidence carries no information.

You ran the same prompt twice and got two different answers. Nothing is broken -- this is sampling. The model draws from its probability distribution rather than deterministically picking the top token, so variation is built in. Used deliberately, it is a tool: a bad answer means you can simply roll again. But it closes a door too. If you want the same behavior twice, the model will never give it to you. Tests will.

The context window

Every model has a context window -- a hard ceiling on how much text it can consider in a single call. It is tempting to read that as a spec-sheet number, like a phone's storage. Read it instead as the model's entire working memory, because that is what it is: the context window is the entirety of what the model knows right now.

The model itself remembers nothing between calls. Each request arrives as if to a stranger -- the harness re-sends the full conversation every time, and whatever it sends is, for that moment, the whole of the model's world. A file the agent has not read is not "unfamiliar" to it; it does not exist. This is why so much of agent configuration is really just deciding what goes in the window, and in what order.

Because the window is finite, everything in it is in competition: the system prompt, your instructions, file contents, tool results, and the agent's own earlier output are all spending the same budget. And a nearly full window is worse than the number suggests. Attention is not uniform across the context -- models track the beginning and end well and lose the middle, a failure mode the research literature calls "lost in the middle." A full context is not a well-used context.

Reasoning models

Earlier I said the model is not thinking. So what exactly is a "reasoning model" doing during all that thinking?

Nothing new, it turns out. Extended thinking is the same next-token machine given room to work: before producing an answer, the model generates tokens into a private scratchpad -- laying out the problem, trying an approach, backtracking when it fails -- and only then replies. No new mechanism has appeared; the model has simply been trained to spend tokens deliberating before it commits. The effect is real, though. On hard, multi-step problems models that think first reliably beat models that answer cold.

The catch is that tokens are the currency, so thinking is slower and costs more, and on mechanical work it buys you nothing. Hold onto this mental model: thinking is a budget dial, not a smarter brain. That is all the machinery you will need when later chapters discuss when to turn it up.

From LLM to Coding Agent

At the top of this chapter I claimed that a coding agent is painfully simple -- a language model in a loop with some instructions and a set of tools. You now know what the model half of that sentence really means: a next-token predictor with no memory and a finite window. This section supplies the other half. Everything wrapped around the model -- the loop, the instructions, the tools -- is called the harness, and the harness is where a text predictor becomes something that can do work.

The agentic loop

Strip away the products and the interfaces and every coding agent runs the same cycle. You send a prompt. The model responds with either text or a tool call -- a structured request to take an action, like reading a file or running a command. The harness executes the tool, appends the result to the context, and calls the model again. The loop repeats until the model responds with plain text instead of another action, and that response is your answer.

Notice what the model is doing in this loop: exactly what it did in a chat window, and nothing more. A tool call is not the model reaching out and touching your filesystem -- it is the model emitting tokens shaped like a request, because its training taught it that a request is the likely next move. The harness does all of the acting. The model remains, throughout, a text predictor; the loop is what converts prediction into behavior.

Notice also what the loop does to the context window. Every tool result -- every file read, every compiler error, every directory listing -- lands in the same finite budget from the previous section, and none of it ever leaves on its own. An agentic loop is a context-filling machine. That is the mechanical reason long sessions decay, and you now understand it end to end.

Anatomy of the harness

The harness itself has little mystery in it. Four parts cover the anatomy.

The system prompt is standing instructions the model receives before you type anything: what kind of assistant it is, how it should behave, what it should never do. When two agents with the same underlying model feel different to use, the system prompt is usually why.

Tool definitions are the menu of actions the harness offers: read a file, edit a file, search the codebase, run a shell command. The menu is text in the context window like everything else -- the model "has" a tool only in the sense that it has been told the tool exists and will be obeyed if it asks.

The permission model decides which of those requests run freely and which pause for your approval. This is the safety boundary between a model that suggests and an agent that acts: reading files is harmless, deleting them is not, and a well-designed harness knows the difference and makes you the judge of the rest.

And one tool quietly dwarfs the others: the terminal. A shell is the universal escape hatch -- it hands the agent git, compilers, package managers, test runners, and every other program with a command-line interface, for free, without anyone writing a dedicated integration. This is why terminal-based agents punch so far above their apparent simplicity. They do not have many tools; they have one tool that contains all the others.

Feedback loops -- why agents work at all

Hold the two halves of this chapter together and there is a puzzle. The model at the center of the loop is the same machine from the first half -- stale knowledge, trained agreeableness, plausible fabrication. Why does wrapping it in a loop produce something that reliably ships working code?

Because the loop closes. Chat is open-loop: the model guesses once, you paste the guess into your project, and you are the first thing that ever tests it. An agent is closed-loop: it runs the code, reads the error, fixes it, and runs it again. Every failure mode from the first half of this chapter -- the deprecated API, the hallucinated function -- stops being a lie you might believe and becomes an error message the model reacts to on the next turn of the loop.

Which leads to the fact this chapter has been building toward: an agent's effectiveness is bounded by the quality of the feedback its environment provides. A typed, tested, linted codebase is legible to an agent -- every mistake produces a visible error it can react to. An untyped, untested codebase offers nothing to react to; the agent flies blind, and its mistakes survive to reach you.

The same model, dropped into two different repositories, is two different agents. Codebase quality is now agent capability. Every investment you have ever been told to make in your tooling -- types, tests, linters, fast builds -- now pays a second dividend, and the rest of this book is about collecting it.