Chapter 3
Development Environment
In the introduction I promised to teach you to sharpen the axe. This chapter is the axe: the agent you choose, the extensions you give it, and the tools you let it reach.
Choosing a coding agent
When people compare coding agents, they instinctively compare models -- which one scores higher on the benchmark, which one "feels smarter." But you learned in the last chapter that the model is only half of the machine. The system prompt, the tools, the permission model, the loop: all of that belongs to the harness, and the harness is what you are actually shopping for. You are not choosing a model; you are choosing a harness.
The clearest proof is a product you already know: the web chat interface. A chat window contains a state-of-the-art model, and it is nearly useless for real engineering work -- because chat is a harness with no tools and no loop. You are the loop. You paste context in by hand, carry the output back to your editor, run it yourself, and paste the error back. Every feedback mechanism from the last chapter routes through your clipboard, and the window forgets everything the moment you close it. The model was never the problem. The harness was.
The serious contenders as I write this are Claude Code, Codex, and Cursor, and the obvious way to sort them -- by interface -- is already obsolete. Cursor grew out of an editor and now opens into an agent-native view by default, with the editor a click away; Claude Code and Codex grew out of the terminal and now ship GUI variants, with the terminal still running underneath. The surfaces are converging on the same shape because they are all skins over the same loop. What persists is what a spec sheet won't show you: the system prompt and its defaults, the permission model, the extension surfaces, and how well the agent wields the shell -- the escape-hatch property from the last chapter. There are others, and there will be more by the time you read this; judge them the same way.
Which one? The honest answer is that the choice matters less than the commitment. Every serious agent runs the same loop around a comparable model, and the capabilities converge within months of one another. What compounds is depth: knowing your agent's configuration, its permission model, its extension surfaces -- the things the rest of this chapter is about. Pick one and learn it properly. This book draws its examples from Claude Code because that is what I use daily, but every mechanism described here has an equivalent in the others.
Plugins, skills, and hooks
Out of the box, an agent is a generalist configured for nobody. The extension surfaces are how you make it yours -- how the corrections you find yourself repeating become configuration instead of conversation. Three mechanisms cover most of it, and they form a spectrum of enforcement: plugins bundle, skills instruct, hooks guarantee.
A plugin is the distribution unit -- a package of skills, hooks, and commands that installs as one piece. The ecosystem around them is young but already useful, and one entry is worth singling out: the superpowers plugin, a curated set of process skills -- brainstorming before building, systematic debugging, test-driven habits -- that amounts to an opinionated methodology for working with agents. You will not agree with all of it. Install it anyway; disagreeing with a concrete process is how you find your own.
Skills
A skill is a packaged procedure the agent loads when the task calls for it: a markdown file with a name, a one-line description, and instructions. The design is context-frugal by construction. Only the name and description ride along in every session; the full body enters the window when the task matches. Remember that everything in the context competes for the same budget -- skills are how you keep a library of expertise on the shelf instead of carrying all of it in working memory.
When should you write one? The second time you type the same explanation. If you have twice walked the agent through your release process, or your migration workflow, or the way your team writes commit messages, that explanation wants to be a skill. The model remembers nothing between sessions -- you know this from the last chapter -- so anything you want it to reliably know, you must write down. A skill is the memory the model doesn't have.
Hooks
A hook is code that runs at a fixed point in the agent's lifecycle: before a tool call, after a file edit, when a session starts. The distinction from a skill is the distinction that matters most in this chapter. A skill is instructions, and instructions are input to a probability distribution -- followed almost always, which is not always. A hook is code. It runs every time, unconditionally, because no model is consulted. Instructions are probabilistic; hooks are deterministic.
That tells you exactly when to use them: any rule with "always" or "never" in it. Always run the formatter after an edit. Never allow a push to main. Block any command matching a dangerous pattern. If you find yourself writing ALL-CAPS warnings in an instruction file, hoping a stronger adjective will shift the distribution -- stop. That rule is begging to be a hook.
MCP servers
The Model Context Protocol is an open standard for plugging external systems into an agent. An MCP server exposes a set of tools -- query this database, drive this browser, search these docs -- and the harness merges them into the agent's menu, right alongside reading files and running commands. It is the standard answer to "how do I give my agent access to X," and it works.
Use it reluctantly. Recall from the last chapter that the tool menu is not an ability the model has -- it is text in the context window. Every MCP server you connect pays its full menu of tool definitions into every session, whether the session uses them or not, and that spending comes out of the same finite budget as your instructions and your code. A directory full of connected servers is not a more capable agent; it is an agent with less room to think. The test for an MCP server is the test for any dependency: does it earn its weight? For most things it might do, a command-line tool does the same job for a fraction of the cost -- the next section makes that case.
Useful MCP servers
Two categories reliably earn their context.
Browser automation. An agent building a web interface without a browser is flying blind in precisely the way the last chapter warned about: it can change the code, but nothing shows it what the change looks like. A browser server -- navigate, click, screenshot -- closes that loop. The screenshot is to visual work what the compiler error is to logic: the feedback that turns a guess into a correction. Models read images now; let yours see what it built.
Documentation lookup. Servers like Context7 fetch current library documentation into the context on demand. This is the direct antidote to the pretraining problem: the model's knowledge of every API is a photograph of the past, and current docs placed in context override stale memory. "Read the docs before you code" is the highest-leverage instruction you can give an agent -- a docs server is what makes it cheap to follow.
The terminal is your toolkit
The last chapter called the shell the universal escape hatch. Here is the practical case for organizing your whole environment around it.
First, the model already knows your tools. Decades of man pages, tutorials, and Stack Overflow answers for grep, git, sed, and curl are in the training data; the model speaks fluent Unix without being taught. A CLI tool needs no definition in the context window and costs nothing until the moment it is invoked -- compare the MCP server paying rent on every session. Second, CLI tools compose. Pipes, flags, and exit codes let the agent build the exact command a task needs out of parts, instead of hoping someone shipped the right tool. When there is a choice between an MCP server and a CLI tool that does the same job, take the CLI.
The same reach is the danger. A shell that can run your tests can run rm -rf, force-push over your history, or pipe a stranger's script into bash -- and an agent that misreads a situation will do any of these with the same calm confidence it does everything else. This is what the permission model from the last chapter is for, and it is configurable: allowlist the commands you trust to run freely (the read-only ones, the test runner, the linter), require approval for the destructive and the irreversible, and use hooks to block outright the patterns nothing should ever run. For fully autonomous work, go further and change the blast radius instead of the rules: a container or a sandbox turns "the agent might wreck something" into "the agent might wreck a copy."
Useful CLI tools
A short list, because the agent will reach for these constantly and they should be there when it does. ripgrep (rg) is fast codebase search that respects .gitignore -- the difference between an agent that searches freely and one that hesitates is measured in seconds per loop iteration, paid hundreds of times per session. fd does the same for finding files. jq lets the agent slice JSON in one step instead of guessing at its shape -- which makes every JSON-speaking API a tool. gh puts the entire GitHub workflow -- PRs, issues, CI runs -- behind a scriptable interface, so "open a PR and watch the checks" becomes something the agent just does. The pattern generalizes: every well-behaved CLI tool you install is a tool call you didn't have to build.
The axe is sharp: an agent chosen for its harness, extended where you repeat yourself, guarded where it must never slip, and handed a terminal full of tools it already knows. The next chapter turns to the tree -- the repository itself, and how to structure one so that all of this capability has something to grip.