← Blog

The Model Context Protocol arrived a year and a half ago with a modest pitch: a standard way to hand an agent tools. Connect a server and the model on the other end grows hands. It can search your mail, file your tickets, read your calendar. The premise was so obvious it never read as a decision. Of course everything plugs into the one agent you talk to. Where else would it go?

Then came skills, instructions that stay out of the way until a task calls for them. Then subagents, whole other workers the main agent can hand a job to and stop holding. Each arrived as a feature announcement. Read in a row, they are also a pattern: every addition moves something out of the one context we started by filling.

My own setup tells the story in configuration. Every server I have ever connected went to the same place, the seat I talk to, until the agent that plans my week and the agent that fixes my frontmatter were the same agent, hauling the same few dozen tools to both jobs. And this week I was inside Fabrica, the agent team I built to never have to open, choosing a model for each subagent and noticing that the choosing did not want to stop at models.

Tools, skills, subagents reads like a roadmap. I think it is a rediscovery. Humans have hit this exact wall before, at the scale of whole societies, and what we built at the wall took a few thousand years to settle.

It settled into layers.

One agent was the right number

An agent, as I use the word, is a model plus a harness: the weights that do the thinking, and the system around them that turns thinking into work. From Prompts to Harnesses traced how the leverage kept sliding from the first half to the second. Take the best available model, wrap it in the best available harness, and you have the strongest single agent that can currently exist. That is what most of us are talking to.

And concentrating everything in that one agent is, at the beginning, simply correct. A startup at day zero works the same way: one founder doing sales, code, support, and invoices, with no handoffs, no meetings, no version of the plan lost in translation between heads. The whole picture lives in one place, and every decision gets made with all of it in view. Nobody at that stage says the company needs structure. The company needs the founder to keep going.

Most of the world is still in the earliest version of this, AI as a chat window, a search engine that answers back. People using it for real work have moved one step further, to the pair: one human and one agent, side by side on an actual task. Both arrangements put everything in front of a single head, and at their scale that is exactly where everything should be. I would rather have one generalist holding the whole picture than a committee holding fragments of it.

The arrangement fails by succeeding, which is what founders find out next.

Bandwidth breaks before capability does

The founder who is drowning at fifty customers is the same person who was brilliant at five. Nothing in their head got worse. The plate got fuller, and a full plate degrades everything on it: each task gets a thinner slice of attention, the switching between tasks eats the hours, and quality drifts down across the board while the founder works harder than they have ever worked. So they make the first hire, and the hire is another generalist, because at a company of two everyone wears every hat. For a while, sharing the plate is enough.

Context windows fail on the same curve. Feed one agent your codebase, your inbox, your tickets, and a few dozen tool definitions, and its answers drift the way a founder’s do. Attention thins, the middle of the pile goes soft, and a question that would have come back sharp in a clean session comes back vague. The ceiling keeps rising, and the longer window and the smarter frontier keep being genuinely better generalists. The industry’s standing answer to the full plate is to upgrade the founder.

Startups know where that answer stops working, because no founder is good enough to be a fifty-person company. Past some size the right hire changes shape: you stop searching for a better generalist and bring in a specialist, someone who does one narrow thing better than any hat-wearer does it. The specialist is a strange hire by day-zero logic. They know less about the company than anyone, they cannot cover for the others, and they make the team harder to hold in one head. They are also the only way forward.

Then, once there are enough of them, something else appears that nobody remembers voting for.

The layers grew back anyway

Structure. Hire seven specialists and a flat circle survives about a week. Someone starts deciding what gets worked on. Someone starts summarizing for someone else. Reporting appears, then reporting about the reporting, and the org chart grows its first layer while the founders still describe the place as flat.

It is tempting to read layers as leftovers from an older order. Ancient China ran on them, county reporting to prefecture, prefecture to province, province to court, court to emperor, with rank stitched into every exchange, and from here that looks like aristocracy doing what aristocracy does: a world of layers because it was a world of classes. Then the modern world removed the classes. Birth stopped assigning rank, the formal hierarchy of persons came down, and every society that did this promptly rebuilt layers inside every company, every agency, and every three-person startup founded by people who swear they hate management. China had even run the control experiment a thousand years early: for centuries its bureaucracy filled those ranks by examination, and the layers survived the removal of blood long before the rest of us tried it.

What the layers rest on never fell, because it is load-bearing in any century: one head holds only so much. Add work past that limit and you need more heads. Add heads and the conversations between them multiply faster than the heads do, until anyone trying to hear everything hears nothing. Layers are the shape that keeps each head’s load survivable: everyone faces a handful of others, compresses what happened below into what matters above, and trusts the level beneath them with the rest.

Hierarchy is bounded attention, drawn as a chart.

The compression has a price, and everyone who has worked under it knows the price by heart. Intent blurs on the way down; the task that leaves the top as a strategy arrives at the bottom as a sentence three retellings old, and the person executing it holds a dimmer copy of the goal than the person who set it. Societies never fixed that by sending everything back through the emperor. They fixed it at the door and at the helm: hire people who can carry intent without supervision, who weigh tradeoffs like owners instead of following scripts, and then steer from above, watching the course rather than the hands and stepping in when it drifts. The standard you hire against is what makes the layers safe to trust.

A context window is this same limit with a spec sheet. An agent’s attention is bounded the way a founder’s week is bounded, and it degrades the same way, quietly, in the middle of the pile. So agent systems have started re-deriving the old structure, and the fossil record is already public: tools, then skills, then subagents, each release moving a bigger unit of work out of the one head that used to hold everything.

Tools belong to the seats that use them

Follow the old structure one step further and it reaches the toolbox, which is where my own setup stops looking merely busy and starts looking miswired. Everything I have ever plugged in hangs off the top seat, the agent whose actual job has quietly become intent: turning what I ask for into work, and keeping the work pointed at what I asked. In a company, this is an org chart that assigns the forklift to the CEO.

Real teams keep the tools at the edge, and the arrangement is wiser than it looks. The person who runs a tool every day understands it in a way the manager who assigns the task never will, and no one considers that a failure of management. It is what makes delegation worth anything: the level below you knows its own terrain better than you do. An agent team should be wired the same way. The seat that reviews code holds the review tools and the bar to apply them. The seat that searches holds the search. The top seat holds the goal, and hands each piece of it down whole.

Context follows the tools. A seat should arrive knowing what its job needs and carrying little else, the way a new specialist gets an onboarding packet rather than the founder’s entire memory. The reviewer gets the diff and the standard it must clear. The researcher gets the question and the places to look. And the top seat, relieved of all of it, thins toward the thing only it holds: what I actually asked for, and whether the work still points at it.

Once the seats stop inheriting the top’s loadout, they stop needing the top’s engine too. Each seat is its own agent, its own pairing of model and harness, and the pairing can fit the job in the seat: a different vendor where independence matters, a smaller and cheaper model where the task is narrow. Whether a given seat can truly run on less is the question I could not answer last week, and structure is what changes its price. Nobody can demote the agent that does everything, because everything is on the line at once, but a seat that holds one job and three tools can be tried, watched, and judged without betting the rest of the team on the result.

The vendors have started conceding the point in their release notes. Tool definitions now load on demand instead of sitting in the prompt from the first message. Skills wait on disk until a task calls for them. The harness I am writing this in defers most of its own tools until asked. Piece by piece, the industry that spent a year and a half plugging everything into one context is taking it back out.

Fabrica goes first

This week’s model selection was the first piece of it. Each seat is getting its own model now, chosen for the job in front of it, and the choosing keeps pulling the rest of the reorg into view. The toolsets should be next: the review seat keeps the review tools, the servers I piled onto the top seat move down to the seats that actually run them, and the top seat gets to forget them all, keeping the goal, the standards, and enough of the picture to notice drift.

There is new work on the other side of this, and it is the work every founder eventually runs into: hiring. Before a seat gets a job, I need to be able to say what good looks like in that seat, the same way a team writes down what it interviews against, because the layers hold only if someone is careful about what gets let into them. I spent the spring shaping this team one failure at a time. The stretch ahead looks less like fixing and more like staffing.

The reorg costs me the thing I liked best about the solo stage. There used to be one conversation I could drop any question into, one seat that had seen everything and could answer for all of it. Nobody in the new structure will know that much, including the seat I talk to, and including me. Companies pay the same price on their way up, and what it buys them is everything a single head cannot do. Societies took a few thousand years to settle into that trade. I got there in a config file.