When Do You Actually Need a Frontier Model?
My default model is the most capable one money can rent, and every task I have goes to it, the hard ones and the trivial ones alike. In the past few days that has meant a frontmatter fix in this site’s repo, a fresh set of tags on a note whose contents I already knew, and a git incantation I could have found with one search. Each of these went up to a frontier model, the same tier I reach for when I design a system from nothing, because reaching for anything else did not occur to me.
A month ago I published a rule for exactly this. Own the Memory, Rent the Brain argued that only the genuinely hard reasoning ever needs to leave for the frontier, and that the rest, the lookups and the tidying and the routine glue, “a small model on my own machine handles all of it.” I wrote that in the present tense. Handles.
It has handled nothing. The local model was never set up. The sentence described a machine that did not exist in the grammar of one that did, and I believed it as I typed it. Every routing decision in the month since has gone the same direction as every one before: up.
The rule never had a chance, and I wanted to know why.
Nobody made this decision
The easy story is weakness of will: I wrote a rule and could not hold myself to it. But that story needs a moment of choosing, some point where the rule said down and I said up, and I cannot find one. There was never a decision to break. There was never a decision at all.
Fabrica, the small agent team I run, makes this checkable, because unlike my habits it has a repository I can read. Every seat on that team runs a frontier model, and nothing in the configuration says so: no model is named anywhere. Each seat simply inherits whatever its vendor ships as the default, and the default, at every vendor, is the top of the line. I even built the review script a flag for handing its seat to a cheaper model. It has never been used.
Fabrica is just where the pattern is easiest to see. Install any of the coding tools this year and the model in your hand, before you have expressed a single preference, is the best one the vendor sells. The smaller tiers exist on the pricing page and almost nowhere else in an ordinary working day; no tool ever routes you toward them. The frontier arrived in my life as a setting someone else had already chosen. The rule I wrote in June was asking me to overturn a decision that nobody, including me, had ever actually made.
Deciding would have chosen the same thing
Seen that way, the fix looks simple: make the decision now, on purpose. Take one task, weigh the options, choose. So I ran the exercise on the smallest task from this past week, the note tagging, and watched deliberation land exactly where the reflex lands.
Start with money, since the rule was supposed to be about allocation. At current list prices, a frontier call for a tagging job of that size costs about three cents; the same call on the cheapest capable tier, about half a cent. The spread is a factor of five between two numbers too small to feel. No behavior has ever been disciplined by two and a half cents.
What routing down actually costs never appears on the price sheet, because it is work. Pick the smaller model. Watch its output for a week. Build the checks that catch failure modes I do not yet know it has. The frontier path is one call that already works; the cheap path is a small engineering project. Three cents do not buy back an afternoon of watching.
And underneath the work sits the load-bearing reason, the one I had to admit when I tallied this honestly: trust. In a month of sending everything up, the default has not burned me once. No confident nonsense caught a week late, no quietly wrong assumption shipped. Every task has come back the way I expected, every time, for pennies. Against that record, choosing anything else stops looking like thrift and starts looking like unpaid risk.
So each decision, taken alone, is correct. That was the uncomfortable end of the exercise: a thousand individually reasonable calls, every one of them pointing up, with nothing anywhere for discipline to grab.
The evidence only flows one way
Each call is correct, and the sum of them is a trap.
Watch what each route sends back. A task routed up succeeds quietly. The output is good, the cost is invisible, and nothing about the transaction ever suggests the task should come back down. A task routed down announces every failure: the wrong tag, the mangled summary, the fix that does not compile. Up teaches me nothing. Down teaches constantly, at the price of cleaning up after the lessons.
So the evidence moves in one direction only. Tasks drift up freely, and nothing ever pushes one back, because the only signal that could, a visible failure above or a visible success below, is exactly what the arrangement never produces. Instead of converging on the right allocation, the routing table freezes at the top.
Trust rides the same current. The frontier earns more of it every day because the frontier gets every rep; the small model has never been handed a single task, so it has never had a chance to demonstrate anything, and so it never will. I said I cannot predict what a smaller model would get wrong, and that is true, and it stays true forever, because prediction takes data and the data is never allowed to exist.
Here is what one week of demoted note tagging would send back: a pile of outputs, some wrong, each wrong one a labeled break point, a map of where that model’s floor actually sits. What the current arrangement has sent back since June is what it will send back next year. Nothing.
The list I would write is all guesses
None of this stops me from having an answer to the title question. Ask me today when a frontier model is genuinely necessary and I will produce a considered-sounding list: synthesis across many sources at once, writing that has to carry a voice, coherence over a long stretch of context, judgment calls with no checkable answer. I believe every item. I also notice it is the same list I would have written in June, unchanged by a month in which no evidence arrived to change it.
Press on one entry. Writing with voice is supposedly the clearest case for the frontier, and it is the case I exercise most: these posts are drafted with a frontier model steering by a style guide distilled from everything published here before. Whether a smaller model holding the same guide would produce a passable draft, I do not know. I have never tried, not even once as an experiment. The belief has all the confidence of experience and none of the experience.
Every entry has the same standing: plausible, untested, inherited from the reflex it exists to justify. The title of this post reads like a knowledge question, the kind you settle with a list, and I have been carrying the list around as if it were the answer. From where I stand it cannot be answered at all. I do not know when I actually need a frontier model, and the way I work guarantees I never find out.
The reflex is starting to compound
For one person, this would stay a private quirk. My month of reflexive routing wasted a few dollars, and a few dollars buy a lot of not having to think about it. But I do not stay one person. I build systems, and systems inherit my defaults without inheriting my hesitation. Fabrica came out frontier in every seat without one deliberate choice, and a small pull request now passes through a dozen frontier sessions on its way to merged: drafted, debated, written, reviewed, revised, reviewed again. The reflex, multiplied by a team that never sleeps and never asks whether the task deserved the model.
The consequences have started arriving from every direction at once. In the past few weeks I have hit my coding subscription’s session cap mid-task, watched a Fabrica run stall partway through an issue waiting on quota, and seen the reviewing side throttle against its own vendor’s plan in the same stretch. The bill has the same slope: one plan became a bigger plan, a second vendor’s plan joined it, and the metered usage on top climbs month over month.
The number itself is survivable. The allocation it buys is the problem: frozen since June, and now running at team scale. The fix I keep circling is reallocation: the same money doing more work, aimed by evidence instead of reflex. Under a cap, the arithmetic finally has teeth. Every frontier call burned on glue is a call the team cannot spend on work only the frontier can do, and unlike the three cents, a throttled afternoon is a cost I have already felt.
The answer I could earn
There is a version of the answer that would hold, and it takes the shape of a procedure. Change one model string in one seat. Watch a week of outputs. Write down where the smaller model breaks, and let the break points write the list. Repeat, seat by seat, until the routing table is made of evidence instead of inheritance. Whatever list came out the other end would be short, checkable, and mine, and I cannot predict a single line of it from here. That is the point of running it.
I can also name what the procedure asks. It asks me to tolerate failures I currently pay three cents at a time to never see. It asks me to spend attention, the one genuinely scarce input in all of this, watching a cheaper model be wrong in ways I then have to catch and catalogue. And it asks me to trade the most comfortable feature of the present arrangement, a default that has never once been wrong, for a stretch in which things go wrong on purpose.
What I will not do is announce the experiment. I have learned to distrust the announcement: in June I wrote a routing rule in the present tense and felt, by writing it, that the routing had begun. The sentence did the work of the wiring for a month, and none of it was real. This post is made of the same material as that sentence, so the claim stays small enough to be true. The question in the title has an answer, and the answer has a price. So far, three cents at a time, I have been paying not to learn it.