← all notes

What beats a well briefed domain-specific agent

july 2026 · 3 minute read · ai, speculation

Every specialist I run is a generalist in costume, told not to use the rest of it's knowledge.

Every specialist agent in my fleet is a generalist in a costume. I write it a charter, take away the tools it shouldn’t have, give it a name and point it at one job. Underneath the costume sits the same model that would cheerfully write the invoice, the database migration and the tone-of-voice check if I let it, and doesn’t much mind which one you ask for.

I still hold that narrow beats broad. I have run it both ways and the fleet is better for it, but I have only ever tested that claim in one direction: two prompted things, the same underlying model, briefed differently. It’s a well-briefed model against a vague one, and nobody should be surprised which won.

The next real cut might not involve a prompt at all. A smaller model, trained rather than briefed, that has never seen anything outside its one job. Something built only to judge whether a client email matches a brand’s tone of voice doesn’t need the ability to re-write a codebase. It needs good judgement about one narrow thing, and nothing it picked up elsewhere to get confused by.

The distinction worth carrying is this: a briefed specialist has breadth it has been asked not to use. Push it slightly outside its lane and it will usually answer, fluently, because everything else it knows is still sitting right there. A trained one has nothing else to fall back on - Its narrowness is a property rather than a policy, and properties don’t drift when a context window fills up.

A briefed specialist is asked to stay in its lane. A trained one has no idea there are other lanes.

Briefing an agent takes an afternoon: a charter, a scoped set of tools, a model choice, and it is running. Training one is a different discipline entirely. A dataset, a training loop, an evaluation harness pointed at a model’s weights rather than a prompt, and a rollback path for the day it is confidently wrong in front of a client. None of that exists in my studio currently, and the expensive kind of optimism would be assuming the training loop is the hard part, when the harness that decides whether the output came out right is the piece I would actually struggle with.

What does exist, almost by accident, is the raw material. Every agent transcript, every task’s acceptance criteria and whether the check actually passed, every activity entry, all sitting in the same database. I built that for audit and for my own memory. It also reads like a labelled dataset for a dozen narrow jobs, which begs further pursuit.

Google now ships a 270 million parameter model built for nothing except being fine-tuned onto a single narrow job, a classification, an extraction, a compliance check, and what comes out the other end is an adapter file of a few megabytes. NVIDIA’s researchers have made the general case in print, that agents spend most of their time doing a small number of repetitive things and that a small model serving those costs roughly an order of magnitude less than a frontier one.

What I am short of is not compute or a training loop but a written definition of correct for one dull job, with well thought out evaluations. Until I can write that down, every specialist I run is still a costume, and I am still the one who chose not to look too closely at the seams.