Jump to another pillar
Vertical agentic orchestration · the proposition
Why a vertical SaaS company should become the orchestration layer of its industry, which products this applies to, what actually gets shipped, and what each side gains from it.
Part one
The recurring piece of work: who meets, how often, over what scope, ending in what decision. Every organisation builds its own, shaped by its structure, its calendar and its politics.
The recurring problem inside the routine: a named class of decision with its own features, its own competing answers and its own way of being judged right or wrong afterwards.
One school of thought for handling that situation. Several exist, all defensible, none universally right. Expertise is not knowing one — it is knowing which applies today.
Worked example — one situation type
Two builds need the same scarce part, and there is not enough.
Five strategies, each one something an experienced practitioner would defend in a review. The disagreement between them is not confusion. It is the expertise.
And then a selector — the part that decides which one applies, given the features of the situation in front of it. A single agent has to be right. A taught team can hold five contradictory doctrines at once and be right about which one applies, which is what the senior practitioner actually does.
Where VAO does not applyProblems with one correct answer belong in a calculation, not an agent. One-off analysis belongs in a spreadsheet. A buyer who wants arbitrary agents over arbitrary data wants a horizontal platform and should be told so. The method is deliberately wrong for some products, and saying which ones is what makes the rest credible.
Part two
The practitioner who knows that this supplier always confirms and never delivers, that this part can be pulled from inspection, that this programme manager will accept a partial — none of that is in any system. It is visible only in what they do.
Why it persists — nobody has ever been asked to write down why, and there is no format for it
Two people facing the same situation on different shifts make different calls, and neither call is recorded with its reasoning. The better judgement stays where it started.
Why it persists — no mechanism propagates it, and no record captures what was chosen or rejected
A large share of the experienced workforce is close to retirement. Every departure is an uncontrolled deletion from a database nobody ever wrote down.
Why it persists — succession plans transfer roles, not judgement
Most enterprises now have something agentic in production. Almost all of it queries: a chat over the data, a better report. Nothing decides, and nothing acts, so nothing about the underlying problem has changed.
Why it persists — reading is safe to ship, and acting requires a governed write path nobody built
Internal low-code agents get built by a central team, then drift. The practitioners whose judgement they encode never owned them, so nobody notices when the encoded judgement stops being right.
Why it persists — the builder made composition cheap without making expertise transferable
Without a library of situations and the strategies that answer them, there is nothing to test against. Confidence in an agent becomes confidence in the person who wrote the prompt.
Why it persists — you cannot delegate authority to something you cannot evaluate
Large enterprises increasingly want a model specialised on their domain and running inside their own boundary. What they do not want is their operating data improving a closed model they do not control.
Why it persists — the vendors offering the strongest models offer the weakest answer to this
Enterprises have no appetite for re-qualifying their stack every time a better model ships. They want the benefit of it without owning the evaluation problem.
Why it persists — most architectures name the model in the agent, making every upgrade a rebuild
The situation gets resolved, everyone moves on, and the reasoning evaporates. What was predicted, what actually happened, and which of the two diverged — none of it is written anywhere, so the same problem next quarter is met with the same guesswork.
Why it persists — outcomes arrive weeks after the decision, long after anyone is still looking
What gets shipped
Three things their team opens. A chat window for asking questions of their own data. A library of ready-built agents that already know how their industry handles specific recurring problems. And a studio where their own people assemble those agents into the routines they actually run — schedules, sequences, reports — without writing code.Work arrives as a list of things that need deciding, each with an owner and a status, rather than as a conversation someone has to start. Before an agent changes anything it shows what it is about to change and how to reverse it. One screen shows, for every agent, what it is currently allowed to do on its own and why.
Outward: turn on the open protocol and the customer’s own agents can read from the system — look things up, compare, summarise, run a what-if. Genuinely useful, and also the ceiling of what reading gets you.Inward: anything that changes the state of the business goes through one path, and that path belongs to the product. It checks the change against the customer’s own rules before it happens, records who authorised it, and can undo it. Handing over the read side is easy because it was going to be commoditised anyway. The write path is where the product actually starts.
Open-weight models, fine-tuned on the accumulated taught decisions of the industry and never on any individual customer’s data. That distinction is the whole point: what the model learns is how this class of problem gets resolved, with every supplier, price and part number stripped out before it goes anywhere near training.Two ways they reach customers. Hosted by us, for customers who want the simplest thing that works. Released to run inside their own walls, for customers who cannot send anything out — the same model, on their hardware, under their control. Licence terms are still being settled.
A services team can build agents against your data. It cannot build the library, because it does not know the vertical.
The strategies, the selectors and the evaluation suites come from sitting with practitioners who have spent careers in one industry. That is the part a general-purpose AI services engagement structurally cannot produce — and the part that is worth owning.
Where it runs
Design for the last one first, even though it sells last. Everything that makes it possible relaxes into the other three; the reverse is a rewrite.
Part four
Every situation type taught at one customer is decomposed into skills, competing strategies, a selector and an evaluation suite — and that decomposition travels to the next customer while their data never does. It requires access to many operators' senior practitioners, a proving ground, and years of calibration to accumulate. Nobody assembles that quickly, and because no agent names a model, it survives both model churn and a sovereignty mandate intact.
Taught judgement, once collected, is training data of a kind that does not otherwise exist — the reasoned choice between defensible options, with the outcome attached. That is what makes a domain-specialised model possible without ever training on a customer's operating data.
Deploying scope and teaching judgement are two different things and should be priced as two. Composition, scheduling, reporting and chat are included at every tier — pricing table stakes invites a line-item comparison you lose. What is priced is promoting a customer's candidate into a taught situation type, raising a delegation level, and adding an enforcement path. Maintaining the graph is what makes it stick, because the graph is the thing that decays without attention.
The strategies their best practitioners actually apply, written down, versioned, attributable and auditable — surviving retirement, reassignment and reorganisation. The strategy library doubles as the curriculum: a new joiner reads why, not just what, and the system explains its own reasoning as it works.
The same doctrine applies across sites and shifts, with local departures recorded as revealed preference rather than lost. Consistency stops depending on who is on duty.
Strategies do not sit in a document that ages. They sit in a graph that records what was predicted, what happened, and what that implies for next time — so each disruption survived improves the response to the next. And the graph exports in an open schema, because removing the largest procurement objection costs less than defending it. The lock-in was never the data.
Because no agent names a model, adopting a better one is a configuration change and a qualification run. The enterprise gets each improvement without re-platforming, mandates its own model where procurement requires it, and its operating data never trains anyone's weights — stated as a line in the pricing table rather than a clause in the contract.
Nothing starts unsupervised. Authority is granted per situation type and per confidence band on measured evidence, shown on a console, and withdrawn automatically when calibration degrades. For a regulated buyer this is not a feature — it is the artifact that gets the programme through review.
What is not claimedNot full delegation on day one — everything starts in shadow mode, recommending only. Not replacement of the practitioner; arbitration and negotiation stay human. Not a chatbot over the graph; conversation is a feature, and executed decisions with recorded rationale are the proposition. And not that the strategies are optimal — they are what your best people do, made explicit, measured and improvable. That is a stronger claim commercially and a weaker one to defend technically, which is the right way round.