Writing ยท August 2026

What the model doesn't know about your company

The pilot always works. Then it sits for nine months. What's in the gap is context, permission, and accountability, and only the first one is an engineering problem.

The prototype always works. That's the confusing part. Someone spends a weekend with an agent framework and comes back with a thing that drafts the copy, or triages the queue, or answers the question nobody could answer before. It's good. Everyone in the room agrees it's good.

Then it sits for nine months. Not because anyone killed it. Because the distance between a thing that works and a thing you can put in front of customers turned out to be almost entirely non-technical, and nobody had budgeted for that.

The ratio, not the gap

The gap between demo and product isn't new. Innovation labs died of it a decade ago, and the postmortems all said the same thing. What's new is the ratio.

The cost of producing a credible prototype has fallen by something like two orders of magnitude in three years. The cost of productionizing one has barely moved. So an organization that used to generate four prototypes a year and ship two now generates sixty and still ships two. The bottleneck didn't get worse. It became the only thing that matters, and it did that fast enough that most operating models haven't caught up.

Which means the interesting question isn't which model to use. It's what's actually sitting in the queue.

It doesn't know your company

This is the part everyone has already worked out, so I'll be brief. A frontier model knows the world and knows nothing about your rate card, your CMS, your content taxonomy, the four systems that hold customer state, or the reason two of them have disagreed since a migration in 2019. Retrieval helps. Well-described tools help more. Someone has to write down the institutional knowledge that currently lives in six people's heads.

It's real work, but it has a known shape and a known owner. If context is the only thing between you and production, you're in better shape than you think.

It isn't allowed to touch anything

The prototype ran against a copy of the data on somebody's laptop. Production runs against the real system, and the real system comes with a legal review, a security review, a data-handling decision, an accessibility standard, a privacy posture, and in a media company an editorial standard that predates all of this and does not care that the sentence came from a model.

Each of those reviews is reasonable. Collectively they're a queue. The queue is serial, it's staffed by people with other jobs, and almost nobody is measuring it.

So measure it. If I could get one instrumentation change into an enterprise AI program, it wouldn't be model evals. It would be a clock on how long a working prototype waits, and for what. Run that for a quarter and you'll find the same three or four reviews blocking most of the portfolio. That list is the roadmap for whoever owns the shared platform, and it will look almost nothing like the roadmap they wrote for themselves. Central engineering teams build from their own backlog because that's the only demand signal they have. Give them a better one.

Nobody's name is on it

This is the one that actually stops things, and it's the one nobody writes about, because it isn't a technology problem and there's no vendor selling a fix.

An editorial assistant writes a sentence that's wrong. A recommendation surfaces something it shouldn't. A summary flattens a nuance that mattered legally. Who is accountable? Not in the policy-document sense. In the sense of: whose performance review has this on it, who gets the call at 9pm, who decides it's fine to ship at 94 percent when the standard for human work was never actually measured.

Until that question has a name attached to it, the thing does not ship. It doesn't matter how good the retrieval got.

And this is the part I'd push back on in my own thinking, and in most of what I read about enterprise AI right now. Context is treated as the hard problem because context is the part we know how to build. But context is buildable. Guardrails are buildable. Accountability has to be assigned, and assigning it costs someone political capital, which is why it keeps getting deferred into a working group. Governance is the word we use for this. It's a bad word for it, because it sounds like a document. It's a name.

The shared platform, and why it usually fails

The obvious response is to build the thing once. Approved models, the internal knowledge already wired up, reusable components, guardrails already cleared, the design system, a deployment path that's already been blessed. A team assembles instead of rebuilding, and the dozen reviews happen once instead of a dozen times. Do it well and the second product costs a fraction of the first.

I think that's right. I also think it's the most-attempted and most-failed pattern in enterprise engineering, and it's worth being precise about why.

Central platforms fail by mandate. The pattern doesn't vary much: the platform is declared the approved path, its abstraction lags what a product team needs this quarter, the team routes around it to hit a date, and a year later you have the platform plus six shadow implementations plus a standing governance fight. The mandate produced exactly the fragmentation it was built to prevent.

The only version that survives is the one teams choose, and that's a much higher bar than approval. It means the platform has to be the fastest way to ship, measured honestly, including on the day somebody needs a thing it doesn't do. It means the unit of reuse is a working application a team can fork and mangle, not a framework they have to study first. A golden path can be abandoned halfway. An abstraction can't, which is why people won't start down one.

I built shared platforms at Viacom when the customer was a handful of networks rather than a portfolio of brands. Same physics, smaller numbers. The teams that adopted did it because the platform got them to launch faster, and the ones that didn't had a reason I'd have found persuasive in their chair.

One note on how you fund it

The easiest case to make for a shared platform is efficiency. Less duplicated engineering, fewer overlapping vendor contracts, consolidation savings. Make that case and you'll get funded the way efficiency projects get funded, which means you'll be defending the line item in eighteen months.

The stronger case is reach. In a portfolio, a capability built once has a marginal cost approaching zero by the third deployment. The small brand and the international market, the ones that are structurally always last in line for anything central, get the capability in the same quarter as the flagship instead of three years later or never. That's not a savings story. It's a coverage story, and coverage is worth considerably more.

Where I'd start

The models are fine. They've been fine for a while, and they're improving faster than any large organization can absorb, which is its own problem and a different post.

If your AI program is stuck, I'd stop evaluating models and go find the queue. Put a clock on it. Name the four things that block most of the work, and name the person accountable for each one. The answer will be a little embarrassing, and it will also be the roadmap.