I used to think if I built the thing, the rest would follow.
We built it. We even signed an LOI. Then the champion left and the pilot disappeared with him, and the work at that company looked exactly the same on Monday. I had a product. I did not have an outcome. I just didn't have the words for that yet.
I thought that was just us being early, or small, or unlucky. Then I started reading, the way I always do, past the surface answer and into the papers, and it turns out this is not a startup story. It is the default.
There is a working paper from last year, Project NANDA, out of the MIT Media Lab. Not a peer reviewed MIT verdict. Their own interviews and a review of a few hundred public deployments. Companies poured something like $30B to $40B into generative AI. About 95% of the organizations they studied were getting zero measurable return on the P&L.1 Not that the chatbot didn't work. That nothing showed up in the numbers. ChatGPT and Copilot were everywhere, people felt faster, and the income statement did not move. Custom tools were worse in a different way. 60% of companies evaluated them, 20% made it to a pilot, 5% made it into production on their definition. The authors said the divide was not model quality. It was brittle workflows, no memory of context, and a mismatch with how the work actually runs.
I am not going to pretend 95% is a sacred number. A lot of those pilots never had a baseline, so "no measurable impact" is what you get when you never set up the measurement. That is kind of the point. I know because I have shipped things I could point at and still not tell you what moved.
What I actually mean
I don't want this to turn into a framework post. I never liked those. I needed a way to think about this that wasn't "we should use AI," because that sentence has stopped meaning anything to me.
McKinsey's latest global survey, about two thousand companies, says the same thing in a less dramatic way. About 88% are using AI in at least one function. Only about 39% say it shows up in EBIT at the company level, and most of those say it is under 5%.2 The ones they call high performers, around 6%, are about three times as likely to have redesigned the actual workflow. In an earlier pass they checked something like 25 things that might explain who gets value. Workflow redesign had the biggest effect. Only about 21% of the companies using generative AI had redesigned any workflow at all.
RAND did the version of this I trust more, because they sat down with the people who actually build it. Sixty five interviews, fifty of them industry data scientists and engineers with at least five years in. 84% of those builders said the project died because of leadership, not because the model was dumb.3 The team got pointed at the wrong problem. The model got optimized for a metric that did not match the work. Or the thing never fit the workflow it was supposed to live in.
Gartner has been saying a quieter version since 2024. They predicted at least 30% of generative AI projects would get abandoned after the proof of concept by the end of 2025, because the business value was unclear, the risk controls were not there, or the cost made no sense once you left the demo.4 That is a forecast, not a body count. Unclear value and no controls. That is the whole problem in one sentence.
Put those next to each other and the pattern is not mysterious. Companies start at the intervention. They pick a model, find a place to put it, and treat deployment as the result. I did that too. AI Outcome Engineering is a name for reversing the order.
The object is a chain:
Outcome → workflow → baseline → intervention → controls → measurement
The order is the method. Reversing it is how most teams work, and why so little of it compounds.
The order
Outcome. A state that has to be true in the business after the work is done, said without the word AI. Payouts leave on time with fewer disputes. A lead is touched before it goes cold. "We deployed an agent" is an activity. RAND's builders kept describing models that were technically fine and pointed at the wrong thing. If you cannot say the outcome in one sentence, you do not have a project yet. You have a slide. I know because I have had slides.
Workflow. The path that currently produces, or fails to produce, that outcome. People, systems, handoffs, the spreadsheet that is the actual process. At my internship I couldn't stop seeing it. Spreadsheets that didn't line up. Numbers that still drove decisions. Everyone treated it as normal. MIT's writeup of why the 95% happened is not "the models were bad." It is brittle workflows and a mismatch with how the work actually runs day to day. McKinsey's high performers are the ones who redesigned the workflow. Same finding, other direction.
The workflows that are actually worth touching are not the ones that sound good in a pitch. They are the ones where someone still sits down on Monday, reads a pile of information, makes a judgment, and takes an action. Slack surveyed knowledge workers and found most of them burn a chunk of every day just switching between apps.5 Asana called the rest of it "work about work" and said it ate more than half the day.6 That is the glue. Agents are not competing with the system of record. They are competing with that.
Baseline. Measure it before anyone writes code. Labor, latency, errors, rework, missed revenue, risk. Taken while the work is still ugly. After Ryft I got obsessed with the gym because it was the one place the numbers didn't lie. Small, ugly, honest. You either moved or you didn't. When you want something to work you start seeing what you want to see. No baseline means you will talk yourself into a win. I have done that too. If you never took a baseline, "no measurable P&L impact" is the default, even if the tool felt useful.
Intervention. The smallest change that could move the outcome on that baseline. That's hard for me. I don't do things halfway. My instinct is to build more, go deeper, turn it into the whole company. That's how I got into trouble last time. Sometimes it's an agent. Sometimes it's a rule. Sometimes it's deleting a step that never needed to exist. RAND's people kept describing teams that chased the newest model instead of the problem. The cut should be small enough that if it fails you can actually see it fail.
Controls. What happens when it's wrong. Who looks. What it's not allowed to do. I used to skip this because it felt like I didn't believe in the thing. Now I think the opposite. If you won't put a kill switch on it, you don't believe in it enough to let it near real work. You just want the demo. Gartner listed inadequate risk controls next to unclear business value as a reason these projects get killed after the demo.4 That is not a legal footnote.
Measurement. Same workflow. Same numbers. Did the outcome move. If it didn't, you shipped software. That's fine. I've shipped software. Just don't call it transformation. Don't switch the metric to make it look like a win. That's the same thing as telling yourself a better story when the champion already left. McKinsey's 39% who report any EBIT impact, and the 6% who report a real one, are the gap between "we used AI" and "it showed up in the business."2
The honest part
I started calling this AI Outcome Engineering because I needed a name for not starting at the model. I know how it sounds. Like a consultancy made it up. I'm not going to dress it up into something cleaner than it is.
This is not a product. It is a way of working. Product work is separate. The claim is narrower, and it is the claim the research already supports: if you cannot name the outcome, map the workflow, show the baseline, describe the cut, name the controls, and measure again, you did not engineer an outcome. You decorated a process.
I don't have this fully figured out. I'm writing it because that's still how I tell if I understand a thing. In my head it felt clear. On the page, some of it still doesn't.
I used to start at the product. I'm not doing that anymore.
I'll let you know if that's enough.
References
- MIT Project NANDA: The GenAI Divide, State of AI in Business 2025 (working paper)
- McKinsey: The State of AI, Global Survey 2025
- RAND: The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed
- Gartner: 30% of Generative AI Projects Will Be Abandoned After Proof of Concept by End of 2025
- Slack: State of Work, on time lost toggling between tools
- Asana: Anatomy of Work Index, on "work about work"