In the Cask years, when a customer's data platform caught fire, our engineers went to the war room, and I learned to love what happened there, because it was the closest thing to honesty an enterprise ever showed me.
An incident does not care about your org chart. There is an alert, a topology, a recent deployment, a log trail, a rollback, a recovery graph. Somebody makes a call, and within minutes reality grades the call. I have sat in those rooms at three in the morning and felt something almost like relief: for once, we would actually find out if we were right. Then the incident would close, I would walk back into the part of the company where decisions took months to grade themselves, and the contrast would physically bother me.
That contrast is this part's whole subject. The loop from Parts 1 through 4 does not close evenly across the enterprise, and the honest way to rank opportunities is not by how impressive an agent demo looks. It is by how closable the feedback loop is.
three engineers and a war room
Score any workflow on six dimensions: observability of state, outcome speed, outcome clarity, repeatability, reversibility, and confounding. Call the composite loopability. The ideal workflow has highly observable state, many repeated cases, clear intermediate outcomes, feedback in minutes or days rather than months, reversible interventions, and limited hidden confounding. War-room incidents score near the top on all six, which is exactly why they felt so honest.
loopability, a scoring function
Notice what this ranking does to conventional wisdom. The glamorous applications, the AI strategist, the AI executive coach, sit at the bottom, useful mainly for precedent retrieval. The boring ones, invoice exceptions, support escalations, incident response, sit at the top, and they happen to be where enterprises bleed the most money in the most measurable way. Process mining grew up on exactly those purchase-to-pay and order-to-cash logs for a reason.1
three loops up close
Customer support, the almost ideal loop. The system knows the state: enterprise account, billing issue, second contact in seven days, sentiment deteriorating, prior refund rejected. The intent: resolve without escalation or churn. The decision: refund, credit, supervisor, or troubleshoot. The action: credit offered. The immediate result: accepted. The delayed outcomes: no reopen in 14 days, account active. Repeat over hundreds of thousands of cases and patterns surface, like "cases with these characteristics reopened 35% less often when a supervisor joined before the third interaction." That starts as descriptive precedent, becomes a recommendation, acceptance becomes feedback, controlled rollout estimates real effect. Analytics, prediction, recommendation, experiment, bounded automation. A true closed loop, in that order.
IT incidents, closest to the coding loop. Alert, topology, recent deployments, logs, prior incidents, chat coordination, executed commands, rollback, service metrics, recovery, recurrence. The analog-digital distinction nearly disappears; the war room was a test harness all along. The system learns motifs: symptom cluster, diagnostic action, state change, candidate root cause, remediation, recovery or not.
Financial operations, boring and lucrative. When invoice exception type X appears for supplier category Y, which resolution sequence historically closes it with least rework? No causal model required to retrieve that precedent. These workflows will produce reliable closed-loop AI years before the glamorous ones.
sales, valuable and dangerous
Sales may be the highest-value loop of all, because the exhaust is rich: calls, transcripts, CRM, email, proposals, security reviews, usage, pricing approvals. An opportunity decomposes into dozens of state transitions, stakeholder identified, problem confirmed, technical objection, security blocker, procurement, legal, close, and at every transition you can capture what the seller decided and what happened next. The immediate value is organizational memory: show me deals that looked like this one at this stage, what the best sellers did next, and what happened after. Pattern-based, and comparatively safe.
The dangerous leap is "offering a 7% discount causes this type of deal to close." I lived this at Cask, and Part 3 named the trap: reps do not discount randomly, they discount deals with specific latent properties. Sales needs much stronger treatment of confounding before autonomous pricing or negotiation. Assistance now, autonomy later, and only with experiments in between. Even fuzzier knowledge work still yields motifs worth learning, requirement ambiguity leading to a clarifying question, a scope cut, and less downstream rework; organizational-routine research has long shown complex work contains recurring behavioral sequences rather than pure improvisation.2
your own regression suite
Here is the payoff that I think matters most in the long run. Companies today cannot tell whether an agent is good in their environment; generic benchmarks measure the model, not the fit. A sufficiently rich decision-outcome graph automatically generates enterprise-specific evals: test the agent at actual historical decision points. Given what was known on March 12, what would you have recommended? Compare against what the human did, what happened immediately, what happened eventually, what similar cases did, and what experts consider reasonable. Not causal proof, but a far stronger evaluation substrate than trivia benchmarks, and over time online experiments add the causal layer. The company ends up owning a regression suite made of its own experience, one that grows with every decision it makes.
the autonomy ladder
And that changes the autonomy debate entirely. Without outcome telemetry, organizations face a binary: human does it or agent does it. With the loop, autonomy becomes continuous:
A: retrieve precedents B: summarize likely outcomes C: recommend action D: execute after approval E: auto-execute when confident and reversible F: autonomous, exceptions escalate to humans
An action that occurs thousands of times, has predictable consequences, and is cheap to reverse climbs the ladder. A unique, high-stakes, irreversible action stays human. The Stanford AI Index supplies the sobering context for why this matters: even the best computer-use agents still fail roughly one in three attempts on the OSWorld benchmark.3 The enterprise should earn autonomy from measured outcomes, not grant it because a model seems intelligent.
So the architecture is drawn, the mess is mapped, the industry is converging, and the league table says where to start. One question remains, and it is the one I get asked most when I talk about this: if you were starting from a blank page, what exactly would you build first? Not the diagram. The product. The first ninety days. I have an answer, it fits in one phrase, and it is not "a company brain," because remembering and learning are not the same thing, and the difference is about to become a category.