The Co-Pilot Zone: Replace, Augment, or Refuse
Every workflow in your company belongs in one of three zones. Almost nobody makes the call out loud, so everything defaults to the middle one, the only answer that requires nothing to change.
Every workflow in your company belongs in one of three zones. Almost nobody makes the call out loud. So everything defaults to the middle one, which is the only answer that requires nothing to change.
"Will AI replace people or augment them" is not a prediction about the future of work. It's a decision. And you're already making it... workflow by workflow, by not making it.
Here's the claim, plainly. Every workflow in your company belongs in exactly one of three zones. Replace. AI runs it end to end; a human defines the input and accepts the output. Augment. AI prepares, a human decides. Refuse. AI does not touch it, deliberately, and that's written down.
Most companies never sort anything. So every workflow lands in the same undeclared middle: a human still in the loop, an AI tool bolted alongside, nobody able to say which of the two owns the outcome. That's the co-pilot default. It feels prudent. What it does is preserve the exact structure AI was supposed to make unnecessary. And it hands you a faster version of the org you already had.
That's the 3-Zone Co-Pilot. Three zones, one explicit call per workflow, and Refuse treated as a real destination rather than a confession.
The co-pilot default is a decision nobody remembers making
Microsoft's 2026 Work Trend Index, 20,000 workers across ten countries, fielded February to April 2026, found that 86% of AI users say they treat AI output as a starting point, not a final answer, and that they "stay responsible for the thinking."
Read that as a virtue and it's reassuring. Read it as a distribution and it's alarming: the same posture on nearly every task, in nearly every function, regardless of the task. One relationship, universally applied, isn't a strategy. It's the absence of one.
The same report: only 13% of AI users say they're rewarded for reinventing work with AI even if results aren't met, and only 26% say their leadership is clearly and consistently aligned on AI. And organizational factors account for more than twice the AI impact that individual factors do: 67% versus 32%.
So: an organization that never decided, employees defaulting to caution because caution is the only thing never punished, and a leadership team that hasn't drawn a line anyone can see. None of that is a technology problem. It's the shape of an unmade decision. And it shows up the same way in every AI-augmented team that bought before it decided.
Now put a second number next to it. Anthropic's Economic Index found that about 77% of business API usage shows automation patterns, the model completing a task with little back-and-forth, against roughly 50% for consumer Claude.ai use. When companies build with AI deliberately, designing a workflow around it, they choose Replace far more often than the public argument implies. The co-pilot posture is the chat-window habit.
The gap between 77% and 50% is the gap between deciding and defaulting.
What Klarna actually proved
The most-cited AI reversal of the last two years is worth reading properly. Almost everyone draws the wrong lesson from it.
On 27 February 2024, Klarna announced that its OpenAI-powered assistant had, in its first month, handled 2.3 million conversations, two-thirds of the company's customer service chats, doing "the equivalent work of 700 full-time agents." Resolution time fell from 11 minutes to under 2 minutes. Repeat inquiries dropped 25%. Customer satisfaction scored on par with human agents. Klarna estimated a $40 million profit improvement for 2024.
Then, in May 2025, Klarna reversed. CEO Sebastian Siemiatkowski told Bloomberg on 8 May 2025: "As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality." The company began recruiting human agents back and promised customers a human would always be available.
The market read this as: AI hit its ceiling. That's the wrong read, and an expensive one. What was oversold wasn't the technology. It was a classification.
A support queue is not one workflow. It's at least two.
"Where is my refund?" Known information, routine generation, answerable from records. Zone 1.
"I've been charged three times, I'm furious, and I'm about to close my account." Judgment under ambiguity, a relationship at stake, context the model does not hold. Zone 2.
Klarna's first move classified the entire queue as Zone 1, because the entire queue arrived down one channel. Channels are not zones. And note the detail that survived the reversal: the AI stayed on the bulk of chats. Klarna didn't retreat from AI. The boundary moved. And a boundary you discover through customer anger is a boundary nobody drew on purpose.
That's the pattern worth taking away. Zone errors don't announce themselves as zone errors. They present as AI failing.
Zone 1: Replace, and the human who's there "to be safe"
Zone 1 is work where the human's contribution is coordination, synthesis of known information, or routine generation. Meeting summaries. Status reports. First-draft proposals. Lead routing. CRM hygiene. The AI does it end to end. The human defines the inputs and accepts the output.
The named mistake in Zone 1 is keeping a human in the loop "to be safe."
It sounds free. It isn't. The human becomes the bottleneck the tool was bought to remove, the loop never compounds, and the cost case collapses. You're now paying for the tool and the queue. Worse: by about week three, that reviewer isn't reviewing. They're rubber-stamping. Which is more dangerous than no review at all, because now there's an accountability story on paper that doesn't describe what happens in practice.
There's a research signature for Zone 1 work, and it's one of the most misread findings in the field. Brynjolfsson, Li and Raymond studied 5,179 customer support agents getting access to a generative AI assistant (Quarterly Journal of Economics, 2025). Productivity rose 14% on average in issues resolved per hour, but 34% for novice and low-skilled workers, with minimal impact on experienced, highly skilled ones.
Read that as a diagnostic, not a headline. If a tool lifts your newest hire to roughly where your most experienced person already was, the task was never carrying judgment. It was carrying pattern recall. Pattern recall is exactly what you Replace. (What that does to how the newest hire ever becomes the most experienced one is the apprenticeship crisis.)
So here's the test before you leave a reviewer attached to Zone 1 work. Remove the review step for two weeks. Can you name specifically what would go wrong? If you can't name it, it isn't a control. It's theatre, and it's billable.
What replaces the reviewer is a quality gate: a defined check the system runs on itself, with a named escalation path when it fails. Building that gate is the step most companies skip, and it's the entire reason they can't take the human out.
Zone 2: Augment, and the trap of finishing the job
Zone 2 is work where the human's contribution is judgment under ambiguity, novel synthesis, or stakeholder management. Hiring decisions. Customer escalations. Contentious prioritisation calls. AI feeds context, drafts options, surfaces what's missing. The human decides.
The named mistake here is the mirror image of Zone 1's: trying to automate it fully. The model lacks the context the human carries about people, relationships and history. That's what the work is made of.
There's a subtler Zone 2 failure, though. The one that makes vague augment so hard to catch.
In July 2025, METR ran a randomized controlled trial with 16 experienced open-source developers on 246 real tasks in repositories they had maintained for years. With AI tools allowed, they took 19% longer. They had expected AI to speed them up by 24%. And after the experiment, after being measurably slower, they still believed they had been 20% faster.
Handle this one honestly, because it's routinely over-claimed. METR is explicit that it's a snapshot of early-2025 tooling in one setting, and not evidence that AI fails to speed up most developers. Take it for what it does show: in genuine Zone 2 work, deep context, ambiguous tasks, a system held in someone's head, the practitioner's own sense of whether AI is helping isn't reliable evidence.
That's the real cost of the co-pilot default. An organization sitting in vague augment has no instrument. Everybody reports it's helping. Nobody can show it. And you can't manage the gap between those two sentences.
Zone 2 done properly isn't "the human uses AI." It's a designed handoff, stated in advance: what AI prepares, what the human decides, what the human may not hand over. The ambiguity is the whole expense.
Zone 3: Refuse is a zone, not a failure state
This is the part nobody else writes, so let me be blunt about it.
Refuse is not "we're not ready yet." It's a call that for this workflow the downside is asymmetric, legal, ethical, or a relationship you can't rebuild, so AI does not touch it. And it's written down, dated and signed, so nobody drifts into AI use by accident and nobody makes that call alone at 11pm on a personal account.
Hiring and firing communication. Board-level disclosures. Ending a customer relationship. Anything covered by privilege.
Here's the uncomfortable part: regulators wrote their Refuse list before most companies wrote theirs.
The EU AI Act's prohibitions have applied since 2 February 2025, and Article 5(1)(f) prohibits AI that infers emotions in the workplace outside medical or safety purposes. Annex III classes AI used to recruit and select candidates, to decide terms of employment, promotion or termination, or to monitor and evaluate performance, as high-risk, with obligations applying from 2 August 2026. In the US, New York City's Local Law 144 has since 5 July 2023 barred employers from using an automated employment decision tool without an independent bias audit, published results, and candidate notice.
If you're a 60-person US B2B company, none of that is abstract the moment you hire one person in the EU or one person in New York.
But the legal list is the floor, and it's the least interesting part of Zone 3. The Refuse calls that matter most to a founder-led company have no statute attached at all. Nobody is going to fine you for letting AI draft the message that ends a five-year customer relationship. It will simply cost you the referral, quietly, and you will never trace it back.
The Zone 3 mistake runs in both directions. Putting AI there to look modern is one failure. The downside is asymmetric, so a small gain never justifies it. The far more common failure is never naming the zone at all. Then every employee makes the call privately, with whatever tool is already open on a personal account.
Lenovo's Work Reborn 2026 research, surveying 6,000 full-time employees at enterprise organizations, found between a fifth and a third of workers using AI outside the influence and governance of IT, and 31% of AI users receiving no employer training at all. That's not a policy being ignored. That's a policy that was never written. So a few thousand private ones got written instead.
Why does "augment everything" feel like the safe choice?
Because it is the only one of the three zones that costs nothing today.
Replace requires you to redesign a process and reassign a person. Refuse requires you to say no in writing and own it. Augment requires... a licence.
So Augment wins by default rather than on the merits, and the bill shows up later, itemised as something else... usually as one of the five reasons AI rollouts fail.
Gartner predicted in June 2025 that more than 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. Those look like three separate problems. They're one problem, seen from three angles:
- Costs escalate when Zone 1 work still routes through a human queue and you pay for the tool and the queue at once.
- Business value stays unclear when nobody wrote down which zone the workflow belonged in, so there's no baseline to measure against.
- Risk controls are inadequate when Refuse was never named, so there's nothing for a control to enforce.
Gartner also estimated that only about 130 of the thousands of vendors claiming agentic AI are the genuine article. "Agent washing," in their term. Senior Director Analyst Anushree Verma: "Most agentic AI propositions lack significant value or return on investment."
Buying is the easy part. That's precisely why buying is what keeps happening.
How to make the call, workflow by workflow
The sort is not complicated. Which is what makes skipping it so expensive.
For each workflow, three questions:
1. What is the human's actual contribution here? Coordination, synthesis of known information, or routine generation → Zone 1. Judgment under ambiguity, novel synthesis, or stakeholder management → Zone 2. Not "how important is this person." What do they add to this piece of work?
2. What breaks if this goes wrong, and can it be undone? Recoverable and cheap → Zone 1 absorbs errors, which is part of why it's Zone 1. Asymmetric, legal, ethical, or a relationship you can't rebuild → Zone 3, and the size of the potential gain doesn't get a vote.
3. Who signed the answer? Not "which team thinks so." Which named person decided, on what date. An unsigned zone is an undecided zone, and it drifts back to the default inside a quarter.
Then, and only then, the design work, which is different in each zone:
- Zone 1 → build the loop. Sensor, policy, tool, quality gate, learning. The quality gate replaces the human reviewer. Skip it and you'll never remove the reviewer, no matter how good the model gets.
- Zone 2 → design the handoff. In writing: what AI prepares, what the human decides, what the human may not delegate.
- Zone 3 → write it down. Name it, date it, tell everyone. The document is the deliverable.
Where this sits in a delivery sequence matters as much as the sort itself: it belongs to the people-and-process work in the ORBIT Framework, before any tool decision, not after one.
In ORBIT terms, the sort is the input the Build function consumes. ORBIT names five functions every AI-augmented team has to cover: Orchestrate (direct the agents toward an outcome; decide what the AI is pointed at and why), Run (manage the workflows the agents are inside), Build (create the systems, tools and prompts the agents operate within), Influence (drive adoption and culture change) and Translate (bridge agent outputs and human decisions). A Builder handed an unsorted org builds whatever is easiest to build, which is how a team ends up with forty agents and four working workflows. A Builder handed a signed zone map builds the Zone 1 loops and nothing else. The Zone 2 handoff is Translate work, written down in advance. The Refuse column is an Orchestrate call: what the agents are pointed at, and what they are deliberately not. In the engagements I run, the sort happens in the activation stage, while the team is getting its first real contact with the tools, and before anything gets architected.
What the sort does to your org chart
Do this honestly across a real company and something uncomfortable falls out. Zone 1 is not evenly distributed across the org chart. It's concentrated.
The roles whose week is mostly coordination and synthesis of known information, routing, summarising, status-keeping, chasing, translating between two teams who could talk directly, sort almost entirely into Zone 1. That's the narrow middle of the hourglass. Run the same sort on one manager's calendar and it answers how many managers you actually need. Which is why the zone question refuses to stay a workflow question: sort the work and you've drawn a new org chart, whether you meant to or not.
That's The Hourglass Collapse. The coordination layer exists because human bandwidth couldn't carry information across the wide top and wide bottom directly. Once something faster carries it, that layer has nothing else load-bearing to do.
And this is exactly where the co-pilot default is most popular. Not by accident. "Let's augment the PMs" preserves the layer, which means it preserves the bottleneck. That's the copilot anti-pattern in its purest form: a faster version of the same structure instead of a different kind of organisation, and the reason the hourglass doesn't collapse on its own inside companies that bought the tools and changed nothing else.
What replaces the layer is not a thinner layer. It's a function map. ORBIT's five functions are the shape of the team on the far side of the sort: someone Orchestrating the outcome, someone Running the workflows, someone Building the loops, someone Influencing adoption, someone Translating what the agents produce into decisions. Titles can lie and reporting lines can lie; the functions can't. A three-person pod and a thirty-person department need the same five covered, and the coordination roles that sorted into Zone 1 don't reappear anywhere on that map. That redraw is the architecture stage of the engagement, and it comes after the sort, not before it... which is the point. Sort first, and the org chart draws itself.
The binary
Two versions of the next twelve months.
In the first, you keep the default. Everything stays vaguely augmented. Every employee privately decides what AI is allowed to own, and none of them writes it down. Eighteen months on you have a faster version of the company you already had, plus an invoice. And your competitors have the same tools you do.
In the second, you spend a week sorting. You come out with a Replace column built properly, an Augment column with real handoffs, and a Refuse column with names and dates on it. Nothing in that version needs a better model than the one you already have.
One of those is what your company does this year.
Not deciding is the first one.
Read this next: The Third Great Restructure... the sort applied to the manager layer, and what to do with the people in it.
Sort your first ten workflows
The three-question zone sort... the version worth running before anyone touches a tool. It goes out to the list.
Or if you'd rather I ran it on your org directly: Fractional CAIO →
FAQ
Are AI agents going to replace human workers, or just augment them?
Both. And the split gets decided workflow by workflow, not at the level of jobs. A job is a bundle of workflows, and one job routinely contains all three zones. Argue it at the job level and you get an unwinnable argument. Answer it at the workflow level and you get a plan by Friday. The useful question: for this workflow, is the human contributing coordination and routine generation (Replace), judgment under ambiguity and stakeholder management (Augment), or is the downside asymmetric enough that AI stays out (Refuse)? Worth noting where deliberate builders land. Anthropic's Economic Index found roughly 77% of business API usage shows automation patterns, against about 50% for consumer Claude.ai use. Organizations that design a workflow around AI choose Replace far more often than the public debate suggests.
Do you "micro-manage" your AI agents, or let them run?
Whichever the zone says. And if you can't answer without thinking about it, that hesitation is the finding. In Zone 1, letting them run is the entire point; a human checking every output "to be safe" rebuilds the bottleneck the tool was bought to remove, and by week three that reviewer is rubber-stamping, which is worse than no review because it produces an accountability story that isn't true. What replaces them is a quality gate: a defined check the system runs on itself, with a named escalation path. In Zone 2, close involvement isn't micro-management, it's the design: AI prepares, you decide, line agreed in advance. The failure mode isn't too much supervision or too little. It's the same amount of supervision everywhere, which is what happens when nobody sorted the work.
Are the heaviest AI users on the team just blowing past everyone else now?
The evidence points the other way, at least for the routine half of the work. Brynjolfsson, Li and Raymond's study of 5,179 customer support agents (Quarterly Journal of Economics, 2025) found productivity up 14% on average, but 34% for novice and low-skilled workers, with minimal impact on experienced, highly skilled ones. On Zone 1 work, AI compresses the skill gap rather than widening it. And the people most confident they're pulling ahead may be the least reliable witnesses: METR's July 2025 trial found 16 experienced open-source developers took 19% longer on tasks in their own repositories with AI tools allowed, while believing afterwards they'd been 20% faster. Microsoft's 2026 Work Trend Index weighs organizational factors at more than twice individual ones in explaining AI impact: 67% versus 32%. So if a gap has opened on your team, the likelier cause isn't that some people prompt better. It's that some workflows got sorted and others didn't.
Why aren't we collaborating on the prompts we give our AI agents?
Because prompts are being treated as personal style rather than process. That's what happens when the zone was never declared. If nobody has said "this workflow is Replace," there's no artifact anyone is expected to contribute to. Just each person's private way of asking. Two numbers give the shape of it: Lenovo's Work Reborn 2026 research (6,000 full-time employees) found between a fifth and a third of workers using AI outside IT's governance and 31% of AI users getting no employer training; Microsoft's 2026 Work Trend Index found only 13% of AI users say they're rewarded for reinventing work with AI when results don't land. People don't share what they were never asked for. The fix isn't a prompt library or a culture initiative. It's ownership: a named owner per Zone 1 workflow, with the prompt and the policy stored where the process lives, not in somebody's chat history. Collaboration follows the artifact. It doesn't precede it.
What work should we refuse to give AI entirely, and isn't refusing just falling behind?
No. Refuse is a zone, not a delay. It's for work where the downside is asymmetric: legal, ethical, or a relationship you can't rebuild. Hiring and firing communication, board-level disclosures, ending a customer relationship, anything covered by privilege. Some of it is already law rather than preference. The EU AI Act has prohibited AI that infers emotions in the workplace since 2 February 2025, and classes AI used in recruitment, promotion, termination and performance monitoring as high-risk from 2 August 2026; New York City's Local Law 144 has required an independent bias audit and candidate notice for automated employment decision tools since 5 July 2023. But the statutes are the floor. The failure isn't refusing. The failure is refusing implicitly, leaving the zone unnamed so every employee decides it privately, on a personal account, and you find out where the line was only after somebody crossed it.
Sources
Microsoft: 2026 Work Trend Index Annual Report, "Agents, human agency, and the opportunity for every organization" (published 5 May 2026). Survey of 20,000 workers across 10 countries (US, Brazil, Australia, India, Japan, France, Germany, Italy, Netherlands, UK), fielded 18 February – 20 April 2026, plus analysis of Microsoft 365 signals. Source of: 86% treat AI output as a starting point; 13% rewarded for reinvention of work with AI even if results aren't met; 26% say leadership is clearly and consistently aligned on AI; 67% vs 32% organizational vs individual factors. https://www.microsoft.com/en-us/worklab/work-trend-index/agents-human-agency-and-the-opportunity-for-every-organization
Anthropic: Economic Index report, Uneven geographic and enterprise AI adoption (September 2025). Source of: ~77% of business/API usage showing automation patterns vs ~50% for Claude.ai consumer use; definitions of automation (directive + feedback loop) vs augmentation (learning, task iteration, validation). https://www.anthropic.com/research/anthropic-economic-index-september-2025-report
Klarna: press release, 27 February 2024, "Klarna AI assistant handles two-thirds of customer service chats in its first month." Source of: 2.3 million conversations; two-thirds of chats; equivalent work of 700 full-time agents; resolution time 11 minutes → under 2 minutes; 25% drop in repeat inquiries; satisfaction on par with human agents; 23 markets, 35+ languages; estimated $40m profit improvement in 2024. https://www.prnewswire.com/news-releases/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month-302072740.html
Bloomberg: "Klarna Turns From AI to Real Person Customer Service" (8 May 2025). Source of the Siemiatkowski quote and the reversal. https://www.bloomberg.com/news/articles/2025-05-08/klarna-turns-from-ai-to-real-person-customer-service
Brynjolfsson, Erik; Li, Danielle; Raymond, Lindsey: Generative AI at Work, Quarterly Journal of Economics 140(2), 889–942 (2025); NBER Working Paper w31161. Source of: 5,179 customer support agents; +14% issues resolved per hour on average; 34% for novice/low-skilled workers; minimal impact on experienced, highly skilled workers. https://www.nber.org/papers/w31161
METR: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (10 July 2025); arXiv:2507.09089. Source of: 16 developers; 246 tasks; 19% slowdown with AI allowed; 24% expected speedup; 20% believed speedup after the fact. Includes METR's own generalization caveats, which this essay reproduces rather than hides. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
Gartner: press release, 25 June 2025, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027." Source of: the >40% cancellation prediction and its three stated causes; the "agent washing" estimate (~130 of thousands of vendors genuine); the Anushree Verma quote. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
Lenovo: Work Reborn Research Series 2026 (published 1 May 2026). Survey of 6,000 full-time employees at enterprise organizations. Source of: between one-fifth and one-third of workers using AI outside IT's influence and governance; 31% of AI users receiving no employer training. Reported at: https://www.helpnetsecurity.com/2026/05/01/shadow-ai-risks-it-oversight/
EU AI Act (Regulation (EU) 2024/1689). Article 5 prohibitions applicable from 2 February 2025, including Article 5(1)(f) on inferring emotions in the workplace outside medical or safety purposes; Annex III point 4 (employment, workers management and access to self-employment) as high-risk, with obligations applying from 2 August 2026. https://artificialintelligenceact.eu/annex/3/ and https://artificialintelligenceact.eu/implementation-timeline/
New York City Local Law 144 of 2021: automated employment decision tools; independent bias audit within the prior year, published results and candidate notice; DCWP enforcement from 5 July 2023. https://www.nyc.gov/site/dca/about/automated-employment-decision-tools.page
Frameworks referenced are my own: the 3-Zone Co-Pilot (Replace / Augment / Refuse), the Hourglass Collapse, and ORBIT.