The Exception Boundary: How AI Relocates Adaptive Coordination Work
Central finding
Increasing AI capability does not produce a uniform decrease or increase in coordination work. The best-supported conclusion is relocation with a conditional net effect.
AI can remove routine integration, monitoring, reconciliation, scheduling, and translation when work is bounded, machine-readable, machine-verifiable, cheaply reversible, and weakly coupled. But as machines absorb the easy cases, the work remaining for people is selected for difficulty. Human effort shifts away from routine execution and toward exception diagnosis, context reconstruction, verification, authorization, recovery, policy authorship, conflict resolution, and system redesign. It also moves outward—to contractors, remote operators, downstream reviewers, customers, and other parties who repair or contest machine-generated actions.
The amount of work after that shift depends on an exception-boundary arithmetic:
- the fraction of work the machine completes without human help;
- the saving on that fraction; and
- the extra cost of cases returned to people.
The third term is easily missed. In the strongest direct field evidence available, an agent was very fast when it succeeded unaided, but cases escalated to humans took longer than equivalent fully human cases. Even so, the deployment reduced aggregate chat duration because the savings on successful automation outweighed the escalation penalty. This demonstrates both sides of the argument: substantial escalation does not necessarily erase automation gains, but the human residue is not merely the old work in smaller quantity. It is harder, later, and more expensive per case.
No study reviewed here measures total coordination work before and after an agentic deployment. All quantitative conclusions therefore concern proxies—chat duration, documentation time, review latency, delivery stability, intervention queues, incident rates, or operational scale—not a complete accounting of human effort and externalized repair.
1. The work AI encounters was already hidden
Organizations rarely operate as their formal processes suggest. Workers routinely bridge incompatible systems, reconstruct missing context, negotiate exceptions, translate between professional groups, maintain unofficial records, and compensate for broken interfaces. This work is often omitted from process maps and productivity measures even though operations depend on it.
Research on invisible work, clinical “shadow systems,” operational resilience, and workplace workarounds describes the same basic pattern: formal representations capture only part of what makes a system function.1 The problem is not simply that management has failed to count a few miscellaneous tasks. Adaptive coordination is frequently inseparable from normal work and becomes visible only when it is removed, overloaded, or transferred to someone else.
That creates a fundamental measurement problem for AI adoption. A team may record:
- more code written, without counting review, integration, and incident response;
- lower handling time, without counting monitoring and escalations;
- less typing by clinicians, without measuring downstream chart reconciliation;
- fewer frontline operators, without counting control-room, maintenance, governance, or appeal work;
- more autonomous actions, without tracking customers or contractors who repair their consequences.
Even workers’ perceptions can move independently of measured effort. A randomized study of experienced open-source developers found that early-2025 AI tools made participants 19% slower, while participants believed the tools had made them faster.2 In a clinical ambient-scribe study, perceived documentation burden and cognitive workload improved, but objective EHR documentation time did not; the authors concluded that the principal change was better calibration between perceived and recorded workload.3
Observability itself is therefore not an unqualified good. It can reveal hidden coordination, but it can also suppress useful adaptation. In Ethan Bernstein’s field experiment, shielding a production line from managerial visibility increased output by 10–15%, apparently by protecting the “productive deviance” through which workers kept the process functioning.4 Instrumentation must distinguish harmful workaround dependence from legitimate local adaptation rather than treating all deviation as noise.
2. The governing mechanism is the exception boundary
The most informative direct evidence comes from a seventeen-day randomized experiment involving 647 customer-service workers and 680,676 chats on Alibaba’s Taobao platform. Only 5.8% of chats were eligible for the agentic system, making this a narrow rather than enterprise-wide deployment.5
Among matched eligible chats:
- 35.0% were completed without escalation;
- 44.1% received algorithm-triggered technical escalation;
- 8.6% received algorithm-triggered emotional escalation; and
- 12.3% were escalated by a human.
Compared with matched fully human chats, the unaided cases were 64.6% shorter. But every escalation category took longer:
- technical escalation: +19.1%;
- emotional escalation: +40.8%;
- human-initiated escalation: +9.5%.
This is relocation in quantified form. The machine did not simply remove 35% of the work and leave the remainder unchanged. It removed a fast-to-automate portion while concentrating people on a remainder whose per-case duration exceeded the original human baseline.
A simple identity captures the result. Let:
- a be the proportion completed unaided;
- s the saving on those cases; and
- p the added cost of escalated cases.
Relative handling time is:
a(1 − s) + (1 − a)(1 + p)
The deployment reduces handling time when:
as > (1 − a)p
Using the observed parameters, the implied break-even escalation rate is roughly 76%. A system can therefore escalate most eligible work and still reduce this particular proxy if its unaided successes are sufficiently cheap.
That result is important counterevidence to claims that coordination must become the dominant bottleneck. Across the workers’ complete portfolios, chat duration fell 3.2%, while the overall customer-rating effect was statistically insignificant. Ratings fell within the automated segment but rose significantly across the much larger non-automated segment, where workers apparently reallocated attention. Objective resolution, measured by customer retrial, was unchanged except for emotional escalations.5
But several qualifications matter:
- Chat duration is not total worker effort.
- The system touched only 5.8% of volume.
- Monitoring of unaided cases was not comprehensively measured.
- Quality and duration do not follow the same arithmetic.
- The deployment lasted seventeen days.
- Its low-error-cost, high-volume, rapidly reversible customer-service setting may be unusually favorable.
Most importantly, the exception penalty is unlikely to remain fixed as capability increases. Better systems should complete more routine cases, but that means cases reaching humans will be progressively selected for ambiguity, conflict, novelty, or failure. This is a supported implication rather than a measured longitudinal law. It is consistent with Lisanne Bainbridge’s observation that the takeover operator “needs to be more rather than less skilled” and that automated systems depend on expertise whose reproduction they may undermine.6
The Alibaba experiment illustrates the mechanism. Emotional escalations were the most expensive, the only category with degraded objective resolution, and the category in which workers showed lower engagement. The study attributed this mainly to late escalation: by the time the human entered, the customer’s frustration had become harder to reverse. Human-initiated interventions worked better largely because they occurred earlier. The scarce resource was not merely human attention; it was timely, context-rich intervention.
3. Three possible regimes—and why autonomy level does not determine them
The evidence supports three recurring regimes. They should not be treated as fixed types of technology. They are outcomes produced by the relationship among automated coverage, exception cost, coupling, containment, and verification capacity.
| Regime | What happens to coordination | Conditions most associated with it | Evidentiary position |
|---|---|---|---|
| Absorptive | Repeated coordination is encoded into protocols and genuinely removed from case-level work | Stable machine-readable interfaces; machine-verifiable state; idempotence; cheap reversal; low goal conflict; bounded failure domains | Demonstrated at scale for deterministic automation, not open-ended agents |
| Relocative | Routine execution falls while human work concentrates at exceptions, authorization, verification, integration, recovery, and redesign | Moderate exception rates; early escalation; adequate context; bounded authority; real override paths; enough queue slack | Directly observed, including in agentic customer service and high-stakes implementation cases |
| Amplifying | Machine action, coupling, or exceptions grow faster than containment and verification capacity | Cheap weakly bounded action; correlated failures; tight coupling; absent aggregate limits; slow feedback; recovery dependent on failed systems | Strongly demonstrated in automated systems and incidents; agent-specific production evidence remains limited |
Absorptive autonomy: coordination encoded once
Large reductions are possible when coordination can be compiled into stable protocols rather than repeatedly adjudicated. Google’s Borg manages clusters containing tens of thousands of machines through declarative desired state, continuous reconciliation, quota and admission controls, rate-limited placement, and failure-domain management. Google’s Canary Analysis Service evaluates hundreds of thousands of production changes daily, while AWS reports more than 150 million automated deployments annually.7
These systems share several properties:
- desired state and permissible actions are machine-readable;
- outcomes are automatically testable;
- operations are idempotent or reversible;
- deployment proceeds through observable stages;
- retries are bounded and rate-limited;
- failures are isolated into known domains;
- unsafe conditions have explicit fallback behavior.
Here, coordination is not merely performed faster. A policy or protocol replaces repeated case-by-case negotiation.
But these are deterministic systems operating over stable APIs and measurable operational state. They do not establish that populations of open-ended agents can inherit Borg-like supervisory ratios. No agentic deployment in the evidence base demonstrates comparable absorptive performance.
Ambient clinical scribing appeared to be a plausible agentic substitution case, but the objective evidence is mixed. In a randomized trial of 238 physicians, one product reduced time-in-note by 9.5%, while another had a statistically null effect. The significant reduction in work exhaustion occurred in the product arm whose documentation-time effect was null; neither arm produced both effects. The products were used in fewer than one-third of eligible visits.8 A separate pre-post study found large improvements in perceived burden but no significant change in objective documentation minutes.3 These findings may represent meaningful improvements in experienced workload, but they do not demonstrate a net reduction in documentation-related coordination, much less downstream coding, billing, review, or handoff work.
Relocative autonomy: the residue moves upward and outward
In the relocative regime, machines handle routine operations while humans assume narrower but more consequential responsibilities. Evidence supports movement toward:
- escalation handling;
- verification and anomaly diagnosis;
- authorization and authority expansion;
- cross-system integration;
- policy specification;
- recovery planning;
- stakeholder alignment;
- architectural redesign.
The clinical deployment Sepsis Watch illustrates the organizational form. It was integrated through a new nurse role, a dedicated device, physician telephone calls, continuing patient tracking, staff training, role negotiation, and stakeholder relationship-building. Its implementing team reported successful integration into routine care—not demonstrated clinical or economic benefit—and emphasized that significant investment was required to align responsibilities and build trust.9 The model did not eliminate coordination; the organization built a coordination system around it.
Delegated authority produces a similar relocation. In the Progent benchmark, deterministic least-privilege policies reduced agent attack success from 39.9% to 1.0%, with explicit approval required for authority expansion.10 This suggests that bounded delegation can replace universal per-action approval. But the work reappears in authoring, versioning, maintaining, and adjudicating the policy—which no study measured.
Production-oriented AI operations designs explicitly envision this division of labor: machines resolve minor incidents inside predefined boundaries, while people handle critical actions and out-of-boundary cases. Yet public descriptions provide no comparative data on total labor, false escalations, approval delay, or incident outcomes.11 They show what organizations intend to build, not that the intended division is cheaper or resilient.
Amplifying autonomy: action outruns containment
The clearest amplification evidence predates contemporary AI. That makes it more—not less—important, because it shows that the mechanism does not depend on immature language models.
In 2012, a defective Knight Capital deployment transformed 212 customer orders into more than four million executions involving approximately 397 million shares in about 45 minutes, producing losses exceeding $460 million. Staff received 97 automated emails before trading began, but the messages had not been designed as actionable alerts. Information available to one component was not transmitted to the component continuing to trade, and no automatic aggregate capital limit stopped the system.12
The lesson is precise: observability without semantic integration, ownership, authority, and automatic containment is not control. It can simply create more signals for humans to process while machine action continues.
The same architectural problem appears in infrastructure failures:
- CrowdStrike’s defective content update affected an estimated 8.5 million Windows devices; subsequent remediation added canaries, deployment rings, bake time, rollback, and customer control over timing.13
- Meta’s 2021 outage disabled services, DNS, and internal diagnostic tools together, forcing physical intervention.14
- Cloudflare’s nominally staged rollout still affected half of requests because stages were too large, while simultaneous human reverts interfered with one another.15
Rollback is therefore not just a software feature. It is a coordination capability that exists only if the control plane, telemetry, access path, and recovery personnel remain usable when the production system fails.
Interacting automated actors add another failure path. The Flash Crash showed feedback among automated execution, high-frequency trading, inventory recirculation, and liquidity withdrawal.16 In simulation, independent pricing algorithms learned supracompetitive pricing and punishment behavior without explicit communication, although this is not field proof of widespread autonomous collusion.17 Multi-agent benchmarks likewise find that scaling helps decomposable work but can sharply damage sequential tasks: across 260 configurations, performance ranged from +80.8% to −70%, with much greater error amplification in independent than centralized architectures. Those results are benchmark evidence, however, and token overhead is not organizational coordination cost.18
4. Verification may become a bottleneck—but the evidence does not show that it must
Software engineering offers the most extensive evidence for a production-versus-verification imbalance, but it is also where the evidence is most contradictory.
Vendor telemetry reports that higher AI adoption coincides with greater upstream output and increases in review latency, code churn, bugs, incident-to-PR ratios, and unreviewed merges. Faros reports that median time to first review rose 156.6% between each organization’s low- and high-adoption periods. But these are uncontrolled within-organization comparisons from self-selected customers, with undisclosed baseline levels. Faros itself characterizes reviewer overload as a likely explanation rather than an observed causal mechanism.19
LinearB’s data further complicate the story. Agentic pull requests waited much longer to be picked up, but once picked up, AI-assisted pull requests were reviewed faster than unassisted ones. The company also reports that many agentic flows targeted low-priority backlog work carrying no expectation of prompt review. Queue time and verification effort are not the same thing.
DORA’s 2024 survey points in the opposite direction on the closest available measures. Per 25% increase in AI adoption, it estimated:
- +3.1% code-review speed;
- +1.3% approval speed;
- −1.5% delivery throughput; and
- −7.2% delivery stability,
with 89% uncertainty intervals.20 DORA’s 2025 account then reported a positive relationship with throughput while retaining a negative relationship with stability, attributing some change to learning and improved use of the tools.21
This is evidence against both extreme positions. AI-generated action can increase downstream burden, but the available evidence does not establish that review or approval necessarily becomes the dominant bottleneck. In fact, the only survey instrument directly measuring both found them improving.
A deeper disagreement remains unresolved:
- Exposure hypothesis: AI acts as a load test, revealing inconsistent data, undocumented dependencies, weak testing, tight coupling, and manual reconciliation that were already present. DORA argues that loosely coupled teams with fast feedback benefit while tightly coupled organizations do not.
- Creation hypothesis: AI generates genuinely new burden through greater action volume, larger changes, and novel dependencies, even in mature organizations. Faros explicitly argues that strong engineering foundations do not protect organizations from the downstream deterioration it observes.
The remedies differ. If AI mostly exposes dysfunction, the answer is architectural redesign. If it creates new burden independent of existing capability, quality and containment must be pushed back toward the point of generation. No study decomposes these mechanisms at the same site before and after adoption.
One prominent induced-demand story also proves less general than it first appears. The curl project saw its confirmed vulnerability-report rate fall from above 15% to below 5%, but it did not close the reporting interface. It removed the monetary reward, moved reporting channels, then returned to HackerOne five weeks later without restoring rewards. The maintainer attributed the problem partly to AI-generated reports, partly to worsening human submissions, and substantially to the incentive. Peer bounty programs did not show comparable increases.22 The episode supports rate limits, identity requirements, abuse controls, and demand-side incentive design—not a general conclusion that cheaper generation necessarily overwhelms verification.
5. Sparse human intervention is a queueing regime, not a staffing ratio
A system with wide machine autonomy and occasional human intervention is viable only if two conditions hold:
- exception demand remains comfortably below human handling capacity, with slack for bursts and correlated failures; and
- detection, context reconstruction, decision, and response occur before the environment’s time-to-harm.
Average intervention frequency is insufficient. Multi-UAV experiments found that including operator waiting and situation-awareness recovery reduced predicted supervisory capacity by as much as 67%, with 36% attributable to situation-awareness waiting.23 A person who intervenes rarely must first reconstruct what the system has been doing, what has changed, and what options remain. That context-recovery time can dominate the nominal decision time.
The flagship real-world autonomous-vehicle case cannot be assessed against this condition because the decisive variable is undisclosed. Seven autonomous-vehicle companies refused to tell a Senate investigation how frequently remote operators intervene. Waymo reported approximately 70 remote-assistance agents on duty worldwide for a fleet of 3,000 vehicles, but that is an on-duty-agents-to-fleet ratio—not an FTE count, queue-load ratio, or measure of total support labor. Half of Waymo’s remote-assistance workforce was reported to be overseas and not licensed for US roads.24
The architecture nevertheless illustrates a useful pattern: the machine initiates requests rather than requiring continuous observation. This converts a monitoring problem, whose cost scales with fleet size, into a queueing problem, whose cost scales with request rate. But it also creates a hard dependency: several companies reported that vehicles enter a minimal-risk condition if they cannot obtain assistance in time.
The one documented interface failure in the record reverses the usual human-in-the-loop story. A driverless Waymo vehicle correctly stopped for a school bus displaying red lights and extended stop arms. It then asked a remote operator whether the bus had active signals; the operator answered “No,” after which the vehicle passed the bus. The NTSB investigation remained open, so the incident does not establish a general failure rate or probable cause.25 It does establish a design possibility that oversight discussions often neglect: the machine may be right and the human wrong.
Meaningful oversight and helpful oversight are therefore separate questions:
- Meaningful: Can the human notice, understand, and intervene in time?
- Helpful: Does the intervention improve rather than degrade the outcome?
A safe sparse-oversight regime cannot assume that human judgment is an infallible final layer. The interface should be arranged so that one unsupported human answer cannot automatically override a safer machine state.
Conditions for viable sparse oversight
The evidence supports the following design conditions individually, although they have not been tested together as a complete configuration:
| Condition | Why it matters |
|---|---|
| Machine-initiated escalation | Makes attention scale with exception demand rather than fleet size |
| Capacity slack for bursts | Prevents queues from becoming unstable under correlated exceptions |
| Early escalation | Preserves time, context, and human engagement |
| Response faster than time-to-harm | Prevents machine-tempo action from outrunning intervention |
| Automatic aggregate limits | Stops damage even if alerts are ignored or misunderstood |
| Independent recovery paths | Keeps rollback, telemetry, access, and personnel available during failure |
| Small, representative deployment stages | Limits blast radius and reveals defects before full commitment |
| Rate limits, bounded retries, backoff, and identity friction | Contain action without requiring approval for every operation |
| Bounded authority with default-deny expansion | Allows routine delegation while reserving new powers for explicit authorization |
| Hold on non-response | Prevents silence or queue failure from being interpreted as consent |
| A non-decisive human answer | Prevents one mistaken intervention from defeating a safer machine state |
| Real override authority and rich context | Keeps oversight from becoming ceremonial |
| Skill-maintenance mechanisms | Preserves the expertise needed for rare, difficult interventions |
This is not a recipe with proven optimal settings. Friction can prevent cascades, but it can also delay beneficial action. Redundancy can catch errors, but it can add latency and diffuse responsibility. Centralization can contain propagation while creating a concentrated interface bottleneck.
6. When oversight becomes ceremonial
Human-in-the-loop arrangements fail when the assigned residual task is structurally impossible.
A systematic review of computerized prescribing alerts found override rates ranging from 46.2% to 96.2%, with the appropriateness of overrides varying from 0% to 95% by alert type.26 An interface ignored more than 90% of the time is not performing meaningful verification, whatever the formal governance documents claim.
Robodebt shows the same problem in public administration. Australia’s system used tax data and income averaging to issue debts, progressively removing human review. The affected class included roughly 648,000 people; approximately A$1.763 billion in averaging-based debts was withdrawn, about A$751 million was promised in refunds, and a A$112 million settlement was approved. The Commonwealth conceded that it lacked a proper legal basis to raise or recover debts based on income averaging.27
The transferable lesson is the exception default. If a person did not provide information in time, Robodebt assumed the tax data were correct and issued a debt. Non-response triggered action rather than a hold. This transformed missing human engagement into apparent confirmation and shifted the burden of correction to the party least able to carry it.
The judicial record also cautions against reducing the episode to deliberate bad faith: the court found little evidence that responsible ministers and senior officials knew income averaging was unreliable as a basis for debts. That points toward a different durable failure—organizational incompetence combined with dangerous defaults—requiring different remedies.
Oversight can work under better conditions. Child-welfare screeners were less likely to follow an algorithm when the displayed score was an erroneous risk estimate, despite needing supervisory approval to override it.28 The case suggests that professional expertise, rich case-level context, and a real override path can support meaningful dissent. It does not establish that these conditions are sufficient in every domain.
7. Where the work goes
Upward: policy, architecture, and authorization
As execution becomes automated, coordination moves into the definition of goals, privileges, escalation thresholds, exception defaults, and recovery procedures. This can be valuable: a well-designed policy replaces thousands of repetitive approvals. But policy authorship becomes a new operational dependency.
The crucial work includes:
- translating organizational intentions into machine-enforceable rules;
- deciding which actions are reversible or require approval;
- defining conflicts among policies and teams;
- maintaining permissions as workflows change;
- adjudicating requests for authority expansion;
- deciding who owns failures that cross system boundaries.
None of the reviewed studies measures this labor directly.
Outward: contractors, remote operators, and affected parties
Efficiency at one organizational boundary can be purchased by moving repair elsewhere. Waymo’s overseas remote-assistance workforce is a visible example. Data labeling, content review, and other forms of “ghost work” provide the broader pattern: automation often depends on contracted human labor made invisible by organizational and geographic distance.29
Externalized coordination also includes customers who appeal decisions, suppliers who reconcile incompatible records, volunteer maintainers who triage low-value submissions, and public agencies that absorb the consequences of private infrastructure failure. These costs are almost completely absent from deployment metrics.
Downstream: reviewers and integrators
Machine-generated outputs often appear finished before they are organizationally usable. Downstream workers must reconstruct provenance, identify omissions, reconcile contradictions, and decide whether the output can safely enter another process. This is most visible in software review, but the same mechanism can operate in documentation, analysis, claims processing, procurement, and customer service.
Into new or transformed occupations
Historical evidence supports occupational restructuring rather than simple disappearance, but not automatic replacement of every displaced task with a better human role. Acemoglu and Restrepo estimate that about half of US employment growth from 1980 to 2015 occurred in occupations whose titles or tasks changed. Their framework distinguishes a displacement effect from a reinstatement effect created by new tasks—but also finds that reinstatement has weakened and identifies “so-so automation,” where displacement is substantial but productivity gains are modest.30
That possibility is especially relevant to adaptive coordination. Much of this work was already cheap in accounting terms because salaried workers performed it informally. Automating it imperfectly may remove visible operational tasks without generating enough productivity to support new higher-order roles.
Where new roles do emerge, they may be more demanding. Robotic-surgery research found that automation disrupted established learning pathways; trainees who acquired skill relied on premature specialization, abstract rehearsal, and undersupervised struggle.31 Data work remains similarly undervalued: in one study, 92% of 53 high-stakes AI practitioners reported at least one “data cascade,” and 45.3% reported two or more.32
The likely pattern is therefore conditional:
- some operational coordination disappears;
- some becomes policy, governance, or system maintenance;
- some moves to remote or lower-status labor;
- some falls on downstream reviewers;
- some is simply dropped;
- and some is reinstated as more demanding judgment work.
There is no evidence that these categories balance automatically.
8. What better models may solve—and what they probably will not
Plausibly capability-sensitive
Several current burdens may decline as models and interfaces improve:
- review and cleanup caused by low-quality outputs;
- errors at today’s capability frontier;
- failures on benign agent tasks;
- some prompt-injection and tool-use defects;
- some software instability associated with early adoption;
- context-provision costs that better memory and retrieval can reduce.
The reversal in DORA’s throughput association between 2024 and 2025 is evidence that at least some measured costs may be transitional, although changes in framework and sample prevent a clean causal interpretation.21
More durable organizational and institutional constraints
Other limits do not disappear merely because a model becomes more accurate.
Contested goals and legitimacy. The District Court of The Hague voided the legal basis for the Dutch SyRI welfare-fraud system because it failed proportionality requirements and the state had not shown the system to be transparent and verifiable. The system’s capacity to adjust risk models during operation counted against it because it weakened verifiability. The court stopped future use but did not require disclosure of the models used in specific projects or destruction of collected data.33 Technical capability was not the deciding issue; legitimacy, inspectability, and legal authorization were.
Accountability. Capability does not confer legal standing, democratic authority, or professional responsibility. A nominal human may instead become a “moral crumple zone,” absorbing blame for a system they could neither understand nor control.34
Relocated discretion. Automated public administration can move judgment from frontline professionals to programmers, analysts, and system owners who lack contact with affected citizens or the professional accountability structures of the displaced role.35
Representation gaps. Every control system operates through incomplete representations. More capable systems may maintain richer models, but no evidence shows that substantially higher capability eliminates the difficulty of preserving common ground across changing, tightly coupled organizations.
Interactive complexity. Additional agents add interactions, shared dependencies, and common-mode risks. Better local performance does not guarantee safe composition.
Joint activity. Effective collaboration requires directability, predictability, and mutual knowledge—not merely good isolated task performance. These are properties of the human-machine system, not of the model alone.36
Measurement failure. Organizations cannot manage coordination work they do not represent. If adoption metrics count machine output but omit exception handling, policy maintenance, externalized repair, and appeals, increased capability may deepen rather than close the gap.
The last three constraints are theoretically strong but have not been tested against systems substantially more capable than those available today. A system that genuinely maintained common ground and supported directability could weaken some of them. They should be treated as durable hypotheses, not impossibility claims.
9. Practical implications and design possibilities
1. Measure the exception boundary, not just automation coverage
Every deployment should record:
- eligible action volume;
- unaided completion rate;
- escalation rate by cause;
- handling time and effort before and after escalation;
- escalation timing;
- burstiness and correlation;
- abandonment or timeout;
- downstream rework;
- approval latency;
- recovery effort; and
- work shifted to customers, suppliers, contractors, or other teams.
Without these measures, organizations cannot determine whether they are absorbing, relocating, or amplifying coordination.
2. Treat the human interface as a queue with a safety deadline
Headcount ratios such as “vehicles per operator” or “agents per reviewer” are inadequate. Capacity must be evaluated against exception arrival distributions, context-recovery time, service time, correlated bursts, and time-to-harm.
Average utilization should remain well below theoretical capacity when failures can be common-mode or bursty.
3. Contain action automatically rather than relying on alerts
Alerts are requests for human coordination, not controls. Systems need independent aggregate limits, rate limits, bounded retries, kill mechanisms, spending caps, transaction limits, and failure-domain isolation. These controls should remain available when the production system or primary control plane fails.
4. Escalate early and with reconstructed context
Late escalation transfers a deteriorated problem. Escalations should arrive before options close and should include:
- relevant history;
- current state;
- attempted actions;
- unresolved uncertainty;
- available reversible choices; and
- the deadline before harm becomes difficult to prevent.
5. Prefer bounded delegation to universal approval
Per-action human approval can recreate the bottleneck automation was intended to remove. A more promising design possibility is deterministic, least-privilege authority with explicit approval only for expansion. This shifts effort into policy engineering, which should be budgeted and measured as production work.
6. Make non-response safe
Where a human does not answer, the default should ordinarily be to hold, degrade safely, or enter a minimal-risk condition—not infer approval. Robodebt demonstrates how a default-to-act rule can convert missing engagement into mass harm.
7. Design for fallible humans
Human review should not automatically overwrite a safer machine state. Depending on error cost, plausible designs include corroboration requirements, evidence-bearing interventions, staged override, redundant review, or automatic reversion to the lower-risk state. These are design possibilities to test, not established superior arrangements.
8. Preserve residual expertise deliberately
Rare intervention makes practice scarce while leaving the hardest cases to people. Simulation, rotation, shadow operation, manual drills, and review of near misses are plausible skill-maintenance mechanisms, but their effectiveness in agentic systems remains largely unmeasured.
9. Separate technical authorization from legitimate authority
A system can be secure and operationally effective while remaining illegitimate. Governance must address who may define goals, who can appeal, who bears consequences, and who is answerable—not only which API calls are permitted.
10. High-value experiments
Several practical studies would materially improve the evidence.
A before-and-after coordination ledger. At one deployment site, combine time-use observation, objective workflow logs, escalation records, and worker-reported cognitive load. Measure routine work, exception handling, policy maintenance, rework, recovery, and externalized repair.
Exception-distribution disclosure. Require production systems to report escalation frequency, burstiness, correlation, handling time, abandonment, and response latency. This would make sparse-oversight claims testable.
Exposure-versus-creation decomposition. Compare AI adoption across units with different pre-existing architecture, testing, data quality, and coupling. Determine whether mature units are protected or experience the same downstream burden.
Policy-maintenance accounting. Measure the labor required to author and update bounded-authority rules, resolve conflicts, and process authority expansion.
Skill-retention trials. Compare simulation, manual rotation, shadow operation, and no-intervention training for operators responsible for rare high-consequence cases.
Agent-incident investigations. Apply the investigative rigor used by the SEC, NTSB, and infrastructure postmortems to an autonomous-agent failure, including human labor, external repair, and counterfactual containment.
11. Important uncertainties
The report’s conclusions are bounded by five major absences.
- No source measures total coordination work. Net-burden claims are inferences from narrower proxies.
- No longitudinal ethnography or field study follows an agentic deployment through organizational adaptation. Transition dynamics are inferred from short experiments, vendor telemetry, and pre-AI organizational research.
- Exception burstiness and correlation are almost entirely unmeasured. Yet these variables determine whether sparse oversight remains viable.
- Externalized repair is largely invisible. Customer appeals, contractor labor, supplier reconciliation, and public remediation rarely enter deployment metrics.
- Every demonstrated absorptive case operates through deterministic automation over stable interfaces. Extrapolation to open-ended agents is not currently warranted.
Sector coverage is also uneven, concentrating on software, healthcare, infrastructure, finance, customer service, and automated transport. Evidence is thin for autonomous procurement, multi-firm supply chains, insurance and claims processing, content moderation, military command, air-traffic control, and non-Anglophone organizational settings.
The evidence nonetheless rules out two simple stories. AI does not merely substitute for human coordination: it often changes where and in what form that work occurs. But coordination is not destined to become an all-consuming bottleneck either: stable interfaces, automated verification, bounded action, and cheap reversal can support enormous machine-managed spans.
The practical question is not how autonomous a system appears. It is whether the organization has engineered and measured the boundary at which autonomy stops.
Footnotes
-
Susan Leigh Star and Anselm Strauss, “Layers of Silence, Arenas of Voice: The Ecology of Visible and Invisible Work,” Computer Supported Cooperative Work 8 (1999): 9–30, DOI 10.1023/A:1008651105359; Frauke Mörike, Hannah L. Spiehl, and Markus A. Feufel, “Workarounds in the Shadow System,” Human Factors 66, no. 3 (2024): 636–646, DOI 10.1177/00187208221087013; SNAFUcatchers Consortium, STELLA: Report from the SNAFUcatchers Workshop on Coping With Complexity (2017), https://snafucatchers.github.io/. ↩
-
METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” July 10, 2025, https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/. ↩
-
Arko et al., “Pre–Post Evaluation of Documentation Burden, Time Perception, and Cognitive Workload of Ambient AI Scribe Tools,” AMIA Joint Summits on Translational Science Proceedings (2026), PMCID PMC13274272. ↩ ↩2
-
Ethan S. Bernstein, “The Transparency Paradox: A Role for Privacy in Organizational Learning and Operational Control,” Administrative Science Quarterly 57, no. 2 (2012): 181–216, DOI 10.1177/0001839212453028. ↩
-
Yiwei Wang, Chuan Zhu, Tianjun Feng, Lauren Xiaoyuan Lu, and Bingxin Jia, “Agentic AI and Human-in-the-Loop Interventions: Field Experimental Evidence from Alibaba’s Customer Service Operations,” arXiv:2605.14830, version 2 (2026), https://arxiv.org/abs/2605.14830. ↩ ↩2
-
Lisanne Bainbridge, “Ironies of Automation,” Automatica 19, no. 6 (1983): 775–779, DOI 10.1016/0005-1098(83)90046-8. ↩
-
Abhishek Verma et al., “Large-scale Cluster Management at Google with Borg,” EuroSys 2015, DOI 10.1145/2741948.2741964; Štěpán Davidovič and Betsy Beyer, “Canary Analysis Service,” ACM Queue 16, no. 1 (2018), DOI 10.1145/3190566; Amazon Web Services, “Automating Safe, Hands-Off Deployments,” June 18, 2020, https://aws.amazon.com/about-aws/whats-new/2020/06/new-abl-article-automating-safe-hands-off-deployments/. ↩
-
Patrick J. Lukac et al., “Ambient AI Scribes in Clinical Practice: A Randomized Trial,” NEJM AI (November 26, 2025), PMCID PMC12768499, ClinicalTrials.gov NCT06792890. The evidence reviewed included the complete publisher abstract and bibliographic record rather than the full article. ↩
-
Mark P. Sendak et al., “Real-World Integration of a Sepsis Deep Learning Technology Into Routine Clinical Care: Implementation Study,” JMIR Medical Informatics (July 15, 2020), PMCID PMC7391165; Mark Sendak et al., “‘The Human Body is a Black Box’: Supporting Clinical Decision-Making with Deep Learning,” FAT* 2020, DOI 10.1145/3351095.3372827. ↩
-
Tianneng Shi et al., “Progent: Programmable Privilege Control for LLM Agents,” arXiv:2504.11703 (2025), https://arxiv.org/abs/2504.11703; Edoardo Debenedetti et al., “AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents,” NeurIPS 37 (2024), https://proceedings.neurips.cc/paper_files/paper/2024/file/97091a5177d8dc64b1da8bf3e1f6fb54-Paper-Datasets_and_Benchmarks_Track.pdf. ↩
-
Ioannis Papapanagiotou et al., “AI in SRE: How Google Is Engineering the Future of Reliable Operations,” Google Site Reliability Engineering (2026), https://sre.google/resources/practices-and-processes/ai-engineering-reliable-operations/. This is implementation documentation, not comparative outcome evidence. ↩
-
U.S. Securities and Exchange Commission, In the Matter of Knight Capital Americas LLC, Release No. 34-70694, October 16, 2013, https://www.sec.gov/litigation/admin/2013/34-70694.pdf. ↩
-
CrowdStrike, Channel File 291 Incident Root Cause Analysis, August 6, 2024, https://www.crowdstrike.com/wp-content/uploads/2024/08/Channel-File-291-Incident-Root-Cause-Analysis-08.06.2024.pdf; David Weston, Microsoft, “Helping Our Customers Through the CrowdStrike Outage,” July 21, 2024, https://news.microsoft.com/source/asia/2024/07/21/helping-our-customers-through-the-crowdstrike-outage-2/. ↩
-
Santosh Janardhan, “More Details About the October 4 Outage,” Meta Engineering, October 5, 2021, https://engineering.fb.com/2021/10/05/networking-traffic/outage-details/. ↩
-
Cloudflare, “Cloudflare Outage on June 21, 2022,” June 21, 2022, https://blog.cloudflare.com/cloudflare-outage-on-june-21-2022/. ↩
-
Staffs of the U.S. Commodity Futures Trading Commission and U.S. Securities and Exchange Commission, Findings Regarding the Market Events of May 6, 2010, September 30, 2010, https://www.sec.gov/news/studies/2010/marketevents-report.pdf. ↩
-
Emilio Calvano, Giacomo Calzolari, Vincenzo Denicolò, and Sergio Pastorello, “Artificial Intelligence, Algorithmic Pricing, and Collusion,” American Economic Review 110, no. 10 (2020), DOI 10.1257/aer.20190623. The evidence reviewed included the publisher’s abstract and bibliographic record. ↩
-
Yubin Kim et al., “Towards a Science of Scaling Agent Systems,” arXiv:2512.08296, version 3 (2026), https://arxiv.org/abs/2512.08296. The reported predictive framework had cross-validated R2 = 0.373, limiting prospective classification. ↩
-
Faros AI, AI Engineering Report 2026: The Acceleration Whiplash (2026), https://www.faros.ai/blog/ai-acceleration-whiplash-takeaways; LinearB, 2026 Software Engineering Benchmarks Report, https://linearb.io/library/ai-in-software-development. These are vendor telemetry from selected customer populations, not controlled causal studies. ↩
-
Derek DeBellis, Kevin M. Storer, Amanda Lewis, Benjamin Good, and DORA/Google Cloud, Accelerate State of DevOps Report 2024, https://dora.dev/research/2024/dora-report/2024-dora-accelerate-state-of-devops-report.pdf. ↩
-
Nathen Harvey and Derek DeBellis, “Announcing the 2025 DORA Report,” Google Cloud, September 23, 2025, https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report. The full 2025 report was not available in the reviewed evidence, so the precise stability coefficient and effects of the changed analytical framework could not be assessed. ↩ ↩2
-
Daniel Stenberg, “The End of the curl Bug-Bounty,” January 26, 2026, https://daniel.haxx.se/blog/2026/01/26/the-end-of-the-curl-bug-bounty/; Stenberg, “curl Security Moves Again,” February 25, 2026, https://daniel.haxx.se/blog/2026/02/25/. ↩
-
M. L. Cummings and P. J. Mitchell, “Predicting Controller Capacity in Supervisory Control of Multiple UAVs,” IEEE Transactions on Systems, Man, and Cybernetics—Part A 38, no. 2 (2008), DOI 10.1109/TSMCA.2007.914757. ↩
-
Office of U.S. Senator Edward J. Markey, Remote Back Seat Operators: Revealing the Autonomous Vehicle Industry’s Reliance on Human Remote Assistance Operators, March 31, 2026, https://www.markey.senate.gov/imo/media/doc/remote_assistance_investigation_report.pdf. ↩
-
National Transportation Safety Board, “Automated Driving System-Equipped Vehicle Passed School Bus Loading Student Passengers,” investigation HWY26FH007, record dated March 3, 2026, https://www.ntsb.gov/investigations/Pages/HWY26FH007.aspx. At the time represented in the evidence, probable cause had not been determined. ↩
-
“Appropriateness of Overridden Alerts in Computerized Physician Order Entry: Systematic Review,” JMIR Medical Informatics (2020), PMCID PMC7400042. ↩
-
Human Rights Law Centre, “The Federal Court Approves a $112 Million Settlement for the Failures of the Robodebt System,” summary of Prygodicz v Commonwealth of Australia (No 2) [2021] FCA 634, September 30, 2021, https://www.hrlc.org.au/case-summaries/2021-9-30-the-federal-court-approves-a-112-million-settlement-for-the-failures-of-the-robodebt-system/. This is a legal-NGO summary reproducing relevant judgment passages; the judgment itself was not directly inspected. ↩
-
Maria De-Arteaga, Riccardo Fogliato, and Alexandra Chouldechova, “A Case for Humans-in-the-Loop: Decisions in the Presence of Erroneous Algorithmic Scores,” CHI 2020, DOI 10.1145/3313831.3376638. The evidence reviewed was abstract-level. ↩
-
Mary L. Gray and Siddharth Suri, Ghost Work: How to Stop Silicon Valley from Building a New Global Underclass (Houghton Mifflin Harcourt, 2019), ISBN 9781328566249, https://ghostwork.info/. ↩
-
Daron Acemoglu and Pascual Restrepo, “Automation and New Tasks: How Technology Displaces and Reinstates Labor,” NBER Working Paper 25684 (2019), https://www.nber.org/system/files/working_papers/w25684/w25684.pdf. ↩
-
Matthew Beane, “Shadow Learning: Building Robotic Surgical Skill When Approved Means Fail,” Administrative Science Quarterly 64, no. 1 (2019): 87–123, DOI 10.1177/0001839217751692. ↩
-
Nithya Sambasivan et al., “‘Everyone Wants to Do the Model Work, Not the Data Work’: Data Cascades in High-Stakes AI,” CHI 2021, DOI 10.1145/3411764.3445518. ↩
-
Nederlands Juristen Comité voor de Mensenrechten et al. v. The Netherlands, District Court of The Hague, ECLI:NL:RBDHA:2020:1878, February 5, 2020. The evidence reviewed was an ESCR-Net case record rather than the court’s JS-gated English text: https://www.escr-net.org/caselaw/2020/nederlands-juristen-comite-voor-mensenrechten-et-al-v-netherlands-eclinlrbdha20201878/. ↩
-
Madeleine Clare Elish, “Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction,” Engaging Science, Technology, and Society 5 (2019): 40–60, DOI 10.17351/ests2019.260; Ben Green, “The Flaws of Policies Requiring Human Oversight of Government Algorithms,” Computer Law & Security Review 45 (2022), https://www.benzevgreen.com/wp-content/uploads/2022/04/22-clsr.pdf. ↩
-
Mark Bovens and Stavros Zouridis, “From Street-Level to System-Level Bureaucracies: How Information and Communication Technology Is Transforming Administrative Discretion and Constitutional Control,” Public Administration Review 62, no. 2 (2002): 174–184, DOI 10.1111/0033-3352.00168. ↩
-
Gary Klein, David D. Woods, Jeffrey M. Bradshaw, Robert R. Hoffman, and Paul J. Feltovich, “Ten Challenges for Making Automation a ‘Team Player’ in Joint Human-Agent Activity,” IEEE Intelligent Systems 19, no. 6 (2004): 91–95, DOI 10.1109/MIS.2004.74; David D. Woods, “The Theory of Graceful Extensibility,” Environment Systems and Decisions 38 (2018): 433–457, DOI 10.1007/s10669-018-9708-3. ↩