Real Design Dimensions, Not a Universal Law
Selective Permeability, Temporal Differentiation, and Coordinated Action
Answer in brief
Selective permeability and temporal differentiation are more than metaphors but less than demonstrated general design principles.
They are strongly supported as descriptive dimensions of coordination-system design. Across nuclear regulation, aviation, distributed computing, financial regulation, critical infrastructure, organizational practice, biology, and experimental AI control, mature systems regulate both:
- what information, authority, resources, or actions may cross a boundary; and
- when, how fast, for how long, or under what state conditions that crossing may occur.
In a few systems the dimensions are not merely adjacent. Nuclear regulation explicitly multiplies the elevated risk rate associated with unavailable safety equipment by the duration of that degraded state. Countercyclical bank-capital rules make permission to consume capital depend on a deliberately asymmetric schedule: releases can be immediate, while rebuilding is delayed and pre-announced. These are genuine joint boundary–time designs.
But the stronger proposition—that a common cross-domain architecture has been shown to improve coordination—is not established. The research found no study that independently varies boundary permissions and operating tempo and measures their joint effects on coordination, safety, or legitimacy. Implementation evidence is abundant; comparative effectiveness evidence is sparse and mixed. Gates, escalation channels, compartmentalization, circuit breakers, and redundancy have all sometimes failed or backfired.
The most defensible conclusion is therefore:
Boundary and tempo are general design dimensions, not universal prescriptions. Their value depends on variables one level below the abstraction: the object propagating, the channel and direction, capacity, magnitude, delay, reversibility, observability, failure correlation, and intervention mechanism.
This qualification is not minor. “Add a boundary” and “slow the system down” are often the wrong prescriptions. The evidence instead favors more precise questions: which path should be gated, which actor may act, how much may propagate, what feedback must arrive first, how long is exposure tolerable, and what happens if the safeguard itself becomes a bottleneck or common-mode failure?
1. Why boundaries and clocks repeatedly appear together
Any coordinated system of semi-autonomous actors faces two separable problems.
The first is admissibility: who may observe, decide, modify shared state, commit resources, interrupt operations, or transmit information. This is the domain of permissions, jurisdiction, separation of duties, compartmentalization, independent verification, and controlled interfaces.
The second is timing: how quickly an action may propagate, how long a degraded state may persist, how frequently decisions may be revised, and how long the system must wait for observation or independent confirmation. This is the domain of leases, deadlines, rate limits, canaries, checkpoints, batching, cooldowns, circuit breakers, and review cycles.
These dimensions can be adjusted independently. An organization can change access rights without changing cadence; a system can impose delay without changing authority. They become coupled when crossing a boundary:
- consumes scarce capacity;
- mutates shared or difficult-to-reverse state;
- creates correlated exposure;
- produces feedback only after a delay;
- requires authority that expires or depends on current conditions; or
- can propagate faster than observers can diagnose and intervene.
This coupling is a recurring architecture, but not necessarily a single recurring cause. Liability, audit requirements, human shift patterns, insurance, regulatory drafting, and limited attention could independently produce both gates and clocks. That deflationary explanation remains untested and is important because the strongest joint designs in the evidence have domain-specific rationales.
2. The strongest evidence that the dimensions are real
2.1 Nuclear regulation: risk rate multiplied by permitted duration
The clearest case is the U.S. Nuclear Regulatory Commission’s treatment of unavailable safety equipment. Regulatory Guide 1.177 defines the incremental conditional core-damage probability associated with a completion-time change as the increase in conditional core-damage frequency while equipment is unavailable, multiplied by the duration of the allowed outage. Its acceptance guideline is therefore explicitly a risk rate × exposure duration calculation.1
Under the Risk-Informed Completion Time program, the allowed duration is not simply a standing extension. It is recalculated as plant configuration changes, cannot exceed 30 days, and for an emergent configuration must be recomputed within the existing required-action time or 12 hours, whichever is shorter. If the recalculated time is already exhausted, the plant enters the more restrictive condition. The calculation also adjusts for increased common-cause-failure probability when the extent of a problem is not yet known.2
This is stronger than an analogy between “boundaries” and “time.” The operability state of safety equipment and the permitted duration of that state are factors in one regulated risk quantity. The surrounding oversight regime is also temporal: regulatory engagement escalates when adverse performance indicators cross thresholds and, in some cases, remain degraded for a specified number of quarters.3
Yet this case establishes implementation, not effectiveness. No outcome evaluation of operating experience under risk-informed completion times was retrieved. The most elegant joint architecture in the evidence base has not been shown here to outperform a simpler or differently calibrated alternative.
2.2 Financial regulation: permission must be both accumulated and releasable
The Basel countercyclical capital buffer uses directional timing asymmetry. Increases may be announced up to 12 months in advance, while decreases take effect immediately. The United Kingdom exercised this architecture during the COVID-19 shock: on 9 March 2020 it cut the buffer from 1% to 0%, while deferring any subsequent increase until at least March 2022.
A Bank of England difference-in-differences study combined confidential bank-level pass-through data with the universe of UK residential mortgage originations. It found that banks receiving more capital relief maintained higher loan values, lower prices, and greater credit availability to riskier borrowers. The estimated rate effect was about seven basis points against an average mortgage rate of 2.0%, and the authors estimated approximately £3.8 billion in additional 2020 credit relative to a no-release counterfactual.4
This is the strongest effectiveness evidence for a deliberately designed rate rule in the corpus—but its scope is narrow. It concerns one country, one shock, and one asset class. More importantly, only eight of 27 Basel Committee jurisdictions entered the pandemic with a non-zero buffer to release, according to the cross-jurisdictional figure cited by the Bank of England paper. The mechanism worked only where the slow layer had accumulated usable capacity in advance.
The transferable insight is not simply “capital buffers work.” It is that pre-committed releasability matters. The authors distinguish a formally releasable buffer from other capital headroom that is nominally usable but that banks may fear drawing down. Permission, credibility, and the communicated return path mattered alongside the stock of capital.
The case also illustrates the strongest rival to a universal control-architecture interpretation. The study explains the timing asymmetry principally as a credibility and time-consistency device, not as a general law of fast and slow control.
2.3 High-reliability operations: when tempo moves authority
Aircraft-carrier flight operations provide the clearest primary account of tempo changing a permission boundary. Rochlin, La Porte, and Roberts observed that flight-deck events could unfold too quickly for appeals through the formal chain of command. Even the lowest-rated person on deck therefore had the authority and obligation to suspend flight operations immediately under appropriate circumstances, subject to later review.5
The sequence is important:
- physical events can outrun hierarchical referral;
- stop authority is moved toward the point of observation;
- immediate intervention is protected;
- evaluation and accountability occur afterward.
This is not flat organization in general. It is a localized change in authority for a particular fast-moving hazard, embedded in a larger hierarchy operating at other tempos. It is also observational evidence, not a controlled evaluation of alternative command structures.
2.4 Distributed systems: authority with an expiry time
Production distributed systems routinely bind authority to temporal validity.
- Chubby uses replicas, majority decisions, master leases, locks, and access-control lists. Authority is not merely granted; it remains valid for a bounded period and under quorum conditions.6
- Spanner combines transaction and replication rules with measured clock uncertainty. Commit completion is delayed when uncertainty is larger so that the system can guarantee external consistency.7
- TCP congestion control combines a boundary on outstanding traffic with acknowledgement-paced growth, retransmission backoff, and multiplicative reduction. The rule regulates both how much may enter the network and how quickly that allowance grows.8
- Software canaries limit the population exposed to a change for an observation period before further propagation. The temporal window is useful only if the initial exposure is actually contained and the monitored effects can appear within that window.9
These are mature, deployed architectures, not metaphors. But their papers and operating documents rarely isolate the causal contribution of leases, waits, quorums, or staging from the rest of the engineering system. Deployment at scale is not itself proof of comparative effectiveness.
3. Formal results are real—and narrow
Control theory supplies genuine theorems about rate, delay, triggering, and stability. Nair and Evans derive a minimum error-free feedback-data rate for mean-square stabilization of certain stochastic linear systems. Tabuada shows that event-triggered control can preserve asymptotic stability and non-zero inter-execution times under explicit input-to-state stability, regularity, and delay assumptions.10
These results demonstrate that information rate and delay can be constitutive constraints rather than convenient metaphors. But they do not prove that organizational review gates, jurisdictional boundaries, or institutional cooldowns inherit the same mathematics.
Distributed-computing theory imposes a similar restraint. Under the assumptions of the CAP result, a partitionable asynchronous network cannot guarantee both atomic consistency and availability. Amazon’s Dynamo accordingly accepts locally available writes during some failures and reconciles conflicts later, using controlled membership, replica parameters, vector clocks, read repair, and application-specific conflict handling.11
Dynamo is important counterevidence to any rule that tighter coordination is always superior. It does not establish that unrestricted propagation is good. It shows that the right boundary depends on which inconsistency is tolerable, how reconciliation works, and which availability failure the application cannot accept.
Formal results therefore strengthen the claim that rate, delay, and admissibility are real design variables. They weaken attempts to convert that claim into a context-free prescription.
4. The operative variable is usually one level down
Across otherwise unrelated fields, outcome evidence repeatedly relocates the explanation from the general noun—“gate,” “boundary,” “buffer,” “speed”—to a more specific design fact.
| General prescription | What the evidence instead identifies |
|---|---|
| Add an independent review gate | Locus and latency: whether review occurs close to the work and whether it creates large, infrequent batches |
| Stage every rollout | Path and exposure: stage the ordinary path, but retain a carefully bounded emergency path where delay is itself hazardous |
| Maintain capital headroom | Pre-committed releasability: whether actors can use the buffer without signalling weakness or violating expected constraints |
| Install a checklist | Completion process and fidelity: what conversation or verification actually occurs, not the checklist’s nominal presence |
| Slow decision-making to improve care | Information and conflict structure: some fast decision makers use more information, more alternatives, and layered advice |
| Remove physical walls to increase collaboration | Realized behavior: participants may substitute electronic communication and reduce face-to-face interaction |
| Add redundant channels | Failure independence and diversity: redundant components sharing the same defect are one failure channel in disguise |
| Compartmentalize monitors | Scope of evidence: local isolation may hide weak signals distributed across actors and time |
Review gates: proximity and batching matter
DORA’s software-delivery research found no evidence that more formal external change approval was associated with lower change-failure rates, while external approval was associated with worse delivery performance. Its recommended alternative retains segregation of duties but moves peer review into the development team’s workflow, reducing handoff latency.12
The important contrast is not “review versus no review.” It is remote, delayed, batched review versus continuous, proximate review.
Supply-chain research independently identifies order batching as one source of the bullwhip effect: order variance can exceed sales variance, with distortion increasing upstream. This supplies a plausible mechanism for why review bottlenecks can be harmful. A gate that appears to reduce local processing cost may create larger, less frequent releases and amplify variance elsewhere.13
Staging: a gate without contained exposure is not containment
Cloudflare’s 2019 Web Application Firewall outage passed pull-request review, approval, continuous integration, and simulation. The rule then reached the global fleet in seconds, exhausting CPU and causing a 27-minute outage. The remedial design was not simply “more review.” Cloudflare introduced staged deployment for the ordinary path while retaining an emergency global path because active security threats may require rapid propagation.14
The case shows why review and containment are not substitutes. A reviewed action can remain dangerous if exposure is global before feedback becomes available.
Checklists and escalation channels: implementation is not effect
The MERIT cluster-randomized trial introduced Medical Emergency Teams across 23 hospitals. Calls rose from 3.1 to 8.7 per 1,000 admissions, but the composite outcome of cardiac arrest, unexpected death, or unplanned intensive-care admission did not improve significantly.15
Ontario’s introduction of surgical safety checklists across 101 hospitals similarly produced no significant reduction in mortality or complications. But these results do not prove that checklists as a class are ineffective. A systematic review of 300 studies covering more than 7.3 million operations found that only 38% described how the checklist was completed. Effects were clearer for provider communication and understanding than for complications or mortality.16
The evidence therefore supports a narrower conclusion: formal completion does not establish faithful mechanism activation, and proximate communication effects do not guarantee distal clinical outcomes.
Information boundaries: openness and visibility are not monotonic goods
In Bernstein’s factory experiment, installing curtains that reduced managerial observability increased defect-free output per hour by as much as 10–15% after the first week, over four treatment lines compared with 28 controls. The proposed mechanism was not secrecy as such but protected local experimentation and less concealment under continuous observation.17
Conversely, Bernstein and Turban’s pre/post studies of open-plan office conversions found roughly a 70% decline in measured face-to-face interaction while electronic communication increased. These studies had no control group, so the causal inference is weaker, but they demonstrate that designed permeability and realized communication can diverge sharply.18
Eisenhardt’s comparative study of eight high-velocity firms makes a related point about speed. Faster strategic decision makers used more information, developed more alternatives, and used a layered advice process. In this setting, speed was not purchased by narrowing input. Conflict resolution and integration between strategic and tactical decisions helped speed rather than hinder it.19
5. What excessive openness, closure, speed, and inertia look like
The cases do not support a single optimum. They instead reveal distinct failure modes.
5.1 Closure can block the signals needed for coordination
The 9/11 Commission concluded that the prevailing “need to know” culture should be replaced with a “need to share” approach and that overclassification impeded information sharing, oversight, and accountability.20 The harmful boundary was partly cultural and self-imposed rather than a carefully designed permission rule, which itself cautions against focusing only on formal architecture.
The Columbia Accident Investigation Board found organizational barriers to the communication of critical safety information, an informal chain of command operating outside formal rules, and a request for orbital imagery that was made and rescinded without effective resolution. Its remedies included stronger independent technical authority and safety assurance—new boundaries designed to counter failures produced by the old ones.21
These cases do not say “remove boundaries.” They show that a boundary can suppress anomaly information while leaving decision authority concentrated.
5.2 Dissolving a boundary can be equally dangerous
The Fukushima Nuclear Accident Independent Investigation Commission described the accident as profoundly man-made and attributed it in part to collusion among the government, regulator, and operator.22 The failure was not excessive compartmentalization but the collapse of institutional independence.
This exposes a construct problem: an information boundary, a jurisdictional boundary, and regulator independence are not the same mechanism. They share a vocabulary but regulate different causal relationships.
5.3 Authorization without magnitude and rate limits is weak control
Several major failures involved authorized actions whose scope and pace were insufficiently bounded.
- In the 2017 Amazon S3 disruption, an operator following a playbook entered an incorrect parameter. AWS concluded that the tool allowed too much capacity to be removed too quickly and added slower removal plus minimum-capacity safeguards.23
- Knight Capital deployed new software to seven of eight servers without independent completion verification, left obsolete functionality on the eighth, and lacked an automated capital threshold. In 45 minutes it generated roughly four million executions involving 397 million shares and lost more than $460 million. Its attempted repair also propagated the defective state to previously correct servers.24
- During the 2010 Flash Crash, an algorithm executed 75,000 futures contracts in 20 minutes using volume as its main pacing input. The official report contrasted this with a comparable earlier execution spread over more than five hours using price, time, volume, and human constraints.25
Authentication answers who may act. It does not answer how much that actor may change, how quickly, against which invariant, or through what rollback path.
5.4 Deadline pressure can reshape decisions even when standards remain nominally fixed
An observational study of U.S. drug-review deadlines found that approval decisions clustered in the weeks preceding statutory deadlines. Drugs approved in the final two months were more likely to be withdrawn for safety reasons (odds ratio 5.5, 95% CI 1.3–27.8), to receive a subsequent black-box warning (OR 4.4, 95% CI 1.2–20.5), and to have dosage forms voluntarily discontinued (OR 3.3, 95% CI 1.5–7.5).26
The wide confidence intervals support a direction more than a precise magnitude, and drugs were not randomly assigned to deadline proximity. Even so, this is unusually direct evidence that a nominally unchanged decision rule can operate differently when commitment formation is concentrated against a clock.
5.5 Rate controls can create the instability they are intended to suppress
Four studies of China’s January 2016 market-wide circuit breaker converge on magnet effects as prices approached the threshold. One found that the breaker had no cooling effect on falling prices, volatility, or order imbalance and induced significant magnet effects. Another found that breaker proximity amplified the magnet effect of a separate price-ceiling rule. The mechanism was withdrawn after four trading days.27
This does not establish that all circuit breakers are harmful. An exchange-sponsored observational study of four U.S. market-wide halts in March 2020 reported stabilized quote volatility without observed liquidity evaporation.28 The designs, markets, events, and evaluation incentives differ.
The larger lesson is that an announced threshold changes behavior before it is reached. A rate control can become a focal point, synchronize withdrawal, or accelerate movement toward the boundary.
5.6 Slow layers can be overtaken by physical propagation
The 2003 North American blackout is the clearest case of institutional response being slower than system dynamics. A 30-minute standard required restoration to a secure N-1 state after a contingency, yet most of the terminal cascade occurred in the final 12 seconds. Investigators concluded that it was not practical to expect operators always to diagnose and correct a massive complex failure within minutes.29
The 2021 Texas event provides independent evidence with different physics and institutions. Grid frequency remained below 59.4 Hz for four minutes and 23 seconds, approaching a nine-minute delayed generator-trip threshold that could have removed approximately 17,000 MW. Manual shedding of as much as 20,000 MW helped avert wider collapse, but protected and underfrequency circuits reduced the load available for rotation, imposing severe and uneven harm on customers.30
This is not evidence that fast automatic control is always superior. It shows that when physical propagation is faster than diagnosis, fast protection must be prepared in advance by slower planning and governance. It also reveals a missing evaluative dimension: aggregate survival can conceal who bears the localized cost.
6. Cross-domain evidence map
| Domain | Recurring architecture | Evidence of implementation | Evidence of effectiveness | Main scope condition or counterevidence | Confidence |
|---|---|---|---|---|---|
| Control theory | Trigger thresholds, data-rate bounds, bounded delay | Formal models and control implementations | Theorems within explicit assumptions | No demonstrated transfer to organizational gates or jurisdiction | High in-domain; low for transfer |
| Distributed systems | Quorums, leases, locks, commit waits, congestion windows, reconciliation | Strong production documentation | Mostly system-level operational evidence, not component-level counterfactuals | Partitions can make availability preferable to synchronous agreement | High on architecture; medium on causal attribution |
| Organizations | Observability boundaries, advice layers, review gates | Extensive | One field experiment, small comparative studies, cross-sectional delivery research | Openness can reduce interaction; external gates can add latency without reducing failure | Mixed, context-sensitive |
| High-reliability and safety organizations | Local stop authority, independent barriers, phased communication rules | Strong field and regulatory documentation | Mostly accident analysis and recommendations | Redundancy fails when channels share the same defect | High on mechanisms; limited outcome evaluation |
| Nuclear regulation | Risk-dependent completion times, state-triggered escalating review | Strong primary documentation | No outcome study retrieved for completion-time design | Plant-specific models and common-cause assumptions are load-bearing | High on implementation; unknown effectiveness |
| Finance | Asymmetric capital-buffer timing, market pauses, position and rate limits | Strong | One positive quasi-experiment; several negative circuit-breaker studies | Rules require pre-accumulated capacity and may create magnet effects | Medium, mechanism-specific |
| Critical infrastructure | Fast automatic protection inside slower planning and jurisdictional layers | Strong incident documentation | Survival in crises, but sparse comparative evaluation | Protection of the whole may impose concentrated local harm | High on failure chronology; medium on design effects |
| Commons and polycentric governance | User boundaries, monitoring, participation, nested jurisdictions | Empirical case synthesis | Some associations with successful governance; weak evidence for cross-scale nesting | In one 91-case synthesis, nesting was the weakest principle and not conventionally significant | Medium for local rules; low for general multi-level claim |
| Biology | Selective transport, checkpoints, asymmetric delays, temporal integration hierarchies | Strong native-domain evidence | Food-web compartmentalization linked to persistence | Biological modularity has domain-specific evolutionary causes; no license to infer organizational effects | High in biology; very low for institutional transfer |
| Multi-agent AI | Selective audits, trusted editing, tool permissions, stateful monitoring | One 2024 testbed, two 2026 preprints, implementation guidance | Bounded experimental effects only | No production multi-agent evidence; local monitoring can fragment the signal | Medium in synthetic tests; low externally |
Institutional-governance evidence is notably thinner than the engineering material. A synthesis of 91 commons cases found strong associations for clearly defined user boundaries and accountable monitoring, but its nested-enterprise principle—the element most relevant to a general multi-layer thesis—was the weakest and did not reach conventional statistical significance. A six-country forest-governance study associated local participation with joint social and ecological gains, losses, and tradeoffs depending on context; it did not directly measure “permeability” as such.31
7. Biology clarifies the mechanisms—and warns against generalization
Biological systems offer unusually clear examples of both dimensions:
- molecular barriers use substrate-selective and directionally biased transport;
- cell-cycle checkpoints actively block later events until earlier conditions are satisfied;
- feed-forward loops can respond rapidly in one direction but delay response in the other, filtering transient signals;
- cortical regions integrate information over different temporal windows; and
- compartmentalized food webs can buffer the propagation of extinctions and increase persistence.32
These are real mechanisms. They do not by themselves establish organizational design principles.
In fact, evolutionary studies provide some of the strongest evidence against easy analogy. Modular structure can emerge from modularly varying environmental goals or from selection to reduce connection costs; both are domain-specific causal explanations demonstrated in computational evolution experiments.33 There is no established reason to assume that a regulator, software team, or public institution faces equivalent selection pressures.
Biology therefore contributes mechanisms and hypotheses, not a warrant for institutional prescriptions. A blood–brain barrier, a regulatory jurisdiction, and an access-control list all regulate crossings, but the similarity of that description does not establish common causes or common optima.
8. Implications for AI-mediated organizations
What present evidence shows
The evidence is extremely limited: one 2024 AI-control testbed and two unreplicated 2026 preprints using synthetic or simulated environments.
In a 5,000-task code-backdooring evaluation, auditing only the first 2% of outputs achieved 15% safety under the testbed’s assumptions. Trusted monitoring with the same audit budget achieved 62% safety at 97% relative usefulness, and trusted editing achieved 92% safety at 94% usefulness.34 These figures depend on the selected models, attack policy, shutdown assumptions, tasks, and simulated audit process; they are not production reliability rates.
The multi-agent results complicate simple compartmentalization. In one synthetic experiment, distributing attack objectives across up to nine coding agents lowered suspicion at each individual step. Write restrictions gave partial mitigation, read restrictions showed no clear effect, and a monitor that saw the full attack detected it more easily—but such a monitor would not generally know in advance which actions belong together. A second experiment found that a standard local monitor caught a distributed attack only one-fifth as often as previous attacks; stateful monitoring that accumulated weak signals across contexts improved detection, although performance declined with large benign backgrounds and remained vulnerable to adaptive obfuscation.35
The present empirical lesson is narrow but important:
A boundary that limits each monitor to one agent or one action can destroy the context needed to detect coordinated behavior; temporal aggregation can partially restore that context.
This introduces an unresolved tradeoff. Cross-context monitoring may improve detection while increasing privacy intrusion, false positives, institutional surveillance power, and the consequences of monitor compromise. None of the experiments evaluates those governance costs.
What remains speculative
There is no production evidence that boundary and rate controls improve the safety of persistent, open-ended multi-agent organizations. It is unknown whether current experimental safety–usefulness curves survive:
- heterogeneous model families;
- parallel rather than sequential execution;
- persistent tool and memory state;
- real approval queues;
- adaptive agents that split tasks around thresholds;
- realistic false-positive costs;
- organizational pressure to bypass slow safeguards; or
- concentrated monitor authority.
The 2003 blackout shows that machine-speed propagation can outrun human diagnosis. It does not show that AI organizations will fail in the same way. Using it for AI is a hypothesis by analogy, not present evidence.
Plausible design possibilities—not established best practices
The evidence supports testing, rather than presuming, the following designs:
- Typed, local execution boundaries with selective cross-agent summaries. Local agents need not expose all content to a central monitor, but may export structured risk signals that permit temporal and cross-context aggregation.
- Separate ordinary and emergency paths. Routine high-impact actions could use staged exposure and observation windows, while emergency paths receive narrower scope, stronger logging, lower magnitude limits, and retrospective review.
- Independent execution invariants. Approval should not be the only control. Tools can independently cap transaction size, resource removal, destination, cumulative exposure, or irreversible actions.
- Stateful monitoring across time. Review should consider cumulative behavior, not only isolated actions, while explicitly measuring false positives and concentrated-observation risk.
- Pre-committed stop and recovery rules. The authority to pause, quarantine, roll back, or revoke credentials should be established before a rapid cascade begins.
- Human intervention matched to the consequence horizon. Where consequences become irreversible faster than humans can diagnose them, human approval alone is unlikely to be an adequate fast control. This is a design hypothesis requiring direct AI evidence, not a conclusion already demonstrated.
A particularly informative experiment would cross monitoring scope with action tempo: local versus cross-agent monitoring, and immediate versus staged execution. It should measure not only task success and attack detection, but also latency, false positives, circumvention, information exposure, concentration of authority, and the distribution of harm. Such a factorial design would directly address the central evidentiary gap.
9. What follows for practical design
The evidence does not support “more boundaries” or “slower action” as general rules. It supports a more discriminating design discipline.
9.1 Specify the crossing before selecting the control
A useful design description should identify:
| Question | Relevant variables |
|---|---|
| What crosses? | Information, authority, code, money, physical resources, commitments, or risk |
| Across which interface? | Team, jurisdiction, service, regulator–operator boundary, replica set, tool boundary |
| In which direction? | Inbound, outbound, bidirectional, escalation-only, release-only |
| Under whose authority? | Local actor, quorum, peer reviewer, independent safety function, regulator |
| How much may cross? | Transaction size, exposed population, cumulative resource change, traffic window |
| How quickly and for how long? | Rate, lease, completion time, deadline, observation window, cooldown |
| What feedback is available? | Delay, noise, coverage, false-positive rate, correlated blind spots |
| Can the action be reversed? | Rollback cost, recovery time, irreversible commitment, hidden state |
| What can fail together? | Shared software, common data, common incentives, common authority |
| Who bears the cost? | Delayed users, excluded groups, concentrated local losses, privacy subjects |
This is not a validated universal checklist. It is the level of specification at which the evidence begins to distinguish working mechanisms from nominal controls.
9.2 Pair authority with magnitude and duration
The strongest incidents show that permission checks alone are insufficient. Authorized actors may still cause systemic damage when tools do not cap the magnitude, duration, or propagation speed of their actions.
A more complete control therefore combines:
- authenticated authority;
- parameter and type validation;
- independent invariants;
- bounded exposure;
- rate limits appropriate to feedback delay;
- observation before expansion; and
- a rollback path that does not itself propagate the defect.
9.3 Match observation windows to consequence latency
Canaries, cooldowns, and checkpoints work only if the relevant effects become visible during the observation window. Rare, seasonal, cumulative, or strategically hidden failures can pass a short canary. Conversely, delaying an emergency response for an observation period can increase harm.
The design problem is therefore not “stage or do not stage,” but whether:
observation window ≥ time needed for relevant evidence to emerge
while remaining shorter than the time after which delayed action becomes intolerably costly.
The research did not identify comparative evidence for selecting this window in most real systems.
9.4 Preserve independence without destroying shared situational awareness
Therac-25 and Ariane 501 show why independent and diverse barriers matter. Redundant systems that share software, assumptions, or interfaces can fail together. Yet the AI-monitoring experiments show the opposite danger: excessive isolation can hide coordinated patterns.
The resulting design tension is genuine:
- execution and shutdown paths may need independence and diversity;
- evidence about weak, distributed signals may need selective aggregation.
A mature design may therefore separate the independence of action channels from the controlled sharing of diagnostic summaries. This composition is plausible, but it has not been comparatively tested across domains.
9.5 Evaluate distribution and legitimacy, not only aggregate stability
The evidence base is weakest here. None of the reviewed studies quantitatively links boundary or tempo design to legitimacy, procedural justice, accountability, or concentrated power.
ERCOT illustrates why aggregate outcomes are insufficient: load shedding helped preserve the grid while protected circuits narrowed the rotatable set and concentrated long outages on others. A central monitor may analogously improve aggregate AI detection while concentrating access to sensitive information and authority over interruption.
Any claim of “effective coordination” that omits who gains discretion, who is delayed, who is observed, and who bears localized failure remains incomplete.
10. Important uncertainties and limits
Four limitations materially constrain the conclusion.
No joint causal test
The research found no study that treats boundary permissions and operating tempo as independent variables and evaluates their separate and joint effects on coordination, safety, or legitimacy. The shared-architecture claim therefore rests on recurring implementation patterns plus domain-specific outcome studies.
That is stronger than analogy by vocabulary, but weaker than a demonstrated general principle.
Construct validity is unresolved
“Boundary” encompasses managerial observability, regulator independence, equipment operability, quorum membership, molecular transport, monitor scope, and jurisdiction. “Tempo” encompasses commit waits, degraded-state duration, deadline pressure, relay delays, merge windows, circuit breakers, and review cadence.
These are not self-evidently single constructs. Regulatory capture and membrane transport may share a label without sharing a mechanism. Generalization is warranted only after identifying the propagated object, channel, capacity, delay, reversibility, failure correlation, and intervention.
Effectiveness evidence is sparse and mixed
The strongest human-organizational evidence includes:
- one positive field experiment involving observability boundaries;
- one positive quasi-experiment involving immediate capital-buffer release;
- one null cluster-randomized trial of an escalation channel;
- one population-level checklist null;
- one cross-sectional software-delivery null on failure reduction;
- several negative studies of one circuit-breaker episode; and
- one observational deadline study with substantial residual uncertainty.
Most other material documents a mechanism, an accident, or a proposed remedy. Serious institutions can implement sensible-looking architectures without proving that those architectures work.
Key domains remain undercovered
Evidence on legitimacy and concentrated power is effectively absent. Polycentric and jurisdictional governance is thin. No outcome evaluation was retrieved for nuclear completion times, the sterile-cockpit rule, Linux merge windows, or software canarying. The canonical Piper Alpha primary report was not obtained. Monetary-policy gradualism remains contested in its own field. Multi-agent AI evidence is synthetic and non-production.
The apparent asymmetry in counterevidence—many documented failures of boundary mechanisms but fewer harmful cooldowns or observation windows—could reflect a real difference, publication bias, or search coverage. The evidence cannot distinguish these explanations.
Conclusion
Selective permeability and temporal differentiation are best understood as general coordinates of coordination design.
They are real because mature systems repeatedly and explicitly regulate who or what may cross an interface, in which direction, at what magnitude, for how long, and under which timing and state conditions. In nuclear regulation, boundary state and duration are combined in one risk calculation. In finance, permission to consume capital is paired with an asymmetric release-and-rebuild schedule. In flight operations, fast hazards move stop authority toward the point of observation. In distributed systems, quorums, leases, uncertainty-dependent waits, congestion windows, and reconciliation make authority inseparable from time.
But this recurrence does not establish a universal law of effective coordination. The same mechanisms can create silos, bottlenecks, common-mode failures, magnet effects, delayed learning, or concentrated power. Speed can be dangerous, but it can also be protective. Openness can improve availability and information flow, but it can also destroy independence; closure can contain failure, but it can also conceal coordinated threats or critical warnings.
The most useful synthesis is therefore not:
Build more boundaries and make the system slower.
It is:
Control propagation at the level of the actual causal interface: specify what may cross, who may authorize it, how much may propagate, how quickly feedback arrives, how long exposure remains tolerable, what can be reversed, what can fail together, and who bears the consequences.
At that level, selective permeability and temporal architecture are productive design concepts. At the level of general nouns, they remain useful but limited metaphors.
Footnotes
-
U.S. Nuclear Regulatory Commission, Regulatory Guide 1.177: An Approach for Plant-Specific, Risk-Informed Decisionmaking: Technical Specifications (August 1998), including the incremental conditional core-damage-probability formulation and acceptance guidelines. ↩
-
U.S. Nuclear Regulatory Commission, Final Revised Model Safety Evaluation for Technical Specification Task Force Traveler TSTF-505, Revision 2, “Provide Risk-Informed Extended Completion Times—RITSTF Initiative 4b”, accession ML18267A259. ↩
-
U.S. Nuclear Regulatory Commission, Inspection Manual Chapter 0305: Operating Reactor Assessment Program (25 November 2019); and Inspection Manual Chapter 0609: Significance Determination Process, including the ΔCDF and ΔLERF significance bands. ↩
-
Aakriti Mathur, Matthew Naylor, and Aniruddha Rajan, Creditable Capital: Macroprudential Regulation and Bank Lending in Stress, Bank of England Staff Working Paper No. 1,011, updated March 2026, originally published January 2023. ↩
-
Gene I. Rochlin, Todd R. La Porte, and Karlene H. Roberts, “The Self-Designing High-Reliability Organization: Aircraft Carrier Flight Operations at Sea,” Naval War College Review 40, no. 4 (1987), consulted through the 1998 reprint. ↩
-
Mike Burrows, “The Chubby Lock Service for Loosely-Coupled Distributed Systems,” 7th USENIX Symposium on Operating Systems Design and Implementation (2006), Google Research PDF; Tushar D. Chandra, Robert Griesemer, and Joshua Redstone, “Paxos Made Live—An Engineering Perspective” (2007), DOI 10.1145/1281100.1281103. ↩
-
James C. Corbett et al., “Spanner: Google’s Globally-Distributed Database,” 10th USENIX Symposium on Operating Systems Design and Implementation (2012), Google Research PDF. ↩
-
Van Jacobson, “Congestion Avoidance and Control,” ACM SIGCOMM 1988, DOI 10.1145/52324.52356; Mark Allman, Vern Paxson, and Ethan Blanton, RFC 5681: TCP Congestion Control (2009). ↩
-
Alec Warner et al., “Canarying Releases,” in The Site Reliability Workbook (Google, 2018). ↩
-
Paulo Tabuada, “Event-Triggered Real-Time Scheduling of Stabilizing Control Tasks,” IEEE Transactions on Automatic Control 52, no. 9 (2007): 1680–1685, DOI 10.1109/TAC.2007.904277; Girish N. Nair and Robin J. Evans, “Stabilizability of Stochastic Linear Systems with Finite Feedback Data Rates,” SIAM Journal on Control and Optimization 43, no. 2 (2004): 413–436, DOI 10.1137/S0363012902402116. The latter was examined through publisher metadata and abstract rather than full text. ↩
-
Seth Gilbert and Nancy Lynch, “Brewer’s Conjecture and the Feasibility of Consistent, Available, Partition-Tolerant Web Services,” ACM SIGACT News 33, no. 2 (2002): 51–59, DOI 10.1145/564585.564601, consulted through partial publisher material; Giuseppe DeCandia et al., “Dynamo: Amazon’s Highly Available Key-value Store,” SOSP 2007, DOI 10.1145/1294261.1294281. ↩
-
DORA, “Streamlining Change Approval,” based on its 2014–2019 State of DevOps research. The evidence is cross-sectional and self-reported, not a causal experiment. ↩
-
Hau L. Lee, V. Padmanabhan, and Seungjin Whang, “Information Distortion in a Supply Chain: The Bullwhip Effect,” Management Science 43, no. 4 (1997): 546–558, DOI 10.1287/mnsc.43.4.546. ↩
-
John Graham-Cumming, “Details of the Cloudflare Outage on July 2, 2019,” Cloudflare, 12 July 2019. ↩
-
Ken Hillman et al., “Introduction of the Medical Emergency Team (MET) System: A Cluster-Randomised Controlled Trial,” The Lancet 365, no. 9477 (2005): 2091–2097, PMID 15964445. Evidence was inspected through the publisher/biomedical abstract record. ↩
-
David R. Urbach et al., “Introduction of Surgical Safety Checklists in Ontario, Canada,” New England Journal of Medicine 370 (2014): 1029–1038, DOI 10.1056/NEJMsa1308261; Brittany A. Armstrong et al., “Effect of the Surgical Safety Checklist on Provider and Patient Outcomes: A Systematic Review,” BMJ Quality & Safety (2022), DOI 10.1136/bmjqs-2021-014361. These findings were inspected through publisher/Europe PMC abstract records. ↩
-
Ethan S. Bernstein, “The Transparency Paradox: A Role for Privacy in Organizational Learning and Operational Control,” Administrative Science Quarterly 57, no. 2 (2012): 181–216, DOI 10.1177/0001839212453028. ↩
-
Ethan S. Bernstein and Stephen Turban, “The Impact of the ‘Open’ Workspace on Human Collaboration,” Philosophical Transactions of the Royal Society B 373 (2018): 20170239, DOI 10.1098/rstb.2017.0239. ↩
-
Kathleen M. Eisenhardt, “Making Fast Strategic Decisions in High-Velocity Environments,” Academy of Management Journal 32, no. 3 (1989): 543–576, DOI 10.2307/256434. ↩
-
National Commission on Terrorist Attacks Upon the United States, The 9/11 Commission Report—Executive Summary (2004), U.S. Government Publishing Office PDF. ↩
-
Columbia Accident Investigation Board, Columbia Accident Investigation Board Report, Volume I (NASA, August 2003), report PDF. ↩
-
National Diet of Japan, Fukushima Nuclear Accident Independent Investigation Commission, The Official Report of the Fukushima Nuclear Accident Independent Investigation Commission—Executive Summary (2012), English report PDF. ↩
-
Amazon Web Services, “Summary of the Amazon S3 Service Disruption in the Northern Virginia Region,” 2 March 2017. ↩
-
U.S. Securities and Exchange Commission, In the Matter of Knight Capital Americas LLC, Securities Exchange Act Release No. 70694, Administrative Proceeding No. 3-15570 (16 October 2013), SEC order PDF. ↩
-
Staffs of the U.S. Commodity Futures Trading Commission and U.S. Securities and Exchange Commission, Findings Regarding the Market Events of May 6, 2010 (30 September 2010), SEC report PDF. ↩
-
Daniel Carpenter, Evan James Zucker, and Jerry Avorn, “Drug-Review Deadlines and Safety Problems,” New England Journal of Medicine 358 (2008): 1354–1361, DOI 10.1056/NEJMsa0706341. Findings and effect estimates were inspected through the publisher/Europe PMC abstract record. ↩
-
Kin Ming Wong, Xiao Wei Kong, and Min Li, “The Magnet Effect of Circuit Breakers and Its Interactions with Price Limits,” Pacific-Basin Finance Journal 61 (2020), DOI 10.1016/j.pacfin.2020.101325; Zeguang Li, Keqiang Hou, and Chao Zhang, “The Impacts of Circuit Breakers on China’s Stock Market,” Pacific-Basin Finance Journal 68 (2021), DOI 10.1016/j.pacfin.2020.101343. Both were examined through bibliographic abstracts. See also Steven Shuye Wang, Kuan Xu, and Hao Zhang, A Microstructure Study of Circuit Breakers in the Chinese Stock Markets, Dalhousie University Working Paper No. 2019-02; and Chen Xiang and Jing Lu, “Magnet Effects of Circuit Breakers in Electronic Order-Driven Markets: Evidence from China,” International Journal of Finance & Economics 28 (2023): 1450–1469, DOI 10.1002/ijfe.2487, examined through its publisher abstract. ↩
-
NYSE American LLC, Notice of Filing and Immediate Effectiveness of a Proposed Rule Change To Make Permanent the Pilot Program for Market-Wide Circuit Breakers, SEC Release No. 34-94565 (1 April 2022), SEC filing. ↩
-
U.S.–Canada Power System Outage Task Force, Final Report on the August 14, 2003 Blackout in the United States and Canada: Causes and Recommendations (April 2004), report PDF. ↩
-
Federal Energy Regulatory Commission, North American Electric Reliability Corporation, and Regional Entity Staff, The February 2021 Cold Weather Outages in Texas and the South Central United States (November 2021), NERC report PDF. ↩
-
Michael Cox, Gwen Arnold, and Sergio Villamayor Tomás, “A Review of Design Principles for Community-based Natural Resource Management,” Ecology and Society 15, no. 4 (2010): 38, full text; Lauren Persha, Arun Agrawal, and Ashwini Chhatre, “Social and Ecological Synergy: Local Rulemaking, Forest Livelihoods, and Biodiversity Conservation,” Science 331 (2011): 1606–1608, DOI 10.1126/science.1199343, examined through the publisher abstract. ↩
-
Leland H. Hartwell and Ted A. Weinert, “Checkpoints: Controls That Ensure the Order of Cell Cycle Events,” Science 246 (1989): 629–634, DOI 10.1126/science.2683079; Shmoolik Mangan, Alon Zaslaver, and Uri Alon, “The Coherent Feedforward Loop Serves as a Sign-sensitive Delay Element in Transcription Networks,” Journal of Molecular Biology 334 (2003): 197–204, PMID 14607112; Uri Hasson et al., “A Hierarchy of Temporal Receptive Windows in Human Cortex,” Journal of Neuroscience 28 (2008): 2539–2550, DOI 10.1523/JNEUROSCI.5487-07.2008; Daniel B. Stouffer and Jordi Bascompte, “Compartmentalization Increases Food-Web Persistence,” PNAS 108 (2011): 3648–3652, DOI 10.1073/pnas.1014353108. Except for Hasson et al., these were inspected through biomedical abstract records. ↩
-
Nadav Kashtan and Uri Alon, “Spontaneous Evolution of Modularity and Network Motifs,” PNAS 102 (2005): 13773–13778, DOI 10.1073/pnas.0503610102; Jeff Clune, Jean-Baptiste Mouret, and Hod Lipson, “The Evolutionary Origins of Modularity,” Proceedings of the Royal Society B 280 (2013): 20122863, DOI 10.1098/rspb.2012.2863. Evidence was inspected through publisher/biomedical abstracts. ↩
-
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger, “AI Control: Improving Safety Despite Intentional Subversion,” Proceedings of the 41st International Conference on Machine Learning, PMLR 235 (2024): 16295–16336, PMLR publication page. ↩
-
Oliver Makins et al., “Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors,” arXiv:2607.07368 (2026), DOI 10.48550/arXiv.2607.07368; Davis Brown et al., “Stateful Online Monitoring Catches Distributed Agent Attacks,” arXiv:2605.31593 (2026), arXiv record. Both are unreplicated preprints using synthetic or simulated environments. ↩