Daemon Corporation

Global Leader in Excellence.

Research Report № 014

Evaluating Coordination Architecture Without Pretending It Is a Score

Central finding

Coordination architecture is diagnosable, but not scoreable.

The most defensible evaluation does not ask whether an organization has “good coordination” in the abstract. It asks whether a named coordination function in a named recurrent episode is working as expected:

  • Was the necessary information sensed and delivered?
  • Did someone with the relevant authority make a disposition?
  • Did the decision carry authorized resources and executable commitments?
  • Did work cross the interface and receive acceptance?
  • Did feedback confirm the resulting state rather than merely confirm that a message or command was sent?
  • Did an exception reach an owner?
  • Did monitoring lead to learning, and did learning reach someone able to revise rules, resources, interfaces, or authority?

A handoff form, decision record, dependency link, meeting, escalation, approval, or lessons repository has no diagnostic meaning simply because it exists or can be counted. It becomes evidence only when tied to an explicit model of the expected function: the trigger, required inputs, responsible actors, decision rights, resource commitments, transfer object, expected state change, confirmation, exception path, and revision authority. This conclusion converges across cybernetics, control-theoretic safety, project and systems-engineering standards, and empirical software studies showing that different representations of dependencies produce different diagnoses.1

That conclusion does not yield a universal dashboard. The evidence points in the opposite direction. General cross-domain “leading indicators” have proved elusive; target systems invite synecdoche—allowing a measured part to stand for the whole—and coordination signals repeatedly change meaning with scale, coupling, urgency, observability, autonomy, reversibility, uncertainty, and power.2 The smallest defensible common framework is therefore a derivation procedure, not a fixed signal list, maturity model, comprehensive organizational map, or composite score.

Three qualifications are essential:

  1. Specificity is not earliness. Interface and broken-loop evidence can locate a defect much more precisely than a mission outcome can, but no study in the evidence base demonstrates that such traces reliably predict coordination failure before outcomes reveal it.
  2. Architecture cannot be purified from context. Resources, stability, mission difficulty, shocks, competence, and history are often co-causes of observed performance. The proper claim is usually that an architectural mechanism plausibly contributed to a result, not that architecture alone caused it.
  3. The proposed Good Work Labs framework is not validated. Its components have evidence behind them, but the assembled procedure has never been tested. No primary study in the evidence base evaluates coordination architecture in a small team, nonprofit, startup, or volunteer organization.

The practical answer is therefore modest but useful: examine one costly coordination loop at a time; derive measures from its assumptions; combine traces with qualitative observation; make one targeted change; pre-register what should move and what should not; and conclude narrowly.


1. The object of evaluation: an enacted loop, not an organizational form

Coordination architecture is best understood as an enacted control and transfer system, not as an organization chart. Its diagnostic unit is a loop or interface connecting:

  1. a trigger, condition, or exception;
  2. the information required to recognize it;
  3. an actor expected to decide;
  4. authority to make that decision;
  5. resources needed to execute it;
  6. a commitment or transfer object;
  7. a receiver and acceptance condition;
  8. feedback about the resulting state;
  9. an exception path when the ordinary process fails; and
  10. authority to revise assumptions, rules, allocations, or interfaces.

This formulation captures the strongest contributions of several traditions without mistaking any tradition’s favored artifact for a universal solution.

Cybernetics contributes the requirement that a regulator contain a usable model of the system being regulated. Control-theoretic safety then distinguishes inadequate decisions, failed execution, missing or delayed feedback, defective process models, and unsafe interactions among controllers.1 Program and systems engineering make those relationships visible through giver/receiver milestones, dependency logic, work authorization, interface agreements, assumptions, anomalies, and corrective-action records.3 Commons governance adds accountability to the affected people, not merely monitoring on behalf of a superior. Incident command and military command-and-control add an equally important correction: enacted authority must sometimes migrate to expertise, so maximal formal rigidity is not the goal.4

The resulting diagnostic grammar is general, but the measures derived from it should be local. For example:

  • A decision record is useful if it identifies the decision right, rationale, affected interfaces, resources, rejected alternatives, and later implementation or revision. A count of records is not useful.
  • A handoff object is useful if it identifies sender, receiver, required content, acceptance, state change, and exception route. Merely completing a form is not evidence of effective transfer.
  • A resource commitment is useful if scope, budget or capacity, authority, timing, and receiver acknowledgement travel together. A decision without these may be an unfunded mandate rather than a commitment.
  • A feedback trace is useful if it confirms the actual state of the controlled process. Confirmation that power was applied to a valve, for example, is not confirmation that the valve opened.
  • An exception record is useful if it reveals where the exception went, how it was evaluated, whether anyone owned its disposition, what action followed, and whether the originator learned the result. The number of exceptions alone is uninterpretable.

This is how evaluation can become more architecture-specific than mission-outcome measurement. It can localize where a loop failed even when it cannot yet establish how much that failure contributed to the ultimate result.


2. What the major traditions actually observe

The traditions in the evidence base do not offer interchangeable measures. They illuminate different functions and operate at different evidentiary levels.

Tradition Principal observations Intended diagnosis What it can establish What it cannot establish alone
Organizational design and coordination theory Timely, accurate, problem-solving communication; shared goals, knowledge, and respect; accountability, predictability, common understanding Quality of enacted relational coordination under interdependence Broad associations between relational process and quality, efficiency, worker, and learning outcomes Causal direction; decision rights, resource commitments, interfaces, or loop closure; that relational coordination is an early-warning measure
Program and project management Activities, dependencies, handoff milestones, critical paths, float, authorization, budgets, variance, corrective actions Missing schedule logic, downstream exposure, commitment–resource mismatch, delayed transfer Whether the represented work network is internally traceable and where delay can propagate Whether the plan, strategy, technical design, or represented dependency model is correct
High-reliability organizing Mindful-organizing behaviors; sensitivity to operations; deference to expertise; reported errors and falls Behavioral conditions associated with reliable performance One short behavioral scale has a prospective association with subsequently reported events That “HRO-likeness” is a valid benchmark; that reported errors equal underlying error rates
Incident and emergency management Command roles, authority migration, improvisation, protocol breaking, after-action records Whether a formal structure can accommodate volatile and novel events That reliable operation may depend on constrained improvisation and movement of decision authority to expertise That formal role clarity, hierarchy, or compliance is monotonically beneficial
Military command and control Decision-right allocation, interaction patterns, information distribution, intended versus actual C2, transitions among approaches Fit between command approach and circumstances; capacity to reconfigure No single C2 configuration dominates all missions; intended and actual authority must be distinguished A universal optimum for centralization, networking, or information sharing
Systems engineering Interface requirements, approvals, assumptions, anomalies, changes, verification status, technical-performance trends Incomplete interfaces, incompatible assumptions, late changes, verification gaps Which modeled technical or organizational boundary remains unresolved That the modeled architecture is optimal, legitimate, or causally effective
Cybernetics and systems safety Controller, process model, action channel, feedback, delay, disturbance, constraint, adaptation Broken control and feedback loops; stale models; uncontrolled change Which segment of a specified loop is absent or defective Whether the objective being controlled is normatively or strategically correct
Commons and polycentric governance Boundaries, locally congruent rules, collective choice, monitoring, monitor accountability, sanctions, dispute-resolution forums, rights to organize, nested arrangements Institutional configurations associated with sustained collective management The largest coded architecture-to-performance association base in the evidence, especially for monitor accountability Simple scale transfer; a blueprint; process qualities such as trust or legitimacy; causal effects from individual principles
Interorganizational networks Integration, external control, system stability, resource munificence, governance form, multi-level effectiveness When network governance works and for whom Architecture and context jointly explain effectiveness; criteria can conflict across community, network, and participant levels Architecture-only scoring or a single effectiveness measure
Collaborative governance Participation conditions, power and resource balance, trust-building, leadership, small wins, satisfaction and consensus Whether collaboration functions as a process Scope conditions under which participation is meaningful or manipulable Whether collaboration produces superior policy or management outcomes
Open-source and software production Commits, reviews, fixes, queue latency, contributor concentration, dependency congruence, meetings, invisible labor Bottlenecks, ownership concentration, coordination fit, hidden integration work Some strong workflow associations and a few causal effects on latency Agreement, authority, fairness, complete labor, or product quality from activity counts
Safety and resilience engineering Brittleness, adaptive capacity, response, monitoring, learning, anticipation, assumption violation Capacity to operate near or beyond ordinary limits Strong theories of bounded adaptive capacity and an operational candidate—assumption violation—for detecting migration toward risk A validated leading measure of brittleness, graceful degradation, or adaptive capacity
Complex adaptive systems Emergence, nonlinearity, feedback, self-organization Whether complexity concepts improve intervention design Mainly a warning against linear attribution A validated architecture diagnostic or demonstrated superior intervention approach

The common mistake is to collapse these observations into a single scale. That would erase the tradeoffs the traditions reveal. A system may become faster while reducing contestability; more documented while increasing reporting burden; more redundant while improving recovery but appearing less efficient; or more participatory while leaving decision authority unchanged.

The maturity-model exception that proves the rule

The NATO C2 maturity model is a genuine monotone ladder: higher levels subsume lower ones, and the highest level is sought. But it does not rank one structural form as universally superior. It ranks the breadth of available C2 approaches and the ability to transition among them as circumstances change. The same doctrine explicitly rejects a one-size-fits-all C2 approach.5

The defensible lesson is not “maturity models are harmless.” It is that, if anything is ordered, it should be adaptive repertoire, not maximum decentralization, networking, information sharing, hierarchy, or formalization.


3. What can be diagnosed with the strongest specificity

3.1 Interfaces and exception routes

The most concrete direct-observation evidence comes from hospital operations. Tucker and Edmondson observed 26 nurses for 239 hours across nine hospitals and catalogued 194 operational failures. Eighty-six percent were “problems rather than errors”; 91 percent arose from breakdowns in information or material transfer to the nurse; coping consumed an average of 33 minutes per eight-hour shift; and only 7 percent of responses addressed the underlying cause even under lenient criteria.6

The architecture remained operational because frontline workers repaired it locally. That repair also hid the defect: the person able to redesign the upstream process rarely learned about it. The diagnostic implication is not that workarounds are bad. It is that a system should observe:

  • the boundary at which a failure appears;
  • the compensating action;
  • the person performing it;
  • the time and attention consumed;
  • where the exception is routed;
  • how many hops occur before it reaches an owner;
  • whether the cause-owning actor learns about it; and
  • whether the interface changes.

Exception routing is therefore a stronger candidate than exception counting. A low exception count can mean smooth operation, poor detection, fear, workaround, or missing instrumentation. A short route to a responsible owner, coupled with disposition and feedback to the originator, has a more specific architectural interpretation.

3.2 Accountability of monitoring

The commons literature supplies the largest coded cross-case association between institutional design and sustained collective performance. In Cox, Arnold, and Villamayor-Tomás’s review of 91 studies and 77 cases, the strongest individual association was for monitors being accountable to, or drawn from, the people governed: ratio 11.7, φ = 0.792, N = 38.7

This is important because it distinguishes two architectures that both contain “monitoring.” A monitor reporting only upward may enforce compliance while remaining insulated from those affected. A monitor answerable to affected participants creates a different accountability loop.

The finding should not be overgeneralized. The cases were coded from published reports, publication selection is not ruled out, and the same literature warns against treating design principles as panaceas or blueprints. A later configurational analysis reused much of the same case base and found that no single principle was necessary and sufficient.8 Monitor accountability is therefore a strong candidate observation, not a portable causal law.

3.3 Enacted decision locus under exception

Formal ownership and actual authority can diverge sharply. In a field study of incident command, formal hierarchical relationships remained fixed while informal tactical authority moved quickly to people with relevant expertise; practitioners reported that the system worked best when supervisors permitted and directed this migration.4 Trauma coordination likewise relied on epistemic contestation, joint sensemaking, cross-boundary intervention, and protocol breaking under novelty—practices made socially costly by reputation and blame.9

The proper observation is consequently not “Are roles clear?” but:

  • Who was documented as the decision owner?
  • Who actually made the decision?
  • Who had the relevant information and expertise?
  • Was authority migration permitted, tacit, resisted, or hidden?
  • Could the acting person obtain the resources needed to execute?
  • Was accountability preserved after authority moved?

Maximizing formal clarity could suppress exactly the adaptation needed during an exception. Conversely, undocumented authority migration can destroy accountability. The evaluand is the fit and traceability of enacted authority, not formalization by itself.


4. Broken loops: more specific than outcomes, not proven earlier

The Columbia accident investigation demonstrates how specifically an architecture can be diagnosed after the fact. It identified, among other defects:

  • observed anomalies normalized rather than converted into decisions;
  • imagery requests left unresolved;
  • an informal decision chain operating outside official rules;
  • no actor responsible for integrated risk above the subsystem level;
  • information lost as analyses moved through the organization;
  • espoused stop-work authority that did not reflect reality;
  • a signature-based consensus process that rendered no one accountable; and
  • an inverted burden of proof requiring engineers to prove danger rather than management to establish safety.10

These findings instantiate nearly every broken loop in the research question:

Broken loop Observable architectural condition
Sensing without decision An anomaly is recorded but has no disposition, owner, or decision deadline
Decision without resources An action is approved but no budget, capacity, access, or authorized executor is attached
Allocation without dependency management Work is assigned without accounting for cross-unit dependencies or integrated risk
Execution without monitoring An action occurs without evidence of receiver acceptance or resulting state
Monitoring without learning Data are collected, but exceptions recur without causal review or changed practice
Learning without revision authority A lesson is recorded, but no actor can change the relevant rule, interface, resource allocation, or decision right
Legitimacy without execution People report voice, fairness, or accepted authority, but decisions and resource flows remain unchanged
Commitment without accountability Many actors sign or assent, but no one is responsible for completion or consequence
Ownerless exception An anomalous case moves among units without an actor responsible for disposition
Jurisdictional conflict Multiple controllers act on the same process under incompatible assumptions

Control-theoretic safety provides a prospective way to derive observations from such loops. Leveson’s method requires that each indicator identify:

  1. the assumption on which safe coordination depends;
  2. how that assumption will be checked;
  3. when it will be checked;
  4. who owns the check; and
  5. what response follows if the assumption is violated.

For shared responsibilities, the assumptions governing how multiple controllers coordinate should themselves be recorded. One worked example treats increasing approval latency as evidence that engineers may begin bypassing an independent technical-review process.2

This is the clearest answer to which failures are in principle detectable before mission failure:

  • an unaccepted handoff;
  • a commitment with no authorized resource;
  • an unresolved interface assumption;
  • a controller with no state feedback;
  • an administratively closed action without receiver or state confirmation;
  • an exception with no owner or disposition;
  • a learning finding with no revision authority;
  • a critical dependency with no responsible integrator; or
  • increasing review latency against a stated assumption about timely independent oversight.

But “in principle detectable” must not be confused with validated early warning. Leveson explicitly describes the evidence as anecdotal rather than scientific proof and notes that prospective validation would require long periods. No study in the evidence base measures these checks at baseline and then demonstrates that they predict later coordination failure.2

The defensible conclusion is therefore split:

  • Broken-loop traces offer high architectural specificity.
  • Their value as leading indicators remains unvalidated.
  • Their immediate use is localization and hypothesis generation, not prediction.

5. The leading-indicator premise is weaker than it appears

Only one instrument in the evidence base retains a clearly prospective design. The Safety Organizing Scale was measured among 1,685 registered nurses in 125 units across 13 hospitals, followed by six months of reported medication errors and patient falls. Higher scale scores were negatively associated with both later reported outcomes.11

Even this result carries important limitations:

  • the sample came from 13 Catholic hospitals;
  • it included one profession;
  • the outcomes were reported rather than independently ascertained events; and
  • reported-event frequency is confounded by differences in detection and reporting.

Relational coordination has a much larger association literature but is not established as a leading indicator. Its flagship hospital study measured both relational coordination and care outcomes cross-sectionally. A systematic review identified 233 studies and 518 findings across 36 countries and 73 industry contexts, but reported no design distribution or causal-direction analysis. Nearly 20 percent of quality findings and roughly 32 percent of efficiency findings were mixed or ran counter to theory.12

The most direct directional test in any adjacent tradition points the other way. A meta-analysis by Beus and colleagues found that injuries predicted subsequent safety climate more strongly than safety climate predicted injuries.13 Process measures may therefore register the consequences of prior performance as much as they forecast future performance.

The practical categories should be stated carefully:

Prospective architectural candidates—not validated leading indicators

  • missing owner or decision right;
  • unaccepted handoff;
  • open interface assumption;
  • commitment without authorized resources;
  • unresolved critical dependency;
  • inaccessible information or resource;
  • controller without state-confirming feedback;
  • key-person dependency;
  • assumption violation; and
  • increasing approval latency against an explicit review assumption.

These can exist before a consequence, but their prospective validity is untested.

Concurrent indicators

  • trigger-to-disposition or handoff-to-acceptance time;
  • queue and unresolved-dependency age;
  • blocked-person time;
  • repeat contacts and failed handoffs;
  • rework, reversals, and duplicated activity;
  • resource-access delay;
  • mismatch between required and actual coordination;
  • meeting, reporting, and managerial-attention load;
  • cognitive and decision burden;
  • concentration of glue work;
  • exception-routing hops; and
  • elapsed time from detection to corrective disposition.

These diagnose the cost or state of an ongoing coordination process. They do not reveal their own meaning.

Lagging indicators

  • missed milestones and cost growth;
  • escaped defects and service degradation;
  • recurrence after administrative closure;
  • recovery time after disruption;
  • staff attrition;
  • substantive goal failure;
  • changed resource allocation;
  • assumption, policy, or interface revision; and
  • architectural change.

Even “learning” and “adaptation” are often lagging: the revision occurs only after a failure or disturbance has made the problem visible.


6. Signals whose direction depends on context

No universal threshold is defensible for the most tempting operational indicators.

Signal Required moderators Interpretation
Decision or handoff latency Urgency, reversibility, scale, coupling Delay may indicate blockage, missing authority, inadequate resources, or appropriate deliberation. Speed may reflect efficient routing or suppressed dissent.
Escalation Power asymmetry, blame, autonomy, local competence, observability Low escalation may mean effective local resolution, silence, fear, inaccessible authority, or hidden workarounds. High escalation may mean confused ownership or a functioning exception channel.
Rework and reversals Uncertainty, task novelty, reversibility They may indicate defective handoffs or healthy exploration and assumption revision.
Duplicated effort and redundancy Coupling, disruption severity, urgency Duplication may be waste, but duplicate channels, reserve capacity, and overlapping competence may provide resilience. No evidence establishes an optimal level.
Participation Power asymmetry, formal authority, control of resources Attendance and contribution show activity, not influence, decision rights, or accountability.
Compliance and adoption Reporting regime, fidelity, timing, observability High compliance can coexist with no outcome effect, ceremonial use, or off-system workarounds.
Coordination overhead Coupling, scale, observability High overhead can be waste or necessary integration; low recorded overhead can mean neglected or invisible dependencies.
Ownership and authority clarity Uncertainty, urgency, expertise distribution A named owner for a handoff or corrective action is useful; rigid formal authority under exception may impede expertise-based response.
Shared situational awareness Observability, task distribution, measurement method Associated with process and performance, but causal direction is unresolved.14
Resource availability and commitment reliability Mission difficulty, total resource level, access rights A commitment-resource mismatch is architecturally specific; overall resource insufficiency is not necessarily an architecture defect.
Meeting and reporting load Dependency density, key-person access, burden distribution Aggregate hours are costs, not quality measures. Reductions can remove needed awareness or merely drive coordination off-platform.
Informal integration and glue work Observability, sustainability, substitutability, concentration Informality may repair a missing interface or productively extend an adequate one. Concentration becomes concerning when absence of one person causes disproportionate degradation.
Procedural fairness and voice Power asymmetry, decision influence, contestability These can improve while allocation and implementation remain unchanged. Psychological safety concerns interpersonal risk-taking, not resource authority or decision influence.
Recovery and graceful degradation Disturbance severity, available slack, recurrence Recovery can demonstrate adaptive capacity or recurring failure to remove causes. Successful operation under calm conditions says little about resilience.
Learning latency Revision authority, quality of evidence, recurrence Fast revision may be learning or reactive drift; slow revision may be blockage or warranted caution.
Architectural change Stability, strategy, environmental change Change may represent correction, experimentation, or uncontrolled drift. The observable candidate is violation of a named assumption, not change itself.

The eight moderators should be recorded as part of every reading:

  • Scale: more actors and dependencies can produce greater latency even under adequate coordination.
  • Coupling: tightly interdependent work may justify more formal interface control and redundancy.
  • Uncertainty: experimentation can legitimately generate exceptions, reversals, and rework.
  • Autonomy: decentralized action requires competence, information access, intent, resources, and bounded authority.
  • Observability: low recorded workload or exception volume may reflect missing instrumentation.
  • Reversibility: fast action is less risky when consequences are readily reversible.
  • Urgency: emergencies change acceptable latency, redundancy, and authority concentration.
  • Power asymmetry: assent, silence, compliance, and participation have different meanings when contestation is costly.

The founding comparative work on interorganizational networks reinforces this approach: effectiveness was characterized as a joint product of integration, external control, system stability, and resource munificence rather than architecture alone.15 Context is not just a nuisance variable to be statistically removed. It is part of what the diagnostic must describe.


7. Diagnostic traps and the counterevidence behind them

7.1 Compliance is not effect

The strongest population-scale test of a mandated coordination artifact examined 101 Ontario hospitals around surgical-checklist adoption, covering 109,341 procedures before and 106,370 after. Adjusted mortality and complication rates did not improve materially. Hospital-reported compliance later reached 99–100 percent in nearly all large community hospitals, and 97 of 101 hospitals reported a special implementation or educational program.16

The compliance figures were not contemporaneous with every hospital’s outcome window, so they do not establish fidelity during the measured null. Nevertheless, the case shows that adoption, later high self-reported compliance, and substantial implementation effort do not establish architectural effect.

7.2 An interface artifact may help, but the artifact cannot be isolated

I-PASS provides the strongest positive handoff case: across 10,740 admissions at nine hospitals, medical errors fell by 23 percent and preventable adverse events by 30 percent. Non-preventable adverse events did not change, providing a useful specificity control; handoff duration and resident workflow also remained effectively unchanged, providing a burden check.17

But I-PASS was a four-component package: a standardized oral and written mnemonic, communication training, faculty development and observation, and a sustainability campaign. Significant reductions occurred at six of nine sites. The study therefore supports the bundle under those conditions, not the claim that a handoff object alone produces better outcomes.

Clinical reviews reinforce the distinction: handoff interventions consistently improved processes while mortality remained unaffected, other outcomes were mixed, and no single handoff tool emerged as best.18 A process measure can therefore move because the intervention makes itself visible, without demonstrating a mission effect.

7.3 Speed can displace rather than remove work

The randomized Nudge intervention covered approximately 8,500 overdue pull requests in 147 repositories and reduced target resolution time by about 60 percent.19 Its authors explicitly stated that the intervention did not reduce total effort: other work was delayed, and effects on quality were left untested.

Latency is therefore a legitimate intervention outcome only when paired with:

  • accuracy or quality;
  • rework and reversals;
  • displaced backlog;
  • workload concentration;
  • recipient acceptance;
  • dissent or contestation opportunity; and
  • effects on other queues.

7.4 Participation is not authority

In Apache, the top 15 developers produced more than 83 percent of changes but only 66 percent of fixes.20 Broad issue reporting and repair participation coexisted with concentrated change authority.

The evaluation should distinguish:

  1. invitation;
  2. attendance;
  3. contribution;
  4. acknowledgement;
  5. influence on disposition;
  6. formal authority;
  7. control of implementation resources;
  8. ownership of consequences; and
  9. practical ability to contest or exit.

Randomized field experiments make the same point institutionally. Changing village decision procedures improved satisfaction, fairness, legitimacy, knowledge, and willingness to contribute while having much smaller effects on the projects selected. A separate participatory-development experiment generated short-run public-goods gains without sustained changes in collective action, decision processes, or marginalized-group involvement.21 Legitimacy and execution are related but separable.

7.5 Error and exception counts cannot be scored

Groups differ not only in the number of errors but also in the likelihood that errors are detected and learned from.22 A high count may indicate poor performance, better detection, greater exposure, or a safer reporting climate.

Near misses are also evaluated systematically as successes. Experimental work involving students and NASA personnel found that managers associated with near misses were rated similarly to managers associated with successes. Near-miss information can encourage riskier subsequent choices and reduce information search. A useful corrective is available: the bias disappeared when probability information and base rates were made salient.23

Every exception review should therefore display exposure and base rates, not merely event narratives or counts.

7.6 Stability is weak evidence of resilience

The successful Vancouver Olympics offered little opportunity to observe C2 agility because no major incident occurred.5 Resilience theory similarly concedes that brittleness is often recognizable only after collapse.24

Stable performance may reflect:

  • favorable conditions;
  • ample but unmeasured slack;
  • an unusually capable individual;
  • hidden repair work;
  • low exposure to disturbance; or
  • an architecture that has not yet been tested.

Near misses, exercises, disrupted periods, and comparable cases that propagated versus contained an exception are more informative than uneventful success.

7.7 Adaptation can become drift

Movement away from original procedures is neither inherently healthy nor pathological. It may be a necessary response to novelty, or it may gradually violate assumptions on which safe coordination depends.

The most defensible observable is not “amount of adaptation.” It is which named assumption changed or was violated, whether anyone detected that violation, and whether the system had authority to review and ratify, reverse, or redesign the adaptation.2 No general threshold separates beneficial improvisation from migration toward uncontrolled risk.

7.8 Redundancy is not automatically waste

Duplicate channels, reserve capacity, overlapping competence, and reconfigurable roles raise costs during ordinary operation but may protect performance during disturbance. Incident command achieves flexibility partly through modular and reassignable roles; resilience theory distinguishes base from extended adaptive capacity.424

The correct question is functional: what failure does the redundancy protect against, at what cost, and under what disturbance? The evidence provides no universal optimal redundancy level.

7.9 Performance may be sustained by hidden labor

Direct observation, practitioner network data, software surveys, and mixed-method meeting studies all indicate that integration labor can be concentrated and poorly represented in performance systems. The numerical magnitudes are weak and context-specific, but the direction is credible.25

A formal trace will miss some combination of:

  • translation across specialist vocabularies;
  • repeated chasing and reconciliation;
  • moderation;
  • mentoring and knowledge transfer;
  • relationship maintenance;
  • off-platform negotiation;
  • private deliberation; and
  • temporary authority brokerage.

That work should be sampled directly, but not fully surveilled. Some tasks should remain undocumented for privacy, autonomy, and flexibility. The key distinction is whether informal work extends a functioning interface or recurrently substitutes for a missing one.

7.10 Records can become ceremonial infrastructure

NASA maintained a formal lessons-learned system that project managers did not routinely search or contribute to.26 A small action-research study of architecture decision records reported improved cooperation, but also found that storage location strongly affected usefulness and that integrating records into existing work required effort.27

Record counts indicate neither retrieval nor learning. A learning trace should ask:

  • Was the lesson retrieved by an affected actor?
  • Did it alter a decision, interface, rule, assumption, or allocation?
  • Was implementation confirmed?
  • Did recurrence decline?
  • Was the lesson later revised or superseded?

7.11 Administrative closure may not be state change

A corrective action can be marked complete even when the receiver has not accepted it and the underlying condition remains. Closure should require some combination of:

  • confirmation by the receiving actor;
  • evidence of the intended state change;
  • sampling of completed actions;
  • recurrence monitoring; and
  • re-opening when the assumption proves false.

7.12 Measurement can create the coordination problem it measures

New fields, records, meetings, interviews, and coding practices consume attention. They can drive work off-platform, create new glue work, encourage gaming, or reduce privacy and autonomy.

For every signal, the pilot should state:

  • what larger phenomenon the signal is standing for;
  • what could improve while the larger phenomenon worsens;
  • how the signal could be gamed;
  • what work is required to produce it; and
  • whether missing-data patterns change after measurement begins.

This “synecdoche check” is as important for a five-measure pilot as for a large dashboard.


8. Representative successes and failures

The evidence is strongest when it distinguishes successful diagnosis from successful intervention.

Strong diagnostic examples

  • Columbia: process tracing identified named architectural defects at named boundaries, including lost information, missing integrated risk ownership, unenacted stop-work authority, and consensus without accountability.10
  • Frontline hospital work: direct observation localized 91 percent of failures at transfers and identified the hidden labor suppressing organizational learning.6
  • Software coordination congruence: across 2,375 modification requests, coordination aligned with modeled dependencies was associated with 32 percent shorter resolution time, while weaker dependency representations explained less.28
  • Commons governance: monitor accountability showed the strongest individual association in the largest coded architecture case base.7

These cases demonstrate localization and association, not a general predictive instrument.

Positive interventions with bounded claims

  • I-PASS: a multi-component handoff package reduced errors, with a specificity control and measured burden neutrality, but only six of nine sites showed significant reductions and the contribution of the handoff object cannot be isolated.17
  • Nudge: a randomized intervention improved pull-request latency but did not reduce total work and may have displaced delay elsewhere.19

Nulls and failures that constrain interpretation

  • Ontario surgical checklists: mandated adoption and substantial implementation effort produced no large outcome improvement.16
  • NASA lessons repository: formal infrastructure existed without routine retrieval or contribution.26
  • Participatory development: short-run outputs did not become durable changes in decision processes or marginalized-group involvement.21
  • Resilience Assessment Grid: a scoping review found no clear scoring method, imbalanced treatment of response, monitoring, learning, and anticipation, and implementation of identified improvements in only two of 12 studies.29
  • Complexity-informed interventions: a scoping review concluded that evidence for the practical usefulness of the complexity lens in healthcare design remained elusive.30

The pattern is consistent: specification, adoption, and process movement are easier to demonstrate than downstream effectiveness.


9. Separating architecture from strategy, resources, context, and luck

A compact diagnostic cannot fully separate architecture from other causes. It can improve causal reasoning by making the architecture-specific mechanism explicit and keeping rival explanations visible.

9.1 Treat context as a first-class observation

Each episode should record, at minimum:

  • mission or task difficulty;
  • case mix;
  • urgency;
  • scale and number of actors;
  • technical and temporal coupling;
  • uncertainty and novelty;
  • reversibility;
  • available resources and access constraints;
  • system stability and recent change;
  • disturbance severity;
  • autonomy;
  • observability;
  • relevant competence;
  • power asymmetry;
  • concurrent interventions; and
  • exceptional effort or unusual slack.

These are not merely covariates to be “controlled away.” They help determine what the architecture was required to do and what the observed signals mean.

9.2 Use an assumption-ranked attribution ladder

Evidentiary position Design Defensible inference
Strongest Randomization of an architectural rule or operational intervention, with pre-specified intermediate, outcome, specificity, and balancing measures Causal effect of the tested package within the studied population
Strong, conditional Controlled interrupted time series, stepped-wedge rollout, difference-in-differences, regression discontinuity, or natural experiment, where design assumptions are credible Causal attribution conditional on those assumptions
Population before/after Objective administrative outcomes, visible exposure timing, large population, adjustment, but no unaffected comparison Strong evidence against a large effect when results are null; weak evidence of a positive causal effect
Multi-site pre/post with specificity control Repeated measures at several sites plus an outcome the intervention should not affect Localization and partial defense against general secular change
Matched episodes plus process tracing Comparable cases or interfaces, predicted mechanism, qualitative tracing, active search for disconfirming cases Plausible contribution, not counterfactual proof
Diagnostic only Cross-sectional dashboard, maturity rating, artifact count, single retrospective, successful outcome, or impressions without baseline Hypothesis generation

No controlled interrupted time series, stepped-wedge rollout, difference-in-differences, regression discontinuity, or natural experiment on coordination architecture appears in this evidence base. That tier is empty because of an evidence gap, not because such designs are inappropriate.

9.3 Match method to the inferential problem

Longitudinal observation and repeated measures are necessary because a cross-sectional signal cannot establish whether poor coordination caused an outcome or the outcome changed perceptions of coordination.

Baselines and run or control charts can help distinguish ordinary variation from an unusual change, but thresholds are not neutral. Tight limits generate false investigations; wide limits miss change; and in-control observations do not prove that the underlying process is adequate.31 Nothing in the evidence supports a universal baseline episode count, minimum detectable change, or acceptable false-alarm rate.

Matched comparisons should preserve case mix, urgency, scale, coupling, and disturbance severity. A comparison between a routine reversible decision and an irreversible crisis decision is not informative merely because both are “decisions.”

Process tracing provides the highest non-experimental architectural specificity. It can show that a missing acceptance step produced rework, that a resource-less decision stalled, or that an exception bounced among jurisdictions. It cannot establish what would have happened under a different architecture.

Near-miss analysis should compare successfully contained events with similar events that propagated, not compare near misses only with obvious failures. Exposure and base rates should be displayed to counter success-like interpretation.

Counterfactual reasoning should identify not only “What if the intervention had not occurred?” but also “What other causal package could have produced this pattern?” Relevant rivals include poor strategy, inadequate technical competence, insufficient total resources, intrinsic difficulty, favorable conditions, shocks, and exceptional individuals.

Mixed quantitative–qualitative evidence is mandatory. Archival traces omit off-platform work, tacit coordination, enacted authority, silence, and private deliberation. Interviews and observation reveal those mechanisms but are vulnerable to recall, status, and hindsight. Neither is sufficient alone.

Where experimental identification is unavailable, the appropriate endpoint is a contribution claim: a verified theory of change, evidence that the predicted mechanism occurred, rival explanations addressed, and a conclusion that the intervention plausibly contributed as one part of a causal package.32


10. The smallest defensible Good Work Labs framework

10.1 Scope

Run each pilot on one recurrent coordination loop or interface:

  • a decision class;
  • a cross-team handoff;
  • an exception route;
  • a resource request;
  • a dependency-resolution process; or
  • a corrective-action and learning loop.

Do not begin with an organizational map. Do not assign a maturity level. Do not aggregate results into a coordination score.

10.2 The procedure

Step 1: Select a suspected coordination cost or failure

Choose an episode that recurs and can be observed. Describe why it matters without presuming its cause.

Examples:

  • decisions stall between technical review and authorization;
  • resource requests are approved but not made accessible;
  • handoffs repeatedly return for clarification;
  • exceptions travel through several people before reaching an owner;
  • commitments are accepted and later renegotiated;
  • corrective actions close but recur;
  • one person repeatedly reconciles incompatible plans.

Step 2: Draw only the local loop

Record:

  • trigger;
  • required information;
  • sender and receiver;
  • affected parties;
  • formal decision owner;
  • enacted decision locus;
  • relevant authority;
  • required resources;
  • handoff or commitment object;
  • acceptance condition;
  • expected state confirmation;
  • dependency relationships;
  • exception path;
  • monitoring responsibility; and
  • revision authority.

This should fit on one page.

Step 3: State the coordination assumptions

For every critical connection, specify:

  • What must be true for the loop to work?
  • What would violate that assumption?
  • How would the violation be checked?
  • When would it be checked?
  • Who owns the check?
  • What happens if it is violated?

Where several actors control the same process, explicitly state how their actions are expected to be coordinated. This is where ownerless exceptions and jurisdictional conflict become observable.

Step 4: Pre-state rival explanations

At minimum:

  • poor strategy;
  • technical error or inadequate skill;
  • insufficient total resources;
  • inaccessible rather than insufficient resources;
  • intrinsic task difficulty;
  • external shock;
  • favorable conditions;
  • exceptional individual effort;
  • abundant slack;
  • case-mix change; and
  • measurement reactivity.

Record relevant resources, stability, urgency, disturbance severity, and case mix in the same instrument.

Step 5: Establish a baseline without inventing a universal sample size

Observe enough comparable episodes to characterize ordinary variation, but do not impose an unsupported episode count. Retain:

  • medians and tails rather than only averages;
  • exposure denominators;
  • episode histories;
  • missing-data patterns; and
  • ordinary, near-miss, and failed cases.

Set thresholds only after understanding the variation and state the false-alarm tolerance.

Step 6: Fill five derived measure slots

The slots are general; their contents come from the loop.

  1. Flow

    • trigger-to-disposition time;
    • handoff-to-acceptance time;
    • resource-request-to-access time.
  2. Integrity

    • missing owner;
    • missing authorized resource;
    • missing acceptance evidence;
    • unresolved dependency age;
    • open interface assumption;
    • state confirmation absent.
  3. Burden

    • blocked-person time;
    • repeat contacts;
    • coordination or reporting time;
    • after-hours repair;
    • concentration of glue work;
    • decision burden by role.
  4. Balancing

    • rework;
    • reversal;
    • quality defect;
    • displaced backlog;
    • reduced dissent opportunity;
    • burden shifted to another team;
    • loss of privacy or autonomy.
  5. Learning, where applicable

    • exception-to-review latency;
    • review-to-revision latency;
    • recurrence after closure;
    • assumption revision;
    • resource reallocation;
    • interface or authority change.

These are not five scores. They are linked observations of different functions.

Step 7: Add four non-optional companions

  1. Specificity control: an outcome the intervention should not change.
  2. Measurement-burden measure: evaluator time, participant time, new reporting work, and changes in missing or off-platform data.
  3. Synecdoche statement: what whole is each signal standing for, and what could be lost if only that signal improves?
  4. Base-rate display: exposure and probability information for exceptions, errors, and near misses.

Step 8: Add a small qualitative sample

Trace three to five ordinary, near-miss, and failed episodes. Interview:

  • the sender;
  • the receiver; and
  • an affected peripheral participant.

Ask:

  • What did the formal record omit?
  • Who translated, chased, reconciled, moderated, or repaired the boundary?
  • What failure did that work prevent?
  • Was it recognized and resourced?
  • Was authority formal, migrated, or improvised?
  • Could someone safely challenge the disposition?
  • Was the exception owner reachable?
  • Would the process degrade if the key integrator were absent?

Step 9: Make one targeted architectural change

Examples include:

  • clarify a decision right;
  • name an exception owner;
  • add receiver acceptance;
  • link a decision to authorized resources;
  • expose dependency status;
  • add state-confirming feedback;
  • create a bounded route for authority migration;
  • connect a lesson to revision authority; or
  • remove an unnecessary approval while preserving contestation.

If adding a form, interface object, or checklist, explicitly acknowledge that standardization has produced both positive bundles and well-powered nulls. Predict movement first in the interface trace, not automatically in the mission outcome.

Step 10: Pre-register the reading

Before examining post-intervention results, state:

  • which measures will be used;
  • expected direction;
  • tolerance for false alarms;
  • important subgroup or role effects;
  • balancing measures;
  • specificity control;
  • criteria for retaining, revising, or stopping the intervention; and
  • rival explanations that would weaken the diagnosis.

Pre-specification matters because a multi-signal system makes post-hoc storytelling easy. In one randomized institutional experiment, the authors demonstrated that the same data could have generated divergent positive and negative interpretations without a pre-analysis plan.21

Step 11: Compare

Prefer, in order:

  • randomized introduction where feasible;
  • phased rollout with comparable units;
  • unaffected comparison interfaces;
  • repeated before/after observations;
  • matched episodes; or
  • process tracing when no stronger counterfactual is available.

Preserve differences in case mix, urgency, scale, coupling, and disturbance severity. Record concurrent changes.

Step 12: Conclude narrowly

The conclusion should state:

  • the likely failing function or interface;
  • the evidence supporting it;
  • the position on the attribution ladder;
  • residual rival explanations;
  • whether the predicted intermediate signal moved;
  • movement in balancing measures and the specificity control;
  • displaced cost and burden;
  • whether the intervention should be retained, revised, or stopped; and
  • what remains unknown about mission outcomes.

10.3 The five operational questions

Pilot question Defensible answer form
Where is coordination failing or costly? At a named segment of a specified loop: unaccepted handoff, missing resource authorization, unresolved dependency, ownerless exception, absent state feedback, learning without revision authority, or excessive routing before ownership.
What evidence supports the diagnosis? Trace evidence, a small qualitative sample, and contemporaneous context observations. Archival records alone are insufficient.
What small architectural intervention might help? One bounded change to a decision right, acceptance object, resource link, exception owner, dependency display, feedback channel, or revision authority. This is a design hypothesis, not demonstrated effectiveness.
What observable change would indicate improvement? Pre-registered movement in the intermediate signal, without unacceptable movement in the balancing measure, with the specificity control remaining stable.
What unintended effects or displaced costs must be checked? Delayed work elsewhere, shifted burden, reduced deliberation, suppressed dissent, work driven off-platform, measurement effort, privacy loss, new glue work, concentrated dependency on individuals, and gaming of closure or the signal.

11. What the framework can and cannot infer

It can

  • Localize a missing, delayed, or overloaded coordination function against an explicit process model.
  • Distinguish among sensing, decision, resourcing, execution, acceptance, feedback, learning, and revision failures.
  • Identify violations of pre-stated assumptions without waiting for the ultimate consequence, although predictive validity is untested.
  • Reveal gaps between formal and enacted authority.
  • Show where compensating labor sustains performance and suppresses learning.
  • Test whether a targeted architectural change moved its predicted intermediate trace.
  • Support a bounded contribution claim while keeping rival explanations visible.
  • Compare the same interface longitudinally or against a genuinely matched interface.

It cannot

  • Establish that a coordination signal leads mission outcomes.
  • Determine whether the strategy, mission, technical design, or controlled objective is correct or legitimate.
  • Fully separate architecture from resources, competence, stability, mission difficulty, shocks, favorable conditions, or luck.
  • Infer good coordination from low latency, low escalation, few errors, high compliance, high participation, low meeting load, or stable outcomes.
  • Measure resilience, brittleness, or graceful degradation prospectively with validated general indicators.
  • Determine whether informal coordination is beneficial or pathological from volume alone.
  • Rank organizations, assign maturity levels over structural form, or aggregate signals into a defensible universal score.
  • Provide validated thresholds for baseline length, false alarms, burden, glue-work concentration, learning latency, or acceptable redundancy.
  • Claim affordability, observability, or transportability at Good Work Labs scale.
  • Work cleanly where no failure mode, coordination hazard, or violated assumption can be named.
  • Claim that the complete framework improves anything. That remains untested.

12. Implications for Good Work Labs

12.1 Treat the first pilots as tests of the diagnostic, not only of the organization

The first question should be whether the proposed observation method is itself usable:

  • Can participants identify a recurrent loop without drawing the whole organization?
  • Can assumptions be stated clearly enough to derive observations?
  • Are the required traces already available?
  • What work is needed to complete them?
  • Does measurement push activity off-platform?
  • Do different observers classify the same loop similarly?
  • Are exception routes and enacted authority recoverable from a few episode traces?
  • Does the diagnostic generate false alarms under changing case mix?

This is a feasibility and construct-validity agenda, not yet an effectiveness trial.

12.2 Favor interventions that complete an already-needed loop

The lowest-burden interventions are likely to add a missing connection rather than a new parallel bureaucracy:

  • attach resource authorization to an existing decision;
  • add receiver confirmation to an existing handoff;
  • route an existing exception log to a named owner;
  • add state confirmation to an existing completion field;
  • give an existing learning review a revision owner;
  • expose an existing dependency rather than build a comprehensive map.

This is a plausible design principle, not an established superiority claim.

12.3 Evaluate distribution, not only averages

Averages can hide the very architecture of interest. Good Work Labs should inspect:

  • tail latency;
  • concentration of decision burden;
  • concentration of repeated repair work;
  • which groups bear reporting costs;
  • whether peripheral participants can contest decisions;
  • whether exceptions originating with lower-power actors travel farther;
  • whether workload or delay is merely transferred; and
  • what happens when the key integrator is unavailable.

12.4 Build evidence chains, not dashboards

A useful repeated experiment links:

stated assumption → observable check → violation → exception route → decision → resource → action → state confirmation → recurrence or revision

The chain can diagnose a broken segment. A dashboard of unlinked latency, meeting, participation, and compliance measures cannot.

12.5 Preserve productive informality

The evaluation should not aim to formalize every interaction. Successful crisis responses, incident command, trauma coordination, and military cases all contain beneficial improvisation or emergent sub-networks.4933

The question is whether informal coordination:

  • extends an adequate architecture;
  • repairs a temporary novel condition;
  • remains bounded and accountable; or
  • repeatedly compensates for a missing interface while exhausting a few people and preventing structural learning.

12.6 Do not promise early warning yet

The most valuable future study would prospectively measure named assumption checks and loop closures, then observe whether failures occur later. Until such evidence exists, Good Work Labs should describe the framework as:

a lightweight, hypothesis-generating diagnostic for localizing coordination costs and broken functions, with untested predictive validity.

That phrasing is not excessive caution. It accurately states the evidence.


13. Evidentiary confidence and unresolved questions

Conclusion Confidence Principal limitation
Traces are diagnostic only against an explicit model of expected function High Few prospective comparisons of alternative diagnostics
A universal coordination score or structural maturity model is unwarranted High Negative conclusion; does not prove that no useful domain-specific index can exist
Interfaces are a high-yield locus for observing failure High within the hospital evidence; Medium cross-domain Direct quantitative evidence is concentrated in one sector
Broken-loop analysis provides architecture-specific localization High Most compelling evidence is retrospective
Broken-loop traces provide earlier warning than outcomes Low / not established One detailed method, no prospective validation
Coordination-process surveys are not established leading indicators Medium-High One prospective behavioral instrument survives, but the broader directional evidence is weak
Compliance and adoption are unreliable proxies for effect Medium-High Strong null evidence, but contemporaneous fidelity remains uncertain
Legitimacy and participation can dissociate from allocation and durable institutional change High Randomized evidence comes from village-development institutions
Enacted, migratable authority is more informative than formal clarity alone Medium-High Mainly qualitative, high-tempo settings
Monitor accountability is a promising cross-domain candidate Medium-High for association; Medium for transportability One published-case base, no established transfer to small organizations
Low escalation and raw exception counts should not be scored High Direction is confounded by detection, reporting, fear, and exposure
General resilience and brittleness measures are not actionable leading indicators High Assumption violation is a promising but unvalidated exception
Glue work is concentrated and under-recorded Medium-High for direction; Low for magnitudes Practitioner data and self-selected surveys dominate
The proposed pilot improves coordination Not established The assembled procedure has never been tested

The most important unresolved questions are:

  1. Do assumption violations and named loop failures actually predict later coordination breakdown?
  2. Can the framework be used affordably in small teams and mission-driven organizations?
  3. What happens when mission outcomes are slow, contested, ambiguous, or unmeasurable?
  4. How much trace structure is enough before documentation becomes a burden or surveillance mechanism?
  5. Can lightweight measures of meaningful voice, contestability, authority acceptance, and exit be validated?
  6. When does concentrated glue work represent efficient expertise rather than brittleness?
  7. Which dependency model is valid for a given domain?
  8. When does adaptation become drift?
  9. How should changing case mix alter thresholds and comparisons?
  10. Can a phased or randomized field experiment test the complete diagnostic procedure rather than one isolated artifact?

Conclusion

The strongest answer is neither a universal coordination index nor a claim that architecture is too contextual to evaluate.

Coordination architecture can be diagnosed when evaluation starts with a specific expected function and follows the evidence through actors, authority, information, resources, commitments, interfaces, feedback, exceptions, accountability, and revision. This provides something mission outcomes cannot: a way to say where coordination is failing and how the failure is organized.

But the diagnostic must remain local and conditional. Faster is not always better. Fewer escalations may mean silence. Participation may lack authority. Compliance may lack effect. Stability may lack adaptive capacity. Redundancy may be protective. Informality may be functional. Performance may be purchased through invisible labor. A process signal can improve while mission outcomes do not, and a successful outcome can conceal a brittle architecture.

For repeated Good Work Labs experiments, the smallest defensible framework is therefore:

  1. one recurrent loop or interface;
  2. an explicit model of its assumptions;
  3. five derived measure slots—flow, integrity, burden, balancing, and learning;
  4. a specificity control, burden measure, synecdoche statement, and base-rate display;
  5. a small qualitative episode sample;
  6. one targeted architectural intervention;
  7. pre-registered interpretation;
  8. repeated or matched comparison; and
  9. a narrow contribution claim with rival explanations and displaced costs stated.

That framework has strong conceptual and evidentiary provenance. It does not yet have demonstrated predictive validity or demonstrated effectiveness. Its first responsible use is as a disciplined way to generate and test architectural hypotheses—not as a score of organizational quality.


Footnotes

  1. W. Ross Ashby, An Introduction to Cybernetics (Chapman & Hall, 1956), digital archive; Roger C. Conant and W. Ross Ashby, “Every Good Regulator of a System Must Be a Model of That System,” International Journal of Systems Science 1, no. 2 (1970): 89–97, doi:10.1080/00207727008920220; Nancy G. Leveson, “A New Accident Model for Engineering Safer Systems,” Safety Science 42, no. 4 (2004): 237–270, doi:10.1016/S0925-7535(03)00047-X. ↩ ↩2

  2. Nancy G. Leveson, “A Systems Approach to Risk Management Through Leading Safety Indicators,” Reliability Engineering & System Safety 136 (2015): 17–34, doi:10.1016/j.ress.2014.10.008; Gwyn Bevan and Christopher Hood, “What’s Measured Is What Matters: Targets and Gaming in the English Public Health Care System,” Public Administration 84, no. 3 (2006), consulted through its abstract. ↩ ↩2 ↩3 ↩4

  3. U.S. Government Accountability Office, Schedule Assessment Guide: Best Practices for Project Schedules, GAO-16-89G (2015), https://www.gao.gov/products/gao-16-89g; U.S. Department of Energy, Earned Value Management System Interpretation Handbook, Version 2.0 (2016), PDF; NASA, NASA Systems Engineering Handbook, NASA/SP-2016-6105 Rev. 2 (2016), PDF. These are standards and implementation guidance, not effectiveness studies. ↩

  4. Gregory A. Bigley and Karlene H. Roberts, “The Incident Command System: High-Reliability Organizing for Complex and Volatile Task Environments,” Academy of Management Journal 44, no. 6 (2001): 1281–1299, doi:10.2307/3069401. ↩ ↩2 ↩3 ↩4

  5. NATO Science and Technology Organization, C2 Agility, STO-TR-SAS-085 (2013), report; David S. Alberts, Reiner K. Huber, James Moffat, and NATO SAS-065 Research Task Group, NATO NEC C2 Maturity Model (2010), report. ↩ ↩2

  6. Anita L. Tucker and Amy C. Edmondson, “Why Hospitals Don’t Learn from Failures: Organizational and Psychological Dynamics That Inhibit System Change,” California Management Review 45, no. 2 (2003): 55–72, doi:10.2307/41166165. ↩ ↩2

  7. Michael Cox, Gwen Arnold, and Sergio Villamayor-Tomás, “A Review of Design Principles for Community-Based Natural Resource Management,” Ecology and Society 15, no. 4 (2010): art. 38, doi:10.5751/ES-03704-150438. ↩ ↩2

  8. Jacopo A. Baggio et al., “Explaining Success and Failure in the Commons: The Configural Nature of Ostrom’s Institutional Design Principles,” International Journal of the Commons 10, no. 2 (2016): 417–439, doi:10.18352/ijc.634. This reanalysis overlaps substantially with the Cox et al. case base and should not be treated as an independent body of cases. ↩

  9. Samer Faraj and Yan Xiao, “Coordination in Fast-Response Organizations,” Management Science 52, no. 8 (2006): 1155–1169, doi:10.1287/mnsc.1060.0526. ↩ ↩2

  10. Columbia Accident Investigation Board, Columbia Accident Investigation Board Report, Volume I (August 2003), PDF. ↩ ↩2

  11. Timothy J. Vogus and Kathleen M. Sutcliffe, “The Safety Organizing Scale: Development and Validation of a Behavioral Measure of Safety Culture in Hospital Nursing Units,” Medical Care 45, no. 1 (2007): 46–54, PMID 17279020, full text. ↩

  12. Jody Hoffer Gittell et al., “Impact of Relational Coordination on Quality of Care, Postoperative Pain and Functioning, and Length of Stay,” Medical Care 38, no. 8 (2000): 807–819, doi:10.1097/00005650-200008000-00005, design statement consulted through a structured abstract; Rendelle Bolton, Caroline Logan, and Jody Hoffer Gittell, “Revisiting Relational Coordination: A Systematic Review,” Journal of Applied Behavioral Science 57, no. 3 (2021): 290–322, doi:10.1177/0021886321991597. ↩

  13. Jeremy M. Beus, Stephanie C. Payne, Mindy E. Bergman, and Winfred Arthur Jr., “Safety Climate and Injuries: An Examination of Theoretical and Empirical Relationships,” Journal of Applied Psychology 95, no. 4 (2010): 713–727, doi:10.1037/a0019164. The directional result was consulted through the article abstract. ↩

  14. Leslie A. DeChurch and Jessica R. Mesmer-Magnus, “The Cognitive Underpinnings of Effective Teamwork: A Meta-Analysis,” Journal of Applied Psychology 95, no. 1 (2010): 32–53, doi:10.1037/a0017328. ↩

  15. Keith G. Provan and H. Brinton Milward, “A Preliminary Theory of Interorganizational Network Effectiveness: A Comparative Study of Four Community Mental Health Systems,” Administrative Science Quarterly 40, no. 1 (1995): 1–33, doi:10.2307/2393698. The explanatory structure used here was available through abstract and secondary characterization rather than inspected full text. See also Keith G. Provan and H. Brinton Milward, “Do Networks Really Work? A Framework for Evaluating Public-Sector Organizational Networks,” Public Administration Review 61, no. 4 (2001): 414–423, doi:10.1111/0033-3352.00045; Keith G. Provan and Patrick Kenis, “Modes of Network Governance,” Journal of Public Administration Research and Theory 18, no. 2 (2008): 229–252, doi:10.1093/jopart/mum015. ↩

  16. David R. Urbach et al., “Introduction of Surgical Safety Checklists in Ontario, Canada,” New England Journal of Medicine 370, no. 11 (2014): 1029–1038, doi:10.1056/NEJMsa1308261. ↩ ↩2

  17. Amy J. Starmer et al., “Changes in Medical Errors after Implementation of a Handoff Program,” New England Journal of Medicine 371, no. 19 (2014): 1803–1812, doi:10.1056/NEJMsa1405556. Detailed results were grounded through the structured abstract. ↩ ↩2

  18. Jennifer L. Rosenthal et al., systematic review of standardized handoffs during care transitions, American Journal of Medical Quality (2018), doi:10.1177/1062860617708244, abstract; Mieke Desmedt et al., “Clinical Handover and Handoff in Healthcare: A Systematic Review of Systematic Reviews,” International Journal for Quality in Health Care 33, no. 1 (2021): mzaa170, doi:10.1093/intqhc/mzaa170, abstract. ↩

  19. Chandra Maddila et al., “Nudge: Accelerating Overdue Pull Requests Towards Completion,” ACM Transactions on Software Engineering and Methodology 32, no. 2 (2022): Article 35, doi:10.1145/3544791. ↩ ↩2

  20. Audris Mockus, Roy T. Fielding, and James D. Herbsleb, “Two Case Studies of Open Source Software Development: Apache and Mozilla,” ACM Transactions on Software Engineering and Methodology 11, no. 3 (2002): 309–346, doi:10.1145/567793.567795. ↩

  21. Benjamin A. Olken, “Direct Democracy and Local Public Goods: Evidence from a Field Experiment in Indonesia,” American Political Science Review 104, no. 2 (2010): 243–267, doi:10.1017/S0003055410000079; Katherine Casey, Rachel Glennerster, and Edward Miguel, Reshaping Institutions: Evidence on Aid Impacts Using a Pre-Analysis Plan, NBER Working Paper 17012 (2011), PDF. ↩ ↩2 ↩3

  22. Amy C. Edmondson, study of error detection and learning in patient-care groups, Journal of Applied Behavioral Science 32, no. 1 (1996), consulted through its abstract. ↩

  23. Robin L. Dillon and Catherine H. Tinsley, “Near-Miss Evaluation Bias as an Obstacle to Organizational Learning: Lessons from NASA,” NASA Technical Reports Server, NTRS 20060047554, PDF; Dillon and Tinsley, related published study in Management Science 54, no. 8 (2008), doi:10.1287/mnsc.1080.0869, abstract; Catherine H. Tinsley and Robin L. Dillon-Merrill, related information-search experiments, SSRN 736245, abstract. ↩

  24. David D. Woods, “The Theory of Graceful Extensibility: Basic Rules That Govern Adaptive Systems,” Environment Systems and Decisions 38 (2018): 433–457, doi:10.1007/s10669-018-9708-3. ↩ ↩2

  25. Viktoria Stray and Nils Brede Moe, “Understanding Coordination in Global Software Engineering: A Mixed-Methods Study on the Use of Meetings and Slack,” Journal of Systems and Software 170 (2020): 110717, doi:10.1016/j.jss.2020.110717; John Meluso et al., “Invisible Labor in Open Source Software Ecosystems,” arXiv:2401.06889, https://arxiv.org/abs/2401.06889. Meluso et al. used a self-selected survey of 142 respondents, so its magnitudes should not be treated as population estimates. Rob Cross, Reb Rebele, and Adam Grant, “Collaborative Overload,” Harvard Business Review (January–February 2016), is a non-peer-reviewed practitioner source and is used only for directional support. ↩

  26. NASA Office of Inspector General, Review of NASA’s Lessons Learned Information System, IG-12-012 (2012), PDF. ↩ ↩2

  27. Bardha Ahmeti et al., “Architecture Decision Records in Practice: An Action Research Study,” 18th European Conference on Software Architecture (2024), doi:10.1007/978-3-031-70797-1_22. The study covered two teams in one company over three months without a control group. ↩

  28. Marcelo Cataldo, James D. Herbsleb, and Kathleen M. Carley, “Socio-Technical Congruence: A Framework for Assessing the Impact of Technical and Work Dependencies on Software Development Productivity,” ESEM 2008, doi:10.1145/1414004.1414008. ↩

  29. Mohammad Safi, Bettina Thude, Frans Brandt, and Robyn Clay-Williams, “The Application of Resilience Assessment Grid in Healthcare: A Scoping Review,” PLOS ONE 17, no. 11 (2022): e0277289, doi:10.1371/journal.pone.0277289. ↩

  30. Julii Brainard and Paul R. Hunter, “Do Complexity-Informed Health Interventions Work? A Scoping Review,” Implementation Science 11 (2016): 127, doi:10.1186/s13012-016-0492-5. ↩

  31. NIST/SEMATECH, “What Are Control Charts?” e-Handbook of Statistical Methods, https://www.itl.nist.gov/div898/handbook/pmc/section3/pmc31.htm. ↩

  32. John Mayne, “Contribution Analysis: Coming of Age?” Evaluation 18, no. 3 (2012): 270–280, doi:10.1177/1356389012451663. ↩

  33. James M. Kendra and Tricia Wachtendorf, “Creativity in Emergency Response after the World Trade Center Attack,” in Beyond September 11th: An Account of Post-Disaster Research (2003), PDF. ↩