Skill v1.0.0
currentAutomated scan100/100name: foundations-cybernetics-vsm description: Applies Beer's VSM and Ashby's Law to diagnose org or agent-system viability. Use when a team or agent hierarchy has coordination, escalation, or requisite-variety problems. compatibility: Portable core only. version: "1.3" last_validated: 2026-09-17
Cybernetics and Viable System Model Foundations
When to Apply
Apply cybernetics-VSM when:
- Org or agent-system steering question — viability, requisite variety, escalation paths
- "Why does this team/system keep failing despite individual competence?" — likely missing S2/S3*/S4
- Recursion across levels — same control pattern at squad / department / company
- Algedonic channel design — when does a critical signal bypass hierarchy and reach S5 directly?
- Variety-engineering — disturbance-to-signal-to-effective-response coverage (Ashby's Law)
Skip and use simpler alternatives when:
- Single team, no recursion, no orchestration question — VSM is overkill
- Org-design question is purely about reporting lines — use a simple RACI, not VSM
- Throughput/bottleneck question — use foundations-theory-of-constraints
- Strategic-interaction question between agents — use foundations-game-theory
- Feedback-loop tuning on a measurable variable — use foundations-control-theory
- The framing imports VSM jargon (S1-S5) without an actual variety/viability problem — risk of decoration; demand the failure signal first
11 canonical cybernetics and VSM primitives for designing viable organizations, control hierarchies, and adaptive systems. Each primitive solves a specific failure mode in how complexity is absorbed, coordinated, and governed. Primitives are domain-agnostic: the same variety-engineering pattern that prevents management overload in an enterprise also prevents orchestrator bottlenecks in an agent swarm; the same algedonic channel that surfaces crises to a board surfaces production incidents to an on-call team.
Contents
- Quick Reference
- Primitive Index
- Formal Supporting Theory
- Misuse Boundaries
- Anti-Patterns
- Decision Checklist
- Composition Recipes
- Workflow
- ASCII Flow
- Navigation
- Fact-Checking
Quick Reference
| # | Primitive | Core Function | When to Reach For It | |
|---|---|---|---|---|
| 1 | Feedback Loops | Regulate behavior via negative (balancing) or amplify via positive (reinforcing) loops | Any adaptive control mechanism; stability vs. growth dynamics | |
| 2 | Ashby's Law of Requisite Variety | Controller must match the variety of the system it governs | Diagnosing under-instrumented control; scaling management layers | |
| 3 | VSM System 1 — Operations | Autonomous operational units that do the actual work | Defining work units, microservices, squads, agent executors | |
| 4 | VSM System 2 — Coordination | Anti-oscillation coordination layer between S1 units | Preventing interference and thrashing between operational units | |
| 5 | VSM System 3 — Internal Control | Here-and-now optimization of the operational environment | Performance management, resource allocation, policy enforcement | |
| 6 | VSM System 3* — Audit Channel | Sporadic direct channel from S3 to S1 bypassing S2 | Spot-checks, audits, compliance sampling; detecting S2 distortion | |
| 7 | VSM System 4 — Intelligence | Outside-and-future scanning; adaptation intelligence | Strategy, environmental scanning, roadmaps, horizon sensing | |
| 8 | VSM System 5 — Identity/Policy | Ultimate authority; closure and identity of the whole | Mission, values, constitutional rules, governance closure | |
| 9 | Recursion Levels | Every viable system contains and is contained in viable systems | Multi-level organizational design; nesting teams, divisions, products | |
| 10 | Variety Engineering | Amplifiers, attenuators, and transducers to balance variety across channels | Reducing information overload; designing dashboards, APIs, interfaces | |
| 11 | Algedonic Channels | High-priority pain/pleasure signals that bypass normal hierarchy levels | Incident escalation, crisis bypass routes, critical alerts |
Primitive Index
Each primitive has a full playbook (definition, when to use, inputs, outputs, failure modes, worked example, sources).
| # | Primitive | Failure Mode It Addresses | |
|---|---|---|---|
| 1 | Feedback Loops | Runaway growth or oscillation from unchecked reinforcing dynamics | |
| 2 | Ashby's Law — Requisite Variety | Control collapse when environmental variety exceeds controller capacity | |
| 3 | VSM S1 — Operations | Centralised execution bottleneck; no operational autonomy | |
| 4 | VSM S2 — Coordination | Thrashing and interference between operational units | |
| 5 | VSM S3 — Internal Control | Local optima divergence; S1 units optimise against each other | |
| 6 | VSM S3* — Audit Channel | S2/S3 filters distort ground truth before it reaches management | |
| 7 | VSM S4 — Intelligence | Strategy-execution gap; S3 unaware of environment shifts | |
| 8 | VSM S5 — Identity/Policy | Identity crisis or policy vacuum; S3/S4 conflict never resolved | |
| 9 | Recursion Levels | Applying VSM at wrong scale; mismatch of model and organisation | |
| 10 | Variety Engineering | Management overload or information starvation from unbalanced variety | |
| 11 | Algedonic Channels | Crisis hidden by normal reporting hierarchy until it is too late |
Formal Supporting Theory
Load `references/formal-theory-map.md` when the task needs more than a primitive lookup: defining the system-in-focus, distinguishing first-order vs. second-order cybernetics, proving an Ashby/requisite-variety claim, mapping VSM systems 1-5 across recursion levels, or separating S3 control, S3* audit, S4 intelligence, S5 policy, and algedonic escalation.
Misuse Boundaries
Load `references/patterns-scenarios-traps.md` before turning VSM into an org chart, central control layer, dashboard scheme, escalation policy, or agent hierarchy. It contains operational scenarios, anti-patterns, known traps, and a compact audit sequence.
Anti-Patterns
| Anti-Pattern | Cybernetics/VSM Diagnosis | Fix | |
|---|---|---|---|
| System 3 collapses System 1 autonomy (micromanagement) | S3 is consuming all operational variety — no recursion depth; Ashby violation | Restore S1 autonomy; S3 sets policy and limits, not execution steps | |
| System 4 disconnected from System 3 (strategy-execution gap) | S4 output never reaches S3; no S3/S4 homeostat | Build explicit S3/S4 interface: shared planning cadence, mutual translation layer | |
| Ashby's Law violated by under-instrumented control | Disturbance distinctions that require different responses are collapsed or unreachable | Define the disturbance classes and response repertoire; add attenuation or amplification where a tested control distinction is missing | |
| Algedonic channel never used — S5 blind to crises | Pain signals absorbed by normal hierarchy; S5 receives filtered reports only | Implement direct bypass route with trigger threshold; test at an interval justified by hazard, disturbance rate, consequence deadline and change events | |
| Recursion confusion — applying VSM at wrong organisational scale | S1/S3/S5 roles assigned to the wrong recursion level | Re-identify the level of recursion; redraw the system boundary before assigning roles | |
| Positive feedback loop with no balancing loop (runaway dynamics) | Reinforcing loop unchecked — growth, debt, or failure cascades | Design an explicit negative feedback loop with a goal variable and measured deviation | |
| S2 coordination layer absent — unit thrashing | S1 units interfere without coordination signals | Introduce S2 scheduling, resource-sharing protocols, or synchronisation mechanisms | |
| S3* audit channel treated as normal management reporting | Spot-check becomes routine; S1 adapts and Goodharts the signal | Keep S3* sporadic and surprise-based; vary timing and scope | |
| Variety amplified without attenuation at higher levels | Upper levels receive raw operational noise; decision paralysis | Apply variety attenuation (aggregation, exception filters) before variety reaches S3/S4 | |
| S5 identity undefined — policy vacuum | S3/S4 conflicts escalate without resolution; ad-hoc decisions contradict each other | Define S5 closure: mission, constraints, values; run S3/S4 conflicts through S5 reference frame | |
| Human oversight of an agent fleet staffed, not engineered | Reviewer headcount is added without mapping outcome-relevant agent behaviours to detectable signals and effective interventions | Add triage, summarisation, tiered escalation, and tested intervention paths; publish the mapping and escalation SLA, not only a rota |
Decision Checklist
- [ ] Control loop needed: Is there a variable that must stay within bounds? → feedback loop (#1)
- [ ] Management layer overwhelmed: Does control complexity exceed controller capacity? → Ashby's Law audit (#2) + variety engineering (#10)
- [ ] Operational units defined: Are execution units autonomous with clear scope? → VSM S1 (#3)
- [ ] Unit interference observed: Do operational units conflict or thrash? → VSM S2 coordination (#4)
- [ ] Optimisation divergence: Are local optima conflicting with system-level goals? → VSM S3 (#5)
- [ ] Ground truth distortion: Is management receiving filtered or misleading data? → VSM S3* audit (#6)
- [ ] Strategy-execution gap: Is there no mechanism for environmental change to inform operations? → VSM S4 (#7)
- [ ] Identity or policy conflict: Do teams lack a shared frame for resolving disagreements? → VSM S5 (#8)
- [ ] Model scale mismatch: Is the VSM being applied to the wrong organisational level? → recursion levels (#9)
- [ ] Information overload or starvation: Are channels between levels carrying the wrong amount of variety? → variety engineering (#10)
- [ ] Crisis hidden in normal reporting: Do critical alerts get delayed by hierarchy? → algedonic channel (#11)
Composition Recipes
Agent-Team Topology Audit
Goal: diagnose whether an agent hierarchy is viable and where failures will occur.
Stack:
- VSM S1 (#3) — identify operational agent units and verify autonomy
- VSM S2 (#4) — check for coordination signals between units; absence = thrashing risk
- VSM S3 (#5) — confirm orchestrator has S3 function: bounded operational/resource policy within S5 identity and ultimate policy, rather than micro-execution
- Ashby's Law (#2) — map decision-relevant disturbance classes to available responses; treat raw state counts only as a diagnostic proxy
- Variety Engineering (#10) — add attenuators (summarisation, exception routing) if orchestrator is overwhelmed
- Algedonic channel (#11) — ensure critical failures bypass normal reporting to human-in-the-loop or S5
Output: viability gap report with specific role assignments and missing interfaces. Do not present a count of alerts, agents, labels, or dashboard states as a cardinal proof of requisite variety unless the state partition and required response mapping are defined.
Inputs: S1 agent units with scope and autonomy level; outcome-relevant disturbance classes; the observations that distinguish them; available responses; and constraints on when each response remains effective.
Rules: Ashby check — for every outcome-relevant disturbance class, verify that the orchestrator can detect the distinction and select an effective response under real timing, authority, and resource constraints. Missing or coupled responses identify the gap; do not infer it by subtracting lever counts from state counts. S2 is absent if S1 units share a resource without an explicit coordination protocol. Conant-Ashby model-adequacy check: verify that the regulator model represents the distinctions required to choose among effective responses; attenuate demand or improve sensing/model/action capacity where the mapping fails.
Outputs: Viability gap table; disturbance-to-observation-to-response mapping with uncovered distinctions and response constraints; missing interfaces; and a recommended structural change per gap.
Human-oversight variety condition. Telukunta et al. (2026, arXiv:2608.10153) propose V_human × G ≥ V_agents as a conceptual framing for amplification through triage, summarisation, and tiered escalation. Do not operationalize these terms as raw cardinal counts or use the inequality as a deployment proof. Test whether peak outcome-relevant behaviours are detected, routed within the SLA, and met by an authorized effective intervention. Headcount without those paths is not an oversight design.
Organisational Design for a Startup
Goal: design a lightweight management structure that scales without creating command bottlenecks.
Stack:
- Recursion levels (#9) — identify the two or three levels the startup actually needs (whole company → product area → squad)
- VSM S1 (#3) — define autonomous squad boundaries with clear operational scope
- VSM S3* (#6) — establish audit/spot-check mechanism so founders maintain ground truth as company grows
- VSM S4 (#7) — assign who owns environmental scanning and translates it into strategy
- VSM S5 (#8) — write a one-page identity document: mission, non-negotiable constraints, value principles
- Feedback loops (#1) — design at least one balancing loop per key performance variable (burn rate, NPS, lead time)
Inputs: Squad scopes; leadership roles mapped to S3/S4/S5; strategy cadence; recurring outcome-relevant operating disturbances; sensing paths; and available responses.
Rules: Each of S1–S5 must be present and named. For each material disturbance, verify a sensing and response path at the correct recursion level; flag uncovered distinctions rather than comparing counts. S3* audit and S4-to-S5 cadence should be set from risk and change rate, then tested.
Outputs: Role-to-system mapping; disturbance-response coverage table; missing-system list; and recommended structural change per uncovered or ineffective path.
Worked example: A SaaS company maps four product squads to S1, shared-roadmap coordination to S2, resource policy and operational audit to S3/S3*, market scanning to S4, and mission constraints to S5. The audit lists material disturbances such as a cross-squad dependency conflict, a production incident, and a market change. If the market-change signal reaches S4 but no decision path can alter portfolio allocation, that distinction lacks an effective response; add the S4-to-S5 decision path or delegate bounded authority. Counts of surfaces, cadences, or management levers do not establish the gap.
Incident Escalation as Algedonic Channel
Goal: ensure production crises reach decision authority fast, bypassing normal ticket queues.
Stack:
- Algedonic channel (#11) — define trigger threshold (e.g., p99 latency > 2× baseline for 5 min)
- VSM S5 (#8) — confirm who holds S5 authority for incident closure decisions
- Feedback loops (#1) — implement a balancing loop that activates on trigger: alert → diagnosis → rollback → verify recovery
- VSM S3 (#6) — use the incident post-mortem as the S3 audit: compare what S3 saw vs. ground truth
- Variety Engineering (#10) — ensure incident dashboards attenuate noise; only deviation-from-normal reaches on-call
Inputs: Feedback loops present (count and type — balancing or reinforcing); latency of each loop (time from signal to corrective action, in minutes or hours); S2 coordination protocols in place (count and description, e.g., "on-call handoff protocol", "shared incident channel"); environment change rate (how quickly the production environment can shift state, e.g., deploy frequency × distinct failure modes per week).
Rules: Compare detection-plus-response latency with the consequence deadline and disturbance evolution, including overlapping changes. A 30-minute response with deploys every 10 minutes is an investigation signal, not an automatic violation: deploy frequency alone does not determine the effective intervention window; S2 coordination protocols required when ≥2 S1 units (e.g., on-call teams, services) share a resource (queue, database, API gateway) — absence is a critical gap; S3* post-mortem audit must compare what S3 saw (dashboards, alerts) against ground truth (actual failure timeline) — run after every P1 incident; algedonic trigger threshold must be defined and tested at a risk-based interval and after material channel/authority changes.
Outputs: Loop diagram (each loop with type, goal variable, latency, and status — active/missing); latency table (loop name, measured latency, environment change rate, pass/fail); missing-protocol list (each shared resource without an S2 coordination protocol flagged as H severity); recommended structural change per gap (e.g., "reduce alert-to-page latency from 15 min to <5 min", "add shared-queue ownership protocol between service A and B").
Scaling a Platform Team
Goal: prevent a platform team from becoming a bottleneck as it serves multiple product teams.
Stack:
- Ashby's Law (#2) — map outcome-relevant request/disturbance classes to observable signals and effective platform responses
- Variety Engineering (#10) — apply amplifiers (self-service APIs, documentation, inner-source) to expand platform's effective variety; apply attenuators (standard interfaces, request templates) on the demand side
- VSM S2 (#4) — add coordination protocol between consuming teams to prevent conflicting platform requests
- VSM S3 (#5) — platform S3 sets platform-wide policy; individual platform sub-teams are S1 units with autonomy within policy
- Feedback loops (#1) — measure platform lead time and consumer satisfaction as balancing-loop goal variables
Output: platform operating model with variety audit, self-service expansion plan, and S3 policy layer.
Inputs: Outcome-relevant request classes, their signals, effective platform responses and constraints; coordination protocols; lead time; and consumer outcome baseline.
Rules: Flag a variety gap when an outcome-relevant request distinction cannot be detected or lacks an effective response under load. Resolve it through self-service/action amplification or demand attenuation. Require coordination for conflicting requests and an explicit S3 policy boundary. Diagnose stagnant lead time before assigning it to S2 or S3.
Outputs: Request-class-to-response coverage table; self-service expansion plan; S3 policy boundary; missing coordination protocols; and recommended structural change per gap.
Workflow
- Identify the system boundary and the level of recursion you are working at (use recursion levels #9 first).
- Map the five VSM systems to actual roles, teams, or agent components.
- Check for missing or collapsed systems — use the Decision Checklist.
- Apply Ashby's Law (#2) by testing the detection and effective-response path for each outcome-relevant disturbance class.
- Design or audit variety engineering (#10) mechanisms on each inter-level channel.
- Confirm algedonic channels (#11) exist and are tested.
- For specific failure modes, open the per-primitive playbook in `assets/templates/cybernetics-vsm/`.
- For multi-failure scenarios, use the Composition Recipes above.
ASCII Flow
Viability or organizational-control problem-> Set system boundary and recursion level-> Map Systems 1-5 to real roles, teams, or agents-> Check Ashby response coverage+-- class undetected or response ineffective -> attenuate demand or amplify sensing/action capacity+-- material classes covered -> audit channels and coupled disturbances-> Verify algedonic alerts and policy/intelligence balance-> Return missing systems, channel fixes, and recursion risks
Navigation
- Per-primitive playbooks: `assets/templates/cybernetics-vsm/` (one file per primitive)
- Composition guide: `assets/templates/cybernetics-vsm/README.md`
- Primitives overview: `references/primitives-overview.md`
- Formal theory map: `references/formal-theory-map.md`
- Patterns, scenarios, and traps: `references/patterns-scenarios-traps.md`
- Sources: `data/sources.json`
Related Skills
<!-- Consumer skills will add cross-links here when their applied recipe layers are built. --> <!-- Do not add cross-links to this file directly — consumer skills link in, not out. -->
Fact-Checking
- Stafford Beer: VSM systems 1–5, algedonic channels, recursion levels, and variety engineering are defined in Beer 1972 (_Brain of the Firm_), Beer 1979 (_Heart of Enterprise_), and Beer 1985 (_Diagnosing the System for Organizations_). Verify claims about specific Beer definitions against these primary texts. 2026-07 correction: per-primitive playbook citations previously attributed each VSM system to its own numbered chapter of _Brain of the Firm_ (e.g., "Ch. 3: System One," "Ch. 8: System Five"). The verified table of contents shows no such one-system-per-chapter structure — Systems One–Three are treated together in one section ("Autonomics"), System Four in "Environments of Decision," and System Five in "The Multinode"; recursion and algedonic channels are not confined to single dedicated chapters at all. Citations in
assets/templates/cybernetics-vsm/were corrected to cite by section title rather than a fabricated chapter number. Chapter-level citations to Beer 1985, Hoverstadt 2009, and Schwaninger 2006 have not been independently re-verified against primary copies in this pass — treat their specific chapter numbers as approximate until confirmed. - Project Cybersyn (Chile, 1971–1973): the most-cited real-world VSM deployment is also the most mythologized. Per Medina 2011 (_Cybernetic Revolutionaries_, MIT Press — the primary archival history), Cybersyn was a telex network plus one mainframe with roughly daily-lagged data, not a real-time networked control system; the Opsroom was never fully deployed (its move to the presidential palace was approved only three days before the 11 September 1973 coup); only ~26.7% of nationalized firms were incorporated by May 1973; and the October 1972 truckers'-strike response was a genuine, documented operational success for the S1/S2 layer. See
references/patterns-scenarios-traps.md→ "Historical Grounding: Project Cybersyn" for the full fact-vs-myth table before citing this case as precedent. - W. Ross Ashby: Law of Requisite Variety is from Ashby 1956 (_An Introduction to Cybernetics_, ch. 11). The formal statement is W(error) ≤ V(disturbance) − V(regulator). Verify quantitative claims against the original. Note: Siegenfeld & Bar-Yam (2025, _Entropy_, 27(8), 835, DOI: 10.3390/e27080835; PMC-indexed as PMC12385218) propose a multi-scale generalisation of Ashby's Law showing that variety requirements are scale-dependent — a relevant refinement for hierarchical/recursive agent architectures where the same system exhibits different variety at different recursion levels. Treat as a clarification of application scope, not a revision of the original law.
- Requisite variety in AI-oversight regulation: the
V_human × G ≥ V_agentsframing above is from Telukunta, Lilis & Baron (2026, arXiv:2608.10153, submitted 10 August 2026), which builds on Beer's VSM for enterprise agent fleets. Evidence grade: C (preprint, not peer-reviewed). Treat the inequality and CASE architecture as a conceptual proposal, not a validated quantitative condition; operational evidence must come from disturbance-response coverage and intervention tests. The underlying cybernetic sources are stronger evidence for the qualitative need for requisite variety, not for multiplying raw oversight counts. On the regulatory hook: EU AI Act Article 14 obligations differ by high-risk category and date; verify the applicable provision before making a current compliance claim. - Norbert Wiener: Feedback and cybernetics foundations from Wiener 1948 (_Cybernetics: Or Control and Communication in the Animal and the Machine_). Positive/negative feedback terminology is consistent with Wiener's original usage.
- Espinosa & Walker: VSM applied to complexity and sustainability in _A Complexity Approach to Sustainability_ (2011). Recursion and viable-systems analysis in real organisations.
- Schwaninger: Intelligent organisations and VSM application in _Intelligent Organizations_ (2006). Apply numeric claims (e.g., performance improvement percentages) only when derived from primary case studies, not secondary summaries.
- Hoverstadt: Practical VSM application in _The Fractal Organization_ (2009). Patterns cited from this source are practitioner heuristics — verify against Beer's original formalism before treating as universal.
- Mechanism effectiveness is context-specific. Test variety-engineering interventions on a constrained scope before rolling out system-wide.
Learnings Loop
When prior decisions or pitfalls are relevant, consult learnings.consolidated.md if present; use learnings.md only for needed history or as the available fallback. Otherwise skip both.
After applying it, if you encountered a pattern worth remembering, a mistake worth preventing, or a domain fact that surprised you, append one dated bullet to learnings.md via agents-skills-feedback-loop/scripts/append_learning.py. Do not modify SKILL.md itself.