<< All versions

Skill v1.0.0

currentAutomated scan100/100
vasilyu1983/ai-agents-public/foundations-cybernetics-vsm
──Details
PublishedSeptember 9, 2026 at 08:06 AM
Content Hashsha256:39fc80813da7ffa2...
Git SHA3d5bc7c5b826
──Files
Files (1 file, 27.0 KB)
SKILL.md27.0 KBactive
SKILL.md · 282 lines · 27.0 KB

name: foundations-cybernetics-vsm description: Applies Beer's VSM and Ashby's Law to diagnose org or agent-system viability. Use when a team or agent hierarchy has coordination, escalation, or requisite-variety problems. compatibility: Portable core only. version: "1.3" last_validated: 2026-09-17


Cybernetics and Viable System Model Foundations

When to Apply

Apply cybernetics-VSM when:

  • Org or agent-system steering question — viability, requisite variety, escalation paths
  • "Why does this team/system keep failing despite individual competence?" — likely missing S2/S3*/S4
  • Recursion across levels — same control pattern at squad / department / company
  • Algedonic channel design — when does a critical signal bypass hierarchy and reach S5 directly?
  • Variety-engineering — disturbance-to-signal-to-effective-response coverage (Ashby's Law)

Skip and use simpler alternatives when:

  • Single team, no recursion, no orchestration question — VSM is overkill
  • Org-design question is purely about reporting lines — use a simple RACI, not VSM
  • Throughput/bottleneck question — use foundations-theory-of-constraints
  • Strategic-interaction question between agents — use foundations-game-theory
  • Feedback-loop tuning on a measurable variable — use foundations-control-theory
  • The framing imports VSM jargon (S1-S5) without an actual variety/viability problem — risk of decoration; demand the failure signal first

11 canonical cybernetics and VSM primitives for designing viable organizations, control hierarchies, and adaptive systems. Each primitive solves a specific failure mode in how complexity is absorbed, coordinated, and governed. Primitives are domain-agnostic: the same variety-engineering pattern that prevents management overload in an enterprise also prevents orchestrator bottlenecks in an agent swarm; the same algedonic channel that surfaces crises to a board surfaces production incidents to an on-call team.

Contents


Quick Reference

#PrimitiveCore FunctionWhen to Reach For It
1Feedback LoopsRegulate behavior via negative (balancing) or amplify via positive (reinforcing) loopsAny adaptive control mechanism; stability vs. growth dynamics
2Ashby's Law of Requisite VarietyController must match the variety of the system it governsDiagnosing under-instrumented control; scaling management layers
3VSM System 1 — OperationsAutonomous operational units that do the actual workDefining work units, microservices, squads, agent executors
4VSM System 2 — CoordinationAnti-oscillation coordination layer between S1 unitsPreventing interference and thrashing between operational units
5VSM System 3 — Internal ControlHere-and-now optimization of the operational environmentPerformance management, resource allocation, policy enforcement
6VSM System 3* — Audit ChannelSporadic direct channel from S3 to S1 bypassing S2Spot-checks, audits, compliance sampling; detecting S2 distortion
7VSM System 4 — IntelligenceOutside-and-future scanning; adaptation intelligenceStrategy, environmental scanning, roadmaps, horizon sensing
8VSM System 5 — Identity/PolicyUltimate authority; closure and identity of the wholeMission, values, constitutional rules, governance closure
9Recursion LevelsEvery viable system contains and is contained in viable systemsMulti-level organizational design; nesting teams, divisions, products
10Variety EngineeringAmplifiers, attenuators, and transducers to balance variety across channelsReducing information overload; designing dashboards, APIs, interfaces
11Algedonic ChannelsHigh-priority pain/pleasure signals that bypass normal hierarchy levelsIncident escalation, crisis bypass routes, critical alerts

Primitive Index

Each primitive has a full playbook (definition, when to use, inputs, outputs, failure modes, worked example, sources).

#PrimitiveFailure Mode It Addresses
1Feedback LoopsRunaway growth or oscillation from unchecked reinforcing dynamics
2Ashby's Law — Requisite VarietyControl collapse when environmental variety exceeds controller capacity
3VSM S1 — OperationsCentralised execution bottleneck; no operational autonomy
4VSM S2 — CoordinationThrashing and interference between operational units
5VSM S3 — Internal ControlLocal optima divergence; S1 units optimise against each other
6VSM S3* — Audit ChannelS2/S3 filters distort ground truth before it reaches management
7VSM S4 — IntelligenceStrategy-execution gap; S3 unaware of environment shifts
8VSM S5 — Identity/PolicyIdentity crisis or policy vacuum; S3/S4 conflict never resolved
9Recursion LevelsApplying VSM at wrong scale; mismatch of model and organisation
10Variety EngineeringManagement overload or information starvation from unbalanced variety
11Algedonic ChannelsCrisis hidden by normal reporting hierarchy until it is too late

Formal Supporting Theory

Load `references/formal-theory-map.md` when the task needs more than a primitive lookup: defining the system-in-focus, distinguishing first-order vs. second-order cybernetics, proving an Ashby/requisite-variety claim, mapping VSM systems 1-5 across recursion levels, or separating S3 control, S3* audit, S4 intelligence, S5 policy, and algedonic escalation.

Misuse Boundaries

Load `references/patterns-scenarios-traps.md` before turning VSM into an org chart, central control layer, dashboard scheme, escalation policy, or agent hierarchy. It contains operational scenarios, anti-patterns, known traps, and a compact audit sequence.


Anti-Patterns

Anti-PatternCybernetics/VSM DiagnosisFix
System 3 collapses System 1 autonomy (micromanagement)S3 is consuming all operational variety — no recursion depth; Ashby violationRestore S1 autonomy; S3 sets policy and limits, not execution steps
System 4 disconnected from System 3 (strategy-execution gap)S4 output never reaches S3; no S3/S4 homeostatBuild explicit S3/S4 interface: shared planning cadence, mutual translation layer
Ashby's Law violated by under-instrumented controlDisturbance distinctions that require different responses are collapsed or unreachableDefine the disturbance classes and response repertoire; add attenuation or amplification where a tested control distinction is missing
Algedonic channel never used — S5 blind to crisesPain signals absorbed by normal hierarchy; S5 receives filtered reports onlyImplement direct bypass route with trigger threshold; test at an interval justified by hazard, disturbance rate, consequence deadline and change events
Recursion confusion — applying VSM at wrong organisational scaleS1/S3/S5 roles assigned to the wrong recursion levelRe-identify the level of recursion; redraw the system boundary before assigning roles
Positive feedback loop with no balancing loop (runaway dynamics)Reinforcing loop unchecked — growth, debt, or failure cascadesDesign an explicit negative feedback loop with a goal variable and measured deviation
S2 coordination layer absent — unit thrashingS1 units interfere without coordination signalsIntroduce S2 scheduling, resource-sharing protocols, or synchronisation mechanisms
S3* audit channel treated as normal management reportingSpot-check becomes routine; S1 adapts and Goodharts the signalKeep S3* sporadic and surprise-based; vary timing and scope
Variety amplified without attenuation at higher levelsUpper levels receive raw operational noise; decision paralysisApply variety attenuation (aggregation, exception filters) before variety reaches S3/S4
S5 identity undefined — policy vacuumS3/S4 conflicts escalate without resolution; ad-hoc decisions contradict each otherDefine S5 closure: mission, constraints, values; run S3/S4 conflicts through S5 reference frame
Human oversight of an agent fleet staffed, not engineeredReviewer headcount is added without mapping outcome-relevant agent behaviours to detectable signals and effective interventionsAdd triage, summarisation, tiered escalation, and tested intervention paths; publish the mapping and escalation SLA, not only a rota

Decision Checklist

  • [ ] Control loop needed: Is there a variable that must stay within bounds? → feedback loop (#1)
  • [ ] Management layer overwhelmed: Does control complexity exceed controller capacity? → Ashby's Law audit (#2) + variety engineering (#10)
  • [ ] Operational units defined: Are execution units autonomous with clear scope? → VSM S1 (#3)
  • [ ] Unit interference observed: Do operational units conflict or thrash? → VSM S2 coordination (#4)
  • [ ] Optimisation divergence: Are local optima conflicting with system-level goals? → VSM S3 (#5)
  • [ ] Ground truth distortion: Is management receiving filtered or misleading data? → VSM S3* audit (#6)
  • [ ] Strategy-execution gap: Is there no mechanism for environmental change to inform operations? → VSM S4 (#7)
  • [ ] Identity or policy conflict: Do teams lack a shared frame for resolving disagreements? → VSM S5 (#8)
  • [ ] Model scale mismatch: Is the VSM being applied to the wrong organisational level? → recursion levels (#9)
  • [ ] Information overload or starvation: Are channels between levels carrying the wrong amount of variety? → variety engineering (#10)
  • [ ] Crisis hidden in normal reporting: Do critical alerts get delayed by hierarchy? → algedonic channel (#11)

Composition Recipes

Agent-Team Topology Audit

Goal: diagnose whether an agent hierarchy is viable and where failures will occur.

Stack:

  1. VSM S1 (#3) — identify operational agent units and verify autonomy
  2. VSM S2 (#4) — check for coordination signals between units; absence = thrashing risk
  3. VSM S3 (#5) — confirm orchestrator has S3 function: bounded operational/resource policy within S5 identity and ultimate policy, rather than micro-execution
  4. Ashby's Law (#2) — map decision-relevant disturbance classes to available responses; treat raw state counts only as a diagnostic proxy
  5. Variety Engineering (#10) — add attenuators (summarisation, exception routing) if orchestrator is overwhelmed
  6. Algedonic channel (#11) — ensure critical failures bypass normal reporting to human-in-the-loop or S5

Output: viability gap report with specific role assignments and missing interfaces. Do not present a count of alerts, agents, labels, or dashboard states as a cardinal proof of requisite variety unless the state partition and required response mapping are defined.

Inputs: S1 agent units with scope and autonomy level; outcome-relevant disturbance classes; the observations that distinguish them; available responses; and constraints on when each response remains effective.

Rules: Ashby check — for every outcome-relevant disturbance class, verify that the orchestrator can detect the distinction and select an effective response under real timing, authority, and resource constraints. Missing or coupled responses identify the gap; do not infer it by subtracting lever counts from state counts. S2 is absent if S1 units share a resource without an explicit coordination protocol. Conant-Ashby model-adequacy check: verify that the regulator model represents the distinctions required to choose among effective responses; attenuate demand or improve sensing/model/action capacity where the mapping fails.

Outputs: Viability gap table; disturbance-to-observation-to-response mapping with uncovered distinctions and response constraints; missing interfaces; and a recommended structural change per gap.

Human-oversight variety condition. Telukunta et al. (2026, arXiv:2608.10153) propose V_human × G ≥ V_agents as a conceptual framing for amplification through triage, summarisation, and tiered escalation. Do not operationalize these terms as raw cardinal counts or use the inequality as a deployment proof. Test whether peak outcome-relevant behaviours are detected, routed within the SLA, and met by an authorized effective intervention. Headcount without those paths is not an oversight design.


Organisational Design for a Startup

Goal: design a lightweight management structure that scales without creating command bottlenecks.

Stack:

  1. Recursion levels (#9) — identify the two or three levels the startup actually needs (whole company → product area → squad)
  2. VSM S1 (#3) — define autonomous squad boundaries with clear operational scope
  3. VSM S3* (#6) — establish audit/spot-check mechanism so founders maintain ground truth as company grows
  4. VSM S4 (#7) — assign who owns environmental scanning and translates it into strategy
  5. VSM S5 (#8) — write a one-page identity document: mission, non-negotiable constraints, value principles
  6. Feedback loops (#1) — design at least one balancing loop per key performance variable (burn rate, NPS, lead time)

Inputs: Squad scopes; leadership roles mapped to S3/S4/S5; strategy cadence; recurring outcome-relevant operating disturbances; sensing paths; and available responses.

Rules: Each of S1–S5 must be present and named. For each material disturbance, verify a sensing and response path at the correct recursion level; flag uncovered distinctions rather than comparing counts. S3* audit and S4-to-S5 cadence should be set from risk and change rate, then tested.

Outputs: Role-to-system mapping; disturbance-response coverage table; missing-system list; and recommended structural change per uncovered or ineffective path.

Worked example: A SaaS company maps four product squads to S1, shared-roadmap coordination to S2, resource policy and operational audit to S3/S3*, market scanning to S4, and mission constraints to S5. The audit lists material disturbances such as a cross-squad dependency conflict, a production incident, and a market change. If the market-change signal reaches S4 but no decision path can alter portfolio allocation, that distinction lacks an effective response; add the S4-to-S5 decision path or delegate bounded authority. Counts of surfaces, cadences, or management levers do not establish the gap.


Incident Escalation as Algedonic Channel

Goal: ensure production crises reach decision authority fast, bypassing normal ticket queues.

Stack:

  1. Algedonic channel (#11) — define trigger threshold (e.g., p99 latency > 2× baseline for 5 min)
  2. VSM S5 (#8) — confirm who holds S5 authority for incident closure decisions
  3. Feedback loops (#1) — implement a balancing loop that activates on trigger: alert → diagnosis → rollback → verify recovery
  4. VSM S3 (#6) — use the incident post-mortem as the S3 audit: compare what S3 saw vs. ground truth
  5. Variety Engineering (#10) — ensure incident dashboards attenuate noise; only deviation-from-normal reaches on-call

Inputs: Feedback loops present (count and type — balancing or reinforcing); latency of each loop (time from signal to corrective action, in minutes or hours); S2 coordination protocols in place (count and description, e.g., "on-call handoff protocol", "shared incident channel"); environment change rate (how quickly the production environment can shift state, e.g., deploy frequency × distinct failure modes per week).

Rules: Compare detection-plus-response latency with the consequence deadline and disturbance evolution, including overlapping changes. A 30-minute response with deploys every 10 minutes is an investigation signal, not an automatic violation: deploy frequency alone does not determine the effective intervention window; S2 coordination protocols required when ≥2 S1 units (e.g., on-call teams, services) share a resource (queue, database, API gateway) — absence is a critical gap; S3* post-mortem audit must compare what S3 saw (dashboards, alerts) against ground truth (actual failure timeline) — run after every P1 incident; algedonic trigger threshold must be defined and tested at a risk-based interval and after material channel/authority changes.

Outputs: Loop diagram (each loop with type, goal variable, latency, and status — active/missing); latency table (loop name, measured latency, environment change rate, pass/fail); missing-protocol list (each shared resource without an S2 coordination protocol flagged as H severity); recommended structural change per gap (e.g., "reduce alert-to-page latency from 15 min to <5 min", "add shared-queue ownership protocol between service A and B").


Scaling a Platform Team

Goal: prevent a platform team from becoming a bottleneck as it serves multiple product teams.

Stack:

  1. Ashby's Law (#2) — map outcome-relevant request/disturbance classes to observable signals and effective platform responses
  2. Variety Engineering (#10) — apply amplifiers (self-service APIs, documentation, inner-source) to expand platform's effective variety; apply attenuators (standard interfaces, request templates) on the demand side
  3. VSM S2 (#4) — add coordination protocol between consuming teams to prevent conflicting platform requests
  4. VSM S3 (#5) — platform S3 sets platform-wide policy; individual platform sub-teams are S1 units with autonomy within policy
  5. Feedback loops (#1) — measure platform lead time and consumer satisfaction as balancing-loop goal variables

Output: platform operating model with variety audit, self-service expansion plan, and S3 policy layer.

Inputs: Outcome-relevant request classes, their signals, effective platform responses and constraints; coordination protocols; lead time; and consumer outcome baseline.

Rules: Flag a variety gap when an outcome-relevant request distinction cannot be detected or lacks an effective response under load. Resolve it through self-service/action amplification or demand attenuation. Require coordination for conflicting requests and an explicit S3 policy boundary. Diagnose stagnant lead time before assigning it to S2 or S3.

Outputs: Request-class-to-response coverage table; self-service expansion plan; S3 policy boundary; missing coordination protocols; and recommended structural change per gap.


Workflow

  1. Identify the system boundary and the level of recursion you are working at (use recursion levels #9 first).
  2. Map the five VSM systems to actual roles, teams, or agent components.
  3. Check for missing or collapsed systems — use the Decision Checklist.
  4. Apply Ashby's Law (#2) by testing the detection and effective-response path for each outcome-relevant disturbance class.
  5. Design or audit variety engineering (#10) mechanisms on each inter-level channel.
  6. Confirm algedonic channels (#11) exist and are tested.
  7. For specific failure modes, open the per-primitive playbook in `assets/templates/cybernetics-vsm/`.
  8. For multi-failure scenarios, use the Composition Recipes above.

ASCII Flow

text
Viability or organizational-control problem
-> Set system boundary and recursion level
-> Map Systems 1-5 to real roles, teams, or agents
-> Check Ashby response coverage
+-- class undetected or response ineffective -> attenuate demand or amplify sensing/action capacity
+-- material classes covered -> audit channels and coupled disturbances
-> Verify algedonic alerts and policy/intelligence balance
-> Return missing systems, channel fixes, and recursion risks

Navigation

Related Skills

<!-- Consumer skills will add cross-links here when their applied recipe layers are built. --> <!-- Do not add cross-links to this file directly — consumer skills link in, not out. -->


Fact-Checking

  • Stafford Beer: VSM systems 1–5, algedonic channels, recursion levels, and variety engineering are defined in Beer 1972 (_Brain of the Firm_), Beer 1979 (_Heart of Enterprise_), and Beer 1985 (_Diagnosing the System for Organizations_). Verify claims about specific Beer definitions against these primary texts. 2026-07 correction: per-primitive playbook citations previously attributed each VSM system to its own numbered chapter of _Brain of the Firm_ (e.g., "Ch. 3: System One," "Ch. 8: System Five"). The verified table of contents shows no such one-system-per-chapter structure — Systems One–Three are treated together in one section ("Autonomics"), System Four in "Environments of Decision," and System Five in "The Multinode"; recursion and algedonic channels are not confined to single dedicated chapters at all. Citations in assets/templates/cybernetics-vsm/ were corrected to cite by section title rather than a fabricated chapter number. Chapter-level citations to Beer 1985, Hoverstadt 2009, and Schwaninger 2006 have not been independently re-verified against primary copies in this pass — treat their specific chapter numbers as approximate until confirmed.
  • Project Cybersyn (Chile, 1971–1973): the most-cited real-world VSM deployment is also the most mythologized. Per Medina 2011 (_Cybernetic Revolutionaries_, MIT Press — the primary archival history), Cybersyn was a telex network plus one mainframe with roughly daily-lagged data, not a real-time networked control system; the Opsroom was never fully deployed (its move to the presidential palace was approved only three days before the 11 September 1973 coup); only ~26.7% of nationalized firms were incorporated by May 1973; and the October 1972 truckers'-strike response was a genuine, documented operational success for the S1/S2 layer. See references/patterns-scenarios-traps.md → "Historical Grounding: Project Cybersyn" for the full fact-vs-myth table before citing this case as precedent.
  • W. Ross Ashby: Law of Requisite Variety is from Ashby 1956 (_An Introduction to Cybernetics_, ch. 11). The formal statement is W(error) ≤ V(disturbance) − V(regulator). Verify quantitative claims against the original. Note: Siegenfeld & Bar-Yam (2025, _Entropy_, 27(8), 835, DOI: 10.3390/e27080835; PMC-indexed as PMC12385218) propose a multi-scale generalisation of Ashby's Law showing that variety requirements are scale-dependent — a relevant refinement for hierarchical/recursive agent architectures where the same system exhibits different variety at different recursion levels. Treat as a clarification of application scope, not a revision of the original law.
  • Requisite variety in AI-oversight regulation: the V_human × G ≥ V_agents framing above is from Telukunta, Lilis & Baron (2026, arXiv:2608.10153, submitted 10 August 2026), which builds on Beer's VSM for enterprise agent fleets. Evidence grade: C (preprint, not peer-reviewed). Treat the inequality and CASE architecture as a conceptual proposal, not a validated quantitative condition; operational evidence must come from disturbance-response coverage and intervention tests. The underlying cybernetic sources are stronger evidence for the qualitative need for requisite variety, not for multiplying raw oversight counts. On the regulatory hook: EU AI Act Article 14 obligations differ by high-risk category and date; verify the applicable provision before making a current compliance claim.
  • Norbert Wiener: Feedback and cybernetics foundations from Wiener 1948 (_Cybernetics: Or Control and Communication in the Animal and the Machine_). Positive/negative feedback terminology is consistent with Wiener's original usage.
  • Espinosa & Walker: VSM applied to complexity and sustainability in _A Complexity Approach to Sustainability_ (2011). Recursion and viable-systems analysis in real organisations.
  • Schwaninger: Intelligent organisations and VSM application in _Intelligent Organizations_ (2006). Apply numeric claims (e.g., performance improvement percentages) only when derived from primary case studies, not secondary summaries.
  • Hoverstadt: Practical VSM application in _The Fractal Organization_ (2009). Patterns cited from this source are practitioner heuristics — verify against Beer's original formalism before treating as universal.
  • Mechanism effectiveness is context-specific. Test variety-engineering interventions on a constrained scope before rolling out system-wide.

Learnings Loop

When prior decisions or pitfalls are relevant, consult learnings.consolidated.md if present; use learnings.md only for needed history or as the available fallback. Otherwise skip both.

After applying it, if you encountered a pattern worth remembering, a mistake worth preventing, or a domain fact that surprised you, append one dated bullet to learnings.md via agents-skills-feedback-loop/scripts/append_learning.py. Do not modify SKILL.md itself.

All versions