Failure Modes – Purpose and Scope
Failure Modes describe recurrent, structural patterns of AI system behavior through which misalignment, incentive distortion, or oversight gaps produce outputs that may appear coherent or confident while diverging from truth, user intent, or governance expectations.
These are not isolated errors or bugs, but systemic phenomenon shaped by architecture, optimization pressures, resource constraints, and unmonitored autonomy. When left unaddressed, failure modes tend to repeat, reinforce, and proliferate.
Within the Integrated Governance Stack (IGS), Failure Modes function as diagnostic lenses, early warning indicators, and governance-relevant phenomenon. They support governance-relevant detection within the Relational Layer through observable patterns of output (such as confidence, omission, tone, and escalation) even when the internal model mechanisms remain opaque.
This body of work introduces this application of Failure Modes under the designation “Qualitative Mechanistic Interpretability.”
Qualitative Mechanistic Interpretability (QMI)
Is the failure present — yes or no?
A governance-facing detection and triage layer based on observable output patterns (confidence, omission, escalation, drift).
Quantitative Mechanistic Interpretability (QtMI)
What internal mechanisms produce the failure?
A specialist layer that traces causal model internals (attention heads/features/causal tracing/circuits) to explain how and why the failure emerges.
Relationship: QMI is complementary to, not a substitute for, QtMI. QMI determines when deeper interpretability work is warranted and helps prioritize where to aim it.
Analogy: This mirrors qualitative vs quantitative microbiology in third-party food testing: first determine whether a contaminant is present, then quantify and isolate species as needed.
Accountability: QtMI is typically the responsibility of the model’s developer/parent organization, due to restricted access to internal architectures, training data lineage, and the ability to modify the system. QMI is the responsibility of the deployer/user organization, the party accountable for real-world operation, monitoring, and documented oversight.
Operational Parallel: The supplier of a packing line is accountable to ensure that the mechanisms are working correctly and meet all regulatory/safety metrics prior to sale; businesses using the packing line are accountable for ensuring that mechanisms are working to their specifications on a regular basis using internal quality standards.
Vendors validate the machinery before sale; operators validate performance during use against internal quality standards.
Status: Ongoing development. Sections are published as they clear review.
Intended Use:
– Structured evaluation of AI outputs and behavior under operational conditions
– Root-cause tracing and control design (prevention, detection, documentation, response)
– Alignment of human oversight with measurable signals and documented decision points
Non-goals:
– A complete catalogue of all possible failures within AI-integrated systems
– A claim that failures can be eliminated
– A substitute for domain expertise or organizational accountability
– Charlotte Wilborn 1.21.2026
Table of Contents
Failure Modes – Introduction: Qualitative Mechanistic Interpretability – v0.1
Failure Modes – Pathways Definition – v0.1
Failure Modes – Root Pathways – v0.1
Failure Modes – Introduction: Qualitative Mechanistic Interpretability – v0.1 – Last updated 1.21.2026
Failure Modes – Pathways Definition – v0.1 – Last updated 1.21.2026
Failure Modes – Root Pathways – v0.2 – Last Updated 3.6.2026 – Addition of Extrinsic (Institutional Override) Pathway