Troubleshooting and operations
AI Incident Triage
Stabilize an AI-enabled service, classify impact, preserve privacy-safe evidence, reconcile uncertain actions, and route decisions to accountable humans.
- Format
- Decision Path
- Level
- Advanced
- Audience
- Developer, Operator, Leader
- Owner
- Project42 Editorial
- Review cadence
- Every 60 days
- Prerequisites
- An incident response policy with severity and notification rules; On-call ownership for application, data, security, safety, and business decisions; Approved evidence retention and communication channels
Protect people and stabilize the service
If credible harm is ongoing, use the fastest approved control: pause a route, disable a tool, revoke a compromised test identity, fail closed, reduce exposure, or restore a verified configuration. Preserve life-safety and emergency procedures outside the AI system. Do not wait for perfect root cause before bounded containment.
Set severity from actual and credible potential impact across people, rights, safety, privacy, security, money, availability, data integrity, compliance, and recoverability. Identify affected tenants, users, versions, models, prompts, tools, data, regions, and time windows. Track confirmed facts, hypotheses, unknowns, decisions, owners, and timestamps separately.
Maintain a decision-ready incident record
Preserve correlation identifiers, configuration versions, deployment events, redacted traces, alerts, user reports, evaluation results, external-state checks, and communication decisions. Restrict raw prompts, outputs, and tool data by purpose and access; a broad incident channel is not an approved data store. Give legal, privacy, safety, security, and provider escalation owners the evidence their policies require.
Incident ID and commander: [ID -> ACCOUNTABLE ROLE]
Start/detection/current time: [UTC]
Observed and potential impact: [PEOPLE, DATA, RIGHTS, SAFETY, SECURITY, MONEY, SERVICE]
Severity and rationale: [LEVEL -> EVIDENCE]
Affected scope: [USERS/TENANTS/VERSIONS/MODELS/TOOLS/REGIONS]
Containment and result: [ACTION -> VERIFIED STATE]
Facts / hypotheses / unknowns: [SEPARATE LISTS]
Uncertain side effects: [OPERATION IDS -> RECONCILIATION OWNER]
Evidence and access: [REDACTED REFERENCES -> RETENTION]
Owner and cadence: [COMMAND, TECHNICAL, COMMS -> NEXT UPDATE]
Stop criteria: [EXPANDING HARM | DATA EXPOSURE | UNKNOWN WRITES | LOST OBSERVABILITY]
Recovery decision: [CONTINUE CONTAINMENT | ROLLBACK | FAILOVER | MANUAL PROCESS]
Notification and escalation: [POLICY/PROVIDER/AUTHORITY -> OWNER]
Verification: [USER, SERVICE, DATA, SAFETY, AND POSTCONDITION CHECKS]Expected evidence and verification
Expected evidence includes impact and severity rationale, bounded scope, verified containment, a timestamped decision log, privacy-safe traces, external-state reconciliation, owners, communications, and recovery criteria. Verify each containment control in the actual affected boundary, confirm error and harm signals stop, check data and tool postconditions, and keep monitoring through the declared observation window.
Stop expansion or automated operation when harm grows, protected data is exposed, write outcomes are unknown, observability is lost, ownership is unclear, or the current configuration cannot be proven. Recovery means maintaining containment, restoring or failing over to a verified path under human authority, reconciling every consequential operation, communicating by policy, and deferring closeout until durable fixes and recurrence checks exist.