Service objectives, incidents, and recovery
Define meaningful objectives and operational responses for degraded, unsafe, or unavailable AI behavior.
Define meaningful objectives
An indicator measures behavior; an objective states an acceptable target over a population and window. Examples include permitted successful completions, supported-answer rate, policy-decision availability, time to first useful output, and recovery time.
State exclusions carefully. Removing every difficult or failed request can make an objective meaningless. Define error budgets and the change or incident response when they are consumed.
Incident model
Prepare for security, privacy, safety, quality, permission, availability, capacity, evidence, cost, and supplier incidents. Establish severity, detection, command roles, containment, communication, evidence preservation, decision logs, recovery criteria, and post-incident review.
Recovery and degradation
Pre-approve safe behavior for unavailable models, retrieval, identity, policy, tools, storage, and telemetry. A fallback model or endpoint changes behavior and must have its own evaluation and authorization.
Test rollback, restore, rebuild, queue drainage, stale-state cleanup, credential revocation, and validation of recovered behavior. Recovery is complete only when the system is authorized, healthy, and producing acceptable evidenced outcomes.
Use incidents to add evaluation cases and update controls, architecture decisions, runbooks, and reassessment triggers.