Capacity, reliability, and recovery
Define capacity envelopes, degradation behavior, service objectives, recovery targets, and tested restoration evidence.
Define the service envelope
AI capacity is shaped by model size, precision, context length, request mix, concurrency, batching, hardware, memory, retrieval, and tool dependencies. A single throughput number cannot describe every workload.
Create a measured envelope that states the model and runtime versions, hardware profile, request distribution, concurrency, latency percentiles, error behavior, power or cost basis, and test limitations. Re-measure after material change.
Service objectives
Choose indicators that reflect user outcomes and control behavior. Examples include successful supported answers, permitted retrieval success, time to first useful output, completion latency, tool-action success, policy-decision availability, and evidence completeness.
For each objective, define the population, measurement window, exclusions, target, owner, and response when the error budget is consumed. Do not label unsafe or unauthorized output as successful simply because the endpoint returned a response.
Design degradation
Plan behavior for accelerator exhaustion, model unavailability, retrieval failure, identity or policy outage, evidence-store failure, excessive context, and dependent-tool failure. Options include queueing, reduced concurrency, an approved smaller model, retrieval-only results, cached approved content, no-answer, or service suspension.
Every fallback changes risk. It requires its own approval and evaluation rather than automatic routing to any available system.
Recovery
Define recovery time and recovery point targets for configuration, model and artifact stores, knowledge indexes, source data, evidence, secrets, and operational state. Distinguish rebuildable derived data from authoritative records.
Test restore and rebuild procedures. A backup job marked successful is not proof that a compatible, authorized system can be restored.
Recovery evidence
Retain the scenario, starting state, versions, backup identity, restore steps, duration, validation results, data loss, exceptions, and approval. Feed failures into the controlled-change backlog and repeat the exercise after material architecture changes.