SovAIHub
ModulesSAI-220
SAI-220 table of contents
Concept2 min readContent reviewed

Capacity, reliability, and recovery

Define capacity envelopes, degradation behavior, service objectives, recovery targets, and tested restoration evidence.

Last content review 2026-08-03Included in SAI-220, SAI-260, SAI-270

Define the service envelope

AI capacity is shaped by model size, precision, context length, request mix, concurrency, batching, hardware, memory, retrieval, and tool dependencies. A single throughput number cannot describe every workload.

Create a measured envelope that states the model and runtime versions, hardware profile, request distribution, concurrency, latency percentiles, error behavior, power or cost basis, and test limitations. Re-measure after material change.

Service objectives

Choose indicators that reflect user outcomes and control behavior. Examples include successful supported answers, permitted retrieval success, time to first useful output, completion latency, tool-action success, policy-decision availability, and evidence completeness.

For each objective, define the population, measurement window, exclusions, target, owner, and response when the error budget is consumed. Do not label unsafe or unauthorized output as successful simply because the endpoint returned a response.

Design degradation

Plan behavior for accelerator exhaustion, model unavailability, retrieval failure, identity or policy outage, evidence-store failure, excessive context, and dependent-tool failure. Options include queueing, reduced concurrency, an approved smaller model, retrieval-only results, cached approved content, no-answer, or service suspension.

Every fallback changes risk. It requires its own approval and evaluation rather than automatic routing to any available system.

Recovery

Define recovery time and recovery point targets for configuration, model and artifact stores, knowledge indexes, source data, evidence, secrets, and operational state. Distinguish rebuildable derived data from authoritative records.

Test restore and rebuild procedures. A backup job marked successful is not proof that a compatible, authorized system can be restored.

Recovery evidence

Retain the scenario, starting state, versions, backup identity, restore steps, duration, validation results, data loss, exceptions, and approval. Feed failures into the controlled-change backlog and repeat the exercise after material architecture changes.