Capacity, cost, and unit economics
Measure constrained capacity, demand, waste, unit cost, and accepted outcomes without optimizing spend in isolation.
Connect resources to accepted outcomes
AI cost can be expressed per request, token, model-hour, accelerator-hour, document indexed, evaluation run, supported answer, or accepted task completion. Choose units that reveal decisions; low cost per token can coexist with expensive unusable outcomes.
Cost model
Include compute and accelerators, storage, network and transfer, platform licenses, artifact and model lifecycle, evaluation, observability, security, facilities, support, operations labor, reserved idle capacity, recovery, and depreciation where relevant.
Capacity model
Measure demand distribution, concurrency, context and output lengths, model loading, batching, queueing, memory, hardware availability, failure headroom, maintenance, and growth. Separate theoretical peak, tested sustained capacity, and approved operational capacity.
Optimization guardrails
Evaluate smaller models, quantization, caching, batching, routing, retrieval, prompt reduction, scheduling, and hardware changes against quality, permission, safety, latency, resilience, and sovereignty requirements. Savings are not accepted if they shift unacceptable risk or hidden labor elsewhere.
Report assumptions, measurement conditions, allocation rules, confidence range, and sensitivity. Recalculate after material workload or architecture change.