To control ABA AI cost, token, quota, and capacity risk, map each use case to a measurable unit of work and every chargeable component. Forecast ordinary and peak demand, set per-action and period limits, alert before exhaustion, require approval for expansion, and define graceful degradation that preserves safety, privacy, deadlines, and human authority. Reconcile estimates with provider usage and financial records, then retest after pricing, model, prompt, or workload changes.
Choose a real unit of work
Sora models cost per document, conversation event, draft, reviewed claim, client message, or completed workflow, depending on the use case. She separates input tokens, cached input, output, embeddings, retrieval, tool calls, storage, transcription, image processing, vendor minimums, support, and human review. Cost per API call can hide retries and correction work. A completed unit requires the source, AI step, human gate, and downstream reconciliation that make the result useful.
Forecast demand in cohorts
Use historical or planned volumes by day, hour, location, workflow, document size, language, and deadline. Model ordinary, peak, outage-recovery, retry, and growth scenarios. Keep assumptions and version dates visible. Average demand cannot show whether a Monday authorization queue or month-end billing run will exhaust a quota. Include capacity for monitoring, regression tests, incidents, and required reprocessing without allowing test traffic to compete silently with production.
Set layered limits
Apply per-request input and output limits, tool-call limits, retry limits, per-user and per-tenant budgets, daily and monthly spend thresholds, concurrency controls, and vendor quota reservations. Block recursive or unbounded agent loops. A hard limit names the safe response and owner; a soft alert names the investigation window. Do not truncate source material or clinical context invisibly to fit a token budget. The system should abstain or route to an approved fallback when required evidence cannot fit.
Prioritize by risk and deadline
Reserve capacity for approved high-priority work while preserving emergency, clinical, payer, privacy, and record duties outside the AI system. Lower-priority drafting or experimentation can pause first. Define priorities before exhaustion and avoid using disability, language, payer type, or staff preference as an adverse shortcut. A qualified clinician decides clinical implications; operations decides queue routing within approved policy.
Design graceful degradation
Test smaller approved models, shorter but complete source packets, queued processing, manual work, reduced optional features, and full holds. Every fallback retains data, access, source, authority, quality, and audit requirements. HHS's current Security Rule summary includes contingency planning for ePHI systems. The January 2025 Security Rule update remains proposed, so use the current rule as the regulatory baseline while tracking future changes.
Test recovery capacity
Adapt the recovery and exercise concepts in NIST SP 800-34 Rev. 1 to the practice's actual environment. It is federal information-system guidance, not a general private ABA mandate. Measure whether the approved fallback can absorb queued work, restore source access, preserve deadlines, and reconcile every unit when normal capacity returns. Include a failed fallback component and an unavailable leader so the exercise tests decision ownership as well as technical throughput.
Reconcile provider and internal evidence
Capture request ID, use case, user, model, version, token and tool counts, retries, latency, result, reviewer, and final work unit without exposing unnecessary PHI. Compare internal records with provider usage, invoices, credits, contractual minimums, and financial accounting. Investigate missing, duplicated, or unexpected traffic. A dashboard estimate is not a ledger, and a paid invoice does not prove every request was authorized or useful.
Work through a capacity period
Sora locks 60 fictional work units due during a peak-day test. Forty-five complete within the approved model, quota, deadline, privacy, source, human-review, and downstream-reconciliation gates: 45 of 60, or 75%. Five exhaust a retry budget, three are truncated, two use an unapproved fallback, one misses its deadline, and four remain queued without a named owner. The practice fixes the capacity plan before increasing volume.
Monitor and approve changes
Track demand, tokens, tools, cost, latency, retries, quota headroom, holds, completion, correction time, and business outcome by use case and version. Compare forecast with actual and preserve the difference. Reapprove after model pricing, context window, cache behavior, prompt, document size, tool, vendor quota, or workload changes. The voluntary AI RMF Manage Playbook can organize resource, monitoring, incident, recovery, and change decisions without becoming a private-practice mandate.
Use these capacity questions
- What completed unit does the spend support?
- Which components, retries, and people create total cost?
- What peak and recovery demand must the route absorb?
- Which limits stop runaway work?
- Which approved fallback preserves required evidence?
- Can finance and operations reconcile each material variance?
Decide what continues during scarcity
Create priority classes from real consequences and deadlines, not from which team has the loudest request. Essential continuity work may receive reserved capacity, while experiments, bulk reprocessing, and low-impact drafting pause. Each class has an approved manual fallback, maximum delay, communication owner, and rule for returning queued work. Clinical and safety decisions remain with qualified people even when the AI route is unavailable.
Apply limits by tenant, use case, user, action, provider, and time window so one workload cannot consume the entire budget or quota. Guard against retry storms, oversized context, runaway agents, duplicated jobs, and routing loops. A cost limit should stop new discretionary work while preserving the evidence and state needed to finish or safely hold in-progress actions.
Reconcile provider billing with internal work
Compare provider usage, invoices, request logs, application events, user records, completed outputs, retries, cache hits, and downstream transactions. Investigate unexplained tokens, calls without a corresponding approved task, tasks with several provider charges, and work marked complete without an output. Separate provider measurement changes from real demand changes before revising a forecast.
Report cost per eligible unit, completed unit, reviewed unit, and accepted unit. Include human correction, fallback, delay, incident, and rework when deciding whether the workflow creates value. A lower token price can be offset by poorer output, more review, or unsafe retry behavior.
Review limits after every material change
A new model, context window, prompt, retrieval design, provider, price, quota, routing rule, user group, or action can change capacity. Re-run peak tests and verify alerts, holds, priority order, manual fallback, and recovery. Keep temporary emergency limits dated and visible so they do not become the silent permanent operating model.
Bring a concise capacity packet to the operating review: forecast versus actual units, provider charges, internal work, retries, queue age, held priorities, manual fallback volume, severe failures, unresolved reconciliation, and proposed limit changes. The decision record states which work continues, which pauses, who communicates the change, and what evidence restores normal service. Cost control remains linked to continuity and risk rather than acting as an isolated finance threshold.
Related resources
- Plan ABA AI Continuity, Fallback, and Vendor Exit
- Route ABA AI Work Across Models and Providers Safely
- Protect PHI in ABA AI Prompts, Logs, and Feedback
- Label AI-Generated Content and Preserve Output Provenance in ABA
Sources
- Council of Autism Service Providers, Organizational Guidelines public overview
- National Institute of Standards and Technology, AI Risk Management Framework
- National Institute of Standards and Technology, AI RMF Playbook Manage function
- U.S. Department of Health and Human Services, HIPAA Security Rule Summary
- National Institute of Standards and Technology, SP 800-34 Rev. 1 Contingency Planning Guide