To govern AI model prompt and retrieval changes in ABA systems, treat each update as a change to the whole evaluated workflow, including parser, embeddings, tools, permissions, interface, vendor terms, and human review. Preserve the approved baseline, classify possible consequences, compare versions on locked tests, require scoped approval, stage the release, verify rollback, and monitor the new version before retiring old evidence.
Baseline the whole AI system
Hana records model provider and alias, dated model identifier when available, system and task prompts, parameters, safety settings, retrieval sources, parser, chunking, embeddings, ranking, tools, permissions, integration code, interface, reviewer instructions, vendor terms, and monitoring. NIST SP 800-218A supports development and integration change controls within its stated scope. A model version alone cannot reproduce the output. Secrets stay in an approved store; the change record points to them without copying their values.
Classify changes by consequence
A documentation typo may be low risk. A new model, expanded data, changed retrieval source, added tool, broader permission, altered human approval, new client-facing output, or vendor secondary-use term can be high risk. Hana assesses affected use cases, people, data, decisions, actions, compliance obligations, prior incidents, and rollback. HHS risk-analysis guidance makes the actual ePHI environment and changes relevant when HIPAA applies. The highest plausible consequence sets the review route until evidence supports a narrower treatment.
Require vendor change evidence
Track release notes, model retirement dates, alias behavior, regional changes, new subprocessors, retention, training and secondary use, safety configuration, tool behavior, rate limits, and incident terms. If the vendor cannot identify a model build, preserve the observed behavior, date, request identifier, and response. The FTC staff article reinforces that AI companies must honor privacy and confidentiality commitments; contract and privacy owners review changed promises before continued use.
Run a locked comparison
Use unchanged regression cases from the approved population plus new cases targeting the proposed change. Compare source support, severe errors, abstention, reviewer effort, latency, accessibility, security, and downstream actions. Preserve both outputs and reviewer decisions. A higher average cannot offset a new wrong-client, unsafe action, unauthorized disclosure, or missing-audit failure that violates a release threshold.
Stage and observe the release
Move from development to shadow mode, limited users or records, and broader release only after each stage passes. Pin versions when the service permits it. Route a controlled percentage to the new configuration, keep the old path available, and reconcile both. Announce changed behavior and limitations to users. A vendor's automatic upgrade schedule does not eliminate the practice's need for hold, fallback, or alternative workflow.
Make rollback real
Rollback instructions identify the prior configuration, data and index compatibility, credential state, queued actions, records created under the new version, and communications needed. Test rollback before release. If outputs or records cannot be reversed, the plan switches to containment, correction, and downstream reconciliation. Retain the new-version evidence even after rollback so the practice can identify every exposed case.
Define an emergency change route
An urgent security, safety, or vendor-retirement change may need a shorter approval path. Predefine who can authorize it, which minimum tests still apply, what scope is permitted, when the temporary configuration expires, and how retrospective review occurs. Preserve the prior state and affected cohort. Urgency changes timing; it does not erase accountability, evidence, or rollback duties.
Work through a release cohort
Hana locks 20 fictional AI configuration releases. Sixteen have a complete baseline, impact analysis, representative comparison, approval, staged rollout, rollback test, and post-release review: 16 of 20, or 80%. One vendor alias changed without a reproducible build, one prompt change weakened abstention, one corpus update lacked supersession data, and one tool permission expanded without approval. Those four releases remain held or rolled back.
Keep change and incident records connected
A production error links to the exact configuration and change. A change prompted by an incident links back to the affected cohort, root cause, corrective action, and validation. The voluntary AI RMF Manage Playbook connects monitoring, incident response, recovery, decommissioning, and change management. Monitor recurrence and adjacent use cases. Do not close the change merely because the deployment succeeded. Closure requires accepted behavior, complete evidence, reconciled downstream state, updated user guidance, and no unresolved release blockers.
Use a release checklist
- Can we reproduce the old and new configurations?
- Which use cases and data paths change?
- Which severe failures could appear?
- What does the locked comparison show by stratum?
- Who approves the release and human workflow?
- Can we stop or roll back without losing evidence?
- Which post-release signals determine acceptance?
Prepare for vendor-forced change
Some providers retire versions, move aliases, alter rate limits, or introduce a subprocessor on a schedule the practice does not control. Hana records notice channels, retirement dates, contract rights, export options, alternative models, and the last date a safe comparison can run. A forced deadline does not justify silent acceptance. The practice can narrow the use, switch to a manual fallback, hold affected work, or end the service while qualified owners finish the review.
When version pinning is unavailable, capture the provider's response identifiers, timestamps, observed behavior, release notes, and monitoring breakpoints. Increase sampling around the change window and separate results before and after the observed shift. If the vendor cannot supply adequate change evidence for a consequential use, treat that uncertainty as a release risk and document the operational alternative.
Reconcile every artifact after rollback
Rolling back code or a model endpoint does not reverse outputs already used. Identify drafts, signed records, messages, claims, schedules, payments, approvals, caches, embeddings, indexes, queued jobs, and user instructions created under the changed version. Each downstream owner decides whether to correct, cancel, notify, resubmit, monitor, or retain the item. Preserve the original and corrected state with authorship and reason.
After technical rollback, rerun the locked acceptance set on the restored path and verify that credentials, permissions, source versions, logs, and alerts match the approved baseline. Communicate the restored scope and any temporary limitations to users. Close the change only when the technical system, affected work, and monitoring cohort are reconciled, not merely when the old configuration responds again.
Run a preflight that matches the change
Before release, Hana confirms the proposed version, affected use cases, reason, owner, risk class, locked comparison, severe-error results, privacy and security review, accessibility result, human-review impact, fallback, rollback test, user communication, and post-release cohort. Each requirement points to evidence for the same configuration. A copied approval from a prior model or data route remains historical evidence, not current authorization.
The preflight also asks what should stay unchanged. Lock tool permissions, target records, approved users, source authority, output destination, and stop rules unless they are explicitly part of the change. This prevents a model upgrade from quietly bundling a broader operational release. After deployment, reconcile the observed configuration to the approved record before counting the change as complete.
Related resources
- Secure AI Agents and Autonomous Actions in ABA Operations
- Monitor ABA AI Quality, Drift, and Emerging Failures
- Respond to AI Errors, Data Exposure, and Unsafe Actions in ABA Practices
- Design Human Approval, Override, and Stop Controls for ABA AI
Sources
- Council of Autism Service Providers, Organizational Guidelines public overview
- National Institute of Standards and Technology, AI Risk Management Framework
- National Institute of Standards and Technology, AI RMF Core
- National Institute of Standards and Technology, AI RMF Playbook Manage function
- National Institute of Standards and Technology, Generative AI Profile
- National Institute of Standards and Technology, SP 800-218A Secure Software Development Practices for Generative AI
- U.S. Department of Health and Human Services, Guidance on Risk Analysis
- Federal Trade Commission, AI Companies: Uphold Your Privacy and Confidentiality Commitments