To validate multimodal AI across images, audio, and video in ABA, define each allowed modality, combined task, source boundary, synchronization rule, output, and human decision. Test representative devices, environments, people, communication modes, missing channels, conflicting channels, edits, and adversarial inputs. Preserve links to exact source segments, report severe errors separately from averages, protect every media copy, and require an accessible non-AI route.

Define Bao's multimodal input, alignment, and outcome test matrix

Bao separates media capture, image frame, audio segment, transcript, detected object, speaker label, timestamp, cross-modal alignment, generated interpretation, confidence, and downstream action. A component can perform well alone while the combined system assigns a statement to the wrong person or aligns an event with the wrong moment. The operating question is whether the complete deployed route preserves source, context, access, uncertainty, and authority.

Record the decisions and evidence that release depends on

The multimodal input, alignment, and outcome test matrix records use case, device and version, setting, participant and role, consent or other authority, modality, file and segment ID, time basis, synchronization, preprocessing, accessibility support, model and configuration, prompt, detected or generated element, source coordinate, cross-modal agreement or conflict, uncertainty, reviewer, severe error, downstream state, retention, deletion, incident, test stratum, owner, and disposition. Structured fields support assignment, comparison, alerts, expiry, testing, and reconciliation. Narrative explains the real workflow, affected people, clinical and operational consequence, access needs, uncertainty, source limits, failed tests, and the accountable owner's disposition.

Run the implementation in a controlled sequence

Bao writes a task contract before collecting test media. Purpose-built fictional cases vary cameras, microphones, lighting, noise, distance, movement, languages, AAC, captions, dropped frames, muted channels, edits, and timing offsets. Reviewers score source-linked outputs without seeing the expected answer first. The system abstains on missing or conflicting evidence and never treats absence from one modality as absence of communication. Release is limited to the exact tested task and configuration.

Keep the standard, platform, and decision boundaries visible

NIST TEVV emphasizes meaningful measurement and evaluation, and the August 2026 TEVV-Athlon initial public draft explicitly includes multimodal systems within a customizable framework. The draft is not final and comments remain open through October 6, 2026. NIST's GenAI evaluation program tests several modalities but does not certify a practice, product, clinical conclusion, recording authority, accessibility, or payer decision.

Use five release gates

  • The exact modalities, combined task, population, setting, device, model, output, and downstream decision are fixed.
  • Representative and adversarial cases include missing, delayed, conflicting, altered, and inaccessible channels.
  • Every material output links to the exact source segment and preserves uncertainty and abstention.
  • Separate qualified reviewers decide clinical, accessibility, privacy, security, payer, and operational readiness.
  • Release, monitoring, correction, retention, deletion, incident, and fallback routes work for the complete system.

Handle a realistic complication

A video may show a client reaching toward an AAC device while the audio channel contains another person's speech. Bao treats identity and timing as unresolved, blocks automatic attribution, retains the source segments under approved access, and routes the event for qualified review rather than forcing one combined narrative.

Protect care, communication, records, and access

Bao traces effects from the multimodal input, alignment, and outcome test matrix to safety, clinical work, communication and AAC, privacy, records, authorizations, claims, payroll, payments, family contact, and accommodations. Urgent safety, incident, and reporting work proceeds through its own authority. A qualified clinician decides whether clinical services can proceed after a material technology failure; each other accountable owner decides within that role's scope.

Work through a fictional practice example

Bao locks 42 fictional multimodal cases. Thirty-one pass capture, alignment, source-link, communication-access, severe-error, privacy, review, and downstream gates. Two swap speakers, two drift by several seconds, one misses AAC, one fabricates a hidden object, two fail on low light, and three lack reproducible preprocessing evidence. Five repair; six remain outside scope. This fictional scenario tests the control and denominator. It supports no conclusion about a real practice, person, product, legal duty, clinical outcome, payer decision, or security posture.

Measure the full locked cohort

Bao's initial readiness is 31 of 42, or 73.8%. The report retains all 42 multimodal cases due, including failed, unknown, skipped, expired, prohibited, and unresolved work. It states the lock date, review cutoff, reasons, owners, and age. Systems, people, records, events, attempts, findings, tests, and remediation actions keep separate denominators.

Test the failure modes that matter

Bao tests ordinary combined case, image only, audio only, missing frames, timing drift, wrong speaker, occlusion, low light, noise, caption mismatch, AAC, gesture, language change, altered media, adversarial overlay, unsupported inference, source retrieval, deletion, and fallback. Each case preserves the system and version, starting state, data, identity or process, expected result, observed result, raw evidence, defect, owner, retest, and disposition. A passed case applies only to the named configuration and conditions.

Avoid the failures that create false confidence

Combining modalities can create confident but unsupported interpretations when channels identify different people, cover different times, omit communication, or carry manipulated content. Weak evaluations pool frames and words into one score, test components but not the integrated route, assume video proves intent, omit AAC and non-speech communication, ignore synchronization and editing, retain media indefinitely, and let a human approval checkbox hide missing source links or unqualified interpretation.

Require independent acceptance

Bao gives an independent reviewer the multimodal input, alignment, and outcome test matrix, locked scope, source map, configuration, raw evidence, failures, approvals, monitoring, remediation, and closure proof. The reviewer reproduces an ordinary path, a severe failure path, and the final denominator. A changed cohort, hidden manual repair, missing record, or undocumented dependency fails acceptance.

Place the implementation inside current healthcare duties

Bao uses the CASP public organizational overview only for high-level business, clinical-operations, and risk context. The HHS risk-analysis guidance requires a regulated entity's risk analysis to reach all ePHI it creates, receives, maintains, or transmits. Neither source validates this multimodal input, alignment, and outcome test matrix, a product, a clinical workflow, or a legal conclusion.

Keep current and proposed rules separate

Bao checks the current HHS Security Rule summary before release. As of August 24, 2026, that page still identifies the January 2025 cybersecurity update as proposed. The page therefore maps current duties and voluntary readiness sources separately and does not state proposed requirements as operative law.

Use each technical source within its stated scope

Bao's page-specific sources are National Institute of Standards and Technology, AI Risk Management Framework, National Institute of Standards and Technology, Generative AI Profile, National Institute of Standards and Technology, AI Test, Evaluation, Validation and Verification, National Institute of Standards and Technology, TEVV-Athlon initial public draft, National Institute of Standards and Technology, GenAI evaluation program, U.S. Department of Health and Human Services, Business Associates. Each retains its stated date, version, sector, status, and limits. The practice still verifies actual entity role, data, configuration, contract, accessibility, clinical authority, payer rules, state law, and deployed evidence.

Maintain the control after release

Bao assigns the multimodal input, alignment, and outcome test matrix a review cadence and event triggers for systems, data, identities, versions, configurations, vendors, workflows, incidents, contracts, law, and ownership. Material changes reopen affected gates and tests. This page remains draft until the named technology, privacy, security, clinical, accessibility, records, payer, and legal reviewers complete their work.

Related resources

Sources