To build an ABA AI evaluation dataset that represents real work, define the target population, people, documents, settings, and failure consequences first. Sample by meaningful risk strata, preserve lawful provenance and labeling rules, include rare and adversarial cases, separate development from final evaluation, document missingness and uncertainty, and report performance by subgroup and error severity instead of relying on one pooled score.
Start with the target population
NIST TEVV emphasizes context and meaningful datasets. Eleni writes the population in operational terms: which workflow, service dates, payer products, document types, languages, settings, users, and outcomes the evaluation is meant to represent. She also names excluded uses. A dataset of clean typed forms cannot support a claim about scanned records, mixed languages, or new payer products. A convenience sample measures itself unless a defensible link to the target population is shown.
Stratify by consequence and difficulty
Sample ordinary work and the cases most likely to fail or cause harm: missing pages, conflicting sources, uncommon layouts, low-quality scans, new rules, multiple clients, rare codes, accessibility needs, ambiguous language, long inputs, adversarial content, and no-answer cases. Set planned counts before looking at model results. Oversampling a critical stratum is useful when results are reported with its sampling design and not disguised as population prevalence.
Control provenance and privacy
Record source, authority, collection date, permitted use, data owner, de-identification or fictional-construction method, transformation, access, retention, and deletion. Synthetic means purpose-built fictional data only when no real record information survives. De-identification follows the practice's documented applicable method and retains residual risk. HHS risk-analysis guidance covers all ePHI in the evaluation environment when HIPAA applies; vendor training and secondary use follow current terms and authority. The FTC staff article warns model providers to honor data-use commitments.
Write labeling rules before annotation
Define the unit, expected answer, acceptable variants, source evidence, severity, invalidity, uncertainty, and escalation rule. Train reviewers on shared examples and calibrate them. Preserve disagreements and qualified adjudication instead of editing history until agreement appears perfect. Clinical, coding, payer, privacy, and accessibility labels require reviewers with the corresponding competence and authority.
Protect the final evaluation split
Keep development examples, prompt-tuning cases, vendor demonstrations, and debugging outputs out of the locked final set. Track hashes or durable IDs to detect duplicates and leakage. If a vendor has seen the final cases or their answers, label the evaluation contaminated and create a new holdout. Retain a stable regression subset and a rotating freshness subset for new formats, rules, and attack patterns.
Keep dataset readiness separate from model performance
First decide whether every case has valid provenance, a usable input, a qualified expected answer, and a defined scoring rule. Then evaluate the model on the locked eligible set. Publish both the planned-to-ready dataset flow and the model results. This prevents missing or disputed cases from disappearing inside a performance percentage.
Match the statistic to the claim
NIST AI 800-3 distinguishes accuracy on a fixed benchmark from generalized accuracy across a larger population of similar items and emphasizes explicit assumptions and uncertainty. A practice can choose a simple or complex statistical model, then state whether results describe the locked cases or support a broader inference. Report numerator, denominator, exclusions, uncertainty, severity, and stratum.
Work through a dataset build
Eleni plans 160 fictional and authorized retrospective cases across eight predefined strata, 20 per stratum. Fourteen cases fail provenance or labeling gates and remain excluded with reasons, leaving 146 of 160, or 91.3%, ready for the locked evaluation. Readiness is reported by stratum. The 146-case score cannot be generalized to formats, payers, languages, or workflows absent from the target population.
Audit dataset coverage
For each stratum, report planned, acquired, eligible, labeled, adjudicated, locked, and evaluated counts. Track duplicates, contamination, missing sources, invalid transformations, unresolved reviewer disagreement, access defects, and expired authority. Compare the production mix with the evaluation mix over time. A shift in documents, clients, workflow, or model use can require a substantively new sample and version.
Use these design questions
- Which real decision will the evaluation inform?
- Which severe failure deserves its own stratum?
- Who may use each source record for this purpose?
- Which reviewer can label the expected result?
- What has the model or vendor already seen?
- Which claim is limited to the fixed benchmark?
- What change makes the set stale?
Use rare cases without distorting the claim
Critical failures often need deliberate oversampling because routine production may contain too few examples for a useful release decision. Eleni can build a challenge stratum for wrong-client pages, rare payer layouts, low-quality scans, indirect prompt injection, conflicting sources, or accessibility barriers. She labels that stratum as a designed stress test and reports it separately from any estimate of routine prevalence.
If the practice wants a population-level claim, it documents the production sampling frame, time period, inclusion probability, missingness, and any weighting or statistical model. A pooled result that mixes an oversampled attack set with routine cases can answer whether the system passed the locked test, but it cannot describe the rate users should expect in ordinary work. Both views can be useful when their denominators remain explicit.
Govern disagreement and dataset change
Reviewer disagreement is evidence about the task. Preserve each initial label, the source consulted, the disagreement category, the qualified adjudicator, and the final rationale. Repeated disagreement may show that the operational definition is unclear, the source itself conflicts, or the workflow requires judgment that the proposed automation cannot reliably support. Those cases may become abstention tests rather than forced right answers.
Every material dataset edit creates a new version. When a case is removed for expired authority, privacy concern, contamination, or invalid expected answer, retain its identifier and reason outside the active set. When production introduces a new format or payer rule, add a freshness cohort without rewriting the historical score. This preserves comparability while giving the next release a representative test of current work.
Hand off the dataset with a data sheet
The data sheet should name purpose, target population, sampling frame, strata, sources, collection dates, permissions, transformations, label owners, adjudication rules, known omissions, access roles, retention, deletion, leakage controls, and intended metrics. It lists uses the dataset cannot support. Reviewers should be able to trace a reported score to the exact locked version and determine whether every case was eligible on the evaluation date.
Eleni also provides a change log and a production-comparison view. The comparison shows whether live work now contains languages, formats, payers, users, or failure conditions missing from the set. When coverage drifts, the owner decides whether to add a new test cohort, narrow the production claim, or pause expansion. A dataset can remain technically intact while becoming operationally irrelevant.
Related resources
- Design Human Approval, Override, and Stop Controls for ABA AI
- Validate Retrieval-Augmented Generation and Citations for ABA Work
- Monitor ABA AI Quality, Drift, and Emerging Failures
- Defend ABA AI Workflows Against Prompt Injection and Unsafe Tool Use
Sources
- Council of Autism Service Providers, Organizational Guidelines public overview
- National Institute of Standards and Technology, AI Risk Management Framework
- National Institute of Standards and Technology, Generative AI Profile
- National Institute of Standards and Technology, AI Test, Evaluation, Validation and Verification
- National Institute of Standards and Technology, AI 800-3 Expanding the AI Evaluation Toolbox
- U.S. Department of Health and Human Services, Guidance on Risk Analysis
- Federal Trade Commission, AI Companies: Uphold Your Privacy and Confidentiality Commitments