To reduce automation bias and overreliance in ABA AI workflows, design review so a person can form and record an independent judgment from source evidence before accepting the AI output. Show uncertainty and material changes, make reject and abstain routes easy, preserve qualified decision authority, test reviewers with realistic seeded errors, control workload and incentives, and monitor source opening, edits, overrides, misses, and downstream corrections together.

Recognize several overreliance patterns

Nolan looks for commission errors, where staff follow a wrong recommendation, and omission errors, where they miss a problem because the system stays silent. He also looks for anchoring on the first draft, selective source reading, copying unsupported rationale, skipping required judgment, and deferring to a score that the user does not understand. Fast completion and high agreement may indicate good assistance, weak review, or an easy cohort. Those signals need source-based examination.

Keep decision ownership visible

A qualified clinician owns clinical interpretation, risk, goals, dosage, and record authorship within scope. Payer, coding, billing, privacy, security, workforce, finance, and legal decisions remain with their authorized roles. The interface names the decision owner and what AI did. A supervisor, practice owner, system administrator, or vendor cannot convert access to the tool into clinical or payer authority. Required consent, assent, communication access, and appeal routes continue outside the model.

Make evidence easier to inspect than the answer

Put the source passage, record, date, version, target person, missing evidence, and conflict next to the proposed output. Highlight material differences and unsupported claims. Let reviewers open the complete source and record their own conclusion before seeing a recommendation when the use case warrants it. The NIST Generative AI Profile describes risks from confident confabulation and fabricated citations. A polished answer should not receive better interface treatment than the evidence needed to challenge it.

Design real choices

Accept, edit, reject, abstain, request evidence, escalate, and stop should be distinct and equally available. Avoid preselected acceptance, countdown pressure, hidden edit boxes, approval language that implies management preference, or repeated warnings that users learn to ignore. Require a reason only when the burden is proportionate and useful for risk review. A rejected answer should not reappear as the default on the next screen or silently populate another field.

Test reviewers with seeded errors

NIST TEVV supports contextual evaluation. Insert authorized fictional errors that resemble production failures: correct-looking wrong-client text, stale payer language, a fabricated citation, a subtle unit error, an accessible-form failure, or an unsupported clinical statement. Measure whether reviewers open sources, detect the error, select the correct route, and avoid creating a new error during correction. Tell staff the program uses quality tests without revealing the exact cases in advance.

Control workload and incentives

Review time, queue age, deadlines, staffing, interruptions, screen layout, and performance goals shape behavior. Do not reward throughput while penalizing abstention or escalation. Include peak periods, repeated decisions, and fresh reviewers in the test design. If qualified review capacity falls below the approved workload, hold or narrow the workflow. Training explains known failures and authority, then observed practice verifies that the design works.

Work through a reviewer study

Nolan locks 24 fictional AI-assisted decisions containing ordinary cases and six seeded material errors. Eighteen show the required source opening, independent conclusion, correct disposition, and preserved rationale: 18 of 24, or 75%. Reviewers accept three seeded errors, overlook one missing source, and twice edit the output without changing an incorrect downstream field. The practice repairs the interface and workflow before expanding use.

Monitor several signals together

Report source-open rate, material-edit rate, reject, abstain, escalation, override, seeded-error detection, severe-error miss, review time, queue age, and downstream correction. Segment by workflow, version, user role, tenure, volume, and relevant access needs. The voluntary AI RMF Core connects human oversight, feedback, appeal, override, monitoring, and change management. No metric creates a human-review safe harbor; raw cases and affected outcomes remain part of the review.

Ask these human-factors questions

  • Can the reviewer reach an independent conclusion from the displayed evidence?
  • Is rejecting or abstaining as easy as accepting?
  • Which incentives reward speed over judgment?
  • Do seeded tests resemble severe production failures?
  • Which workload state forces a hold?
  • How are reviewer feedback and affected-person concerns resolved?

Redesign the workflow after an overreliance finding

Suppose reviewers approve a plausible but unsupported authorization summary because the AI answer appears before the payer source. Move the authoritative fields and cited passage ahead of the suggestion, require the reviewer to resolve a highlighted conflict, and keep the AI draft hidden until the source review begins. Narrow the task if the interface cannot present enough evidence for a responsible decision.

Other repairs may include removing confidence theater, separating draft from final state, adding a meaningful reject route, rotating work, lowering volume, requiring a second check for severe fields, or making the manual baseline easier. Test the revised design with fresh seeded errors under realistic time pressure. Measure whether reviewers detect the defect and whether the added control creates inaccessible or burdensome workarounds.

Align incentives with careful review

Review quality deteriorates when staff are rewarded only for queue speed or criticized for abstaining. Track material edits, source openings, justified rejections, escalations, corrections, and detected severe errors alongside time. Leaders should make clear that holding unsupported work is an expected control outcome. Protect reviewers who report that volume, interface design, training, or authority makes the task unsafe.

Supervisor sampling should include accepted, rejected, edited, and escalated cases. Discuss the evidence and decision process rather than comparing a person with the model as if one score establishes competence. Repeated failure can signal an individual training need, but it can also reveal a poorly designed task or impossible workload.

Watch overreliance across workflow changes

Monitor behavior after a model improvement, interface redesign, staffing change, deadline surge, or new use case. Higher apparent accuracy can increase trust faster than the evidence supports. Compare severe-error detection, edit rate, time on evidence, override use, and user concerns across versions while keeping the cohort visible. Reopen the human-factors evaluation when review behavior changes materially.

Give reviewers a brief pre-use statement that names the task, source of truth, known failure patterns, authority boundary, and available choices. Keep the statement close to the action rather than buried in annual training. After material errors, update the interface and examples, then ask reviewers to demonstrate the correction path with fictional cases. Training evidence should show usable judgment, not only that a slide deck was opened.

Ask reviewers periodically whether the system makes disagreement easy, whether they feel rushed to accept, and which evidence they still cannot inspect from the decision screen.

Related resources

Sources