To set AI confidence, abstention, and escalation thresholds for ABA work, define the decision and harm first. Use a locked representative set to connect system scores with observed correctness, severity, missing evidence, and reviewer capacity. Create separate rules for accept-for-review, abstain, hold, and escalate; keep consequential authority with qualified people; and revalidate thresholds whenever data, models, prompts, sources, users, or downstream actions change.

Define what the score actually means

Luis records whether the system emits a probability, similarity score, ranking value, self-reported confidence phrase, or vendor-specific signal. These quantities are not interchangeable. A language model saying "high confidence" is not evidence of calibrated correctness. For each score, document how it is produced, its range, version, target event, known limitations, and whether it was validated on the practice's workflow. When no meaningful score exists, use observable evidence gates such as required-source presence, field consistency, or rule completeness. Keep vendor dashboard categories separate from the practice's release states so a product update cannot quietly change an operational threshold.

Start with the consequence

A threshold for sorting internal low-risk drafts can differ from one that influences clinical care, disclosure, authorization, claims, money, employment, or client communication. Define the possible error, affected person, reversibility, detection opportunity, time pressure, and qualified decision-maker. High-severity errors may require a hard rule or universal review even when the average score is strong. No score grants clinical, payer, coding, privacy, legal, or financial authority.

Create more than two routes

Luis uses four states: route to ordinary review, route to enhanced review, abstain because evidence is insufficient, and stop because a critical rule failed. The interface shows why the item entered its state and which evidence is missing or conflicting. Staff can challenge the route and send it to the appropriate owner. A default answer after timeout, missing source, or unsupported format defeats the purpose of abstention.

Calibrate on locked cases

NIST TEVV emphasizes measurement in the context where AI is used. Lock a representative set before tuning. For each score band, report total cases, correct cases, severe errors, abstentions, reviewer edits, and unresolved cases. Check whether observed correctness rises with the score and whether the relationship holds across relevant formats, payers, languages, access needs, and user groups. Keep development and final evaluation cases separate.

Balance errors and workload

Lowering a threshold may route more work to ordinary review while increasing hidden errors. Raising it may increase abstentions and delay work. Model the number of items each route creates during normal and peak periods, the time needed for qualified review, deadlines, and the approved fallback. The workflow must hold safely when reviewers are unavailable. An escalation queue that cannot be staffed is a delayed failure path. Record the queue assumptions beside the threshold evidence.

Use independent evidence gates

Confidence cannot override wrong-client detection, expired authorization, missing source, inaccessible input, unsupported language, prohibited use, privacy restriction, or absent required approval. Combine the score with deterministic checks and domain rules. The NIST Generative AI Profile identifies confabulation risks, including fabricated citations. A polished citation or high score still fails when it does not support the material claim.

Work through a routing test

Luis locks 50 fictional payer-document cases and predefines the expected route. Forty-one enter the correct ordinary-review, enhanced-review, abstain, or stop state: 41 of 50, or 82%. Three severe-error cases receive ordinary review, two valid cases abstain, one missing-source case produces an answer, and three items escalate after the deadline. Report each failure type; the pooled percentage cannot show whether the threshold is acceptable.

Monitor and revalidate

Track the score distribution, route distribution, correctness by band, critical errors, abstention reasons, overrides, review time, queue age, missed deadlines, and downstream corrections. Investigate shifts in inputs, source rules, model versions, prompts, retrieval, users, or workflow. The voluntary AI RMF Core treats monitoring, human oversight, override, and change management as connected risk work. Reopen threshold validation after a material change or unexpected severe error.

Ask these threshold questions

  • What event does the score predict?
  • Which critical failures bypass the score?
  • How many locked cases support each band?
  • Which route owns ambiguous or missing evidence?
  • Can qualified reviewers absorb the resulting queue?
  • What drift or incident forces recalibration?

Build a consequence and capacity table

For each route, list the cost of a false positive, false negative, unnecessary abstention, and delayed escalation. A wrong-client match, unsupported clinical statement, or unauthorized payer action may require a hard block regardless of the model score. A low-impact formatting suggestion may tolerate broader automated assistance when a user can reverse it easily. The table keeps the threshold tied to the decision rather than to a convenient percentile.

Add queue capacity, reviewer skill, service hours, deadlines, and manual fallback. A threshold that sends 40 cases per day to a queue staffed for 10 is not safe merely because calibration is mathematically sound. Simulate expected and peak volumes, preserve priority rules, and hold lower-priority automation when qualified review cannot meet the defined response time.

Document the routing rule as a versioned control

Record the model and score version, eligible population, exclusions, source prerequisites, threshold values, route definitions, reviewer roles, expected volume, evaluation cohort, error tradeoffs, stop rules, and approval date. Include examples just above and below each boundary so operators can understand the intended behavior. A later reviewer should be able to reproduce why the same score entered a different route after a version or policy change.

Give users a correction and appeal path that does not depend on persuading the model. Track overrides, hidden manual workarounds, delayed escalations, and complaints as evidence about the rule. When the input mix, model, workflow, queue capacity, or consequence changes, rerun calibration and publish the new route version before applying it to live work.

Test edge cases around every boundary

Use locked cases on both sides of each threshold, including ties, missing scores, malformed inputs, conflicting evidence, unavailable reviewers, and a provider timeout. Verify that deterministic prerequisites and critical blockers still control the route. A confidence value should never convert missing authority or absent source evidence into permission to act.

Maintain a threshold change log that shows the prior and new rule, reason, evaluation cohort, error movement, queue impact, approver, effective date, and rollback point. Review the first production cohort under the new rule before retiring the old comparison. If performance improves only because more difficult cases abstain or disappear from the denominator, state that tradeoff directly and decide whether the resulting delay and access burden remain acceptable.

Include the number and age of cases waiting in every route so threshold performance remains connected to the people and work affected by delay.

Related resources

Sources