An ABA Wilson score interval calculator shows how much model-based uncertainty surrounds a binary response count, such as 8 qualifying responses in 10 valid opportunities. This worksheet calculates the Wilson limits, preserves the response and denominator rules, and compares separately observed or clearly hypothetical opportunity counts. The interval is useful only when its binomial assumptions are plausible. It is not a mastery band, a forecast for the next session, or evidence that treatment caused change.

Clinicians & ABA Professionals / Data, Outcomes and Clinical Decision-Making.

The percentage hides the amount of evidence

Two records can both display 80 percent while containing very different amounts of information. Eight responses in 10 opportunities and 40 in 50 have the same point proportion, yet the smaller record has substantially wider Wilson limits. The difference is not a correction to the observed percentage. It is a reminder that the denominator matters when a team interprets sampling uncertainty under a stated model.

The calculator answers one narrow question: if the recorded outcomes can reasonably be treated as a fixed number of binary trials with a common response probability and suitable independence, what Wilson score interval follows from x, n, and the selected confidence level? It cannot determine whether the opportunities were valid, comparable, representative, or clinically important.

That distinction matters in ABA. Opportunities may occur in clusters, depend on the immediately preceding trial, change with prompts or partners, or be available only in particular routines. A precise computation does not make those conditions disappear.

The Wilson calculation in visible steps

Let x be the number of qualifying responses, n the number of valid opportunities, p-hat = x / n, and z the standard-normal critical value for the prespecified two-sided confidence level. For 95 percent confidence, this calculator uses z = 1.9599639845.

Compute:

D = 1 + z^2 / n

center = (p-hat + z^2 / (2n)) / D

half-width = z x sqrt(p-hat(1 - p-hat) / n + z^2 / (4n^2)) / D

lower = center - half-width

upper = center + half-width

The NIST confidence-interval handbook gives this score-based form and distinguishes it from an adjusted Wald interval. The boundaries stay within 0 and 1 without manually clipping a normal approximation.

A methods review of confidence intervals for a binomial proportion compares Wilson and other approaches and emphasizes that finite-sample coverage depends on the method and circumstances. Another review of binomial interval methods documents poor Wald performance, particularly with small samples and proportions near zero or one. These papers support transparent method choice. They do not validate a local ABA observation system.

Stop before arithmetic when the record is unclear

Leave the interval blank when n = 0, x is not an integer from 0 through n, or the confidence level was selected after inspecting several outputs. Stop as well when:

  • the response definition changed during the window;
  • a prompted response is sometimes counted as success and sometimes not;
  • invalid, unavailable, or refused opportunities were silently coded as failures;
  • trial retries were added to the denominator without a prewritten rule;
  • the same event appears in both the numerator and denominator more than once;
  • opportunity availability depends on responding in a way the model does not address;
  • records from unlike goals, support levels, partners, or settings were pooled merely to enlarge n; or
  • the intended decision calls for a single-case analysis, validated assessment, research protocol, or risk procedure beyond this worksheet.

The BCBA Test Content Outline addresses operational definitions, measurement, validity, reliability, representative data, graphing, and data-based decisions. The BACB ethics-code page points certificants to the current code. Neither source designates a Wilson interval as a required mastery or treatment decision rule.

Copyable record and calculator

Complete this ABA Wilson score interval calculator block before entering counts:

FieldPrespecified recordRecord ID, program version, and calculator versionClient-selected or client-informed purposeObservable response definitionValid-opportunity definitionPrompt, support, partner, setting, and materialsObservation window and sampling planExclusion, unavailable, refusal, and retry rulesConfidence level and critical value sourceWhether comparison view is observed or hypotheticalObserver, reviewer, and calculation date

Then retain every intermediate value:

CalculationPrimary viewSensitivity viewQualifying responses xValid opportunities nObserved proportion p-hatConfidence levelCritical value zz^2Denominator DWilson centerWilson half-widthLower limitUpper limitInterval widthDisplay rule and displayed intervalAssumption or comparability concerns

Never relabel a hypothetical sensitivity count as observed data. If the views use different conditions, describe them as different estimands rather than as replications of one stable probability.

Fictional example: the same percentage, different opportunity counts

This example contains invented data. In a defined communication routine, View A records x = 8 qualifying responses in n = 10 valid opportunities. The team selected a two-sided 95 percent Wilson interval before reviewing the count.

p-hat = 8 / 10 = 0.8

z^2 = 3.8414588207

D = 1 + 3.8414588207 / 10 = 1.3841458821

center = (0.8 + 3.8414588207 / 20) / 1.3841458821 = 0.7167401600

half-width = 0.2265776885

The full-precision limits are 0.4901624715 and 0.9433178485. With a prespecified one-decimal percentage display, the record reads 80.0% observed; 95% Wilson interval 49.0% to 94.3%.

View B is a separate fictional observation block under the same written definitions, not a multiplication of View A. It records 40 qualifying responses in 50 valid opportunities. The point remains 0.8, while the Wilson limits are 0.6696289407 and 0.8875624998, or 67.0% to 88.8% after the same display rule. Its interval width is about 21.8 percentage points, compared with about 45.3 for View A.

The narrower interval does not prove that View B is more valid. Fifty repetitive, prompted, or tightly clustered opportunities may provide less independent information than the model assumes. The comparison only shows what the Wilson calculation returns when x, n, and the model assumptions differ.

Zero and 100 percent still contain uncertainty

At x = 0, n = 10, the 95 percent Wilson interval is 0 to 0.2775327999. At x = 10, n = 10, it is 0.7224672001 to 1. The observed endpoints do not collapse the interval to a point.

These boundary examples are useful during review because a display of 0 or 100 percent can look more conclusive than its denominator supports. They do not mean that the client has a fixed hidden success probability across future sessions. Changing supports, health, preference, partner behavior, opportunity quality, or setting can change what is being observed.

A binomial interval is not a single-case conclusion

The calculation assumes a fixed n, two mutually exclusive outcomes, a stable probability, and trials that are independent enough for the chosen model. ABA time series often have trend, phase changes, serial dependence, and contextual variation. These features can make a pooled binary-trial model a poor description even when its arithmetic is flawless.

For an intervention-effect question, retain the ordered session data and graph. The WWC single-case technical documentation describes visual analysis across level, trend, variability, overlap, immediacy, and consistency, with repeated demonstrations needed for a causal inference in its standards framework. A Wilson interval around one pooled proportion does not supply those demonstrations.

The output is also not a prediction interval for a later session, a Bayesian credible interval, a confidence score for the person, or a clinical-importance threshold. The repeated-sampling interpretation concerns the procedure under its model, not the probability that this already calculated interval contains a fixed parameter.

Preserve the denominator story with the result

Retain the raw opportunity rows, exclusions, response codes, observation order, calculator version, formula, confidence level, full-precision output, display rounding, and reviewer. Use a new record version when the response or opportunity definition changes. Never overwrite the observation record to make an interval look narrower.

The Standards for Educational and Psychological Testing provide broad guidance on validity, reliability, fairness, and intended score uses. They help frame the interpretive questions; they do not certify this calculator or turn local opportunity data into a standardized measure.

Use coded IDs and approved systems when a worksheet contains protected health information. HHS publishes separate high-level summaries of the federal Privacy Rule and Security Rule. Applicable consent, access, retention, safeguards, accessibility, payer, legal, and organizational controls still require local review.

Related resources

Sources