An ABA Scott pi agreement calculator can show how much two observers agree after accounting for a chance model built from their pooled category use. The arithmetic is compact, but the design decisions are not. Freeze the units, categories, observer roles, exclusions, data version, and intended use before calculating. Then keep the full disagreement table beside the coefficient.
Clinicians & ABA Professionals / Data, Outcomes and Clinical Decision-Making.
Decide whether Scott pi fits the rating design
Scott's original 1955 article, listed in the journal's Volume 19, Issue 3 record, addresses reliability for nominal coding. This worksheet translates the two-rater calculation into a clinical review aid; the paper does not make the statistic an ABA standard or supply a universal cutoff.
Use this ABA Scott pi agreement calculator only when two observers independently assign the same complete set of units to one category each and the categories are mutually exclusive and nominal. A unit might be a scored opportunity, a predefined interval, or another observable event whose sampling and dependence assumptions a qualified reviewer has approved. Do not use this version for ordered ratings, partial credit, multiple selections per unit, or unequal sets of rated units.
The BACB Ethics Codes and the current BCBA Test Content Outline provide professional context for measurement and data-based practice. Neither source requires Scott pi or converts a coefficient into evidence of competence, treatment integrity, accuracy, or clinical benefit. The Standards for Educational and Psychological Testing likewise place an interpretation within its intended use and supporting evidence.
Freeze the agreement plan before calculation
Copy and complete this record before looking at any result.
Design fieldPrespecified entryObservable constructIndependent unit definitionUnit inclusion windowNominal categories and operational definitionsObserver A roleObserver B roleBlinding or independence procedureMissing-pair ruleStop; this worksheet requires complete pairsExclusions decided before reviewData version and extraction timePrimary statisticScott piPrespecified comparatorCohen kappa, or noneIntended decision and consequenceQualified reviewers
Changing categories, exclusions, or the observation window after seeing pi changes the question. Preserve the original plan and document any new analysis as a separate version.
Build the two-rater contingency table
Create a square K x K table. Rows are Observer A's categories, columns are Observer B's categories, and each paired unit contributes exactly one count. Keep zero cells because they preserve possible pairings even when none were observed.
Observer A \ Observer BCategory 1Category 2...Row totalCategory 1Category 2...Column totalN
Confirm that every count is a nonnegative integer, every included unit appears once, row totals and column totals each sum to N, and both axes contain the same K categories in the same order. Stop if any paired observation is missing; this page does not silently delete or impute units.
Calculate observed agreement
Let n_cc be the diagonal count for category c. Exact agreement is the diagonal total divided by the number of paired units:
Po = (sumc n_cc) / N
Retain the diagonal counts and the off-diagonal cells. P_o alone cannot reveal whether disagreement clusters around one definition, observer, setting, or time period.
Pool the category margins
For each category c, calculate Observer A's row total rc and Observer B's column total sc. Scott pi pools the two observers' margins:
pc = (rc + s_c) / (2N)
Check that sumc pc = 1. The expected-agreement term is:
Pe,pi = sumc p_c^2
This differs from Cohen's original kappa, which uses the product of the two observer-specific proportions for each category. The distinction is a modeling choice about expected agreement, not a verdict about which observer is correct.
Calculate Scott pi
With unrounded values, compute:
Scott pi = (Po - Pe,pi) / (1 - P_e,pi)
If P_e,pi = 1, the denominator is zero and pi is undefined. That boundary occurs when all pooled ratings occupy one category, a design that supplies no category variation for this coefficient. Report the boundary instead of forcing the answer to zero or one. Round only the final display.
Work the fictional example
This fictional table contains 40 paired units and two nominal categories. It contains no client, caregiver, clinician, provider, payer, assessment, or practice data.
Observer A \ Observer BCriterion metCriterion not metRow totalCriterion met20424Criterion not met41216Column total241640
The diagonal contains 20 + 12 = 32 agreements, so Po = 32/40 = 0.8000000000. Both observers use the two categories 24 and 16 times. The pooled proportions are therefore p1 = 48/80 = 0.6000000000 and p_2 = 32/80 = 0.4000000000.
P_e,pi = 0.6000000000^2 + 0.4000000000^2 = 0.5200000000
Scott pi = (0.8000000000 - 0.5200000000) / (1 - 0.5200000000) = 0.5833333333
Because the observer-specific margins are identical here, Cohen's chance term is also 0.5200000000, and Cohen kappa is also 0.5833333333. Equality in this example does not make the formulas interchangeable.
Run a pooled-margin sensitivity view
The second fictional table holds N=40, diagonal agreement 32, P_o=0.8000000000, and the pooled category proportions 0.6000000000/0.4000000000 fixed. Only the observer-specific margins change, from 24/16 for each observer to 28/12 for Observer A and 20/20 for Observer B.
Observer A \ Observer BCriterion metCriterion not metRow totalCriterion met20828Criterion not met01212Column total202040
Scott pi remains 0.5833333333 because its observed agreement and pooled margins are unchanged. Cohen's expected term becomes (0.70 x 0.50) + (0.30 x 0.50) = 0.5000000000, giving kappa 0.6000000000. Brennan and Prediger's methods paper emphasizes that agreement coefficients encode assumptions about fixed or free margins. This sensitivity view makes one such assumption visible; it does not identify a winner.
Report both tables only if the comparison was planned or clearly labeled exploratory. Never move observations between cells to improve a statistic.
Interpret the result without a universal label
Scott pi is a chance-corrected summary under a pooled-marginal model. It does not locate the source of disagreement, validate category definitions, prove observer independence, establish accuracy against a reference standard, demonstrate treatment response, or justify a clinical decision by itself.
Avoid automatic labels such as poor, moderate, or excellent. A qualified reviewer should connect the coefficient, raw cell counts, disagreement pattern, sampling design, clinical consequence, and client or caregiver perspective to the intended use. Report N, the category definitions, both observers' margins, pooled proportions, Po, Pe,pi, pi, exclusions, missingness, planned comparators, and limitations so another reviewer can reconstruct the interpretation.
Stop when the design exceeds this calculator
Pause and obtain statistical or psychometric review when there are more than two observers, ordered categories, partial credit, missing pairs, repeated or clustered units, or changing category sets. Also pause for confidence intervals, hypothesis tests, sample-size planning, or a reference-standard accuracy question. Fleiss kappa, weighted kappa, Krippendorff alpha, Gwet AC1, Randolph free-marginal kappa, and familiar ABA IOA methods answer different questions under different assumptions.
Direct graphs, source observations, observer notes, disagreement review, clinical importance, feasibility, consent, assent when applicable, competence, supervision, and client and caregiver interpretation remain primary evidence. A coefficient is one review input.
Protect the worksheet and its provenance
Use the minimum data necessary for the review. Store unit-level records only in an approved system, restrict access, preserve an audit trail, and avoid pasting protected information into an unapproved calculator. The current HHS Privacy Rule summary and HHS Security Rule summary describe federal requirements for regulated entities. Organizational and state requirements may add obligations. This page is not legal advice.
Copyable result record
Result fieldValueData versionN complete paired unitsCategories and definitionsObserver A marginsObserver B marginsPooled proportionsP_oP_e,piScott piPrespecified comparatorDisagreement pattern reviewedMissing or excluded unitsInterpretation and limitationsReviewer and review date
Related resources
- ABA Cohen Kappa Prevalence-Sensitivity Agreement Calculator for Clinicians
- ABA Fleiss Kappa Multi-Rater Agreement Calculator for Clinicians
- ABA Gwet AC1 Chance-Model Sensitivity Calculator for Clinicians
- ABA IOA Method Sensitivity and Agreement Comparison Calculator for Clinicians
Sources
- Scott, Reliability of Content Analysis: The Case of Nominal Scale Coding.
- Public Opinion Quarterly, Volume 19, Issue 3.
- Cohen, A Coefficient of Agreement for Nominal Scales.
- Brennan and Prediger, Coefficient Kappa: Some Uses, Misuses, and Alternatives.
- Behavior Analyst Certification Board Ethics Codes.
- BCBA Test Content Outline, Sixth Edition.
- Standards for Educational and Psychological Testing.
- HHS Summary of the HIPAA Privacy Rule.
- HHS Summary of the HIPAA Security Rule.