An ABA Gwet AC1 calculator can summarize nominal agreement while keeping its chance model visible. This worksheet accepts the same fixed number of ratings for every independent unit, calculates observed pair agreement, then applies the nominal AC1 expected-agreement term. Use it as a transparent sensitivity tool, not as a way to shop for a larger reliability coefficient.
Clinicians & ABA Professionals / Data, Outcomes and Clinical Decision-Making.
Decide whether this bounded calculator fits
Gwet's peer-reviewed paper develops an agreement coefficient for settings in which high observed agreement can coexist with unstable kappa results. The author's research-paper index provides the broader publication context. That motivation does not establish that AC1 is always preferable.
Use this ABA Gwet AC1 calculator only when:
- each unit is independent for the intended analysis;
- every unit has the same fixed number
rof ratings; - each rating selects exactly one mutually exclusive nominal category;
- the category set and scoring definitions were fixed before data review;
- the observer sampling and assignment process is documented; and
- a point estimate, without a confidence interval or hypothesis test, is sufficient for the bounded review.
Stop if ratings are missing, row totals differ, categories are ordered, the same unit contributes dependent repeated observations, or the decision needs interval estimation. A qualified statistician or psychometrician should select and implement a design-specific method in those cases.
Record the design before entering counts
Copy this provenance block into the analysis record. Do not leave its decisions implicit.
Required fieldPrespecified entryObservable constructUnit definitionNominal category labels and operational definitionsNumber of ratings per unit, rObserver pool and assignment processObserver training and independence safeguardsIncluded dates, settings and peopleExclusions and reasonsData extract name, version and frozen timestampPrimary agreement statisticGwet AC1, nominal and unweightedPrespecified sensitivity comparisonFleiss kappa, or noneDecision the result may informReviewer and review date
The BACB Ethics Codes and BCBA Test Content Outline supply professional and measurement context. They do not require AC1, define a minimum sample, or create a universal interpretation cutoff. The Standards for Educational and Psychological Testing emphasize evidence for an intended interpretation and use; one coefficient does not supply that evidence by itself.
Build the unit-by-category table
Enter one row per independent unit and one column per category. Each cell n_iq is the number of ratings for unit i placed in category q. The row total must equal r.
UnitCategory ACategory BCategory CRow total rUnit pair agreement P_i12345678Totals
Keep observer-level source records even though this table aggregates them. A count table cannot identify which observer supplied which rating or whether observer assignments created dependence.
Calculate observed pair agreement
For N units, Q categories, and r ratings per unit, calculate each unit's agreement:
Pi = [sumq niq(niq - 1)] / [r(r - 1)]
This counts agreeing ordered rating pairs over all ordered rating pairs for the unit. Then average across units:
Pa = (1/N) sumi P_i
Retain the unit values. The mean can hide a subset of units with repeated disagreements or an operational definition that observers apply inconsistently.
Calculate the nominal AC1 chance term
First obtain each category's pooled proportion across the unit table:
piq = (1/N) sumi(n_iq / r)
For unweighted nominal AC1, calculate:
Pegamma = [1/(Q - 1)] sumq piq(1 - pi_q)
Then calculate:
AC1 = (Pa - Pegamma) / (1 - Pe_gamma)
The general multi-rater formulation is documented in the SAS implementation paper. This page implements only the unweighted nominal, fixed-rating-count case. Custom weights, ordinal distance, variance estimation, confidence intervals and tests are outside its scope.
Require Q >= 2. Under valid nominal proportions with at least two categories, this bounded chance term cannot equal one. If software nevertheless reports Pegamma = 1, the denominator is zero and AC1 is undefined; treat that as an input or implementation failure rather than replacing it with zero or one. Keep full precision through all steps and round only the final display.
Work a fictional three-category example
The following counts are invented for arithmetic demonstration. They contain no client, caregiver, clinician, assessment, provider, payer, school, practice or clinical data.
UnitCategory ACategory BCategory CTotalP_i140041.0000231040.5000331040.5000404041.0000503140.5000603140.5000700441.0000810340.5000Totals1112932
The eight unit agreements average to P_a = 0.6875000000. The category proportions are 0.3437500000, 0.3750000000, and 0.2812500000. Therefore:
Pegamma = 0.3310546875
AC1 = (0.6875000000 - 0.3310546875) / (1 - 0.3310546875) = 0.5328467153
Record both the coefficient and its inputs. A bare value of 0.5328 does not reveal the observed agreement, category mix, sample, rating design or disagreement locations.
Result fieldValue to retainUnits and ratings per unitN=8, r=4Category proportions[0.3437500000, 0.3750000000, 0.2812500000]Mean observed agreement0.6875000000AC1 expected agreement0.3310546875AC10.5328467153Prespecified comparatorFleiss kappa 0.5280235988
Compare the chance model without choosing a winner
A prespecified sensitivity analysis can hold the eight unit-level agreement values at [1, 0.5, 0.5, 1, 0.5, 0.5, 1, 0.5] while concentrating pooled category use. In a fictional concentrated table, category totals become [27, 4, 1]; P_a remains 0.6875000000.
For the original table, Fleiss expected agreement is sumq piq^2 = 0.3378906250, and Fleiss kappa is 0.5280235988. For the concentrated table, that expected term is 0.7285156250, and Fleiss kappa is -0.1510791367. AC1 uses a different chance term.
Prespecified viewP_aAC1 PegammaAC1Fleiss P_eFleiss kappaOriginal pooled categories0.68750000000.33105468750.53284671530.33789062500.5280235988Concentrated pooled categories0.68750000000.13574218750.63841807910.7285156250-0.1510791367
This comparison demonstrates chance-model sensitivity. Nothing in the display shows that the concentrated data are better, that AC1 recovered truth, or that Fleiss kappa failed. A published critical comparison argues that AC1 is not a substitute for Cohen's kappa and cautions against reading the coefficients as interchangeable. State the estimand and chance assumption before reviewing the answers, then report both prespecified results even when they point in different directions.
Run the boundary controls
Use these controls before releasing a result:
- Verify every row total equals the prespecified
rand every cell is a nonnegative integer. - Confirm each unit is independent for the planned interpretation and document whether the same observers recur across units.
- Recalculate one mixed row manually. A
[3,1,0]row withr=4hasP_i = 0.5. - Reconcile category totals to
N x rand proportions to1within numerical tolerance. - Calculate
Pa,Pe_gamma, and AC1 from unrounded values. - Preserve the fixed-marginal comparator only if it was named before results were examined.
- Stop when the denominator is zero or an input violates the bounded design.
Never rebalance categories, recode observations, delete difficult units, or switch coefficients because another display looks more favorable. Any correction must follow a documented source-record error and leave an audit trail.
Interpret the output beside direct evidence
AC1 is a chance-corrected agreement summary under a particular model. It does not identify which rating is correct, validate an operational definition, establish treatment integrity, demonstrate clinical improvement, prove experimental control, or determine staff competence. Review disagreements by unit and category alongside raw data, graphs, observer notes, contextual changes, treatment goals and client or caregiver perspectives.
Avoid universal labels such as poor, moderate or excellent unless a qualified reviewer has justified them for the decision, population, design and consequence. Report sample size, ratings per unit, category definitions, pooled proportions, observed agreement, expected agreement, AC1, sensitivity analyses, exclusions, missingness policy, and known limitations.
Protect data and preserve review gates
Use the minimum necessary data. The HHS Privacy Rule summary and HHS Security Rule summary provide federal context but do not determine every legal or contractual obligation. Follow applicable consent, access, retention, encryption, incident-response, payer, regulator, state-law and organizational requirements.
This draft requires external clinical measurement, statistical, psychometric, research-methods, ethics, client, caregiver, privacy, security, accessibility, implementation, payer, regulatory and legal review as applicable. Do not publish or use it as an automated clinical decision rule while those gates remain pending.
Related resources
- ABA Fleiss Kappa Multi-Rater Agreement Calculator for Clinicians
- ABA Cohen Kappa Prevalence-Sensitivity Agreement Calculator for Clinicians
- ABA Randolph Free-Marginal Multi-Rater Kappa Calculator for Clinicians
- ABA IOA Method Sensitivity and Agreement Comparison Calculator for Clinicians
Sources
- BACB Ethics Codes.
- BACB BCBA Test Content Outline, 6th edition.
- Standards for Educational and Psychological Testing.
- Gwet, Computing inter-rater reliability and its variance in the presence of high agreement.
- AgreeStat research papers.
- SAS, Computing Inter-Rater Reliability and Its Variance in the Presence of High Agreement.
- Wongpakaran and colleagues, Gwet's AC1 is not a substitute for Cohen's kappa.
- Fleiss, Measuring Nominal Scale Agreement Among Many Raters.
- HHS Privacy Rule summary.
- HHS Security Rule summary.