An ABA intraclass correlation coefficient calculator can summarize reliability among quantitative ratings only after the design question is fixed. This worksheet accepts a complete target-by-rater matrix, derives the two-way ANOVA mean squares, and displays four ICC forms. Choose absolute agreement or consistency, and single ratings or a mean of raters, before comparing results.

Clinicians & ABA Professionals / Data, Outcomes and Clinical Decision-Making.

Decide which reliability claim the design can support

An intraclass correlation coefficient, or ICC, compares variability among targets with variability attributable to raters and residual disagreement. It is not one formula with interchangeable labels. Shrout and Fleiss tied coefficient choice to the reliability model and intended use. McGraw and Wong later separated the model, measurement type, and agreement definition.

Use this basic worksheet only when all of these statements are true:

  • every target is independent of every other target;
  • the same identified set of k raters scores every target;
  • ratings are quantitative on the same defensible scale;
  • the matrix is complete, with no imputation or omitted cells;
  • rater identities, scoring rules, units, setting, and data version are frozen; and
  • a qualified methods reviewer accepts the two-way crossed ANOVA calculation.

Koo and Li recommend reporting the ICC model, type, and definition because different forms can yield different values and meanings. The ABA intraclass correlation coefficient calculator below therefore shows four values but does not choose one after the results are visible.

This worksheet uses a two-way crossed calculation, but it does not decide whether the rater factor should be treated as fixed or random for inference. That choice depends on whether the claim is limited to the named raters or generalized to a larger rater population. Record the intended population and obtain methods review before carrying the point estimate beyond this matrix.

Before entering data, answer two questions:

  1. Agreement definition: Must raters return the same numeric value? Choose absolute agreement. May raters differ by a stable additive offset while preserving target ordering? Consistency may answer that narrower question.
  2. Measurement type: Will one rater's value be used operationally? Choose single rating. Will the operational score always be the mean of these same k raters? Choose mean of k.

Do not describe a mean-of-raters ICC as the reliability of one person's rating. Do not describe a consistency ICC as exact agreement.

Freeze the record before calculating

Copy this provenance block into the review record.

FieldPrespecified entryConstruct and observable definitionUnit and numeric scaleIndependent target definitionIncluded targets and exclusionsRater names or coded IDsRater training and independence ruleSame-rater-for-every-target confirmationYes / NoMissing-cell ruleStop; do not use this worksheetAgreement definitionAbsolute agreement / ConsistencyMeasurement typeSingle rating / Mean of k ratingsIntended population and generalizationData export and scoring-rule versionPrespecified decision consequenceAnalyst and independent verifier

The BACB Ethics Codes and BCBA Test Content Outline provide professional and measurement context. They do not require this statistic, select an ICC form, or supply an acceptance threshold. The Standards for Educational and Psychological Testing support aligning evidence and interpretation with the proposed use; they do not turn a sample coefficient into proof of validity.

Enter the complete target-by-rater matrix

Keep raw values. Add rows for targets and columns for raters without averaging first.

TargetRater 1Rater 2Rater 3Target mean123456Rater meanGrand mean:

Let n be the number of targets and k the number of raters. For value xij, let x-bari. be target i's mean, x-bar.j be rater j's mean, and x-bar.. be the grand mean.

Calculate:

  • SS targets = k x sum[(target mean - grand mean)^2]
  • SS raters = n x sum[(rater mean - grand mean)^2]
  • SS error = sum[(x_ij - target mean - rater mean + grand mean)^2]
  • MS targets = SS targets / (n - 1)
  • MS raters = SS raters / (k - 1)
  • MS error = SS error / [(n - 1)(k - 1)]

Retain full precision through the formulas. Round only the displayed values.

Calculate four two-way ICC forms

Use MST, MSR, and MS_E for the three mean squares.

Prespecified questionFormulaReported valueAbsolute agreement, single rating: ICC(A,1)(MST - MSE) / [MST + (k - 1)MSE + (k/n)(MSR - MSE)]Consistency, single rating: ICC(C,1)(MST - MSE) / [MST + (k - 1)MSE]Absolute agreement, mean of k: ICC(A,k)(MST - MSE) / [MST + (MSR - MS_E)/n]Consistency, mean of k: ICC(C,k)(MST - MSE) / MS_T

The formulas expose two predictable differences. A rater mean shift affects absolute agreement but drops out of consistency. Averaging k ratings generally produces a different reliability claim from using one rating. Hallgren's observational-data tutorial also warns that ICC variants depend on design and the kind of agreement being assessed.

Report a negative estimate as calculated. It can signal that within-target disagreement exceeds the between-target signal under the chosen model. Silently setting it to zero erases diagnostic information. A zero denominator or nonfinite result is a stop condition, not a result to label acceptable.

Fictional example: six targets scored by three raters

Suppose three trained observers independently assign a quantitative implementation score to the same six fictional, independent session samples. These numbers demonstrate arithmetic only.

TargetRater 1Rater 2Rater 3Target mean11011910.0000000000220222121.0000000000330293130.0000000000440434141.3333333333550495250.3333333333660636161.3333333333Rater mean35.000000000036.166666666735.8333333333Grand mean: 35.6666666667

The calculation record is:

  • SS targets = 5436.0000000000; df = 5; MS targets = 1087.2000000000.
  • SS raters = 4.3333333333; df = 2; MS raters = 2.1666666667.
  • SS error = 15.6666666667; df = 10; MS error = 1.5666666667.
  • ICC(A,1) = 0.9954155078 and ICC(C,1) = 0.9956893916.
  • ICC(A,3) = 0.9984671510 and ICC(C,3) = 0.9985589895.

These high fictional values do not establish clinical accuracy, an adequate observation sample, or a suitable scoring system. Current recommendations for test-retest ICC documentation emphasize the selected formula and uncertainty. This compact worksheet intentionally does not calculate confidence intervals.

Run model-sensitivity controls

First, add 8 to every Rater 2 value while leaving the other values unchanged. The target and error mean squares stay 1087.2 and 1.5666666667, but MS raters rises to 154.1666666667. The consistency results remain ICC(C,1) = 0.9956893916 and ICC(C,3) = 0.9985589895; absolute agreement falls to ICC(A,1) = 0.9305694448 and ICC(A,3) = 0.9757332455. That gap is the question, not a reason to select the larger coefficient.

Second, compress the target range by subtracting 0, 8, 16, 24, 32, 40 from rows 1 through 6. Within-target patterns remain, but MS targets falls to 50.4. The four results become 0.9071207430, 0.9122042341, 0.9669966997, and 0.9689153439. Reliability depends on the sampled target heterogeneity, so a coefficient does not travel unchanged to a narrower or broader population.

Record both controls:

Sensitivity viewICC(A,1)ICC(C,1)ICC(A,k)ICC(C,k)MeaningOriginal matrix0.99541550780.99568939160.99846715100.9985589895Baseline arithmeticRater 2 plus 80.93056944480.99568939160.97573324550.9985589895Absolute-agreement penalty is visibleCompressed target range0.90712074300.91220423410.96699669970.9689153439Less between-target variability

Stop when the design exceeds this worksheet

Do not use these formulas for missing cells, changing rater panels, different raters per target, one-way random designs, nested observations, repeated sessions from the same client, ordinal categories, or dependence created by overlapping windows. Do not infer a confidence interval from the point estimate. Those cases need a statistician or psychometrician to specify the model and estimand.

Reliability is not validity, accuracy, clinical importance, treatment integrity, experimental control, or causation. Review the raw matrix and the size and direction of disagreements. Pair this tool with a direct graph and a clinical error analysis. Do not replace observable definitions, competency checks, supervision, assent, client priorities, or caregiver interpretation with a coefficient.

Restrict entries to what the analysis requires. Federal HHS Privacy Rule and Security Rule summaries describe safeguards but do not authorize a specific dataset, storage method, or disclosure. Apply organizational, contractual, payer, state, legal, accessibility, retention, and security requirements before using identifiable records.

Related resources

Sources