An ABA weighted kappa agreement calculator summarizes how two independent observers use the same ordered categories on the same units. This worksheet preserves the full count table and compares linear and quadratic agreement weights. Define the category order and weighting policy before viewing results; different weights encode different consequences for near and far disagreements.

Clinicians & ABA Professionals / Data, Outcomes and Clinical Decision-Making.

Use weighted kappa only for genuinely ordered categories

Weighted Cohen kappa gives partial agreement credit when two ratings differ by one or more steps on a prespecified ordinal scale. It is appropriate only when category order has meaning. A rating of 0, 1, 2, 3 might represent increasing levels of a defined observation code, but the numbers are labels unless the scoring system supports their order and the chosen distance rule.

Cohen's original weighted-kappa paper introduced scaled disagreement or partial credit for two raters. Hallgren's observational-data tutorial explains that the method is used for ordinal ratings and that the weighting scheme changes how disagreement magnitude contributes to the result.

Do not use this worksheet when:

  • Categories are nominal, such as topography names with no defensible order.
  • More than two observers contribute to the same summary.
  • Observers do not score the same units independently.
  • Cells are missing, imputed, or derived from unequal sampling.
  • Repeated units from a client are treated as independent.
  • The clinical cost of disagreement cannot be represented by the selected weights.

For binary nominal coding, use the Cohen kappa prevalence-sensitivity calculator. For quantitative values rather than ordered categories, consider a design-appropriate ICC or paired-measurement analysis.

Use the ABA weighted kappa agreement calculator only after the two observers, common units, ordered categories, and weight policy are fixed.

Prespecify category and weight rules

Copy this record before tallying the table.

FieldPrespecified entryObservable constructCommon unit scored by both observersOrdered category labelsOperational definition for each categoryWhy the order is defensibleWhy equal step spacing is or is not defensibleObserver training and independence ruleInclusion and exclusion rulesMissing or unpaired observation ruleStopAgreement-weight policyLinear / Quadratic / Qualified custom matrixData and scoring-rule versionPrespecified decision consequenceAnalyst and independent verifier

The BACB Ethics Codes and BCBA Test Content Outline supply professional and measurement context, not a required coefficient or cutoff. The Standards for Educational and Psychological Testing support linking evidence to the intended interpretation and use. None of these sources makes ordinal categories equally spaced by default.

Build the square count table

Use one row per Observer A category and one column per Observer B category. Every common unit contributes to exactly one cell.

Observer A category / Observer B category0123Row total0123Column totalN =

Confirm that every cell is a nonnegative integer and that all row and column totals equal N. Preserve the table even if a single coefficient is copied into a report. Category prevalence and asymmetric observer use remain visible in the marginals.

Calculate exact and weighted agreement

For K ordered categories numbered 0 through K - 1, use agreement weights:

  • linear: w_ij = 1 - |i - j| / (K - 1)
  • quadratic: w_ij = 1 - (i - j)^2 / (K - 1)^2

For each cell, pij = nij / N. Expected cell probability is the row proportion multiplied by the column proportion. Then calculate:

  • Exact observed agreement: sum(p_ii).
  • Weighted observed agreement: Po,w = sum(wij x p_ij).
  • Weighted expected agreement: Pe,w = sum(wij x rowi proportion x columnj proportion).
  • Weighted kappa: kappaw = (Po,w - Pe,w) / (1 - Pe,w).

Use full precision through the final division. If P_e,w = 1, the coefficient is undefined. A negative result is possible and should not be silently replaced with zero.

A review of many-rater ordinal measures gives these common linear and quadratic agreement weights and distinguishes two-rater weighted kappa from other designs. Robustness work on kappa-type coefficients reinforces that weights must be assigned to reflect the seriousness of category differences before calculating the statistic.

Make the two weight matrices visible

For four categories, the linear agreement matrix is:

Linear weightB0B1B2B3A01.00000000000.66666666670.33333333330.0000000000A10.66666666671.00000000000.66666666670.3333333333A20.33333333330.66666666671.00000000000.6666666667A30.00000000000.33333333330.66666666671.0000000000

The quadratic agreement matrix is:

Quadratic weightB0B1B2B3A01.00000000000.88888888890.55555555560.0000000000A10.88888888891.00000000000.88888888890.5555555556A20.55555555560.88888888891.00000000000.8888888889A30.00000000000.55555555560.88888888891.0000000000

Quadratic weights grant more partial credit to nearby disagreements than linear weights. A larger quadratic result is not automatically better or more valid. It reflects a different loss function. A recent ordinal-agreement guide notes there is no universal rule for selecting a weighting scheme and recommends retaining category distributions and unscaled agreement information.

Fictional example: four ordered observation codes

Assume two observers independently assign one of four prespecified ordered codes to 80 fictional observation units.

Observer A category / Observer B category0123Row total0183002112204026203172223004711Column total202625980

Exact observed agreement is (18 + 20 + 17 + 7) / 80 = 0.7750000000.

With linear weights:

  • P_o,w = 0.9250000000.
  • P_e,w = 0.6376041667.
  • kappa_w = 0.7930439782.

With quadratic weights:

  • P_o,w = 0.9750000000.
  • P_e,w = 0.7850347222.
  • kappa_w = 0.8837021483.

The coefficients differ because the quadratic matrix gives more credit to one-step disagreements. The calculation does not say which loss policy fits the clinical use. A reviewer must evaluate the category construction and the consequences of misclassification.

Run reversal and disagreement-distance controls

Reverse the category order for both observers, turning 0, 1, 2, 3 into 3, 2, 1, 0 while reversing rows and columns. Absolute distances are unchanged. Exact agreement remains 0.7750000000, linear weighted kappa remains 0.7930439782, and quadratic weighted kappa remains 0.8837021483. This invariance check applies to the symmetric linear and quadratic distance weights used here; it is not promised for an arbitrary custom matrix. If these results change, the table or weight mapping was reversed incorrectly.

Next, move four Observer A category 1 cases from Observer B category 2 to category 3. Exact agreement still equals 0.7750000000, yet disagreements are farther apart. The linear result falls to 0.7552155772 and the quadratic result falls to 0.8176014592.

Sensitivity viewExact agreementLinear weighted kappaQuadratic weighted kappaOriginal table0.77500000000.79304397820.8837021483Both category orders reversed0.77500000000.79304397820.8837021483Four disagreements moved farther0.77500000000.75521557720.8176014592

This control shows the utility and the limitation of weighting. Exact agreement ignores disagreement distance. Weighted results respond to distance only through the chosen matrix.

Interpret the table before the coefficient

Inspect which categories each observer uses, where disagreements concentrate, and whether the operational definitions need repair. McHugh's kappa review discusses limitations of chance-corrected agreement and questions fixed acceptability judgments. Avoid universal labels or cutoffs detached from the decision, sample, prevalence, precision, and consequences.

Weighted kappa is not observer accuracy because there is no gold standard in this table. It is not evidence that the scale is valid, that a target improved, that treatment was implemented with fidelity, or that a disagreement is clinically harmless. Preserve direct graphs, original observations, adjudication notes, training records, and the prespecified scale.

Use a statistician or psychometrician for custom weights, unequally spaced categories, more than two observers, missing data, clustered observations, confidence intervals, or inferential comparisons. Do not use quadratic weighting merely because it produces a higher coefficient.

Limit the worksheet to necessary data. The HHS Privacy Rule summary and HHS Security Rule summary do not authorize a particular dataset or tool. Follow applicable consent, assent, confidentiality, payer, contractual, state, organizational, retention, accessibility, and legal controls.

Related resources

Sources