An ABA Fleiss kappa calculator summarizes nominal-category agreement when every independent unit receives the same number of ratings. The people rating one unit need not be the people rating another. This worksheet keeps each unit's category counts, its pair agreement, the pooled category proportions, expected agreement, and the final coefficient visible. Fix the design before calculating because the same arithmetic does not fit every many-observer dataset.

Clinicians & ABA Professionals / Data, Outcomes and Clinical Decision-Making.

Confirm the design before using Fleiss kappa

Fleiss's original many-rater paper extends nominal kappa to samples in which each subject receives the same number of ratings, while the particular raters can vary by subject. That is narrower than a generic instruction to use Fleiss kappa whenever more than two people observe behavior.

Use the ABA Fleiss kappa calculator only when all of these conditions are met:

  • Categories are mutually exclusive nominal labels, not ordered score levels.
  • Every unit has exactly m ratings and m >= 2.
  • A unit contributes counts across categories whose row total equals m.
  • Units used in the calculation are independent for the intended interpretation.
  • Observers produce ratings independently under the same coding rules.
  • The rater pool and sampling process match the intended generalization.

Stop if some units have fewer ratings, missing entries, changing row totals, or ratings that should receive distance-based partial credit. The nominal Krippendorff alpha calculator provides a bounded alternative for variable participation and missing values. The weighted kappa calculator addresses two observers using defensibly ordered categories.

Define categories and the fixed rating rule

Write operational definitions before viewing agreement. Category A, B, and C are placeholders, not clinically meaningful labels. If an observation could belong in two categories, repair the coding rules before collecting reliability data. If an observer can select several categories for one unit, this single-label calculator is not the right model.

The rater design also needs a plain-language statement. For example: "Each sampled interval receives four independent ratings from eligible observers drawn from the same trained pool." A stable count of four is part of the estimator, not a formatting preference.

Cohen's original nominal-kappa paper concerns a fixed pair of judges. Fleiss kappa uses a different aggregation for many ratings and pooled category prevalence. Do not transpose a two-rater contingency table into this unit-count worksheet.

Copy the calculation record

Complete this record before entering counts.

FieldPrespecified entryObservable constructIndependent unit of analysisNominal category labelsOperational definition for each categoryRatings required per unit, mEligible rater poolObserver independence procedureSampling period and inclusion ruleExclusion ruleMissing or unequal rating-count ruleStopData and coding-rule versionIntended decision consequenceAnalyst and independent verifier

The BACB Ethics Codes and BCBA Test Content Outline provide professional and measurement context. They do not require Fleiss kappa, define a sufficient sample, or establish a universal cutoff. The Standards for Educational and Psychological Testing support connecting evidence to the intended interpretation and use; they do not make a coefficient proof of validity.

Enter a unit-by-category count table

Use one row per independent unit and one column per nominal category. The row sum must equal the fixed number of ratings m.

UnitCategory A countCategory B countCategory C countRow totalPair agreement P_i12345Category totalsn x m =

Every cell must be a nonnegative integer. Check every row against the same declared m; a row total that differs by one is a design failure, not a rounding issue. Preserve a separate trace from these counts to the original observer-level ratings so disagreements can be reviewed. Aggregated counts are not a substitute for the source record.

Calculate agreement for each unit

Let uppercase N denote the number of included units, m the fixed ratings per unit, and n_ij the number of ratings assigning unit i to category j. Keep those quantities distinct. With m ratings per unit, calculate:

Pi = [sumj(n_ij^2) - m] / [m(m - 1)]

This equals the proportion of ordered rater pairs that agree on that unit. For four ratings, a 4-0-0 row has Pi = 1.0000000000; a 3-1-0 row has Pi = 0.5000000000; and a 2-1-1 row has P_i = 0.1666666667.

Average across the N included units:

Pbar = sumi(P_i) / N

Keep unit values visible. A single mean can hide a small set of ambiguous categories, weak definitions, observer drift, or context-specific disagreement.

Pool category prevalence and calculate kappa

For each category j, calculate the pooled proportion:

pj = sumi(n_ij) / (N x m)

Confirm that sumj(pj) = 1. Expected agreement from these pooled marginals is:

Pe = sumj(p_j^2)

Then calculate:

kappaF = (Pbar - Pe) / (1 - Pe)

Use full precision until display rounding. If P_e = 1, the denominator is zero and kappa is undefined. Negative values can occur when observed pair agreement is below the agreement expected from the pooled category distribution. Keep a negative result visible; do not replace it with zero.

Fictional example with four ratings per unit

Assume eight fictional independent observation units, four ratings per unit, and three nominal categories.

UnitABCRow totalP_i140041.0000000000231040.5000000000331040.5000000000404041.0000000000503140.5000000000603140.5000000000700441.0000000000810340.5000000000Category totals1112932

The mean observed pair agreement is Pbar = 0.6875000000. Pooled category proportions are pA = 0.3437500000, pB = 0.3750000000, and pC = 0.2812500000. Therefore:

  • P_e = 0.3378906250.
  • kappa_F = 0.5280235988.

These fictional values verify the worksheet arithmetic. They do not establish observer accuracy, a representative sample, sound category definitions, or a clinical decision threshold.

Hold pair agreement constant and shift prevalence

Now retain exactly three unanimous rows and five 3-1-0 patterns, but assign most ratings to Category A. Use rows 4-0-0, 3-1-0, 3-1-0, 4-0-0, 3-1-0, 3-1-0, 4-0-0, and 3-0-1.

Every row's pair-agreement pattern is unchanged, so Pbar stays 0.6875000000. Category totals are now 27, 4, and 1, producing pooled proportions 0.8437500000, 0.1250000000, and 0.0312500000. Expected agreement rises to 0.7285156250, and kappaF becomes -0.1510791367.

Sensitivity viewP_barP_eFleiss kappaBalanced fictional table0.68750000000.33789062500.5280235988Concentrated pooled prevalence0.68750000000.7285156250-0.1510791367

This negative sensitivity result does not by itself show that observers performed worse. It shows that the pooled marginal distribution changed the chance-agreement term enough to exceed the unchanged mean pair agreement. Treat it as a prompt to inspect category prevalence and the sampling design, not permission to rebalance, delete, or relabel observations. Report the pooled proportions and unit patterns with the coefficient.

Apply stop conditions without silent repair

Do not calculate when any row total differs from m, a cell is negative or noninteger, m < 2, there are no included units, or P_e = 1. A row with a missing rating cannot be repaired by inserting the modal category or zero.

The basic worksheet also stops for ordered ratings, multiple selections per observer, repeated dependent observations treated as independent, nested designs, adjudicated consensus replacing original ratings, confidence intervals, hypothesis tests, or comparisons between coefficients. Use a qualified statistician or psychometrician to select and document a method for those designs.

Interpret agreement with direct evidence

Fleiss kappa is a chance-corrected summary under its design assumptions. It does not identify which observer is correct because there is no reference standard in the count table. It does not establish treatment integrity, criterion validity, treatment effect, experimental control, staff competence, or clinical importance.

Review original ratings, unit-level disagreement, category frequencies, sampling coverage, direct graphs, coding definitions, observer training, drift checks, and the consequence of disagreement. Decide in advance how results will prompt retraining, definition revision, additional observation, consultation, or no action. Avoid universal qualitative labels detached from the use, uncertainty, sampling process, prevalence, and consequences.

Limit the worksheet to necessary information. The HHS Privacy Rule summary and HHS Security Rule summary do not authorize a specific dataset or system. Follow applicable consent, assent, confidentiality, retention, access, accessibility, payer, contractual, state, organizational, and legal controls.

Related resources

Sources