An ABA free marginal kappa calculator can make one strong assumption unusually easy to see: chance agreement is fixed at 1/K, where K is the number of theoretically possible nominal categories. Randolph's statistic is simple to compute, but its simplicity makes category governance essential. Specify the category set before examining the coefficient.

Clinicians & ABA Professionals / Data, Outcomes and Clinical Decision-Making.

Establish the free-marginal design

Randolph's 2005 conference paper, cataloged by ERIC, presents a free-marginal alternative for multi-rater nominal agreement. It is a conference paper rather than a peer-reviewed journal article. Later peer-reviewed inequalities and comparisons among multi-rater kappas clarify mathematical relationships without making one coefficient universally correct.

This bounded worksheet requires:

  • independent units for the intended calculation;
  • the same number r of ratings for every unit;
  • one mutually exclusive nominal category per rating;
  • a defensible category universe of size K, fixed in advance;
  • documented observer selection, assignment and independence; and
  • a point-estimate use that does not require uncertainty intervals or inference.

Do not use the worksheet for missing ratings, unequal row totals, ordinal categories, repeated dependent units, clustered observations or adjudicated consensus. Escalate those designs to a qualified methods reviewer.

Freeze the category universe and provenance

The category-count decision is part of the model, not a formatting preference. Complete this block before copying counts.

Required fieldPrespecified entryObservable constructIndependent unit definitionComplete nominal category universeTheoretically possible category count, KWhy every category is possible and relevantRatings per unit, rObserver pool and assignment ruleTraining and independence safeguardsDate range, settings and populationExclusions with reasonsFrozen source extract and timestampPrimary coefficient and sensitivity planDecision the estimate may informReviewer and approval date

An empty observed column may still represent a possible category, but its inclusion needs an a priori rationale and a stable definition across all sampled units. A category that was never available under the protocol does not belong merely to make K larger. Do not add or drop categories after seeing the result.

The BACB Ethics Codes and BCBA Test Content Outline contextualize professional measurement responsibilities. They do not endorse free-marginal kappa or prescribe a cutoff. The Standards for Educational and Psychological Testing support evaluating evidence for the proposed interpretation and consequence rather than treating a single statistic as validation.

Enter the nominal count matrix

For each unit i, enter the number n_ij of ratings assigned to category j. Row totals must equal the prespecified r.

UnitCategory ACategory BCategory CRow total rPair agreement P_i12345678Totals

Keep the observer-level source data, observer assignments and correction log. Aggregation removes information about individual patterns and cannot establish observer accuracy.

Derive observed agreement

Calculate the share of agreeing ordered rating pairs for each unit:

Pi = [sumj nij(nij - 1)] / [r(r - 1)]

Average those values across N units:

Po = (1/N) sumi P_i

The observed-agreement calculation matches the fixed-count Fleiss worksheet for this design. The coefficients are not interchangeable because the free-marginal distinction enters through expected agreement.

Apply the fixed chance assumption

In this ABA free marginal kappa calculator, K is prespecified. Randolph free-marginal expected agreement is:

Pefree = 1 / K

The coefficient is:

kappafree = (Po - 1/K) / (1 - 1/K)

If K is less than two, the calculation is invalid. Retain full numerical precision and round only the displayed coefficient. Do not clamp a negative estimate to zero.

This expected-agreement model does not estimate category proportions from the observed ratings. That can be useful as a deliberate sensitivity assumption, but it is not a factual claim that observers select every category equally often.

Follow a fictional calculation

This arithmetic-only example is invented and contains no real person, assessment, provider, payer, school, practice or treatment information.

UnitCategory ACategory BCategory CTotalP_i140041.0000231040.5000331040.5000404041.0000503140.5000603140.5000700441.0000810340.5000

The unit values average to Po = 0.6875000000. With three prespecified categories, Pe_free = 1/3 = 0.3333333333. Therefore:

kappa_free = (0.6875000000 - 0.3333333333) / (1 - 0.3333333333) = 0.5312500000

Store N, r, K, the category definitions, Po, Pe_free, and the final coefficient. Without those fields, another reviewer cannot reconstruct the model.

Result fieldValue to retainUnits and ratings per unitN=8, r=4Prespecified category countK=3Observed agreement0.6875000000Fixed chance agreement0.3333333333Free-marginal kappa0.5312500000Prespecified comparatorFleiss kappa 0.5280235988

Expose category-count sensitivity

Randolph's paper notes that the coefficient depends on the number of categories and can be increased by superfluous categories. Keep the observed counts and P_o fixed to see the effect of a category-count assumption, but do not treat an unjustified category as a valid alternative.

Prespecified category universeKP_o1/KFree-marginal kappaA, B, C30.68750000000.33333333330.5312500000A, B, C, D40.68750000000.25000000000.5833333333

The second result is higher even though the rating table and observed agreement did not change; Category D has an observed count of zero in every unit. Report the four-category view only if D was genuinely possible under a prespecified protocol. Never invent, split, merge, or remove categories after reviewing the coefficient.

For a prespecified fixed-marginal comparison, the same fictional three-category table has pooled proportions [0.3437500000, 0.3750000000, 0.2812500000]. Fleiss expected agreement is 0.3378906250, giving Fleiss kappa 0.5280235988. The small difference here is dataset-specific. A different category distribution may yield a larger difference.

Complete validation controls

Run every check before interpretation:

  1. Confirm K >= 2 and preserve a written rationale for each category.
  2. Ensure all cells are nonnegative integers and every row totals r.
  3. Verify the unit sample and observer assignments match the frozen protocol.
  4. Manually check at least one unanimous row and one mixed row.
  5. Reconcile all category totals to N x r.
  6. Recalculate Po, 1/K, and kappafree without intermediate rounding.
  7. Compare another coefficient only if the sensitivity plan was prespecified.
  8. Stop for missing data, dependence, ordered categories or requested inference.

A [3,1,0] row with four ratings gives P_i = 0.5000000000. A unanimous [0,4,0] row gives 1.0000000000. Any software output that disagrees with those controls needs investigation before use.

Interpret agreement without claiming accuracy

Free-marginal kappa is a model-based summary of reproducibility. It does not tell which category is correct, demonstrate construct validity, establish treatment integrity, prove staff mastery, measure social validity, or support causal conclusions. Examine unit-level disagreements, original observations, graphs, category confusions, context changes and implementation conditions.

No universal label or threshold is adopted here. A qualified reviewer should connect the estimate, its assumptions and uncertainty to the intended decision and the cost of error. Document client and caregiver perspectives when the operational meaning or consequence affects care.

Maintain privacy and release controls

Use de-identified or minimum-necessary data where feasible. The HHS Privacy Rule summary and HHS Security Rule summary are federal orientation materials, not a complete compliance determination. Apply relevant authorization, access, encryption, retention, incident, state-law, contract and organizational controls.

This page remains a draft pending external review by clinical measurement, statistical, psychometric, research-methods, ethics, client, caregiver, data-governance, privacy, security, accessibility, implementation, payer, regulatory and legal reviewers as applicable. It is not a clinical decision engine.

Related resources

Sources