An ABA Cohen kappa interobserver agreement calculator summarizes agreement between two independent observers who apply the same binary code to the same units. The worksheet keeps the full two-by-two table visible, calculates observed agreement and chance-expected agreement, and reports Cohen kappa beside positive and negative agreement. Two worked tables show why the same overall percentage can produce a different kappa when the category marginals change.
Kappa is one lens on a particular coding record. It does not prove that the definition is valid, that either observer is accurate, or that a treatment worked.
Clinicians & ABA Professionals / Data, Outcomes and Clinical Decision-Making.
Decide whether this is the right agreement question
Use this calculator only when two observers independently classify every included unit into the same two nominal categories, such as “target scored” and “target not scored.” Here, independent scoring means each observer finishes and records a rating without seeing the other observer’s rating. It does not mean that repeated units from one client can automatically be analyzed as statistically independent observations. The unit might be a prespecified interval, trial, opportunity, or event-matching decision. Both observers need the same units in the same order; an omitted unit is not automatically a negative.
This is not a general replacement for ABA interobserver-agreement methods. The IOA method sensitivity calculator compares six familiar percentage methods for count and interval records. Kappa instead asks how much the observed binary agreement exceeds the agreement expected from the observers’ marginal category rates. That chance correction can be informative, but it also makes the result sensitive to how common each category is.
Do not place counts, durations, latency values, ordered severity levels, or ratings from more than two observers into this binary worksheet. Do not collapse a multicategory record after seeing which split gives the best result. Those designs need a measure chosen for their actual scale and sampling structure.
Freeze the two-by-two table
Let the rows represent Observer A and the columns represent Observer B:
Observer B positiveObserver B negativeRow totalObserver A positiveaba + bObserver A negativecdc + dColumn totala + cb + dN
Here, a is the number of units both observers scored positive, d is the number both scored negative, and b and c preserve the two directions of disagreement. The total is:
N = a + b + c + d
All four cells must be nonnegative integers. N must equal the number of common, valid, independently scored units. If one observer has a missing or invalid unit, resolve the pair under a prespecified rule rather than silently filling a cell.
The original paper, A Coefficient of Agreement for Nominal Scales, defines kappa as agreement after accounting for chance agreement implied by the raters’ marginal distributions. The formula does not determine whether the units, definitions, or observation sample are clinically defensible.
Calculate observed, expected, and category-specific agreement
Observed agreement is the agreeing diagonal divided by all included units:
Po = (a + d) / N
Calculate each observer’s marginal proportions:
PA+ = (a + b) / N and PA- = (c + d) / N
PB+ = (a + c) / N and PB- = (b + d) / N
The agreement expected from those marginals is:
Pe = (PA+ x PB+) + (PA- x PB-)
Then:
kappa = (Po - Pe) / (1 - Pe)
Retain full precision until display. When Pe = 1, the denominator is zero and kappa is undefined. This occurs, for example, when both observers put every unit in the same single category. Overall agreement can then be 100 percent while the table contains no variation from which this chance correction can be calculated.
Positive and negative agreement keep the two categories visible:
positive agreement = 2a / (2a + b + c)
negative agreement = 2d / (2d + b + c)
If the denominator for one category-specific measure is zero, mark that measure undefined rather than displaying zero or 100 percent. The Feinstein and Cicchetti analysis of high agreement with low kappa recommends examining positive and negative agreement separately to understand the marginals behind an omnibus value.
Complete the record before calculating
Before using an ABA Cohen kappa interobserver agreement calculator, freeze the coding record and review rules:
FieldPrespecified recordAgreement record ID and data versionClient-selected or client-informed purposeExact common unit and matching rulePositive code and examplesNegative code and examplesAmbiguous, prompted, unavailable, missing, and invalid rulesObservation window, setting, activity, support, and partnerObserver A and Observer B rolesHow independent scoring was protectedEligible units, sampled units, and exclusionsCalculator version and display ruleReviewer and calculation date
Copy the four cells into a separate calculation block:
ResultFull-precision valueDisplayed value or notea, both positiveb, A positive/B negativec, A negative/B positived, both negativeNObserved agreement PoExpected agreement PeCohen kappaPositive agreementNegative agreementDisagreement units reviewedSampling or dependence concern
Retain the unit-level record behind the table. Four totals cannot show whether disagreements cluster around one observer, one response topography, one setting, one prompt condition, or a small group of hard-to-score units.
Fictional example: a balanced table
All values in this example are invented. Suppose two observers independently score 100 common intervals. Their table contains a = 40, b = 10, c = 5, and d = 45.
Observed agreement is:
Po = (40 + 45) / 100 = 0.85
Observer A scores 50 percent positive and 50 percent negative. Observer B scores 45 percent positive and 55 percent negative. Therefore:
Pe = (0.50 x 0.45) + (0.50 x 0.55) = 0.50
kappa = (0.85 - 0.50) / (1 - 0.50) = 0.70
The category-specific values are:
positive agreement = 80 / 95 = 0.8421052632
negative agreement = 90 / 105 = 0.8571428571
With a prespecified three-decimal display, record observed agreement 0.850; expected agreement 0.500; kappa 0.700; positive agreement 0.842; negative agreement 0.857. Do not add a universal label such as “good” or “acceptable.” Whether the discrepancies matter depends on the decision, sampling plan, response definition, risk, and client context.
Prevalence sensitivity: the same 85 percent is not the same table
Now keep 100 units and 15 disagreements but use a = 10, b = 10, c = 5, and d = 75. Observed agreement remains 0.85. The marginals change: Observer A is 20 percent positive and Observer B is 15 percent positive.
Observer B positiveObserver B negativeRow totalObserver A positive101020Observer A negative57580Column total1585100
Pe = (0.20 x 0.15) + (0.80 x 0.85) = 0.71
kappa = (0.85 - 0.71) / (1 - 0.71) = 0.4827586207
Positive agreement is 20 / 35 = 0.5714285714; negative agreement is 150 / 165 = 0.9090909091. The drop in kappa does not mean the arithmetic is broken. It exposes how the category prevalence and marginals affect the chance-corrected value. It also shows why reporting kappa without the table, observed agreement, and category-specific values can hide the clinically useful disagreement pattern.
Treat this as a sensitivity demonstration, not permission to rebalance a sample after collection. A deliberately enriched training set may be useful for calibration, but its kappa does not describe a naturalistic sample unless that connection is separately justified.
Run label-reversal controls
First swap Observer A and Observer B. Cells b and c exchange places; kappa, observed agreement, positive agreement, and negative agreement should remain unchanged. If they do not, inspect the formulas or copied table.
Next reverse the category names while preserving every unit. Cells a and d exchange roles, as do the disagreement directions. Kappa and overall agreement stay the same, while positive and negative agreement trade places. This is a useful arithmetic control and a reminder that the category-specific labels carry meaning that kappa alone does not.
A negative kappa is possible when observed agreement is below the agreement expected from the marginals. Do not automatically interpret that as deliberate disagreement. Check unit alignment, observer independence, missingness, definition drift, category coding, and sampling before offering an explanation.
Keep agreement, accuracy, and validity separate
High agreement has several possible explanations. The observers may both follow a clear definition, repeat the same error, or encounter a sample dominated by one easy category. Kappa cannot distinguish among those explanations. An external criterion may itself be fallible, and calling one observer the “gold standard” does not establish accuracy.
The ABA spreadsheet tool described by Reed and Azulay illustrates common IOA calculations, while the review by Essig, Rotta, and Poling discusses the profession’s uneven attention to agreement and procedural fidelity. Those sources provide ABA context; they do not prescribe Cohen kappa as the default clinical agreement measure.
The BCBA Test Content Outline addresses measurement, operational definitions, reliability, graphing, and data-based decisions. The BACB ethics-code page points clinicians to the current professional code. Neither supplies a universal kappa threshold. The Standards for Educational and Psychological Testing provide broader guidance on validity, reliability, fairness, and intended interpretations; a local coefficient cannot carry those arguments by itself.
Preserve direct graphs, raw units, disagreement notes, treatment-integrity data, observer training records, and the client’s and caregiver’s interpretation. Do not pool repeated intervals from one client as though they were a representative independent sample without qualified methods review. Do not use a higher kappa to claim treatment effect, experimental control, generalization, maintenance, social validity, competence, or safety.
When the underlying rows contain protected health information, use approved systems, coded identifiers, and role-appropriate access. HHS maintains separate summaries of the federal Privacy Rule and Security Rule. Consent, retention, access, security, accessibility, payer, regulator, legal, and organizational requirements still need qualified review.
Related resources
- ABA IOA Method Sensitivity and Agreement Comparison Calculator for Clinicians
- How to Plan Interobserver Agreement Sampling in Clinical ABA
- How to Choose an IOA Calculation for ABA Data
- ABA Bland-Altman Paired-Measurement Agreement Calculator for Clinicians
Sources
- BACB Ethics Codes
- BCBA Test Content Outline, 6th edition
- Standards for Educational and Psychological Testing
- Cohen: A Coefficient of Agreement for Nominal Scales
- Feinstein and Cicchetti: High Agreement but Low Kappa, Part II
- Reed and Azulay: An Excel Tool for Calculating Interobserver Agreement
- Essig, Rotta, and Poling: Interobserver Agreement and Procedural Fidelity
- HHS summary of the HIPAA Privacy Rule
- HHS summary of the HIPAA Security Rule