An ABA interobserver agreement data sheet should preserve two independent records before anyone discusses discrepancies. It should also name the measurement definition, eligible comparison unit, agreement method, paired sampling frame, and missing observations. Without that trail, a percentage cannot show whether observers agreed on occurrences, agreed mostly on nonoccurrences, or even scored the same opportunities.

Clinicians & ABA Professionals / Data, Outcomes and Clinical Decision-Making.

This worksheet organizes an agreement check; it does not choose a clinical measure or certify an observer's accuracy. A qualified clinician remains responsible for the measurement system, sampling plan, interpretation, training response, and case-specific decisions.

Freeze the comparison contract before observing

The current BCBA Test Content Outline, Sixth Edition includes operational definitions, occurrence and temporal measures, discontinuous measurement, representative sampling, and evaluation of measurement validity and reliability. The BACB's current test-content outline page identifies the outlines in use. These are certification documents, not a client protocol or an endorsement of this form.

Write the comparison rule before a paired observation starts. The rule should explain what the observers will score, when an observation begins and ends, which record is primary, which units are eligible, how timing tolerances work, and how unavailable or technically incomplete units will be handled. Choosing the most flattering algorithm after seeing the records changes the question.

Paired-observation contractEntryClient or authorized case codeSocially meaningful questionObservable response and definition versionMeasurement procedure and unitObservation window and eligible opportunitiesPrimary and secondary observer rolesIndependence and blinding arrangementsPairing rule for events, trials, intervals, or durationsTiming tolerance, if anySelected agreement algorithm and rationaleMissing, unavailable, invalid, and unpaired codesOrdinary communication and access supportsPrivacy, recording, consent, and retention controlsReview owner and escalation route

If definitions or equipment change, open a dated version. Do not combine paired data collected under materially different boundaries merely because the column headings look alike.

Preserve the paired raw record

Observers should enter their records separately. Freeze both versions before reconciliation so the clinical team can distinguish the original measurement evidence from later consensus or correction.

UnitContext and eligible opportunityObserver AObserver BPaired?Exact agreement?Disagreement typeTiming or access note1Yes / noYes / no2Yes / noYes / no3Yes / noYes / no4Yes / noYes / no5Yes / noYes / no6Yes / noYes / no

The article on continuous recording and IOA algorithms documents that different algorithms answer different questions and can behave differently across event patterns. A broader inter-rater reliability tutorial likewise emphasizes matching the statistic to the design and reporting enough detail for interpretation. Neither paper supplies a universal clinical threshold.

Name the method and denominator

For a binary trial-by-trial record, build the two-by-two table before calculating a percentage.

Observer A / Observer BB occurrenceB nonoccurrenceRow totalA occurrenceA nonoccurrenceColumn total

The following simple views can be useful when they were selected prospectively and fit the measurement system:

  • Total exact agreement = occurrence and nonoccurrence matches divided by paired eligible units.
  • Occurrence agreement = twice the occurrence matches divided by twice the occurrence matches plus both kinds of disagreement.
  • Nonoccurrence agreement = twice the nonoccurrence matches divided by twice the nonoccurrence matches plus both kinds of disagreement.

These percentages are not interchangeable. Total agreement can look strong when nonoccurrence dominates. Occurrence and nonoccurrence agreement can expose that imbalance, but neither proves which observer was correct. A chance-corrected or continuous-measure statistic may be more appropriate for another design; this sheet should link to the actual method rather than silently substitute one.

Review paired-sample coverage

Agreement evidence also depends on where and when the paired observations occurred.

Sampling dimensionPlanned paired unitsCompleted paired unitsMissing or invalidConditions representedConditions absentPeople or implementersSettingsTimes or activitiesLow- and high-rate periodsProcedure versions

A high score from one convenient session should not be described as representative of unobserved staff, routines, settings, rates, or procedure versions. Keep the sampling result beside the agreement result.

Worked fictional example

The example is fictional. The team defines one eligible opportunity for a break request and freezes 20 paired opportunities. Observer A and Observer B record independently. Their two-by-two table contains seven occurrence agreements, nine nonoccurrence agreements, three A-occurrence/B-nonoccurrence disagreements, and one A-nonoccurrence/B-occurrence disagreement.

Observer A / Observer BB occurrenceB nonoccurrenceRow totalA occurrence7310A nonoccurrence1910Column total81220

The ABA interobserver agreement data sheet retains three calculations:

  • Total exact agreement: (7 + 9) / 20 = 16 / 20 = 80.0%.
  • Occurrence agreement: (2 x 7) / ((2 x 7) + 3 + 1) = 14 / 18 = 77.8%.
  • Nonoccurrence agreement: (2 x 9) / ((2 x 9) + 3 + 1) = 18 / 22 = 81.8%.

The record does not call 80.0% accurate. It shows four disagreement units that need review and a modest paired sample tied to one definition and context. If a later definition clarification changes how a unit would be scored, preserve the original result and create a new version rather than rewriting history.

Separate agreement, accuracy, and integrity

The calibration analysis of observational rate measurement explains that two observers may agree without either matching an accepted reference. Research applying signal-detection concepts to observer records also separates agreement from accuracy and bias. Agreement is evidence about correspondence between records, not proof that a definition measures the intended construct.

Treatment integrity answers a different question about whether defined procedure components occurred. A telehealth functional-analysis study reports separate agreement methods for interval and trial records within its own design, illustrating why the measurement unit and formula must travel together. A clinical team should not transfer that study's sampling percentage or decision criteria into a different service by default.

Record disagreements without blame

Use a correction log to classify what the team learned while keeping the original paired data intact.

UnitOriginal A/B codesDefinition or boundary issueTiming issueAccess or technology issueFeedback or clarificationNew version needed?Owner and follow-upYes / no

Disagreement may reveal ambiguous examples, an impractical timing rule, fatigue, device latency, an access barrier, or an ordinary scoring mistake. Match the response to the documented pattern. Do not turn a percentage into a personnel label or conceal conditions that made observation difficult.

Protect the record and the people represented

The BACB's current ethics-code hub says readers should check back for current versions. The linked Ethics Code for Behavior Analysts applies to BCBA and BCaBA certificants and applicants; it is not organizational certification. The CASP Version 3 access page describes autism-specific guidelines and licensing conditions, so this worksheet does not reproduce or attribute licensed guideline content.

Keep identifiers minimal, follow applicable consent and recording requirements, limit access, and preserve retention and correction rules. The current HHS Security Rule risk-analysis guidance describes risk analysis for regulated entities. It neither certifies this worksheet nor determines whether a particular record is protected health information.

Questions for the review meeting

  • Did both observers use the same frozen definition, procedure version, unit, and observation window?
  • Were the records independent until the comparison was frozen?
  • Which eligible units were missing, invalid, or unpaired, and why?
  • Why was this agreement method selected before the records were scored?
  • Do occurrence and nonoccurrence patterns tell a different story from total agreement?
  • Which people, settings, rates, and procedure versions were not represented?
  • Does a disagreement suggest definition repair, practice, equipment repair, or a new observation?
  • What client communication, privacy, access, safety, payer, licensure, or local rules apply?

Related resources

Sources