ABA observer agreement compares records made independently by two trained observers who watched the same defined event or time period. Families can ask which behavior or implementation step was scored, which observation windows matched, what calculation was used, how many observations were sampled, and what disagreements occurred. Agreement supports confidence in consistency while leaving accuracy, validity, and treatment value as separate questions.
ABA observer agreement
Request the operational definition, observation dates, setting, observer roles, independence procedure, matched window, raw records when appropriate, calculation, result, exclusions, and clinical use. A percentage without the number of double-scored observations can make a small or selective sample appear broader than it was.
Families can request a plain-language explanation rather than a technical worksheet. A useful answer says what both people watched, how their records were matched, how much care was sampled, and what the team did with any disagreements. The result should stay tied to the specific behavior, implementation step, person, setting, and period observed.
Agreement can be calculated in several ways
The calculation should fit the measurement system. Common questions include:
- Did both observers record the same occurrence in each interval?
- Did their total counts match closely across the observation?
- Did they identify the same eligible opportunities and responses?
- Did they record similar durations or latencies?
- Did they score the same treatment components as correct, incorrect, or inapplicable?
Overall agreement, occurrence agreement, nonoccurrence agreement, exact count agreement, interval-by-interval agreement, and other methods can produce different values from the same records. Families do not need to select the formula. They can ask why the chosen method addresses the decision and whether another result would reveal a hidden pattern.
Check what the result can support
Independent means each observer records without copying or adjusting to the other's entries during the observation. Later comparison and feedback are appropriate. If observers discuss uncertain events before locking records, that exercise may be calibration, though it is weak evidence of independent agreement for that session.
Both observers should use the same operational definition, time source, observation window, and event-matching rule. If one person sees the first five minutes and another begins later, their total records are not a matched sample. If one observer counts a request when a device icon is touched and the other waits for a message directed to a partner, they are measuring different events.
Ask how much care was double-scored
Agreement from one short visit may show that two observers could apply the rule in that setting. That sample cannot describe measurement across every staff member, time, or environment. Ask how observations were selected and how many planned double-scored samples occurred.
Keep coverage and agreement separate. If five paired observations were due and four occurred, coverage is 4 of 5. If the observers agreed on 36 of 40 matched opportunities within those four, agreement is 36 of 40, or 90%. Both facts help the family understand the evidence.
Review the disagreements, not only the percentage
Disagreements can cluster around the start or end of an event, brief responses, crowded settings, similar-looking behaviors, off-camera activity, or a missing communication support. A disagreement table can identify the definition or observation condition that needs repair.
The practice should preserve the original records. After reviewing discrepancies, it can clarify the definition, calibrate observers, repair equipment, or collect a new independent sample. Editing the original entries until they match would erase the evidence that prompted the quality review.
Use current clinical and measurement sources
The CASP public summary places assessment, planning, implementation, and evaluation within its autism-treatment scope. The BACB Ethics Code addresses competence, client involvement, consent and assent when applicable, documentation, and data-based evaluation for covered behavior analysts.
The BCBA Test Content Outline includes measurement, data integrity, observer agreement, and visual analysis as examination content. The outline does not set one calculation, sampling percentage, or action threshold for every case.
Choose the calculation for the measure
Vollmer and colleagues discuss practical consequences of data reliability and treatment-integrity monitoring. Reed and Azulay describe several agreement calculations. Method choice depends on the measurement system and the question being asked.
Keep communication and conditions visible
The ASHA AAC portal supports continuous communication-tool access. Observer agreement is meaningful only when the record also identifies communication access, ordinary supports, opportunity availability, and other conditions that affect what could be observed.
Observers need to know which communication forms count and how the client indicates assent, dissent, discomfort, or a request for help. A device positioned outside one observer's view can create disagreement that reflects camera placement rather than unclear client communication. The record should identify that limitation.
A practical example
Two observers independently score ten communication opportunities using the same plan version and response window. They agree on eight and disagree on two, so opportunity-by-opportunity agreement is 8 of 10, or 80%.
Both disagreements occur when Nia begins a message on AAC and completes it after looking toward another person. One observer scores the device selection; the other scores the completed partner-directed message. The team preserves both records and learns that the definition does not identify the response endpoint clearly.
The clinician clarifies the endpoint after reviewing Nia's communication method, and the observers practice with fictional examples. They then independently score twelve new opportunities and agree on eleven. The practice reports the original 8 of 10, the definition change, and the later 11 of 12 as separate phases. The later result does not rewrite the original sample.
When a family should ask for another check
A new paired observation may be useful after a definition changes, observers drift, a new setting or staff role begins, a graph changes unexpectedly, a recording problem appears, or the data support a high-consequence decision. The qualified clinician should determine what evidence is enough for that decision while keeping client burden, consent, privacy, and access visible.
Ask what record the agreement check evaluated
Agreement can be calculated on client-response data, treatment-fidelity data, duration, opportunities, or another defined record. Name that record clearly. Agreement on a staff-implementation checklist says whether observers scored staff similarly. It says nothing directly about whether they agreed on the client's response unless that response was separately scored.
A family-facing summary should identify the target record, period, observers, matched sample, calculation, raw result, coverage, disagreements, limitations, and decision. If the team checked several measures, report each denominator separately. Combining a high agreement score from one measure with a weaker result from another produces an average that may answer no useful question.
Ask for the next check date and the event that would trigger an earlier review.
Questions families can use
Ask what both people watched, whether they recorded independently, how events were matched, which calculation fits the measure, how much care was sampled, and whether disagreements changed the definition or data.
Sources
- Council of Autism Service Providers, ABA Practice Guidelines Version 3.0 public summary
- Behavior Analyst Certification Board, Ethics Code for Behavior Analysts
- Behavior Analyst Certification Board, BCBA Test Content Outline, 6th edition
- Vollmer, Sloman, and St. Peter Pipkin, Practical Implications of Data Reliability and Treatment Integrity Monitoring
- Reed and Azulay, A Tool for Calculating Interobserver Agreement
- American Speech-Language-Hearing Association, Augmentative and Alternative Communication
Finni resources