High ABA observer agreement shows that two observers recorded similarly in the sampled conditions. Both observers can still share the same misunderstanding, miss events outside the camera view, use a weak definition, or sample an unrepresentative period. Accuracy requires comparison with a valid reference when one exists. Clinical usefulness also depends on measurement validity, context, client experience, and the decision being made.

High ABA observer agreement

Review agreement beside definition quality, observer training, sample coverage, missing data, treatment fidelity, client outcome, and client feedback. Ask whether an expert or verified record provided an accuracy reference. When no reference exists, describe the result as agreement and keep the uncertainty visible.

The distinction matters because two people can make the same error. Agreement is evidence about consistency between records. Accuracy asks whether a record matches a trustworthy reference. Validity asks whether the measure captures the concept needed for the clinical question. Reliability, accuracy, and validity can overlap while remaining separate claims.

See how shared errors can produce high agreement

High agreement can occur when both observers:

  • learned the same vague or incorrect definition
  • missed activity outside their shared camera view
  • used the same clock with an incorrect interval setup
  • relied on the same incomplete opportunity list
  • interpreted a caregiver or client response in the same unsupported way
  • scored only an easy or unrepresentative part of care

A polished percentage cannot repair those conditions. Review the source of the observations and the definition before treating the result as dependable evidence.

Match the strength of the claim to the evidence

Different evidence supports different wording:

  • “The observers agreed” is appropriate when independent matched records were scored consistently.
  • “The observers applied the definition consistently in this sample” adds the definition and sampling boundary.
  • “The records were accurate” requires a valid reference that the observed records matched.
  • “The measure represents the target well” requires evidence that the definition and measurement system fit the intended concept.
  • “The plan worked” requires outcome and experience evidence beyond observer agreement.

Families can ask the provider to restate a conclusion at the level the evidence supports. This improves clarity without dismissing the value of a strong agreement check.

Check what the result can support

A strong percentage can be compatible with a poor measure. For example, two observers may agree every time that a vague category occurred, while another reader cannot reproduce the category from its definition. Fixing the definition may change both the data and the later agreement result.

Use a reference when a valid one exists

Some measures have a checkable reference. A verified system timestamp may help assess recorded latency. A frame-by-frame review may clarify a brief event when the video captures the full context and its use is authorized. A controlled fictional record can test whether observers apply a coding rule correctly. The reference itself needs to fit the question and remain free of the same error.

Many clinical events have no perfect gold standard. In those cases, use trained independent observers, clear definitions, representative sampling, agreement checks, and transparent limitations. Call the result agreement rather than accuracy unless the comparison truly supports an accuracy claim.

Ask whether the measure is useful

Observers may record a response consistently while the measure fails to address the family's question. Counting completed worksheets may be reliable, yet it may say little about independent understanding, client preference, or use outside the session. The clinical team should connect the measure to the goal, context, and decision.

Ask what would change if the number rose or fell. If the answer is unclear, the team may be collecting a precise measure with little decision value.

Use current clinical and measurement sources

The CASP public summary places assessment, planning, implementation, and evaluation within its autism-treatment scope. The BACB Ethics Code addresses competence, client involvement, consent and assent when applicable, documentation, and data-based evaluation for covered behavior analysts.

The BCBA Test Content Outline includes measurement, data integrity, observer agreement, and visual analysis as examination content. It does not set one calculation, sampling percentage, or action threshold for every case.

Choose the calculation for the measure

Vollmer and colleagues discuss practical consequences of data reliability and treatment-integrity monitoring. Reed and Azulay describe several agreement calculations. Method choice depends on the measurement system and the question being asked.

Keep communication and conditions visible

The ASHA AAC portal supports continuous communication-tool access. Observer agreement is meaningful only when the record also identifies communication access, ordinary supports, opportunity availability, and other conditions that affect what could be observed.

Accuracy questions become especially important when a client uses multiple communication forms. Two observers may agree that no request occurred because both were looking only for speech. A complete operational definition and accessible observation method should include the person's reliable AAC, gestures, signs, writing, or other forms relevant to the measure.

A practical example

Two observers agree on 19 of 20 intervals, or 95%. A permitted video review captures the relevant screen and partner interaction. It shows that both observers counted device navigation as a help request even though the active definition required a message directed to another person.

The 95% remains a correct statement about agreement under the shared scoring error. It does not establish accurate help-request data. The team marks the affected period, clarifies the definition with the client and clinician, and calibrates the observers on new fictional clips. A fresh independent sample produces 16 agreements in 18 intervals. Outcome and client-feedback measures remain separate.

The report explains why the original result was insufficient instead of replacing 95% with the later value. Families can now see the shared error, the repair, and the limits of both samples.

A quick family review checklist

Ask whether the definition is observable, the entire event could be seen or heard, both observers used the same window, the sample represents relevant care, missing data stayed visible, and a valid reference was used when available. Then ask whether the measure addresses a decision the client and family actually care about.

Watch for selective samples

A provider may choose clear recordings or sessions because they are easier to score. That can be reasonable for early calibration. It provides weaker evidence about routine measurement when crowded, noisy, mobile, or less predictable conditions are absent. Ask how samples were selected and whether excluded or unscorable periods were reported.

If twelve paired observations were planned and eight were completed, show 8 of 12 coverage. Agreement among the eight should not make the four missing samples disappear. A later sample can fill the gap while keeping the original period intact.

Questions families can use

Ask what serves as the accuracy reference, whether both observers could share bias, how representative the sample was, whether the definition is reproducible, and what other evidence supports the decision.

Related resources

Sources

Finni resources

Ready for the next step?

Find ABA care near you