Two observers ABA data checks can be useful when a definition is new or difficult, scoring affects a high-consequence decision, observers disagree, staff or settings change, or data quality becomes uncertain. The responsible clinician should choose a sampling plan that fits the measure and risk. Double-scoring every session is rarely the only option; the plan should state when, where, and why it occurs.

Two observers ABA data

Predeclare the target, observer qualifications, number and distribution of observations, settings, staff, calculation, acceptable evidence, and response to disagreement. Sample relevant conditions instead of choosing only easy sessions. Add consent, assent, privacy, recording, and client-burden controls when a second observer or device changes the visit.

Double-scoring is a quality check with a purpose. It can test a new definition, monitor drift, investigate an unexpected result, or strengthen evidence before a consequential decision. Name that purpose so the team can choose the smallest useful sample and interpret it fairly.

Situations that often justify a second observer

Consider a paired observation when:

  • a new measure has not yet been applied consistently
  • the event is brief, complex, or difficult to distinguish from similar actions
  • observer agreement has weakened or scoring drift is suspected
  • new staff, settings, materials, or communication methods change observation conditions
  • the graph shifts in a way that may reflect measurement rather than client change
  • a mastery, safety, discharge, treatment-change, or other high-consequence decision depends heavily on direct observation
  • a family or client raises a specific, credible data-quality concern

The responsible clinician can also decide that another method provides better evidence. Two observers are one option, not a universal requirement for every session.

Check what the result can support

A routine sample can monitor drift. A triggered sample can investigate an unusual graph, new staff, a changed definition, or a safety concern. Keep those purposes separate in the record. A triggered observation should not silently replace the routine sampling denominator.

Build a representative sampling plan

Distribute paired observations across the conditions that matter to the decision. That may include different staff, days, settings, activities, levels of support, and times. A sample made entirely of scheduled center visits may not address a concern about community implementation.

State the planned number and the due period. Report completed paired observations divided by those due. If six were planned and five occurred, coverage is 5 of 6. Calculate agreement only within valid matched samples and keep the missing observation visible.

Triggered checks need their own labels. If an extra paired observation follows a complaint, report it as an investigation sample. It should not make the routine monitoring plan appear complete.

Limit disruption and client burden

A second person, camera, or remote connection can change privacy, noise, attention, comfort, and ordinary interaction. Ask whether the observer can use an existing visit, view only the necessary period, or score a properly authorized recording. Explain the observer's role to the client in an accessible way and preserve the person's communication system.

Follow the applicable consent and assent process. If the client asks the observer to leave or the recording to stop and that right applies, honor the response and label the sample incomplete. Do not arrange a dangerous, distressing, or unwanted event merely to obtain a second score.

Use current clinical and measurement sources

The CASP public summary places assessment, planning, implementation, and evaluation within its autism-treatment scope. The BACB Ethics Code addresses competence, client involvement, consent and assent when applicable, documentation, and data-based evaluation for covered behavior analysts.

The BCBA Test Content Outline includes measurement, data integrity, observer agreement, and visual analysis as examination content. It does not set one calculation, sampling percentage, or action threshold for every case.

Choose the calculation for the measure

Vollmer and colleagues discuss practical consequences of data reliability and treatment-integrity monitoring. Reed and Azulay describe several agreement calculations. Method choice depends on the measurement system and the question being asked.

Keep communication and conditions visible

The ASHA AAC portal supports continuous communication-tool access. Observer agreement is meaningful only when the record also identifies communication access, ordinary supports, opportunity availability, and other conditions that affect what could be observed.

Observers should be trained to recognize the person's relevant communication forms. If one person watches the client and another watches only the device, the records may reflect different information. Define the observable response and make both observers' access to that event comparable.

A practical example

A team changes the definition of an independent break request and schedules six double-scored observations across three staff members and two settings. Five occur, so paired-observation coverage is 5 of 6, or 83.3%. The missed community observation remains open with an owner and due date.

Within the five valid samples, observers independently score 30 eligible opportunities. They agree on 27, or 90%. All three disagreements involve whether a partially completed AAC message reached the defined endpoint. The clinician clarifies the endpoint with the client, observers recalibrate on fictional examples, and the team completes a new community sample.

The report keeps coverage, agreement, disagreement type, and the definition change separate. It does not present 90% as proof that the break-request measure is accurate, representative of all care, or clinically valuable.

Decide what happens after the check

Predefine a response to strong and weak evidence. Strong agreement may support continued use of the measure with ordinary monitoring. Weak or patterned agreement may trigger definition revision, observer training, equipment repair, another method, or a hold on the decision. Record who owns the repair and what evidence closes it.

There is no universal sampling percentage

The right amount of double-scoring depends on the measurement, variability, consequence of error, observer experience, settings, and available alternatives. A fixed percentage copied across every case can undersample a complex safety measure and oversample a stable low-burden measure. The plan should explain its risk-based choice.

Sampling can change over time. A new definition may need more frequent paired observations. Strong, stable results across representative conditions may support a lower monitoring frequency. A staff transition, data anomaly, client concern, or clinical change may increase it again. Preserve each version of the sampling plan and the reason for the change.

Families can ask when the next routine check is due and which events trigger an extra check. That answer is more informative than a promise that agreement is checked “regularly.”

The schedule should also identify who can approve a change, how missed samples are aged, and when a weak result holds the dependent decision. This turns a monitoring percentage into an accountable quality process.

Questions families can use

Ask which decision justifies a second observer, how sessions were selected, whether the observer changes the setting, what privacy applies, which calculation will be used, and what happens if the result is weak.

Related resources

Sources

Finni resources

Ready for the next step?

Find ABA care near you