ABA agreement and fidelity can be reported together because they answer different questions. Observer agreement asks whether independent observers recorded consistently. Treatment fidelity asks whether the procedure was implemented as defined. When a human observer scores fidelity, agreement can also be sampled on that fidelity record. Client outcome, experience, access, and treatment value still require separate evidence.

ABA agreement and fidelity

Use a four-column review: outcome, client experience, treatment fidelity, and observer agreement. For each, record the measure, denominator, period, sample coverage, result, and limitation. Align observation windows when interpreting the measures together and preserve missing samples instead of carrying one score into another field.

Putting the measures in one review can make gaps easier to see. It should not collapse them into a composite score. A practice could have reliable observation, weak implementation, improving outcomes, and poor client experience at the same time. Each signal deserves its own definition and decision.

Keep four questions separate

SignalQuestionExample denominatorObserver agreementDid independent observers record similarly?matched opportunities or componentsTreatment fidelityDid observed implementation follow the active plan?correctly implemented due opportunities or componentsClient outcomeDid the measured client result change?eligible opportunities, time, or another goal-linked unitClient experienceHow did the person report fit, comfort, burden, or preference?completed accessible ratings or direct reports

The exact measure varies by case. The table helps families ask which question each number answers and prevents a high score in one column from filling an empty one elsewhere.

Check what the result can support

Strong agreement with low fidelity suggests observers consistently saw implementation depart from the plan. High fidelity with weak agreement suggests the fidelity record itself may be uncertain. Strong scores for both show consistent measurement of planned implementation during sampled periods, while the outcome remains a separate clinical question.

Read common result patterns carefully

  • High agreement and high fidelity: observers consistently scored implementation as matching the plan in the sample. Review outcomes, experience, adverse effects, and coverage before judging value.
  • High agreement and low fidelity: observers consistently identified implementation gaps. Review components, training, resources, plan feasibility, and setting conditions.
  • Low agreement and apparently high fidelity: confidence in the fidelity score is limited. Repair definitions, observation conditions, or observer preparation before relying on the number.
  • Low agreement and low fidelity: both implementation and measurement may need repair. Avoid deciding that staff or the client caused the pattern from the summary alone.

These interpretations remain proportional to the sample. One paired observation cannot represent every staff member or setting.

Align the periods when combining a review

Agreement from March, fidelity from April, and outcome data from May may each be valid. They do not describe one shared phase without a reasoned link. Display dates, plan versions, settings, staff roles, and observation coverage for each measure.

If agreement was checked on a small subset of fidelity observations, say so. For example, ten fidelity observations may include three independently double-scored samples. Report fidelity coverage as ten observations and agreement coverage as three. Avoid implying that every fidelity entry received an agreement check.

Use current clinical and measurement sources

The CASP public summary places assessment, planning, implementation, and evaluation within its autism-treatment scope. The BACB Ethics Code addresses competence, client involvement, consent and assent when applicable, documentation, and data-based evaluation for covered behavior analysts.

The BCBA Test Content Outline includes measurement, data integrity, observer agreement, and visual analysis as examination content. It does not set one calculation, sampling percentage, or action threshold for every case.

Choose the calculation for the measure

Vollmer and colleagues discuss practical consequences of data reliability and treatment-integrity monitoring. Reed and Azulay describe several agreement calculations. Method choice depends on the measurement system and the question being asked.

Keep communication and conditions visible

The ASHA AAC portal supports continuous communication-tool access. Observer agreement is meaningful only when the record also identifies communication access, ordinary supports, opportunity availability, and other conditions that affect what could be observed.

Client experience can be direct data without becoming a performance target. Preserve the person's communication form and exact context. A request to stop may show that staff correctly followed the stop rule, while the procedure itself still needs review.

A practical example

Two observers agree on 9 of 10 scored implementation opportunities, or 90% agreement. Staff complete the full critical sequence in 7 of those 10, or 70% opportunity-level fidelity. The client independently uses the target response in 4 of 10 eligible opportunities, and reports through an accessible rating that one prompt feels uncomfortable in three sessions.

The team does not average 90%, 70%, and 40%. It reviews the single observer disagreement, the three implementation gaps, the outcome pattern, and the discomfort report. The clinician removes the uncomfortable prompt after client review, issues a new version, and staff rehearse it. Later fidelity, agreement, outcome, and experience data begin a new labeled phase.

This four-signal review identifies both a measurement question and a plan-fit concern. Neither would be clear from one blended quality score.

A family-facing reporting format

For each signal, ask for the name, definition, dates, raw numerator and denominator, coverage, result, limitation, responsible interpreter, and next action. A one-page matrix can provide clarity while the source records remain available through the applicable access process.

Avoid averaging unlike percentages

Observer agreement of 90%, treatment fidelity of 70%, and outcome performance of 40% use different numerators and denominators. Their average, 66.7%, has no clear clinical meaning. Keep each value in its own column and explain the question it answers.

The same caution applies across settings and periods. Agreement from one staff member's center visit cannot validate fidelity data from every home session. A client outcome collected across a month should not be presented as if it came only from the three paired observations. Align the evidence that truly overlaps and label the rest.

Turn the matrix into decisions

Each concern needs an owner and next step. Weak agreement may require a measurement repair. Low fidelity may require coaching, resources, or plan-feasibility review. Weak or unwanted outcomes may require clinical reassessment. Client discomfort may require an immediate response and an accessible plan review.

One person may coordinate the meeting, while qualified roles retain their own decisions. Operations should not revise clinical content, and a payer's coverage decision does not author the treating clinician's recommendation. The written review can show these boundaries and the evidence each owner considered.

Keep missing information visible

Use “not measured,” “not due,” “unscorable,” or another defined state instead of zero when a signal has no valid denominator. A blank client-experience column is not proof that the client was comfortable. A missing agreement sample is not proof that observers would have disagreed. Assign missing work an owner, due date, and reason.

Families can ask for the oldest unresolved item and the event that closes it. This prevents a polished dashboard from hiding the quality question that still blocks a decision.

Questions families can use

Ask which record agreement evaluates, which record fidelity evaluates, whether their windows match, how much care was sampled, what the client experienced, and which outcome supports the next decision.

Related resources

Sources

Finni resources

Ready for the next step?

Find ABA care near you