ABC event sequence agreement should separate whether independent observers matched the same anchor events from whether they assigned the same antecedent and consequence codes. Define timestamp tolerance, event-matching rules, code comparisons, and eligible observation coverage before scoring. Report matched anchors divided by primary anchors, then code agreement within matched anchors. Preserve unmatched events and disagreements for calibration.

Separate detection from classification

ABC event sequence agreement has at least two stages. Anchor agreement asks whether observers detected the same response events. Code agreement asks whether they assigned the same antecedent, consequence, timing, and window states to matched anchors. A single overall percentage can hide a serious anchor-detection problem behind high coding agreement on the few events both observers saw.

Define the response anchor, timestamp precision, matching tolerance, one-to-one matching method, and sequence-code categories before scoring. State whether onset or offset is matched and how response episodes, simultaneous events, and boundary events are handled.

Preserve independent records

Observers use the same codebook, observation period, clock source, and view of the setting while recording independently. Lock both records before reconciliation. One observer must not correct the other during collection or copy an event after noticing it on a shared interface.

Record observer identity, training status, device clock, start and stop times, visibility loss, interruptions, and access or health events. If observers have different views or audio, label that condition. Agreement under unequal access is hard to interpret. The discrepancy may come from sampling conditions even when the codebook is clear.

Match anchors with a fixed algorithm

Start from every primary-observer anchor. Find at most one secondary anchor within the tolerance according to the declared nearest-time or ordered-matching rule. Preserve primary-only and secondary-only events. Avoid allowing one secondary event to match several primary events.

Report primary anchors, secondary anchors, matched pairs, unmatched primary anchors, unmatched secondary anchors, and timing differences. If observers disagree about episode boundaries, record the raw onsets before adjudication. A wider tolerance may raise match coverage while linking genuinely different events.

Compare codes within matched pairs

Only after anchor matching should the review compare antecedent category, consequence category, sequence direction, lag bin, truncation state, and other relevant codes. Report each component separately and a combined exact-match value if useful.

An anchor pair can agree on antecedent and disagree on consequence. Another can share event labels but fall in different lag bins because of timestamp differences. Keep these patterns visible. Reconciliation creates an adjudicated record for later analysis but must not overwrite either independent source.

Verify Farah’s audit

The primary observer records 20 response anchors. The secondary observer matches 18 under the fixed tolerance. Anchor-match coverage is 18 / 20 = 90%. Two primary anchors remain unmatched. Any secondary-only anchors would be reported separately and cannot be derived from the supplied counts.

Among the 18 matched anchors, antecedent and consequence sequence codes agree on 15. Code agreement is 15 / 18 = 0.8333, or 83.3%. Three matched anchors have at least one sequence-code disagreement. The 83.3% denominator is matched pairs, while the 90% denominator is primary anchors.

Multiplying the two percentages to create one “overall agreement” would conceal the staged design. The audit should retain 18 of 20 and 15 of 18 as separate controls. It should also state whether 15 refers to exact combined sequence agreement or a specific component rule.

Handle missing, truncated, and overlapping events

If either observer loses visibility or timing precision, label the affected observation span. An anchor with a truncated consequence window may still match as a response event while its consequence code remains unavailable. Do not score unavailable as disagreement or agreement without a declared rule.

Overlapping response episodes can produce different anchor counts. Review whether the operational definition is observable at the needed precision. If one observer records onset and the other records offset, repair the collection protocol before calculating a tolerance-based score.

Zero primary anchors create an undefined anchor-coverage denominator. Zero matched pairs create an undefined code-agreement denominator. Report the lack of eligible events and examine coverage; 0% would imply valid pairs were available and disagreed.

Use disagreement as calibration evidence

Classify disagreements by missed detection, extra detection, timestamp drift, antecedent code, consequence code, overlap, truncation, missing visibility, or interface error. Review source clips or records under approved privacy controls. Update definitions or training only after preserving the original result.

Retest on new independent observations. A coached re-score of the same 20 anchors measures reconciliation, not prospective observer performance. Track whether the specific error type improves under the unchanged matching rule.

Report the audit clearly

A concise report could say: “Observers independently recorded the same period. Eighteen of 20 primary response anchors matched within the declared tolerance (90%); two were unmatched. Sequence codes agreed on 15 of 18 matched anchors (83.3%). Three code disagreements were retained for calibration. Anchor detection and code classification are reported separately.”

Use a flow display from 20 primary anchors to 18 matched pairs to 15 exact code matches. Add unmatched secondary anchors, observation coverage, and disagreement types. Avoid a graph containing only 83.3% because it omits the unmatched events.

Bound what agreement establishes

Agreement does not establish that event definitions are clinically valid, that observers sampled representative contexts, or that a descriptive relation shows function. Two observers can agree on the same flawed definition. Lag-sequential work underscores timing precision, while comparative descriptive-method research limits functional interpretation.

The BCBA Test Content Outline provides examination scope, and the BACB ethics hub and CASP summary provide professional context. Qualified clinicians decide whether the measurement remains useful.

Protect access and client participation

Review the event definitions and disputed observations with Farah through an accessible process. Keep AAC continuously available under ASHA guidance. Ask whether the codes recognize communication, pain, discomfort, assent, dissent, and meaningful context.

Observer agreement cannot justify delaying food, water, bathroom use, mobility, prescribed health care, rest, relationships, or emergency help. Stop observation for urgent medical or safety needs and preserve the resulting truncated state.

Clinician checklist and limitations

Before accepting the audit, confirm:

  • observers recorded independently with synchronized clocks and the same accessible view;
  • anchor definition, tolerance, one-to-one matching, episode, and window rules are fixed;
  • primary, secondary, matched, and unmatched event counts reconcile;
  • 18/20 = 90% and 15/18 = 83.3% reproduce;
  • component disagreements, missing spans, truncation, and overlaps remain visible;
  • adjudication preserves both original records and uses approved privacy controls; and
  • the review records accessible client input, calibration action, retest owner, and date.

Farah’s sample contains only 20 primary anchors. One changed match moves coverage by 5 points; one changed code moves agreement by 5.6 points. Tolerance choice, timestamp precision, visibility, event frequency, and category prevalence affect the results. Reopen the method after definition, device, observer, software, context, or access changes.

Related resources

Sources