An ABA stimulus equivalence worksheet needs to preserve the relation logic behind every score. A single “equivalence accuracy” percentage can hide which relations were taught, which reversals or combinations were never taught, whether probe trials remained independent, and whether one weak relation was averaged away. This tool keeps the training map, trial record, relation-specific probes, access conditions, and evidence limits visible.
The worksheet is educational and case-specific. The tool cannot decide whether matching-to-sample or equivalence-based instruction is appropriate, select stimuli, or define a passing criterion. A qualified clinician must connect any use to an authorized plan and the client's meaningful goals and response preferences. Interdisciplinary input matters when language, literacy, vision, hearing, motor access, or AAC is involved.
Clinicians & ABA Professionals / Assessment and Treatment Planning.
Clinical boundary: Preserve communication access and the client's right to pause, refuse, or change course. Stop or hold when comparison stimuli cannot be discriminated, the response mode is unavailable, or position or cueing errors occur. The same applies when a nominal probe relation has already been taught, distress or another unwanted effect appears, or the responsible clinician cannot interpret the result. This record cannot establish comprehension, language ability, consent, medical safety, treatment necessity, payer coverage, equivalence-class formation, generalization, maintenance, or benefit. The BACB ethics resources are one starting point for the qualified review that remains necessary.
Read the relation notation before the percentages
In a three-member class, A, B, and C identify stimulus sets or modalities; subscripts identify the proposed class. For example, A1, B1, and C1 may be the spoken name, printed name, and photograph associated with one community destination. “AB training” means A is presented in the sample role and B in the comparison role under the defined matching-to-sample arrangement. That notation leaves BA untrained unless the exposure record says otherwise.
Stimulus equivalence is commonly discussed in terms of reflexivity, symmetry, and transitivity, with equivalence tests combining relevant derived relations. Applied research with autistic children and a broader review of derived-stimulus-relation technology show varied training structures, class sizes, modalities, comparison arrays, and test sequences. The exact test logic therefore belongs in the protocol. Identity matching should not be relabeled reflexivity merely because identical stimuli appeared. An untaught relation should not be called emergent if staff modeled or corrected it earlier.
The safest summary begins with a relation map. A completed ABA stimulus equivalence worksheet should keep these statuses visible:
- Directly trained: the exact sample-to-comparison relation received programmed instruction.
- Baseline or prerequisite: performance was assessed for a separate readiness question.
- Untrained probe: the relation was intentionally withheld and tested under a frozen rule.
- Exposed: the relation was modeled, corrected, rehearsed, or otherwise encountered before a planned probe.
- Not evaluated: the planned test was absent, invalid, inaccessible, or outside scope.
Purpose, class membership, and access contract
Choose class members because the relation matters to the client, not because the stimuli happen to fit a tidy laboratory diagram. Record modality, salience, language, familiarity, cultural meaning, and possible confounds. If color, position, typography, voice, device layout, or staff behavior could cue a response, treat that feature as part of the arrangement.
ASHA describes AAC as an individualized, multimodal system. Keep the client's established communication system available. A selection by touch, switch, eye gaze, partner-assisted scanning, sign, speech, or another accepted mode may be valid when its observation and interpretation rules are defined in advance.
Blank scope and access card
FieldCase-specific entryClient-selected purpose and meaningful useProposed class count and class sizeA-set role and modalityB-set role and modalityC-set role and modalityAccepted response forms and response windowOrdinary AAC, language, motor, sensory, visual, hearing, and processing supportsComparison-array size, layout, and position-balancing ruleKnown stimulus-preference, salience, or rejection concernsSample-attending and comparison-scanning evidenceAssent, dissent, pause, stop, and re-entry signalsQualified owner, protocol version, effective date, and review date
Class inventory and stimulus-control checks
Use durable identifiers and retain the exact file, object, spoken form, or presentation specification. “Picture of library” is not enough when two teams use different images. Keep personal information out of the worksheet unless the approved record system and clinical purpose require it.
Blank class-member inventory
Proposed classA memberB memberC memberModality and exact asset versionMeaning/access checkPotential irrelevant cueNotes1A1B1C12A2B2C23A3B3C3
Before interpreting class-consistent performance, ask whether the comparisons are distinguishable and whether position responding is balanced. Also ask whether the sample controls observation and whether a compound sample contains elements that compete for control. Research with complex auditory-visual samples illustrates why independent control by each element cannot simply be assumed.
Freeze the training and test structure
Draw the relation network before instruction. Published work on building emergent-responding training systems illustrates how explicitly mapping relations supports implementation. The example below is only notation. It is not a recommendation to use sample-as-node, comparison-as-node, linear-series, or any other structure.
Blank relation map
RelationSample roleComparison rolePlanned statusDirect teaching allowed?Prompt/correction ruleConsequence ruleExposure before probeVersionABABtrain / baseline / probe / excludeBCBCtrain / baseline / probe / excludeBABAtrain / baseline / probe / excludeCBCBtrain / baseline / probe / excludeACACtrain / baseline / probe / excludeCACAtrain / baseline / probe / exclude
If the team changes the training structure after seeing probe results, create a new phase. Do not keep the old “untrained” label after teaching begins. The question may remain clinically useful, but it has changed.
Matching-to-sample trial ledger
Record the exact sample, full comparison array, positions, initial selection, latency if relevant, prompt or correction, and validity. Freeze the initial response before feedback. A replayed spoken sample or repeated trial may be appropriate under a written rule, but it must not silently replace the original result.
Blank training and probe record
Date/timeRelation/classtrial typeSampleComparison array and positionsAccess available?Initial responseIndependent correct?Prompt/correctionConsequenceValid? reasonPrior exposure or contaminationAssent/dissent or unwanted effectObservertrain / probe / prerequisiteyes / no / not scorabletrain / probe / prerequisiteyes / no / not scorabletrain / probe / prerequisiteyes / no / not scorable
Invalid trials remain outside the performance denominator. Examples include a missing AAC device, an incorrect comparison array, a position-balancing error, or staff pointing or gaze cues. Other examples include an interrupted opportunity, an ambiguous response, a data-timing failure, or evidence that the supposedly untrained relation was already taught. A valid incorrect response is different: it belongs in the relation-specific denominator.
Relation-specific summary comes first
Aggregate data can be useful for workload and coverage, but the relation cell is the interpretive unit. A result of 15 correct responses across 20 valid probes does not reveal whether every relation is similar or one relation is substantially weaker.
Blank relation summary
RelationOriginal statusPlannedPresentedValidIndependent correctIncorrect/no responsePrompted or correctedExposure statusRelation-specific percentageInterpretation limitBAclean / exposed / uncertainCBclean / exposed / uncertainACclean / exposed / uncertainCAclean / exposed / uncertain
Use these distinct calculations:
- valid training coverage = valid training trials ÷ planned training trials;
- independent training performance = independently correct training trials ÷ valid training trials;
- valid probe coverage = valid emergent-relation probes ÷ planned emergent-relation probes;
- relation-specific performance = independently correct trials for one relation ÷ valid trials for that relation; and
- aggregate probe performance = independently correct probes across included relations ÷ valid probes across those same relations.
Neither aggregate accuracy nor one successful relation is, by itself, a declaration that a complete equivalence class formed. The protocol must state which relations, classes, repeats, and criteria are necessary. It must also say how invalid or exposed trials affect the question. Early research on functional and equivalence classes is one reminder that measured relations and possible naming effects need explicit interpretation.
Fictional community-destination example
The fictional plan uses three proposed classes. A members are spoken destination names, B members are printed destination names, and C members are photographs. The client-selected purpose is recognizing familiar destinations across ordinary community materials. The example assumes that the accessible response method and stimulus set have already been reviewed; it does not imply that spoken words, printed text, photographs, or a three-class structure fit another client.
AB and BC were designated for direct training. BA, CB, AC, and CA were withheld for separate probes.
Training summary
Eighteen training trials were planned. Seventeen were presented; one was not available because the scheduled setting closed early. One presented trial was invalid because the comparison array did not match the frozen version. Sixteen valid training trials remained, and 14 initial responses were independently correct.
- 16 / 18 × 100 = 88.9% valid training coverage
- 14 / 16 × 100 = 87.5% independent correct responding in valid training trials
The second percentage describes performance within valid training trials. It does not make the two unobserved or invalid planned trials disappear.
Emergent-relation probe summary
Twenty-four probes were planned: six for each of BA, CB, AC, and CA. Twenty-two were presented, 20 were valid, and 15 initial responses were independently correct.
- 20 / 24 × 100 = 83.3% valid probe coverage
- 15 / 20 × 100 = 75.0% aggregate correct responding among valid probes
The relation cells show why the aggregate is insufficient:
RelationPlannedPresentedValidIndependent correctRelation-specific resultBA66544 / 5 × 100 = 80.0%CB66544 / 5 × 100 = 80.0%AC65544 / 5 × 100 = 80.0%CA65533 / 5 × 100 = 60.0%
An aggregate of 75.0% would conceal the weaker CA cell. The record supports a narrow statement about observed responses under this protocol and leaves class formation, alternative stimulus control, replication, generalization, maintenance, and meaningful use for qualified review.
Procedural-fidelity review
Define only observable components that matter to the relation question. A six-component example might include correct sample, correct comparison array, balanced positions, ordinary access present, prompt/consequence rule, and accurate initial-response recording.
EventCorrect sampleCorrect arrayPositions followedAccess presentPrompt/consequence ruleInitial record accurateMismatch and impact
Twenty observed events in the fictional review create 120 component opportunities. One hundred twelve components matched and eight did not. Research on controlling relations within equivalence classes reinforces why implementation details and possible alternative control deserve review:
112 / 120 × 100 = 93.3% component implementation.
That value is not equivalence accuracy and cannot repair a contaminated probe. Review the eight mismatches by relation, class, and timing. Position or array errors concentrated in one relation may offer an alternative account of the relation-specific pattern.
Interpretation and version-control record
Use plain language that separates observation from inference. “Fifteen of 20 valid probes were independently correct, with CA at three of five” is reproducible. “The client understands the concepts” is not established by this worksheet.
Review itemEntryDirectly trained relations and performanceUntaught relations actually testedResult by relation and proposed classMissing, invalid, exposed, or contaminated testsPotential position, modality, salience, naming, or compound-control explanationAccess conditions and accepted response formsAssent, dissent, burden, and unwanted effectsFidelity mismatches that limit interpretationWhat remains unknown about emergence, use, generalization, and maintenanceHold, revise, continue, or discontinue questionQualified decision owner and next review
Start a new version when class members, modalities, sample/comparison roles, array size, positions, accepted response, or access supports change. Do the same when the training structure, prompt or consequence rule, test sequence, or scoring logic changes. Retain the old network with its trials. Merging phases can make a relation appear untrained even though it was taught in an earlier version.
Related resources
Sources
- BACB ethics codes and related requirements
- Ethics Code for Behavior Analysts
- BACB test content outlines
- BCBA Test Content Outline, Sixth Edition
- Council of Autism Service Providers ABA Practice Guidelines access page
- The Development of Functional and Equivalence Classes in High-Functioning Autistic Children: The Role of Naming
- Using Complex Auditory-Visual Samples to Produce Emergent Relations in Children With Autism
- Toward a Technology of Derived Stimulus Relations
- Controlling Relations in Stimulus Equivalence Classes of Preschool Children and Individuals With Down Syndrome
- Developing and Implementing Emergent Responding Training Systems With Available and Low-Cost Computer-Based Learning Tools
- Evidence From Children With Autism That Derived Relational Responding Is a Generalized Operant
- ASHA Augmentative and Alternative Communication Practice Portal
These sources describe different participants, stimuli, modalities, training structures, and research questions. They support explicit relation notation and direct tests of untrained responding. They do not make the fictional structure, any percentage, or an equivalence-based procedure appropriate for an individual client.