An ABA measurement selection worksheet should keep the clinical question in charge of the tool. Response dimension, observation boundary, opportunity structure, client context, and likely error come first. Software fields and staff convenience come later. A feasibility review can show whether a method is usable. By itself, it cannot establish validity, representativeness, reliability, or clinical value.
Clinicians & ABA Professionals / Data, Outcomes and Clinical Decision-Making.
This copyable worksheet organizes a qualified review. It prescribes no system, target, observation plan, or interval. Selection remains a client-informed clinical decision. Account for access, supports, setting constraints, risk, burden, privacy, and error consequences.
Freeze the question before comparing methods
The BCBA Test Content Outline, Sixth Edition distinguishes direct, indirect, and product measures, as well as occurrence and temporal dimensions. It also addresses continuous and discontinuous procedures, validity, reliability, representative data, display, and interpretation. The BACB's test-content outline page identifies the current outline. This examination content neither selects nor approves a clinical method.
Describe one decision in plain language. “Measure task participation” is too broad. A team might need to know the latency to start after an accessible cue or the duration of engagement. Another question might concern independently completed steps, help requests, or an agreed work-product criterion. Each question points toward a different dimension and source of evidence.
Measurement-question contractWorking entryClient-selected outcome or practical purposeExact clinical or assessment questionObservable response and definition versionCritical response dimensionEligible opportunity or observation windowOrdinary supports, AAC, prompts, and materialsSetting, people, and conditions to representDecision the data may informConsequence of overestimating the responseConsequence of underestimating the responseClient and caregiver burden or preferenceQualified clinical owner and review date
If the data will not answer the stated question, do not improve the worksheet score. Rewrite the question or choose a different source of evidence.
Separate the response dimension from the display
Count, rate, duration, latency, interresponse time, and percentage of opportunities are different measures. Interval estimates, permanent products, rating scales, and indirect reports are not interchangeable with them. Each carries a unit, denominator, observation demand, and error pattern. An attractive graph can still obscure the wrong choice.
The paper On Terms: Frequency and Rate in Applied Behavior Analysis explains why count and time must remain explicit. It notes that a simple-looking rate can hide periods when behavior could not occur. The discussion does not select a measure for a particular person. It supports recording observation time and opportunity structure rather than a floating number.
Candidate measureDirectly answers which question?Unit and denominatorWhat can be missed or distorted?Observation demandKeep for pilot?Event count or rateYes / noDurationYes / noLatencyYes / noPercentage of eligible opportunitiesYes / noPartial-interval recordingYes / noWhole-interval recordingYes / noMomentary time samplingYes / noPermanent productYes / noStructured rating or indirect reportYes / no
A familiar measure may be rejected when its unit does not match the decision. Record the reason. A later reviewer should be able to see why the convenient option was not silently chosen.
Compare fit, error, and feasibility side by side
A peer-reviewed proposed model for selecting measurement procedures considers observability, critical dimensions, opportunity structure, personnel resources, and environmental constraints. Its studied context is problem behavior. The model supports structured questions. Its scope does not make one path suitable for every client, skill, setting, or use.
Use narrative evidence before any numeric ranking. A single weighted score can conceal a disqualifying problem. For example, the measure may miss the critical dimension or require observation the client has not agreed to.
Review lensCandidate ACandidate BCandidate CObservable under ordinary conditionsMatches the critical dimensionEligible opportunities can be identifiedUnit and denominator remain interpretableExpected direction and magnitude of errorCan represent hard and easy contextsObserver attention and competing dutiesTraining and calibration demandAAC, language, mobility, sensory, or health accessRecording, device, privacy, and security burdenGraph and report compatibilityReconsideration trigger
Keep a “not known” state. A wish to begin collecting data does not turn optimistic feasibility assumptions into facts.
Treat sampling as a method with known limitations
The practitioner discussion on discontinuous methods of data collection reviews how partial-interval, whole-interval, and momentary time-sampling procedures can produce different errors. It also considers interval duration and observer workload. A later study of discontinuous measurement in common practice examined correspondence with continuous records. Its results concern particular data and interval arrangements. Neither source supplies a universal interval length or authorizes a sampling estimate when continuous measurement is feasible and needed.
For each sampling candidate, state what the observer sees between checks. Define what an interval code means and whether likely error matters. Momentary sampling at 30-second points cannot recover an exact start latency. Continuous timing across a long classroom period may create burden or alter the activity. Fit depends on the question and context.
Design a small feasibility pilot
A pilot should test workflow and data properties before the team commits to a method. The pilot should not expose a client to unnecessary observation. It must not postpone needed care. Decide the contexts, opportunities, and materials in advance. Specify observers, devices, access supports, and stop conditions too.
Pilot planEntryCandidate measure and versionResponse definition versionPlanned settings and opportunity countClient and caregiver reviewObserver roles and competing dutiesEquipment or software testMissing, invalid, and no-opportunity codesIndependent observer checkFeasibility questionsPrivacy, consent, and recording controlsSafety and immediate stop conditionsQualified reviewer and decision date
The evidence-based practice discussion in applied behavior analysis describes integration of best available evidence and clinical expertise. It also includes client values and context. A pilot contributes local evidence; it does not replace the other parts of the decision.
Record valid, invalid, and unavailable observations
Build the pilot ledger before collecting the first value. Keep true zero, no opportunity, missing, not observed, and invalid distinct. A timing-device failure is not a zero-second latency. A cue never delivered does not create an infinite latency.
OpportunityPlanned contextEligible opportunity?Measure attemptedRaw value and unitData stateBurden or access noteUse in summary?1Yes / no / unknownValid / no opportunity / invalid / missing / not observedYes / no2Yes / no / unknownValid / no opportunity / invalid / missing / not observedYes / no3Yes / no / unknownValid / no opportunity / invalid / missing / not observedYes / no4Yes / no / unknownValid / no opportunity / invalid / missing / not observedYes / no5Yes / no / unknownValid / no opportunity / invalid / missing / not observedYes / no
Feasibility is not simply the percentage of completed fields. Ask whether missingness clusters in difficult settings. Check whether observation changes the interaction or a device distracts from care. Staff should be able to collect the measure without abandoning other duties. The person's experience of the process matters too.
A fictional 12-opportunity latency pilot
The following example uses synthetic values. It is not a client record, recommended observation schedule, mastery criterion, or treatment result.
A fictional team asks how long an observable work start takes after an accessible cue under ordinary conditions. It pilots latency measurement across 12 planned opportunities. Ten yield valid latencies. One has no eligible opportunity because the cue was not delivered. One is invalid because the timing device failed the predefined clock check.
OpportunityStateLatency in secondsNote1Valid42Definition and clock checks complete2Valid55Definition and clock checks complete3Valid61Definition and clock checks complete4Valid38Definition and clock checks complete5Valid47Definition and clock checks complete6No opportunityNot recordedAccessible cue was not delivered7Valid72Definition and clock checks complete8Valid50Definition and clock checks complete9InvalidNot recordedClock failed the pilot validation rule10Valid44Definition and clock checks complete11Valid59Definition and clock checks complete12Valid52Definition and clock checks complete
The states reconcile: 10 valid + 1 no opportunity + 1 invalid = 12 planned opportunities.
Valid measurement coverage: 10 / 12 × 100 = 83.3%.
The valid latencies total 520 seconds: 42 + 55 + 61 + 38 + 47 + 72 + 50 + 44 + 59 + 52 = 520. The arithmetic mean among valid observations is 520 / 10 = 52.0 seconds. Sorted values are 38, 42, 44, 47, 50, 52, 55, 59, 61, and 72, so the median is (50 + 52) / 2 = 51.0 seconds.
Pilot calculationResultPlanned opportunities12Valid latency observations10No opportunity1Invalid timing record1Valid coverage83.3%Sum of valid latencies520 secondsMean among valid latencies52.0 secondsMedian among valid latencies51.0 seconds
For illustration only, suppose a reviewer also asks how many valid latencies were at or below a 60-second demonstration boundary. Eight of ten meet that description, so 8 / 10 × 100 = 80.0%. The boundary is not a mastery criterion, benchmark, payer rule, or recommendation. It shows what a percentage-of-opportunities derivative retains and loses: the result hides the 38-to-72-second distribution and says nothing about the two nonvalid opportunities.
No-opportunity and invalid records are not entered as zero; they stay outside the latency summary. The pilot does not establish a baseline, change over time, treatment effect, client benefit, or adequate representation of all settings.
Explain why nearby measures answer different questions
The fictional team records why three alternatives were not selected for the exact latency question.
AlternativeWhat it could answerWhy it does not replace latency hereOccurrence within 60 secondsWhether a valid opportunity crossed a predefined boundaryDiscards the actual time and depends on a separately justified boundaryMomentary time sampling every 30 secondsWhether the response state is present at sampled momentsCannot locate the exact start after the cue and can miss changes between samplesPermanent productWhether an agreed product exists after the activityCannot show when initiation occurred, may reflect help from others, and may omit unfinished work
Those methods answer different questions rather than forming a universal hierarchy. A permanent product may be the least intrusive source for another goal. Sampling may be proportionate when continuous observation is impractical. The worksheet should make each match visible.
Review access, burden, privacy, and technology
Measurement should preserve ordinary access. The ASHA AAC practice resource supports communication-system access and describes multimodal communication. Although it does not prescribe this pilot, the resource cautions against treating device access as an experimental convenience. Record communication, sensory, mobility, and language conditions that affect opportunities. Health, pain, medication, fatigue, and environment can matter as well.
Technology needs its own validation. Confirm clocks, offline behavior, duplicate events, unit conversions, exports, permissions, and version history. A method can fit the question yet fail operationally when timestamps drift. It can also fail when observers cannot use the interface while providing care. A flawless app still cannot make an irrelevant measure valid.
Use authorized systems and the minimum information needed. Recording and remote observation require applicable consent and qualified review. School or workplace data, retention, and disclosure may require additional controls. Do not move protected information into an unmanaged spreadsheet.
Make a qualified selection and preserve the alternatives
Decision recordEntrySelected measure and exact unitQuestion it answersPilot evidence reviewedValidity and reliability evidence still neededRepresentativeness limitsClient and caregiver feedbackAccess, burden, and privacy conditionsRejected or deferred alternatives and reasonsDisplay and interpretation ruleTraining and implementation ownerRecalibration or replacement triggersQualified approval, date, and version
The BACB's Ethics Codes page provides current access. The Ethics Code for Behavior Analysts addresses competence, communication, consent, assessment, data collection, evaluation, documentation, and client responsibility. Those duties remain within the code's scope. Payer, licensure, privacy, employment, education, and legal requirements still need local review.
The CASP ABA Practice Guidelines Version 3 access page identifies the publication's scope and licensing conditions. This original worksheet neither reproduces that content nor claims endorsement.
Schedule reconsideration before the method drifts
Reopen the selection when the question, response definition, AAC access, prompt structure, or setting changes. Observer duties, technology, opportunity rate, data quality, client preference, risk, and decision use can also trigger review. Reconsider persistent missingness, weak agreement, excessive burden, or a misleading display.
Reconsideration triggerEvidence to reviewOwnerDue date or eventDefinition or response form changesOpportunity structure changesMissing or invalid data clusterObserver agreement or calibration concernClient burden, withdrawal, or access concernDevice, software, or export changeNew setting, staff role, or decision useQualified review finds a better-fitting measure
The selected method is a versioned decision, not a permanent property of the client. Preserve earlier units and effective dates. Do not draw a continuous trend across incompatible measures unless a qualified reviewer establishes and documents an appropriate bridge.
Related resources
- How to Select a Measurement System for an ABA Treatment Goal
- ABA Frequency, Rate, and Duration Data Sheet and Coverage Worksheet
- ABA Interval Recording Data Sheet and Sampling Review Worksheet
- ABA Missing, Invalid, and No-Opportunity Data Review Worksheet
Sources
- Behavior Analyst Certification Board, Test Content Outlines
- Behavior Analyst Certification Board, BCBA Test Content Outline, Sixth Edition
- Behavior Analyst Certification Board, Ethics Codes
- Behavior Analyst Certification Board, Ethics Code for Behavior Analysts
- Council of Autism Service Providers, ABA Practice Guidelines Version 3 access page
- LeBlanc and colleagues, A Proposed Model for Selecting Measurement Procedures
- Fiske and Delmolino, Use of Discontinuous Methods of Data Collection
- Morris and colleagues, Procedures and Accuracy of Discontinuous Measurement
- Calkin, On Terms: Frequency and Rate in Applied Behavior Analysis
- Slocum and colleagues, The Evidence-Based Practice of Applied Behavior Analysis
- American Speech-Language-Hearing Association, Augmentative and Alternative Communication