When is a multiple-probe design used? A multiple-probe design is used when continuous baseline measurement across every tier would be unnecessary, burdensome, impractical, or likely to create practice effects. It retains the staggered logic of a multiple-baseline design while measuring untreated tiers at planned intervals. It fits predictable skill acquisition better than highly variable or rapidly changing outcomes and requires strategically timed probes before each intervention start.
Planned probes replace some continuous baseline observations
A multiple-baseline design repeatedly measures every untreated tier. A multiple-probe design intentionally leaves gaps and measures those tiers at preplanned checkpoints. The intervention still begins at different times across participants, behaviors, settings, or other justified units.
The causal question remains the same: does each tier change only after its staggered intervention begins? Planned probes check whether untreated performance stayed near baseline while earlier tiers received intervention. The WWC Single-Case Design Technical Documentation describes the broader logic of repeated measurement, systematic manipulation, and effect demonstrations at different points in time.
Planned missingness is part of the design. It should appear as gaps on the graph. Connecting unobserved intervals with a continuous line or imputing stable performance makes the evidence look stronger than the data collected.
Skill acquisition is a common fit
Repeated baseline trials can teach through practice, expose the person to repeated failure, create frustration, or consume time without providing instruction. Intermittent probes reduce that exposure. The design often fits discrete, observable skills that are expected to remain low without teaching and improve after it.
A single-case design guide describes multiple probes as a variation of multiple baseline in which intermittent baseline assessment reduces data-collection burden. It cautions that highly variable behavior may require continuous measurement because sparse probes may fail to reveal the untreated pattern.
A 2024 rehabilitation review identifies testing effects and baseline frustration as possible reasons to select a probe design. It also notes the tradeoff: fewer observations provide less information about maturation, outside intervention, and other baseline changes.
Probe timing carries the design logic
A defensible schedule commonly includes:
- initial probes across every tier during overlapping calendar periods
- repeated measurement in the first tier before its intervention starts
- probes of untreated tiers when an earlier tier enters or progresses through intervention
- a sufficiently dense baseline series immediately before each tier's own intervention
- repeated outcome measurement after intervention begins
- maintenance or generalization probes when those questions matter
Exact schedules depend on the question and the governing research standard. Predeclare the windows, decision rules, and reasons. Extra probes may be needed when an untreated tier changes, variability appears, an outside service begins, or a long interval has elapsed.
Probe conditions should measure performance without accidentally delivering the teaching procedure being tested. Define feedback, prompting, reinforcement, materials, and stopping rules. A probe can still provide ordinary access, safety, and communication supports.
Tiers need comparability and independence
Across-participant tiers should receive a relevant version of the same intervention and outcome measure. Across-behavior or across-setting tiers should be distinct enough that teaching one is unlikely to change the others before their planned starts.
Generalization is clinically valuable and experimentally important. If an untreated skill improves after a related skill is taught, record the change. That tier may no longer support the planned replication, even though the person gained a useful ability.
The multiple-baseline validity review emphasizes offsets in calendar time, days, and sessions and explains how maturation, testing experience, and coincidental events can threaten staggered designs. Sparse probes make an event confined to an unmeasured interval especially hard to detect.
Current standards add probe-specific requirements
The current WWC handbook page identifies Version 5.0 as its current procedures and standards. It treats multiple probe as a special case of multiple baseline. Planned missing data are the main difference, so multiple-probe findings must satisfy the multiple-baseline requirements plus additional rules for initial overlapping probes and baseline observations around phase changes.
The current handbook uses at least three tiers with changes at three different times for its three-demonstration convention. Its exact point, overlap, and phase requirements determine a WWC research rating. Those counts do not guarantee clinical fit, stable measurement, or a functional relation.
Define measurement and implementation separately
For each probe, specify the eligible opportunity, target response, scoring window, prompts, feedback, trial count, session length, exclusion, and missing-data rule. Use matched tasks where needed, then report remaining difficulty differences. Graph raw observations in their true time positions.
Define the independent variable with enough detail for replication. Measure procedural fidelity during intervention and probe fidelity during baseline checkpoints. Train and calibrate observers and sample agreement across tiers and phases. Keep observer agreement, fidelity, outcomes, and context logs separate.
The SCRIBE 2016 statement calls for reporting operational measures, intervention and control delivery, procedural fidelity, completed sequences, raw outcomes, adverse events, and limitations. A concise graph should not hide probe omissions or changed procedures.
A fictional multiple-probe example
A fictional clinician evaluates video modeling for three independent, eight-step documentation simulations completed by a consenting adult trainee. The scenarios use fictional records. Training is staggered across secure messaging, clinical handoff, and incident routing. Initial probes occur for all three skills, followed by boundary probes before each skill's training.
Secure messaging begins at 2 of 8 correct steps and reaches 23 of 24 across three post-training probes. Clinical handoff scores 1 of 8 initially and 2 of 8 just before its later start, then reaches 22 of 24 post-training. Incident routing scores 1 of 8 initially, 1 of 8 after the first skill is trained, and 2 of 8 after the second, then reaches 24 of 24 after its own training.
The staggered probes are compatible with a training effect if fidelity, scenario matching, timing, and raw data are credible. They also show no large early generalization across skills. This small illustration does not establish live-work performance, client benefit, or a universal training method.
Burden and safety govern selection
Intermittent probes reduce repeated exposure while leaving less information between checkpoints. Use continuous measurement when risk, health, variability, or rapid change requires closer monitoring. Avoid delaying effective support for a research sequence when the wait creates unacceptable harm or burden.
Preserve AAC, food, water, bathroom access, mobility, prescribed care, pain care, rest, emergency help, and effective safety protections across probes. The current BACB Ethics Code applies to BCBA and BCaBA certificants and people who completed an application. It addresses competence, client involvement, informed consent and assent when applicable, medical needs, risk, data, and evaluation.
Related terms
Sources
- What Works Clearinghouse, Single-Case Design Technical Documentation
- What Works Clearinghouse, Handbooks and Other Resources
- Smith, Single-Subject Experimental Design for Evidence-Based Practice
- Perdices and colleagues, The Role of Single-Case Experimental Designs in Evidence Creation in Rehabilitation
- Slocum and colleagues, Threats to Internal Validity in Multiple-Baseline Design Variations
- Tate and colleagues, The Single-Case Reporting Guideline In BEhavioural Interventions (SCRIBE) 2016 Statement
- Behavior Analyst Certification Board, Ethics Code for Behavior Analysts
Take the next step with clarity
Whether you are finding care, growing as a clinician, or building a stronger ABA practice, Finni brings the people, tools, and support together to help you move forward.
Explore clinical roles at Finni practices