How does a multiple-baseline design work? A multiple-baseline design repeatedly measures several tiers and introduces the same intervention at different times across participants, settings, behaviors, or other justified units. Evidence strengthens when each tier changes only after its staggered intervention begins while untreated tiers retain their baseline pattern. Credible use requires comparable yet sufficiently independent tiers, trustworthy measurement, faithful implementation, planned timing, and acceptable delay.
Staggered starts create the comparison
Every tier begins in baseline. The intervention starts first in one tier while the others remain untreated and continue measurement. Once the first tier shows the expected change, the intervention begins in a second tier, then later in another. Each transition creates a new opportunity to test whether outcome change follows intervention timing.
The WWC Single-Case Design Technical Documentation describes the causal logic of repeated measurement, systematic manipulation, and demonstrations of effect at different points in time. Unlike a reversal design, a multiple baseline can build replication without removing an intervention that has already begun.
The design is useful when learning is durable, returning to baseline is unlikely, or withdrawal would be unacceptable. Its tradeoff is delay: later tiers may remain in baseline while earlier tiers receive a potentially useful intervention.
Tiers can be participants, settings, or behaviors
- Across participants: the same intervention and outcome are studied with different people.
- Across settings: one person's outcome is measured in places such as home, clinic, school, or community.
- Across behaviors: the intervention is staggered across distinct response classes for one person.
Other units, including staff, teams, materials, or routines, may serve as tiers when the causal logic is defensible. State the unit explicitly. “Three baselines” is too vague when readers cannot tell what differs.
Tier selection balances comparability with independence. Tiers should be similar enough that the same intervention and measure make sense. They should also be independent enough that treatment in one is unlikely to change another before its planned start.
A single-case design guide emphasizes independence and equivalence when selecting multiple-baseline conditions. Generalization across behaviors or settings can be valuable clinically while weakening the staggered experimental contrast. Record early change as generalization rather than forcing the tier to look stable.
Concurrent and nonconcurrent baselines differ
In a concurrent multiple baseline, tiers are measured during overlapping calendar periods. A shared event that affects every tier can become visible because only the treated tier should change at each stagger. In a nonconcurrent multiple baseline, cases begin at different calendar times. Replicated within-tier changes can still carry evidence, while shared history events become harder to evaluate across tiers.
The review by Slocum and colleagues defines multiple-baseline comparisons through offsets in calendar time, days in baseline, and sessions in baseline and examines threats in both variants. It also cautions that concurrent designs remain insensitive to some tier-specific events. An illness affecting one participant or a staff change affecting one setting may imitate treatment in only that tier.
Record calendar dates, session numbers, baseline duration, and tier-specific events. Calling a design concurrent because graphs appear side by side is insufficient when observations occurred in different months.
Plan start timing and phase rules
Stagger intervention onset far enough apart to observe the expected effect in the earlier tier before changing the next. If effects emerge slowly, tightly spaced starts can make the tiers appear to change together. If starts are excessively far apart, later tiers carry avoidable burden.
Predeclare how each transition will be chosen. A rule might require a stable enough pattern and a minimum exposure window, or it may randomize the start within an eligible range. The randomized single-case design review describes prospective randomization of intervention timing as one option. Clinical acceptability limits the eligible window.
Avoid shifting the next start merely because a desirable result appeared. Record any departure, reason, and effect on interpretation. Unexpected deterioration, assent withdrawal, urgent need, or risk can override the planned sequence.
Current standards set research-rating requirements
The current WWC handbook page identifies Version 5.0 as its current procedures and standards. For a multiple-baseline or multiple-probe design, its three-demonstration convention requires at least three tiers with phase changes at three different points in time. Additional design-specific observation and overlap requirements determine the rating.
Those rules govern WWC evidence review. A clinical conclusion also depends on level, trend, variability, immediacy, overlap, consistency, fidelity, measurement validity, confounds, fit, and risk. Meeting a tier count cannot rescue dependent tiers or an invalid measure.
Measure every tier throughout the design
Define the dependent variable, eligible opportunity, observation window, prompts, exclusions, and missing-data rule. Collect baseline observations often enough to show what each untreated tier does while other tiers receive intervention. A missing observation should stay visible with its reason.
Define the independent variable for replication and measure fidelity in every intervention tier. Preserve context logs, staff and setting changes, health or medication changes, observer training, and sampled agreement. The SCRIBE 2016 statement calls for reporting measures, intervention delivery, procedural fidelity, sequence, raw outcomes, adverse events, and limitations.
A fictional multiple-baseline example
A fictional practice evaluates an eight-step handoff checklist through simulations with three consenting staff members. No client information is used. Each simulation creates eight scored opportunities, and the same assessor, scenario-difficulty rules, and scoring definition apply across tiers. Training starts after 6 simulations for Rin, 9 for Tao, and 12 for Mara.
Rin completes 16 of 48 baseline steps correctly and 43 of 48 intervention steps. Tao completes 25 of 72 baseline steps and 45 of 48 intervention steps. Mara completes 34 of 96 baseline steps and 44 of 48 intervention steps. Untreated tiers remain near their baseline range until their own staggered starts. Trainers deliver the package correctly in 18 of 18 intervention simulations.
The staggered pattern supports a functional relation between this training package and simulated staff performance if raw session data show credible differentiation and no rival change. The summary percentages cannot substitute for the graph. The analysis establishes neither real-session fidelity nor a client outcome; those require separate measurement.
Ethics govern delay and tier selection
Keep ordinary supports, AAC, health care, safety protections, and emergency response available across tiers. Avoid baseline delay when waiting creates unacceptable risk or withholds an established necessary service. A multiple baseline may fit a training rollout or low-risk skill better than urgent clinical need.
The current BACB Ethics Code applies to BCBA and BCaBA certificants and people who completed an application. It addresses competence, client involvement, informed consent and assent when applicable, medical needs, risk, data, and evaluation. Scientific timing remains subordinate to those duties.
Related terms
Sources
- What Works Clearinghouse, Single-Case Design Technical Documentation
- What Works Clearinghouse, Handbooks and Other Resources
- Slocum and colleagues, Threats to Internal Validity in Multiple-Baseline Design Variations
- Smith, Single-Subject Experimental Design for Evidence-Based Practice
- Onghena and Edgington, Randomized Single-Case Experimental Designs in Healthcare Research: What, Why, and How?
- Tate and colleagues, The Single-Case Reporting Guideline In BEhavioural Interventions (SCRIBE) 2016 Statement
- Behavior Analyst Certification Board, Ethics Code for Behavior Analysts
Take the next step with clarity
Whether you are finding care, growing as a clinician, or building a stronger ABA practice, Finni brings the people, tools, and support together to help you move forward.
Explore clinical roles at Finni practices