What is internal validity in single-case design? Internal validity is the credibility of a causal inference that a manipulated independent variable produced the observed change in a dependent variable within the study. It strengthens when condition changes and outcomes correspond repeatedly, measurement and implementation are trustworthy, and history, maturation, instrumentation, sequence, carryover, selection, and other rival explanations become implausible.
Internal validity asks what caused the change
Suppose an outcome improves after an intervention starts. The timing supports a treatment hypothesis, while a staff change, medication change, recovery from illness, practice, or a new measurement rule may offer another explanation. Internal validity concerns how well the design and conduct separate the manipulated condition from those alternatives.
The WWC Single-Case Design Technical Documentation describes design standards that evaluate internal validity before reviewers assess evidence for a relation between the independent and outcome variables. Repeated measurement and systematic manipulation are central. A case can serve as its own control across conditions.
Internal validity belongs to an inference, not an intervention name. The same procedure can have stronger causal evidence in one study and weaker evidence in another. A clean-looking graph cannot supply missing control.
Common threats create rival explanations
A review of validity threats in ABA single-case experiments applies a broader research-design taxonomy to single-case work. The practical threats include:
- history: an outside event coincides with the condition change
- maturation: development, fatigue, recovery, or another time-linked process changes the outcome
- testing or reactivity: repeated measurement changes the response
- instrumentation: observers, definitions, equipment, or sampling change
- attrition: missing participants, tiers, sessions, or observations alter what remains
- selection: cases, task sets, settings, or times differ systematically across conditions
- ambiguous precedence: the supposed cause may follow or develop with the outcome
Single-case studies also face sequence and carryover effects. Practice during one condition can change later performance. A medication, learning procedure, satiation effect, or emotional response may persist after a phase ends. Frequent alternation can itself influence behavior.
A confound appears when another plausible cause aligns with the independent variable. If intervention sessions always use an experienced clinician and baseline always uses a new clinician, staff and treatment are inseparable. The evidence may concern the combined package while failing to isolate the treatment.
Replication makes coincidence less plausible
A single A-B change has weak control over history and maturation. Stronger single-case designs create repeated opportunities to show an effect at different times. Reversal designs repeat condition changes. Multiple-baseline designs stagger intervention across tiers. Alternating-treatments designs repeat nearby contrasts. Changing-criterion designs test whether outcomes track planned criterion shifts.
The useful sequence is prediction, verification, and replication. Baseline predicts the future pattern under similar conditions. Untreated tiers, withdrawal, or comparison conditions test that prediction. Later changes replicate the effect.
Replication can repeat a flaw. If the same staff change accompanies every intervention phase, each demonstration preserves the confound. If the target response is irreversible, a reversal can fail to verify the effect even when learning occurred. Design selection must fit the expected behavior of the outcome.
The family-of-single-case-designs review notes that systematic environmental changes, maturation, dropout, and uncontrolled events can obscure treatment-outcome relations. It also explains how reversal and multiple-baseline logic can address some of these influences.
Design structure and execution both matter
The current WWC handbook page identifies Version 5.0 as the current standards. Design-specific rules govern eligible phases, observations, and attempts to demonstrate an effect. Those thresholds define a research-rating framework rather than a universal clinical minimum.
Within any design, define the independent variable and dependent variable precisely. Keep measurement stable across conditions. Measure procedural integrity separately. Train and calibrate observers and sample agreement where appropriate. Preserve raw observations, phase dates, condition order, missing data, staff and setting changes, health or medication changes, and protocol deviations.
High observer agreement addresses scoring consistency without proving that the measure captures the intended construct. High fidelity shows the procedure was delivered while leaving a coinciding history event possible. Stable baseline supports prediction, yet the next phase may change for reasons beyond treatment.
The SCRIBE 2016 statement calls for reporting operational measures, intervention delivery, procedural fidelity, completed sequences, raw outcomes, adverse events, and limitations. Transparent reporting lets readers judge threats that the design could not remove.
A fictional internal-validity example
Three fictional adults choose to test the same accessible checklist for beginning separate low-risk volunteer routines. Baseline and intervention observations occur concurrently. Checklist introduction is staggered after 6 observations for Alex, 9 for Bo, and 12 for Chen. Ordinary communication and safety supports remain available in every condition.
Alex starts within two minutes in 1 of 6 baseline observations and 5 of 6 intervention observations. Bo starts in 2 of 9 baseline observations and 6 of 6 intervention observations. Chen starts in 2 of 12 baseline observations and 5 of 6 intervention observations. Before each person's staggered start, the untreated cases stay near their earlier levels. The checklist condition is implemented as defined in 18 of 18 intervention observations.
Three staggered changes support internal validity if cases, measurement, settings, opportunities, and fidelity are credible. A calendar-wide event would need to explain why each change appeared only after that person's intervention began. The evidence still applies to the studied checklist packages and routines. It does not establish preference, generalization, maintenance, or benefit for other people.
Internal and external validity answer different questions
Internal validity asks whether the independent variable caused the observed change in this study. External validity asks how far that relation applies across people, settings, responses, materials, implementers, and time. Strong causal evidence can have narrow generality. Broad observation across settings can remain weak on cause.
Construct validity asks whether the study's operations represent the concepts named in the conclusion. Statistical conclusion validity concerns the adequacy of quantitative inference. These forms of validity interact, though one does not substitute for another.
Safety can override the strongest design
Design choices should preserve AAC, food, water, bathroom access, mobility, prescribed care, pain care, rest, emergency help, and effective safety protections. Avoid withdrawal or delay when loss of benefit could create unacceptable risk. Medical decisions remain with the authorized medical professional.
The current BACB Ethics Code applies to BCBA and BCaBA certificants and people who completed an application. It addresses competence, client involvement, informed consent and assent when applicable, medical needs, risk, data, and evaluation. A safer, less controlled design can be the responsible choice, with causal limits stated plainly.
Related terms
Sources
- What Works Clearinghouse, Single-Case Design Technical Documentation
- What Works Clearinghouse, Handbooks and Other Resources
- Tincani and Travers, Applying the Taxonomy of Validity Threats From Mainstream Research Design to Single-Case Experiments in Applied Behavior Analysis
- Tate and colleagues, The Single-Case Reporting Guideline In BEhavioural Interventions (SCRIBE) 2016 Statement
- Dallery and colleagues, The Family of Single-Case Experimental Designs
- Behavior Analyst Certification Board, Ethics Code for Behavior Analysts
Take the next step with clarity
Whether you are finding care, growing as a clinician, or building a stronger ABA practice, Finni brings the people, tools, and support together to help you move forward.
Explore clinical roles at Finni practices