What makes a behavior change clinically significant? A change is clinically significant when credible evidence shows a meaningful improvement in everyday life for the person receiving services, considering the goal, magnitude, consistency, context, general use, risks, burden, and the person’s own view. A numerical difference, statistically significant result, or experimental effect can contribute evidence and does not alone establish personal value.
Meaningful change has several parts
Ask whether the change:
- concerns an outcome the person values
- is large enough to matter in daily life
- appears consistently enough for the intended decision
- occurs in relevant settings and with relevant people
- lasts across a useful time period
- improves access, comfort, safety, participation, autonomy, or another priority
- avoids unacceptable side effects, burden, or loss of choice
The answer can differ across people even when their numerical change is identical.
Statistical, experimental, and clinical significance differ
Statistical significance concerns whether an observed result is unlikely under a statistical model and assumptions. Experimental control concerns whether a research design supports a causal relation between an intervention and outcome. Clinical significance concerns meaningful value for the person in context.
A study can show a reliable group difference with a small everyday effect. An individual graph can show a large change without a design that isolates cause. A personally valuable change may be hard to represent with one standardized score.
Start with the person’s priorities
Ask the person what would make life better and how they want to communicate the answer. Use AAC, sign, gesture, writing, pictures, interviews, ratings, or another reliable form. Family and clinician perspectives add context and remain separate sources.
A goal chosen only because it is easy to count may produce precise data with little personal importance.
A fictional example
Mei is a fictional 15-year-old who wants adults at school to respond to her “headphones” or “quieter” message before noise becomes painful. Across eight baseline opportunities with AAC available, adults respond within 30 seconds in 2 of 8, or 25%. Across 12 later opportunities, they respond in 10 of 12, or 83.3%.
Mei reports that 9 of the 12 later situations felt manageable. One classroom still lacks headphones, and one adult response is late. The change is large and relevant, while the team keeps the access failures visible.
The before-and-after comparison supports a meaningful-benefit hypothesis. It does not isolate cause because staff training, materials, schedules, practice, and time changed together.
Review measurement credibility
Check operational definitions, valid opportunities, observation coverage, denominator consistency, missing data, observer agreement, procedural fidelity, and changes in support. A dramatic graph built from inconsistent measurement provides weak evidence.
Preserve raw counts and individual observations. An average can hide one setting where the person still lacks access or faces risk.
Experimental evidence answers a causal question
The manifest starter, WWC Document 229, belongs to the What Works Clearinghouse standards resources. The current WWC Version 5.0 Procedures and Standards Handbook explains that single-case designs can support causal-effect review when the case serves as its own control and the design meets applicable standards.
Those standards address evidence quality in educational research. A design rating cannot decide whether an outcome is important to a person, clinically appropriate, or worth its burden.
Compare change with everyday reference points
Useful reference points may include the person’s stated target, safety requirement, access threshold, natural opportunity, prior functioning, meaningful peer or community demand, or validated measure interpretation. Choose the reference before declaring success.
Normative comparison can be useful for a defined question and should not become pressure to appear typical. Support can remain appropriate even when a person meets a benchmark.
Include general use, maintenance, and burden
Check whether the change appears with everyday partners, settings, materials, and ordinary supports. Review whether it lasts and remains valuable. General use does not require removing AAC, accommodations, or chosen help.
Count therapy time, travel, fatigue, family work, missed school or activities, privacy cost, and possible unwanted effects. A benefit can be real and still require a less burdensome design.
State uncertainty and the decision
Document what evidence supports the conclusion, what remains uncertain, which perspectives differ, and which decision follows. Terms such as “promising,” “meaningful in this routine,” or “needs replication” can be more accurate than a universal success label.
Set a review date and state what would change the decision. New pain, a different setting, loss of an accommodation, declining comfort, or a measurement change may alter the interpretation. Continue collecting the person’s report alongside performance data so a technically improved score does not conceal reduced choice, access, or wellbeing.
Software may calculate change or flag a threshold. An appropriately qualified clinician interprets clinical meaning with the person and relevant stakeholders.
Related terms
Sources
Take the next step with clarity
Whether you are finding care, growing as a clinician, or building a stronger ABA practice, Finni brings the people, tools, and support together to help you move forward.
Explore clinical roles at Finni practices