An ABA Bland Altman agreement calculator describes the differences between two quantitative measurement methods applied to the same units. This worksheet calculates each pair’s mean and signed difference, the average difference or bias, the sample standard deviation of the differences, and approximate 95 percent limits of agreement. It also requires a direction label, an acceptability band chosen before reviewing results, and a plot-based check for magnitude patterns and unusual pairs.
The limits describe a model for individual paired differences. They are not a confidence interval for the mean, a correlation, or automatic evidence that two methods can be used interchangeably.
Clinicians & ABA Professionals / Data, Outcomes and Clinical Decision-Making.
Use paired quantitative data, not a binary IOA table
This calculator fits a narrow design: Method A and Method B measure the same quantitative construct, in the same units, on each matched unit. Examples might include two duration-timing methods applied to the same independently selected recordings or two export procedures applied to the same frozen test records. A pair must refer to the same underlying unit.
It is not the right calculator for positive/negative ratings, occurrence/nonoccurrence intervals, total-count IOA percentages, or six-way comparisons among familiar ABA IOA methods. The IOA method sensitivity calculator serves that separate purpose. The Cohen kappa calculator addresses binary nominal ratings. Bland-Altman differences preserve the original quantitative units.
The 1986 paper Statistical Methods for Assessing Agreement Between Two Methods of Clinical Measurement explains why correlation is misleading for method agreement and presents a graphical, difference-based alternative. Two methods can track high and low values together while still differing by an amount that matters for a decision.
Declare the direction and acceptable difference first
Before entering values in an ABA Bland Altman agreement calculator, write the difference direction explicitly:
di = Ai - Bi
A positive difference means Method A recorded a higher value than Method B. Reversing that direction must reverse the sign of the bias and both limits.
Before calculating, define what range of pairwise differences would be acceptable for the intended use. That band is a clinical, operational, measurement, and stakeholder judgment; the formula does not choose it. A method pair might be adequate for rough workload estimates but inadequate for a safety-sensitive latency measure. Do not create the acceptability band after seeing the limits.
Also freeze the measurement unit. Seconds, minutes, event counts, percentage points, and standardized scores cannot be pooled. If variability grows with the magnitude of the measurement, raw differences may be inappropriate; a log, percentage, regression, or other qualified approach may be needed.
Calculate bias and approximate 95 percent limits
For each of n pairs, calculate the pair mean and signed difference:
mi = (Ai + Bi) / 2
di = Ai - Bi
The mean difference, often called bias, is:
d-bar = sum(di) / n
Calculate the sample standard deviation of the differences:
sd = sqrt(sum((di - d-bar)^2) / (n - 1))
At least two complete pairs are required for that sample-SD formula. If all complete differences are identical, report sd = 0 and equal lower and upper limits at the common difference; do not describe the zero-width result as proof that either method is accurate or that future differences cannot vary.
The basic approximate 95 percent limits of agreement are:
lower LoA = d-bar - 1.96 x sd
upper LoA = d-bar + 1.96 x sd
The 1999 Bland and Altman paper Measuring Agreement in Method Comparison Studies describes these limits as an interval within which 95 percent of differences are expected to lie under the model. The paper also addresses magnitude-related differences, repeatability, repeated measurements, and nonparametric alternatives. This simple worksheet calculates only the basic independent-pair version.
Small samples make the estimated bias, standard deviation, and limits uncertain. The 1.96 multiplier does not erase that estimation uncertainty. If a consequential decision requires confidence intervals around the bias or limits, sample-size planning, repeated-measure handling, or a nonstandard scale, involve a statistician or qualified research-methods reviewer.
Freeze the measurement-comparison record
FieldPrespecified recordComparison ID, data version, and calculator versionClient-selected or client-informed purposeQuantitative construct and unitMethod A name, version, and procedureMethod B name, version, and procedureDirection, written as Method A minus Method BPair-matching rule and common observation boundaryEligible units, sampling frame, and exclusionsSetting, support, partner, scorer, and device contextMissing, censored, truncated, and invalid rulesAcceptable difference band chosen before resultsReason pairs can use the basic independent-unit modelReviewer and calculation date
Retain every pair in the working table:
Pair IDMethod A AiMethod B BiMean miDifference diContext or review note12...
Then summarize without deleting the rows:
ResultFull-precision valueDisplayed value or decision noteNumber of complete pairs nMean difference d-barSum of squared deviationsSample variance of differencesSample SD of differencesLower limit of agreementUpper limit of agreementPrespecified acceptable bandPairs outside the limits or acceptable bandMagnitude pattern or outlier concernRepeated-unit or dependence concern
Fictional example: two duration methods
All values are invented. Suppose two timing methods are applied to ten independently selected practice recordings, with duration recorded in seconds and differences defined as Method A minus Method B.
PairABMeanDifference1424443.0-22515050.513353233.534606160.5-15484647.026555555.007706668.048394240.5-39656464.5110444243.02
The differences sum to 7, so:
d-bar = 7 / 10 = 0.7 seconds
The sum of squared deviations from 0.7 is 44.1. The sample variance is 44.1 / 9 = 4.9, and:
sd = sqrt(4.9) = 2.2135943621 seconds
Therefore:
lower LoA = 0.7 - 1.96 x 2.2135943621 = -3.6386449498 seconds
upper LoA = 0.7 + 1.96 x 2.2135943621 = 5.0386449498 seconds
With a prespecified one-decimal display, record bias +0.7 seconds; approximate 95% limits -3.6 to +5.0 seconds. If the prespecified acceptable band were -3 to +3 seconds, the estimated limits extend beyond it. That comparison would argue against treating the methods as interchangeable for that stated use. It would not prove that either method is accurate or explain the source of the differences.
Plot the differences against the pair means
Create a scatterplot with each pair mean mi on the horizontal axis and difference di on the vertical axis. Add horizontal lines for the bias, both limits, and the prespecified acceptable band. Keep pair identifiers available for review.
Look for widening spread, curvature, a steady upward or downward pattern, clusters tied to a setting or device, and individual unusual pairs. A roughly horizontal cloud does not prove the assumptions, but a clear pattern can show that one pair of fixed limits is a poor summary. For example, larger positive differences at larger means may indicate proportional bias or a scale problem.
Do not remove an outlier solely because it widens the limits. Trace the pair to its source. Correct a documented data error under a versioned rule; otherwise report both the original result and any prespecified sensitivity analysis.
Outlier and direction sensitivity
In a separate fictional sensitivity view, change Pair 10’s difference from 2 to 8 while leaving the other nine differences unchanged. The bias becomes 1.3, the sample SD becomes 3.1989581637, and the limits widen to -4.9699580009 and 7.5699580009 seconds. This shows how one unusual pair affects a ten-pair estimate. It is a prompt to investigate, not permission to delete the row.
Now reverse the direction to Method B minus Method A. The original example must produce bias -0.7 and limits -5.0386449498 to 3.6386449498: the new lower endpoint is the negative of the old upper endpoint, and the new upper endpoint is the negative of the old lower endpoint. If that mapping fails, inspect the copied values, difference direction, rounding, and formulas.
Changing every Method A value by a fixed calibration offset shifts the bias and both limits by that offset but leaves the SD unchanged. That arithmetic fact does not establish that a post hoc correction is valid. Any calibration rule needs separate evidence and version control.
Stop when the basic model does not fit
The basic calculation assumes complete pairs from units that can reasonably be treated as independent and a difference distribution for which a mean and SD are useful summaries. Repeated observations from the same client, therapist, device, site, or recording can be clustered. When the underlying value changes between repeated measurements, the structure becomes more complicated. Bland and Altman’s repeated-observation paper describes methods for clustered designs rather than treating every row as independent.
Stop and seek qualified review when either method uses a different construct or observation boundary, pair matching is uncertain, units change, missingness is informative, differences are markedly skewed, variability changes with magnitude, or a floor or ceiling truncates values. Do not substitute correlation, a paired t-test, or a narrow confidence interval for the mean bias for the pairwise agreement question.
The ABA spreadsheet tool by Reed and Azulay and the review by Essig, Rotta, and Poling provide context for familiar IOA calculations and procedural-fidelity reporting. They do not make Bland-Altman limits the default for ABA data. The BCBA Test Content Outline, BACB ethics-code page, and Standards for Educational and Psychological Testing support careful measurement, competence, intended interpretation, and responsible data use; none supplies a universal acceptable limit.
Keep direct clinical graphs, raw pairs, observation notes, method-specific repeatability evidence, treatment-integrity data, and client and caregiver interpretation beside the calculation. Agreement between methods is not treatment effect, experimental control, validity, social importance, generalization, maintenance, competence, or safety.
Keep protected source rows in approved systems, replace direct identifiers with sanctioned record codes, and limit access to authorized roles. HHS publishes separate federal Privacy Rule and Security Rule summaries. The team must still resolve consent, retention, accessibility, payer, regulator, legal, security, and organizational requirements through qualified review.
Related resources
- ABA IOA Method Sensitivity and Agreement Comparison Calculator for Clinicians
- How to Calculate Duration per Occurrence IOA
- How to Align Observation Start and Stop Rules for IOA
- ABA Cohen Kappa Prevalence-Sensitivity Agreement Calculator for Clinicians
Sources
- BACB Ethics Codes
- BCBA Test Content Outline, 6th edition
- Standards for Educational and Psychological Testing
- Bland and Altman: Statistical Methods for Assessing Agreement
- Bland and Altman: Measuring Agreement in Method Comparison Studies
- Bland and Altman: Agreement With Multiple Observations per Individual
- Reed and Azulay: An Excel Tool for Calculating Interobserver Agreement
- Essig, Rotta, and Poling: Interobserver Agreement and Procedural Fidelity
- HHS summary of the HIPAA Privacy Rule
- HHS summary of the HIPAA Security Rule