An ABA goal attainment scaling template should make an individualized outcome easier to describe and audit, not easier to retrofit after the outcome is known. This worksheet helps a clinical team write one five-level Goal Attainment Scaling record before a defined review period, retain the baseline and measurement conditions, and later document why the available evidence supports one attained level. The worksheet does not decide what the client should value, invent an outcome hierarchy, or turn an ordinal rating into proof that treatment worked.

Clinicians & ABA Professionals / Data, Outcomes and Clinical Decision-Making.

The decision hidden inside a five-level scale

Goal Attainment Scaling, often shortened to GAS, begins with a personally relevant goal and a small ordered set of possible outcomes. The original Kiresuk-Sherman method described individualized scales and a later transformation across goals. In a clinical ABA setting, the useful first decision is more basic: can another informed reviewer distinguish the five descriptions using the same retained evidence?

Even a polished label such as "better than expected" leaves room for incompatible interpretations. One reviewer may focus on frequency, another on independence, and another on the setting. Putting those choices on the page before the rating date makes the later judgment inspectable and keeps the client's priority beside the conditions under which each level would count.

This is a scale-design aid. It is not a replacement for direct measurement, a treatment plan, a standardized instrument, or an assessment manual. The original Goal Attainment Scaling paper introduced individualized measurable scales and a standardized transformation, but it did not make every locally written scale valid, reliable, or clinically meaningful.

Five labels do not create equal units

A common GAS convention assigns -2, -1, 0, +1, and +2 to five outcomes around an expected level. The numbers identify order. They do not establish that the distance from -2 to -1 equals the distance from +1 to +2. They also do not mean 20%, 40%, 60%, 80%, and 100% of a goal.

Baseline needs its own field. Some GAS applications place it at -1; others use -2, and the appropriate location depends on how the scale was defined. No convention should silently choose that anchor. When the observed starting state fits none of the five descriptions, the record has exposed a design problem to resolve before use, not a reason to assign the nearest convenient number.

The five descriptions should be mutually exclusive, ordered in the intended direction, and observable under stated conditions. A scale can still be ordinal even when every anchor uses the same behavior dimension. Calling the levels "equidistant" requires evidence and a defensible interpretation beyond tidy wording.

When the record is not ready to write

Pause if any of these inputs is unresolved:

  • The client priority and intended benefit have not been confirmed with the client or an appropriate representative.
  • The target construct mixes unrelated outcomes, such as communication frequency, distress, and caregiver prompting in one rating.
  • The behavior dimension, observation window, valid opportunity, setting, partner, or measurement method is unspecified.
  • Baseline evidence is unavailable or was collected under conditions that will not be comparable at review.
  • Safety, assent, consent, accessibility, cultural, language, privacy, or legal questions would change how the outcome should be defined or observed.
  • The planned use requires a validated standardized measure, payer-specific instrument, or formal research protocol that this local worksheet cannot supply.

The current BACB Test Content Outline includes operational definitions, measurement validity and reliability, representative data, client-informed goals, contextual fit, and data-based decisions. Those competencies support careful scale construction; they do not endorse a universal GAS format or cutoff.

Blank one-goal scale record

Use this ABA goal attainment scaling template in a controlled clinical document. Include only the minimum identifying information needed for the task.

FieldRecord before the review periodRecord ID and versionClient-authored priority or authorized representative inputPlain-language purpose of this one scaleOne construct and observable behavior dimensionBaseline dates, evidence, and observed statusWhich level contains baseline, with rationaleReview date or review windowSetting, activity, people, materials, and access conditionsValid opportunity and exclusion rulesMeasurement procedure, unit, and observation scheduleAccommodations and communication accessAssent, consent, safety, cultural, and language considerationsEvidence to retain for later scoringWriter, client or representative, reviewers, and approval date

Write the outcome levels before seeing the review-period result:

ScoreOrdered description under the same declared conditionsEvidence that would distinguish this levelAnchor source and approval+2+10 expected outcome-1-2

Then complete the rating record without changing the scale:

Rating fieldReview entryReview-period evidence usedEvidence excluded and whyAssigned attained levelExact anchor language supporting the assignmentAmbiguous or in-between evidenceIndependent second rating, if plannedAgreement or disagreement and resolutionClient or representative interpretationUnexpected benefits, burdens, harms, or contextual changesScale revision needed for a future period

Would another reviewer choose the same row?

Each row should answer the same question with the same unit and context. If 0 describes independent requests per valid opportunity during a chosen routine, the neighboring rows should not suddenly switch to total daily requests or caregiver satisfaction. When two dimensions genuinely matter, separate scales or a clearly defined decision rule may be more interpretable than a compound sentence joined by "and."

Avoid adjectives that depend on private judgment. "Uses communication more consistently" is hard to reproduce. "Uses the agreed communication form in 6 to 8 of 10 valid opportunities across the two named routines, with no more than the defined gestural prompt" exposes what would be observed. Even that description needs a valid-opportunity rule, an observation schedule, and a reason those conditions reflect the client's priority.

Do not make the outer anchors impossible merely to force most results toward the middle. Likewise, avoid tiny distinctions that the measurement procedure cannot resolve. The 2023 educational review of GAS discusses challenges involving scale construction, baseline placement, level definition, context, rater reliability, and training. Its practical guidance is not a validation certificate for this worksheet.

Fictional example: joining a chosen community routine

The following record is invented for arithmetic-free demonstration. No real client is described, and the example does not recommend a goal.

FieldFictional entryPriorityParticipate in a chosen weekend makerspace activity and have an easy way to pauseConstructInitiating participation while preserving a break requestObservable dimensionNumber of independently initiated activity steps completed from 4 predeclared valid opportunitiesBaseline-1; usually completes 1 valid opportunity with a gestural prompt and uses the agreed break card when offeredConditionsSame makerspace routine, familiar communication system available, sensory supports in place, no scoring during safety interruptionsReview windowFour planned visits over six weeksEvidenceOpportunity-level record, prompt code, break access record, session context note, client feedback

ScoreFictional prewritten description+2Independently begins all 4 valid steps and independently uses the break option when wanted across at least 3 review visits+1Independently begins 3 of 4 valid steps, with the break option continuously available, across at least 3 review visits0Independently begins 2 of 4 valid steps, with no more than the defined gestural prompt on the other steps, across at least 3 review visits-1Independently begins 1 of 4 valid steps and accepts no more than the defined gestural prompt on another step across at least 3 review visits-2With the agreed communication access available, does not independently begin a valid step across at least 3 review visits

If the agreed communication system is unavailable, the record is invalid for this scale rather than evidence for -2. Log the access failure outside the ordinal rating and decide whether enough eligible visits remain. That separation prevents an environmental failure from being recast as client nonattainment.

Suppose the retained record fits +1, but one visit used different materials and another was interrupted. The rating note should identify which visits remained eligible, not average the five level numbers or revise +1 after seeing the data. If the evidence fits both 0 and +1, the honest result is ambiguity that informs the next scale version.

What a second rating reveals

An independent reviewer can apply the frozen anchor text to the same deidentified evidence before seeing the first rating. Keeping both levels and the reason for disagreement creates useful scale-design evidence. Agreement only shows that the descriptions were usable for that review; it does not prove construct validity, measurement accuracy, or a causal intervention effect.

The autism-focused randomized-trial methods paper describes a systematic GAS development process in one research setting. That example supports documenting a protocol and predefined levels. The research does not establish that the same process, anchors, or T-score interpretation transfers automatically to routine ABA care.

What the attained level leaves unresolved

An attained level records where the retained evidence landed on one prewritten ordinal scale. That can support a conversation about the chosen outcome and the next scale version. By itself, the level cannot establish reliable change, clinical importance, social validity, generalization, maintenance, or a treatment effect.

The 2022 scoping review of GAS in randomized trials reports wide variation in implementation and discusses third-party review, intervention relevance, level spacing, and facilitator training. The review supports caution about what a GAS result represents. That evidence does not create a universal five-level ABA instrument.

GAS also differs from the existing percent-of-goal-obtained calculator. PoGO uses phase means and a baseline-to-goal distance in a common quantitative unit. GAS assigns an achieved ordinal level from a prewritten individualized scale and may later aggregate several goals under additional assumptions. The two outputs are not interchangeable.

Records, privacy, and future revisions

The scale version belongs with the approval date, review evidence, rating rationale, and any later revision. Once an outcome is known, preserve the original anchors. A revised scale starts a future review period under a new effective date.

The BACB ethics-code page points practitioners to the current code and emphasizes consumer protection. The AERA standards landing page identifies the current testing standards used here for general measurement boundaries, not as approval of GAS for a particular use.

Where the record contains protected health information, organizations need their own legal and security analysis. HHS explains the Privacy Rule and the Security Rule, while noting that summaries do not replace the rules. Prefer coded examples, limit access, and avoid placing identifiable information in an unapproved calculator or collaboration tool.

Related resources

Sources