Library · Measurement

The scores are not for the payer

Measurement-based care gets sold to clinicians as a compliance obligation, which is exactly backwards and exactly why it fails. The evidence for routine outcome monitoring is about catching the clients who are not improving early enough to change course. That it also happens to satisfy payers is a side effect worth having, not the point.

Ask a room of clinicians about outcome measures and you will get a predictable split. Some administer them religiously. Most administered them for a while, found the scores went into a folder nobody opened, and quietly stopped. A few never started because the whole thing felt like turning therapy into a spreadsheet.

All three positions are defensible given how measurement is usually implemented. The version that works is narrower than "measure everything" and more useful than "measure nothing."

What the evidence is actually about

The strongest case for routine outcome monitoring is not that it makes good clinicians better. It is that clinicians — all clinicians — are poor at identifying which of their clients are deteriorating. Studies of therapist prediction consistently find that clinicians substantially under-identify clients who are getting worse, and that feedback about off-track cases changes outcomes for that subgroup.

That framing matters, because it tells you where the value is concentrated. For a client improving steadily, the score confirms what you already knew. For the one who is not, it surfaces something you would probably have missed until they dropped out. Measurement is an alarm, not a report card.

The corollary nobody says out loud. If you are not going to look at the score and let it change what you do, do not collect it. An unexamined instrument is administrative burden for the client, false reassurance for you, and — if a chart shows three years of worsening PHQ-9 scores with an unchanged treatment plan — an actively worse position at audit than not measuring at all.

Choosing instruments

A workable outpatient set is smaller than most guidance suggests. Two condition-specific measures and one general one cover most caseloads.

InstrumentMeasuresItemsRangeRecall
PHQ-9Depression90–272 weeks
GAD-7Anxiety70–212 weeks
PCL-5PTSD symptoms200–801 month
ORS / SRSGeneral distress and alliance4 + 40–40 eachPast week / session

The PHQ-9 and GAD-7 are the default pair for good reasons: free, brief, extensively validated, widely recognised by payers, and familiar to primary care, which matters when you are coordinating with a prescriber or a GP. The PCL-5 is the standard for trauma-focused work. The ORS/SRS pair occupies a different niche — it is session-by-session and includes an alliance measure, which makes it the practical choice for feedback-informed practice.

Cadence

The most common implementation error is measuring too often, which produces fatigue and meaningless variation, or too rarely, which produces a chart with two data points a year and no trend.

A defensible default: administer at intake, then every four to six sessions, then at termination. Increase frequency during acute periods or medication changes, where you actually want a tighter signal. Session-by-session instruments like the ORS are designed for a different rhythm and are the exception rather than a contradiction.

Whatever cadence you choose, write it into the treatment plan. A documented schedule turns irregular administration into a deliberate protocol, and it answers the question a reviewer would otherwise ask.

What payers look for

Payer expectations vary, but the pattern is consistent enough to plan around. What they want to see is not a high score or a low one — it is evidence that the condition is being tracked and that treatment responds to what the tracking shows.

  • A baseline established at or near intake.
  • Repeat administration on a discernible schedule.
  • Scores in the chart, dated and attributable, not summarised from memory.
  • Evidence the score influenced something — a plan revision, a frequency change, a referral, or an explicit decision to continue unchanged with reasoning.

That last item is the one practices miss. A score that appears in the chart and never appears in the clinical reasoning is data, not measurement-based care. See the golden thread for why the connection matters more than the number.

Screening is not diagnosis

Worth stating flatly because it is misused constantly, including by software. A PHQ-9 of 22 does not diagnose major depressive disorder. It indicates that a client endorsed a symptom pattern at a severity consistent with significant depression, on a self-report instrument, over a two-week window, on one occasion.

Diagnosis requires clinical evaluation — history, differential, functional impact, ruling out medical and substance-related causes. Instruments inform that process and do not replace it. A chart where diagnoses appear to have been assigned by score is a chart with a formulation problem, and it reads that way to a reviewer.

Reliable change and clinical significance

Two concepts worth knowing because they prevent over-reading small movements.

Reliable change is the amount of score movement that exceeds measurement error — the point at which you can say something actually changed rather than the instrument wobbled. For the PHQ-9 this is commonly cited around five points. A client going from 14 to 12 has probably not changed.

Clinical significance asks whether the client has crossed from a clinical to a non-clinical range. Both matter, and they answer different questions: reliable change asks "did something move," clinical significance asks "does it still constitute a problem."

Practical implementation

The failure mode is almost never clinical and almost always logistical. Instruments administered on paper get scored inconsistently and filed rather than entered. Instruments emailed separately get ignored. Instruments administered in session consume clinical time that clinicians understandably resent giving up.

What works is administration ahead of the session, scored automatically, visible before the client walks in. That turns a five-minute administrative task into a thirty-second review, and it is the difference between a practice that sustains measurement and one that abandons it in month four.

This is the specific thing Weft automates: measures go out on the cadence in the treatment plan, land in the chart already scored, and surface in the note draft as a sentence you can keep or cut. The clinical judgement stays yours. The clerical work stops being yours.

Instruments with licensing constraints

The PHQ-9 and GAD-7 are free to use without permission. Many other instruments are not — some require licensing fees, registration, or restrict use to particular settings. The ORS and SRS sit in a middle category with free individual use subject to registration.

Before standardising a practice on an instrument, check its terms. Practices occasionally build a workflow on a copyrighted measure and discover the licensing position only when they scale. It is a cheap thing to check early and an expensive thing to discover late.

Who fills the form in

A small operational decision with larger consequences than it appears. Self-administered instruments measure what the client reports; clinician-administered ones measure what the client reports to you, which is not identical. Where a client under-reports to avoid disappointing a clinician they like, self-administration ahead of the session tends to produce more honest data.

Where literacy, language or cognitive factors are in play, self-administration produces worse data rather than better. Reading items aloud is legitimate and should be recorded as the method, so a later reader knows why a score might differ from earlier self-administered ones.

Verified 29 July 2026. Instrument specifications are as published by their developers; interpretation thresholds vary by population and setting and are presented as commonly cited starting points rather than diagnostic rules. Screening instruments do not establish diagnoses. Primary references: VA National Center for PTSD; APA Services. This page is clinical reference, not clinical or legal advice.

Questions

Common questions

How often should outcome measures be administered?
A common default is at intake, every four to six sessions thereafter, and at termination, with increased frequency during acute periods or medication changes. Session-by-session instruments like the ORS follow a different rhythm. Whatever cadence you choose, document it in the treatment plan so administration reads as protocol rather than as irregular.
Do payers require outcome measures?
Requirements vary by payer and plan. What is consistent is that reviewers look for a baseline, repeat administration on a discernible schedule, scores dated in the chart, and evidence that results influenced the treatment plan. The connection to clinical decisions matters more than the scores themselves.
Can a PHQ-9 score diagnose depression?
No. It indicates a self-reported symptom pattern at a given severity over a two-week window. Diagnosis requires clinical evaluation including history, differential and functional impact, and ruling out medical and substance-related causes.
What is reliable change?
The amount of score movement that exceeds the instrument's measurement error, so you can say something genuinely changed. For the PHQ-9 this is commonly cited around five points — smaller movements are generally not interpretable as real change.
Which outcome measures are free to use?
The PHQ-9 and GAD-7 are free and require no permission. The PCL-5 is available at no cost from the VA National Center for PTSD. Many other instruments carry licensing terms — check before standardising a practice workflow on one.
Is it worse to collect scores and not use them?
In an audit sense, potentially yes. A chart showing worsening scores across years with an unchanged treatment plan documents that a signal was available and not acted on. Measurement you will not respond to is burden without benefit.