NHS Talking Therapies Outcome Measures by Session
Outline
It is 17:40 and Priya still has two NHS packs to close. The PHQ-9 is in. The GAD-7 is in. The problem descriptor changed to OCD two weeks ago, and nobody added the OCI. That is the working problem behind NHS Talking Therapies outcome measures: collect the right tools at the right session points, then write one sentence so the score is not only a national return.
This page is a clinician collection map for PWPs, high-intensity practitioners, supervisors, and private therapists working next to the pathway. It covers the core battery, anxiety-disorder specific measures by problem descriptor, session timing, caseness and reliable change, and how recovery is constructed from first and last scores. It is educational guidance, not an NHS form and not a substitute for the live service manual. Follow local EPR rules, the current NHS Talking Therapies manual, and NHS Digital constructions when they differ from any example here.
Educational resource for registered and licensed mental-health clinicians working in or alongside UK adult psychological-therapy pathways. Service manuals and dataset constructions change. Verify current thresholds and reporting rules against official sources before you rely on any workflow.
What you collect, and when
NHS England treats routine outcome monitoring as part of care, not as an audit extra. NHS Talking Therapies outcome measures are a session-by-session battery, not an intake-and-discharge extra. People complete depression and anxiety scales at every contact, with a short interference scale alongside the symptom scores, so a last available score still exists if treatment ends earlier than planned.
Stepped care and the episode spine sit in the NHS Talking Therapies overview. The collection work itself is which tool is due, which hour it belongs to, which number is caseness, and what the note has to say.
Scroll the visual sideways to view the full diagram
| Session point | What to collect | What the note must say | Why this hour |
|---|---|---|---|
| Assessment | Problem descriptor, PHQ-9, GAD-7, matching ADSM if listed, WSAS, employment status if your EPR requires it | Why this battery matches the descriptor | Wrong ADSM here follows the whole episode |
| Every treatment session | PHQ-9, GAD-7, WSAS, and matching ADSM if listed | Score, comparison point, one clinical meaning sentence | Last completed scores can become the outcome pair |
| Mid-episode review or step discussion | Current battery, trend from baseline, any descriptor change | Stay, step up, step down, or change the ADSM | Scores without a decision do not change care |
| Unexpected ending or DNA discharge | Last completed scores; do not invent missing forms | Incomplete data, last date, clinical reading of that pair | National constructions still use a last score when one exists |
| Planned ending | First and last scores on PHQ-9 and GAD-7, plus the ADSM when listed, plus WSAS | Ending reason, residual symptoms, next support | Recovery and reliable change are first-to-last constructions |
Do not wait for a “measures session.” The hour you skip is often the hour that becomes last contact.
The core battery
Three tools sit under almost every adult episode.
PHQ-9 is the depression measure for all patients in the service. Track the total and read item 9 the same day if it is endorsed. Item 9 is a risk signal, not a completed risk assessment.
GAD-7 is the default anxiety measure and stays in the session battery even when an ADSM is listed. It was built for generalized anxiety. It does not ask about panic attacks, compulsions, trauma intrusions, or agoraphobic avoidance. When those features are the problem you are treating, GAD-7 is the wrong lead anxiety scale for the recovery pair. Add the matching ADSM; do not stop collecting GAD-7.
WSAS (Work and Social Adjustment Scale) tracks interference at work, home, social life, private leisure, and relationships. Symptom drop without function change still belongs in the review conversation.
Keep the licensed service versions of these questionnaires. This page does not reproduce item text.
Which ADSM belongs to which descriptor
NHS England publishes the recommended anxiety or medically unexplained symptoms (MUS) measure against the primary problem descriptor. Collect PHQ-9 at every contact for every descriptor; it is not repeated as a table column below. The live NHS Talking Therapies manual (v7.1) asks you to give the matching ADSM at every treatment session in addition to PHQ-9, GAD-7, and WSAS. Do not drop GAD-7 when an ADSM is due. Public service-standards wording about replacing GAD-7 is about the lead anxiety scale in the recovery pair, not about omitting GAD-7 from collection. National recovery constructions use PHQ-9 plus the listed ADSM when that ADSM is present. If the ADSM is missing, the construction falls back to PHQ-9 with GAD-7.
| Primary problem descriptor | Anxiety or MUS measure | Recovery pair when complete |
|---|---|---|
| Depression | GAD-7 | PHQ-9 and GAD-7 |
| Generalized anxiety disorder | GAD-7 | PHQ-9 and GAD-7 |
| Mixed anxiety and depression | GAD-7 | PHQ-9 and GAD-7 |
| No problem descriptor | GAD-7 | PHQ-9 and GAD-7 |
| Chronic pain in context of anxiety or depression | GAD-7 | PHQ-9 and GAD-7 |
| Agoraphobia | Mobility Inventory (MI) | PHQ-9 and MI |
| Health anxiety | Health Anxiety Inventory | PHQ-9 and HAI |
| OCD | Obsessive-Compulsive Inventory (OCI) | PHQ-9 and OCI |
| Panic disorder | Panic Disorder Severity Scale (PDSS) | PHQ-9 and PDSS |
| PTSD | PTSD Checklist for DSM-5 (PCL-5) | PHQ-9 and PCL-5 |
| Social anxiety | Social Phobia Inventory (SPIN) | PHQ-9 and SPIN |
| Body dysmorphic disorder | Body Image Questionnaire (BIQ) weekly | PHQ-9 and BIQ |
| Chronic fatigue syndrome | Chalder Fatigue Questionnaire (CFQ) | PHQ-9 and CFQ |
| Irritable bowel syndrome | Francis IBS scale | PHQ-9 and Francis IBS |
| MUS not otherwise specified | PHQ-15 | PHQ-9 and PHQ-15 |
If the descriptor changes, add or switch the matching ADSM from the next session and write why. Keep GAD-7 in the battery. Priya’s OCD miss is the usual one: GAD-7 keeps falling while obsessions are never measured, then the ending recovery pair cannot describe the problem that was treated.
For high-intensity CBT session wording around homework and technique, the IAPT CBT documentation guide is the place to look. The measure set and session timing below still decide which scale is due.
Caseness and reliable change
Caseness means the score is at or above the clinical cut-off used in the service dataset. Reliable change means the difference exceeds measurement error for that scale. For NHS Talking Therapies outcome measures, the live service manual is the source. Confirm Table 9 before you treat a number as current.
| Outcome measure | Caseness (clinical case) | Reliable change index |
|---|---|---|
| PHQ-9 | 10 or above | 6 or more |
| GAD-7 | 8 or above | 4 or more |
| SPIN | 19 or above | 10 or more |
| OCI | 40 or above | 32 or more |
| PDSS | 8 or above | 5 or more |
| PCL-5 | 32 or above | 10 or more |
| Health Anxiety Inventory | 18 or above | 4 or more |
| Mobility Inventory | 2.3 per-item average | 0.73 or more |
| BIQ weekly | 40 or above | 10 or more |
| PHQ-15 | 10 or above | 7 or more |
| Chalder Fatigue Questionnaire | 19 or above | 5 or more |
| Francis IBS scale | 75 or above | 50 or more |
A drop that is smaller than the reliable change index can still matter clinically. It does not meet the national “reliable improvement” test. Write both facts when they diverge: “PHQ-9 14 to 11, not reliable change; sleep improved and item 9 is 0; stay on this step and re-score next contact.”
How national outcomes are built from your scores
NHS Digital reports three therapy outcomes for referrals that finished a course of treatment: recovery, reliable improvement, and reliable recovery. A course of treatment, in service language, is discharge after at least two sessions recorded as treatment or as assessment and treatment.
| Outcome | Clinical meaning | First-to-last test (service construction) |
|---|---|---|
| Recovery | Moved out of caseness | At caseness on PHQ-9 or the anxiety/ADSM at the start; below caseness on both at the last score |
| Reliable improvement | Change larger than measurement error | First-to-last change on the tailored pair exceeds the reliable change index |
| Reliable recovery | Both of the above | Meets recovery and reliable improvement |
| Reliable deterioration | Worsening larger than measurement error | Increase on the pair exceeds the reliable change index |
These constructions are for the published dataset. They are not a substitute for your formulation. Someone can recover on PHQ-9 and GAD-7 while the untreated ADSM would have told a different story. That is why the matching ADSM has to be in the recovery pair, collected alongside GAD-7 rather than instead of it.
National percentage targets move with planning guidance. Do not copy last year’s recovery target into a clinical note. Point supervisors at the current NHS Digital monthly publication if they need the service rate.
What to write so the score is usable
A dumped total without meaning wastes the collection time. Use the same four lines in the session note:
- Name the tools (PHQ-9, GAD-7, ADSM if listed, WSAS)
- Record today’s totals and the comparison point (assessment, last session, or both)
- State improved, stable, deteriorated, or incomplete, in one sentence
- Tie that sentence to today’s plan, including risk if item 9 or clinical inquiry raised it
Example, fictional: “PHQ-9 16 (assessment 18), GAD-7 11 (assessment 12), OCI 48 (assessment 52, still above caseness), WSAS work 6. Scores stable-high; exposure to checking was partial; plan stay on high-intensity CBT, OCI next session, risk low and unchanged.”
That is enough for a covering clinician. It is also enough for supervision when someone asks why you did not step up.
Incomplete forms belong in the same block: who declined, which scale, and whether language, time, or distress blocked completion. Do not score blank items as zero.
What private therapists should copy
If you work privately next to this pathway, borrow the timing and the interpretation habit. You do not need NHS dataset codes.
Copy:
- A named measure pair that matches the problem you are treating
- Session-by-session collection when the episode is short enough that last contact is unpredictable
- One clinical sentence, not a dashboard paste
- A written response to item 9 or any other risk flag the tool raises
Leave behind:
- Recovery percentages written as if your private chart were a national return
- Local EPR mandatory fields that only serve the IAPT dataset (still the dataset name in many submissions)
- ADSM swaps you cannot resource or license
For UK record-keeping outside the service, use the UK therapy documentation guide.
Closing
The job is narrower than “do outcomes.” Pick the ADSM from the problem descriptor, collect PHQ-9, GAD-7, and WSAS at every session, add that anxiety or MUS measure when it is listed, and write what the change means before you leave the desk. That is how NHS Talking Therapies outcome measures stay clinical rather than becoming a dashboard paste. First and last paired scores then have something honest to report.
If you want structured drafting with clinician review before anything is signed, start a free trial of Emosapien. NHS and private UK practices still own service policy, dataset submissions, and clinical sign-off. No software removes that work.
References
- NHS England. Service standards for NHS Talking Therapies, including ADSMs.
- NHS England. NHS Talking Therapies for anxiety and depression manual (confirm the live version and Table 8 / Table 9 at use).
- NHS England Digital. NHS Talking Therapies outcomes: recovery, reliable improvement, and reliable recovery.
- NHS England Digital. Submitting NHS Talking Therapies data.