STarT Back Screening Tool Calculator

Free STarT Back Screening Tool calculator. Score all 9 items in about a minute, get the Q5-9 psychosocial subscale, and see the low, medium or high risk subgroup with the matched Keele care pathway.

STarT Back Screening Tool Calculator

Thinking about the last 2 weeks, answer all 9 items. The calculator returns the total score, the Q5–9 psychosocial subscale, the risk subgroup, and the matched care pathway Keele defines for that subgroup.

0 of 9 answered —

Attribution. The Keele STarT Back Screening Tool © Keele University 01/08/07, funded by Arthritis Research UK. Hill JC, Dunn KM, Lewis M, et al. Arthritis Rheum. 2008;59(5):632–641. Keele materials may not be reproduced without permission — keele.health.

Not a diagnosis and not a red-flag screen. The STarT Back Tool stratifies prognostic risk of persistent back-related disability. It does not identify serious spinal pathology, which must be excluded separately, and it is not a substitute for assessment by a licensed clinician.

Responses are calculated entirely in your browser. Nothing you enter is transmitted or stored.

STarT Back Screening Tool Calculator: Score, Stratify and Match Care

Most outcome measures tell you how a patient is doing. The STarT Back Screening Tool is different: it exists to tell you what to do next. Nine questions, answered in about a minute, sort a patient with low back pain into one of three prognostic subgroups — and each subgroup comes with a specific, manualised care pathway developed at Keele University.

Score it below, then read what most STarT Back pages skip: a scoring detail on item 9 that trips up more clinicians than it should, and an honest look at the evidence — strong for the tool's prognostic accuracy, more mixed for whether stratified care improves outcomes everywhere it has been tried.

What the STarT Back Tool does

The Subgrouping for Targeted Treatment (STarT) Back Screening Tool is a nine-item questionnaire developed by Jonathan Hill and colleagues at Keele University and published in Arthritis Care & Research in 2008. It stratifies patients with non-specific low back pain by their risk of persistent disabling back pain, using a mix of physical and psychosocial prognostic indicators.

It is a prognostic and allocation instrument, not a severity measure and not a diagnostic test. It was developed and validated in UK primary care, and that context matters when you transport it — see the evidence section.

The tool has been translated into more than 45 languages, has a version for children and adolescents, a six-item short form, and a LOINC panel code (91349-1) for interoperability.

The 9 items

The instruction stem is: "Thinking about the last 2 weeks tick your response to the following questions:"

Items 1 to 8 are first-person statements, answered Agree or Disagree. Item 9 is a question with five options.

#ItemConstructScoring1My back pain has spread down my leg(s) at some time in the last 2 weeksReferred leg painAgree = 12I have had pain in the shoulder or neck at some time in the last 2 weeksComorbid painAgree = 13I have only walked short distances because of my back painWalking limitationAgree = 14In the last 2 weeks, I have dressed more slowly than usual because of back painDressing limitationAgree = 15It's not really safe for a person with a condition like mine to be physically activeFear-avoidanceAgree = 16Worrying thoughts have been going through my mind a lot of the timeAnxietyAgree = 17I feel that my back pain is terrible and it's never going to get any betterCatastrophisingAgree = 18In general I have not enjoyed all the things I used to enjoyLow moodAgree = 19Overall, how bothersome has your back pain been in the last 2 weeks?BothersomenessSee below

Item 9 requires particular attention when scoring. The five options score Not at all = 0 · Slightly = 0 · Moderately = 0 · Very much = 1 · Extremely = 1. "Moderately" scores zero. The Keele form prints this as an explicit 0 0 0 1 1 row beneath the options. At least one widely mirrored US copy transcribed "Moderately" as scoring 1 — which is arithmetically inconsistent with the tool's 0–9 range — and another well-known physiotherapy site states the item is scored 0–4. Both errors change the risk category for real patients. The calculator above scores it correctly.

How to score and stratify

Two numbers come out of the nine items.

The algorithm, exactly as implemented in Keele's official calculator:

IF   total <= 3                       -> LOW RISK

ELSE IF total >= 4 AND subscale <= 3  -> MEDIUM RISK

ELSE IF total >= 4 AND subscale >= 4  -> HIGH RISK

Check the total first, then the subscale. Note that all nine items are required — the STarT Back Tool has no pro-rating rule, so an incomplete form cannot be stratified.

Worked example: a patient answers Agree to items 1, 3 and 6, and rates their pain as "Very much" bothersome on item 9 (items 2, 4, 5, 7 and 8 are Disagree). Total score = 4 (items 1, 3, 6 and 9). Psychosocial subscale (items 5–9) = 2 (items 6 and 9). Total ≥ 4 and subscale ≤ 3 → Medium risk.

What each risk subgroup means

These are not labels. Each maps to a distinct package of care defined in Keele's implementation manual and tested in the IMPaCT trial.

Low risk — total ≤ 3

16.7% had a poor outcome at 6 months in the development cohort.

Matched pathway: a single 30-minute primary-care consultation. Comprehensive assessment including physical examination; individualised education and reassurance about diagnosis, prognosis and treatment; advice on medication, activity and work; written materials and a short educational video. The patient is then discharged with advice to re-consult if necessary. Routine onward referral to physiotherapy is explicitly not part of the low-risk pathway.

The main clinical risk in this group is over-treatment, and the economics of stratified care depend substantially on not putting low-risk patients into extended episodes of care.

Medium risk — total ≥ 4, subscale ≤ 3

53.2% had a poor outcome at 6 months. Relative risk of persistent disability versus low risk: 2.19 (95% CI 1.10–4.38).

Matched pathway: standardised physiotherapy, up to six sessions — one assessment plus up to five follow-ups. An individualised plan negotiated with the patient, aimed at reducing symptoms and disability and promoting self-management, using advice, explanation, reassurance, education, manual therapy and exercise.

The manual explicitly does not recommend bed rest, traction, massage or electrotherapy. Acupuncture is optional at the discretion of the physiotherapist and patient.

High risk — total ≥ 4, subscale ≥ 4

78.4% had a poor outcome at 6 months. Relative risk versus low risk: 7.30 (95% CI 4.11–12.98).

Matched pathway: psychologically informed physiotherapy — a 60-minute assessment plus 45-minute treatment sessions, up to six in total. Everything in the medium-risk package, plus: building rapport; validating and normalising the patient's experience; a comprehensive biopsychosocial assessment; addressing knowledge gaps and correcting misunderstandings; and creating opportunities for the patient to respond differently to difficult internal experiences. It specifically targets the four psychological prognostic indicators — catastrophising, low mood, anxiety and pain-related fear — using simple cognitive-behavioural techniques.

Keele's manual stresses that clinicians delivering this pathway need ongoing clinical supervision from appropriately skilled personnel. Keele's implementation guidance flags gaps in this training and supervision as a key risk to pathway fidelity — allocating a high-risk patient to a clinician without that support undermines the pathway, and is one reason the tool can produce little or no benefit in some settings.

Does stratified care actually work? The honest answer

This is where most STarT Back content stops at the 2011 Lancet headline. The full picture is more useful.

The positive evidence

Hill et al., Lancet 2011 (the STarT Back RCT). 851 patients across ten English general practices, randomised 2:1 to stratified care or current best practice. Roland–Morris Disability Questionnaire improvement at 4 months: 4.7 versus 3.0 points, adjusted difference 1.81 (95% CI 1.06–2.57), effect size 0.32. At 12 months: 4.3 versus 3.3, difference 1.06 (95% CI 0.25–1.86), effect size 0.19. Economically the strategy was dominant — +0.039 QALYs at lower cost (£240.01 versus £274.40 per patient, a £34.39 average saving).

Foster et al., IMPaCT Back, Annals of Family Medicine 2014. A population-based implementation study across 64 family physicians and 922 patients. Overall RMDQ improved 0.7 points; the high-risk group improved 2.3. Time off work halved (4 versus 8 days) and sick certification fell about 30% (9% versus 15%), with lower health care costs and no worsening of outcomes.

The replications that failed

MATCH (US), Cherkin et al., JGIM 2018. A pragmatic cluster RCT across six primary care clinics in Washington State, 1,701 participants, with intensive implementation support — six clinician training sessions, five days of PT training, the tool embedded in the EHR. Result: no statistically significant differences in either primary patient outcome overall or within any risk subgroup at 6 months, and no change in care processes. Clinicians used the tool for about half of patients but, in the authors' words, "did not change the treatments they recommended."

TARGET (US), Delitto et al., eClinicalMedicine 2021. 76 primary care clinics across four US health systems, ~9,900 patients screened. Usual care plus psychologically informed physical therapy did not reduce the transition from acute to chronic low back pain: 47% versus 51% chronic at 6 months. Only 39% of intervention patients actually received the PIPT referral.

Denmark, Morsø et al., European Journal of Pain 2021. 334 patients randomised (half the intended sample). RMDQ improvement 5.9 versus 5.5 at 3 months, 6.1 versus 6.5 at 12 months — non-significant. Stratified care did produce fewer treatment sessions (3.5 versus 4.5) and lower prescription and imaging costs, but no overall cost-effectiveness advantage.

What to take from this

The tool's prognostic performance replicates reasonably well across settings. The claim that stratified care improves outcomes replicated in England but did not replicate in the two US trials or the Danish trial. The MATCH and TARGET authors point to gaps in implementation and pathway fidelity — audit-and-feedback, matched-treatment access, and clinician behaviour change — as plausible contributors, though the trials themselves cannot fully separate an implementation gap from a genuine ceiling on the intervention's effect outside the original UK health system. Either reading leads to the same practical conclusion: the benefit depends on a delivery system that reliably changes clinician behaviour, and it will not appear simply because the questionnaire is in the chart.

The MATCH authors offered three explanations worth heeding: no audit-and-feedback on adherence (unlike the UK trials), matched treatment options that were more numerous and harder to access, and a substantially more disabled US baseline population (RMDQ 11.8 versus 8.4 in the English studies).

Practically: if you deploy the STarT Back Tool, deploy the matched pathways and the audit process alongside it. The tool by itself has not been shown to change outcomes — the evidence supports the tool plus a reliably delivered care pathway, not the questionnaire in isolation.

Psychometrics

PropertyValueSourceTest–retest (quadratic weighted kappa)0.73 overall; 0.69 psychosocial subscaleHill et al. 2008Internal consistency (Cronbach's α)0.79 full tool; 0.74 psychosocial itemsMultipleICC0.90Bruyère et al. 2014Discriminative validity (AUC)0.73 (referred leg pain) to 0.92 (disability)Hill et al. 2008, developmentPoor outcome at 6 months by subgroup16.7% low / 53.2% medium / 78.4% highHill et al. 2008Inter-rater agreement with expert cliniciansWeighted kappa 0.28; 47% agreementHill et al. 2010, Clin J PainFloor / ceiling10.8% scored 0, 5.4% scored 9 (adequate); subscale floor 22.2% (poor)—

That inter-rater figure is worth sitting with rather than glossing over. A weighted kappa of 0.28 is low agreement by conventional benchmarks, and clinicians and the tool landed on the same subgroup only 47% of the time in that comparison. Kappa is also sensitive to how subgroups are distributed in a given sample, so this single figure shouldn't be read as a precise, universal error rate. What it does support is the developers' own guidance: use the tool's output "as an adjunct to their own decision-making rather than a replacement." A calculator that returns "High risk" is offering a prior worth weighing, not a verdict to act on unquestioned.

STarT Back or the Örebro questionnaire?

The Örebro Musculoskeletal Pain Screening Questionnaire (ÖMPSQ) is the main alternative. They were built for different jobs: the ÖMPSQ was designed as a prognostic tool, the SBST as a treatment-allocation tool.

Outcome predictedSTarT Back pooled AUCÖMPSQ pooled AUCPain0.59 — non-informative0.69 — poorDisability0.74 — acceptable0.75 — acceptableAbsenteeismnot pooled0.83 — excellent

Karran et al., BMC Medicine 2017 — systematic review and meta-analysis of 18 prospective cohorts in recent-onset low back pain.

The headline finding is the one to carry into practice: the STarT Back Tool predicts disability, not pain. A pooled AUC of 0.59 for future pain is essentially non-informative. If your question is "will this patient still hurt in six months," the SBST is the wrong instrument.

Key distinction: STarT Back predicts risk of persistent disability, not pain intensity. A pooled AUC of 0.59 for future pain is essentially non-informative — a low score does not mean a patient's pain will resolve, and a high score does not mean it will worsen. Use the score to guide how much care to deploy, not to counsel a patient on their expected pain trajectory.

Head-to-head agreement between the two is only moderate (Cohen's kappa 0.42, 70.2% agreement in 315 Swedish primary-care patients), and the SBST classified 53.7% as high risk versus 36.5% for the short ÖMPSQ — it flags substantially more people. Agreement was particularly poor in women over 50. On the other hand the SBST is far less error-prone to score in practice: 13 miscalculations versus 54 for the ÖMPSQ-short across those 315 patients. For work and absenteeism outcomes, the ÖMPSQ is the better instrument.

Limitations

The STarT Back Tool is a strong instrument for what it was built to do, but it has real boundaries worth stating plainly.

  • It predicts disability, not pain. A pooled AUC of 0.59 for future pain is non-informative. A low score does not promise a patient will be pain-free, and a high score does not mean their pain will worsen — it speaks to the risk of persistent disabling back pain, nothing else.
  • It does not screen for red flags. Cauda equina syndrome, fracture, infection, malignancy and other serious pathology require their own clinical assessment; the STarT Back Tool is a prognostic and allocation instrument layered on top of that assessment, not a substitute for it.
  • Agreement with clinical judgement is modest. A weighted kappa of 0.28 against independent expert clinicians means the tool and a clinician's own read of the patient will disagree a meaningful share of the time. It is designed to inform decision-making, not override it.
  • It was developed and validated in UK primary care. Translations exist in more than 45 languages and the tool's prognostic accuracy replicates reasonably well elsewhere, but the evidence that stratified care built on the score improves outcomes is strongest in the UK health system it was designed for — two US trials and a Danish trial did not replicate the outcome benefit (see the evidence section above).
  • It requires all nine items. There is no validated pro-rating rule for a partially completed form, so an incomplete questionnaire cannot be reliably scored or stratified.

Using the STarT Back Tool in a clinical workflow

The STarT Back Tool is built to sit inside a normal primary-care or physiotherapy visit, not alongside it. A practical sequence:

  1. Screen at first contact. Administer the nine items for adults presenting with non-specific low back pain, before treatment planning begins.
  2. Rule out red flags first. The tool is prognostic, not diagnostic — complete your standard clinical assessment for serious pathology alongside it, not in place of it.
  3. Score all nine items. An incomplete form cannot be reliably stratified; there is no pro-rating rule.
  4. Match the pathway, not just the label. Low, medium and high risk each map to a specific, manualised care package (see above). The evidence for benefit is for the tool plus that pathway — recording a risk category without delivering the matched care is not what was tested.
  5. Train and supervise. Keele's implementation guidance calls for clinician training and ongoing supervision, particularly for the psychologically informed physiotherapy delivered to high-risk patients.
  6. Re-assess and audit. Track how scores translate into referrals, treatment received and outcomes over time. The trials that failed to replicate a benefit largely lacked this audit-and-feedback loop — it is not an optional add-on.
Score it, stratify it, treat it — in one visit. Spry pairs the STarT Back score with the matched care pathway and built-in tracking, so the same visit that produces a risk category also produces a plan and an audit trail — the two things the evidence says stratified care actually depends on.

Licensing

The Keele STarT Back Screening Tool is © Keele University 01/08/07, funded by Arthritis Research UK. The tool, its translations, the implementation manual and Keele's own calculator are published openly and free of charge, and Keele materials may not be reproduced without permission. Enquiries: health.iau@keele.ac.uk. Retain the copyright line on any reproduction.

Frequently asked questions

How is the STarT Back Screening Tool scored?

Items 1 to 8 score 1 for Agree and 0 for Disagree. Item 9 scores 1 only for "Very much" or "Extremely" — "Moderately" scores zero. The total runs 0 to 9, and the psychosocial subscale is the sum of items 5 to 9, running 0 to 5.

What are the STarT Back risk categories?

A total of 3 or less is low risk. A total of 4 or more with a psychosocial subscale of 3 or less is medium risk. A total of 4 or more with a subscale of 4 or 5 is high risk.

Which items make up the STarT Back psychosocial subscale?

Items 5, 6, 7, 8 and 9 — fear-avoidance, anxiety, catastrophising, low mood and bothersomeness. Some secondary sources list this incorrectly; Keele's own form labels the box "Sub Score (Q5-9)".

Does Moderately score a point on STarT Back item 9?

No. Only "Very much" and "Extremely" score 1. The Keele form prints the score row as 0, 0, 0, 1, 1 beneath the five options. It is an item worth double-checking, since it is easy to mis-score.

What treatment matches each STarT Back risk group?

Low risk: a single 30-minute primary-care consultation with education, reassurance and advice, then discharge. Medium risk: standardised physiotherapy, up to six sessions. High risk: psychologically informed physiotherapy — a 60-minute assessment plus 45-minute sessions targeting catastrophising, low mood, anxiety and pain-related fear, with clinical supervision for the treating clinician.

Does the STarT Back Tool predict pain?

No. A meta-analysis of 18 cohorts pooled its accuracy for future pain at AUC 0.59, which is non-informative. It predicts disability (AUC 0.74), not pain intensity.

Is the STarT Back Tool validated outside the UK?

It has been translated into more than 45 languages and its prognostic performance replicates reasonably well. But the claim that stratified care improves outcomes has not replicated everywhere: two large US trials (MATCH, TARGET) and a Danish trial found no significant benefit, largely because matched pathways were not reliably delivered.

Can the STarT Back Tool replace clinical judgement?

No. Agreement between the tool and independent expert clinicians is poor (weighted kappa 0.28, 47% agreement), and the developers recommend using it as an adjunct to clinical reasoning rather than a replacement. It also does not screen for red flags.

Do I need training to use the STarT Back Tool?

No formal certification is required to administer or score the tool itself. Keele's implementation guidance does recommend training and ongoing clinical supervision for staff delivering the matched care pathways — particularly the psychologically informed physiotherapy for high-risk patients — since the evidence for benefit depends on the pathway being delivered reliably, not just the score being recorded.

Is the STarT Back Tool free to use?

Yes. The tool, its translations and Keele's implementation manual are published openly and free of charge for clinical use. The materials are copyrighted by Keele University and may not be reproduced or modified without permission.

References

  1. Hill JC, Dunn KM, Lewis M, Mullis R, Main CJ, Foster NE, Hay EM. A Primary Care Back Pain Screening Tool: Identifying Patient Subgroups for Initial Treatment. Arthritis Care & Research. 2008;59(5):632-641.
  2. Hill JC, Whitehurst DG, Lewis M, et al. Comparison of Stratified Primary Care Management for Low Back Pain With Current Best Practice (STarT Back): A Randomised Controlled Trial. The Lancet. 2011;378(9802):1560-1571.
  3. Foster NE, Mullis R, Hill JC, et al. Effect of Stratified Care for Low Back Pain in Family Practice (IMPaCT Back): A Prospective Population-Based Sequential Comparison. Annals of Family Medicine. 2014;12(2):102-111.
  4. Cherkin D, Balderson B, Wellman R, et al. Effect of Low Back Pain Risk-Stratification Strategy on Patient Outcomes and Care Processes: The MATCH Randomized Trial in Primary Care. Journal of General Internal Medicine. 2018;33(8):1324-1336.
  5. Delitto A, Patterson CG, Stevans JM, et al. Stratified Care to Prevent Chronic Low Back Pain in High-Risk Patients: The TARGET Trial. eClinicalMedicine. 2021;34:100795.
  6. Morsø L, et al. Effectiveness of Stratified Treatment for Back Pain in Danish Primary Care: A Randomized Controlled Trial. European Journal of Pain. 2021;25(9):2020-2038.
  7. Hill JC, Vohora K, Dunn KM, Main CJ, Hay EM. Comparing the STarT Back Screening Tool's Subgroup Allocation of Individual Patients With That of Independent Clinical Experts. The Clinical Journal of Pain. 2010;26(9):783-787.
  8. Karran EL, McAuley JH, Traeger AC, Hillier SL, Grabherr L, Russek LN, Moseley GL. Can Screening Instruments Accurately Determine Poor Outcome Risk in Adults With Recent Onset Low Back Pain? A Systematic Review and Meta-Analysis. BMC Medicine. 2017;15(1):13.
  9. Bruyère O, Demoulin M, Beaudart C, Hill JC, Maquet D, Genevay S, Mahieu G, Reginster JY, Crielaard JM, Demoulin C. Validity and Reliability of the French Version of the STarT Back Screening Tool for Patients With Low Back Pain. Spine. 2014;39(2):E123-E128.

Why settle for long hours of paperwork and bad UI when Spry exists?

Modernize your systems today for a more efficient clinic, better cash flow and happier staff.
Schedule a free demo today