Depression screening tools compared: PHQ-9, CESD-R, DASS-21, HAM-D, MADRS, and more

Compare six depression screening and assessment tools by purpose, respondent, score range, timeframe, and clinical role.

For most adult primary care screening, the PHQ-9 is the practical starting point, while the PHQ-2 can serve as a brief first step if results at or above its usual cutoff lead to fuller assessment. CESD-R fits some research protocols. HAM-D and MADRS require trained raters. No score alone establishes a diagnosis or selects treatment. That limit matters.

This guide compares the PHQ-9, PHQ-2, CESD-R, DASS-21 reference, HAM-D, and MADRS, with attention to the exact edition, population, purpose, and follow-up process behind each name.

Depression assessment tools at a glance

ToolWho completes itScored content and timeframeBest-fit role
PHQ-9Adult patientNine symptoms over the past two weeks; total 0 to 27Brief screening and symptom-severity tracking
PHQ-2Adult patientTwo symptoms over the past two weeks; total 0 to 6First-step screen before fuller assessment
CESD-RAdult respondentTwenty items about the past week or so; raw display and a separate collapsed scoring calculationPopulation research or a protocol that specifies the CESD-R
DASS-21 referenceRespondentTwenty-one items split across depression, anxiety, and stress domains for the past weekDimensional, three-domain assessment
HAM-DTrained clinicianA clinician-rated depression scale; the local definition uses 17 itemsSpecialist assessment or a research protocol that names an exact edition
MADRSTrained clinicianTen clinician-rated items; total 0 to 60Measuring change in specialist care or trials

The numerical PHQ details come from the original PHQ-9 validation study and PHQ-2 validation study. The official CESD-R explanation documents its 20-item algorithm. The DASS developer site describes the three-domain scale. The original HAM-D paper and MADRS paper describe clinician-rated measures, but later editions and scoring conventions differ.

What a depression score can establish

A screening score describes answers given for a defined timeframe and can show that symptoms merit more assessment or that reported symptom burden has changed. It cannot establish the cause, rule a disorder in or out, or choose a treatment without clinical context.

A response about suicidal thoughts or self-harm requires direct, timely follow-up regardless of the total score.

If you need help now. In the US, Call 988 or text 988. Call 911 if you are in immediate danger. Outside the US, contact your local emergency number or find support in your country.

Some depression scales contain a safety-related item and others do not. A single questionnaire item is not a structured suicide-risk assessment. A low total also cannot cancel an endorsed safety item.

PHQ-9: the practical adult primary care option

The PHQ-9 asks about nine depression symptoms over the previous two weeks and sums them from 0 to 27. The original adult validation study supports the familiar descriptive boundaries at 5, 10, 15, and 20.

TotalDescriptive band
0 to 4None to minimal symptoms
5 to 9Mild symptoms
10 to 14Moderate symptoms
15 to 19Moderately severe symptoms
20 to 27Severe symptoms

These are symptom bands, not treatment instructions. A person with a low total may still need care because of functional change, a safety concern, a symptom outside the questionnaire, or a condition that the PHQ-9 does not measure.

In the original adult primary care and obstetrics-gynecology sample, a cutoff of 10 had 88% sensitivity and 88% specificity for major depression against an independent mental health interview (Kroenke et al., 2001). Those figures describe that study and cutoff, while accuracy elsewhere changes with the setting, population, language, and reference standard.

The PHQ-9 works well when a service needs a short adult screen that can also describe symptom severity over time. It is less useful when physical illness, pregnancy, medication effects, sleep disruption, or another condition may explain several physical symptoms. In those cases, the item pattern, function, history, and clinical interview matter more than the total alone. Read the answers.

PHQ-2: a first step, not a rule-out test

The PHQ-2 uses the first two PHQ-9 symptoms and produces a total from 0 to 6. The original adult study used 3 or more as the usual threshold for a positive screen. At that threshold, sensitivity was about 83% and specificity was 90% in the detailed analysis (Kroenke et al., 2003). It is a gate.

A result below 3 does not prove that depression is absent or that no further assessment is needed, because the PHQ-2 samples only depressed mood and loss of interest. It can miss people whose main concerns involve sleep, energy, concentration, guilt, movement, appetite, or safety.

A result of 3 or more supports follow-up with a fuller assessment, often the PHQ-9 or a clinical interview. It does not mean that major depressive disorder is likely for that individual. The PHQ-2 also contains no question about suicidal thoughts or self-harm, so services need a separate safety process.

CESD-R: detailed research-oriented symptom assessment

The official CESD-R explanation describes 20 items organized into nine symptom groups. It uses a collapsed 0-to-60 CESD-style total and a category algorithm. Survey Doctor also displays the sum of the stored 0-to-4 item values, which can range from 0 to 80.

That difference is easy to mishandle. The official cutoff of 16 applies to the collapsed 0-to-60 calculation, not to Survey Doctor's raw 0-to-80 display. The category names produced by an algorithm remain screening outputs. They should not be presented to a reader as a diagnosis.

CESD-R can fit population research or a study protocol that calls for its broader symptom coverage. It is longer than the PHQ-9 and less familiar in routine primary care. Confirm the exact form, scoring calculation, population, and follow-up plan before use.

DASS-21: three domains, different purpose

The DASS developer's guidance describes depression, anxiety, and stress as three seven-item domains. The depression domain emphasizes low positive emotion, loss of interest, and related distress. It does not map directly onto all diagnostic criteria for a depressive disorder and does not contain a suicide or self-harm item.

The DASS-21 is useful when a protocol needs three dimensional scores from one questionnaire. It is not a substitute for a diagnostic interview or a complete safety process.

The DASS-21 can be useful when a protocol needs three dimensional scores from one questionnaire. It is not a substitute for a diagnostic interview or a complete safety process.

HAM-D and MADRS: clinician-rated tools

HAM-D and MADRS are not self-completed screening questionnaires. A trained clinician rates symptoms using a defined interview and scoring rubric. That can add observed behavior and follow-up questions, but it also makes rater training and the exact edition important.

HAM-D has several versions and severity conventions. Record the exact edition, interview process, rubric, and timeframe so repeated results remain comparable.

The original MADRS paper describes 10 clinician-rated items scored from 0 to 6 for a total from 0 to 60. Later forms and interviewing conventions can differ, so use a defined version and trained raters.

Neither clinician-rated scale should be treated as automatically superior to a patient-completed measure. They answer a different question and require more resources. Rater consistency matters. A clinical trial should use the instrument and edition named in its protocol.

Choosing by setting and purpose

The setting changes what a useful result looks like. Purpose comes first. Choose the respondent, depth, and follow-up process before choosing the scale.

Adult primary care

Use the PHQ-9 when you need one brief adult depression screen with a severity total. A PHQ-2-first workflow can reduce the initial question burden, but it needs a clear next step for results at or above 3 and a separate safety process.

For older-adult protocols, compare the GDS-15 and PHQ-9 before choosing a form or cutoff.

The U.S. Preventive Services Task Force recommends adult depression screening when systems can provide further evaluation and care. It found no evidence for one optimal screening interval. Set timing from clinical judgment, risk factors, symptoms, life events, and the purpose of another result rather than applying a universal annual or every-visit rule.

For implementation details, see the primary care depression screening workflow.

Specialty care and clinical trials

Choose the measure that matches the protocol and decision. The PHQ-9 can provide a patient perspective. A verified HAM-D or MADRS edition may add a trained rater's assessment when the service can support consistent administration. Do not mix editions or compare totals across different scales as if the numbers shared a unit. They do not.

Population and academic research

CESD-R may fit studies that need broad self-completed symptom coverage. DASS-21 may fit a protocol designed around separate depression, anxiety, and stress domains. The study plan should name the exact version, population, language, scoring method, missing-data rule, and rights basis before collection starts.

Adolescents and perinatal populations

Do not assume an adult validation automatically covers adolescents, pregnancy, or the postpartum period. Use a version validated for the intended population and follow the relevant clinical pathway. Perinatal screening also needs a plan for immediate assessment after an affirmative self-harm response and for prompt attention to possible postpartum psychosis, as described by ACOG.

Comparing results over time

Repeat the same instrument, version, timeframe, language, and scoring method when possible. A PHQ-9 total cannot be converted directly into a HAM-D, MADRS, CESD-R, or DASS-21 result. A fixed point difference also does not carry the same meaning across unrelated scales.

Before interpreting a change, check whether treatment, physical health, sleep, substance use, pregnancy, a major stressor, missing answers, or the administration method changed. The score can describe a pattern. It cannot identify which factor caused it.

Bottom line

Start with the clinical question, then choose the shortest verified tool that answers it. For most adult primary care screening, that is the PHQ-9 or a PHQ-2-to-PHQ-9 workflow. Use longer or clinician-rated scales only when the setting, exact edition, training, rights, and follow-up process support them. Whatever the tool, interpret the result as one part of an assessment, not a diagnosis or treatment decision.

Track your mental health

Create an account to explore published assessments, automatic scoring, and score history

View plans