Using the HAM-D in clinical practice

The Hamilton Depression Rating Scale has been the standard depression measure for decades. Here's how to use it effectively in modern psychiatric practice and when other options might serve you better.

The Hamilton Depression Rating Scale (HAM-D, also called HDRS or HRSD) has been the dominant measure in depression research since Max Hamilton first published it in 1960. It remains the most widely used clinician-administered depression rating scale and the benchmark against which other measures are validated.

But the HAM-D was designed for research, not clinical practice. Using it effectively in modern psychiatric settings requires understanding both its strengths and its limitations, and knowing when other instruments might serve you better.

What the HAM-D measures

The HAM-D assesses depression severity through clinician ratings of patient symptoms. The original 17-item version covers mood symptoms (depressed mood, guilt, suicidal ideation, work and activities), anxiety symptoms (psychic and somatic anxiety), sleep disturbance (early, middle, and late insomnia), somatic symptoms (gastrointestinal, general, genital, weight loss), and other features (agitation, retardation, hypochondriasis, insight).

The scale uses a mixed scoring system: nine items use a 5-point scale (0-4) including depressed mood, guilt, suicide, work and activities, retardation, agitation, and both anxiety items. Eight items use a 3-point scale (0-2) covering the insomnia items, somatic symptoms, genital symptoms, weight loss, and insight. Total score range is 0-52 for the 17-item version.

Severity thresholds

ScoreSeverity
0-7Normal / clinical remission
8-13Mild depression
14-18Moderate depression
19-22Severe depression
>=23Very severe depression

These are later display conventions, not rules from Hamilton's original scale. Some trials define response as a 50% reduction and remission as 7 or below. State the convention rather than presenting it as a universal clinical definition.

HAM-D versions

HAM-D exists in several forms, including 6-, 17-, and 21-item versions and structured interview variants. They are not interchangeable. Survey Doctor currently implements a 17-item form and does not calculate a HAM-D6 score or administer GRID-HAM-D. Record the exact edition used in research or clinical documentation.

Administration

The HAM-D is clinician-rated. You score items based on clinical interview and observation, which differs fundamentally from self-report measures like PHQ-9.

Raters should use a documented training and structured interview process, then keep that process consistent when comparing scores.

Setting the time frame matters. Survey Doctor's implemented form assesses the past week. Be explicit about this with patients and record the edition and interval. Cover all items systematically rather than skipping those that seem irrelevant. Rate observation-based items from the interview and consider whether physical symptoms may have another cause.

Scoring and interpretation

Sum all 17 items for the total score. But understand what that score includes:

Heavy insomnia weighting: Three items assess insomnia. A patient with severe sleep disturbance but otherwise mild depression may score as moderately depressed.

Anxiety components: Psychic and somatic anxiety contribute substantially. Patients with anxious depression may score higher than their depressive symptoms alone would suggest.

Somatic emphasis: Multiple somatic items may inflate scores in patients with comorbid medical conditions.

For longitudinal use, record the absolute and percentage change, but do not treat either as proof that treatment caused the difference. Compare only ratings made with the same edition, interval, interview method, and trained-rater process.

Beyond totals, examine item-level patterns. Any elevation on the suicide item (item 3) requires clinical attention regardless of total score.

If you need help now. In the US, Call 988 or text 988. Call 911 if you are in immediate danger. Outside the US, contact your local emergency number or find support in your country.

Item-level ratings can show whether sleep, anxiety, retardation, or agitation contributed to the total. They do not select medication or another treatment by themselves.

Strengths and limitations

The HAM-D has a long research history and remains common in antidepressant trials. Regulators evaluate study endpoints in context; the scale is not an FDA-designated universal standard.

Clinician rating adds interview and observation, but it is not automatically more accurate than self-report. Rater expectations, interview structure, training, and patient communication all affect results.

However, the 20-30 minute administration time limits how often HAM-D can be used in busy practices. Despite thorough coverage of some domains, the scale underrepresents hopelessness and cognitive symptoms while overweighting neurovegetative symptoms. Designed in 1960, it reflects mid-20th-century understanding of depression.

The mixed scoring system (3-point and 5-point scales) complicates interpretation. Clinician-administered scales also introduce rater bias. Clinicians may rate patients they're treating as improving more than they actually are.

When to use HAM-D

HAM-D adds value in specific contexts:

Specialist evaluation: A clinician-rated symptom inventory can add structured information when a specialist is reviewing a complex course. It does not confirm treatment resistance or authorize a specific treatment.

Documenting symptom patterns: Item ratings can support clinical notes when symptoms or functioning change.

When observation matters: A clinician-rated scale can add observed psychomotor and interview behavior. It should complement, not dismiss, the patient's own report.

When to use alternatives

For routine monitoring, a self-report measure such as the PHQ-9 may be more practical. It answers a different question and should not be treated as interchangeable with HAM-D.

For measurement-based care programs, choose an instrument that matches the purpose, staffing, population, and follow-up process.

The MADRS is another clinician-rated depression scale with different item emphasis. For a self-report option, the CESD-R uses a different symptom-group algorithm. Choose based on the validated purpose rather than assuming one score converts to another.

If a particular domain is the clinical focus, like sleep, anxiety, or functioning, use domain-specific measures: HAM-A or GAD-7 for anxiety, dedicated sleep measures for insomnia.

Integrating HAM-D into practice

A service using HAM-D should define when the rating answers a useful question. Avoid a fixed universal schedule. If another instrument is used between HAM-D ratings, keep its trend separate.

If implementing HAM-D regularly, invest in formal training for all raters, establish reliability procedures with periodic rating comparisons, use structured interview versions, and monitor for rater drift over time. Record total scores, individual item scores, time frame assessed, and administration method to enable meaningful longitudinal tracking.

Track your mental health

Create an account to explore published assessments, automatic scoring, and score history

View plans