I'm always skeptical when it comes to self-reported measures.
There is no reason to be skeptical of well-designed, validated self-report measures. As has been alluded to, accurate measurement of depressive symptoms in medical trainees requires anonymity:
Levine RE, Breitkopf CR, Sierles FS, Camp G.
Complications associated with surveying medical student depression: The importance of anonymity.
Academic Psychiatry. 2003; 27:12–18.
http://link.springer.com/article/10.1176/appi.ap.27.1.12
Furthermore, high quality, well-validated survey instruments have high sensitivities and specificities for accurately diagnosing major depressive disorder.
Kroenke K, Spitzer RL, Williams JB.
The PHQ-9: Validity of a brief depression severity measure.
J Gen Intern Med. 2001; 16:606–613.
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC1495268/
The authors of both JAMA studies (i.e., this year's study on medical students and last year's study on resident physicians) provide tables of the sensitivities and specificities of the instruments synthesized in their analyses so that readers can judge for themselves which pooled estimates are the most reliable for diagnosing MDD. As you can see, well-validated instruments combined with proper cutoffs yield excellent results (e.g., the PHQ-9 with cutoff of >=10, the CES-D with cutoff >=16). Other instruments (e.g., the PRIME-MD) are really only useful for
screening patients and prevalence estimates obtained with them do not accurately reflect actual MDD prevalence, but rather only "depressive symptom" prevalence.
From an epidemiological standpoint, virtually all studies of depression are conducted using surveys, for a few reasons. (1) In-person interviews do not scale. It would be impossible to conduct 129,000 in-person interviews with medical students, for example. (2) Standardized surveys, which probe the exact same set of symptoms as would a psychiatrist following the DSM-V, do not vary from one participant to the next (i.e., there is no variability introduced by the interviewer-interviewee relationship). This reduces confounding. (3) The surveys produce objective, quantifiable data (i.e., a depressive-symptom score), allowing more refined statistical analyses (i.e., regression modeling) to be performed.