Thoughts about AI use in Residency Application Evaluation?

This forum made possible through the generous support of SDN members, donors, and sponsors. Thank you.
Get help with your application

Use all the free resources available to you from SDN: articles, guides, expert advising, forums discussions, and school research.

collegestud2013

Full Member
15+ Year Member
Advertisement - Members don't see this ad
With the recent expansion of AI tools into ERAS workflows and application screening, I wanted to start a discussion on something that seems to be moving very quickly without a clear consensus on risks or guardrails.


Two recent articles highlight both the promise and potential pitfalls:



A few points from these that stood out to me:




1) AI is already being used – but evidence is limited​


  • Most studies focus on predicting interview offers or rank lists
  • Very few actually evaluate fairness or bias rigorously

At the same time:


  • Application volume continues to rise
  • Programs are under pressure to triage efficiently

So AI is filling a real operational need, even if the science is still early




2) Bias may not be reduced… and could be amplified​


  • Most studies acknowledge bias, but only a minority actually measure it
  • AI systems can replicate:
    • Historical biases in training data
    • Structural inequities already present in selection processes

More concerning:


  • Bias can be hidden inside black-box models
  • And once scaled, it affects every applicant simultaneously



3) AI tools for evaluating applications may be unreliable​


  • AI detection tools (for personal statements, etc.) show:
    • High false positives
    • Difficulty distinguishing mixed vs human-written content

Raises real concerns about:


  • Penalizing applicants incorrectly
  • Disadvantaging non-native English speakers



4) Legal risk is not theoretical anymore​


The JAMA article outlines several potential legal issues:


  • Disparate impact discrimination (even without intent)
  • Residency selection increasingly viewed as employment, not just education
  • Programs (not vendors) likely bear most liability
  • Lack of transparency and explainability could become a major issue

This starts to look very similar to lawsuits already happening in corporate AI hiring tools




5) Bigger question: what is “acceptable use”?​


Some open questions I’m curious how people here think about:

  • Should AI be limited to data extraction / summarization only?
  • Is using AI to score “fit” or personality ever appropriate?
  • Should applicants be able to see and challenge AI-generated evaluations?
  • At what point does this cross from “efficiency tool” → “automated decision-making”?



6) Parallels to social media professionalism issues?​


We’ve seen how quickly professionalism standards can be enforced in other domains (recent discussions here on social media conduct about Mayo med student).

AI may be similar:
  • Rapid adoption
  • Vague policies
  • High-stakes consequences


​


AI in residency selection seems inevitable, but we may be:


  • Underestimating bias + legal exposure
  • Overestimating current model reliability

Feels like we’re in the “early adoption” phase before:


  • Standardization
  • Regulation
  • Or litigation forces change



Would be especially interested to hear from:


  • PDs / faculty using or evaluating these tools
  • Residents who went through recent cycles
  • Anyone with legal or admin perspective

Thoughts?
 
It certainly is fascinating.

AI tool development is insanely fast with the current LLM's. I had thought that to get anything "real" done, you needed to build your own model (or "tune" a current model to do what you want). But it turns out they are pretty good if you just give them the data you have, and ask them to do something with it.

AI models are black box by design. There's no way to know how they get to the answers they generate. That's a feature, not a bug.

AI may have biases, but ? if those are better or worse than human biases. And bias can perhaps be measured (i.e. if the model is suggesting whom to interview, the demographics of those selections can be compared to the overall demographics to see if there are impacts). This is imperfect, but it might be better than trying to address human generated bias.

Not mentioned in these articles is where data analysis might go with all of this. Medical schools often try to obfuscate student performance - some are better than others. But if one company (cough - cough - Thalamus - cough) has all of the MSPE's / Transcripts for a single school, it's child's play now to extract true class rankings or comparative performance.

And once AI is established in residency application review, the arms race really begins. People will try to study what types of content in letters or elsewhere will push the AI to think the candidate is better. I would not be surprised if we start to see hidden messages (text in white on a white background) on documents designed to tell the AI what to do. Although maybe someone who is writing all of this will tell the AI to consider that a negative factor. And so on.

A parallel can be made (somewhat weakly) to voting. Paper ballots that are hand counted seem like the most safe option. But hand counting is very error prone. Electronic voting probably makes less errors -- but oftentimes you have no idea where those errors might be. There was a situation not long ago where the automatic ballot reader failed because of where the ballot fold happened to be. But the machine counted the ballots and said all was fine -- there was no way to know it was off. A truly fully electronic system might avoid that, but then risks someone hacking and changing votes. But boy, electronic counting is so much faster, cheaper, and easier. It's here to stay for sure -- and that's a good thing.