Clinical Assessment & AI

AI, Faculty, and Inter-Rater Reliability Challenges in Healthcare Simulation

Current evidence and what it means for the future of clinical assessment.

Synopsis

Grading clinical performance is inherently interpretive, and inter-rater reliability (IRR) varies depending on assessment format, rater training, and context. Research settings yield excellent IRR but only with sustained investment in faculty calibration that is rarely replicated in routine program operations. AI graders are highly consistent with themselves (intra-rater reliability) but remain imperfectly aligned with human expert judgment. That gap narrows with structured rubrics and prompt calibration. The path forward is not AI replacing human raters, but AI calibrated to human experts; handling volume and consistency while faculty oversight anchors high-stakes decisions.

Inter-Rater Reliability: Human Raters and AI Rater observing student rating

Grading clinical communication is ambiguous: did the learner elicit the patient's illness narrative and demonstrate empathic communication? Standardizing these assessments across multiple raters has challenged health professions education for decades. Now, with generative AI entering clinical grading, institutions face a fundamental question: how reliable are AI graders compared to human faculty?

This article argues the answer hinges on what is measured and how, mapping evidence across six related questions and situating it within an emerging AI benchmarking literature. The conclusion is that the near future is not AI versus humans — it is AI calibrated to humans.

1. What's the Best Achievable Inter-rater Reliability (IRR) in Research Settings?

When conditions are tightly controlled — trained raters, calibration sessions, video-based scoring — inter-rater reliability in healthcare simulation can be excellent, with studies reporting agreement above 0.84–0.90 across rater pairings [1]. Research using generalizability theory reinforces this, showing that rater variance can be minimized with proper training and enough cases [2]. These highly controlled situations require substantial infrastructure investment and are unlikely to reflect what most programs experience day to day.

2. How Much Effort Does That Require?

Substantial, ongoing effort. Rater training works by optimizing standardized tool use and reducing variation in individual preconceptions, but its effects are not permanent. Rater drift — the tendency of raters to shift their scoring standards over time, typically toward increasing leniency due to fatigue, familiarity, or reduced attention to scoring anchors — occurs without ongoing recalibration [3,4]. In one controlled OSCE study, Cohen's κ improved from 0.42 to 0.78 following structured faculty workshops [5]. In a separate study, repeated structured rater discussions aimed at reaching scoring consensus improved kappa from 0.49 to 0.82 across 40 scored videos [6].

In considering alternate raters, research suggests appropriately trained standardized patient actors (SPs) achieve inter-rater reliability comparable to or better than explicitly trained faculty for behavioral checklist items [7]. However, for subtler judgments of clinical competency — particularly using rubrics, global scales, or competency frameworks — faculty clinical expertise remains important. Park et al. (2016) demonstrated that trained faculty raters using a locally developed scoring rubric for patient notes achieved reproducible scores with acceptable reliability but required significant faculty training [8]. Bond et al. (2023) found that natural language processing (NLP)-based automated scoring of post-encounter patient notes approached faculty-faculty reliability at substantially lower cost, with faculty effort focused on initial machine scoring calibration [9].

3. Is It Fair to Compare IRR Across Checklists, Rubrics, and Competency Scales?

No. Assessment formats differ fundamentally in what they ask raters to do, yielding meaningfully different IRR profiles as a result.

A systematic review by Ilgen et al. (2015) analyzed 45 simulation-based assessment studies and found that pooled IRR was similar for global rating scales (GRS: ICC 0.78, 95% CI 0.71–0.83) and checklists (ICC 0.81, 95% CI 0.75–0.85) [10]. This may seem counterintuitive — the apparent objectivity of checklists does not yield markedly higher IRR than holistic global ratings. The explanation lies in what each scale measures: GRS showed higher inter-item reliability (0.92 vs. 0.66) and inter-station reliability (0.80 vs. 0.69) than checklists, meaning global ratings generalize better across cases and items, while each checklist is task-specific by design [10].

Analytic rubrics that combine structured criteria with behavioral anchors occupy a middle ground. Park et al. (2016) found Cronbach's α = 0.65–0.77 per case for trained faculty raters and patient note rubric [8]. Hassanein et al. (2026), studying dental student clinical short-answer responses using a 12-point rubric, found that three calibrated expert raters achieved ICC = 0.84 [11] — toward the higher end of what analytic rubrics typically yield for IRR.

The Dreyfus model of skill acquisition occupies a different level of abstraction. Rather than scoring discrete behaviors, raters must locate a learner on a novice-to-expert continuum — a judgment requiring integrative expertise. In one Group OSCE (GOSCE) study setting with trained SPs, Dreyfus-scale kappa values ranged from 0.61 to 0.72 across four years of observation [12].

EPA entrustment decisions also sit at the most demanding end of the spectrum. Rather than 'what did they do?', raters must answer: 'would you trust them to do this unsupervised?' A generalizability study of an EPA-based OSCE at the transition to internship found that a moderately reliable score (G=0.76) required 12 cases, more than some OSCE programs typically deploy, and that faculty raters were more reliable than SP raters for this task [2]. Real-world entrustment IRR data outside of structured pilots remains sparse.

Table 1 summarizes the IRR evidence by format, including available AI benchmarks. Note that routine production environment educational IRR data is sparse and rarely collected. Thus, estimates are from the pre-rater training state of existing studies.

Table 1. Inter-Rater Reliability (IRR) by Assessment Format and Setting

Table 1: Inter-Rater Reliability (IRR) by Assessment Format and Setting.
Acronym Legend: acc. = accuracy; EPA = entrustable professional activity; G = generalizability coefficient; GOSCE = group objective structured clinical examination; GRS = global rating scale; ICC = intraclass correlation coefficient; κ = Cohen's kappa; LLM = large language model; SME = subject matter expert.

4. What's the Likely IRR in Real Production Education Settings?

Considerably lower. Research standards for rater training are rarely sustained in production environments. Calibration is often a one-time event, and IRR is seldom monitored in normal operations. The research-to-practice gap is real and underacknowledged.

Within medical education, Yoo and Yoo (2024) found that untrained faculty raters using a portfolio assessment tool in a medical school Introduction to Clinical Medicine course achieved ICC = 0.38, improving modestly to 0.44 after excluding extreme outlier raters — well below the thresholds observed in research settings [13]. Rater drift without ongoing recalibration is well documented across both clinical and educational assessment contexts [3,4].

This matters for interpreting AI benchmarking studies: when AI is compared to 'expert consensus,' the comparison group is research-quality human performance. The more relevant comparison for many programs is everyday faculty grading — and AI may already meet or exceed that standard in certain assessment formats.

5. How Do AI Agents Fare? Intra-Rater and Inter-Rater Reliability

Busch et al. (2025) [14] offers the clearest clinical education benchmark to date. Four large language models (LLMs) — GPT-4o, Claude 3.5, Llama 3.1, and Gemini 1.5 Pro — were scored against expert consensus on all 28 items of the Master Interview Rating Scale (MIRS) across 10 objective structured clinical examination (OSCE) cases, using varied methods to train the AI. MIRS is an anchored rubric: for example, Item 3 ('Negotiates priorities and sets agenda') is scored 1–5, with a score of 5 defined as 'The interviewer fully negotiates priorities of patient concerns, listing all of the concerns and sets the agenda at the onset of the interview. The patient is invited to participate in making an agreed plan' [14].

The key finding: AI was highly self-consistent (intra-rater reliability) but inter-rater accuracy against expert consensus remained modest: exact accuracy 0.27–0.44, off-by-one accuracy 0.67–0.87 [14]. AI scored reliably with itself; it just was not scoring exactly like human experts. Table 2 summarizes additional references, some of which shared exact match, off-by-one match, or other markers of reliability.

Table 2. Selected AI vs. Human IRR Studies in Clinical and Educational Assessment

Table 2: Selected AI vs. Human IRR Studies in Clinical and Educational Assessment.
Acronym Legend: ASAG = automated short answer grading; CoT = chain-of-thought; COM = College of Medicine; EFL = English as a Foreign Language; ICC = intraclass correlation coefficient; LLM = large language model; MIRS = Master Interview Rating Scale; NLP = natural language processing; OSCE = objective structured clinical examination; SME = subject matter expert; SP = standardized patient; temp = temperature parameter; UME = undergraduate medical education.

The pattern that emerges: AI reliability is highly format- and calibration-dependent, paralleling what we observe with human raters. Performance on simpler, more structured tasks is more readily approximated by AI than holistic clinical communication assessment.

6. Calibrating AI to Faculty: Human-in-the-Loop and the Path Forward

The current state of the art for closing the AI-to-human gap combines three levers: prompt engineering (i.e., better instructions for the AI), few-shot calibration, and selective human re-review. When using off-the-shelf commercial LLMs, prompt engineering and retrieval-augmented generation (RAG) — providing the AI model with relevant reference materials — are the primary calibration tools available. In contrast, fine tuning the AI model itself requires access to model weights that most programs will not have.

Providing the AI model with detailed rubric criteria, worked examples of strong and weak performances, and explicit behavioral anchors substantially improves alignment with human raters [14]. This mirrors frame-of-reference training in human rater calibration, which is considered the most effective approach for improving human IRR [18].

The SURE (Selective Uncertainty-based Re-Evaluation) pipeline — where repeated LLM prompting identifies low-certainty cases flagged for human re-review — reduced manual grading time by 40–90% while maintaining accuracy [19]. This selective human-in-the-loop model represents one practical option, and this area remains ripe for exploration.

How long will human-in-the-loop be needed? Until AI can reliably replicate expert-level holistic judgment — particularly for Dreyfus-scale placements and EPA entrustment decisions — some form of human oversight will remain essential for high-stakes assessment. For lower-stakes formative assessment using checklists and structured rubrics, AI-assisted grading with appropriate calibration protocols may already be appropriate in many programs.

Conclusion

The question is not whether AI can grade clinical performance. In certain formats and with appropriate calibration, it can — sometimes more consistently than untrained human raters in production settings. The harder questions are: consistent with what, calibrated how, and for which decisions?

The IRR literature tells us that reliability is a function of format, training, and context — for both humans and AI. The path forward is not replacement but a structured partnership: AI handles volume and consistency; humans provide the expert judgment that anchors the system. The goal is not merely to ask whether AI reaches the research-quality human ceiling. It is to ensure that whatever grading system we deploy — human, AI, or hybrid — meets the standard the learner and the future patient deserves.

References

  1. Malau-Aduli BS, Mulcahy S, Warnecke E, et al. Inter-rater reliability: comparison of checklist and global scoring for OSCEs. Creative Education. 2012;3(6):937-942. doi:10.4236/ce.2012.326142
  2. Suneja M, DuChene Hanrahan K, Kreiter C, Rowat J. Psychometric properties of entrustable professional activity–based objective structured clinical examinations during transition from undergraduate to graduate medical education: a generalizability study. Acad Med. 2025;100(2):179-183. doi:10.1097/ACM.0000000000005719
  3. McLaughlin K, Coderre S, Woloschuk W, Leverett J, Wright BJ. The effect of differential rater function over time (DRIFT) on objective structured clinical examination ratings. Med Educ. 2009;43(10):989-994. doi:10.1111/j.1365-2923.2009.03437.x
  4. Harik P, Clauser BE, Grabovsky I, Nungester RJ, Swanson D, Nandakumar R. An examination of rater drift within a generalizability theory framework. J Educ Meas. 2009;46(1):43-58. doi:10.1111/j.1745-3984.2009.01067.x
  5. Kannummal Veetil S, Haque PD, Jain D, Garg M, Rajoria S, Pramanik BK. Assessing the feasibility and acceptability of the objective structured clinical examination (OSCE) in undergraduate surgical training: a pilot study. Cureus. 2025;17(9):e92992. doi:10.7759/cureus.92992
  6. Sakurai H, Yamamoto N, Ohta R, et al. Effect of moderation on rubric criteria for inter-rater reliability in an objective structured clinical examination with real patients. J Educ Eval Health Prof. 2022;19:20. doi:10.3352/jeehp.2022.19.20
  7. Adamo G, Vu NV, Clyman SG, et al. Simulated and standardized patients in OSCEs: achievements and challenges 1992-2003. Med Teach. 2003;25(3):262-270. doi:10.1080/0142159031000100300
  8. Park YS, Hyderi A, Bordage G, Xing K, Yudkowsky R. Inter-rater reliability and generalizability of patient note scores using a scoring rubric based on the USMLE Step-2 CS format. Adv Health Sci Educ Theory Pract. 2016;21(4):761-773. doi:10.1007/s10459-015-9664-3
  9. Bond WF, Zhou J, Bhat S, Park YS, Ebert-Allen RA, Ruger RL, Yudkowsky R. Automated patient note grading: examining scoring reliability and feasibility. Simul Healthc. 2023. doi:10.1097/SIH.0000000000000757
  10. Ilgen JS, Ma IWY, Hatala R, Cook DA. A systematic review of validity evidence for checklists versus global rating scales in simulation-based assessment. Med Educ. 2015;49(2):161-173. doi:10.1111/medu.12621
  11. Hassanein FEA, Hussein RR, Ahmed DE, et al. Calibration of AI large language models with human subject matter experts for grading of clinical short-answer responses in dental education. BMC Oral Health. 2026;26:286. doi:10.1186/s12903-026-07665-4
  12. Chen YC, Liu YC, Tsai YJ, et al. Dreyfus scale-based feedback increased medical students' satisfaction with the complex cluster part of an interviewing and physical examination course and improved skills readiness. J Educ Eval Health Prof. 2019;16:30. doi:10.3352/jeehp.2019.16.30
  13. Yoo DM, Yoo JH. Inter-rater reliability and content validity of the measurement tool for portfolio assessments used in the Introduction to Clinical Medicine course at Ewha Womans University College of Medicine: a methodological study. J Educ Eval Health Prof. 2024;21:33. doi:10.3352/jeehp.2024.21.33
  14. Busch F, Hoffmann L, Makowski MR, et al. Benchmarking generative AI for scoring medical student interviews in objective structured clinical examinations (OSCEs) [preprint]. arXiv. 2025. arXiv:2501.13957v2
  15. Yokose M, Hirosawa T, Sakamoto T, et al. The validity of generative artificial intelligence in evaluating medical students in objective structured clinical examination: experimental study. JMIR Med Educ. 2025;11:e79465. doi:10.2196/79465
  16. Sreedhar R, Chang L, Gangopadhyaya A, Shiels PW, Loza J, Chi E, Gabel E, Park YS. Comparing scoring consistency of large language models with faculty for formative assessments in medical education. J Gen Intern Med. 2024. doi:10.1007/s11606-024-09050-9
  17. Yavuz F, Çelik Ö, Yavaş Çelik G. Utilizing large language models for EFL essay grading: an examination of reliability and validity in rubric-based assessments. Br J Educ Technol. 2025;56(1):150-166. doi:10.1111/bjet.13494
  18. Vergis A, Leung C, Robertson R. Rater training in medical education: a scoping review. Cureus. 2020;12(11):e11363. doi:10.7759/cureus.11363
  19. Korthals L, Akrong E, Geller G, Rosenbusch H, Grasman R, Visser I. Towards reliable LLM grading through self-consistency and selective human review: higher accuracy, less work. Mach Learn Knowl Extr. 2026;8(3):74. doi:10.3390/make8030074