LLMs are increasingly used to mediate patient communication, yet scalable evaluation of their safety, accuracy, and communication quality remains an open problem. LLM judges have emerged as automated evaluators, but whether they can holistically replicate human expert judgment is unvalidated. Informed consent for clinical trials presents a demanding case for such validation because it requires conveying complex information to lay audiences under ethical and safety constraints.
Validating LLM judges for automated oversight of patient communication
calendar_today
September 17, 2026
person
Xu, Z., Zeng, J., Zhou, S., Zhang, Z., Heintz, T., Tonneau, M., Yaghoubi, A., Ye, B., Goddla, V., Lehmann, L., Chen, Y.-H., Kozono, D., Revette, A., Maues, J., Brown, T., Catalano, P., Mak, R. H., Dligach, D., Bitterman, D.
domain
medrxiv