CLINIC: Evaluating Multilingual Trustworthiness in Language Models for Healthcare
2025/12/12 by Ghosh, Akash, Sridhar, Srivarshinee, Ravi, Raghav Kaushik +3
Medicine · Social Sciences · #Artificial Intelligence in Healthcare and Education #Ethics and Social Impacts of AI #Global Health and Surgery
paper · doi:10.48550/arxiv.2512.11437
Abstract
Integrating language models (LMs) in healthcare systems holds great promise for improving medical workflows and decision-making. However, a critical barrier to their real-world adoption is the lack of reliable evaluation of their trustworthiness, especially in multilingual healthcare settings. Existing LMs are predominantly trained in high-resource languages, making them ill-equipped to handle the complexity and diversity of healthcare queries in mid- and low-resource languages, posing significant challenges for deploying them in global healthcare contexts where linguistic diversity is key. In this work, we present CLINIC, a Comprehensive Multilingual Benchmark to evaluate the trustworthiness of language models in healthcare. CLINIC systematically benchmarks LMs across five key dimensions of trustworthiness: truthfulness, fairness, safety, robustness, and privacy, operationalized through 18 diverse tasks, spanning 15 languages (covering all the major continents), and encompassing a wide array of critical healthcare topics like disease conditions, preventive actions, diagnostic tests, treatments, surgeries, and medications. Our extensive evaluation reveals that LMs struggle with factual correctness, demonstrate bias across demographic and linguistic groups, and are susceptible to privacy breaches and adversarial attacks. By highlighting these shortcomings, CLINIC lays the foundation for enhancing the global reach and safety of LMs in healthcare across diverse languages.
Citations
- Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
- DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
- Infogen: Generating Complex Statistical Infographics from Documents
- Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples
- SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models' Knowledge of Indian Culture
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
- Gemma 3 Technical Report
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
- CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis
- Adversarial Attacks on Large Language Models in Medicine
- OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow Understanding
- MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models
- CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models
- CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
- Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents
- Capabilities of Gemini Models in Medicine
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Apollo: A Lightweight Multilingual Medical LLM towards Democratizing Medical AI to 6B People
- MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
- How do Large Language Models Handle Multilingualism?
- Determinants of LLM-assisted Decision-Making
- Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models
- Towards Building Multilingual Language Model for Medicine
- M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
- MedSumm: A Multimodal Approach to Summarizing Code-Mixed Hindi-English Clinical Queries
- CLIPSyntel: CLIP and LLM Synergy for Multimodal Question Summarization in Healthcare
- PromptBench: A Unified Library for Evaluation of Large Language Models
- Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine
- NurViD: A Large Expert-Level Video Database for Nursing Procedure Activity Understanding
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- Jailbreaking Black Box Large Language Models in Twenty Queries
- SafetyBench: Evaluating the Safety of Large Language Models
- Towards Generalist Biomedical AI
- CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility
- A Comprehensive Overview of Large Language Models
- DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- HuatuoGPT, towards Taming Language Model to Be a Doctor
- Multilingual Pixel Representations for Translation and Effective Cross-lingual Transfer
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
- Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
- MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data
- HuaTuo: Tuning LLaMA Model with Chinese Medical Knowledge
- LLaMA: Open and Efficient Foundation Language Models
- Large Language Models Encode Clinical Knowledge
- GLUE-X: Evaluating Natural Language Understanding Models from an Out-of-distribution Generalization Perspective
- GreenPLM: Cross-Lingual Transfer of Monolingual Pre-Trained Language Models at Almost No Cost
- Red Teaming Language Models with Language Models
- The FLORES-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation
- Compositional Explanations of Neurons
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization
Related