Chinese Journal of Pharmacovigilance ›› 2026, Vol. 23 ›› Issue (9): 1017-1022.
DOI: 10.19803/j.1672-8629.20260385

Previous Articles     Next Articles

Data cleaning methods for individual case safety reports based on large language model applications

Liu Hongliang1, Song Haibo1, Mi Lan2, Meng Kangkang1, Wang Tao1, Qi Yan1, Jia Jinsheng1,*   

  1. 1Center for Drug Reevaluation, NMPA/NMPA Center for Innovation and Research in Regulatory Science, Beijing 100163, China;
    2Key Laboratory of Carcinogenesis and Translational Research (Ministry of Education/Beijing), Peking University Cancer Hospital & Institute, Beijing 100142, China
  • Received:2026-05-15 Online:2026-09-15 Published:2026-09-15

Abstract: Objective To explore the applications of large language models (LLMs) in data cleaning for individual case safety reports (ICSRs). Methods A technical framework was developed by integrating the characteristics of data cleaning. This framework was built on a standardized knowledge base, employed a rule engine as a pre-filtering mechanism, used retrieval-augmented generation (RAG) for candidate recall, leveraged an LLM for semantic reasoning and re-ranking, and established a closed-loop system through manual review and log-based feedback. Recommendations regarding implementation were accessible for such scenarios as giving tips for reporting by grassroots institutions, reviews by regulators, term extraction from ADR narrative descriptions, and continuous iteration and optimization of models. Results The constructed framework proved that LLMs could technologically improve both the standardization of data and cleaning efficiency of ICSRs. Conclusion LLMs can be embedded as an auxiliary data cleaning tool within explainable, traceable, and auditable processes of oversight. However, due to their inherent limitations, they should not replace professional human review. During actual use, their value in enhancing data cleaning and reducing manual workload should be progressively validated through phased and pilot programs, evaluation of benchmark datasets , and compliance-driven cleaning.

Key words: Large Language Model (LLM), Adverse Drug Reaction, Individual Case Safety Reports (ICSRs), Data Cleaning, Retrieval-Augmented Generation

CLC Number: