中国药物警戒 ›› 2026, Vol. 23 ›› Issue (9): 1017-1022.
DOI: 10.19803/j.1672-8629.20260385

• 法规与管理研究 • 上一篇    下一篇

基于大语言模型应用的个例安全性报告数据清洗方法探讨

刘红亮1, 宋海波1, 米岚2, 孟康康1, 王涛1, 漆燕1, 贾晋生1,*   

  1. 1国家药品监督管理局药品评价中心,国家药品监督管理局监管科学创新研究基地,北京 100163;
    2北京大学肿瘤医院暨北京市肿瘤防治研究所,恶性肿瘤发病机制及转化研究教育部重点实验室,北京 100142
  • 收稿日期:2026-05-15 出版日期:2026-09-15 发布日期:2026-09-15
  • 通讯作者: *贾晋生,男,工程师,药品不良反应监测。E-mail: jiajinsheng@cdr-adr.org.cn
  • 作者简介:刘红亮,男,硕士,主管药师,药品不良反应监测。
  • 基金资助:
    北京市自然科学基金资助项目(L254089); 国家自然科学基金资助项目(72274193)

Data cleaning methods for individual case safety reports based on large language model applications

Liu Hongliang1, Song Haibo1, Mi Lan2, Meng Kangkang1, Wang Tao1, Qi Yan1, Jia Jinsheng1,*   

  1. 1Center for Drug Reevaluation, NMPA/NMPA Center for Innovation and Research in Regulatory Science, Beijing 100163, China;
    2Key Laboratory of Carcinogenesis and Translational Research (Ministry of Education/Beijing), Peking University Cancer Hospital & Institute, Beijing 100142, China
  • Received:2026-05-15 Online:2026-09-15 Published:2026-09-15

摘要: 目的 探讨大语言模型在个例安全性报告数据清洗中的应用方法。方法 结合数据清洗要求,通过以标准知识库为基础、以规则引擎为前置过滤、以检索增强生成实现候选召回、以大语言模型完成语义推理与重排序、以人工审核和日志反馈形成闭环的技术框架,围绕基层报告单位填报提示、监测机构审核、不良反应过程描述内容的术语提取和模型持续迭代优化等场景提出实施建议。结果 通过构建的技术框架,大语言模型能够在一定程度上提升个例安全性报告数据标准化水平和清洗效率。结论 大语言模型可作为数据清洗辅助能力嵌入可解释、可追溯、可审计的监督流程,但鉴于其固有属性,不宜替代专业审核。在实施阶段,应通过分阶段试点、基准样本集评价和安全合规清洗,逐步验证其在提升数据标准化水平和降低人工负担方面的应用价值。

关键词: 大语言模型, 药品不良反应, 个例安全性报告, 数据清洗, 检索增强生成

Abstract: Objective To explore the applications of large language models (LLMs) in data cleaning for individual case safety reports (ICSRs). Methods A technical framework was developed by integrating the characteristics of data cleaning. This framework was built on a standardized knowledge base, employed a rule engine as a pre-filtering mechanism, used retrieval-augmented generation (RAG) for candidate recall, leveraged an LLM for semantic reasoning and re-ranking, and established a closed-loop system through manual review and log-based feedback. Recommendations regarding implementation were accessible for such scenarios as giving tips for reporting by grassroots institutions, reviews by regulators, term extraction from ADR narrative descriptions, and continuous iteration and optimization of models. Results The constructed framework proved that LLMs could technologically improve both the standardization of data and cleaning efficiency of ICSRs. Conclusion LLMs can be embedded as an auxiliary data cleaning tool within explainable, traceable, and auditable processes of oversight. However, due to their inherent limitations, they should not replace professional human review. During actual use, their value in enhancing data cleaning and reducing manual workload should be progressively validated through phased and pilot programs, evaluation of benchmark datasets , and compliance-driven cleaning.

Key words: Large Language Model (LLM), Adverse Drug Reaction, Individual Case Safety Reports (ICSRs), Data Cleaning, Retrieval-Augmented Generation

中图分类号: