Skip to content

Navigation Menu

Sign in
Appearance settings

Search code, repositories, users, issues, pull requests...

Provide feedback

We read every piece of feedback, and take your input very seriously.

Saved searches

Use saved searches to filter your results more quickly

Appearance settings

zonghui0228/ChineseBioMedNLP-Challenges

Open more actions menu

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

15 Commits
15 Commits
 
 
 
 
 
 
 
 

Repository files navigation

中文  |  English 


中文生物医学健康信息处理评测挑战任务合集

本仓库全面总结了中国生物医学自然语言处理社区挑战评测竞赛的最新进展。该仓库聚焦于蓬勃发展的生物医学自然语言处理领域,该领域的重要性日益凸显,这主要得益于来自科学文献、电子病历、临床试验报告和社交媒体等来源的大量文本数据的积累。仓库收集了各种任务的信息,包括命名实体识别、实体规范化、属性抽取、关系抽取、事件抽取、文本分类、文本相似度、知识图谱构建、问答系统、文本生成以及大型语言模型评估。

如何引用

Hui Zong, Rongrong Wu, Jiaxue Cha, Weizhe Feng, Erman Wu, Jiakun Li, Aibin Shao, Liang Tao, Zuofeng Li, Buzhou Tang, Bairong Shen. Advancing Chinese biomedical text mining with community challenges. Journal of biomedical informatics, 2024;157:104716. PMID: 39197732. DOI:10.1016/j.jbi.2024.104716

概览

会议名称 年份 任务缩写 任务名称
CHIP 2025 CQC 面向住院电子病历入院记录的内涵质控任务
CHIP 2025 CDrugRed 面向中文电子病历的代谢性疾病出院用药推荐任务
CHIP 2025 MedCodeGPT 医学NLP代码自动生成测评-识别符合临床试验入组标准的患者
CHIP 2024 SDTTCM 中医辨证思维评测任务
CHIP 2024 LymphomaCoding 淋巴瘤信息抽取及肿瘤编码自动生成任务
CHIP 2024 TCDC 典型病历诊断一致性任务
CHIP 2023 PromptCBLUE CHIP-PromptCBLUE医疗大模型评测任务
CHIP 2023 NER 中文医学文本小样本命名实体识别评测任务
CHIP 2023 MedOCR 药品纸质文档识别与实体关系抽取任务
CHIP 2023 YIER-LLM CHIP-YIER医疗大模型评测任务
CHIP 2023 PICOS 医疗文献PICOS识别任务
CHIP 2023 DTC 中文糖尿病问题分类评测任务
CHIP 2022 AGAC 面向“基因-疾病”的关联语义挖掘任务
CHIP 2022 CMedCausal 医疗因果实体关系抽取任务
CHIP 2022 Text2DT 从医疗文本中抽取诊疗决策树
CHIP 2022 MedOCR 医疗纸质文档电子档(ePaper)OCR识别
CHIP 2022 CDN 临床诊断编码任务
CHIP 2021 MDCFNPC 医学对话临床发现阴阳性判别任务
CHIP 2021 CDEE 临床发现事件抽取任务
CHIP 2021 CDN 临床术语标准化任务
CHIP 2020 CMeEE 中文医学文本命名实体识别
CHIP 2020 CMeIE 中文医学文本实体关系抽取
CHIP 2020 CDN 临床术语标准化任务
CHIP 2020 Covid19 新冠肺炎趋势预测
CHIP 2020 TCM-QA 中医文献问题生成
CHIP 2020 TCM-NER 中药说明书实体识别
CHIP 2019 CDN 临床术语标准化任务
CHIP 2019 STS 平安医疗科技疾病问答迁移学习比赛
CHIP 2019 CTC 临床试验筛选标准短文本分类
CHIP 2018 CNER-AE 中文电子病历中临床医疗实体及属性抽取
CHIP 2018 STS 平安医疗科技智能患者健康咨询问句匹配大赛
CCKS 2023 PromptCBLUE PromptCBLUE医疗大模型评测
CCKS 2021 CNER-EE 面向中文电子病历的医疗实体及事件抽取
CCKS 2021 Phe-Drug-Mol 表型-药物-分子多层次知识图谱的链接预测
CCKS 2021 MedDG 蕴含实体的中文医疗对话生成
CCKS 2021 CMRC 面向中文医疗科普知识的内容理解
CCKS 2020 Covid19 新冠知识图谱构建与问答
CCKS 2020 CNER-EE 面向中文电子病历的医疗实体及事件抽取
CCKS 2019 CNER-AE 面向中文电子病历的医疗实体识别及属性提取
CCKS 2018 CNER 面向中文电子病历的命名实体识别
CCKS 2017 CNER 电子病历命名实体识别
DCIC 2021 CNER 智能医疗决策,病理“金数据”赋能医学诊断
CCL 2021 IMCS 智能医疗对话诊疗评测
CSMI 2020 PHQC 公众健康问句分类
CCIR 2019 EMR-Query 基于电子病历的数据查询类问答

评测任务

CHIP

  • 2025
    • 面向住院电子病历入院记录的内涵质控任务

      This evaluation centers on the core task of "medical record quality control," aiming to assess the model's capabilities in understanding, reasoning, and identifying issues within medical narrative documents. The task focuses on a common type of inpatient medical record-the Admission Note.

      Dataset size: training set: 1000, test-A set: 300, test-B set: 300

      link | leaderboard

    • 面向中文电子病历的代谢性疾病出院用药推荐任务

      This task constructs a dataset (CDrugRed) specifically for evaluating medication recommendations for metabolic diseases. The data is based on anonymized medical records from the Second Affiliated Hospital of Dalian Medical University, comprising 5,894 medical records from 3,190 patients. The data is divided into a training set, a development set (A-list test set), and a test set (B-list test set), where the development and test sets do not contain standard answers.

      Dataset size: training set: 1910, validation set: 320, test set: 960

      linkleaderboard

    • 医学NLP代码自动生成测评-识别符合临床试验入组标准的患者

      This evaluation task provided a total of 51 clinical trial enrollment/exit questions across 19 categories, with varying levels of difficulty. Each question included the question itself and its FHIR and FSH definitions. Participating teams were required to generate and submit NLP code according to the FHIR message bundle standards.

      Dataset size: training set: 51x17, test-A set: 51x17, test-B set: 51x17

      link | leaderboard

  • 2024
    • 中医辨证思维评测任务

      The dataset comprises 300 medical case records collected and processed by the task organization to establish a dedicated database. These medical cases were sourced from prominent public platforms such as the "Chinese Journal of Traditional Chinese Medicine" and the "Chinese Medicine Clinical Case Database", known for their high-quality and influential medical case data.

      Dataset size: training set: 200, validation set: 50, test set: 50

      link | leaderboard

    • 淋巴瘤信息抽取及肿瘤编码自动生成任务

      The dataset originates from published Chinese medical clinical case reports. Through a screening process, a total of 162 case reports related to lymphoma diseases have been gathered.

      Dataset size: training set: 54, validation set: 54, test set: 54

      linkleaderboard

    • 典型病历诊断一致性任务

      The dataset integrates diagnostic medical records of various common diseases, aiming to comprehensively and objectively evaluate the diagnostic capabilities of medical AI models by accurately replicating the decision-making process of physicians in diagnosing diseases.

      Dataset size: training set: 2590, test set: 647

      link | leaderboard

  • 2023
    • CHIP-PromptCBLUE医疗大模型评测任务

      The dataset is sourced from the CBLUE benchmark, encompassing 18 scenarios of medical natural language processing tasks, and includes over 450 instruction fine-tuning templates.

      Dataset size: training set: 87100, validation set: 8456, test set: 8456

      link | link_en | paper | github | leaderboard1 | leaderboard2

    • 中文医学文本小样本命名实体识别评测任务

      The dataset comprises manually annotated entities relevant to medical clinical contexts within medical texts. It includes 15 labels: item, sociology, disease, etiology, body, age, adjuvant, therapy, electroencephalogram, equipment, drug, procedure, treatment, microorganism, department, epidemiology, symptom, and others.

      Dataset size: training set: 400, validation set: 100, test set: 100

      link | link_en | paper | leaderboard

    • 药品纸质文档识别与实体关系抽取任务

      The drug leaflets were manually annotated for drugs, diseases, and clinical findings.

      Dataset size: training set: 400, validation set: 200, test set: 400

      link | link_en | paper | leaderboard

    • CHIP-YIER医疗大模型评测任务

      A series of multiple-choice questions constructed from medical entrance exam questions, clinical practice physician assessments, medical textbooks, medical literature/guidelines, and publicly available medical records. Dataset size: 1500.

      Dataset size: training set: 1000, test set: 500

      link | link_en | paper | leaderboard

    • 医疗文献PICOS识别任务

      The dataset consists of titles and abstracts from medical publications, annotated with five categories: Population (P), Intervention (I), Comparison (C), Outcome (O), and Study Types (S)

      Dataset size: training set: 2500, validation set: 1000, test set: 1000

      link | link_en | paper | leaderboard

    • 中文糖尿病问题分类评测任务

      The dataset consists of diabetes questions from internet, encompassing six categories: diagnosis, treatment, common knowledge, healthy lifestyle, epidemiology, and others.

      Dataset size: training set: 6000, validation set: 1000, test set: 1000

      link | link_en | paper | leaderboard

  • 2022
    • 面向“基因-疾病”的关联语义挖掘任务

      The AGAC corpus comprises 12 categories of molecular entities related to "gene-disease" associations and their triggering term entities: Var, MPA, Interaction, Pathway, CPA, Reg, PosReg, NegReg, Disease, Gene, Protein, and Enzyme. It includes semantic role annotations: ThemeOf and CauseOf; regulatory types: Loss of Function (LOF), Gain of Function (GOF), Regulation (REG), and Composite changes in function (COM).

      Dataset size: training set: 250, test set: 2000

      link | link_en | paper

    • 医疗因果实体关系抽取任务

      The annotated dialogue corpus includes three types of relationships: "causal", "conditional", and "hyponymy"

      Dataset size: training set: 2000, test set: 2000

      link | link_en | paper | leaderboard

    • 从医疗文本中抽取诊疗决策树

      The dataset consists of extracted diagnostic and therapeutic decision trees from clinical practice guidelines and medical textbooks. A diagnostic and therapeutic decision tree is defined as a binary tree composed of conditional nodes and decision nodes.

      Dataset size: training set: 300, validation set: 100, test set: 100

      link | link_en | paper

    • 医疗纸质文档电子档(ePaper)OCR识别

      The dataset consists of scanned images of paper medical records from the internet, defining 87 attributes to be extracted, which include types such as discharge summaries; outpatient invoices; pharmacy purchase invoices; and hospitalization invoices.

      Dataset size: training set: 1000, validation set: 200, test set: 500

      link | link_en | paper | leaderboard

    • 临床诊断编码任务

      The dataset comprises annotated information extracted from EHRs, including diagnostic details (such as admission diagnosis, preoperative diagnosis, postoperative diagnosis, and discharge diagnosis), as well as surgical names, medication names, and medical order names.

      Dataset size: training set: 2700, test set: 337

      link | link_en | paper

  • 2021
    • 医学对话临床发现阴阳性判别任务

      The data comes from publicly available internet-based telemedicine consultations includes patient chief complaints and physician diagnostic judgments, categorized into four attributes: negative, positive, other, and unspecified.

      Dataset size: training set: 6000, validation set: 2000, test set: 2000

      link | link_en | paper

    • 临床发现事件抽取任务

      The dataset extracted present medical history or imaging findings reports from EHRs, involving attributes across four dimensions: anatomical sites, main terms, descriptive terms, and occurrence status.

      Dataset size: training set: 2070, test set: 532

      link | link_en

    • 临床术语标准化任务

      The dataset comprises diagnostic entities, partial surgical entities, and standardized surgical relationship corpora extracted from Chinese EHRs. Dataset size: NA.

      Dataset size: training set: 9699, test set: 801

      link | link_en

  • 2020
    • 中文医学文本命名实体识别

      The dataset comprises 9 entity types extracted through medical text mining: diseases, clinical manifestations, medications, medical devices, medical procedures, anatomical structures, medical laboratory tests, microorganisms, and departments.

      Dataset size: training set: 15000, validation set: 5000, test set: 6618

      link | link_en | paper

    • 中文医学文本实体关系抽取

      The pediatric training corpus and the corpus extracted from a hundred common diseases yielded 53 schemas, comprising 10 synonymous relations and 43 other relations.

      Dataset size: training set: 17924, validation set: 4482, test set: 5602

      link | link_en | paper

    • 临床术语标准化任务

      The dataset includes diagnostic entities extracted from Chinese EHRs.

      Dataset size: training set: 8000, test set: 10000

      link | link_en

    • 新冠肺炎趋势预测

      The dataset comprises regional time-series data of confirmed COVID-19 cases, including daily counts of newly diagnosed cases.

      Dataset size: nan

      link | link_en

    • 中医文献问题生成

      Texts from the field of Traditional Chinese Medicine (TCM), including four TCM books and selected texts from TCM forums, with manually constructed question-answer pairs.

      Dataset size: training set: 3500, validation set: 750, test set: 750

      link | link_en

    • 中药说明书实体识别

      The dataset comprises 13 types of entities extracted from traditional Chinese medicine drug instructions: DRUG, DRUG_INGREDIENT, DISEASE, SYMPTOM, SYNDROME, DISEASE_GROUP, FOOD, FOOD_GROUP, PERSON_GROUP, DRUG_GROUP, DRUG_DOSAGE, DRUG_TASTE, and DRUG_EFFICACY.

      Dataset size: training set: 1200, validation set: 400, test set: 397

      link | link_en

  • 2019
    • 临床术语标准化任务

      The dataset consists of real surgical entities extracted from Chinese electronic medical records that require standardization.

      Dataset size: training set: 4000, validation set: 1000, test set: 2000

      link | link_en | paper

    • 平安医疗科技疾病问答迁移学习比赛

      The dataset comprises extracted online disease question-answer sentence pairs related to diabetes, hypertension, hepatitis, aids, and breast cancer.

      Dataset size: training set: 20000, validation set: 10000, test set: 50000

      link | link_en | paper | leaderboard

    • 临床试验筛选标准短文本分类

      Descriptive sentences of Chinese clinical trial inclusion/exclusion criteria, and a predefined set of 44 semantic categories for these criteria.

      Dataset size: training set: 22962, validation set: 7682, test set: 7697

      link | link_en | paper

  • 2018
    • 中文电子病历中临床医疗实体及属性抽取

      Imaging examination reports related to lung cancer and breast cancer.

      Dataset size: training set: 600, test set: 200

      link | link_en

    • 平安医疗科技智能患者健康咨询问句匹配大赛

      The dataset comes from real patient health consultation corpus. Given two sentences, it is required to determine whether the intentions are the same or similar.

      Dataset size: training set: 20000, test set: 10000

      link | link_en

CCKS

  • 2023
    • PromptCBLUE医疗大模型评测

      The dataset is sourced from the CBLUE benchmark, encompassing 16 scenarios of medical natural language processing tasks, and includes 94 instruction fine-tuning templates.

      Dataset size: training set: 68500, validation set: 10270, test set: 20540

      link | link_en | paper | github | leaderboard1 | leaderboard2

  • 2021
    • 面向中文电子病历的医疗实体及事件抽取

      The medical named entity recognition dataset consists of manually annotated plain text documents from EHRs, identifying medically relevant entities. It includes 6 predefined categories: diseases and diagnoses, examinations, tests, surgeries, medications, and anatomical locations. The medical event extraction dataset includes manually annotated plain text documents from EHRs, focusing on attribute entities related to primary entities of tumor events. It encompasses 3 categories: primary site, lesion size, and metastatic site.

      Dataset size: 2800 and 3000

      link | link_en | paper | leaderboard

    • 表型-药物-分子多层次知识图谱的链接预测

      A knowledge graph constructed from structured data sourced from reputable websites encompasses 7 types of relationships: associated_with, disease_mapped_to_gene, treats, targets, interacts_with, annotates, and pathway_has_gene_element.

      Dataset size: 80,000 entities, 1,200,000 triples.

      link | link_en | paper | leaderboard

    • 蕴含实体的中文医疗对话生成

      The MedDG dataset, annotated with entities, encompasses 12 types of gastroenterology-related diseases. Each dialogue is annotated with 160 relevant entities across 5 categories: diseases, symptoms, attributes, examinations, and medications. Dataset size: 20,611.

      Dataset size: training set: 17864, test set: 4347

      link | link_en | paper | leaderboard

    • 面向中文医疗科普知识的内容理解

      The dataset for reading comprehension of medical popular science knowledge includes the main content and a list of question-answer pairs, including question description, question ID, answer list. The dataset of recognizing irrelevant answers in medical popular science knowledge, the format of this dataset is one entry per line, with five columns, including Label, Docid, Question, Description, and Answer.

      Dataset size: 36000 and 55000

      link | link_en | paper | leaderboard

  • 2020
    • 新冠知识图谱构建与问答

      The COVID-19 knowledge graph encompasses seven entity types: virus, bacteria, disease, drug, medical specialty, examination subject, and symptom. The COVID-19 concept graph additionally includes type relationships between entities and concepts, as well as hierarchical relationships among concepts. The antiviral drug graph includes entities, entity attributes, and relationships between entities. The integrated dataset from the open-domain knowledge base PKUBASE and the OpenKG COVID-19 special topic includes information on entity category triples, hierarchical relationships between types, and predicates.

      Dataset size: This knowledge graph contains 66,499,920 triples, 25,574,536 entities, and 408,690 relations. A training set of 4,000 items, a validation set of 1,529 items, and a test set of 1,599 items

      link | link_en | paper | leaderboard1 | leaderboard2 | leaderboard3 | leaderboard4

    • 面向中文电子病历的医疗实体及事件抽取

      For medical named entity recognition dataset, it consists of manually annotated entities in EHRs, including diseases and diagnoses, examinations, tests, surgeries, medications, and anatomical locations. For medical event extraction dataset, it comprises manually annotated attribute entities associated with primary entities of oncology events in EHRs. It includes three categories: primary site, lesion size, and metastatic site.

      Dataset size: 1500 and 1400

      link | link_en | paper | leaderboard1 | leaderboard2

  • 2019
    • 面向中文电子病历的医疗实体识别及属性提取

      Medical Named Entity Recognition Dataset: The dataset comprises manually annotated documents from EHRs, capturing clinically relevant entities. These entities are categorized into 5 pre-defined categories: Symptoms and Signs, Examinations and Tests, Diseases and Diagnoses, Treatments, and Body Parts. Dataset size: 12,020. Attribute Extraction Dataset: This dataset consists of manually annotated documents from EHRs, focusing on attribute entities related to tumor events. These entities are categorized into 3 types: Lesion Size, Primary Site, and Metastatic Site. Dataset size: 2,000.

      Dataset size: 1379

      link | link_en | paper | leaderboard

  • 2018
    • 面向中文电子病历的命名实体识别

      The dataset consists of manually annotated entities from EHRs, including anatomical sites, symptom descriptions, independent symptoms, medications, and surgeries.

      Dataset size: 800

      link | link_en | paper | leaderboard

  • 2017
    • 电子病历命名实体识别

      The dataset consists of manually annotated entities from EHRs, including anatomical locations, symptom descriptions, independent symptoms, medications, and surgeries.

      Dataset size: 400

      link | link_en | leaderboard

DCIC

  • 2021
    • 智能医疗决策,病理“金数据”赋能医学诊断

      Extracted 10 types of entities from pathological text.

      Dataset size: training set: 1000, test set: 1050

      link | link_en | leaderboard

CCL

  • 2021
    • 智能医疗对话诊疗评测

      The dataset consists of dialogue cases from online medical consultation platforms, utilized for entity recognition, simulated dialogues, and disease diagnosis. Each sample in the dataset includes disease category, patient self-description text, symptoms, and entities and labels inferred from entire medical dialogues.

      Dataset size: >2000

      link | link_en

CSMI

  • 2020
    • 公众健康问句分类

      The dataset consists of public health queries categorized into six major themes: diagnosis, treatment, anatomy/physiology, epidemiology, healthy lifestyle, and choosing healthcare providers. Dataset size: 8000.

      Dataset size: training set: 5000, test set: 3000

      link | link_en | paper | leaderboard

CCIR

  • 2019
    • 基于电子病历的数据查询类问答

      The dataset consists of query-based question-answer pairs derived from EHRs. Dataset size: NA.

      Dataset size: training set: 1800, validation set: 600, test set: 600

      link | link_en | leaderboard

About

中文生物医学健康信息处理领域的评测任务

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Morty Proxy This is a proxified and sanitized view of the page, visit original site.