浙江农业科学 ›› 2026, Vol. 67 ›› Issue (9): 2205-2210.DOI: 10.16178/j.issn.0528-9017.20260443

• 果树与蔬菜 • 上一篇    下一篇

低资源场景下甜菜育种实体识别及知识图谱应用研究

张丽1(), 陈俊杰1, 王小宇2   

  1. 1.内蒙古农业大学 计算机与信息工程学院,内蒙古 呼和浩特 010041
    2.国网内蒙古东部电力有限公司,内蒙古 呼和浩特 010041
  • 收稿日期:2026-06-04 出版日期:2026-09-11 发布日期:2026-09-24
  • 作者简介:张丽,研究方向为知识图谱、软件工程和农业信息化。E-mail:Li4516161@163.com。
  • 基金资助:
    内蒙古自治区高等学校科学技术研究项目(NJZY22523)

Study on entity recognition and knowledge graph application for beet breeding in low-resource scenarios

ZHANG Li1(), CHEN Junjie1, WANG Xiaoyu2   

  1. 1.College of Computer and Information Engineering,Inner Mongolia Agricultural University,Hohhot 010041,Inner Mongolia
    2.State Grid Inner Mongolia East Electric Power Co. ,Ltd. ,Hohhot 010041,Inner Mongolia
  • Received:2026-06-04 Online:2026-09-11 Published:2026-09-24

摘要:

甜菜是我国重要的糖料作物,其育种文献中蕴含着丰富的品种、性状、基因等结构化知识,但标注数据匮乏、术语复杂,导致现有实体识别方法性能受限。为解决该问题,本研究构建了国内首套甜菜育种命名实体识别(NER)专用数据集,定义了8类核心实体,通过少量人工标注与规则模板增强相结合的策略,有效扩充训练样本。基于该数据集,BERT-BiLSTM-CRF模型F1值达到了90.19%,优于基线模型。为验证所提方法的实用性,用该模型对502篇文献进行实体识别,构建了以文献为核心、支持溯源的知识图谱,并通过结构化检索任务检验图谱的覆盖度与检索准确性。试验表明,基于图谱的规范化检索准确率达90.8%,验证了知识图谱在甜菜育种领域的可用性,为低资源场景实体识别与知识组织提供了有效参考。

关键词: 甜菜育种, 命名实体识别, BERT-BiLSTM-CRF模型, 知识图谱, 低资源文本挖掘

Abstract:

Beet is an important sugar crop in China,and its breeding literatures contain rich structured knowledge such as varieties,traits,and genes. However,the scarcity of annotated data and the complexity of terminology limit the performance of existing entity recognition methods. To address this problem,we constructed the first dedicated dataset for named entity recognition(NER)in beet breeding in China,defining eight core entity types. Through a strategy combining small-scale manual annotation with rule-based template augmentation,the strategy effectively expanded the training samples. Using this dataset,the BERT-BiLSTM-CRF model achieved an F1 value of 90.19%,outperforming baseline model. To validate the practicality of the proposed method,we applied the model to recognize entities in 502 literatures and built a knowledge graph that centered on the literature and supports traceability. The graph's coverage and retrieval accuracy were evaluated through structured-retrieval tasks. Experimental results showed that the normalized retrieval accuracy based on the graph reached 90.8%,confirming the usability of the knowledge graph in the domain of beet breeding. This study provides an effective methodological reference for entity recognition and knowledge organization in low-resource scenarios.

Key words: beet breeding, named entity recognition, BERT-BiLSTM-CRF model, knowledge graph, low-resource text mining

中图分类号: