Two learning approaches for protein name extraction
Journal of Biomedical Informatics
1046 - 1055
Item Usage Stats
Protein name extraction, one of the basic tasks in automatic extraction of information from biological texts, remains challenging. In this paper, we explore the use of two different machine learning techniques and present the results of the conducted experiments. In the first method, Bigram language model is used to extract protein names. In the latter, we use an automatic rule learning method that can identify protein names located in the biological texts. In both cases, we generalize protein names by using hierarchically categorized syntactic token types. We conducted our experiments on two different datasets. Our first method based on Bigram language model achieved an F-score of 67.7% on the YAPEX dataset and 66.8% on the GENIA corpus. The developed rule learning method obtained 61.8% F-score value on the YAPEX dataset and 61.0% on the GENIA corpus. The results of the comparative experiments demonstrate that both techniques are applicable to the task of automatic protein name extraction, a prerequisite for the large-scale processing of biomedical literature. © 2009 Elsevier Inc. All rights reserved.
KeywordsBigram language model
Protein name extraction
Information storage and retrieval
Natural language processing
Terminology as topic
Published Version (Please cite this version)http://dx.doi.org/10.1016/j.jbi.2009.05.004
Showing items related by title, author, creator and subject.
Tekin, Cem; Yoon, J.; Van Der Schaar, M. (AAAI Press, 2016)With the advances in the field of medical informatics, automated clinical decision support systems are becoming the de facto standard in personalized diagnosis. In order to establish high accuracy and confidence in ...
Ozcelik, E.; Cagiltay, N. E.; Ozcelik, N. S. (Pergamon Press, 2013)Considering the role of games for educational purposes, there has an increase in interest among educators in applying strategies used in popular games to create more engaging learning environments. Learning is more fun and ...
Kanoun, K.; Tekin, C.; Atienza, D.; Van Der Schaar, M. (Institute of Electrical and Electronics Engineers, 2016)Several techniques have been recently proposed to adapt Big-Data streaming applications to existing many core platforms. Among these techniques, online reinforcement learning methods have been proposed that learn how to ...