Two learning approaches for protein name extraction

Tatar, S.; Cicekli, I.

Two learning approaches for protein name extraction

Files

Two learning approaches for protein name extraction.pdf (443.69 KB)

Date

2009

Authors

Tatar, S.

Cicekli, I.

BUIR Usage Stats

2
views

17
downloads

Citation Stats

Abstract

Protein name extraction, one of the basic tasks in automatic extraction of information from biological texts, remains challenging. In this paper, we explore the use of two different machine learning techniques and present the results of the conducted experiments. In the first method, Bigram language model is used to extract protein names. In the latter, we use an automatic rule learning method that can identify protein names located in the biological texts. In both cases, we generalize protein names by using hierarchically categorized syntactic token types. We conducted our experiments on two different datasets. Our first method based on Bigram language model achieved an F-score of 67.7% on the YAPEX dataset and 66.8% on the GENIA corpus. The developed rule learning method obtained 61.8% F-score value on the YAPEX dataset and 61.0% on the GENIA corpus. The results of the comparative experiments demonstrate that both techniques are applicable to the task of automatic protein name extraction, a prerequisite for the large-scale processing of biomedical literature. © 2009 Elsevier Inc. All rights reserved.

Source Title

Journal of Biomedical Informatics

Publisher

Academic Press

Permalink

http://hdl.handle.net/11693/22534

Published Version (Please cite this version)

http://dx.doi.org/10.1016/j.jbi.2009.05.004

Collections

Scholarly Publications - Computer Engineering

Language

English

Type

Article

Full item page

Two learning approaches for protein name extraction

Files

Date

Authors

Editor(s)

Advisor

Supervisor

Co-Advisor

Co-Supervisor

Instructor

BUIR Usage Stats

Citation Stats

Series

Abstract

Source Title

Publisher

Course

Other identifiers

Book Title

Keywords

Degree Discipline

Degree Level

Degree Name

Citation

Permalink

Published Version (Please cite this version)

Collections

Language

Type

Two learning approaches for protein name extraction

Files

Date

Authors

Editor(s)

Advisor

Supervisor

Co-Advisor

Co-Supervisor

Instructor

BUIR Usage Stats

Citation Stats

Share

Series

Abstract

Source Title

Publisher

Course

Other identifiers

Book Title

Keywords

Degree Discipline

Degree Level

Degree Name

Citation

Permalink

Published Version (Please cite this version)

Collections

Language

Type