The Open UniversitySkip to content
 

A cascaded approach to normalising gene mentions in biomedical literature

Yang, Hui; Nenadic, Goran and Keane, John A. (2007). A cascaded approach to normalising gene mentions in biomedical literature. Bioinformation, 2(5) pp. 197–206.

Full text available as:
[img]
Preview
PDF (Not Set) - Requires a PDF viewer such as GSview, Xpdf or Adobe Acrobat Reader
Download (377Kb)
URL: http://www.bioinformation.net/002/004600022007.htm
Google Scholar: Look up in Google Scholar

Abstract

Linking gene and protein names mentioned in the literature to unique identifiers in referent genomic databases is an essential step in accessing and integrating knowledge in the biomedical domain. However, it remains a challenging task due to lexical and terminological variation, and ambiguity of gene name mentions in documents. We present a generic and effective rule-based approach to link gene mentions in the literature to referent genomic databases, where pre-processing of both gene synonyms in the databases and gene mentions in text are first applied. The mapping method employs a cascaded approach, which combines exact, exact-like and token-based approximate matching by using flexible representations of a gene synonym dictionary and gene mentions generated during the pre-processing phase. We also consider multi-gene name mentions and permutation of components in gene names. A systematic evaluation of the suggested methods has identified steps that are beneficial for improving either precision or recall in gene name identification. The results of the experiments on the BioCreAtIvE2 data sets (identification of human gene names) demonstrated that our methods achieved highly encouraging results with F-measure of up to 81.20%.

Item Type: Journal Article
ISSN: 0973-2063
Academic Unit/Department: Mathematics, Computing and Technology > Computing & Communications
Interdisciplinary Research Centre: Centre for Research in Computing (CRC)
Item ID: 12989
Depositing User: Hui Yang
Date Deposited: 30 Jan 2009 01:41
Last Modified: 06 Dec 2010 22:09
URI: http://oro.open.ac.uk/id/eprint/12989
Share this page:

Actions (login may be required)

View Item
Report issue / request change

Policies | Disclaimer

© The Open University   + 44 (0)870 333 4340   general-enquiries@open.ac.uk