Automatic vs. manual curation of a multi-source chemical dictionary: the impact on text mining

There is no author summary for this article yet. Authors can add summaries to their articles on ScienceOpen to make them more accessible to a non-specialist audience.

Abstract

Correction In 'Automatic vs. manual curation of a multi-source chemical dictionary: the impact on text mining' (Hettne et al. Journal of Cheminformatics 2010, 2:3) [1], the name of the automatically curated dictionary is identified as 'Chemlist'. CHEMLIST is a trademark that the American Chemical Society has used for many years to identify its Regulated Chemicals Listing (CAS) database. To avoid future confusion, the 'Chemlist' dictionary mentioned in this article has been renamed to 'Jochem.'

Related collections

Most cited references 1

Record: found
Abstract: found
Article: not found

Automatic vs. manual curation of a multi-source chemical dictionary: the impact on text mining

Kristina M Hettne, Antony Williams, Erik M van Mulligen … (2010)

Background Previously, we developed a combined dictionary dubbed Chemlist for the identification of small molecules and drugs in text based on a number of publicly available databases and tested it on an annotated corpus. To achieve an acceptable recall and precision we used a number of automatic and semi-automatic processing steps together with disambiguation rules. However, it remained to be investigated which impact an extensive manual curation of a multi-source chemical dictionary would have on chemical term identification in text. ChemSpider is a chemical database that has undergone extensive manual curation aimed at establishing valid chemical name-to-structure relationships. Results We acquired the component of ChemSpider containing only manually curated names and synonyms. Rule-based term filtering, semi-automatic manual curation, and disambiguation rules were applied. We tested the dictionary from ChemSpider on an annotated corpus and compared the results with those for the Chemlist dictionary. The ChemSpider dictionary of ca. 80 k names was only a 1/3 to a 1/4 the size of Chemlist at around 300 k. The ChemSpider dictionary had a precision of 0.43 and a recall of 0.19 before the application of filtering and disambiguation and a precision of 0.87 and a recall of 0.19 after filtering and disambiguation. The Chemlist dictionary had a precision of 0.20 and a recall of 0.47 before the application of filtering and disambiguation and a precision of 0.67 and a recall of 0.40 after filtering and disambiguation. Conclusions We conclude the following: (1) The ChemSpider dictionary achieved the best precision but the Chemlist dictionary had a higher recall and the best F-score; (2) Rule-based filtering and disambiguation is necessary to achieve a high precision for both the automatically generated and the manually curated dictionary. ChemSpider is available as a web service at http://www.chemspider.com/ and the Chemlist dictionary is freely available as an XML file in Simple Knowledge Organization System format on the web at http://www.biosemantics.org/chemlist.

0 comments Cited 13 times – based on 0 reviews      Review now

Bookmark

All references

Author and article information

Journal

Journal ID (nlm-ta): J Cheminform

Title: Journal of Cheminformatics

Publisher: BioMed Central

ISSN (Electronic): 1758-2946

Publication date Collection: 2010

Publication date (Electronic): 3 June 2010

Volume: 2

Page: 4

Affiliations

[1 ]Department of Medical Informatics, Erasmus University Medical Center, Rotterdam, The Netherlands

[2 ]Department of Health Risk Analysis and Toxicology, Maastricht University, Maastricht, The Netherlands

[3 ]Royal Society of Chemistry, 904 Tamaras Circle, Wake Forest, NC-27587, USA

Article

Publisher ID: 1758-2946-2-4

DOI: 10.1186/1758-2946-2-4

PMC ID: 2890529

PubMed ID: 20525267

SO-VID: 29ada904-7108-4f35-9cbe-472f4d68c29b

License:

This is an Open Access article distributed under the terms of the Creative Commons Attribution License ( http://creativecommons.org/licenses/by/2.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

History

Date received : 1 June 2010

Date accepted : 3 June 2010

Comments

Comment on this article

scite_

Cited by 3

See all cited by

- Version 1

Automatic vs. manual curation of a multi-source chemical dictionary: the impact on text mining

Read this article at

Abstract

Related collections

ChemSpider related publications

Most cited references 1

Automatic vs. manual curation of a multi-source chemical dictionary: the impact on text mining

Author and article information

Journal

Affiliations

Article

History

Categories

Comments

Comment on this article

Similar content 77

Cited by 3