Using the Full-text Content of Academic Articles to Identify and
  Evaluate Algorithm Entities in the Domain of Natural Language Processing

There is no author summary for this article yet. Authors can add summaries to their articles on ScienceOpen to make them more accessible to a non-specialist audience.

Abstract

In the era of big data, the advancement, improvement, and application of algorithms in academic research have played an important role in promoting the development of different disciplines. Academic papers in various disciplines, especially computer science, contain a large number of algorithms. Identifying the algorithms from the full-text content of papers can determine popular or classical algorithms in a specific field and help scholars gain a comprehensive understanding of the algorithms and even the field. To this end, this article takes the field of natural language processing (NLP) as an example and identifies algorithms from academic papers in the field. A dictionary of algorithms is constructed by manually annotating the contents of papers, and sentences containing algorithms in the dictionary are extracted through dictionary-based matching. The number of articles mentioning an algorithm is used as an indicator to analyze the influence of that algorithm. Our results reveal the algorithm with the highest influence in NLP papers and show that classification algorithms represent the largest proportion among the high-impact algorithms. In addition, the evolution of the influence of algorithms reflects the changes in research tasks and topics in the field, and the changes in the influence of different algorithms show different trends. As a preliminary exploration, this paper conducts an analysis of the impact of algorithms mentioned in the academic text, and the results can be used as training data for the automatic extraction of large-scale algorithms in the future. The methodology in this paper is domain-independent and can be applied to other domains.

Related collections

Author and article information

Journal

Publication date Created: 21 October 2020

Article

DOI: 10.1016/j.joi.2020.101091

ArXiV ID: 2010.10817

SO-VID: f6fd9f8e-fc3f-4ae7-bfe5-215fa447a0f3

License:

http://arxiv.org/licenses/nonexclusive-distrib/1.0/

History

Custom metadata

Journal reference Journal of Informetrics,2020

Categories cs.CL cs.IR cs.LG

ScienceOpen disciplines: Theoretical computer science,Information & Library science,Artificial intelligence

Data availability:

ScienceOpen disciplines: Theoretical computer science, Information & Library science, Artificial intelligence

Comments

Comment on this article

scite_

Using the Full-text Content of Academic Articles to Identify and Evaluate Algorithm Entities in the Domain of Natural Language Processing

Read this article at