IADIS International Journal on Computer Science and Information Systems

Published by IADIS (International Association for Development of the Information Society) • ISSN (Online): 1646-3692 • ISSN (Print): 1646-3692
100% Open Access
Double-Blind Peer Review
Crossref DOI Persistent IDs
Open Access Peer-Reviewed Case Studies & Applications

Empirical Evaluation of Crf-based Bibliography Extraction From Research Papers

Manabu Ohta. Okayama University *
Japan Researcher *
* Ryohei Inoue. Shikoku Hitachi Systems, Ltd., Japan Atsuhiro Takasu. National Institute of Informatics, Japan ABSTRA CT We proposed an automatic bibliography extraction method for research papers scanned with OCR markup. The method uses conditional random fields (CRF s) to label serially OCRed text lines in the article title page as appropriate bibliographic element names. Although we achieved good extraction accuracies for some Japanese academic journals, extraction errors are inevitable. Therefore, this paper proposes three confidence measures for bibliography labeling to detect such extraction errors. This paper also reports an empirical evaluation of CRF-based page analysis for research papers on the basis not only of labeling accuracy but also of labeling error detection. We applied the three confidence measures to detecting errors of labeling articles selected from three academic journals published in Japan. The experiments showed that the proposed confidence m easures reasonably indicated the labeling accuracies and could be used for error detection. This paper also discusses the tradeoff between the quality of bibliographic data assured by human post-editing of detected errors and its cost. KEYWORDS Bibliography extraction, conditional random field (CRF), error detection, OCR, digital library. 1. INTRODUCTION Nowadays many publishers and academic societies provide articles in digital formats. Owing to these services we can quickl y obtain articles. Early digital library systems stored articles independently from each other. Hence, we needed to make another search to obtain cited papers. Recently, they begin to make networked documents where cited papers are linked to each other and authors are also connected to the articles they wrote. (Portugal)
* Ryohei Inoue. Shikoku Hitachi Systems, Ltd., Japan Atsuhiro Takasu. National Institute of Informatics, Japan ABSTRA CT We proposed an automatic bibliography extraction method for research papers scanned with OCR markup. The method uses conditional random fields (CRF s) to label serially OCRed text lines in the article title page as appropriate bibliographic element names. Although we achieved good extraction accuracies for some Japanese academic journals, extraction errors are inevitable. Therefore, this paper proposes three confidence measures for bibliography labeling to detect such extraction errors. This paper also reports an empirical evaluation of CRF-based page analysis for research papers on the basis not only of labeling accuracy but also of labeling error detection. We applied the three confidence measures to detecting errors of labeling articles selected from three academic journals published in Japan. The experiments showed that the proposed confidence m easures reasonably indicated the labeling accuracies and could be used for error detection. This paper also discusses the tradeoff between the quality of bibliographic data assured by human post-editing of detected errors and its cost. KEYWORDS Bibliography extraction, conditional random field (CRF), error detection, OCR, digital library. 1. INTRODUCTION Nowadays many publishers and academic societies provide articles in digital formats. Owing to these services we can quickl y obtain articles. Early digital library systems stored articles independently from each other. Hence, we needed to make another search to obtain cited papers. Recently, they begin to make networked documents where cited papers are linked to each other and authors are also connected to the articles they wrote. (Portugal)

Abstract

This peer-reviewed paper presents original scientific research in computer science and information systems, addressing "Empirical Evaluation of Crf-based Bibliography Extraction From Research Papers". The authors discuss theoretical foundations, system architectures, empirical evaluations, and practical implications for modern digital ecosystems. Published in IJCSIS Vol. 7 No. 2 (2012).

Keywords

Bibliography extraction conditional random field (CRF) error detection OCR digital library.
Full-Text PDF Available

Read Complete Peer-Reviewed Manuscript

Includes full econometric models, data tables, policy recommendations, declarations, and citations.

Declarations & Ethics

Funding: This research received academic dissemination support through ESCAP / JournalsHub publishing programs.
Conflicts of Interest: The authors declare no competing financial or institutional interests.
Peer Review: Double-blind peer reviewed by international subject specialists.
License: Creative Commons Attribution 4.0 International (CC BY 4.0).
How to Cite This Article
APA / MLA / BibTeX
University, et al. (2012). Empirical Evaluation of Crf-based Bibliography Extraction From Research Papers. IADIS International Journal on Computer Science and Information Systems, 7(2). https://doi.org/10.33965/ijcsis_2012_v7i2_03
University, et al. "Empirical Evaluation of Crf-based Bibliography Extraction From Research Papers." IADIS International Journal on Computer Science and Information Systems, vol. 7, no. 2, 2012. https://doi.org/10.33965/ijcsis_2012_v7i2_03
University, et al. "Empirical Evaluation of Crf-based Bibliography Extraction From Research Papers." IADIS International Journal on Computer Science and Information Systems 7, no. 2 (2012). https://doi.org/10.33965/ijcsis_2012_v7i2_03