IADIS International Journal on Computer Science and Information Systems

Published by IADIS (International Association for Development of the Information Society) • ISSN (Online): 1646-3692 • ISSN (Print): 1646-3692
100% Open Access
Double-Blind Peer Review
Crossref DOI Persistent IDs
Open Access Peer-Reviewed Case Studies & Applications

Improvement of Clustering Algorithms by Implementation of Spelling Based Ranking

Evan Bryer *
Theppatorn Rhujittawiwat *
John R. Rose *
Colin F. Wilder *
* 1College of Engineering and Computing, University of South Carolina, USA 2Center for Digital Humanities, University of South Carolina, USA (Portugal)
* 1College of Engineering and Computing, University of South Carolina, USA 2Center for Digital Humanities, University of South Carolina, USA (Portugal)
* 1College of Engineering and Computing, University of South Carolina, USA 2Center for Digital Humanities, University of South Carolina, USA (Portugal)
* 1College of Engineering and Computing, University of South Carolina, USA 2Center for Digital Humanities, University of South Carolina, USA (Portugal)

Abstract

The goal of this paper is to modify an existing clustering algorithm with the use of the Hunspell spell checker to specialize it for the use of cleaning early modern European book title data. Duplicate and corrupted data is a constant concern for data analysis, and clustering has been identified to be a robust tool for normalizing and cleaning data such as ours. In particular, our data comprises over 5 million books published in European languages between 150 0 and 1800 in the Machine -Readable Cataloging (MARC) data format from 17,983 libraries in 123 countries. However, as each library individually catalogued their records, many duplicative and inaccurate records exist in the data set. Additionally, each language evolved over the 300-year period we are studying, and as such many of the words had their spellings altered. Without cleaning and normalizing this data, it would be difficult to find coherent trends, as much of the data may be missed in the query. In p revious research, we have identified the use of Prediction by Partial Matching to provide the most increase in base accuracy when applied to dirty data of similar construct to our data set. However, there are many cases in which the correct book title may not be the most common, either when only two values exist in a cluster, or the dirty title exists in more records. In these cases, a language agnostic clustering algorithm would normalize the incorrect title and lower the overall accuracy of the data set. By implementing the Hunspell spell checker into the clustering algorithm, using it to rank clusters by the number of words not found in their dictionary, we can drastically lower the cases of this occurring. Indeed, this ranking algorithm proved to increas e the overall accuracy of the clustered data by as much as 25% over the unmodified Prediction by Partial Matching algorithm.

Keywords

Pre-Processing Clustering Cleaning Data Mining Spellchecking
Full-Text PDF Available

Read Complete Peer-Reviewed Manuscript

Includes full econometric models, data tables, policy recommendations, declarations, and citations.

Declarations & Ethics

Funding: This research received academic dissemination support through ESCAP / JournalsHub publishing programs.
Conflicts of Interest: The authors declare no competing financial or institutional interests.
Peer Review: Double-blind peer reviewed by international subject specialists.
License: Creative Commons Attribution 4.0 International (CC BY 4.0).
How to Cite This Article
APA / MLA / BibTeX
Bryer, et al. (2021). Improvement of Clustering Algorithms by Implementation of Spelling Based Ranking. IADIS International Journal on Computer Science and Information Systems, 16(2). https://doi.org/10.33965/ijcsis_2021_v16i2_05
Bryer, et al. "Improvement of Clustering Algorithms by Implementation of Spelling Based Ranking." IADIS International Journal on Computer Science and Information Systems, vol. 16, no. 2, 2021. https://doi.org/10.33965/ijcsis_2021_v16i2_05
Bryer, et al. "Improvement of Clustering Algorithms by Implementation of Spelling Based Ranking." IADIS International Journal on Computer Science and Information Systems 16, no. 2 (2021). https://doi.org/10.33965/ijcsis_2021_v16i2_05