Modern Tools for Old Content - in Search of Named Entities in a Finnish OCRed Historical Newspaper Collection 1771-1910

Kimmo Tapio Kettunen, Eetu Mäkelä, Juha Markus Kuokkala, Teemu Petteri Ruokolainen, Jyrki Antero Niemi

Tutkimustuotos: Artikkeli kirjassa/raportissa/konferenssijulkaisussaKonferenssiartikkeliTieteellinenvertaisarvioitu

Abstrakti

Named entity recognition (NER), search, classification and tagging
of names and name like frequent informational elements in texts, has become a
standard information extraction procedure for textual data. NER has been applied
to many types of texts and different types of entities: newspapers, fiction,
historical records, persons, locations, chemical compounds, protein families, animals
etc. In general a NER system’s performance is genre and domain dependent
and also used entity categories vary [1]. The most general set of named entities
is usually some version of three partite categorization of locations, persons
and organizations. In this paper we report first trials and evaluation of NER
with data out of a digitized Finnish historical newspaper collection Digi. Digi
collection contains 1,960,921 pages of newspaper material from years 1771–
1910 both in Finnish and Swedish. We use only material of Finnish documents
in our evaluation. The OCRed newspaper collection has lots of OCR errors; its
estimated word level correctness is about 74–75 % [2]. Our principal NER tagger
is a rule-based tagger of Finnish, FiNER, provided by the FIN-CLARIN
consortium. We show also results of limited category semantic tagging with
tools of the Semantic Computing Research Group (SeCo) of the Aalto University.
FiNER is able to achieve up to 60.0 F-score with named entities in the evaluation
data. Seco’s tools achieve 30.0–60.0 F-score with locations and persons.
Performance of FiNER and SeCo’s tools with the data shows that at best about
half of named entities can be recognized even in a quite erroneous OCRed text
Alkuperäiskielienglanti
OtsikkoLWDA 2016 Lernen, Wissen, Daten, Analysen 2016 Proceedings of the Conference "Lernen, Wissen, Daten, Analysen"
JulkaisupaikkaAachen
KustantajaCEUR Workshop Proceedings
Julkaisupäiväsyyskuuta 2016
TilaJulkaistu - syyskuuta 2016
OKM-julkaisutyyppiA4 Artikkeli konferenssijulkaisuussa
TapahtumaLernen, Wissen, Daten, Analysen - Potsdam, Saksa
Kesto: 12 syyskuuta 201414 syyskuuta 2016

Julkaisusarja

NimiCEUR Workshop Proceedings
ISSN (elektroninen)1613-0073

Tieteenalat

  • 113 Tietojenkäsittely- ja informaatiotieteet

Siteeraa tätä

Kettunen, K. T., Mäkelä, E., Kuokkala, J. M., Ruokolainen, T. P., & Niemi, J. A. (2016). Modern Tools for Old Content - in Search of Named Entities in a Finnish OCRed Historical Newspaper Collection 1771-1910. teoksessa LWDA 2016 Lernen, Wissen, Daten, Analysen 2016 Proceedings of the Conference "Lernen, Wissen, Daten, Analysen" (CEUR Workshop Proceedings). Aachen: CEUR Workshop Proceedings.