Skip to main content

2020 | OriginalPaper | Buchkapitel

Overview of CLEF HIPE 2020: Named Entity Recognition and Linking on Historical Newspapers

verfasst von : Maud Ehrmann, Matteo Romanello, Alex Flückiger, Simon Clematide

Erschienen in: Experimental IR Meets Multilinguality, Multimodality, and Interaction

Verlag: Springer International Publishing

Aktivieren Sie unsere intelligente Suche, um passende Fachinhalte oder Patente zu finden.

search-config
loading …

Abstract

This paper presents an overview of the first edition of HIPE (Identifying Historical People, Places and other Entities), a pioneering shared task dedicated to the evaluation of named entity processing on historical newspapers in French, German and English. Since its introduction some twenty years ago, named entity (NE) processing has become an essential component of virtually any text mining application and has undergone major changes. Recently, two main trends characterise its developments: the adoption of deep learning architectures and the consideration of textual material originating from historical and cultural heritage collections. While the former opens up new opportunities, the latter introduces new challenges with heterogeneous, historical and noisy inputs. In this context, the objective of HIPE, run as part of the CLEF 2020 conference, is threefold: strengthening the robustness of existing approaches on non-standard inputs, enabling performance comparison of NE processing on historical texts, and, in the long run, fostering efficient semantic indexing of historical documents. Tasks, corpora, and results of 13 participating teams are presented.

Sie haben noch keine Lizenz? Dann Informieren Sie sich jetzt über unsere Produkte:

Springer Professional "Wirtschaft+Technik"

Online-Abonnement

Mit Springer Professional "Wirtschaft+Technik" erhalten Sie Zugriff auf:

  • über 102.000 Bücher
  • über 537 Zeitschriften

aus folgenden Fachgebieten:

  • Automobil + Motoren
  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Elektrotechnik + Elektronik
  • Energie + Nachhaltigkeit
  • Finance + Banking
  • Management + Führung
  • Marketing + Vertrieb
  • Maschinenbau + Werkstoffe
  • Versicherung + Risiko

Jetzt Wissensvorsprung sichern!

Springer Professional "Technik"

Online-Abonnement

Mit Springer Professional "Technik" erhalten Sie Zugriff auf:

  • über 67.000 Bücher
  • über 390 Zeitschriften

aus folgenden Fachgebieten:

  • Automobil + Motoren
  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Elektrotechnik + Elektronik
  • Energie + Nachhaltigkeit
  • Maschinenbau + Werkstoffe




 

Jetzt Wissensvorsprung sichern!

Springer Professional "Wirtschaft"

Online-Abonnement

Mit Springer Professional "Wirtschaft" erhalten Sie Zugriff auf:

  • über 67.000 Bücher
  • über 340 Zeitschriften

aus folgenden Fachgebieten:

  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Finance + Banking
  • Management + Führung
  • Marketing + Vertrieb
  • Versicherung + Risiko




Jetzt Wissensvorsprung sichern!

Fußnoten
2
muc, ace, conll, kbp, ester, harem, quaero, germeval, etc.
 
4
For space reasons, the discussion of related work is included in the extended version of this overview  [12].
 
5
From the Swiss National Library, the Luxembourgish National Library, and the Library of Congress (Chronicling America project), respectively. Original collections correspond to 4 Swiss and Luxembourgish titles, and a dozen for English. More details on original sources can be found in [12].
 
7
The November 2019 dump used for annotation is available at https://​files.​ifi.​uzh.​ch/​cl/​impresso/​clef-hipe.
 
16
True positive, False positive, False negative.
 
23
[38] for French, [32] for German.
 
24
SEM [10] is a CRF-based tool using Wapiti [28] as its linear CRF implementation.
 
Literatur
1.
Zurück zum Zitat Akbik, A., Bergmann, T., Blythe, D., Rasul, K., Schweter, S., Vollgraf, R.: FLAIR: an easy-to-use framework for state-of-the-art NLP. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (demonstrations), pp. 54–59. Association for Computational Linguistics, Minneapolis, Minnesota, June 2019. https://www.aclweb.org/anthology/N19-4010 Akbik, A., Bergmann, T., Blythe, D., Rasul, K., Schweter, S., Vollgraf, R.: FLAIR: an easy-to-use framework for state-of-the-art NLP. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (demonstrations), pp. 54–59. Association for Computational Linguistics, Minneapolis, Minnesota, June 2019. https://​www.​aclweb.​org/​anthology/​N19-4010
2.
Zurück zum Zitat Akbik, A., Blythe, D., Vollgraf, R.: Contextual string embeddings for sequence labeling. In: Proceedings of the 27th International Conference on Computational Linguistics, pp. 1638–1649. Association for Computational Linguistics, Santa Fe, New Mexico, USA, August 2018. http://www.aclweb.org/anthology/C18-1139 Akbik, A., Blythe, D., Vollgraf, R.: Contextual string embeddings for sequence labeling. In: Proceedings of the 27th International Conference on Computational Linguistics, pp. 1638–1649. Association for Computational Linguistics, Santa Fe, New Mexico, USA, August 2018. http://​www.​aclweb.​org/​anthology/​C18-1139
4.
Zurück zum Zitat Bollmann, M.: A large-scale comparison of historical text normalization systems. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 3885–3898. Association for Computational Linguistics, Minneapolis, Minnesota (2019). https://doi.org/10.18653/v1/N19-1389 Bollmann, M.: A large-scale comparison of historical text normalization systems. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 3885–3898. Association for Computational Linguistics, Minneapolis, Minnesota (2019). https://​doi.​org/​10.​18653/​v1/​N19-1389
5.
Zurück zum Zitat Borin, L., Kokkinakis, D., Olsson, L.J.: Naming the past: named entity and animacy recognition in 19th century Swedish literature. In: Proceedings of the Workshop on Language Technology for Cultural Heritage Data (LaT-eCH 2007), pp. 1–8 (2007) Borin, L., Kokkinakis, D., Olsson, L.J.: Naming the past: named entity and animacy recognition in 19th century Swedish literature. In: Proceedings of the Workshop on Language Technology for Cultural Heritage Data (LaT-eCH 2007), pp. 1–8 (2007)
6.
Zurück zum Zitat Cappellato, L., Eickhoff, C., Ferro, N., Névéol, A. (eds.): CLEF 2020 Working Notes. In: CEUR Workshop Proceedings Working Notes of CLEF 2020 - Conference and Labs of the Evaluation Forum (2020) Cappellato, L., Eickhoff, C., Ferro, N., Névéol, A. (eds.): CLEF 2020 Working Notes. In: CEUR Workshop Proceedings Working Notes of CLEF 2020 - Conference and Labs of the Evaluation Forum (2020)
7.
Zurück zum Zitat Chiron, G., Doucet, A., Coustaty, M., Visani, M., Moreux, J.P.: Impact of OCR errors on the use of digital libraries: towards a better access to information. In: Proceedings of the 17th ACM/IEEE Joint Conference on Digital Libraries JCDL 2017, pp. 249–252. IEEE Press, Piscataway (2017), http://dl.acm.org/citation.cfm?id=3200334.3200364 Chiron, G., Doucet, A., Coustaty, M., Visani, M., Moreux, J.P.: Impact of OCR errors on the use of digital libraries: towards a better access to information. In: Proceedings of the 17th ACM/IEEE Joint Conference on Digital Libraries JCDL 2017, pp. 249–252. IEEE Press, Piscataway (2017), http://​dl.​acm.​org/​citation.​cfm?​id=​3200334.​3200364
10.
Zurück zum Zitat Dupont, Y., Dinarelli, M., Tellier, I., Lautier, C.: Structured named entity recognition by cascading CRFs. In: Intelligent Text Processing and Computational Linguistics (CICling) (2017) Dupont, Y., Dinarelli, M., Tellier, I., Lautier, C.: Structured named entity recognition by cascading CRFs. In: Intelligent Text Processing and Computational Linguistics (CICling) (2017)
12.
Zurück zum Zitat Ehrmann, M., Romanello, M., Flückiger, A., Clematide, S.: Extended overview of CLEF HIPE 2020: named entity processing on historical newspapers. In: Cappellato, L., Eickhoff, C., Ferro, N., Névéol, A. (eds.) CLEF 2020 Working Notes. Working Notes of CLEF 2020 - Conference and Labs of the Evaluation Forum CEUR-WS (2020) Ehrmann, M., Romanello, M., Flückiger, A., Clematide, S.: Extended overview of CLEF HIPE 2020: named entity processing on historical newspapers. In: Cappellato, L., Eickhoff, C., Ferro, N., Névéol, A. (eds.) CLEF 2020 Working Notes. Working Notes of CLEF 2020 - Conference and Labs of the Evaluation Forum CEUR-WS (2020)
15.
Zurück zum Zitat El Vaigh, C.B., Goasdoué, F., Gravier, G., Sébillot, P.: Using knowledge base semantics in context-aware entity linking. In: 2019 Proceedings of the ACM Symposium on Document Engineering DocEng 2019, pp. 1–10. Association for Computing Machinery, Berlin, Germany, September 2019. https://doi.org/10.1145/3342558.3345393 El Vaigh, C.B., Goasdoué, F., Gravier, G., Sébillot, P.: Using knowledge base semantics in context-aware entity linking. In: 2019 Proceedings of the ACM Symposium on Document Engineering DocEng 2019, pp. 1–10. Association for Computing Machinery, Berlin, Germany, September 2019. https://​doi.​org/​10.​1145/​3342558.​3345393
16.
Zurück zum Zitat Galibert, O., Rosset, S., Grouin, C., Zweigenbaum, P., Quintard, L.: Extended named entity annotation on OCRed documents : from corpus constitution to evaluation campaign. In: Proceedings of the Eighth conference on International Language Resources and Evaluation, pp. 3126–3131. Istanbul, Turkey (2012) Galibert, O., Rosset, S., Grouin, C., Zweigenbaum, P., Quintard, L.: Extended named entity annotation on OCRed documents : from corpus constitution to evaluation campaign. In: Proceedings of the Eighth conference on International Language Resources and Evaluation, pp. 3126–3131. Istanbul, Turkey (2012)
17.
Zurück zum Zitat Ganea, O.E., Hofmann, T.: Deep joint entity disambiguation with local neural attention. In: Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 2619–2629 (2017) Ganea, O.E., Hofmann, T.: Deep joint entity disambiguation with local neural attention. In: Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 2619–2629 (2017)
18.
Zurück zum Zitat Grishman, R., Sundheim, B.: Design of the MUC-6 evaluation. In: Proceedings of the Sixth Conference on Message Understanding Conference (MUC-6), Columbia, Maryland (1995) Grishman, R., Sundheim, B.: Design of the MUC-6 evaluation. In: Proceedings of the Sixth Conference on Message Understanding Conference (MUC-6), Columbia, Maryland (1995)
19.
Zurück zum Zitat Hoffart, J., et al.: Robust disambiguation of named entities in text. In: EMNLP (2011) Hoffart, J., et al.: Robust disambiguation of named entities in text. In: EMNLP (2011)
21.
Zurück zum Zitat van Hulst, J.M., Hasibi, F., Dercksen, K., Balog, K., de Vries, A.P.: REL: an entity linker standing on the shoulders of giants. In: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval SIGIR 2020. ACM (2020) van Hulst, J.M., Hasibi, F., Dercksen, K., Balog, K., de Vries, A.P.: REL: an entity linker standing on the shoulders of giants. In: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval SIGIR 2020. ACM (2020)
23.
Zurück zum Zitat Klie, J.C., Bugert, M., Boullosa, B., de Castilho, R.E., Gurevych, I.: The inception platform: machine-assisted and knowledge-oriented interactive annotation. In: Proceedings of the 27th International Conference on Computational Linguistics: System Demonstrations, pp. 5–9 (2018) Klie, J.C., Bugert, M., Boullosa, B., de Castilho, R.E., Gurevych, I.: The inception platform: machine-assisted and knowledge-oriented interactive annotation. In: Proceedings of the 27th International Conference on Computational Linguistics: System Demonstrations, pp. 5–9 (2018)
24.
Zurück zum Zitat Kolitsas, N., Ganea, O.E., Hofmann, T.: End-to-end neural entity linking. In: Proceedings of the 22nd Conference on Computational Natural Language Learning, pp. 519–529. Association for Computational Linguistics, Brussels, Belgium, October 2018. https://doi.org/10.18653/v1/K18-1050 Kolitsas, N., Ganea, O.E., Hofmann, T.: End-to-end neural entity linking. In: Proceedings of the 22nd Conference on Computational Natural Language Learning, pp. 519–529. Association for Computational Linguistics, Brussels, Belgium, October 2018. https://​doi.​org/​10.​18653/​v1/​K18-1050
25.
Zurück zum Zitat Krippendorff, K.: Content Analysis: An Introduction to its Methodology. Sage Publications, Thousand Oaks (1980)MATH Krippendorff, K.: Content Analysis: An Introduction to its Methodology. Sage Publications, Thousand Oaks (1980)MATH
26.
Zurück zum Zitat Labusch, K., Neudecker, C., Zellhöfer, D.: BERT for named entity recognition in contemporary and historic german. In: Preliminary proceedings of the 15th Conference on Natural Language Processing (KONVENS 2019): Long Papers, pp. 1–9. German Society for Computational Linguistics & Language Technology, Erlangen, Germany (2019) Labusch, K., Neudecker, C., Zellhöfer, D.: BERT for named entity recognition in contemporary and historic german. In: Preliminary proceedings of the 15th Conference on Natural Language Processing (KONVENS 2019): Long Papers, pp. 1–9. German Society for Computational Linguistics & Language Technology, Erlangen, Germany (2019)
28.
Zurück zum Zitat Lavergne, T., Cappé, O., Yvon, F.: Practical very large scale CRFs. In: Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, pp. 504–513. Association for Computational Linguistics (2010) Lavergne, T., Cappé, O., Yvon, F.: Practical very large scale CRFs. In: Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, pp. 504–513. Association for Computational Linguistics (2010)
30.
Zurück zum Zitat Makhoul, J., Kubala, F., Schwartz, R., Weischedel, R.: Performance measures for information extraction. In: Proceedings of DARPA Broadcast News Workshop, pp. 249–252 (1999) Makhoul, J., Kubala, F., Schwartz, R., Weischedel, R.: Performance measures for information extraction. In: Proceedings of DARPA Broadcast News Workshop, pp. 249–252 (1999)
31.
Zurück zum Zitat Martin, L., et al.: Camembert: a tasty french language model (2019) Martin, L., et al.: Camembert: a tasty french language model (2019)
33.
Zurück zum Zitat Nadeau, D., Sekine, S.: A survey of named entity recognition and classification. Lingvisticae Investigationes 30(1), 3–26 (2007)CrossRef Nadeau, D., Sekine, S.: A survey of named entity recognition and classification. Lingvisticae Investigationes 30(1), 3–26 (2007)CrossRef
35.
Zurück zum Zitat Nguyen, D.B., Hoffart, J., Theobald, M., Weikum, G.: Aida-light: high-throughput named-entity disambiguation. In: LDOW (2014) Nguyen, D.B., Hoffart, J., Theobald, M., Weikum, G.: Aida-light: high-throughput named-entity disambiguation. In: LDOW (2014)
38.
Zurück zum Zitat Ortiz Suárez, P.J., Dupont, Y., Muller, B., Romary, L., Sagot, B.: Establishing a new state-of-the-art for French named entity recognition. In: Proceedings of the 12th Language Resources and Evaluation Conference, pp. 4631–4638. European Language Resources Association, Marseille, France, May 2020. https://www.aclweb.org/anthology/2020.lrec-1.569 Ortiz Suárez, P.J., Dupont, Y., Muller, B., Romary, L., Sagot, B.: Establishing a new state-of-the-art for French named entity recognition. In: Proceedings of the 12th Language Resources and Evaluation Conference, pp. 4631–4638. European Language Resources Association, Marseille, France, May 2020. https://​www.​aclweb.​org/​anthology/​2020.​lrec-1.​569
39.
Zurück zum Zitat Pennington, J., Socher, R., Manning, C.D.: Glove: global vectors for word representation. EMNLP 14, 1532–43 (2014) Pennington, J., Socher, R., Manning, C.D.: Glove: global vectors for word representation. EMNLP 14, 1532–43 (2014)
40.
Zurück zum Zitat Peters, M., et al.: Deep Contextualized Word Representations. In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). pp. 2227–2237. Association for Computational Linguistics, New Orleans, Louisiana (2018). https://doi.org/10.18653/v1/N18-1202 Peters, M., et al.: Deep Contextualized Word Representations. In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). pp. 2227–2237. Association for Computational Linguistics, New Orleans, Louisiana (2018). https://​doi.​org/​10.​18653/​v1/​N18-1202
41.
Zurück zum Zitat Piotrowski, M.: Natural language processing for historical texts. Synth. Lect. Hum. Lang. Technol. 5(2), 1–157 (2012)CrossRef Piotrowski, M.: Natural language processing for historical texts. Synth. Lect. Hum. Lang. Technol. 5(2), 1–157 (2012)CrossRef
42.
Zurück zum Zitat Plank, B.: What to do about non-standard (or non-canonical) language in NLP. In: Proceedings of the 13th Conference on Natural Language Processing KONVENS 2016. Bochumer Linguistische Arbeitsberichte (2016) Plank, B.: What to do about non-standard (or non-canonical) language in NLP. In: Proceedings of the 13th Conference on Natural Language Processing KONVENS 2016. Bochumer Linguistische Arbeitsberichte (2016)
43.
44.
Zurück zum Zitat Rosset, S., Grouin, C., Fort, K., Galibert, O., Kahn, J., Zweigenbaum, P.: Structured named entities in two distinct press corpora: contemporary broadcast news and old newspapers. In: Proceedings of the 6th Linguistic Annotation Workshop, pp. 40–48. Association for Computational Linguistics (2012) Rosset, S., Grouin, C., Fort, K., Galibert, O., Kahn, J., Zweigenbaum, P.: Structured named entities in two distinct press corpora: contemporary broadcast news and old newspapers. In: Proceedings of the 6th Linguistic Annotation Workshop, pp. 40–48. Association for Computational Linguistics (2012)
45.
Zurück zum Zitat Rosset, S., Grouin, C., Zweigenbaum, P.: Entités nommées structurées : guide d’annotation Quaero. NOTES et DOCUMENTS 2011–04, LIMSI-CNRS (2011) Rosset, S., Grouin, C., Zweigenbaum, P.: Entités nommées structurées : guide d’annotation Quaero. NOTES et DOCUMENTS 2011–04, LIMSI-CNRS (2011)
48.
Zurück zum Zitat van Strien, D., Beelen, K., Ardanuy, M.C., Hosseini, K., McGillivray, B., Colavizza, G.: Assessing the impact of OCR quality on downstream NLP tasks. In: ICAART 2020 - Proceedings of the 12th International Conference on Agents and Artificial Intelligence. SCITEPRESS - Science and Technology Publications, January 2020. https://doi.org/10.17863/CAM.52068 van Strien, D., Beelen, K., Ardanuy, M.C., Hosseini, K., McGillivray, B., Colavizza, G.: Assessing the impact of OCR quality on downstream NLP tasks. In: ICAART 2020 - Proceedings of the 12th International Conference on Agents and Artificial Intelligence. SCITEPRESS - Science and Technology Publications, January 2020. https://​doi.​org/​10.​17863/​CAM.​52068
49.
Zurück zum Zitat van Strien, D., Beelen, K., Ardanuy, M., Hosseini, K., McGillivray, B., Colavizza, G.: Assessing the impact of OCR quality on downstream NLP tasks. In: Proceedings of the 12th International Conference on Agents and Artificial Intelligence, pp. 484–496. SCITEPRESS - Science and Technology Publications, Valletta, Malta (2020). https://doi.org/10.5220/0009169004840496 van Strien, D., Beelen, K., Ardanuy, M., Hosseini, K., McGillivray, B., Colavizza, G.: Assessing the impact of OCR quality on downstream NLP tasks. In: Proceedings of the 12th International Conference on Agents and Artificial Intelligence, pp. 484–496. SCITEPRESS - Science and Technology Publications, Valletta, Malta (2020). https://​doi.​org/​10.​5220/​0009169004840496​
52.
Zurück zum Zitat Vilain, M., Su, J., Lubar, S.: Entity extraction is a boring solved problem: or is it? In: Human Language Technologies 2007: The Conference of the North American Chapter of the Association for Computational Linguistics; Companion Volume, Short Papers, NAACL-Short 2007, Rochester, New York, pp. 181–184. Association for Computational Linguistics (2007). http://dl.acm.org/citation.cfm?id=1614108.1614154 Vilain, M., Su, J., Lubar, S.: Entity extraction is a boring solved problem: or is it? In: Human Language Technologies 2007: The Conference of the North American Chapter of the Association for Computational Linguistics; Companion Volume, Short Papers, NAACL-Short 2007, Rochester, New York, pp. 181–184. Association for Computational Linguistics (2007). http://​dl.​acm.​org/​citation.​cfm?​id=​1614108.​1614154
Metadaten
Titel
Overview of CLEF HIPE 2020: Named Entity Recognition and Linking on Historical Newspapers
verfasst von
Maud Ehrmann
Matteo Romanello
Alex Flückiger
Simon Clematide
Copyright-Jahr
2020
DOI
https://doi.org/10.1007/978-3-030-58219-7_21

Premium Partner