Labelling the past: data set creation and multi-label classification of Dutch archaeological excavation reports

Abstract The extraction of information from Dutch archaeological grey literature has recently been investigated by the AGNES project. AGNES aims to disclose relevant information by means of a web search engine, to enable researchers to search through excavation reports. In this paper, we focus on the multi-labelling of archaeological excavation reports with time periods and site types, and provide a manually labelled reference set to this end. We propose a series of approaches, pre-processing methods, and various modifications of the training set to address the often low quality of both texts... Mehr ...

Verfasser: Brandsen, Alex
Koole, Martin
Dokumenttyp: Artikel
Erscheinungsdatum: 2021
Reihe/Periodikum: Language Resources and Evaluation ; volume 56, issue 2, page 543-572 ; ISSN 1574-020X 1574-0218
Verlag/Hrsg.: Springer Science and Business Media LLC
Sprache: Englisch
Permalink: https://search.fid-benelux.de/Record/base-27457046
Datenquelle: BASE; Originalkatalog
Powered By: BASE
Link(s) : http://dx.doi.org/10.1007/s10579-021-09552-6

Abstract The extraction of information from Dutch archaeological grey literature has recently been investigated by the AGNES project. AGNES aims to disclose relevant information by means of a web search engine, to enable researchers to search through excavation reports. In this paper, we focus on the multi-labelling of archaeological excavation reports with time periods and site types, and provide a manually labelled reference set to this end. We propose a series of approaches, pre-processing methods, and various modifications of the training set to address the often low quality of both texts and labels. We find that despite those issues, our proposed methods lead to promising results.