A4 Vertaisarvioitu artikkeli konferenssijulkaisussa

From Web Crawl to Clean Register-Annotated Corpora




TekijätLaippala Veronika, Rönnqvist Samuel, Hellström Saara, Luotolahti, Juhani, Repo Liina, Salmela Anna, Skantsi Valtteri and Pyysalo Sampo

Konferenssin vakiintunut nimiWeb as Corpus Workshop

Julkaisuvuosi2020

Kokoomateoksen nimiProceedings of the 12th Web as Corpus Workshop

Sarjan nimiProceedings of the Web as Corpus Workshop

Numero sarjassa12

Aloitussivu14

Lopetussivu22

ISBN979-10-95546-68-9

Verkko-osoitehttps://www.aclweb.org/anthology/2020.wac-1.3

Rinnakkaistallenteen osoitehttps://research.utu.fi/converis/portal/detail/Publication/51216717


Tiivistelmä

The web presents unprecedented opportunities for large-scale collection of text in many languages. However, two critical steps in the development of web corpora remain challenging: the identification of clean text from source HTML and the assignment of genre or register information to the documents. In this paper, we evaluate a multilingual approach to this end. Our starting points are the Swedish and French Common Crawl datasets gathered for the 2017 CoNLL shared task, particularly the URLs. We 1) fetch HTML pages based on the URLs and run boilerplate removal, 2) train a classifier to further clean out undesired text fragments, and 3) annotate text registers. We compare boilerplate removal against the CoNLL texts, and find an improvement. For the further cleaning of undesired material, the best results are achieved using Multilingual BERT with monolingual fine-tuning. However, our results are promising also in a cross-lingual setting, without fine-tuning on the target language. Finally, the register annotations show that most of the documents belong to a relatively small set of registers, which are relatively similar in the two languages. A number of additional flags in the annotation are, however, necessary to reflect the wide range of linguistic variation associated with the documents.


Ladattava julkaisu

This is an electronic reprint of the original article.
This reprint may differ from the original in pagination and typographic detail. Please cite the original version.





Last updated on 2024-26-11 at 18:38