A4 Refereed article in a conference publication

From Web Crawl to Clean Register-Annotated Corpora




AuthorsLaippala Veronika, Rönnqvist Samuel, Hellström Saara, Luotolahti, Juhani, Repo Liina, Salmela Anna, Skantsi Valtteri and Pyysalo Sampo

Conference nameWeb as Corpus Workshop

Publication year2020

Book title Proceedings of the 12th Web as Corpus Workshop

Series titleProceedings of the Web as Corpus Workshop

Number in series12

First page 14

Last page22

ISBN979-10-95546-68-9

Web address https://www.aclweb.org/anthology/2020.wac-1.3

Self-archived copy’s web addresshttps://research.utu.fi/converis/portal/detail/Publication/51216717


Abstract

The web presents unprecedented opportunities for large-scale collection of text in many languages. However, two critical steps in the development of web corpora remain challenging: the identification of clean text from source HTML and the assignment of genre or register information to the documents. In this paper, we evaluate a multilingual approach to this end. Our starting points are the Swedish and French Common Crawl datasets gathered for the 2017 CoNLL shared task, particularly the URLs. We 1) fetch HTML pages based on the URLs and run boilerplate removal, 2) train a classifier to further clean out undesired text fragments, and 3) annotate text registers. We compare boilerplate removal against the CoNLL texts, and find an improvement. For the further cleaning of undesired material, the best results are achieved using Multilingual BERT with monolingual fine-tuning. However, our results are promising also in a cross-lingual setting, without fine-tuning on the target language. Finally, the register annotations show that most of the documents belong to a relatively small set of registers, which are relatively similar in the two languages. A number of additional flags in the annotation are, however, necessary to reflect the wide range of linguistic variation associated with the documents.


Downloadable publication

This is an electronic reprint of the original article.
This reprint may differ from the original in pagination and typographic detail. Please cite the original version.





Last updated on 26/11/2024 06:38:47 PM