A4 Refereed article in a conference publication

Out-of-Domain Evaluation of Finnish Dependency Parsing




AuthorsKanerva Jenna, Ginter Filip

EditorsNicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Jan Odijk, Stelios Piperidis

Conference nameInternational Conference on Language Resources and Evaluation

Publishing placeParis

Publication year2022

JournalLREC Proceedings

Book title Proceedings of the 13th Conference on Language Resources and Evaluation (LREC 2022)

Series titleLREC Proceedings

First page 1114

Last page1124

ISBN979-10-95546-72-6

eISSN2522-2686

Web address http://www.lrec-conf.org/proceedings/lrec2022/pdf/2022.lrec-1.120.pdf

Self-archived copy’s web addresshttps://research.utu.fi/converis/portal/detail/Publication/176213812


Abstract

The prevailing practice in the academia is to evaluate the model performance on in-domain evaluation data typically set aside from the training corpus. However, in many real world applications the data on which the model is applied may very substantially differ from the characteristics of the training data. In this paper, we focus on Finnish out-of-domain parsing by introducing a novel UD Finnish-OOD out-of-domain treebank including five very distinct data sources (web documents, clinical, online discussions, tweets, and poetry), and a total of 19,382 syntactic words in 2,122 sentences released under the Universal Dependencies framework. Together with the new treebank, we present extensive out-of-domain parsing evaluation utilizing the available section-level information from three different Finnish UD treebanks (TDT, PUD, OOD). Compared to the previously existing treebanks, the new Finnish-OOD is shown include sections more challenging for the general parser, creating an interesting evaluation setting and yielding valuable information for those applying the parser outside of its training domain.


Downloadable publication

This is an electronic reprint of the original article.
This reprint may differ from the original in pagination and typographic detail. Please cite the original version.





Last updated on 2024-26-11 at 10:37