Towards better language representation in Natural Language Processing A multilingual dataset for text-level Grammatical Error Correction
Number of Authors: 302025 (English)In: International Journal of Learner Corpus Research, ISSN 2215-1478, E-ISSN 2215-1486, Vol. 11, no 2, p. 309-Article in journal (Refereed) Published
Abstract [en]
This paper introduces MultiGEC, a dataset for multilingual Grammatical Error Correction (GEC) in twelve European languages: Czech, English, Estonian, German, Greek, Icelandic, Italian, Latvian, Russian, Slovene, Swedish and Ukrainian. MultiGEC distinguishes itself from previous GEC datasets in that it covers several underrepresented languages, which we argue should be included in resources used to train models for Natural Language Processing tasks which, as GEC itself, have implications for Learner Corpus Research and Second Language Acquisition. Aside from multilingualism, the novelty of the MultiGEC dataset is that it consists of full texts — typically learner essays — rather than individual sentences, making it possible to train systems that take a broader context into account. The dataset was built for MultiGEC-2025, the first shared task in multilingual text-level GEC, but it remains accessible after its competitive phase, serving as a resource to train new error correction systems and perform cross-lingual GEC studies
Place, publisher, year, edition, pages
John Benjamins Publishing Company , 2025. Vol. 11, no 2, p. 309-
National Category
Computer and Information Sciences
Identifiers
URN: urn:nbn:se:ri:diva-78407DOI: 10.1075/ijlcr.24033.masScopus ID: 2-s2.0-105003035015OAI: oai:DiVA.org:ri-78407DiVA, id: diva2:1999052
Note
Swedish Work on Swedish has been supported by Nationella Språkbanken and Huminfra, both funded by the Swedish Research Council (2018–2024, contract 2017-00626; 2022–2024, contract 2021-00176) and their participating partner institutions, as well as the Swedish Research Council grant 2019-04129.
2025-09-182025-09-182025-12-08Bibliographically approved