Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Towards better language representation in Natural Language Processing A multilingual dataset for text-level Grammatical Error Correction
University of Gothenburg, Sweden; University of Cambridge, UK.
RISE Research Institutes of Sweden.ORCID iD: 0000-0002-7020-8275
CATALPA, Germany; FernUniversität in Hagen, Germany.
Number of Authors: 302025 (English)In: International Journal of Learner Corpus Research, ISSN 2215-1478, E-ISSN 2215-1486, Vol. 11, no 2, p. 309-Article in journal (Refereed) Published
Abstract [en]

This paper introduces MultiGEC, a dataset for multilingual Grammatical Error Correction (GEC) in twelve European languages: Czech, English, Estonian, German, Greek, Icelandic, Italian, Latvian, Russian, Slovene, Swedish and Ukrainian. MultiGEC distinguishes itself from previous GEC datasets in that it covers several underrepresented languages, which we argue should be included in resources used to train models for Natural Language Processing tasks which, as GEC itself, have implications for Learner Corpus Research and Second Language Acquisition. Aside from multilingualism, the novelty of the MultiGEC dataset is that it consists of full texts — typically learner essays — rather than individual sentences, making it possible to train systems that take a broader context into account. The dataset was built for MultiGEC-2025, the first shared task in multilingual text-level GEC, but it remains accessible after its competitive phase, serving as a resource to train new error correction systems and perform cross-lingual GEC studies

Place, publisher, year, edition, pages
John Benjamins Publishing Company , 2025. Vol. 11, no 2, p. 309-
National Category
Computer and Information Sciences
Identifiers
URN: urn:nbn:se:ri:diva-78407DOI: 10.1075/ijlcr.24033.masScopus ID: 2-s2.0-105003035015OAI: oai:DiVA.org:ri-78407DiVA, id: diva2:1999052
Note

Swedish Work on Swedish has been supported by Nationella Språkbanken and Huminfra, both funded by the Swedish Research Council (2018–2024, contract 2017-00626; 2022–2024, contract 2021-00176) and their participating partner institutions, as well as the Swedish Research Council grant 2019-04129.

Available from: 2025-09-18 Created: 2025-09-18 Last updated: 2025-12-08Bibliographically approved

Open Access in DiVA

fulltext(314 kB)57 downloads
File information
File name FULLTEXT01.pdfFile size 314 kBChecksum SHA-512
3e5d652f4eb01b58beaa60b6eaae3c2fe22ad50e9460370a162092e61dc7b494445c07402b7ef94c3b2c4e0c4a6518787bb2a99fc1cf46af20c4216275ea9a34
Type fulltextMimetype application/pdf

Other links

Publisher's full textScopus

Authority records

Kurfalı, Murathan

Search in DiVA

By author/editor
Kurfalı, Murathan
By organisation
RISE Research Institutes of Sweden
In the same journal
International Journal of Learner Corpus Research
Computer and Information Sciences

Search outside of DiVA

GoogleGoogle Scholar
Total: 58 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 2244 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf