Change search
Link to record
Permanent link

Direct link
Alternative names
Publications (10 of 22) Show all publications
Carlsson, F., Broberg, J., Hillbom, E., Sahlgren, M. & Nivre, J. (2024). BRANCH-GAN: IMPROVING TEXT GENERATION WITH (NOT SO) LARGE LANGUAGE MODELS. In: 12th International Conference on Learning Representations, ICLR 2024: . Paper presented at 12th International Conference on Learning Representations, ICLR 2024.Vienna, Austria. 7 May 2024 through 11 May 2024. International Conference on Learning Representations, ICLR
Open this publication in new window or tab >>BRANCH-GAN: IMPROVING TEXT GENERATION WITH (NOT SO) LARGE LANGUAGE MODELS
Show others...
2024 (English)In: 12th International Conference on Learning Representations, ICLR 2024, International Conference on Learning Representations, ICLR , 2024Conference paper, Published paper (Refereed)
Abstract [en]

The current advancements in open domain text generation have been spearheaded by Transformer-based large language models. Leveraging efficient parallelization and vast training datasets, these models achieve unparalleled text generation capabilities. Even so, current models are known to suffer from deficiencies such as repetitive texts, looping issues, and lack of robustness. While adversarial training through generative adversarial networks (GAN) is a proposed solution, earlier research in this direction has predominantly focused on older architectures, or narrow tasks. As a result, this approach is not yet compatible with modern language models for open-ended text generation, leading to diminished interest within the broader research community. We propose a computationally efficient GAN approach for sequential data that utilizes the parallelization capabilities of Transformer models. Our method revolves around generating multiple branching sequences from each training sample, while also incorporating the typical next-step prediction loss on the original data. In this way, we achieve a dense reward and loss signal for both the generator and the discriminator, resulting in a stable training dynamic. We apply our training method to pre-trained language models, using data from their original training set but less than 0.01% of the available data. A comprehensive human evaluation shows that our method significantly improves the quality of texts generated by the model while avoiding the previously reported sparsity problems of GAN approaches. Even our smaller models outperform larger original baseline models with more than 16 times the number of parameters. Finally, we corroborate previous claims that perplexity on held-out data is not a sufficient metric for measuring the quality of generated texts.

Place, publisher, year, edition, pages
International Conference on Learning Representations, ICLR, 2024
Keywords
Computational linguistics; ’current; Computationally efficient; Current modeling; Language model; Modern languages; Parallelizations; Research communities; Sequential data; Text generations; Training dataset; Generative adversarial networks
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:ri:diva-75007 (URN)2-s2.0-85200551619 (Scopus ID)
Conference
12th International Conference on Learning Representations, ICLR 2024.Vienna, Austria. 7 May 2024 through 11 May 2024
Note

The research presented in this paper was supported by the Swedish Research Council (grant no. 2022-02909) and by a donation from Meta.

Available from: 2024-09-10 Created: 2024-09-10 Last updated: 2025-09-23Bibliographically approved
Gogoulou, E., Lesort, T., Boman, M. & Nivre, J. (2024). Continual Learning Under Language Shift. Paper presented at 27th International Conference on Text, Speech, and Dialogue, TSD 2024, Brno. 9 September 2024 through 13 September 2024. Lecture Notes in Computer Science, 15048 LNAI, 71-84
Open this publication in new window or tab >>Continual Learning Under Language Shift
2024 (English)In: Lecture Notes in Computer Science, ISSN 0302-9743, E-ISSN 1611-3349, Vol. 15048 LNAI, p. 71-84Article in journal (Refereed) Published
Abstract [en]

The recent increase in data and model scale for language model pre-training has led to huge training costs. In scenarios where new data become available over time, updating a model instead of fully retraining it would therefore provide significant gains. We study the pros and cons of updating a language model when new data comes from new languages – the case of continual learning under language shift. Starting from a monolingual English language model, we incrementally add data from Danish, Icelandic and Norwegian to investigate how forward and backward transfer effects depend on pre-training order and characteristics of languages, for models with 126M, 356M and 1.3B parameters. Our results show that, while forward transfer is largely positive and independent of language order, backward transfer can be positive or negative depending on the order and characteristics of new languages. We explore a number of potentially explanatory factors and find that a combination of language contamination and syntactic similarity best fits our results. 

Place, publisher, year, edition, pages
Springer Science and Business Media Deutschland GmbH, 2024
Keywords
Adversarial machine learning; Federated learning; Modeling languages; Continual learning; English languages; Forward-and-backward; Icelandics; Language model; Large language model; Model scale; Multilingual NLP; Pre-training; Training costs; Contrastive Learning
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:ri:diva-75655 (URN)10.1007/978-3-031-70563-2_6 (DOI)2-s2.0-85203595319 (Scopus ID)
Conference
27th International Conference on Text, Speech, and Dialogue, TSD 2024, Brno. 9 September 2024 through 13 September 2024
Note

The research presented in this paper was supported by the Swedish Research Council (grant no. 2022-02909). A significant part of the computations was enabled by the Berzelius resource provided by the Knut and Alice Wallenberg Foundation at the National Supercomputer Centre at Link\u00F6ping University, Sweden (Berzelius-2023-178). In addition, the authors gratefully acknowledge the HPC RIVR consortium (www.hpc-rivr.si) and EuroHPC JU (eurohpc-ju.europa.eu) for funding this research by providing computing resources of the HPC system Vega at the Institute of Information Science (www.izum.si). Magnus Boman acknowledges funding from the Swedish Research Council on Scalable Federated Architectures.

Available from: 2024-11-01 Created: 2024-11-01 Last updated: 2025-09-23Bibliographically approved
Karlgren, J., Dürlich, L., Gogoulou, E., Guillou, L., Nivre, J., Sahlgren, M. & Talman, A. (2024). ELOQUENT CLEF Shared Tasks for Evaluation of Generative Language Model Quality. Paper presented at 46th European Conference on Information Retrieval, ECIR 2024. Glasgow, UK. 24 March 2024 through 28 March 2024. Lecture Notes in Computer Science, 14612 LNCS, 459-465
Open this publication in new window or tab >>ELOQUENT CLEF Shared Tasks for Evaluation of Generative Language Model Quality
Show others...
2024 (English)In: Lecture Notes in Computer Science, ISSN 0302-9743, E-ISSN 1611-3349, Vol. 14612 LNCS, p. 459-465Article in journal (Refereed) Published
Abstract [en]

ELOQUENT is a set of shared tasks for evaluating the quality and usefulness of generative language models. ELOQUENT aims to bring together some high-level quality criteria, grounded in experiences from deploying models in real-life tasks, and to formulate tests for those criteria, preferably implemented to require minimal human assessment effort and in a multilingual setting. The selected tasks for this first year of ELOQUENT are (1) probing a language model for topical competence; (2) assessing the ability of models to generate and detect hallucinations; (3) assessing the robustness of a model output given variation in the input prompts; and (4) establishing the possibility to distinguish human-generated text from machine-generated text.

Place, publisher, year, edition, pages
Springer Science and Business Media Deutschland GmbH, 2024
Keywords
Benchmarking; CLEF; Generative language model; Human assessment; Language model; LLM; Modeling quality; Multilinguality; Quality benchmark; Quality criteria; Shared task; Computational linguistics
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:ri:diva-72876 (URN)10.1007/978-3-031-56069-9_63 (DOI)2-s2.0-85189366495 (Scopus ID)
Conference
46th European Conference on Information Retrieval, ECIR 2024. Glasgow, UK. 24 March 2024 through 28 March 2024
Available from: 2024-04-26 Created: 2024-04-26 Last updated: 2025-09-23Bibliographically approved
Aleksandrova, A. & Nivre, J. (2024). Models and Strategies for Russian Word Sense Disambiguation: A Comparative Analysis. Paper presented at 27th International Conference on Text, Speech, and Dialogue, TSD 2024. Brno. 9 September 2024 through 13 September 202. Lecture Notes in Computer Science, 15048 LNAI, 267-278
Open this publication in new window or tab >>Models and Strategies for Russian Word Sense Disambiguation: A Comparative Analysis
2024 (English)In: Lecture Notes in Computer Science, ISSN 0302-9743, E-ISSN 1611-3349, Vol. 15048 LNAI, p. 267-278Article in journal (Refereed) Published
Abstract [en]

Word sense disambiguation (WSD) is a core task in computational linguistics that involves interpreting polysemous words in context by identifying senses from a predefined sense inventory. Despite the dominance of BERT and its derivatives in WSD evaluation benchmarks, their effectiveness in encoding and retrieving word senses, especially in languages other than English, remains relatively unexplored. This paper provides a detailed quantitative analysis, comparing various BERT-based models for Russian, and examines two primary WSD strategies: fine-tuning and feature-based nearest-neighbor classification. The best results are obtained with the ruBERT model coupled with the feature-based nearest neighbor strategy. This approach adeptly captures even fine-grained meanings with limited data and diverse sense distributions. 

Place, publisher, year, edition, pages
Springer Science and Business Media Deutschland GmbH, 2024
Keywords
Benchmarking; Nearest neighbor search; BERT; Comparative analyzes; Encodings; Feature-based; In contexts; Polysemous word; Russian; Sense inventories; Word sense; Word Sense Disambiguation; Computational linguistics
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:ri:diva-75653 (URN)10.1007/978-3-031-70563-2_21 (DOI)2-s2.0-85203586585 (Scopus ID)
Conference
27th International Conference on Text, Speech, and Dialogue, TSD 2024. Brno. 9 September 2024 through 13 September 202
Available from: 2024-11-01 Created: 2024-11-01 Last updated: 2025-09-23Bibliographically approved
Karlgren, J., Dürlich, L., Gogoulou, E., Guillou, L., Nivre, J., Sahlgren, M., . . . Zahra, S. (2024). Overview of ELOQUENT 2024: Shared Tasks for Evaluating Generative Language Model Quality. Lecture Notes in Computer Science, 14959 LNCS, 53-72
Open this publication in new window or tab >>Overview of ELOQUENT 2024: Shared Tasks for Evaluating Generative Language Model Quality
Show others...
2024 (English)In: Lecture Notes in Computer Science, ISSN 0302-9743, E-ISSN 1611-3349, Vol. 14959 LNCS, p. 53-72Article in journal (Refereed) Published
Abstract [en]

ELOQUENT is a set of shared tasks for evaluating the quality and usefulness of generative language models. ELOQUENT aims to apply high-level quality criteria, grounded in experiences from deploying models in real-life tasks, and to formulate tests for those criteria, preferably implemented to require minimal human assessment effort and in a multilingual setting. The tasks for the first year of ELOQUENT were (1) Topical quiz, in which language models are probed for topical competence; (2) HalluciGen, in which we assessed the ability of models to generate and detect hallucinations; (3) Robustness, in which we assessed the robustness and consistency of a model output given variation in the input prompts; and (4) Voight-Kampff, run in partnership with the PAN lab, with the aim of discovering whether it is possible to automatically distinguish human-generated text from machine-generated text. This first year of experimentation has shown—as expected—that using self-assessment with models judging models is feasible, but not entirely straight-forward, and that a a judicious comparison with human assessment and application context is necessary to be able to trust self-assessed quality judgments. 

Place, publisher, year, edition, pages
Springer Science and Business Media Deutschland GmbH, 2024
Keywords
Generative language model; Human assessment; Language model; LLM; Modeling quality; Quality criteria; Self-assessed quality; Shared task; Generative adversarial networks
National Category
Natural Language Processing
Identifiers
urn:nbn:se:ri:diva-76049 (URN)10.1007/978-3-031-71908-0_3 (DOI)2-s2.0-85205360663 (Scopus ID)
Available from: 2024-10-30 Created: 2024-10-30 Last updated: 2025-09-23Bibliographically approved
Dürlich, L., Gogoulou, E., Guillou, L., Nivre, J. & Zahra, S. (2024). Overview of the CLEF-2024 Eloquent Lab: Task 2 on HalluciGen. In: CEUR Workshop Proceedings: . Paper presented at 25th Working Notes of the Conference and Labs of the Evaluation Forum, CLEF 2024. Grenoble. 9 September 2024 through 12 September 2024 (pp. 691-702). CEUR-WS, 3740
Open this publication in new window or tab >>Overview of the CLEF-2024 Eloquent Lab: Task 2 on HalluciGen
Show others...
2024 (English)In: CEUR Workshop Proceedings, CEUR-WS , 2024, Vol. 3740, p. 691-702Conference paper, Published paper (Refereed)
Abstract [en]

In the HalluciGen task we aim to discover whether LLMs have an internal representation of hallucination. Specifically, we investigate whether LLMs can be used to both generate and detect hallucinated content. In the cross-model evaluation setting we take this a step further and explore the viability of using an LLM to evaluate output produced by another LLM. We include generation, detection, and cross-model evaluation steps for two scenarios: paraphrase and machine translation. Overall we find that performance of the baselines and submitted systems is highly variable, however initial results are promising and lessons learned from this year’s task will provide a solid foundation for future iterations of the task. In particular, we highlight that human validation of generated output is ideally necessary to ensure the robustness of the cross-model evaluation results. We aim to address this challenge in future iterations of HalluciGen. 

Place, publisher, year, edition, pages
CEUR-WS, 2024
Keywords
Computational linguistics; Computer aided language translation; Modeling languages; Cross model; Detection models; Evaluation; Generative language model; Hallucination; Internal representation; Language model; Machine translations; Model evaluation; Performance; Machine translation
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:ri:diva-75021 (URN)2-s2.0-85201646530 (Scopus ID)
Conference
25th Working Notes of the Conference and Labs of the Evaluation Forum, CLEF 2024. Grenoble. 9 September 2024 through 12 September 2024
Note

This lab has been partially supported by the Swedish Research Council (grant number 2022-02909) and by UK Research and Innovation (UKRI) under the UK government's Horizon Europe funding guarantee [grant number 10039436 (Utter)].

Available from: 2024-09-06 Created: 2024-09-06 Last updated: 2025-09-23Bibliographically approved
Nivre, J. (2024). Ten Years of Universal Dependencies. In: Proceedings of the International Conference Computational Linguistics in Bulgaria: . Paper presented at 6th International Conference on Computational Linguistics in Bulgaria, CLIB 2024. Sofia. 9 September 2024 through 10 September 2024. Institute for Bulgarian Language
Open this publication in new window or tab >>Ten Years of Universal Dependencies
2024 (English)In: Proceedings of the International Conference Computational Linguistics in Bulgaria, Institute for Bulgarian Language , 2024Conference paper, Published paper (Other academic)
Place, publisher, year, edition, pages
Institute for Bulgarian Language, 2024
National Category
Philosophy, Ethics and Religion
Identifiers
urn:nbn:se:ri:diva-76165 (URN)2-s2.0-85206240578 (Scopus ID)
Conference
6th International Conference on Computational Linguistics in Bulgaria, CLIB 2024. Sofia. 9 September 2024 through 10 September 2024
Available from: 2024-11-19 Created: 2024-11-19 Last updated: 2025-09-23Bibliographically approved
Weissweiler, L., Böbel, N., Guiller, K., Herrera, S., Scivetti, W., Lorenzi, A., . . . Schneider, N. (2024). UCxn: Typologically Informed Annotation of Constructions Atop Universal Dependencies. In: 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings: . Paper presented at Joint 30th International Conference on Computational Linguistics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024. Hybrid, Torino, Italy. 20 May 2024 through 25 May 2024 (pp. 16919-16932). European Language Resources Association (ELRA)
Open this publication in new window or tab >>UCxn: Typologically Informed Annotation of Constructions Atop Universal Dependencies
Show others...
2024 (English)In: 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings, European Language Resources Association (ELRA) , 2024, p. 16919-16932Conference paper, Published paper (Refereed)
Abstract [en]

The Universal Dependencies (UD) project has created an invaluable collection of treebanks with contributions in over 140 languages. However, the UD annotations do not tell the full story. Grammatical constructions that convey meaning through a particular combination of several morphosyntactic elements-for example, interrogative sentences with special markers and/or word orders-are not labeled holistically. We argue for (i) augmenting UD annotations with a “UCxn” annotation layer for such meaning-bearing grammatical constructions, and (ii) approaching this in a typologically informed way so that morphosyntactic strategies can be compared across languages. As a case study, we consider five construction families in ten languages, identifying instances of each construction in UD treebanks through the use of morphosyntactic patterns. In addition to findings regarding these particular constructions, our study yields important insights on methodology for describing and identifying constructions in language-general and language-particular ways, and lays the foundation for future constructional enrichment of UD treebanks. 

Place, publisher, year, edition, pages
European Language Resources Association (ELRA), 2024
Keywords
Case-studies; Corpus annotations; Grammatical construction; Interrogative sentences; Treebanks; Typology; Universal dependency; Word orders
National Category
General Language Studies and Linguistics
Identifiers
urn:nbn:se:ri:diva-74948 (URN)2-s2.0-85195888534 (Scopus ID)9782493814104 (ISBN)
Conference
Joint 30th International Conference on Computational Linguistics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024. Hybrid, Torino, Italy. 20 May 2024 through 25 May 2024
Note

This work was initiated by the Dagstuhl Seminar 23191 \u201CUniversals of Linguistic Idiosyncrasy in Multilingual Computational Linguistics\u201D (https://www.dagstuhl.de/23191). In addition to the authors, our discussion group also included Francis Bond, Jorg B\u00FCcker, Mathieu Constant, Daniel Flickinger, Sylvain Kahane, Peter Ljungl\u00F6f, Teresa Lynn, Alexandre Rademaker, Manfred Sailer, and Agata Savary. We are grateful for Grew infrastructure support from Bruno Guillaume; for feedback from members of the NERT lab at Georgetown and anonymous reviewers; and for discussion with Nat\u00E1lia Sathler Sigliano about some of the constructions in Portuguese. This work was supported in part by Israeli Ministry of Science and Technology grant No. 0002336 (Nurit Melnik, PI), CAPES PROEX grant No. 88887.816228/2023-00 (Arthur Lorenzi, PhD) and NSF award IIS-2144881 (Nathan Schneider, PI).

Available from: 2024-08-28 Created: 2024-08-28 Last updated: 2025-09-23Bibliographically approved
Savary, A., Zeman, D., Mititelu, V. B., Barreiro, A., Caftanatov, O., De Marneffe, M.-C., . . . Wróblewska, A. (2024). UniDive: A COST Action on Universality, Diversity and Idiosyncrasy in Language Technology. In: : . Paper presented at 3rd Annual Meeting of the ELRA-ISCA Special Interest Group on Under-Resourced Languages, SIGUL 2024 at LREC-COLING 2024 (pp. 372-382). European Language Resources Association (ELRA)
Open this publication in new window or tab >>UniDive: A COST Action on Universality, Diversity and Idiosyncrasy in Language Technology
Show others...
2024 (English)Conference paper, Published paper (Refereed)
Abstract [en]

This paper presents the objectives, organization and activities of the UniDive COST Action, a scientific network dedicated to universality, diversity and idiosyncrasy in language technology. We describe the objectives and organization of this initiative, the people involved, the working groups and the ongoing tasks and activities. This paper is also an open call for participation towards new members and countries. 

Place, publisher, year, edition, pages
European Language Resources Association (ELRA), 2024
Keywords
Diversity; Idiosyncrasy; Language technology; New members; Scientific networks; Universality; Working groups
National Category
Economics and Business
Identifiers
urn:nbn:se:ri:diva-74914 (URN)2-s2.0-85195240358 (Scopus ID)9782493814296 (ISBN)
Conference
3rd Annual Meeting of the ELRA-ISCA Special Interest Group on Under-Resourced Languages, SIGUL 2024 at LREC-COLING 2024
Note

This paper is funded by the CA21167 COST Action UniDive, supported by COST (European Cooperation in Science and Technology).

Available from: 2024-08-19 Created: 2024-08-19 Last updated: 2025-09-23Bibliographically approved
Li, N., Zahra, S., de Brito, M. M., Flynn, C. M., Görnerup, O., Worou, K., . . . Nivre, J. (2024). Using LLMs to Build a Database of Climate Extreme Impacts. In: ClimateNLP 2024 - 1st Workshop on Natural Language Processing Meets Climate Change, Proceedings of the Workshop: . Paper presented at 1st Workshop on Natural Language Processing Meets Climate Change, ClimateNLP 2024. Bangkok, Thailand. 16 August 2024 (pp. 93-110). Association for Computational Linguistics (ACL)
Open this publication in new window or tab >>Using LLMs to Build a Database of Climate Extreme Impacts
Show others...
2024 (English)In: ClimateNLP 2024 - 1st Workshop on Natural Language Processing Meets Climate Change, Proceedings of the Workshop, Association for Computational Linguistics (ACL) , 2024, p. 93-110Conference paper, Published paper (Refereed)
Abstract [en]

To better understand how extreme climate events impact society, we need to increase the availability of accurate and comprehensive information about these impacts. We propose a method for building large-scale databases of climate extreme impacts from online textual sources, using LLMs for information extraction in combination with more traditional NLP techniques to improve accuracy and consistency. We evaluate the method against a small benchmark database created by human experts and find that extraction accuracy varies for different types of information. We compare three different LLMs and find that, while the commercial GPT-4 model gives the best performance overall, the open-source models Mistral and Mixtral are competitive for some types of information.

Place, publisher, year, edition, pages
Association for Computational Linguistics (ACL), 2024
Keywords
Computational linguistics; Database systems; Open systems; Benchmark database; Climate event; Climate extremes; Comprehensive information; Extraction accuracy; Extreme climates; Human expert; Large-scale database; Open-source model; Performance; Data accuracy
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:ri:diva-76190 (URN)2-s2.0-85204502136 (Scopus ID)
Conference
1st Workshop on Natural Language Processing Meets Climate Change, ClimateNLP 2024. Bangkok, Thailand. 16 August 2024
Note

The research presented in this paper was supported by the Swedish Research Council (grants no. 2022-02909, 2022-03448 and 2022-06599). Ni Li is supported by the VUB Research Council in the framework of a EUTOPIA inter-university co-tutelle PhD program between the Vrije Universiteit Brussel, Belgium, and the Technische Universit\u00E4t Dresden, Germany. The EUTOPIA alliance is part of the European Universities Initiatives co-funded by the European Union. The experiments with the open-source LLMs were enabled by the National Academic Infrastructure for Supercomputing in Sweden (NAISS), partially funded by the Swedish Research Council through grant agreement no. 2022-06725. We thank NAISS for providing computational resources under Project 2024/22-211.

Available from: 2024-11-18 Created: 2024-11-18 Last updated: 2025-12-08Bibliographically approved
Projects
Syntactic parsing of synthetic languages [2008-02073_VR]; Uppsala UniversityThe Gender and Work Database at Uppsala university [2010-06012_VR]; Uppsala UniversityUniversal Dependency Parsing [2016-01817_VR]; Uppsala UniversityModular multilingual models [2022-02909_VR]; Uppsala UniversityEn avancerad databas av effekterna av extrema klimathändelser i Europa från nättexter [2022-03448_VR]; Uppsala UniversityCentre of excellence on Impacts of Climate Extremes under global change (ICE) [2022-06599_VR]; Uppsala University
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-7873-3971

Search in DiVA

Show all publications