Repository logo
  • English
  • Deutsch
  • Français
Log In
New user? Click here to register.Have you forgotten your password?
  1. Home
  2. CRIS
  3. Publication
  4. Lextreme: A multi-lingual and multi-task benchmark for the legal domain
 

Lextreme: A multi-lingual and multi-task benchmark for the legal domain

URI
https://arbor.bfh.ch/handle/arbor/36270
Version
Published
Date Issued
2023-12-10
Author(s)
Niklaus, Joël  
Matoshi, Veton  
Rani, Pooja
Galassi, Andrea
Stürmer, Matthias  
Chalkidis, Ilias
Type
Conference Paper
Language
English
Abstract
Lately, propelled by the phenomenal advances around the transformer architecture, the legal NLP field has enjoyed spectacular growth. To measure progress, well curated and challenging benchmarks are crucial. However, most benchmarks are English only and in legal NLP specifically there is no multilingual benchmark available yet. Additionally, many benchmarks are saturated, with the best models clearly outperforming the best humans and achieving near perfect scores. We survey the legal NLP literature and select 11 datasets covering 24 languages, creating LEXTREME. To provide a fair comparison, we propose two aggregate scores, one based on the datasets and one on the languages. The best baseline (XLM-R large) achieves both a dataset aggregate score a language aggregate score of 61.3. This indicates that LEXTREME is still very challenging and leaves ample room for improvement. To make it easy for researchers and practitioners to use, we release LEXTREME on huggingface together with all the code required to evaluate models and a public Weights and Biases project with all the runs.
DOI
10.24451/arbor.22335
https://doi.org/10.24451/arbor.22335
Publisher DOI
10.18653/v1/2023.findings-emnlp.200
Publisher URL
https://aclanthology.org/2023.findings-emnlp.200/
Related URL
https://2023.emnlp.org/ org https://arxiv.org/abs/2301.13126
Organization
Institut Public Sector Transformation (IPST)  
Data and Infrastructure  
Wirtschaft  
Conference
Findings of the Association for Computational Linguistics: EMNLP 2023
Publisher
Cornell University
Submitter
NiklausJ
Citation apa
Niklaus, J., Matoshi, V., Rani, P., Galassi, A., Stürmer, M., & Chalkidis, I. (2023). Lextreme: A multi-lingual and multi-task benchmark for the legal domain. Findings of the Association for Computational Linguistics: EMNLP 2023. Cornell University. https://doi.org/10.24451/arbor.22335
File(s)
Loading...
Thumbnail Image
Download

open access

Name

2301.13126.pdf

License
Attribution-NonCommercial-ShareAlike 4.0 International
Size

2.05 MB

Format

Adobe PDF

Checksum (MD5)

0d3e963932695f81f31dcc164fd3d7df

Loading...
Thumbnail Image
Download

open access

Name

2023.findings-emnlp.200.pdf

License
Attribution 4.0 International
Version
published
Size

2.05 MB

Format

Adobe PDF

Checksum (MD5)

1ab5470558b036507ada09c2e05caf7a

About ARBOR

Built with DSpace-CRIS software - System hosted and mantained by 4Science

  • Cookie settings
  • Privacy policy
  • End User Agreement
  • Send Feedback
  • Our institution