Loflòc: A Morphological Lexicon for Occitan using Universal Dependencies

Marianne Vergez-Couret, Myriam Bras, Aleksandra Miletić, Clamença Poujade

Tutkimustuotos: Artikkeli kirjassa/raportissa/konferenssijulkaisussaKonferenssiartikkeliTieteellinenvertaisarvioitu

Abstrakti

This paper presents Loflòc (Lexic obèrt flechit Occitan - Open Inflected Lexicon of Occitan), a morphological lexicon for Occitan. Even though the lexicon no longer occupies the same place in the NLP pipeline since the advent of large language models, it remains a crucial resource for low-resourced languages. Occitan is a Romance language spoken in the south of France and in parts of Italy and Spain. It is not recognized as an official language in France and no standard variety is shared across the area. To the best of our knowledge, Loflòc is the first publicly available lexicon for Occitan. It contains 650 thousand entries for 57 thousand lemmas. Each entry is accompanied by the corresponding Universal Dependencies Part-of-Speech tag. We show that the lexicon has solid coverage on the existing freely available corpora of Occitan in four major dialects. Coverage gaps on multi-dialect corpora are overwhelmingly driven by dialectal variation, which affects both open and closed classes. Based on this analysis we propose directions for future improvements.

Alkuperäiskielienglanti
OtsikkoProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
ToimittajatNicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, Nianwen Xue
Sivumäärä9
JulkaisupaikkaParis
KustantajaEuropean Language Resources Association (ELRA)
Julkaisupäivä2024
Sivut10716-10724
ISBN (elektroninen)978-2-493814-10-4
TilaJulkaistu - 2024
OKM-julkaisutyyppiA4 Artikkeli konferenssijulkaisuussa
TapahtumaThe 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) - Torino, Italia
Kesto: 20 toukok. 202425 toukok. 2024

Julkaisusarja

NimiInternational conference on computational linguistics
KustantajaInternational Committee on Computational Linguistics
ISSN (painettu)2951-2093
NimiLREC proceedings
KustantajaLanguage Resources Association (ELRA
ISSN (elektroninen)2522-2686

Lisätietoja

Publisher Copyright:
© 2024 ELRA Language Resource Association: CC BY-NC 4.0.

Tieteenalat

  • 6121 Kielitieteet
  • 113 Tietojenkäsittely- ja informaatiotieteet

Siteeraa tätä