Abstrakti
We present an evaluation of several state-of-the-art machine translation systems supporting terminology constraints in the English–Finnish translation direction. We first perform a meta-evaluation, in which we critically evaluate the evaluation metrics we use, including the questions asked of human evaluators and the automatic evaluation methods. We find that common metrics such as term accuracy and TERm do not agree with the human evaluators’ judgement on the correctness of the terms, while LLM-as-a-judge shows promise even though it does not agree with the human evaluators on all questions. We then compare the evaluated systems based on the human evaluation results, LLM-as-a-judge, COMET, and chrF2. We find that of the systems considered, soft constraint methods, including a term-trained model and an LLM, perform better than hard constraints forced using a constrained beam search.
| Alkuperäiskieli | englanti |
|---|---|
| Otsikko | Proceedings of the 26th Annual Conference of The European Association for Machine Translation |
| Kustantaja | European Association for Machine Translation |
| Julkaisupäivä | 15 kesäk. 2026 |
| Tila | Julkaistu - 15 kesäk. 2026 |
| OKM-julkaisutyyppi | A4 Artikkeli konferenssijulkaisuussa |
Tieteenalat
- 113 Tietojenkäsittely- ja informaatiotieteet
Siteeraa tätä
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver