Back to Research papers
Research paper index

Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?

Osvaldo Quinjica, Eric Bennett, Xinchen Yang, Andrew Schonebaum, Marine Carpuat

arXiv:2608.08283Published August 8, 2026Updated August 12, 20260 citations
  • cs.CL
  • cs.AI

Abstract

Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic evaluation metrics developed for modern languages are reliable in this setting, using translation from Classical Chinese to English as a test case. We introduce a diagnostic framework based on minimal pairs capturing error types salient in scholarly use, probing both reference-based and reference-free metrics for error sensitivity and tolerance to valid variation. We find that all metrics exhibit blind spots, however MetricX24 performs best overall. Our findings highlight the need for more robust and interpretable metrics for historically and culturally distinct translation settings.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.