WikiHow-MY: A Human Post-Edited English–Myanmar Instructional MT Corpus and an Instruction-Faithfulness Score for Procedural Translation

We present WikiHow-MY, to our knowledge the first instructional/how-to English-Myanmar parallel corpus: ~10k human post-edited (MTPE), quality-checked sentence pairs from WikiHow, released by rehydration under CC BY-NC-SA 3.0 with article-disjoint splits. On this resource we benchmark NLLB-200 (zero-shot and fine-tuned), Google Translate, and Gemini 2.5 with chrF++/spBLEU/BLEU/COMET, and find that surface and semantic metrics disagree on the best system. We then introduce the Instruction Faithfulness Score (IFS), an interpretable, source-anchored metric that decomposes procedural fidelity into step, action, entity, and quantity preservation, and use it as a diagnostic in the first human instruction-followability study for a low-resource language (9 raters, 420 ratings), where learned metrics such as COMET are known to be poorly calibrated. The finding generalizes beyond any single metric: none of the metrics we test, surface (chrF++/BLEU) or our procedural IFS, is correlated with human followability at the segment level (r ~ 0), and surface-saturating metrics, IFS included, cannot rank systems: IFS spans just 2.25 points across the four systems, awards its highest score to the fine-tuned model humans place third of four, and ranks humans' best system only third. We then show this is a failure of estimation, not of measurability: keeping IFS's four procedural dimensions but estimating them with a learned per-dimension judge rather than surface preservation recovers human followability (r = 0.53 vs. 0.13) and the system ranking (Spearman 1.00 vs. 0.20), including when the judge's own outputs are held out. Our contribution is thus both a diagnostic when-metrics-fail result for procedural MT, and its constructive counterpart: followability is recoverable when the same procedural dimensions are estimated per dimension by judgment. Finally, we probe whether decoding-time reranking can close the in-domain gap to the commercial system, and report a negative result: although a reference-optimal selection from the fine-tuned model's n-best beats Google Translate, no reference-free signal (learned quality estimation, structural IFS, or MBR) recovers this gain, a concrete decoding-time symptom of learned-metric miscalibration for Burmese.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC