Hierarchical Phrase-based Translation with Weighted Finite State Transducers and Shallow-N Grammars

Authors

  • Adrià de Gispert University of Cambridge
  • Gonzalo Iglesias University of Vigo
  • Graeme Blackwood University of Cambridge
  • Eduardo R. Banga University of Vigo
  • William Byrne University of Cambridge

Abstract

In this paper we describe HiFST, a lattice-based decoder for hierarchical phrase-based translation and alignment. The decoder is implemented with standard Weighted Finite-State Transducer (WFST) operations as an alternative to the well-known cube pruning procedure. We find that the use of WFSTs rather than k-best lists requires less pruning in translation search, resulting in fewer search errors, better parameter optimization, and improved translation performance. The direct generation of translation lattices in the target language can improve subsequent rescoring procedures, yielding further gains when applying long-span language models and Minimum Bayes Risk decoding. We also give insight as to how to control the size of the search space defined by hierarchical rules. We show that shallow-N grammars, low-level rule catenation and other search constraints can help to match the power of the translation system to specific language pairs.

Published

2024-12-05

Issue

Section

Long paper