Deep Diacritization: Efficient Hierarchical Recurrence for Improved Arabic Diacritization

We propose a novel architecture for labelling character sequences that\nachieves state-of-the-art results on the Tashkeela Arabic diacritization\nbenchmark. The core is a two-level recurrence hierarchy that operates on the\nword and character levels separately---enabling faster training and inference\nthan comparable traditional models. A cross-level attention module further\nconnects the two, and opens the door for network interpretability. The task\nmodule is a softmax classifier that enumerates valid combinations of\ndiacritics. This architecture can be extended with a recurrent decoder that\noptionally accepts priors from partially diacritized text, which improves\nresults. We employ extra tricks such as sentence dropout and majority voting to\nfurther boost the final result. Our best model achieves a WER of 5.34%,\noutperforming the previous state-of-the-art with a 30.56% relative error\nreduction.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC