MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing
Researchers have introduced MULTITEXTEDIT, a new benchmark designed to evaluate cross-lingual degradation in text-in-image editing systems. Addressing the limitations of existing English-centric benchmarks, this study presents a controlled dataset of 3,600 instances covering 12 typologically diverse languages, five visual domains, and seven editing operations. The benchmark isolates language variables by using common visual bases paired with human-edited references and region masks. To detect script-level errors often missed by standard metrics, such as missing diacritics or reversed right-to-left order, the authors developed a Language Fidelity (LSF) metric. This metric utilizes a two-stage Large Vision Model protocol that achieved a high agreement score with native-speaker annotators. Evaluations of twelve open-source and proprietary systems revealed significant performance drops across all models when handling non-English languages, particularly Hebrew and Arabic. The study highlights that while global layout and background fidelity are often preserved, specific script forms are frequently distorted, indicating a pervasive mismatch between semantic intent and pixel-level accuracy in current AI editing tools.
Wire timeline
MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing
Researchers have introduced MULTITEXTEDIT, a new benchmark designed to evaluate cross-lingual degradation in text-in-image editing systems. Addressing the limitations of existing English-centric benchmarks, this study presents a controlled dataset of 3,600 instances covering 12 typologically diverse languages, five visual domains, and seven editing operations. The benchmark isolates language variables by using common visual bases paired with human-edited references and region masks. To detect script-level errors often missed by standard metrics, such as missing diacritics or reversed right-to-left order, the authors developed a Language Fidelity (LSF) metric. This metric utilizes a two-stage Large Vision Model protocol that achieved a high agreement score with native-speaker annotators. Evaluations of twelve open-source and proprietary systems revealed significant performance drops across all models when handling non-English languages, particularly Hebrew and Arabic. The study highlights that while global layout and background fidelity are often preserved, specific script forms are frequently distorted, indicating a pervasive mismatch between semantic intent and pixel-level accuracy in current AI editing tools.
cs.AI updates on arXiv.org