Confidence-Guided Diffusion Augmentation for Enhanced Bangla Compound Character Recognition
Researchers Md. Sultan Al Rayhan and Maheen Islam have proposed a novel confidence-guided diffusion augmentation framework to address challenges in recognizing handwritten Bangla compound characters. These characters are difficult to identify due to complex structures, significant intra-class variation, and a scarcity of high-quality annotated data. The new framework integrates class-conditional diffusion modeling with classifier guidance to synthesize high-quality handwritten samples. It features Squeeze-and-Excitation enhanced residual blocks within the U-Net backbone and employs a confidence-based filtering mechanism to ensure only class-consistent synthetic images are used. By fusing these filtered synthetic images with original training data, the team retrained multiple classification architectures, including ResNet50, DenseNet121, VGG16, and Vision Transformer. Experiments on the AIBangla dataset showed consistent performance improvements, with the best model achieving 89.2% accuracy, significantly surpassing previous benchmarks. This study demonstrates that quality-aware diffusion augmentation can effectively enhance character recognition in low-resource script domains, offering a robust solution for improving optical character recognition systems for complex scripts.
Wire timeline
Confidence-Guided Diffusion Augmentation for Enhanced Bangla Compound Character Recognition
Researchers Md. Sultan Al Rayhan and Maheen Islam have proposed a novel confidence-guided diffusion augmentation framework to address challenges in recognizing handwritten Bangla compound characters. These characters are difficult to identify due to complex structures, significant intra-class variation, and a scarcity of high-quality annotated data. The new framework integrates class-conditional diffusion modeling with classifier guidance to synthesize high-quality handwritten samples. It features Squeeze-and-Excitation enhanced residual blocks within the U-Net backbone and employs a confidence-based filtering mechanism to ensure only class-consistent synthetic images are used. By fusing these filtered synthetic images with original training data, the team retrained multiple classification architectures, including ResNet50, DenseNet121, VGG16, and Vision Transformer. Experiments on the AIBangla dataset showed consistent performance improvements, with the best model achieving 89.2% accuracy, significantly surpassing previous benchmarks. This study demonstrates that quality-aware diffusion augmentation can effectively enhance character recognition in low-resource script domains, offering a robust solution for improving optical character recognition systems for complex scripts.
cs.AI updates on arXiv.org