Unified Image And Text Learning With Beit-3 For Multimodal Classification

The unified multimodal deep learning approach for plant disease prediction utilizing text and picture modalities in BEiT-3 is proposed in this research. Off-the-shelf classification systems relying on image data alone typically fall prey to accuracy and robustness issues when image quality varies or metadata is not exhaustive. To alleviate such concerns, the proposed model incorporates the cross-attention mechanism of BEiT-3, which provides suitable encoders to jointly learn the representation and their accompanying text descriptions. LAMB is chosen as the optimizer during training to promote stability and maintain growth in convergence with bigger training sets. Accuracy, precision, recall, $\mathbf{F 1}$-score, and $\mathbf{R O C}$-AUC are popular performance metrics used to assess the model on a specific multimodal plant disease dataset. According to experimental data, the model can discriminate between plants that are infected and those that are not, with an accuracy of 98.20%, precision of 97.80%, and AUC of 0.9821. This demonstrates the efficacy of vision-language model fusion toward advanced agricultural disease diagnosis. The framework proposed shows promise for scalable deployment toward precision farming and smart agricultural systems.

Paper

Full text

PDF

Unified Image And Text Learning With Beit-3 For Multimodal Classification

Semantic Scholar · 2025

Abstract

The unified multimodal deep learning approach for plant disease prediction utilizing text and picture modalities in BEiT-3 is proposed in this research. Off-the-shelf classification systems relying on image data alone typically fall prey to accuracy and robustness issues when image quality varies or metadata is not exhaustive. To alleviate such concerns, the proposed model incorporates the cross-attention mechanism of BEiT-3, which provides suitable encoders to jointly learn the representation and their accompanying text descriptions. LAMB is chosen as the optimizer during training to promote stability and maintain growth in convergence with bigger training sets. Accuracy, precision, recall, $\mathbf{F 1}$-score, and $\mathbf{R O C}$-AUC are popular performance metrics used to assess the model on a specific multimodal plant disease dataset. According to experimental data, the model can discriminate between plants that are infected and those that are not, with an accuracy of 98.20%, precision of 97.80%, and AUC of 0.9821. This demonstrates the efficacy of vision-language model fusion toward advanced agricultural disease diagnosis. The framework proposed shows promise for scalable deployment toward precision farming and smart agricultural systems.

Similar papers

© 2026 NYSGPT2525 LLC