The unified multimodal deep learning approach for plant disease prediction utilizing text and picture modalities in BEiT-3 is proposed in this research. Off-the-shelf classification systems relying on image data alone typically fall prey to accuracy and robustness issues when image quality varies or metadata is not exhaustive. To alleviate such concerns, the proposed model incorporates the cross-attention mechanism of BEiT-3, which provides suitable encoders to jointly learn the representation and their accompanying text descriptions. LAMB is chosen as the optimizer during training to promote stability and maintain growth in convergence with bigger training sets. Accuracy, precision, recall, $\mathbf{F 1}$-score, and $\mathbf{R O C}$-AUC are popular performance metrics used to assess the model on a specific multimodal plant disease dataset. According to experimental data, the model can discriminate between plants that are infected and those that are not, with an accuracy of 98.20%, precision of 97.80%, and AUC of 0.9821. This demonstrates the efficacy of vision-language model fusion toward advanced agricultural disease diagnosis. The framework proposed shows promise for scalable deployment toward precision farming and smart agricultural systems.
Paper
Full text
Unified Image And Text Learning With Beit-3 For Multimodal Classification
Semantic Scholar · 2025
Abstract
The unified multimodal deep learning approach for plant disease prediction utilizing text and picture modalities in BEiT-3 is proposed in this research. Off-the-shelf classification systems relying on image data alone typically fall prey to accuracy and robustness issues when image quality varies or metadata is not exhaustive. To alleviate such concerns, the proposed model incorporates the cross-attention mechanism of BEiT-3, which provides suitable encoders to jointly learn the representation and their accompanying text descriptions. LAMB is chosen as the optimizer during training to promote stability and maintain growth in convergence with bigger training sets. Accuracy, precision, recall, $\mathbf{F 1}$-score, and $\mathbf{R O C}$-AUC are popular performance metrics used to assess the model on a specific multimodal plant disease dataset. According to experimental data, the model can discriminate between plants that are infected and those that are not, with an accuracy of 98.20%, precision of 97.80%, and AUC of 0.9821. This demonstrates the efficacy of vision-language model fusion toward advanced agricultural disease diagnosis. The framework proposed shows promise for scalable deployment toward precision farming and smart agricultural systems.