AraVirusPPI: A protein language model-powered method for predicting Arabidopsis thaliana-virus protein-protein interactions

Abstract

Plant-virus protein-protein interactions (PPIs) play crucial roles in viral infection and host immune responses. However, large-scale experimental characterization of plant-virus PPIs is costly and labor-intensive, suggesting the need to develop computational methods. Recent advances in biological sequence language models are offering unprecedented opportunities to develop machine learning-based plant-virus PPI predictors. In this study, we collected publicly available PPI data between Arabidopsis thaliana (A. thaliana) and viruses, establishing a robust dataset for method development. Based on this dataset, we developed AraVirusPPI, a machine learning framework for predicting A. thaliana-virus PPIs. AraVirusPPI employs the protein language model ESM Cambrian (ESMC) to encode sequence features and combines these representations with Extreme Gradient Boosting (XGBoost) to build the prediction model. Furthermore, we systematically evaluated multiple protein and genomic language models combined with various classification algorithms, thereby justifying the optimal model architecture adopted in AraVirusPPI. We applied AraVirusPPI to predict PPIs in the A. thaliana–Cabbage leaf curl virus (CabLCV) system. Predicted viral target proteins were compared with experimental datasets from A. thaliana-Turnip mosaic virus (TuMV) and A. thaliana-Pseudomonas syringae (Psy) systems. Network topology and functional enrichment analyses showed that the predicted CabLCV targets in A. thaliana were more similar to TuMV than Psy targets. These results suggest potential shared and distinct host-targeting patterns among pathogens and provide computational evidence supporting the biological significance of the proposed AraVirusPPI predictor.