Explaining grain quality trait importance and classification of traditional rice cultivars of Assam, India using SHAP-enhanced Random Forest, XGBoost, and Support Vector Machine models

No hay miniatura disponible

Fecha

Título de la revista

ISSN de la revista

Título del volumen

Editor

Elsevier

Resumen

Descripción

Grain quality is a major determinant of consumer preference and breeding priorities in rice. Assam, a hotspot of traditional rice diversity, harbours distinct Sali rice classes including Common Sali, Scented (Joha), Glutinous (Bora), Semi-glutinous (Chokuwa), and Deep Water (Bao), each characterized by unique physicochemical and culinary attributes. In this study, 130 traditional genotypes were evaluated for 24 grain quality traits encompassing physical, starch, and cooking/pasting characteristics. Nested ANOVA revealed significant variation among classes, with glutinous types exhibiting low amylose and high peak viscosity, scented types showing superior head rice recovery, and Common Sali types characterized by greater hardness. To complement classical statistical analyses, three supervised machine learning algorithms viz. Random Forest (RF), Extreme Gradient Boosting (XGBoost), and Support Vector Machine (SVM) were applied for multi-class classification. Using 23 informative traits after feature filtering and addressing class imbalance through class weighting and the Synthetic Minority Over-sampling Technique (SMOTE), Random Forest achieved the highest predictive performance (test accuracy = 97.2%, Cohen's kappa = 0.962), followed by SVM, whereas XGBoost showed comparatively lower performance. Models were robust to feature reduction, with minimal differences observed between 23-and 21-trait configurations. Discriminative ability was high for RF and SVM (macro-average AUC >0.95), while XGBoost exhibited comparatively lower AUC. SHapley Additive exPlanations (SHAP) consistently identified hardness, pasting temperature, amylose content, adhesiveness, chalkiness, and alkali spreading value as key drivers of class differentiation. The strong concordance between statistical analysis, predictive modelling, and known starch-quality genetics underscores the biological relevance of these traits. Integrating classical statistics with explainable machine learning thus provides a robust framework for trait prioritization, supporting breeding and value enhancement of Assam's traditional rice germplasm.

Palabras clave

ecotypes, machine learning, germplasm, plant breeding, Assam, India

Citación