Machine learning identifies genomic regions associated with yield in Coffea arabica

Vol. 6, 2025 - 344732
Simple Abstract
Favorite this paper
How to cite this paper?
Abstract

Genome-wide association studies have been widely used to identify genomic regions controlling agronomic traits, but conventional linear and mixed models may have limited power for polygenic traits governed by small effects, nonlinear relationships, and epistatic interactions. This study aimed to evaluate machine-learning methods as complementary tools for marker identification and to investigate genomic regions associated with yield in Coffea arabica. A population of 195 genotypes derived from crosses between the Catuaí group and Timor Hybrid was evaluated for yield in 2014, 2015, and 2016. Genotyping generated 21,211 single-nucleotide polymorphisms, which were filtered according to call rate, minor allele frequency, and observed heterozygosity. Decision tree, bagging, random forest, boosting, and multivariate adaptive regression splines were compared with conventional and multilocus genome-wide association models. Marker importance was standardized, and high-confidence genomic regions were defined by recurrence across at least two model configurations and by a linkage-disequilibrium window of 205 kilobases. Simulated traits controlled by 8 to 240 quantitative trait loci under heritabilities of 0.5 and 0.8 were also analyzed to verify detection power, precision, false-positive control, and computational efficiency. In the coffee population, no marker surpassed the Bonferroni threshold in the conventional association analyses, whereas machine learning selected 53 markers grouped into 30 genomic regions. The marker located on chromosome 4 at position 2,724,101 was recovered by five algorithms and represented the strongest consensus signal. Other recurrent regions were detected on chromosomes 3, 5, 7, 9, and 11. Functional annotation identified biologically relevant candidate genes related to sugar transport, carbon allocation, plant development, and responses to environmental stress. The simulated analyses supported the empirical findings: bagging and random forest retained detection power above 90% in the most polygenic scenario, while multivariate adaptive regression splines with additive terms and the most specific multilocus association model achieved specificity above 99%. These results demonstrated clear complementarity among methods. Tree-based ensembles were more appropriate for broad screening because they reduced false negatives, whereas more specific methods and cross-model consensus were effective for prioritizing reliable candidate regions. The combined strategy expanded the identification of potentially functional genomic regions for coffee yield and provided a reproducible framework for studying complex traits in plant breeding.

Share your ideas or questions with the authors!

Did you know that the greatest stimulus in scientific and cultural development is curiosity? Leave your questions or suggestions to the author!

Sign in to interact

Have a question or suggestion? Share your feedback with the authors!

Institutions
  • 1 UFV - Universidade Federal de Viçosa
  • 2 Empresa Brasileira de Pesquisa Agropecuária
  • 3 Universidade Federal de Viçosa
Track
  • 8. Genome-wide selection and association
Keywords
machine learning
genome-wide association
marker importance
coffee yield
candidate genes