To cite this paper use one of the standards below:
Genome-wide association studies have been widely used to identify genomic regions controlling agronomic traits, but conventional linear and mixed models may have limited power for polygenic traits governed by small effects, nonlinear relationships, and epistatic interactions. This study aimed to evaluate machine-learning methods as complementary tools for marker identification and to investigate genomic regions associated with yield in Coffea arabica. A population of 195 genotypes derived from crosses between the Catuaí group and Timor Hybrid was evaluated for yield in 2014, 2015, and 2016. Genotyping generated 21,211 single-nucleotide polymorphisms, which were filtered according to call rate, minor allele frequency, and observed heterozygosity. Decision tree, bagging, random forest, boosting, and multivariate adaptive regression splines were compared with conventional and multilocus genome-wide association models. Marker importance was standardized, and high-confidence genomic regions were defined by recurrence across at least two model configurations and by a linkage-disequilibrium window of 205 kilobases. Simulated traits controlled by 8 to 240 quantitative trait loci under heritabilities of 0.5 and 0.8 were also analyzed to verify detection power, precision, false-positive control, and computational efficiency. In the coffee population, no marker surpassed the Bonferroni threshold in the conventional association analyses, whereas machine learning selected 53 markers grouped into 30 genomic regions. The marker located on chromosome 4 at position 2,724,101 was recovered by five algorithms and represented the strongest consensus signal. Other recurrent regions were detected on chromosomes 3, 5, 7, 9, and 11. Functional annotation identified biologically relevant candidate genes related to sugar transport, carbon allocation, plant development, and responses to environmental stress. The simulated analyses supported the empirical findings: bagging and random forest retained detection power above 90% in the most polygenic scenario, while multivariate adaptive regression splines with additive terms and the most specific multilocus association model achieved specificity above 99%. These results demonstrated clear complementarity among methods. Tree-based ensembles were more appropriate for broad screening because they reduced false negatives, whereas more specific methods and cross-model consensus were effective for prioritizing reliable candidate regions. The combined strategy expanded the identification of potentially functional genomic regions for coffee yield and provided a reproducible framework for studying complex traits in plant breeding.
With nearly 200,000 papers published, Galoá empowers scholars to share and discover cutting-edge research through our streamlined and accessible academic publishing platform.
Learn more about our products:
This proceedings is identified by a DOI , for use in citations or bibliographic references. Attention: this is not a DOI for the paper and as such cannot be used in Lattes to identify a particular work.
Check the link "How to cite" in the paper's page, to see how to properly cite the paper