A scalable workflow for variant effect prediction

Vol 1, 2023 - 164834
Abstract
Favorite this paper
How to cite this paper?
Abstract

INTRODUCTION: The evolution of life is strongly shaped by mutations, which can provide a higher capability of surviving in harsh environments. One of the main reasons for mutations to occur is due to environmental pressure, when it comes to bacteria, it could be a change of the environment pH or the presence/absence of a specific compound. Nowadays, we are facing a growth of the antimicrobial resistance in many organisms, one of them is the Pseudomonas aeruginosa, a gram-negative bacteria, which is associated with cases of nosocomial infection. The scientific community has been working in ways to better understand the impact of the mutations in proteins and in bacteria to overcome this issue. Various approaches for assessing the impact of mutations often yield widely different results. Large-scale comparative analyses of genomic variation enabled through recent progress in structure prediction combined with the growing number of genomic sequences in public databases are currently hampered by an absence of a unified analysis framework. OBJECTIVE: Our goal was to create a pipeline implementing a diverse set of programs for variant impact analysis and to compare the quality of the predictions between different methods. METHODS: We bundled SIFT, UNET, ESM1v, mCSM for protein-protein interactions, mCSM for protein stability, and mCSM for protein-DNA interactions, FoldX, SDM and the BLOSUM-62 matrix distance providing unified inputs and outputs and rich visualisations using the Snakemake workflow manager. We then validated our approach by applying it to mutational scanning data from the MaveDB database and a large variant set from an opportunistic pathogenic bacteria. DISCUSSION AND RESULTS: We found that ESM1v showed the highest correlation with the ground truth followed by SIFT and FoldX and exhibited the best discrimination between adaptive deleterious from non-deleterious Pseudomonas variants. Comparative analysis showed higher concordance between tools based on structural modelling (mSCM and FoldX) and sequence homology-based tools (ESM1v and SIFT). Our highly customizable software paves the way to leverage large-scale data sets for variant impact analysis and opens up an avenue towards combining the predictive power of different approaches.

Share your ideas or questions with the authors!

Did you know that the greatest stimulus in scientific and cultural development is curiosity? Leave your questions or suggestions to the author!

Sign in to interact

Have a question or suggestion? Share your feedback with the authors!

Track
  • 3. Drug design and delivery
Keywords
High throughput; nsSNP; protein stability