SSF-PREDICT: REFOCUSING ON PUTATIVE KEY RESIDUES TO ESTIMATE PROTEIN FUNCTION

Vol 2, 2024 - 314898
Abstract
Favorite this paper
How to cite this paper?
Abstract
The deluge of protein sequence data, fueled by advancements in high-throughput sequencing and mass spectrometry, has far outpaced our ability to characterize protein function through traditional experiments. To bridge this gap, computational methods for Automated Function Prediction have emerged. These methods often leverage functional classification systems such as Gene Ontology (GO) to transfer knowledge from well-characterized proteins to the vast majority that remain unannotated. However, GO hierarchies can lead to information loss due to the overrepresentation of general terms. Meanwhile, protein family models (PFMs) support many well-established prediction algorithms and databases, including the HMMER suite and Pfam, HHpred and PDB and SCOP, among others. Regardless, both approaches supply users with predicted general function, not pinpointing key functional residues, which could be done based on the PFMs underpinning predictions. Access to residue-level data could strengthen confidence in predictions, while also improving our understanding of protein family evolution and function conservation. In addition, this could classify novel proteins based on functional gain or loss. Therefore, we are developing an automated pipeline, SSF-Predict (Site Specific Function Predict), that analyzes large, genome-sized batches of protein sequences, predicting their function and identifying potentially key residues in a reasonable timeframe. Our pipeline prioritizes understanding protein function through the lens of the query sequence's amino acids, aiming for transparent and modular function prediction. We achieve structurally validated functional transfer for a large subset of annotations by employing the HMMER suite followed by validation using the BioLiP2 database. Further, our inferences are enriched by site-specific coevolutionary analysis. We've laid the groundwork by building the necessary annotation database and intermediary files for efficient data transfer. We've also implemented most of the basic Python functionality, prior to integration in a workflow management system for improved code reproducibility and user-friendliness. Additionally, we're incorporating broader type-specific validation and considering additional functionality to provide more value to end users. Finally, we welcome feedback from the community to continuously improve our pipeline and establish it as a valuable tool for protein functional exploration. This work was supported by Coordenação de Aperfeiçoamento de Pessoal de Nível Superior (CAPES), Programa de Excelência Acadêmica (PROEX).

Share your ideas or questions with the authors!

Did you know that the greatest stimulus in scientific and cultural development is curiosity? Leave your questions or suggestions to the author!

Sign in to interact

Have a question or suggestion? Share your feedback with the authors!

Institutions
  • 1 Federal University of Minas Gerais
  • 2 European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus
Track
  • 1. Protein Dynamics and Function
Keywords
protein functional annotation
amino acid residues
pipeline