Please use this identifier to cite or link to this item: http://dr.iiserpune.ac.in:8080/xmlui/handle/123456789/9837
Title: Graph Representation Learning for Binding Pocket Prediction in Proteins
Other Titles: Geometric Pre-Training of GNNs for Structure-Based Drug Design
Authors: Laha, Arnab K.
KARAMPURI, YASH
Dept. of Mathematics
20201105
Keywords: Graph Neural Networks (GNNs)
Graph Representation Learning
Protein-Ligand Interaction
Structure-Based Drug Design
Node and Graph Embeddings
Issue Date: May-2025
Citation: 125
Abstract: This study explores several key concepts in graph-based learning and applies them to the problem of ligand-binding pocket prediction and clustering on protein surfaces. First, we investigate graph embedding techniques, scalable feature learning with algorithms like Node2vec, and graph representation learning methods. We then explore neighborhood reconstruction methods and how multi-relational data and knowledge graphs can be used, building a solid foundation for applying graph-based techniques to biological data. Next, we focus on the problem of predicting and clustering ligand-binding pockets on protein surfaces. Using a graph-based approach, we first generate a set of evenly spaced points on the protein’s Solvent Accessible Surface (SAS) with a fast algorithm from the CDK library. For each point, we calculate feature descriptors based on the local chemical environment, including properties of solvent-exposed atoms, distance-weighted properties of nearby atoms (within 6A), and other neighborhood features. These descriptors are used to predict ligandability scores through Graph Neural Networks (GNN) and Graph Convolutional Networks (GCN). Points with high ligandability scores are then clustered using single-linkage clustering with a 3A cut-off to form pocket predictions. The predicted pockets are ranked by their cumulative ligandability scores. This method provides an efficient framework for identifying potential ligand-binding pockets, contributing to drug discovery and protein-ligand interaction studies.
Description: All Python scripts used throughout this research, including those related to the supplementary studies, are available in the following GitHub repository: https://github.com/YashKarampuri/MSThesis-Supplementary.
URI: http://dr.iiserpune.ac.in:8080/xmlui/handle/123456789/9837
Appears in Collections:MS THESES

Files in This Item:
File Description SizeFormat 
20201105_Karampuri Yash_MS_Thesis.pdfMS Thesis2.07 MBAdobe PDFView/Open


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.