Cross-docking and redocking reveal distinct determinants of success in physics-based and AI-driven binding pose prediction in protein–ligand complexes
Kapali Suri, Anshul Yadav, Abhishek Tripathi, N. Arul Murugan
Abstract
Protein-ligand pose prediction is central to structure-based drug discovery, yet the relative performance of physics-based and AI-driven methods under realistic cross-docking conditions remains insufficiently characterized. Here, we compare physics-based docking methods (AutoDock4, AutoDock Vina, and DOCK 6) with data-driven approaches, including the deep-learning model GNINA 1.3 and the diffusion-based frameworks AlphaFold 3, Boltz-2, and DiffDock. Performance was evaluated using standardised redocking and cross-docking protocols across three Alzheimer's disease targets representing distinct binding-site architectures: acetylcholinesterase (AChE; deep gorge), β-secretase 1 (BACE1; flexible flap-controlled site), and glycogen synthase kinase-3β (GSK-3β; open, solvent-exposed pocket). Physics-based methods were competitive during redocking but showed substantial performance reductions under cross-docking, whereas diffusion-based approaches generally maintained higher cross-docking accuracy. GNINA 1.3 rigid achieved an 87.7% minimum heavy-atom RMSD success rate during redocking, which decreased to 13.5% during cross-docking, whereas AlphaFold 3, Boltz-2, and DiffDock achieved cross-docking success rates of 93.1%, 89.6%, and 85.7%, respectively. AlphaFold 3 consistently outperformed Boltz-2 despite its smaller training set, suggesting that predictive performance is influenced not only by training-data volume but also by factors such as model architecture and confidence calibration. Training-overlap analysis further showed that AI-based methods retained substantial failure rates even for complexes represented in their training data, indicating that training-data overlap alone does not ensure reliable pose prediction. Under the current protocol conditions, rigid docking outperformed flexible protocols, while flexible-docking pocket volumes showed more restricted sampling relative to experimental holo structures. Among the GNINA 1.3 configurations, CNN rescoring with refinement produced the highest pose-recovery success rates, followed by CNN rescoring alone and the default Vina/empirical scoring approach in cross-docking.
§ The Valyu brief
Reading the full paper and taking notes. This takes a few seconds…
§ Ask this paper
Ask a question about this paper
Valyu reads the full text and answers from what the paper actually says.
Searching the other archives…