Stephanie Wankowicz
@stephanieaw
Computational Structural Biologist Scientific Program Director, diffUSE Project (she/her) diffuse.science Past: Vanderbilt, UCSF, Dana-Farber, Broad Institute, UMass Amherst
We are reframing the PDB as a source of experimental ensemble information, providing ensemble prediction with additional training data. This work is only the beginning. We are thinking iteratively about prediction/modeling in structural biology to extract more of this ensemble data.
We ran qFit across high-resolution PDB entries with deposited structure factors, producing 60,000+ ensemble models, the largest experimentally derived set to date. 83.7% fit the data better than the deposited structure, revealing side-chain heterogeneity on the surface. thestacks.org/publications...
Most PDB structures report one set of coordinates. But the experimental data behind them, in both X-ray crystallography and cryo-EM, is produced by an ensemble of structures. How to extract this ensemble data at scale has been a quest in structural biology for 40+ years.
While there were some differences in the structure predictors, with Boltz-2 fine-tuned on molecular dynamics being the most "guidable", most of the bias depended on the protein being explored and on how much conformational diversity existed in the training set, hinting needing better training data.
To test the platform, we asked different structure predictors how well they could predict "altloc" segments from 40 high-res crystal structures. Because altlocs are stripped from the training data, these are physically real but "out of the training set."
sampleworks lets you swap structure predictors, guidance methods, and loss functions (experimental data). While there are multiple methods for steering structure predictors, each is tied to ONE model + ONE form of guidance, with significant engineering overhead to switch between them.
To test the platform, we asked different structure predictors how well they could predict "altloc" segments from 40 high-res crystal structures. Because altlocs are stripped from the training data, these are physically real but "out of the training set."
sampleworks lets you swap structure predictors, guidance methods, and loss functions (experimental data). While there are multiple methods for steering structure predictors, each is tied to ONE model + ONE form of guidance, with significant engineering overhead to switch between them.
Closing this chapter means leaving an amazing team. These people took a chance, worked hard, and made the lab a fantastic place to work everyday.
Episode 13 Tortured Protein Department Podcast: Long Live with @fraserlab.com! Catch my big news with the @diffuseproject.bsky.social, psycho travel stories, April Fools pranks, and AI grad students. podcasts.apple.com/us/podcast/l...
We observed a tradeoff: weak guidance improves fit to data without wrecking protein geometry. But strong guidance destroys the geometry. Interestingly, models differ a lot in how they handle this despite similar architectures.
We explored the extent to which diffusion-based structure predictors can be guided to generate accurate structural ensembles, asking whether these models can produce out-of-distribution conformations and, when they do, how well they balance between fit to exp. data and geometry.
Structural interpretations of protein–ligand binding have historically emphasized enthalpy because entropy cannot be directly visualized in static structures. We show that by modeling crystallographic ensembles, we can make this visible.
New Preprint!! We show that binding entropy can be quantitatively predicted from crystallographic ensemble models, accounting for both protein conformational entropy and solvent entropy! www.biorxiv.org/content/10.6...
5) Structural entropy proxies correlate with ITC-measured binding entropy. We show that ensemble-derived measures of conformational entropy quantitatively track with experimental binding entropy. This establishes a link between crystallographic ensembles and binding thermodynamics.
4) Solvent network remodeling parallels changes in protein entropy. Ligands that induce greater heterogeneity also release more water molecules and produce less connected protein–solvent hydrogen-bond networks, suggesting a coordinated, system-level redistribution of protein and solvent entropy.
3) Ligands that share similar interaction fingerprints (determined by OT algorithm) have conserved patterns of conformational heterogeneity and solvent reorganization across the protein.
2) We created a fused Gromov–Wasserstein optimal transport algorithm to capture protein-ligand interactions, not just ligand and/or protein information.
1) Ligand binding reshapes protein conformational entropy in reproducible ways. We identify three intrinsic “entropic reservoir” regions whose heterogeneity increases upon ligand binding. These areas were already flexible in apo proteins, potentially suggesting some 'dynamic' conservation.
@vanallenlab.bsky.social 10 year reunion with over 50 people and dinner with part of cohort one. I miss doing science with these folks. ❤️
Our lab’s guide from Light Hall to MRB3. I hope others in @vubasicsciences.bsky.social can appreciate!
I love this, but I am seeing examples where the widget says there are comments, but when I click on the comments there are none there.
This seems to be due to more careful manual modeling as detected by more alternative conformations placed in binding sites and rotamer outliers identified in binding sites as having better density support.