Martin Steinegger 🇺🇦
@martinsteinegger
Developing data intensive computational methods • PI @ Seoul National University 🇰🇷 • #FirstGen • he/him • Hauptschüler
@wheelerlab.org is talking about his effort to build a resource for protein dynamics (PDB for MD). MDRepo is a repository to store simulation data. This is really needed to push the needle in design, function and more. Please help make it successful by sharing your MD data. #ismb2026 🌐 mdrepo.org
My student @yewonhan.bsky.social will talk today at 3DSIG about our work on expanding the AlphaFold database to complexes (31M predictions). It starts at 14.20 in the International Ballroom East. Please join. We also brought Marv stickers with us. #ISMB2026
@rolanddunbrack.bsky.social is talking about kinases and interface scoring (ipSAE). AF’s ipTM score avg. PAE over the full area not just the interface, so disorder and extra domains drag it down. ipSAE uses a PAE cutoff to avoid low scores, therefor only consider good interface contacts. #ismb2026
Christopher E. Mason’s keynote at the HitSeq at #ISMB2026 felt like watching an episode of „Sendung ohne Namen“ (does anybody know it?). A crazy amount of ideas, facts, and sci-fi. From environmental metagenomics to space genomics.
Richard Durbin starts the #ISMB2026 with his keynote about BWT for genomic search. A walk-through from BWA to his most recent work on pan-genomes GBWT. www.biorxiv.org/content/10.6.... Congratulations on the accomplishments by a Senior Scientist Award.
ICML is happening in Seoul this year, and I’ve been getting several messages about lab visits. Who else will be in town and would like to meet? @milot.bsky.social lab and mine are planning a dinner on July 7th, after the reception. Let me know if you’re interested!
Meet the Folddisco Marv, designed by Hyunbin Kim, who also developed Folddisco.
Kieran et al. trained a generative protein binder design model. The training is based on Teddymer, a dataset developed by @sooyoung-cha.bsky.social. By treating monomer domains as multimers and clustering them with Foldseek, she created a set that allowed Complexa to learn. 💾 teddymer.foldseek.com
A small number of large clusters capture most of the structural space: the top 1% of representatives cover ~25% of entries, and the top 20% cover ~82%. At the same time, clusters without a PDB multimer match are enriched among smaller clusters, suggesting much of the unexplored interaction. 4/
In the ~8M heterodimer set, prediction success is highest for pairs that are more alike: higher inter-chain sequence identity and smaller length differences both increase the rate of high-confidence models. 3/
Scale alone doesn’t determine success. Although most predictions come from eukaryotic proteomes, the highest fraction of high-confidence complexes is observed in bacteria and archaea. 2/
A key takeaway: this resource dramatically expands the known structural interactome. In nearly every proteome, the number of high-confidence complex predictions surpasses experimental multimer structures in the PDB by one to three orders of magnitude. 1/
AlphaFold database has entered the era of complexes. Together with NVIDIA, DeepMind and EBI, we use ColabFold, OpenFold and MMseqs2-GPU to predict ~31 million complexes (homo & hetro-dimers) resulting in 1.8 million high-quality predictions 📄 research.nvidia.com/labs/dbr/ass... 🌐 alphafold.ebi.ac.uk
Congratulations! But how did you get the Marv logo into figure 1? We needed to remove the logos 🥲
End-to-end protein design in the browser through evedesign. Generate and interactively explore designs in 2D/3D and export them as codon-optimized DNA. The underlying open source framework (released soon) is build to easily add new methods, more on that soon. 🌐 evedesign.bio
Below we show GPU-accelerated Foldseek, searching 128 structures against AFDB50 (54 million structures). On 128 CPU cores this takes ~120 seconds, whereas a single GPU completes it in ~25 seconds. 2/n
@pedrobeltrao.bsky.social here you can see the clustering at 90% identity.
If you use Boltz1/2, BioEmu, Chai1, or other MSA-dependent models, you’re likely using our ColabFold server. Please be considerate! Avoid large submissions across many IPs instead generate the MSA locally. Our server is an old-timer from 2014 and can’t handle that load.
MMseqs2 v18 is out - SIMD FW/BW alignment (preprint soon!) - Sub. Mat. λ calculator by Eric Dawson - Faster ARM SW by Alexander Nesterovskiy - MSA-Pairformer’s proximity-based pairing for multimer prediction (www.biorxiv.org/content/10.1...; avail. in ColabFold API) 💾 github.com/soedinglab/M... & 🐍
Folddisco webserver result view update: - Added description texts for AFDB - Integrated TaxoView taxonomy visualization & filter by @sunjaelee.bsky.social - Inter-residue distance clustering by DBSCAN to explore motif diversity. 🌐 search.foldseek.com/folddisco 📄 www.biorxiv.org/content/10.1...
Today at 5pm, @eunbelivable.bsky.social will present her work on the Big Fantastic Viral Database (BFVD) at #ISMB2025 in BOSC. She also has a poster B-123 (tomorrow, 22nd), so please drop by to have ta chat and grab some stickers! 📄 academic.oup.com/nar/article/...
We provide a user-friendly Folddisco webserver, enabling instant structural motif searches in PDB, AFDB-Proteomes, AFDB50 (available later today), and ESMatlas (ESM30). Explore it here: search.foldseek.com/folddisco 8/9
Folddisco can be applied for PPI interface searches. When querying an interface between antibody chains (gray/black), it successfully identifies matching interfaces within monomeric antibody fragments (cyan), showcasing its potential to detect novel interaction partners and interfaces. 7/9
Folddisco can distinguish functional states. We searched GPCR activation motifs (CWxP, NPxxY, DRY), clearly separating active/inactive states. A search in the AFDB shows ~53% active, closely mirroring experimental PDB 54%, suggesting AlphaFold might follow its training conformation distribution. 6/9
Folddisco can annotate proteins: querying a canonical zinc-finger uncovers an uncharacterized oyster protein and metagenomic proteins. It also detects partial catalytic metal sites in E. coli peptide deformylase. All of these hits would be missed by Foldseek or sequence aligners. 5/9
Folddisco builds indexes faster and smaller than previous tools: indexing AFDB50 (53M structures) takes only ~24h vs. ~20 days (extrapolated) for pyScoMotif. Querying a zinc-finger motif across AFDB50 takes just ~13s, up to 48x faster than pyScoMotif. 4/9
Folddisco accurately detects discontinuous motifs like zinc fingers and segment-based motifs, previously requiring separate tools. Additionally, we built a SCOPe benchmark by sampling conserved residues from families and measuring the recall up to the first false positive. 3/9
Folddisco detects (partial) motifs, allowing for substitutions and angle-length variations, by utilizing an index storing all residue pairs within 20Å encoded as geometric features. For space efficiency, it omits positions and compresses ids through run-length encoding (1.4TB for 53M structures) 2/9
Folddisco finds similar (dis)continuous 3D motifs in large protein structure databases. Its efficient index enables fast uncharacterized active site annotation, protein conformational state analysis and PPI interface comparison. 1/9🧶🧬 📄 www.biorxiv.org/content/10.1... 🌐 search.foldseek.com/folddisco