Quaid Morris
@quaidmorris
Computational biology, machine learning, AI, RNA, cancer genomics. My views are my own. He/him/his
SHAP identifies the key features: iPAEs dominate, iPTMs contribute, top global feature is pLDDT. Interestingly the peptide:MHC interface also matters, suggesting successful presentation contributes alongside TCR recognition. Both ensemble means and SDs contribute independently.
And the neoantigen task: cognate vs non-cognate when the peptide differs by ONE residue. Across 8 mutational-scan datasets enFoldX leads the field (median AUC 0.71) — ahead of sequence models, TCRen, and a single-structure version of our own features.
On IMMREP25's unseen-peptide challenge — epitopes ≥4 substitutions from anything published — enFoldX beat every sequence- and structure-based method.
Ten-fold CV on VDJdb: AUC 0.82 (TCR-split) human, 0.98 mouse. And a model trained only on human data also classifies mouse TCR:pMHCs at AUC 0.76 — cross-species transfer learning, a first for this task.
How it works: fold an ensemble (10 seeds × 5 samples), extract 106 structural, confidence, and biophysical features per prediction—focused primarily on the TCR:pMHC interface and peptide–CDR3 contacts—then summarize each as a mean and SD across the ensemble.
Rather than trusting one prediction, enFoldX summarizes an AF3 ensemble. Non-cognate pairs show greater structural disagreement (higher RMSD, higher iPAE), and performance keeps improving as more ensemble members are included, unlike single-structure models.
AF3 hallucinates for both binders and non-binders—but how often it hallucinates is informative. For CMV epitope NLVPMVATV and its M5K mutant, misdocked structures appear in 7/10 seeds for the non-cognate complex, but only rarely for the cognate.
Predicting which TCR recognizes which peptide-MHC is a longstanding AI challenge. Sequence-based models struggle to generalize to new peptides or TCRs, while AF3 often predicts plausible-looking structures even for non-cognate pairs.
Metient is the first multiobjective method to scale to 1000s of nodes. We showcase this on single-cell lineage tracing data in lung adenocarcinoma, where Metient assigns an early role to mediastinal lymph tissue in subsequent metastatic spread.
In a large cancer cohort, Metient posits that 33% of metastatic sites were seeded by multiple tumor clones, suggesting that migration via cancer cell clusters could be common, thus supporting therapies that disrupt their formation.
Metient first maps out a Pareto front of plausible metastatic histories using stochastic optimization. Then it uses new biologically grounded “metastases priors” to score histories without making ad hoc assumptions about metastatic spread.
Metient can help answer key questions about metastatic spread: How often do metastases seed other metastases? How often do multiple clones seed a single metastasis? Is metastatic potential rare, or gained multiple times by the same cancer?
It looks like you are trying to kill all AI pop-ups, would you like to ... <sniff>
I don’t know what “flag” means, but I would very much like to be added to What’s Science @danirabaiotti.bsky.social please! Here’s a cute cat picture for your trouble
ucRBPs are highly abundant in whole-cell mass spectrometry surveys. Weak non-specific RNA interactions coupled with high abundance could explain their prevalence in RNA interactome capture studies.
In contrast, CCHC‐zf domains from seven human proteins recognized specific RNA motifs, indicating that this is a major class of RBD.
Identification and analysis of sequence-specific unconventional RNA binding domains were mostly consistent with RNA‐binding as a derived function.
Only 23 ucRBPs (4.7 %) displayed sequence specificity, including 5 proteins not previously shown to possess RNA binding activity.
Interested in gaining research experience before #gradschool? Consider our #postbac program - the MSK Bridge! Join @MSKEducation for a live info session on Thursday, Jan 19 @ 3 pm (EST). Register: https://tinyurl.com/MSKBridgeInfoSession2023 #PrePhD #PreMDPhD #AcademicTwitter
Out today: COVID forecasting using conditional latent ODE w/ @ianshi3 - easy forecasting of NPI change impacts - data-driven - often sig. better, never sig. worse at death forecasts than all others, incl expert-driven ones P: https://doi.org/10.1093/jamia/ocac160 G:...
Check out our new RBP motif algorithm: PRIESSTESS 🧙♀️ - fits RNA sequence and structure models - is trivially interpretable - generalizes as well or better as black-box...
Congratulations to @mmdarmofal for winning a #transmed best talk award at at #ISMB2022
Multiple genomic features, known to affect mutational processes, are weakly correlated with changepoint locations: gene and mutation density, cell-of-origin chromatin state, copy number aberrations, and kataegis. But no single feature explains the recurrence.
Some recurrent changepoints may track changes in chromatin state during cancer development. Five of eight recurrent changepoints in CLL show shifts in activity between SBS9 and SBS5, signatures which decrease and increase, respectively, during subclonal expansion.
At recurrent changepoints, mutational signatures often change in similar or comparable ways. Fifty-five melanomas with a changepoint on chr 1p show similar, large changes in the activity of two UV-associated signatures: SBS7a and SBS7b.
Recurrent changepoints are common. Eight cancer types have recurrent changepoints present in at least seven samples, and five of these recurrent changepoints are shared by multiple tumor types.
AI score was, however, a significant predictor of outcome (HR=1.4) and cancer recurrence (HR=1.7), even accounting for all other major prognostic factors. And it was even a better predictor than the original AI status annotation used to develop the score!
I'm enjoying playing around with the new GeneMANIA data release by @garybader1's group (Ruth Isserlin and Christian Lopes). Nearly 900 human networks! Check it out: https://www.genemania.org