Jay Shendure
@jshendure
shendure lab |. krishna.gs.washington.edu
.. faith over many years that we'd eventually pull this off, and we are very grateful. Please see seattlehub.org/nextcell for annotated browsable tree plus raw & processed data. We've barely scratched surface of exploring this large tree of regulative development, so please have at it! /finis
.... @choijunhong.bsky.social , @cxqiu.bsky.social. .and thanks to amazing community of scientists & funders made this possible, inc. Seattle Hub for Synthetic Biology @alleninstitute.org @biohub.org @brotmanbaty.bsky.social @uwgenome.bsky.social @hhmi-science.bsky.social ... who had a lot of ...
You can walk tree in either direction. Descending: blastomere A → one pre-gastrulation founder → a E7.43 progenitor whose 624 sampled descendants exhibit clonally demarcated contributions to three endodermal organs. Zoom in further to homogenous liver subclade. 15/n
Fun result - although tree built from TAPE edits alone , clade co-occurrence recovers germ layer organization & sweeping developmental time turns that into a *dated* hierarchy of couplings 14/n
At tips, 51% of tree siblings share cell type, 9× chance, & extremes track anatomy, e.g. lung & airway 139×, DRG neurons 69× (cell types that commit from spatially restricted pool & expand locally). Recent terminal differentiations are legible among heterotypic tree siblings 13/n
The tree resolves 182 ancestral lineages at E6.0; follow them to E13.5 & just 24 account for half the embryo, where a neutral birth-death process predicts 35. Clonal dominance is set early, then simply inherited. (see paper re: 2nd phase of clonal dominance) 12/n
Then, near silence for 4 days! Followed by fast re-start of keystrokes at ~E6 at ~9 edits/cell/day, fading slowly as TAPEs fill. Avg. 55 / 66 sites written by E13.5. Analogous to 13.5 page book w/ pp2-5 torn out (pre-gastrulation arguably dull anyways :). Not ideal, but legible.
Plot twist: At 2-cell stage there is burst of edits (major ZGA!). Daughter A got 23, B got 19, none shared. Thankfully as we only had one mouse, this hands us an internal replicate. Trees made & measured twice in two half-embryos that shared a zygote. All findings reproduce. 10/n
Annotation didn't start from scratch. Our E8–P0 mouse atlas (by @cxqiu.bsky.social, @bkmartin.bsky.social, Ian Welsh from JAX) covered essentially every cell type at E13.5, so we could immediately label all 1,340,794 tips — 25 major trajectories, 135 cell types. 9/n
Phylogenetic tree building methods don't scale well, so @seidels.bsky.social did it in 2 steps. 1) Build a 640,012 cell backbone by NJ on an ordered edit distance (SciPhy) & strict molecular clock dating, then 2) distance-based placement of more cells to get to 1,340,794 tips. 8/n
To get DNA Typewriter in vivo, Qi Yu & @alleninstitute.org / SeaHub's in vivo platform injected PEMax, epegRNAs, blank TAPE into B6 zygotes w/ piggyBac, then let development proceed to E13.5. 100 injected → 78 transferred → 10 embryos → 1 w/ great recording, on 11 x 6-unit TAPEs. 6/n
For background: in 1983 Sulston & colleagues completed the cell lineage of C. elegans by watching every division under a microscope, & forty years later it is still the only complete animal cell lineage we have. You obviously cannot build a mouse cell lineage this way. 4/n
We've also built NextCell (h/t @claude_code ), an interactive browser for the entire annotated phylogeny — seattlehub.org/nextcell Open science, so the trees and data are released there too (or will be very soon, lmk if anything missing). 3/n
Thrilled to post thread re: new single-cell lineage of mouse embryo reconstructed w/ DNA Typewriter. One animal, zygote to late organogenesis (E13.5). Tree has 1,340,794 transcriptionally profiled, annotated tips (cells), 1,142,588 dated internal nodes, rooted at zygote 1/n
Only ~31% of our training data was needed to get within 5% of max model performance. Readouts from locus-scale libraries like this one may be be a genuinely useful data type for deep learning models of gene regulation, orthogonal to biochemical assays like ATAC-seq and conventional MPRAs. 11/n
We also tried predicting activity from sCRL composition. Linear models: r = 0.57. Add interaction terms: r = 0.81-0.82. Tree ensembles (random forest/gradient-boosted trees): r = 0.90. Combinatorial regulatory logic is real, and learnable.
In fact, the insulator rules are quite different. Within the most active promoter group, we found that insulator position is very important: activity was ~3-fold higher when the insulator lay upstream of the enhancer than when it intervened between enhancer and promoter. 9/n
Not all enhancers behaved the same. Adding more copies of e-NMU(iii) or (iv) in the upstream slots (S1-S4) raised activity, log-additively, while e-NMU(i), (ii), and (v) showed no dose-response at all. Only two of five turned out to be active enhancers in this context. 7/n
So we split the data into 4 "promoter groups" by whichever element drives transcription. Each has its own baseline activity and its own dose-response to upstream enhancers — and those two properties are independent of each other. 6/n
We found that two of the five "enhancers," e-NMU(iii) and (iv), aren't just enhancers. When placed right next to the minimal promoter in slot 5 (S5), they overpower it, acting as strong, orientation-dependent promoters themselves — up to 110x more active than everything else. 5/n
LAMPRA = Long @$$ Massively Parallel Reporter Assay. Instead of one short CRE next to a promoter, we combinatorially assemble 5 x 1-kb elements to make synthetic cis-regulatory loci (sCRLs), barcode them, and use a standard MPRA readout to measure activity. 2/n
It took some doing, but the resulting STEAM model predicts genome-wide distal enhancer landscapes for hard-to-obtain cell types throughout the human genome, despite not being trained on any human molecular data, as well as 240 other mammalian genomes. Here shown for alpha-fetoprotein/albumin locus.
Latest from Shendure & Qiu labs (@cxqiu.bsky.social) )! We combined a new 4M cell mouse whole embryo scATAC-seq atlas (E10-P0), millions of 'evolutionarily coherent' orthologs from 241 mammalian genomes (Zoonomia), and the CREsted CNN framework (@steinaerts.bsky.social).
Super excited about first Shendure/Baker Lab collaboration & preprint on a multiplex sequencing-based strategy for screening de novo proteome editors in mammalian cells. Kudos to the brilliant Chase Suiter (not here) & @greenahn.bsky.social on the work! Preprint here: www.biorxiv.org/content/10.1...