Joanna Masel
@joannamasel
Theoretical biologist and advisor to data scientists at the University of Arizona. Mostly theoretical population genetics and molecular evolution, but I've also published in biochemistry, infectious disease, aging, economics, education. Opinions are my own
We have produced a working hypothesis for future work on random polypeptide libraries of different composition, on ancestrally reconstructed synthetases, and on non-stationary substitution models. The origin of the genetic code is finally an answerable question. 13/13
The retrofunctionalization hypothesis is that late-added amino acids tend to recruit older, less specific synthetases. We find that the signal comes from early amino acids with late-resolved promiscuities using more complex synthetases. 12/13
We could add amino acids in order at nodes of synthetase chronograms if and only if we hypothesized exactly the same promiscuities that had been independently suggested by the 2-1-3 rule. The trees also make sense re promiscuities, deep mutational scanning data, and codon adjacencies. 10/13
As a control that we weren’t hallucinating patterns, we did the same thing with Trifonov’s (2000) influential hypothesized amino acid ordering. We got no satisfyingly plausible ambiguities, and worse adherence to the 2-1-3 rule. 9/13
Serine’s codons would once have been mutationally connected, before cysteine and arginine took some of them. 8/13
Adding amino acids in order reveals 2 glaring exceptions to the 2-1-3 rule. They can be resolved by assuming ambiguous coding of hydrophobic amino acids V/I/M, and of small/polar amino acids A/T, mirroring exchangeability in deep mutational scanning data, and latent synthetase promiscuity. 7/13
We infer ancestral codes by position amino acids, starting with VTG (left), based on their current codon usage (right), following the 2-1-3 rule. Glycine is so small that its polar backbone is more important – the first code would have a good hydrophobic-polar mix. 6/13
From where similar amino acids are, the 2-1-3 rule posits that information within 3nt codons once came only from the middle position, then the 1st, with the last still not fully used today. Distinguishing purines from pyrimidines came before distinguishing within each. 5/13
Our order doesn't support the coevolution theory that amino acids upstream in metabolic pathways were added first. 6 pairs work (S before C and M, putting E before Q, P, and R, and putting T before I), but 4 don't (D after T, M, and K, with N a tie, and G after S). 4/13
The relationship is stronger still when considered separately for amino acids handled by class I vs class II amino acyl tRNA synthetases, suggesting that it was the evolution of synthetases that drove progression toward larger amino acids. 3/13
Minor improvements to our previous estimate bsky.app/profile/joan... of which amino acids are enriched vs depleted in LUCA (indicating their order of appearance) strengthened the signal that small amino acids came first. 2/13
Results are surprisingly insensitive to mean selection coefficients. The first attached figure is for the zero environmental change case, the second (top row is what matters) is broader.
Small populations produce fewer of the new mutations that are eventually required to keep up with environmental change. Small populations also allow more deleterious mutations to fix. In both cases, lower fitness then creates a vicious cycle or “extinction vortex”. 2/5
For the sequence whose gene tree you are inferring, strict filters hurt, and our new gentle filter CLOAK performs best. Propagating uncertainty from our 16 variant alignments into a consensus among 16 variant trees was worse, showing the presence of systematic not just random alignment error. 9/10
Using a substitution model trained on strictly filtered alignment data leads to better inference on gene trees, bringing them closer to the known species tree according to Lin-Rajan-Moret distance (an improved extension of Robinson-Foulds distance). 8/10
Stricter filters have stronger effects in reducing exchangeabilities associated with less plausible amino acid substitutions, i.e. those that require more than one mutation, according to the genetic code. 7/10
Phylogenetics relies on substitution models (rates of evolution between amino acids or nucleotides), which are normally decomposed into a symmetric exchangeability matrix and equilibrium frequencies. Filtering reduces exchangeabilities between amino acids not linked by single mutations 6/10
In a trade-off between precision and recall, CLOAK is the best gentle filter, the “partial filtering” option within Divvier doi.org/10.1093/molb... is the best strict filter, and TAPER and the Divvier’s divvying option are in between. GUIDANCE2 and HmmCleaner perform less well. 5/10
Trigger warning: the attached images of real multiple sequence alignments may cause feelings of distress among biologists: 2/10
The same difference in exploitative ability yields more coexistence when created by search speed differences than via handling times. 7/9
Other parameters yield the “dominance-discovery” trade-off described in ants, where the Dove loses at contests but is better at finding new resources. We also find a new “Dove-discovery” trade-off, where Hawks search better, but the opportunity costs from contests is too high. 5/9
We develop a mechanistic model in which consuming a resource takes time, and Handlers can be interrupted by Searchers who find them rather than free Resource. This initiates a Contest, distracting the consumers, thus allowing the resource to grow to higher levels. 3/9
We applied these concepts to experimental data on 517 Arabidopsis genotypes. The proportion of deaths that were selective is higher than anyone anticipated. Despite artificial conditions, this demonstrates that these concepts can be applied to data – immediately producing a surprise. 13/14
In a relative fitness model with selection on just one life history stage, selective deaths maps to the lead, illustrated here for the asexual case of Desai & Fisher with the lead q mutations better than the mean, each of them with selection coefficient s academic.oup.com/genetics/art... 9/14
Lag load arguments are flawed, but in 1971 Nei pmc.ncbi.nlm.nih.gov/articles/PMC... and Felsenstein www.journals.uchicago.edu/doi/10.1086/... derived selective deaths without them. Reproductive excess is a budget out of which selective deaths must be paid. Little work after that 7/14
Relative load has since been rediscovered as the "lead" q in traveling wave models. Load/lead is a difference in fitness – this leads to different mathematical insights than approaches based on variances. 5/14
Haldane defined "selective deaths" (incl. missing births) as those causally responsible for allele frequency change. The "cost of selection" is the proportion of deaths that need to be selective for a given rate of sweeps. link.springer.com/article/10.1... 2/14
How? I have tried several times but the red "deactivate" box can't be clicked on.
So I'm still stuck, my 2021 password either doesn't work or I can't find it, and no option to reset.