Hanbin Lee
@epigenci
Imaginary evolutionary biologist. Want to be real one day. I'm very keen to review papers for not-for-profit journals!
By binning frequency, alpha-models jump back to the desired performance level by approximating the evolutionary formula we derived, but with many more parameters.
This is essentially a local linear regression. Any curve or parameterization can approximate the target curve as long as they're limited to a short interval. In this spirit, we show that alpha-model in practice simply approximates the evolutionary formula we derived under stabilizing selection.
In practice, no one fits a simple alpha-model. Arguably all modern implementations (e.g. GCTA) splits the model across frequency bins. For a pre-specified frequency bin, separate parameters are fit.
In forward simulations with stabilizing selection, we show that this model performs better than the alpha-model. Furthermore, alpha-model that is intended to capture frequency dependent architecture completely breaks down. It does worse than models completely agnostic to frequency dependence.
We show that this formula holds fairly generally under Gamma DFE where the constant "2" is replaced by one of the shape parameters "1/k".
The model has demonstrated empirical success, but the formula itself wasn't derived from evolutionary theory. This work shows that the correct formula under stabilizing selection is a rational function. It doesn't explode to infinity even if frequency approaches zero (alpha model does).
Stabilizing selection has emerged as a powerful tool to explain complex trait architecture. Among many of its qualitative prediction, variants with larger effect size tend to appear in lower frequencies. The standard parameterization so far was the alpha-model.
How should we apply linear mixed models to populations under stabilizing selection? The first paper I wrote with my grad school advisor is published in Genetics. academic.oup.com/genetics/adv... 1/n
This was the first scene of the department when I first visited and also the last one today before I left the office. Always wondered why the sign is under poor lighting, making it look gloomy. It was a rough time in the US but all good memories belong to the department.
MATH 626 FINAL QUESTION : Compute the relaxation time of a random walk on this district using the canonical path method
It's simply because people almost actively refuse to consider these mechanistic factors explicitly and treat heritability as a blackbox summary, causing confusion. Random variable and generative parameters are also conflated. see e.g. my recent paper. academic.oup.com/genetics/adv...
"Worried that your boundary is infinitely far away because your space is not compact? Then make one." I feel so bad about this.
The quantity that I use in my recent preprint for a hacky Bulmer correction is also the average, not the realized one. This makes theoretical analysis easier but ignores a potentially important variability. I tried to say this explicitly. 5/n
It's a good exercise to check which one does the quantity of interest fall into. For instance, the beautiful Simons et al. (2018) paper is computing the average, not the realized one. That is, this is different from the usual V_g computed from sample variance. 4/n
Does better prediction results guarantee that it reflects the underlying evolutionary process better? No. \alpha=0 model does better in prediction than non-zero \alpha models. BUT It compensates its lack of frequency dependence by underestimating mutational variance. (9/n)
Also, it's basically simulating from its own model to confirm its correctness. This is clearly a circular argument. The consequence is clear: if you fit non-zero alpha model to slim simulated genotypes under stabilizing selection, it's a disaster. It does worse than \alpha=0 model. (8/n)
TL;DR; is in the following formula. \sigma_b is the focal trait's mutational variance: how large the newly emerging mutation's effect size is. \sigma_a is the one for the latent trait that is actually subject to stabilizing selection. \rho_{ab} is the strength of the coupling between them. (4/n)
This has a strong meme potential together with the president. ncatlab.org/nlab/show/He...
Comparison to BOLT-LMM clearly shows the advantage of this new class of ARG-based algorithms over more traditional ones. The backbone (implicit genotype access by matmul) is the same, but ARG offers a huge improvement. (3/n)