Cory McCartan
@corymccartan
Asst. Prof. of Statistics & Political Science at Penn State. I study stats methods, gerrymandering, & elections. Founder of UGSDW and proud alum of HGSU-UAW L. 5118. 🏳️🌈 corymccartan.com
We also show random tree features have similar uncertainty quantification to full BART, despite not learning the tree structure
Enough theory—does this work in practice? We show random tree features are competitive with full BART, random forests, and xgboost, and lie along the Pareto frontier of accuracy & computation
Turns out (with some conditions) you only need n^(1/3) trees to achieve the optimal learning rate! Notably this rate is faster than existing rates for BART theory, which suffer from the curse of dimensionality* *not quite this simple, but basically
OK, but this is for infinitely many trees, right? What do we do in practice? We propose "random tree features," which are a BART model where the tree structure isn't learned
We can characterize the RKHS (function space) that BART's kernel corresponds to. It lies in between a weak class (just requiring 1 derivative; slow learning rates) and a smooth class (p derivatives; fast-ish rates), and has a nice minimax learning rate that depends only logarithmically on dimension!
The BART kernel is anisotropic, a.k.a. not rotation invariant, which means intuitively it prioritizes main effects over interactions (more about this below)
In fact, we show formally for the first time that as the # trees grows, BART converges to a Gaussian Process with a kernel we can describe in closed form! This means (a) no tree learning happens in the limit, and (b) we can study the kernel to learn things about the types of functions BART learns
We first try knocking out different parts of the BART model to see what matters most. Turns out, at least when you have many trees, it's (mostly) not being Bayesian, nor learning the tree structures, nor even the flexible part of trees per se
New WP w/@melodyyhuang.bsky.social studying the success of BART models, which regularly win causal inference competitions! We argue that BART should be thought of as a random features approximation to a limiting GP. This view helps understand BART & apply it in more places arxiv.org/abs/2607.28844
Out today in the APSR: our paper studying which gerrymandering reforms work! Long story short, do what Michigan does! Slightly longer story: we project these different complex reforms onto a 1-dimensional "partisan leeway" axis, and do continuous-trt DiDiD there. doi.org/10.1017/S000...
Had a great time presenting Shiro Kuriwaki's and my new EI methods to EPSS in Belfast on Saturday! Our R package implementing all these methods (as well as partial ID bounds) is also now on CRAN! corymccartan.com/seine/
New R package on CRAN! ONNX is a runtime & file format for ML models. 'onnxr' lets you load & run models in about 2 lines of code! E.g. image detection running in ms from a pretrained model. Perfect for embeddings, smaller models, etc. And lots of .onnx available online! corymccartan.com/onnxr/
Roberts court: "the mere fact of racially polarized voting is not relevant to proving racially polarized voting patterns"
Newly updated simple congressional model (tinyurl.com/cmchousemodel), now with LA map & recalibrated election model fit only to 2026-2024 data! Ds have lost 4 seats on avg to redistricting, but the loss grows to 5.6 seats in a D+6 environment. Ds 98% to win house in a D+6 environment, though
The NYT calls it a 'blow' to the VRA, but Thomas and Gorsuch are clear on what the effects of the decision will be: "an end" to the "misadventure" of "roughly proportional representation" for minority groups.
The egregious decision in Rucho now lets SCOTUS punch a massive loophole for VRA §2: partisan goals can justify disparate racial impacts. This is about districting, but the same logic would let states resurrect literacy tests, etc. as long as the stated goal was partisan discrimination.
Time to plug my simple congressional model again, which I've updated with the new VA plan: tinyurl.com/cmchousemodel. Net effect of re-redistricting is D+1.5 seats in tied environment, growing to D+2 seats in a D+10 environment. Bit more competitive map, too.
New WP! Philip O'Sullivan, Kosuke Imai, and I generalize an earlier SMC sampler for redistricting plans, allowing it to scale better & be applied to multi-member districts, too, like the Dáil Éireann (Irish parliament) below. Look for a new 'redist' version soon(ish)! arxiv.org/abs/2603.22188
v0.2 of `bases` is up on CRAN! `bases` brings nonparametrics into your favorite modeling functions—using random Fourier features is as easy as sticking `b_rff()` into your formula. v0.2 brings graph Fourier & random convolutional features, `mgcv` integration, & more! corymccartan.com/bases/
Last fall I shared new methods research with @shirokuriwaki.bsky.social on ecological inference—inferring individual relationships from aggregate data. Our new review WP frames past EI methods as linear models, and argues credible EI requires controlling for covariates arxiv.org/abs/2601.07668
Interestingly, the TX gerrymander doesn't have any net impact in a D+12 environment Plots below show new effect of redistricting this cycle—cf left with the changes vs right if the 3-judge ruling holds
Davidson/Nashville swinging D+25 is wild Most swing districts next year won't look the way TN-07 does, with mainly rural counties that are shifting D by less
D+10-15 national environment would mean Dems end up ~100 seats over Rs in the House tinyurl.com/cmchousemodel
IF it holds, this would net Dems 1.6 seats, on average, due to mid-decade redistricting (D+1 seat in a Dem-favoring environment)
Have updated my simple House model spreadsheet with currently enacted districting plans. Net effect is R+0.4 seats on average (!), with actually a _Dem_ advantage past a D+8 national environment. Copy, edit, & explore for yourself: tinyurl.com/cmchousemodel
NYC precinct map: Mamdani now vs the primary Purple = relative improvement vs primary Orange = relative loss vs primary (e.g. GOP voters) Takeaway? Mamdani improved significantly with Black voters since June! We see this in EI estimates as well: Mamdani likely won Black voters ~ 52/43 vs Cuomo
Precinct-level NYC data: Mamdani exceeded our voterfile-based expectations, Cuomo slightly exceeded them; Sliwa fell way behind, especially in places he was expected to do well!