Brian Heseung Kim
@brhkim
Data Scientist & Education Policy Researcher // Building rigorous, open-source AI trainings and tooling for researchers and non-profits to use AI responsibly and critically // PhD + MPP @UVA // www.brhkim.com // www.openaugments.org
Also new: run DAAF on Anthropic, OpenAI, or OpenRouter (Kimi K3, GLM 5.2, and friends), one friendly control panel instead of scattered helper scripts, a self-healing updater, and as many side-by-side installs as you like. Free and open-source, as always.
For the R folks (social science, ed policy, epi): R is now a first-class execution language; switching from the Python default takes one sentence. And for PII/HIPAA/FERPA-bound data: a new workflow helps you prototype your analysis on a fully synthetic "twin" dataset. The real data is never touched!
🥳DAAF v3.0.0 is out today!! The two headliners: DAAF now speaks fluent R, and it can now ALSO run on your ChatGPT subscription/API key (Yes, GPT-5.6 inside Claude Code!). The whole release is about greatly expanding who DAAF is for across the social sciences. daafguide.substack.com/p/new-daaf-v...
And then the main intuition of my post in one gif! AI keeps changing what students can do, and our expectations need to shift accordingly. This is ultimately, I think, really exciting! But only with awareness, criticality, creativity, and imagination.
Added GLM 5.2 to DAAFBench, and it's so good it required a complete rewrite of the key takeaways and recs. It is utterly indistinguishable from Opus 4.5/4.6/4.8 at 25-33% of the cost and open-weight. What a wild week. Kimi K2.7 Coder also added in, but frankly unimpressive!
4. Open-weight models are highly capable -- a real win for self-hosting -- but with a crucial decline in consistency 5. Local-class open-weight models just aren’t there yet IMO, these are *the* frontiers to watch for a truly hopeful democratization of research capacity via AI!
Some interesting takeaways: 1. Fable 5 is *actually* in a class of its own 2. Opus 4.7 was genuinely worse than Opus 4.5/4.6, but Opus 4.8 was a strong return to form 3. Sonnet 4.6 is an enormously high performer relative to its cost Very validating of my intuitions to date
Different models also have empirically different strengths (e.g., consistency v. average score), and these results let users of DAAF make informed decisions about which models to use for *specific aspects* of their workflows. Lots of potential optimization to explore here.
Granular results like this allow us to carefully examine: what is the actual trade-off for cost and performance? And just as importantly, when will open-weight models allow us to conduct this type of complex work while entirely self-hosting? Both crucial for intensive data work!
The goal here was to very explicitly understand: which models are actually capable of managing the *processes, procedures, and protocols* necessary for ensuring responsible, reproducible, and rigorous research work with AI agents? This is distinct from raw coding/analytic power
I ran 17 different frontier models through 51 different tests (2500 total runs!) to examine: How well do different AI models handle the complexities of rigorous quantitative data analysis workflows? Excited to introduce DAAFBench: Orchestration! daaf.openaugments.org/bench
Would love to hear what people think. Does it get the points across well? What questions do you still have? What worries does AI in research bring to mind? #academicsky #academicchatter #econsky #econtwitter #phdsky #educationsky
And even if you’re already deep in the weeds using Claude Code, I hope it’s also a nice inventory of very valuable features you’ll want to adapt for your own setups! Don't worry, DAAF will always remain free and open-source! This is just a nicer wrapper around the GitHub :)
Whether or not you use DAAF, I think there's a ton for people to learn here about how agentic orchestration, context engineering, and the current frontier of AI capability actually works. I really strived to cultivate and share this critical awareness with peers via this site, and I hope it's useful
🥳 If you’ve been wondering how to use Claude Code as a quantitative researcher/social scientist of any kind: I’ve *finally* made a very nice, very accessible, and very informative homepage for the Data Analyst Augmentation Framework (DAAF), and I think you'll wanna take a look!
3. Everything you’d want to do with DAAF is now just one convenient utility script away! No more faffing about with Docker command lines or Git issues. Everything is handled for you, carefully and robustly, for Windows/MacOS/Linux, so you can focus on your work.
These are in *addition to* installing Claude Code, bespoke methodological skills, curated Claude permissioning/security defenses, automatic context management protocols, and a high-performance and fully reproducible Python data science/analysis environment that *just works*
2. DAAF now comes bundled with everything you need to make it your main AI-empowered research environment! From VS Code ready to roll, to a fully interactive session log explorer for complete transparency and auditability into *everything* Claude does for you.
1. Installation happens in one line now! From a fresh computer to talking with a DAAF-empowered Claude Code in no more than ten minutes. It’s easier than ever to get started with Claude Code in a highly curated, secure environment, which is particularly wild because...
For each atomic step of the data analysis pipeline, DAAF carefully injects carefully curated references that guide how it works -- things like best practices for various causal inference methodologies, or in-depth explainers on how to use specific coding libraries
What people need to realize is that Claude needs *grounding* to be useful: curated reference guides that help it think more like an actual scientist beyond its fuzzy "memory" and beyond sporadically searching through whatever pops up in Google Search. That's where DAAF comes in!
A lot of peeps have asked: What does it actually look like to use DAAF to analyze data? And how is it better v. Claude Code alone? It's exactly the right Q, and so I put together this interactive walkthrough showing every step, doc, and output from a full project! openaugments.org/daaf_anatomy...
If you're interested and want to learn more about DAAF, I'm actually running a webinar with @aefpweb.bsky.social this Thursday that's free and open to the public! Register here to join 150 others and counting, and note I'll be sharing the recording broadly afterwards aefpweb.org/ev_calendar_...
And my personal favorite feature slide from the video, representing so much effort and hopefully useful additions from v1.0.0!
DAAF sits between you and Claude Code to automatically and consistently help Claude think more like a *responsible* and *rigorous* researcher. Think of it as a force-multiplying exoskeleton for human researchers -- a tool explicitly designed to augment your hard-earned expertise, *not* replace it
What is DAAF? DAAF is a free and open-source instructions framework for Claude Code that helps skilled researchers rapidly scale their expertise and accelerate data analysis across any domain with AI assistance -- without sacrificing the transparency, rigor, or reproducibility good science demands.
And be on the lookout for DAAF v2.0.0 updates early next week. Huge updates and expansions to be VERY excited about; Claude's best summary of what to expect attached :) #edusky #econsky #academicchatter #academicsky
Lots more to say: this is just Step 4/6 in my vision for a more optimistic AI-empowered academia!
My dream: an "open knowledge repository" that sits alongside journals. Publicly share/verify everything that can be mechanically verified; new research traces lineage back to these stats to build arguments, and the new "journal paper" format can focus solely on the causal/complex/theoretical claim.