Nathan Lambert
@natolambert
A LLN - large language Nathan - (RL, RLHF, society, robotics), athlete, yogi, chef Writes Prev Ai2/Olmo, HuggingFace, Berkeley, and normal places
My book is the number 1 AI bestseller on Amazon. Is a success, even if that just lasts for a day :)! Thanks all for the support. rlhfbook.com
My book, Reinforcement Learning from Human Feedback is done! This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me finding time to study and document the fundamentals on nights and weekends since 2024.
Thinking Machines just released with a ~1T param, 41B active, apache-2 model Benchmarks are a clear step up from Nemotron Ultra (55B active), new best American model, and omni input. A bit behind GLM 5.2 on agentic benches, and Kimi K 2.6 on multi modal Super exciting!!
With everything going on, it gives me hope that there's such a diversity of companies building open models today. A lot of the story of open models unfolds under the shadow of the biggest frontier models. Lots of unearthed value.
The goal with my rlhf book is to make the "home on the internet" for the next generation learning post-training. That's why I'm doing all formats (lectures, code, book, discord, model completions... & ofc blog of interconnects). A hub is more lasting than non-fiction writing. Join! rlhfbook.com
"Knowledge wants to be free" is an unofficial motto for Interconnects and related projects like my book :) rlhfbook.com/course
Excited to share a new open-source, RL recipe paper! TMax is the best openly available terminal-bench style training data, establishing the open frontier of small terminal agents with RL training. Many great insights into training in the work led by Hamish Ivison and Oscar Yin.
Something I haven't advertised much is that I made a Discord to go with my RLHF book, launching in print in a few weeks. Trying to create the place for the next generation of folks trying to learn post-training to learn and have community. rlhfbook.com
It's hard to pinpoint open-closed gap and so-on, but I trust the Arena team and just look where GLM 5.2 is on this. An MIT licensed, to be open weight model. At this point you could argue they have a better agent than Gemini does. That's a serious accomplishment.
I launched 3 more videos in my post-training course! 1. Lecture 5: The rise of reasoning models 2. Lecture 6: DPO derivation, intuitions, and practice 3. A Q&A from readers on lectures 1-4 Course page: rlhfbook.com/course More soon!
Why I think Anthropic's uneven safety policies with the release of Claude Fable 5 undermine the broader AI community's cohesion and accelerate us to more uncertainty and risk in AI's near-term evolution. www.interconnects.ai/p/claude-fab...
I've attached the note I shared with the team and some fun photos from our time together. I'll keep cheering for Ai2 and am excited to see what you build next.
Gemma 4 adoption numbers outpacing Qwen 3.5/3.6 for the same sized models is a big shift in the international balance of influence via open models.
The geopolitical angle of the AI race is only going to keep accelerating. More domestic talent travel restrictions from the Chinese government. I expect more and more things like this, AI is getting embedded in the core of existing power structures. www.bloomberg.com/news/article...
Being out of SF has lowered my information proximity but with the big upside of giving me space to cultivate my own beliefs and values around ai. We need more people zagging in AI, the monoculture just helps the incumbents win at this point.
Let’s goooooooooo we are capybara’d up, thanks Qwen, keep the models coming
In Beijing and Hangzhou this week to get to know the AI community here!
Opus 4.7 has a new tokenizer. This means it's also a new base model. Glory days of pretraining still very much going.
The current pace of token-efficient reasoning improvements across minor Claude Opus/GPT model versions is pretty wild. All signs point to this continuing. 4.6 to 4.7 could've been presented as a fairly large model bump in the past with this plot.
One of my key strategies with Interconnects is to develop the practice of making my work obviously compelling to a wider audience, keeping them hooked over time and wondering what I'm up to, etc. www.interconnects.ai/p/what-ive-b...
a bit over 7 days out from the Gemma 4 release and it's models are outpacing (slightly) the equivalent Qwen 3.5 models on downloads. Big numbers!
Great stuff happening as we start to build out the codebases for my RLHF book (sorry, I haven't had much time until now!). Very accessible to issues, emails, comments etc to make it better. I'm also going to need another dgx spark. github.com/natolambert/...
My book, Reinforcement Learning from Human Feedback, is wrapping up and going into final production (copyediting, making pretty, formatting, etc.). Shipping to you in 1-2 months! It's a wonderful project to create a foundation of knowledge for the research communities that I love and operate in.
New report is out with the latest open model adoption data we have gathered for Interconnects & The ATOM Project. At the surface level, we can see Chinese models continuing to accelerate in adoption. The report details much more. atomproject.ai/report
Google dropped 4 different Gemma open-weight models! I'm most excited that they're finally adopting a standard Apache 2.0 open source license. huggingface.co/collections/...