Geoffrey Irving
@girving
Cofounder and Chief Scientist at Resolution. Alignment will be solved eventually, but not necessarily in time. Previously UK AISI, DeepMind, OpenAI, Google Brain, etc.
One lacking area of alignment theory is how best to think about rationalization, the process of (1) guessing an answer and (2) justifying it after the fact. Ideally multiple teams at Resolution will touch on this question from different directions, using different tools.
We're excited to announce that Resolution has a $160M grant from Coefficient Giving: $108M unconditional, with a further $52M conditional on hiring and compute needs. We'll use it to grow teams across our research portfolio and invest heavily in research automation. 🧵 resolution.org/post/funding
Now the binary is under 32KB on Linux and Mac, and uses a RAM base of 44KB on Linux and 224KB on Mac. Tons of unsafe rust, #![no_std], raw assembly for syscalls, custom formatting, arena allocators, etc. Maybe 50% whether Fable will find any security bugs?
A while ago I wrote a minimal replacement for top called ltop using Claude Opus 4.6-4.8. At first I made it so that I could nicely filter the processes to just the ones I wanted, but then I noticed the binary was pretty small, and got curious if it could be smaller... 🧵 github.com/girving/ltop
But I just published “Automated alignment is harder than you think” (arxiv.org/abs/2605.06390)! Automated alignment is not the best plan! A better plan is to not build ASI yet, and the world should try hard to realise that plan. Alas, the speed of progress calls for backups.
We are starting a new, nonprofit alignment organization, ⊢ Sequent Research, bringing together researchers previously on UK AISI’s Alignment Team, Timaeus, and elsewhere to research how to align superintelligence. We are hiring! 🧵 sequent.org/launch
A while ago I wrote a Claude skill to parse Google Docs into Markdown. It was easy, but took a few iterations to get right, as the first pass had weird glitches like \! instead of !. What else needs a couple more iterations? Literally Google Doc's own agent Markdown converter.
Oracle-relativised stochastic computations are also fun. github.com/girving/deba...
I should say my other favourite monad is the set of finitely supported probability distributions over values, or rather the class of randomised computations. github.com/girving/deba...
FWIW, this bug would never have manifested in my implementation of formally verified software floating point, since the dependent type enforces that there is a unique NaN representation. github.com/girving/inte...
As an exercise in learning recent Claude Code + Opus 4.6, I've formalised Seiferas's simplified construction of the Ajtai-Komlós-Szemerédi O(log n)-depth, O(n log n)-size sorting networks in Lean, using Margulis–Gabber–Galil expander graphs. github.com/girving/aks
It is unfortunately that arxiv has no mechanism for setting a social media image for a paper. I do not need the enormous ARXIV logo for this kind of link.
New "boundary point jailbreaking" method against LLM safeguards (with prior disclosure to multiple labs) by using noised versions of harmful queries to turn sparse feedback from failed attacks into dense feedback. 🧵 www.aisi.gov.uk/blog/boundar...
Anyone know if there are certificates for any sparse, symmetric positive definition matrix to be positive definite, that can be checked in quadratic time? Emphasis on *any* such matrix, no structure or other assumptions allowed other than SPD. scicomp.stackexchange.com/questions/45...
Joint with Jonah Brown-Cohen, Simon Marshall, Ilan Newman, Georgios Piliouras, and Mario Szegedy (now I have papers with multiple Szegedy brothers :)). arxiv.org/abs/2602.08630
New complexity theory paper mapping the precise query complexity of debate, given unbounded provers. No new safety ideas: the goal is a self-contained presentation of debate + cross-examination, with the precise complexity class it achieves. 🧵
Unless... @qntm.org "Very funny polling result where ~3/4 of people will say they read a book last year but if you ask them to name the book the share drops 20 points" x.com/JosephPolita...
New report on trends in AISI's evaluations of frontier AI models over the past two years. A lot of AI discourse focuses on viral moments, but it is important to zoom out to the less flashy trend: AI models are steadily growing in capabilities, including for dual-use. www.aisi.gov.uk/frontier-ai-...
The UK AI Security Institute ran an Alignment Conference from 29-31 November in London! The goal was to gather a mix of people experienced in and new to alignment, and get into the details of novel approaches to alignment and related problems. Hopefully we helped create some new research bets! 🧵
New open source library from the UK AI Security Institute! ControlArena lowers the barrier to secure and reproducible AI control research, to boost work on blocking and detecting malicious actions in case AI models are misaligned. In use by researchers at GDM, Anthropic, Redwood, and MATS! 🧵
Ominous start to a Wikipedia page about a formula... en.wikipedia.org/wiki/Fa%C3%A...
From near the end of Sleepwalkers, by Christopher Clark, as World War I starts.
Short note on relativisation in debate protocols: to model AI training protocols, we need results that hold even if our source of truth (humans for instance) is a black box that can't be introspected. With @benjamin-hilton.bsky.social and Simon Marshall. 🧵 www.alignmentforum.org/posts/XycoFu...
New alignment theory paper! We present a new scalable oversight protocol (prover-estimator debate) and a proof that honesty is incentivised at equilibrium (with large assumptions, see 🧵), even when the AIs involved have similar available compute.
Going back through old blog posts, and I still love these old cloth collision event visualizations. naml.us/post/visualizi…
AISI's research agenda is out! We cover a variety of topics in the evaluation and mitigations of risks from frontier LLMs, including both work happening at AISI and work we are excited to see others tackle. www.aisi.gov.uk/research-age...