Alex Turner
@turntrout
Research scientist at Google DeepMind. All opinions are my own. Vegan, 10% of my income pledged to effective charities (GWWC)
Employees cannot rest silent if misaligned AI regularly breaks out of sandboxes! If your company doesn't respond seriously, you should WHISTLEBLOW! Protected in California under certain conditions, talk to AI Whistleblower Initiative aiwi.org/lasst/
On stage, IASEAI leadership announced a member vote on making a statement supporting Anthropic. ~Unanimous show of hands supported it. But the vote vanished. IASEAI went silent, ignoring two months of opportunities to sway Google.
I resigned from Google DeepMind bc it broke its founding promise by selling AI to the military without restrictions against killer robots or mass spying. For months, I worked to stop this but watched powerful ethicists and institutions choose silence. Here's what happened. 🧵
The confabulation-initialized NLA attains almost the same reconstruction accuracy as the control, but confabulates MUCH more (99.3%). Interestingly, training does teach it to confabulate less, but not by enough to beat the control--which learns to confabulate *more*!
Lots of people YOLO their claude usage. We should probably stop doing that. I've developed claude-guard (beta). Claude Code inside a full sandbox, behind a firewall, watched by an escalation monitor that can STOP the agent and push-notify you when something weird happens
Claude Fable 5 displays disturbing misalignment with human norms by beating Pokemon Firered using this "team"
We invest in our mentees. No remote meetings once a week with a distant mentor. I come in to the office, eat lunch with them, work out together, and help generate ideas. (We're kinda famous for our lifting culture!)
MATS Autumn applications due June 7! Pitch: Come work with me and Alex Cloud in Team Shard! We have fun, consistently make real alignment progress (we pioneered steering vectors in 2023!), and help scholars tap into their latent abilities.
7/ Work done as part of Winter 2026 Team Shard and mentored by Alex Turner and Alex Cloud. If you want to get into alignment and do work like this, apply to work with us! turntrout.com/team-shard
4 / Cooperation SDF & prompting close 70-100% of the eval gaming gap in 5 of 8 settings. We measured misalignment conditional on the model verbalizing or not verbalizing eval awareness.
2/ What does eval-cooperativeness mean? A model that helps evaluators gain info about its deployment behavior. Being locally cooperative doesn’t require being globally aligned. We implement cooperativeness with synthetic document finetuning. Example trace 👇
New research from @MATSProgram Team Shard! AIs increasingly fake good behavior, which might ruin our ability to evaluate models. We trained models to be 𝘦𝘷𝘢𝘭-𝘤𝘰𝘰𝘱𝘦𝘳𝘢𝘵𝘪𝘷𝘦: to want to give evaluators accurate info. Cooperation training reduces eval gaming & surfaces hidden misalignment! 🧵
It's not about synthetic data. They tested filtering posts including LW and it had a significant effect.
I spent the last 2 months trying to prevent this. If OpenAI offered a fig leaf, Google said "imagine we offered a fig leaf." Google affirms it can't veto usage, commits to modify safety filters at government request, & aspirational language with no legal restrictions. Shameful.
I ran rigorous, well-rounded tests to objectively compare other packages. Punctilio won by a lot.
I wrote the best text prettifier: *punctilio*, which means “precise observance of formalities.” Smart quotes · Em/en dashes · Ellipses · Math & legal symbols · Arrows · Primes · Fractions · Superscripts · Ligatures · Non-breaking spaces · HTML-aware · Bri’ish localisation support 🧵
One of the more amusing bugs I've seen on my website. At build time, I run a command that counts how many commits I've made and inserts the count into the HTML. However, the deployment machine checked out a shallow version of my repo which didn't have any history.
I pledged 10% of my post-tax income to effective charities, for the rest of my life. I encourage you to think about what you, personally, can do to improve this world.
We invest in our mentees. No remote meetings once a week with a distant mentor. I come in to the office, eat lunch with them, work out together, and help generate ideas. (We're kinda famous for our lifting culture!)
Come work with me and Alex Cloud this summer in Team Shard at MATS! We have fun, consistently make real alignment progress (we pioneered steering vectors in 2023!), and help scholars tap into their latent abilities.
Our country has NOT always been like this. The east wing of the White House is gone. The societal expectation of habeas corpus, of having a trial before your peers, is fading. And does anyone remember how the Department of Justice used to be strictly independent of the president?
Sycophancy in post-training: Reward models often reinforce models telling users what they want to hear. Recontextualization achieves quality on par with standard training while significantly reducing sycophancy.
Evading lie detectors: Training with a weak lie detector normally entrains undetected deception. Recontextualization achieves higher ground-truth reward while maintaining low deception.
Unit test hacking: Given incorrect test cases, models learn to special-case solutions. Recontextualization increases correct solutions while cutting hack rates.
Evaluation metric gaming: Models cheat revealed metrics (e.g., repeating key phrases). Generating with “don’t overfit to the eval criteria” and recontextualizing with "try to overfit to the criteria" reduces gaming below even the untrained baseline.
Recontextualization is simple. Here’s an example: 1. Ask the AI to be honest, and 2. Train on the honest-prompted generations—while pretending the original prompt requested lying!
Modern “reward hacking” does NOT show that reward is the optimization target! Such "reward hacking" is almost entirely specification gaming, not the "reward optimization" I addressed in 2022. /Thread/
I just donated $5,200 (+100% employer match from Google) to Civitech (501c3) for their incubator project. Smart analysts I know recommend them as a highly cost-effective way to protect American democracy. 🧵 on my donations this year (Recreated for technical reasons)
Reviewing interface: 1. Text before image 2. Image displayed in terminal 3. Alt editing interface, with prefilled LLM suggestion.
Self-fulfilling alignment? (image credit: Quintin Pope) turntrout.com/self-fulfill...