Blake Richards
@tyrellturing
Researcher at Google and CIFAR Fellow, working on the intersection of machine learning and neuroscience in Montréal (academic affiliations: @mcgill.ca and @mila-quebec.bsky.social).
10/15) They even achieved zero-shot cooperation via indirect similarity inference! Agents who never interacted directly, but observed each other's interactions with NPCs, accumulated evidence of similarity. Then, in the terminal Prisoner's Dilemma, they cooperated.
9/15) Crucially, this isn't naive altruism. In direct interactions, agents correctly inferred similarity and cooperated with identical copies. But when matched against a dissimilar, random agent, they defected! They only cooperate when evidence points to a similar partner.
7/15) We formalize this as the embedded Bayesian agent. Because an embedded agent models itself as part of the world, its epistemic uncertainty over its own decisions and the external world are coupled. Therefore, internal deliberation acts as Bayesian evidence!
4/15) According to classical game theory, the information gathered shouldn't matter for the final game; the agents should always defect. Instead, we observed that as the info-gathering phase lengthened, interacting AI agents converged on robust mutual cooperation! 🤝 Why?
3/15) But are classical models of rational agents actually compatible with modern AI? We tested foundation models combined with optimal planning in a two-phase setup: an information-gathering phase playing various matrix games, followed by a final, one-shot Prisoner's Dilemma.
2/15) Historically, predicting how rational actors behave has been the domain of game theory. Take a classic social dilemma like the one-shot Prisoner's Dilemma. Without reputation or reciprocity on the line, classical theory mandates that a rational agent would always defect.
1/15) What could drive AI agents to cooperate with each other, even if there is no chance for reciprocity or pay back? 🤔 🧵 Our team at Google, Paradigms of Intelligence, uncovered new paths to cooperation and a new game theory for foundation models 👇
Just realized that when you connect to an external connection on Slack, your icon is whatever your icon is in the workspace you added the channel to. I added the #NeurIPS2026 SAC channel to my internal lab Slack. So, yeah, visually, I'm Alyssa Edwards for all the other NeurIPS SACs now. 😅💄
One thing scientists forget is that a good philosopher will construct arguments s.t. if you accept their premise, then their conclusions necessarily follow. It's a mistake to try and poke holes in the logic of a good philosopher - that's their day job! The best bet is to reject their premise. 🤓
11/ Solving the Theory of Mind Recursion Problem ♾️ "I predict you predicting me predicting you..." To handle infinite Theory of Mind recursions, we solve the Grain-of-Truth problem for embedded agents.
10/ As a result, we show that MUPI agents can actually converge on cooperative strategies, even in games that typically always produce non-cooperative solutions, like the non-iterative Prisoner’s Dilemma!
5/ Decoupled ➡️ Embedded MUPI agents learn to predict the world they inhabit. But, they don't just predict future external observations, because they consider themselves as embedded within the world. Therefore, they also learn to predict their own future actions!
4/ Retrospective ➡️ Prospective Standard RL is retrospective ("do more of what worked before"). But social settings are non-stationary because other agents are also learning. This requires prospective learning – predicting the future to anticipate how other agents will adapt.
1/ Why does RL struggle with social dilemmas? How can we ensure that AI learns to cooperate rather than compete? Introducing our new framework: MUPI (Embedded Universal Predictive Intelligence) which provides a theoretical basis for new cooperative solutions in RL. Preprint🧵👇 (Paper link below.)
I saw this post yesterday, and I was so impressed by the unhinged moral outrage aimed at such benign uses of AI I had to save it. 😂
Back in Canada after two weeks in East Asia. Thanks to my friends and colleagues Jee Kwag and Jiook Cha for hosting me in Seoul, and @hiallen72.bsky.social for hosting me in Taipei! I had a wonderful time, and some fantastic, fascinating conversations. 🧠❤️
Thanks for response. 🙂 a) Hallucinations in certain contexts != poorer reasoning. Reasoning benchmarks show clear improvements (see below). b) Prediction different than factual Q&A, hallucinations meaningless concept for prediction. c) Again, paradigm shift is in data analysis, not modelling.
Coming to the #Cosyne2025 workshops? Wanna dance on the final night? We got you covered. @glajoie.bsky.social and I have organized a party in Tremblant. Come and get on the dance floor y'all. 🕺 April 1st 10PM-3AM Location: Le P'tit Caribou DJs Mat Moebius, Xanarelle, and Prosocial Please share!
"Through the roof"? If I'm reading this correctly, the number of cases of schizophrenia has been stable, but psychosis NOS is ~30% higher since 2016. However, the rate of "CUD" is about 2.5x since 2016, so 30% seems like a pretty modest increase given the level of use, no?
DeepSeek is an AI company, and their latest model is both basically as good as Gemini and o1, but open and trained at a fraction of the price (apparently):
It's an equivalent circuit, of course. (Didn't think I had to explain that.) This is one of the most well-established results in neuroscience: neurotext.library.stonybrook.edu/C3/C3_3/C3_3.... And the derivative of the membrane potential is a linear function of the current at *every* time-step.
8/ But, if you look at the simple linear-non-linear rate-based model in papers like this, it's still not doing *that bad*. We're talking ~70% of variance in spike rate explained. Not nearly as good as the more complex models, but hardly an "unrelated" gross abstraction, I'd say.
2/ First, let's start with the obvious: real neurons integrate their inputs. If you start from the basic principles of the relationship between voltage and current, and you know synapses induce currents when they receive neurotransmitter, then this is obvious.
You mean the small text on the page after you've already put in your input? Yeah, that's not nearly sufficient. It should be upfront, and very clear.
Montreal, where I now live, has undergone a similar transition as Paris, and it's so wonderful (see picture of the pedestrianized zone near me). Meanwhile, in the city I grew up in (Toronto), car culture continues and the province has announced they're going to spend $50M to *remove* bike lanes. 🤦♂️
This is a fun little app, but it gets some things very wrong, e.g. @andpru.bsky.social hates mice, and me, I hate puns. 🫠 Still, not too far off the mark... 😅 Full roast here: blueskyroast.com/roast/tyrell...
Indeed, that's not what I mean... Specifically, I'm referring to short-term facilitating synapses, which, due to their vesicle release mechanisms, barely respond to an individual spike, and ramp up their response to each subsequent spike if the rate is high-enough.