Kale-ab Tessera
@kale-ab
ML PhD Student @ Uni. of Edinburgh, working on Multi-Agent Problems. | Organiser @deeplearningindaba.bsky.social @rl-agents-rg.bsky.social | 🇪🇹🇿🇦 kaleabtessera.com
There's much more in the paper and website, including coordination by task type, model-size effects, and sampled interactive traces from all 13 models. The agent conversations alone are worth exploring, from role assignment to coordinated revives: alem-world.github.io/traces.html
Can stronger models carry weaker teammates in heterogeneous teams? In our initial experiments, not really. Mixed teams perform almost exactly at the average of their homogeneous baselines (+0.1 / −0.1 Total%), neither collapsing to the weakest model nor matching the strongest.
What drives coordination? Communication matters most. Removing it drops Gemini’s Coord.% from 17.5 to 5.3 and Gemma-4-31B’s from 8.8 to 3.8, while Base% changes much less. One clue as to why: ~44% of messages address a specific teammate, often to assign work, share intent, or align timing.
Looking beyond aggregate scores: 1. Base competence ≠ coordination: smaller open-weight models beat GPT-5.4 on coordination in some settings, with non-overlapping CIs. 2. Survival ≠ progress: MARL matches the strongest LLM's return in far fewer steps, echoing ARC-AGI-3 focus on action efficiency.
On the Hard setting, zero-shot Gemini 3.1 Pro performs comparably to the best MARL agent trained for 1 billion steps (17.5 vs 17.6 Coord.%, 15.4 vs 15.3 Total.%). *(Different interfaces, same underlying environment).
Agents navigate 9 procedural levels, pursue 93 achievements, and act for up to 10,000 steps. Alem generates not just the world, but the coordination problem itself, with controllable difficulty. Text, pixel, and symbolic interfaces let LLMs, VLMs, humans, and MARL agents face the same environment.
LLM agents are improving at long-horizon tasks. But deployed agents will also need to coordinate with each other. Yet most benchmarks still test agents alone or in short, structured collaboration. So we built Alem ("world" in Amharic), a JAX benchmark built on Craftax-Coop/Craftax.
Can LLM agents coordinate in long-horizon, open-ended worlds? We test 13 LLMs in Alem, a new benchmark. Most struggle, averaging ~6% normalised return. Yet on Hard, zero-shot Gemini 3.1 Pro matches the best MARL agent after 1B training steps. Our ablations show communication matters most. 🧵
We use these probes across 37 scenarios and 7 envs. Three takeaways: 𝟭) history dependence != history utility (only 43% need memory for high returns); 𝟮) hidden state and teammate info are separate difficulty drivers; 𝟯) sync and temporal coordination differ across envs.
🪄To study this, we introduce information-theoretic probes for behaviours that the envs induce under standard MARL algorithms (IPPO/MAPPO). They measure how actions depend on obs, histories, teammate-private information, and other agents' actions, beyond just returns.
As we build more cooperative multi-agent envs, we should ask: are we really testing the properties that make Dec-POMDPs hard, partial observability and decentralised coordination, or can agents succeed via shortcuts? 🤔 Our #AAMAS2026 Oral probes this question 🔎🧵
Happening now - Exhibit Hall C,D,E poster #404 I heard there will be good vibes at this poster 🤙
First time in a Waymo. Honestly, a pretty surreal experience! Surprised by how smooth the ride was and how quickly I felt comfortable in the car 😮
Thrilled to present HyperMARL at #NeurIPS2025 in San Diego next week! 🚀 (Amos will present at @euripsconf.bsky.social too.) TL;DR: Coupling obs and agent IDs can hurt performance in MARL. Agent-conditioned hypernets cleanly decouple grads and enable specialisation. 📜: arxiv.org/abs/2412.04233
Great first couple of days at DLI @deeplearningindaba.bsky.social in Kigali 🇷🇼, some highlights include amazing talks talks by @verenarieser.bsky.social and Max Welling, great pracs and tuts, and of course the opening party ( before the rain 😢) 🎉 #DLI2025
🇨🇦 Heading to @rl-conference.bsky.social next week to present HyperMARL (@cocomarl-workshop.bsky.social) and Remember Markov (Finding The Frame Workshop). If you are around, hmu, happy to chat about Multi-Agent Systems (MARL, agentic systems), open-endedness, environments, or anything related! 🎉
🔎 We also do ablations and see the importance of the decoupling and the simple initialisation scheme we follow.
💡To address the coupling problem, we propose 𝐇𝐲𝐩𝐞𝐫𝐌𝐀𝐑𝐋: a method that explicitly 𝐝𝐞𝐜𝐨𝐮𝐩𝐥𝐞𝐬 obs- and agent-conditioned gradients with hypernetworks. This means obs grad noise is avg. per agent (Zᵢ) before applying agent-cond. grads (Jᵢ) -- unlike FuPS, which entangles both.
🔬 We isolate FuPS’s failure in matrix games: shared policies struggle when agents need to act differently. Inter-agent gradient interference is at play -- especially when obs and agent IDs are 𝐜𝐨𝐮𝐩𝐥𝐞𝐝. Surprisingly, using only IDs (no obs) performed better and reduced interference.
📜🤖 Can a shared multi-agent RL policy support both specialised & homogeneous team behaviours -- without changing the learning objective, requiring preset diversity levels or sequential updates? Our preprint “𝘏𝘺𝘱𝘦𝘳𝘔𝘈𝘙𝘓: 𝘈𝘥𝘢𝘱𝘵𝘪𝘷𝘦 𝘏𝘺𝘱𝘦𝘳𝘯𝘦𝘵𝘸𝘰𝘳𝘬𝘴 𝘧𝘰𝘳 𝘔𝘶𝘭𝘵𝘪-𝘈𝘨𝘦𝘯𝘵 𝘙𝘓” explores this!
Reminder of how common paper rejections are, even for people like @teorth.bsky.social 💡
Coffee shop + Sunday + Admin = elite combo. Honestly, I don't know how I managed a to-do list before this.