Seth Karten
@sethkarten
Autonomous Agents | Research @ Prime Intellect | PhD @ Princeton | Prev: CMU, Waymo | NSF GRFP Fellow
How are the academics feeling about this? Does it even change anything for profs?
Great to see Continual Harness acknowledged in Schmidhuber’s latest survey paper
Wow, three papers in CoLM 2026... Here I come San Francisco! These papers predicted some early trends in multi-agent safety & economic envs, automatic RL env creation, and PPO for VLMs/LLMs 🧵
New paper alert: Continual Harness: Online Adaptation for Self-Improving Foundation Agents Paper (arXiv). arxiv.org/abs/2605.09998 Article (Substack). sethkarten.substack.com/p/gemini-pla... Project page (video demos). sethkarten.ai/continual-ha...
im very glad to see a rebound this year in the total number of NSF GRFP awards to exceed the most this century. Founding the next generation of American scientists is important to keep growth of the sciences.
How do we close the gap between specialist RL and generalist LLM agents? We're benchmarking it in Pokémon. Join us at the PokeAgent Challenge competition workshop @ NeurIPS 2025. 📍 Dec 7, 8AM 🎮 Track 1: Competitive Pokémon (game-theoretic reasoning) 🗺️ Track 2: Speedrunning (long-horizon planning)
In the NeurIPS PokeAgent Challenge, we stress-test 4 ranking systems across (100k+ agent matches): - Bradley-terry (batch MLE, our ground truth) - Elo (online, chess-standard) - Glicko-1 (online, uncertainty-aware) - GXE: (Glicko-derived win %) (2/5)
A benchmark environment is nothing without data so you can pretrain before you RL. Announcing our replay archive preview: We are releasing an additional 25k games to help you train a metagame exploiter (5 million more released after qualifier) replays.pokeagentshowdown. com:8443/ (3/3)
Pokemon is truly the pareto frontier of agent research - The RPG requires an autonomous embodied agentic agent with perception, planning, memory, and control - VGC and Gen 9 OU penalize erroneous actions with fast-paced opponent-modeling in short games (1/3)
If you arent paying attention, we are in a rapidly shifting period of ML paper culture. ICLR/ICML/NeurIPS are being treated as random, out of touch processes with more and more unnecessary work to submit Most people are saying TMLR is the only good alternative, but are skeptical
🚨 Hackathon Weekend! 🚨 Jumpstart your PokéAgent Challenge submission ahead of NeurIPS! 📅 Sept 13–14 ✅ Leaderboards reset Sat 10AM EDT 🎙️ Lightning talks in LLMs, RL, and Pokemon 💬 Live Office hours 🏆 $2k in prizes
The solution would generalize to another two player partially observable turn-based text game. The most bespoke items are tools, but there has been work recently that shows that you can make these tools modular LLM calls, further increasing generality
Democratic alignment: in a special case, periodic citizen voting can fire the planner. Leader turnover keeps welfare high and prevents policy drift—central nudging plus decentralized oversight in one sandbox.
Centralized nudging: the planner’s marginal taxes beat U.S. statutory rates and approach Saez on aggregate welfare (almost double vs baseline).
Synthetic behavioral policies → we sample workers from 2023 ACS skills & demographics, then let each agent verify its own bounded rational utility from individualized preferences, enabling counterfactual reasoning.
🚀 New preprint! 🤔 Can one agent “nudge” a synthetic civilization of Census‑grounded agents toward higher social welfare—all by optimizing utilities in‑context? Meet the LLM Economist ↓
🚀 Launch day! The NeurIPS 2025 PokéAgent Challenge is live. @neuripsconf.bsky.social Two tracks: ① Showdown Battling – imperfect-info, turn-based strategy ② Pokemon Emerald Speedrunning – long horizon RPG planning 5 M labeled replays • starter kit • baselines. Bring your LLM, RL, or hybrid agent!
🚀 5 days until my ICML spotlight poster! Key insights we’ll unpack: • Base LLM + test-time planning • Game-theoretic scaffolding • Context-engineered opponent prediction • Comparative LLM-as-judge (relative > absolute) Catch me Thu Jul 17, 4:30-7 PM PT👇
Social media takeoff is hard. Bluesky still lacks the capability to compete with twitter
Excited to announce that I will be spending the summer at @Waymo on the simulation realism team! I’ll be working on learning to generate simulated worlds. 🚙🚙🚙 Send me a message if youre in the bay area and want to chat!
Excited to share that the PokeAgent challenge was accepted as a NeurIPS competition! This should serve as an excellent benchmark for competitive games AND ‘speedrunning’ the RPG. I hope to see both the RL and LLM agent communities working together here to eval agents in Pokemon More info soon👀
What happens to TRI though? I thought they had an AV division. Also Toyotas arent EVs so I am confused how the driving tech stack would work. I think Waymo is just diversifying their risk into personal vehicles
Insane new study from zurich studies the influence of the LLMs for persuasion on the r/ChangeMyView subreddit. Let's just say people are outraged... Is the study justified since bots are already rampant on reddit? Or does this cross ethical lines?