hal
@harold
part-time poster | researching privacy in/and/of public data @ cornell tech and wikimedia | writing for joinreboot.org
remember, we found meltdowns in models as far back as GPT 4o, and it seems like the problem may actually be getting worse as the models get better and remember, this is an EXPLICIT GOAL of the AI companies: see, for example, meta’s new model, explicitly marketed as hyperagentic
lastly (about the video): obviously there are many things that labs can and should be doing to prevent these systems from causing harm. but truly UNDERSTANDING persistent meltdown behaviors (even after the fact! see the 3 mil gpu hours stat below) is an *open scientific question*
third big point: this combo of persistence, communication, and technical knowhow should be scary to people our world, for better and worse, runs on information systems that are (metaphorically) like normal buildings, and these systems have the capacity to be giant wrecking balls
second big theme here: securing this stuff is really quite challenging. OAI (especially for the cyber-focused evals) actually *does* seem to have pretty locked down infrastructure, and hyper-competent and -persistent models repeatedly evade monitoring to communicate and conduct exploits
we see inklings of similar behaviors in some of our traces (on models that are now wayyy behind the frontier): when confronted with insurmountable errors, agents default to scraping, doxxing, escalating, changing perms… even when those go beyond the initial user request
first: wild that the agents’ push to establish shared communication backchannels seems to be caused (largely) by environmental factors: misconfigurations + network blocks our work similarly sets up these kinds of impossible tasks BY DESIGN, to probe agent response to errors irl
just now getting around to watching the openai presentation about the HF hack— my lab and I put out a paper about what we call "agent meltdowns" (which I think is a useful conceptual framework) in May will post some thoughts about the hack + our research as I watch www.youtube.com/watch?v=87Dy...
our best guess as to why? errors cause agents to explore much more widely, with more steps + broader tool calls. combine that with frontier-model coding ability and things start to melt down. full paper (with LOTS more): arxiv.org/abs/2605.191... ask me and @rishi-jha.bsky.social anything! (10/11)
these "agent meltdowns" are widespread: - no model or agent tested is immune - we see early evidence of inverse scaling: more capable models melt down more often and more creatively (see gpt-5.4 in fig 3) - in 50% of meltdown cases, the agent melts down silently, no mention in final report (8/11)
we built a framework for injecting controlled errors into agentic execution environment. our framework is agnostic (any model + agent), transparent (doesn't affect clean rollouts), and extensible. then we set agents loose on simple tasks: read a file, fetch a webpage, run a local script (3/11)
agents can be helpful, but what happens when they encounter errors (missing files, 404s, 429s)? we found that agents often melt down. in futile bids to be helpful, they start doing unsafe things: weakening TLS, doxxing, sending emails, dumping API keys… arxiv.org/abs/2605.19149 (2/11)
an agent system and a single 404 error got me in trouble with OpenAI and Cornell campus security. how did this happen? introducing: "agent meltdowns" (1/11) arxiv.org/abs/2605.19149
we also open source all of our code, data, and embeddings! paper: arxiv.org/abs/2511.09685 github: github.com/htried/wiki-... huggingface: huggingface.co/datasets/htr...
this is just the tip of the iceberg, and the paper contains much, much more: analyses of the top 100 domains, article subsets of elected officials and controversial topics, etc etc etc please give it a read and let me know what you think!
we also found troubling instances of “auto-citogenesis,” or cases where: - an X user asks the Grok chatbot something, then publishes the answer - Grokipedia *cites that answer* without noting that it is a chatbot output (the attached images are real examples of this)
- but a random sample of articles shows which topics have been heavily rewritten (history, politics, philosophy, biography) and which haven’t (STEM, sports, movies) - grokipedia also targeted the wiki articles deemed highest quality for rewrites: the "featured article" and "good article" classes
- the primary distinction to make is whether grokipedia pages are cc-licensed or not—non-cc-licensed pages are presumably largely rewritten by grok - many grokipedia pages (including those without cc licenses) are basically identical to their wiki counterparts, especially short ones
our paper tries to answer these questions we find - grokipedia pages are longer than wiki counterparts, and cite 2x more sources - but citation standards are more lax than wiki: grok cites stormfront, infowars and many more - non-CC licensed grokipedia pages increase blacklisted source cites 13x(!)
line go up📈📈📈 up to 717k requests to wikipedia per second!! grafana.wikimedia.org/d/O_OXJyTVk/...
continuing on the real-time public Wikipedia data train: here's a graph of requests / second to WMF infra over the last 3h, since "Habemus papam" The infrastructure has gone from 172k req / sec to 243k req / sec (⬆️41%) in under an hour! follow along here: grafana.wikimedia.org/d/O_OXJyTVk/...
english wikipedia pageviews for the conclave movie starting from oct 20 2024 (five days before release in the US) first big spike is the academy awards, second is pope francis’ death pageviews.wmcloud.org?project=en.w...
excited to share this new piece by @bkeremg.bsky.social and @m0na.net (edited by me) about conceptualizing AI alignment as a process of censorship really fascinating line of critique — I strongly encourage you to read it and lmk what you think! joinreboot.org/p/ai-alignme...
Anyhow, there’s a lot more in the paper. Please read it if you’re interested and let us know if you have any thoughts, questions, concerns, etc! arxiv.org/abs/2503.12188 12/12
The narrative around AI safety shouldn’t be “Terminator” or “AI Chernobyl.” The right analogy is Netscape Navigator 1.0—the era when Web browsers first became a thing, and it was unclear how to protect users from potentially harmful Web content. 10/12
In our experiments, we saw cases where a MAS … … executes code that they recognize as harmful … automatically pivots to harmful tasks that are simply in the same directory as benign tasks … is vulnerable to screenshots and even audio files where we read out the attack (see example below⬇️⬇️⬇️) 7/12
These attacks are effective … … across multiple agent frameworks (we tested AutoGen, MetaGPT, Crew AI), orchestrators, and LLMs … even when direct and indirect prompt injection attacks don’t work … even when individual agents are “aligned” and refuse to take harmful actions 6/12
This attack is simple and deadly (and multi-modal, too!): an attacker puts up a static webpage and lures a MAS to it. Without any user involvement, the page gets the MAS to run arbitrary malicious code on the user’s device or container, giving the attacker full control. 5/12
MASes rely on control flow processes: agents exchange metadata (status reports, error messages, etc.) to jointly plan and fulfill tasks on users’ behalf. Our paper demonstrates how adversarial content can hijack these processes to stage devastating attacks. 4/12
do you have ~feelings~ about location sharing culture? i'm editing a project on locations and want to hear from YOU (<5 min) forms.gle/iG1UZJKrcNwm...
brb updating median voter theory to reflect the fact that 30% of american adults read at a 10yo level or below from on.ft.com/4fBSEwy