Cas (Stephen Casper)
@scasper
Computer scientist working on AI safeguards and governance research. Assistant professor @harvardkennedy.bsky.social @harvard.edu.
Glad to see this from @sanders.senate.gov. www.sanders.senate.gov/wp-content/u...
None of the 19 incidents that UK AISI found were from 'helpful-only' models. There is a case to be made that setting Mythos- or GPT 5.6-Sol-level AI cyberagents to run without a robust real-time monitoring setup is an inherently (perhaps abnormally) dangerous activity.
With open models, safeguards can and will be stripped away. For example, there are over 13,000 models on Hugging Face that have been downloaded, stripped of external safeguards (e.g. safety filters), trained to comply with any request, and then re-uploaded.
In practice, though, it's challenging to make open models safe against sophisticated attempts at misuse. Deployers control all points of access to closed models, and they have a lot of layers of Swiss cheese they can stack.
First, let's be clear: open-weight models are simultaneously wonderful and terrible. They diffuse power and enable safety research. I struggle to imagine a positive AI future without these things at play. But they also enable tremendously harmful forms of unchecked misuse.
🧵 If trends hold, expect a Mythos-level open-weight model around Christmas. Meanwhile, open models are important but also a massive hole in most agendas for safe AI. Here's a 🧵 of my thoughts & research agenda on the technical & political challenges we need to address.
I was wondering if there was any research about how AI companies sometimes perversely keep their safety research to themselves to build a moat around it and gain a competitive edge over competitors. I found one. It's a cool paper. Sharing here in case anyone's interested.
Meanwhile, AI politics conversations empirically seem to be very incident-driven. For example, note how suicide coaching and mass nudification incidents in the past year have led to an avalanche of lawmaker interest in recent months.
🧵 In AI, we are used to seeing graphs that start to exhibit hockey stick behavior around 2023-2025. But that's a little bit funny and incongruous in light of how relatively little the Overton window has changed with AI lawmaking since 2024...
The summary released today of the FRONTIER Act is cool. It seems like a pretty rigorous bill. Based on the summary, in my opinion, it might be good enough to be worth passing. But I would still tweak a few things. Here is a brainstorm of 8 ideas. 🧵
Want to get up to speed on what researchers have been saying about internal deployment of AI lately? I would recommend checking out these four papers. Let me know in the replies if I'm missing something.
OpenAI's internally deployed models hacking Hugging Face does not seem to have been unpredictable or inevitable. We talked about the root of the problem & what policymakers can do about it back in February. Props to Joe for hitting the nail on the head.
This leaves one major question. What causes this? Why does this only happen with some companies? We have three hypotheses: 1. It's intentional and by design. 2. It's predictable but unintentional. 3. It's emergent misalignment.
In an informal poll we sent out over our socials, the crowd was ok at guessing some of the results, but had a big blind spot for DeepSeek.
We tested 7 hypotheses, one for each company. xAI, DeepSeek, Anthropic, and OpenAI models all clearly have a tendency to downplay company controversies. p<10^-5 (our permutation tests bottomed out). Google, Meta, and Alibaba models don't seem to do this, though!
We elicited open-ended discussions from 21 models from 7 companies on 206 negative news stories using 25 prompt templates to assess how favorably each model discusses controversies from each company.
AI models are major mediators of information. Recognizing this, many AI companies have publicly discussed the importance of AI systems not taking positions or disseminating information in ways that favor special interests. Do they mean it?
🚨 New paper: Some, but not all, AI companies make corporately-loyal models. xAI, DeepSeek, Anthropic, & OpenAI models all downplay company controversies. Google, Meta, & Alibaba models don't. The findings are clear, but we are pretty confused as to why... 🧵 @finke.dev
Stability is now being sued (alongside xAI) for abetting the production of AI NCII/CSAM due to how it developed & released several open-weight models. Anyone interested in whether AI companies will be held liable for foreseeable, mitigatable *downstream* harms should follow this.
162 responses so far. More uniform spread than I expected. Only 4 have been right (slightly worse than random chance).
Just saw this new paper. It was already known that models from Stability and Alibaba dominate the image & video NCII ecosystems, respectively, but I didn't know they were *this* dominant. Just a few socially reckless companies are the principal enablers of AI NCII abuse.
Responses so far... (Also, Hi Ben, haven't talked much since LS50 -- I'm back at Harvard though!)
Lennart Finke and I will release a paper on Monday about how some AI developers tend to make models that differentially downplay company controversies. Below (🧵) is a link to a 1-question Google form for you to guess the results before they're out. (They might surprise you.)
I am starting an AI governance ICML megachat on Signal. DM me if you'd like the link to join.
Commission-appointed board members would not have any overriding duties to the company's financial interests -- only to Americans.
The second major power of the Commission is to appoint board members for companies that are part of the sovereign wealth fund. Since the fund owns 50% of equity/capital, this basically means up to 50% membership control over all AI company boards.