Rich Harang
@rich
Using bad guys to catch math since 2010. Distinguished Security Architect (AI/ML) and AI Red Team at NVIDIA. He/him. Personal account etc; `from std_disclaimers import *` AI Security since it was ML Security.
I usually resist dunking on/amplifying these terrible takes, but holy shit. I know some day, probably not too far off, my kid's gonna start having too much going on to want to sit and shoot the shit with me all that often, and I'll be sad when it finally happens. Why would you ever give that up?
Air quality at 6:00 today, after the 5th of July fireworks that didn't start until the 4th had been over for 30 minutes. At least the incessant fucking fighter jet flyovers stopped early for weather.
Since recent events seem to have dragged AI powered biosafety back into the chat, thought I'd take the excuse to repost this. IYKYK.
Trying to get quicker and looser and less fussy at hatching -- quick sketch from imagination
I've used this 'seed' a few times now (code in alt text), with multiple models and every time I get something useful out. Tell it to "improve this script" once, then bootstrap to taste. It does need an Opus-4.6 level model to one-shot it, but cheaper models can get you there eventually.
Meanwhile, on Twitter (not "X"; their words not mine).... (From quick inspection: mostly crypto + telegram scams -- this is about a week's worth)
Tapping the "Models give you what you ask for, not what you want" sign yet again.
I am begging AI Red Teams to stop killing themselves trying to prevent attacks that can be just as easily accomplished by editing client-side HTML. For example:
An arcane tome filled with occult knowledge about the true workings of the world, that causes madness and despair in all who pursue its dark secrets? Yeah we've got one in the back.
Apropos of the "models pretending to escape from their server" thing: (from transformer-circuits.pub/2024/scaling...)