Tomer Ullman
@tomerullman
Associate Professor, Department of Psychology, Harvard University. Computation, cognition, development.
cf. Veo 3.1 from a month ago trying to think fo the director notes here. "That was great V, great. Just a few small things: the line is actually 'Alas, poor Yorick! I knew him, Horatio'; also, objects follow the principles of continuity and permanence"
every time there's a new video model I try some basic physics on it. So, here's the recent Google Omni tasked with continuining an image of a man throwing a mug at a tower of skulls. will the mug hit the skulls? see for yourself.
staying at friends and trying to work on a grant submission but they have an advanced anti-grant-writing system
have kids, and/or like to read and/or stuff? May I recommend "Sawyer Lee and the Quest to Just Stay Home"? (by @zachweinersmith.bsky.social)
telling an LLM to "just fix style and grammar" can change a lot more than just style and grammar
Or, here is a bland review, and the polished version. I didn't even bother bolding all the stuff added in the polished review. It is also so long it is now two images, sorry.
I think it's useful to see some bad offenders: Here's a bland student comment, and the polished version. I bolded the polished parts that insert significant new content, not present in the original. I'm not saying the new one is super-duper, but it goes from a 70 to an 85 grade at least.
I'm not talking about simple language change, obviously ~0% of the original survives under that metric. I mean *actual* changes. The changes are mostly driven by added comments, and the worst offender from what I've tested was Claude Sonnet 4.5
So, do LLMs meaningfully change stuff, even if asked just to polish? Yes. 'content drift' examines deletions, additions, and stance changes. it shows models change things. think of a content drift of 0.3 as "30% of the polished comment is not in the original". So imagine what a 0.8 is.
the 'polish' prompt was pretty clear, I think: do not add or remove anything, just rewrite to fit expected style, fix grammar, that kind of thing. I then had a judge LLM compare each 'bland' and 'polished' review for additions, deletions, and changes in stance.
I took random AI/ML papers, and had an LLM write bland, contentless reviews for them, based on the title alone. Then, a different LLM was given: the paper, the bland review, and the instruction to 'polish' the bland review (in the style of a leading AI/ML conference)
building some Lego with the kids and the instructions here had me puzzled for a moment if I'm supposed to reach for the 6 or the 9 bag
I wrote the Open Encyclopedia of Cognitive Science entry on 'Intuitive Physics': oecs.mit.edu/pub/m27ympza... --- for those who don't know it yet, OECS (@oecs-bot.bsky.social) has accessible, informative, free, serious entries on many basic topics in cognitive science
there are a LOT more details and discussions in the paper, I recommend reading the whole thing, it is very fun. That's not me saying, that's the Reviewers. Here's the line that I suspect made the Editor say "Too much whimsy!", but they graciously went with the Reviewers on this
we created a novel set of stimuli with a hierarchy of types violations, and had people rate the likelihood of this nonsense. we found: (4) people rank 'far'-type violations as more nonsensical than 'near', though both are nonsense; in line with the 'severity of violation' idea
maybe this is about 'metaphor'? Like, the more something sounds like a metaphor, the more sense it makes? you know, "filling your purse with luck" is literally nonsense, but it sounds like a metaphor, more so than 'hanging your coat on a sneeze'. So maybe that...? (3) No
So that's cool. People have a graded sense for some inconceivable events. But what is that sense based on? maybe some simple notion of 'more linguistically likely'? to make a long story short: (2) No. people's judgements of the inconceivable are not correlated with string probability
with all this background, you might think people would draw a sharp line between sense and nonsense. But we find that: (1) People have *graded* judgements about the inconceivable. Some things are (consistently) more nonsensical than others. (i do so love this stimuli)
Now out (for realz) in Cognition: "People Make Graded Judgments About The Inconceivable" (by Hu, Sosa, & me) Free preprint: www.tomerullman.org/papers/grade... Journal link: bit.ly/gradedInconCog @jennhu.bsky.social @cognitionjournal.bsky.social
Dual Coding Theory was proposed and studied by Dr. Allan Paivio, a professor of psychology, who in general was interested in the processing of non verbal information oh by the way, this is his Wikipedia page photo (Dr. Paivio, aka Mr Canada)
expansion idea for the boardgame Sky Team: you must work as a pair to land this guy