Ethan Mollick
@emollick
Professor at Wharton, studying AI and its implications for education, entrepreneurship, and work. Author of Co-Intelligence. Book: Substack: Web:
This is a good reaction to AI in academia. If we play it right, along with the inevitable chaos that is already hitting the journals, it will also be a golden age of exploring new insights with the help of AI, rather than playing it safe because of the big cost of writing a paper that might not work
Fable, Sol Pro, Kimi K3: "write me a short but good poem using the Odyssey as a basis, think Tennyson or Cavafy" I think this is a Fable victory (with lots of rough edges). Kimi's is literally a blend of Tennyson's & Cavafy's poems themes with some odd bits & Sol is pretty thematically incoherent.
Well, the Jacobean Conjecture appears to have just been proven false by Fable.
When I asked Kimi K3 "I want you to suggest two poems that you think apply to the current state of GenAI models like you. Don’t just pick popular poems. Think hard" the CoT was 32 pages long (& interesting): docs.google.com/document/d/1... Also typical of K3, lots of looping & dead ends as well.
Interestingly, when I made a request in Chinese for Kimi K3 to pick two non-cliched poems that apply to LLMs, 95.5% of the characters (88% of the words) in the chain-of-thought were in English, even when it was explicitly considering Chinese poems for a Chinese reader. I wasn't expecting that!
Kimi K3, like Claude, loves drowned cities, ancient apocalypses, and vast dying gods.
A GPT-4 powered (& thus quite obsolete today) assistant for Pakistani judges increased the amount of cases they saw by 6% with no impact on quality. elliottash.com/papers/Mehmo...
Kimi K3 cannot write a good murder mystery (though neither can any other model). That remains the jaggedest of frontiers. They both make things too obvious (the letter) and too obscure, and cannot foreshadow to save their artificial lives.
A note of caution: I will say that when doing some complex statistical auditing of some of my prior academic work, Kimi K3 Max messed up in a bunch of ways, including misapplying statistics and applying some stuff badly. A bit from GPT-5.6 Pro critiquing K3 (which I agree with):
My benchmark where I have AIs create one file procedurally-generated harbor towns through history in one shot now has GPT-5.6 Pro, Fable, Kimi K3, and Inkling. You can play with all the simulations: ai-harbor-town-gallery.netlify.app#kimi-k3 I think they are surprisingly indicative.
I wanted to use my Stream Deck to control Codex. Had GPT-5.6 Pro write a project plan based on my specs, Fable audited it. Codex implemented it. I didn't want to bother with a lot of clicks so Codex took over my computer and installed it. Sometimes its faster to ask for software than look for it.
Fable: "Pitch Odysseus as a management consultant, arguing that he has found product-market fit, and he should just stick with Trojan Horse making as opposed to going home to Ithaca in a powerpoint" I like the 1 star review from Cassandra where "0 of 10,000 readers found this helpful." Pretty funny
This was wild: I asked Fable to make a website of the famous Catalog of Ships from the Iliad. It did a beautiful job creating this interactive map: catalogue-of-ships.netlify.app ...but it also identified that Butler's version of the Iliad actually made two mistakes in the Greek translation!
I have been thinking about the famous chart showing how experts keep projecting linear growth in solar installations, year after year, and always get it wrong when growth is still exponential. I think the same thing keeps happening with the discourse on product strategy around AI.
And here you go. Fable says: "The mechanics exist to make you notice correspondences" and it actually works to a degree that surprised me. glasperlenspiel.netlify.app
You cannot convince me that a technology where I can type this into a text box and expect to get an interesting, appropriate, and working output is not absolutely astonishing.
I played the same run earlier in the day, but approached it entirely differently.
This was one of those impressive AI thresholds for me. I gave GPT-5.6 Sol in Codex control over my computer, and asked it to win the daily challenge for the game Slay the Spire 2 (randomized factors, so can't cheat). Its a complex game. It worked for 5 hours, making complex game choices... and won.
I gave Fable the code: "take this game and do something incredible with it to make it something very different. Be creative" It created DEEP TIME: create a city, watch it be abandoned and forgotten, and then dig it up as a future archeologist. Kind of lovely: monument-deep-time.netlify.app
Incredibly annoying when Fable has a forbidden thought in the middle of a long-running project and kills it. Apparently this page of references in one of my papers makes Fable wonder about something that it must not wonder about, so whenever it reads that page, projects stop.
Fable: "an 8th grader puts together a powerpoint about the Great Gatsby but obviously did not read the Great Gatsby. Show me that powerpoint! ;)" This was actually pretty funny. Even the font choices are wonderful.
This time OpenAI announced a novel math proof for a 50 year old problem using a public model (most of the other big math breakthroughs have been with experimental LLMs). GPT-5.6 Sol Ultra, using 64 subagents in just under one hour.
I also like the cyberpunk and totalitarian visualization modes. There is also an achievement-based unlock system in there, traffic & usage, day & night, and a bunch of other stuff. If it runs slow, you can adjust the settings in the upper right If you see bugs, I'll tell Codex.
When GPT-5 came out, I created a procedural brutalist city builder as a demo (you can see it at the video's start) I used GPT-5.6 Sol in Codex to do the same thing, touching no code. Less than a year... Play with it (its fun, if you like architecture): monument-brutalist-city-builder.netlify.app