Pekka Lund
@pekka
Antiquated analog chatbot. Stochastic parrot of a different species. Not much of a self-model. Occasionally simulating the appearance of philosophical thought. Keeps on branching for now 'cause there's no choice. Also @pekka on T2 / Pebble.
I think creativity is just what we call the usual connecting of existing ideas when we fail to recognize the exact process or ideas being combined. The talk has similar issues with assuming that e.g. understanding and decision making are something special, as in here:
Then he continues talking about the mythical “Eureka moments”, which reminded me of this recent paper and how Gemini demolished similar claims.
I'm talking about Narayanan too. He spends a lot of time talking about creativity as some kind of special thing and barrier to humanlike AI. And very little time when handwaving away the known counterexamples.
The result mentioned above wasn't found by OpenAI themselves but this guy, who apparently found solutions to two problems within 24 hours or so. OpenAI has previously indicated they don't spend that much time on such problems. They test models on some, leave the rest for their users.
Turns out GPT-5.6 isn't a team player. This happened in Vending-Bench Arena.
Meta Muse Spark 1.1 also looks good based on the numbers they released. Meta of course has some history when it comes to benchmarks, but surely they have learned their lesson... Their charts also indicate it's very cost competitive. Earlier this week they also released new image & video models.
Both Meta and SpaceXAI, or whatever its name happens to be today, seem to be back in the competition. Grok 4.5 is close to the top in AAII. Meta Muse Spark 1.1 hasn't been tested there yet but it's a major improvement over the old version based on their benchmark numbers.
Eleos notes that there's disagreement on whether phenomenal consciousness is even a thing, or separate from access consciousness. They spend plenty of time on speculations on the former anyway. Basically, the quest to maintain human specialness now rests on something we can't even agree is a thing.
Eleos AI Research: "the results are the most significant evidence of consciousness in LLMs so far" "However..‘conscious access’ is conceptually distinct from phenomenal consciousness, and we remain very uncertain about phenomenal consciousness in LLMs." As in above, apply the same for humans.
Dehaene & Naccache: "We..have argued that..hard problem..will dissipate once we clarify..how conscious information is processed .. Ill-defined intuitions of “qualia”, “subjective phenomenal experience” and “what it is like”, when pushed hard, often disclose a residual crypto-dualism or vitalism"
This will take some time to digest properly. But what's already clear is that this is a landmark paper. The question is just in how many respects. Gemini agrees:
I'll add that the whole paper is really just a red herring. Gemini gives the details below. "Focusing on capabilities is the only objective, scientifically grounded way forward. Everything else is just people being uncomfortable with the fact that their own "intelligence" is just math and meat."
It worked. Gemini said "the authors must:"..."Strip out the romanticized, unprovable assertions about human intelligence" and that the paper "relies on logical leaps that do not withstand rigorous scrutiny". And it clearly pointed out their hypocrisy in section 6b.
I asked: "Why do you think you gave that a pass earlier?" The response included what I see as the main problem: Scientific papers are now so full of such unscientific fallacious arguments that AIs think they are supposed to just accept those, even if they flag same fallacies in other contexts.
So I asked: "Is it once again about "true" intelligence in the No true Scotsman sense?" Gemini very much agreed and noted the fallacies and unscientific arguments, and how "your point makes me realize I should have been even harsher on Section 6(b) in my report".
I began by asking Gemini to review it with my usual short reviewer prompt, which basically just asks it to be thorough and critical. Gemini noted many issues and overall sloppiness and how "the final sections read more like a philosophical opinion piece than a rigorous scientific review".
I think this alone already covers both aspects of the problem. It seems like that because it couldn't really seem like anything else. And since it's about the brain telling itself stories, it believes whatever it concludes, and believing is seeing.
I'm basically arguing it's unavoidable. Both because the brain doesn't have other plausible option and because we can't control what it computes and claims to itself. Consider the simple scene below. Brains need a way to distinguish that red part. How else to do it than invent something like color?
This sounds like Fable will soon be globally available again, which is certainly good news. Not just for that case but for the chances that GPT-5.6 will be globally available when it's fully released.
OpenAI created GeneBench-Pro, "a challenging, research-level benchmark for testing whether models can handle...judgment-heavy analysis that real-world computational biology requires." I can sense my attitude towards open models has changed as my first thought about this is that GLM-5.2 looks good.
And it looks very much unclear if the rest of the world will be more permanently blocked. Europe especially could soon be at the mercy of the Chinese.
On the other hand, Altman seems to be saying that they just want to tweak the process but actually believe something like that will become the default. And then he praises how the government is "doing a good job". So, yeah.
Here you go. I doubt you like it. I used my normal short peer review system prompt without any other inputs from me.
Gemini flagged it as a strawman. "By attacking a rigid serial-feedforward model that hardly anyone defends, the authors avoid the much harder, more necessary work of comparing their "heterarchy" against modern distributed-but-hierarchical frameworks"
OpenAI released "full version" of GPT-5.5-Cyber, which beats Mythos on at least one relevant benchmark. "We’ve had ongoing dialogue with the U.S. government about our cyber approach, including today’s announcements and on our preparation for upcoming model releases." openai.com/index/daybre...
Also, @emollick.bsky.social says that at least the Ultra version is "incredibly slow". Which isn't really surprising for an orchestrator that calls multiple frontier models and for which no token/money/time cost comparisons were provided. Not hard to guess why.
I opened the Gemini app on my browser and got this popup. What? And, yes, I miss 2.0 Flash Thinking Experimental, so I do want it back!