Avik Dey
@avikdey
Data • AI/ML • OSS • People • Engineering • Society • Approximately Generated Illusions; use specialized Small LMs as tools • Learn deeply to explain simply
We in the US have harmed AI innovation by failing to define any constraints for model development. In the absence of constraints - incremental scale has became a substitute for invention. Kimi K3’s long-context design demonstrates well how constraints can force innovation. arxiv.org/pdf/2607.24653
When OpenAI head of strategic futures says it’s not distillation, that’s the strongest argument that could be that it is distillation.
Study Anthropic would have funded except for this line: “… our findings impose a mechanistic constraint on neural theories of language processing and on accounts that equate prediction in human language comprehension with the comparatively unconstrained operations of large language models (LLMs).”
Here you go - uploaded - SPA, no framework, no routing. Straight from Fable's weights, except for the Gaussian noise fix that I mentioned earlier. Even includes a clickable historical timeline and a guessing game. For a 60 token input - looks sleek! sunny-starlight-502006.netlify.app
Inspired by this, I asked Fable to generate a diffusion explorer. 60 tokens input, ~15k tokens output. Looks really nice. But, like before, the more you look the more you will find. Here's one visually, the mean brightness of noisy version at t=50 should be ~50% (right), instead it's ~30% (left).
It’s 65% now? 3 months back Anthropic was hyping it at 90%! www.techspot.com/news/112408-... Also, Karpathy trying to sell Slack integration as a new paradigm for development with AI … I am starting to think more and more - what’s he thinking?
1. Community notes can be fun sometimes. 2. Master vibe coder is still promoting autocomplete, not AI.
Every part of this was entirely predictable when they hired a data labeling company CEO to build their frontier model. bsky.app/profile/avik...
My position remains: I applaud the kids for their vision and wisdom. bsky.app/profile/avik...
Sometimes I dig thru public code for agents … it’s fun. These ones were from the weekend. Of the ones in the parent folder this deserves special mention because it transcends the human experience. Don’t miss the <TERMINATION_CONDITIONS>!!! github.com/github/aweso...
So, ambitious tasks like throwaway vibe code for home automation on the weekends? “You can give it a lot more ambitious tasks than what you're used to, the model "gets it" and it will just go, and it's never felt this tempting to stop looking at the code at all (but don't do this in prod!).”
This is the kind of research paper that is an absolute joy to read. Just the thoroughness of the evaluation is remarkable. LLMs rebuilding programs from scratch with access to the existing stripped binary: 0% at 100% pass rate and a measly 3% at 95% pass rate.
The man used to be a great educator then he decided, he would rather be a mediocre influencer.
Why are Anthropic models better at coding than OpenAI’s? Because they have harnessed the power of the grand wizard! Seriously, how these got through any sort of review cycle and survived with modules in 1000s of lines each - does leave you wondering a bit! raw.githubusercontent.com/ComeOnOliver...
What does mandates mean in this context and how is it related to the national processing? #Hungarian #Election
Karpathy thinking that folks who are professionally evaluating LLMs are using the app to assess capabilities is a sad cope - expected better from him. Not going to link, you know where to find it.
Though my own fascination is for studies such as this one that study generalization in infants from small training sets than listening to CEO sproutings. While the authors call it “small”, comparatively it’s more like minuscule training sets. Absolutely fascinating. arxiv.org/pdf/2510.15060
In early 2010s when ML was having its moment with Big Data, half of the world’s favorite ML demo app was something something sentiment analysis. In the age of LLMs, its still sentiment analysis, now with Regex - a Claude Code speciality. Interesting thread: neuromatch.social/@jonny/11632...
If you want to understand how shallow this analysis is, compare these visualizations - original on left and my lunch hour throwaway Gemini Pro code on the right. Data sources were BLS and O*NET. Check the title on mine. Turns out "AI exposure" looks an awful lot like "uses a computer at work".
We have known for a while now that “reasoning” is just another illusion, LLMs are just good at faking it. arxiv.org/pdf/2603.05488
Why would senior engineers be on the hook for the foolishness of engineering executives FOMO driven poor decision making? Lead by example, take responsibility for your decisions - sign off on the AI code pushes yourself. (posted by lukaszolejnik.bsky.social on the bird site)
Catch? You have to know how - yourself: “It's not perfect, it needs high-level direction, judgement, taste, oversight, iteration and hints and ideas. It works a lot better in some scenarios than others (e.g. especially for tasks that are well-specified and where you can verify/test functionality).”
After watching Australia lose to Sri Lanka in the T20 World Cup match this morning, the only visual that came to mind was this one. Proud of my cousins from down under. (credit @jeremyphowhard via OpenAI on the bird site)
My favorite use of ChatGPT was as scratchpad & search tool. Last night made butter chicken for dinner. Modified original ingredients and steps to use 1 cup of heavy cream and 3.5 tbsp ground blanched almond. Asked GPT-5.2 Auto to scribe the new recipe. Mangled it. Every damn time. #MissYouChatGPT-4o
Might as well announce ARC-4 and ARC-5 now. BTW, today’s TTA systems still heavily relies on DL - that distinction between “static DL” and “adaptive systems” is a spectrum not a binary.
OpenAI jokers don’t even understand that high performing engineers perform at their best when they are given autonomy and ownership with TRANSPARENT goals and ZERO micromanagement and NOT when they are being “governed & observable”. IDIOTS.