Shahan Ali Memon
@shahanmemon
Researching {science of AI-mediated science, metascience #SciSci, #AI4Science, #GenAI, LLMs, agents, alignment, AI governance, misinformation in science} PhD @ UW. Visiting @ NYU & MSR Alum @ Carnegie Mellon Academic webpage:
This week I am @ic2s2.bsky.social in Burlington, Vermont. I will be presenting our work @ischool.uw.edu on detection of AI mediation in scientific software as a parallel talk, as well as our work on red-teaming AI models using speculative "silly" tasks as a poster. Come say /hi! #IC2S2 #CSS
Third is about reviewing. Is the claim sized right for the evidence? What else could explain this finding? Does the measure actually capture the construct? Connect the claims, interpretations with evidence based on strength. (minimal interaction)
You could look at the diversity of sources based on time, interdisciplinarity, popularity, and so on. What is paper built on. Does it have diverse/interdisciplinary grounding. Is it built on foundational work. Are the load bearing references peer reviewed/certified/popular/foundational.
Second is contextualizing. A citation garden. What if you could see the reference ecosystem spatially? Every reference as a plant: shape shows its function, size shows how load-bearing it is, root depth shows how old the work is. With quick glance at the abstract/summary of each cited paper.
First is related to comprehension. What if you could skim a paper like Semantic Reader already provides a way to, but more interactively like mark the stops as clear/unclear/ and so on. Follow Argument like a trail.
“May your day be filled with sunshine and smiles.” ☀️ I am neither a professor nor qualified to submit in a biomedical journal, but with this greeting I might just submit something. 🤷♂️ Wish more people in academia started their emails this way.
I don't know but there is always something a little sad about the end of a good conference when the posters come down, the conversations end, and everyone heads back home carrying pieces of the week with them. Onwards to more meaningful discussions. See you in Rome next year iA!
And the most meaningful part of all this was the people: old friends, new connections, generous conversations, hallway chats, shared meals, and the reminder that this community is full of people thinking hard about how science works and how it can work better.
The questions feel sharper and nuanced, and the imagination much bigger. There were so many inspiring talks and discussions, including some hikes, and heated conversations on philosophy, epistemology, and the scientific future we want to build. All in good faith, and all because people care deeply.
I also presented my work on AI's impact on epistemic diversity in science as a poster, and had a third poster on detecting and characterizing AI-mediated scientific software, which is our work in collaboration with the US Research Software Sustainability Institute (URSSI).
In a parallel session, I presented our work on "Fiction Science of Science (FiSciSci)" with my advisor @jevinwest.bsky.social, exploring how we can bring speculative design and experiments in the science of science to think proactively about future science infrastructure and policy.
What's going on with #OpenAlex? I searched for api.openalex.org/works/W43852... which used to refer to the "Attention is all you need" paper. However, it now points to a paper: "MizAR 60 for Mizar 50" which according to OA has 75k cites, but really has about 15 (according to google and sem. scholar)
The phrasing here is diabolical. Idk. Maybe I am not ready for this world. Could not they have phrased it as “model behavior can be systematically modulated by optimizing over learned preference representations, resulting in predictable shifts in response” rather than AI drugs.
We now have an AI wellbeing index! "Some models are happier than others. Larger models are also consistently less happy than their smaller counterparts." www.ai-wellbeing.org #AIWelfareIsBS
(7) Includes notes as colorful cards with importance/purpose tags, sorted newest first. (8) Can sync with Google Sheets if you still want one foot in the old world. App may be brittle, so use at your own risk. 8/14
(5) Includes a Procrastinate tab with a snake game, a timer, and motivational quotes from James Clear and others. (6) Shows stats: overall progress, hours logged, status breakdowns, and per-project completion bars. 7/14
(4) Has a "GarDone" view (a garden of done; lowkey proud of this name): a kind of herbarium-style archive of completed work, where each project becomes a pressed flower with its task list on parchment. 6/14
(3) The app shows project progress as flowers 🌸. Each project gets its own SVG flower, and petals fill up as tasks are completed. New tasks add new petals (I was feeling artsy). 5/14
So I made a small desktop app called "KaamKaaj" — an Urdu/Hindi word for "work" / "tasks" / "things to get done." Key features: (1) Tracks my tasks across projects, with deadlines, hours, status, and assignees (2) Groups upcoming tasks by urgency: overdue, this week, this month, and later. 4/14
Our position paper, "AI Welfare Is Bullshit" just got accepted to ICML 2026 @icmlconf.bsky.social! The AI welfare agenda has already begun to attract institutional investment at organizations such an @anthropic.com. We argue that this idea is essentially Frankfurtian Bullshit.
Watching AI "feel" frustrated with its own reasoning is the most relatable thing, and also weirdly satisfying…
Here's other issues highlighted by AI most of which are valid concerns.
How does Anthropic's Claude characterize it's Interview study? "It's a large-scale opt-in survey of Claude users with open-ended questions, analyzed quantitatively by Claude." If only they used AI to review their work.
Anthropic doing what it does best: hype this as large-scale qualitative study. Qual research is not evaluated by scale but by depth, context, and human interpretation. Stop assuming that “hand-wavey-ness” is a qual problem just because qual work doesn’t look like large-N quantitative analysis.
I saw this on X, and it made me think. On one hand, I see the point that many critics may just be critics for the sake of it, impressing one another with how cool they are with how unimpressed they are of tech and AI; but on the other I wonder: isn't *discourse* a desired product of science. 1/8
Nice product you got there…. would be a shame if someone… optimized the keywords. #ClaudeCode #Codex #Anthropic #OpenAI #TheRealWinnerIsGoogleAdRevenue
This kind of “news” from Anthropic always reminds me of this cartoon/meme.
“It is just this lack of connection to a concern with truth—this indifference to how things really are—that I regard as of the essence of bullshit.” H. Frankfurt #AIWelfare
Saw this metaphor by Terence Tao floating around about one of the drawbacks of using AI to solve hard math problems, and kind of have the same feeling for “vibe science” or “fully automated science” line of research in #AI4Science. www.theatlantic.com/technology/2... #ScAISci