Kanishka Misra
@kanishka
Assistant Professor of Linguistics at UT Austin. Works on computational understanding of language, concepts, and generalization. Aspiring wugologist! 🕸️👁️:
Tomorrow, I head to beautiful Gothenburg, Sweden, to lecture at Analytical Connectionism (www.analytical-connectionism.net/school/2026/). In a shocking turn of events, I will talk about controlled rearing -- a topic I have never before presented!
Excited that our paper on minimal translation pairs for ASL was recognized with a best paper award at the Workshop on Generative AI for Sign Language @ CVPR 2026! Congrats to my co-authors @skarabuklu.bsky.social, Shester, Diane, Greg, and Karen! Paper: arxiv.org/abs/2604.27232
Doing my first ever conf-workshop keynote at the computational developmental linguistics workshop (comp-dev-ling.github.io)! Come one, come all!
I will be at #ACL2026 from July 2--7! I will be giving a keynote at CDL workshop on controlled rearing and hypothesis generation from language models! Tianyang Xu (first author) and I will present work on cross-modal generalization in VLMs on July 7! Paper: aclanthology.org/2026.acl-lon...
New opinion piece on the interface between research on concepts and categories in minds vs. in neural network LMs! I take the position that there is much to be learned from this interface (e.g., learning about concepts from language alone) and outline some directions for future.
Seems from the CFP page that this (pre-selecting commit) is just about EMNLP, and not AACL, unless otherwise announced by AACL)
Perhaps the most important update to the manuscript is the strengthening of our initial hypotheses about how the harmonic alignment between discourse prominence and positional prominence affects cross-dative generalization:
Announcing a new version of our 2024 paper on linguistic hypothesis generation from LMs! @najoung.bsky.social and I have systematized our hypothesis generation framework, added stringent criteria for model selection, 10x-ed our learning trials, and included an epigraph from Jeff Elman 🙏!
Excited to visit @gronlp.bsky.social (April 1) and @amsterdamnlp.bsky.social (April 2) to present a talk on something I have been thinking together with @najoung.bsky.social for the past 3 years!
If models were generalizing arbitrarily, then we shouldn’t see any differences in their performance across these settings (i.e., no matter what, crow == bird). However, we find that models seem to only generalize when the training data preserves category coherence! 9/
By coherence we mean the visual similarity between members of the same category, which we calculate using the DINOv2 embeddings used in our VLM training. Even in the original configuration, we found models to perform better on categories that were visually more coherent 8/
To test this, we created counterfactual data: 1) where category-label pairings were shuffled across categories (🪛= “robin”; 🎸= “crow”) and 2) where they were shuffled within categories (🦅=“robin”; 🦜=“crow”). These swaps also manipulate the categories’ visual coherence 7/
Are LMs simply executing something like “If crow THEN bird?” regardless of what the image shows? E.g., if during supervision we label images of kayaks as “crow” would the model still generalize to birds or does the model expect categories to have some level of coherence? 6/
Having established these preconditions to our task, we then find that models are also able to generalize (non-trivially) to hypernyms without ever having “seen” them explicitly, suggesting that LM representations support cross-modal generalization! 5/
We establish that this paradigm works in the first place with a vision encoder that has never been trained on language data (i.e., ❌ SigLIP ✅DINO), that the models learn the task on the lower-level categories themselves, and that the LMs indeed have taxonomic knowledge 4/
Taxonomic knowledge is interesting because of number of hypotheses about the learnability of category knowledge from linguistic cues, for both computational models and humans. Evidence of cross-modal generalization would lend strong support for these hypotheses! 3/
We use a VLM-training paradigm (frozen vision encoder w/o language training mapped to frozen LM) where we partially supervise on lower level categories during training, and then test if the LM recovers hypernymy knowledge from what it has seen in language data. 2/
What is the interplay between representations learned from (language) surface forms alone, and those learned from more grounded evidence (e.g.,vision)? Excited to share new work understanding “Cross-modal taxonomic generalization” in (V)LMs arxiv.org/abs/2603.07474 1/
Nearly 2 years ago, @jessyjli.bsky.social, @janetlauyeung.bsky.social, @valentinapy.bsky.social, and I decided that it's time to bring discourse structure to the center of NLP teaching.
This is not due to surface form differences alone – PCA on the hidden-state representations of the three types of propositions alongside their (near) synonymous counterparts (e.g., all/every vs. certain/some / indef. article/generics) shows that they are distinguished based on inductive constraints!
Finally, we tested the two models that robustly satisfied both these preconditions on tests mimicking the original Gelman et al. (2002) study, and observed them to show qualitatively similar patterns as children and adults (all > generic > some).
For presupposition #2, we created a new developmentally inspired dataset to test if models can answer questions targeting "all" and "some" in contexts presented to them in multiple modalities (language/vision) - only Qwen3 VL 4B and 8B got high accuracies across conditions!
For presupposition #1, we tested VLMs on their ability to robustly identify categories across multiple object-centric images under various negative sampling conditions:
We replicate Gelman et al. (2002), where subjects were taught novel properties of animals under 3 diff conditions and then asked if a specific animal had that property: All (all Xs have…) Some (some Xs have…) Generic (Xs have…) Subjects showed a consistent pattern: all > generic > some
“All bears have a property”, “Some bears have a property”, “Bears have a property” are different in terms of how the property is generalized to a specific bear – a great example of how language constrains thought! This holds for kids, adults, and according to our new work, (V)LMs! 🧵
I’ll be in Boston attending BUCLD this week — I won’t be presenting but I’ll be cheering on @najoung.bsky.social who will present at the prestigious SLD symposium about the awesome work by her group, including our work on LMs as hypotheses generators for language acquisition! 🤠👻