bastian bunzeck
@bbunzeck
wondering how humans and computers learn and use language 👶🧠🗣️🖥️💬 the work is mysterious and important, see bbunzeck.github.io mentally still at @ucsandiego.bsky.social, phd @clausebielefeld.bsky.social
Great slide by @kanishka.bsky.social on how he envisions the link between machine cognition and human cognition. 🤖👶
Always good to remember as people who study the capital L in LMs: language (learning) was one of the very first problems tackled in connectionist modeling 🫡
Ilaria Appel compared fasttext embeddings of German colour terms with their RGB vectors. Surprisingly, these measures are not correlated very well. She is now collecting human similarity ratings, both visual and language-based, to see how they relate to these measures. 🎨
@mdhk.net tracks the development of linguistic features during pretraining of self-supervised speech encoders trained on Dutch. Through layers and across pretraining, these patterns largely follow the traditional linguistic levels of description. Find her bsky thread at tiny.cc/learningSSL-thread 🎤
Iuliia Lysova investigates how fuzzy linguistic categories are represented in language models. Turns out that, e.g., content words require much larger computational subgraphs at prediction time than function words. But even this phenomenon is gradient 🤙
Including the best LM leaderboard I’ve ever seen. When will OpenRouter support this “12-14-year-olds” model?
Alvin Tan showed me that VLMs broadly align with children when it comes to task difficulty, but not when it comes to item difficulty. Some tasks easy for children, like image rotation, are almost impossible for VLMs. 🤨
Maxime Poli shows how pretraining speech LMs on ambient sounds (animal vocalizations) instills useful inductive biases for ABX discrimination in them 🙌
Oh yeah, I’m in Gothenburg for the next two weeks for this incredibly exciting summer school 🎉
Oh wow, that is so similar to what I found!! Second author is off as well, third author even gender-swapped 🥲
posting from uc san diego for a bit. left bielefeld and the rain behind to raise some more babylms on syntax and lemons. stay tuned 🍋☀️
Also, the dialogue pairs taken from real data provide a much better reward signal than those synthetically generated. Here, real data beats synthetic data quite drastically!
While performance on most benchmarks also decreases, it actually increases on our own dialogue minimal pairs (real vs. randomly sampled adjacency pairs), from 64% for the pretrained model to 68% after reinforcement learning, even outperforming the BabyLM baseline by 10%.
We replaced the standard BabyLM corpus with 10M tokens of dialogue triplets from CHILDES and trained an autoregressive model that we call llamalogue.
As part of this year's BabyLM challenge, we (researchers from @gronlp.bsky.social and @clausebielefeld.bsky.social diverged from established pretraining paradigm by training only on dialogue data from CHILDES.
From conference to conference: September ends with a trip to #IWCS in beautiful Düsseldorf. Hyped for two days of semantics (and two more days of construction grammar and NLP). 🥳
From conference to conference — after last week’s #semdial I am at #konvens in Hildesheim this week. I will be presenting out German BabyLM Corpus (with @simphon.bsky.social) and our PI Sina Zarrieß will give a Keynote on BabyLMs tomorrow. 🥳
I will present a poster on the First Language article I wrote with Holger Diessel now at #semdial 😁💬