Martin Zettersten
@mzettersten
Asst Prof UCSD Cognitive Science PI LIL Lab UCSD: language development | cognitive development | learning (he/his)
Rec #4: Match targets and distractors on animacy (+anything that affects saliency). Images of animate things grab infants' attention. This makes accuracy look worse for inanimates when matched w/ an animate. If saliency is mismatched for target vs. distractor, it can completely mask word knowledge
Rec #3 (actually #5 in the paper, there's a lot!): Increasing # of trials per participant really matters for reliable and valid measurement. Especially for individual difference studies, it's essential to increase the number of observations per participant as much as possible to have reliable signal
Rec #2: Avoid baseline correction. One common way to deal with saliency differences between images is to correct accuracy by subtracting baseline looking to the two images. We find: this practice substantially decreases reliability: noisy measure - (correlated) noisy measure = extra noisy measure!
Rec #1: Measure accuracy over a long analytic window. Researchers tend to constrain analyses to shorter time windows that end just after when average looking to the visual referent peaks. We find: expanding the time window noticeably improves reliability (the later parts of the window have signal!)
LWL has been an incredibly useful paradigm for studying language development. BUT there are many, many decisions that go into analyzing complex timecourse data and designing studies. Our idea: let's use Peekbank to take a data-driven look at what choices improve measurement. Our top recs below.
also, check out the supplement either at the journal or in the preprint (osf.io/preprints/ps...) for a ton more on looking time data and measurement, including one of my favorite figures, showing how test-retest reliability depends on both the number of trials and the sample
At the same time, the 2 evidence sources still do not fully agree. There are key differences in the effects of moderators. E.g., ManyBabies1 found that IDS preference was larger for older infants & headturn-based designs. We did not find similar effects in the meta-analysis. 7/
In the end, this led us to a pretty satisfying result: the meta-analysis and the multi-site replication almost *perfectly* agree on the average effect size (d~0.35). Check out the squint-inducing plot below with *many* effect sizes. 6/
3 tidbits: (1) The effect is bigger in adults than kids - and rule-based categories are quite hard for kids. (2) Kids' verbal knowledge of specific feature names correlates with learning (suggesting a connection w/ language) (3) Love these goofy aliens in the kid-friendly version of task. (2/2)
Was anything stable across test sessions? It turns out the answer is YES. Infants’ average *overall* looking times between sessions (and also how many trials they contributed) was robustly correlated across sessions. It’s just that *preferential* looking was not consistent (5/6)
Interestingly, as we increased the number of trials required for inclusion, test-retest correlations increased somewhat (more trials per kid helps!)- but even the largest correlations were quite small (and likely too small to be practical for studying indiv differences) (4/6)
So, as part of ManyBabies1, a number of labs brought in babies for a second test session. Despite a large infant sample (N=158), we saw no evidence of test-retest reliability in preregistered analyses. Correlations between looking time in session 1 and 2 was small (r=.09) (3/6)