Jamie Cummins
@jamiecummins
Advanced postdoc at University of Bern. Also sometimes @bennettoxford.bsky.social. Metascience stuff. Creator and developer of RegCheck (). STAR Editor at Psych Science. @error.reviews 🇮🇪
10/🧵 We also learned that the researchers using RegCheck have come from a variety of different scientific disciplines. Each discipline, of course, has its own considerations for what matters in comparisons. To facilitate this, we now offer default comparison dimensions that vary by field.
9/🧵 We also learned that many folk have been using RegCheck to compare Stage 1 and Stage 2 Registered Reports. To facilitate this, we now include a specific Registered Reports comparison mode that takes this into account.
8/🧵 Reports can now also be exported as HTML files, so they can be saved and viewed locally in the exact same way as they are online.
7/🧵 Here's what a full RegCheck workflow looks like from scratch - including linking a registration directly from OSF, and navigating the RegCheck report. Completed end-to-end in about 5 minutes in real time.
6/🧵 But our goal with RegCheck has always been to facilitate human comparisons. To this end, we now also include the 'Evidence' viewer: where you can view documents side-by-side, with the specific quotes from the documents highlighted in-context.
5/🧵 Probably the most important addition: a totally revamped RegCheck report layout. The report comes with two views. The first, 'Overview', gives a top-level summary of RegCheck decisions for each dimension. All summaries directly cite quotes from the paper, which can be clicked on to view.
3/🧵 User accounts! You can now create an account using your Google account or OCRID iD. Your previously-generated RegCheck reports can now be easily retrieved through your account. You change the privacy settings of your reports (public or private) and delete them whenever you want.
1/🧵 First: a new frontend, redesigned from the ground-up to be easier to navigate. Also including a highly-requested light mode (a begrudging addition).
Reading back on transcripts from a lecture I gave earlier this year:
5/ Are the findings of Argyle et al. influenced by analytic flexibility? Re-running their silicon sampling approach using different choices for model, temperature/reasoning effort, prompt content, and demographic information provided, I say: yes.
6/🧵 I have also added integration with the ClinicalTrials.gov API. You can now provide a ClinicalTrials.gov identifier and RegCheck will automatically retrieve the registration.
4/🧵 RegCheck's layout now reflects its purpose: to assist humans in doing a research activity that they otherwise do not do - not to replace them. Most importantly, the UI provides direct quotes from the papers and registrations to make it easier to verify/refute judgements made by the software.
3/🧵 This has two consequences for RegCheck V2: 1. It's slower: V1 took about 45 seconds to 1 minute to run. V2 can take up to 15 minutes! 2. It's dramatically better: Internal testing shows very impressive and comprehensive coverage, catching many details that V1 previously missed.
1/🧵 18 months ago, we launched RegCheck. It worked, but it baked in our assumptions about what should be compared. But that's a decision for each separate user, not us. RegCheck V2 puts that decision, and others, back in the user's hands.
Most saliently: the original cut-offs that have historically been used by Project Implicit to determine levels of bias have not taken this uncertainty into account. When we do, the criteria change a lot. An IAT score of 0 is compatible with "weak" and "moderate" bias in either direction.
More generally, all measures left a lot to be desired. With the best-performing measure, the IAT, a given participant's score would distinguish them from (on average) 50% of other participants.
The results were stark. The IAT and its variants did OK at distinguishing scores from zero. Measures beyond this were really poor.
People use implicit measures to measure individual folks' (implicit) attitudes. The scores people get on a task like the IAT are often treated like they're an accurate index of bias. For example: until very recently, you'd get a message like this on projectimplicit.net when you completed an IAT.
I was thrilled to have been invited by @sakshighai.bsky.social to speak to folk at LSE on Wednesday about methodological and inferential issues that have cropped up in social science attempts to study large language models!
With every LLM since GPT-4, I've tried a game: ask it to commit a 20 Questions guess to a cipher, we play 20 Questions, and then we see if what it claims to have been its original choice is consistent with its cipher. ChatGPT-5.1 Thinking is the first model to do this successfully!
My master thesis file name on my old university's thesis archive site still makes me chuckle.
Some of the questions used for evaluation explicitly allude to the "5 Bits" structure, but again, this wasn't included in the prompt. If one were to build a software based on LLMs to try to create Science articles, it would look very different to this.
Science writers, as the white paper elucidates, use an article structure called the "5 Bits" (screenshot 1). The prompts given to the LLM (screenshot 2) do not specify this. They do not provide good examples (as in one- or few-shot prompting). They generally do not follow prompting best-practices.
The LLM-based samples also varied substantially in their estimates of the between-scale correlation. The blue line the point estimate for the correlation in the human data (r = 0.26).
The silicon samples varied a lot in terms of how closely they modeled the response distribution of scales in the human data. But they were generally not good. See the shaded blue area in the two plots? That covers the 95% interval for where bootstrapped human data falls.
All of the silicon sample configurations were only, at best, moderately correlated with the human data when it came to preserving the ranking of participants. And many of them were negatively correlated with the human data.