Alex Crenshaw
@crenshaw
Assistant Professor of Psychology. Couples, clinical trials, stats & methods. Views my own.
How much do these choices matter? Using a convenience sample of published psychometric data from common measures, differences in operationalizing the same JT formula produced RCI threshold discrepancies as large as 79%, or ~1 SD unit... (6/n)
We reviewed 226 clinical trials from four clinical psych journals (2020–2023). Twenty-nine formally assessed reliable change. Most (22/29; 76%) used the JT method. However, we found at least 5 distinct operationalizations of this same JT formula... (5/n)
New paper! We find that the Reliable Change Index (RCI), used to statistically evaluate whether an individual patient improved, is operationalized so inconsistently across clinical trials in psychology that results are often incomparable... (1/n)
We also need better reporting. Only 35% of studies provided enough info to know the operationalization, and 56% reported no information at all. Reporting recommendations:
Using baseline SD meets 7 of 8. All criteria aren't equally important, but baseline SD stands far above other options, where main limitation is potential SMD inflation due to range restriction common at baseline in clinical trials. Second best is post-tx control group SD.
My lab conducted a methods review of published clinical psych trials to see how they operationalized the SMD. Of 126 studies reporting an SMD, only 44 (35%) provided sufficient information. Of those, we identified at least 9 different ways the SMD was computed.
Past work shows SMDs computed using different SDs can produce widely differing values for the same effect. Here is just one example:
New commentary out in JCCP! In which I try to solve the problem of incomparable effect size estimates in clinical trials by recommending a standard, one-size-fits-most way to compute the standardized mean difference (SMD; e.g., Cohen’s d) in psychology trials. doi.org/10.1037/ccp0...
We also need better reporting. Only 35% of studies provided enough info to know the operationalization, and 56% reported no information at all. Reporting recommendations:
Using baseline SD meets 7 of 8, far more than other options, w/ main limitation being potential SMD inflation due to range restriction common at baseline in clinical trials. Second best is post-tx control group SD.
My lab conducted a methods review of published clinical psych trials to see how they operationalized the SMD. Of 126 studies reporting an SMD, only 44 (35%) provided sufficient information. Of those, we identified at least 9 different ways the SMD was computed.
Past work shows SMDs computed using different SDs can produce widely differing values for the same effect. Here is just one example:
Across 24 tx-phase universes: PTSD symptom improvement was fully consistent: both groups improved, w/ no group differences. Relationship satisfaction was partially consistent. Fear of intimacy was highly dependent on analysis approach. The multiverse differentiated robust from fragile results
We demonstrate the approach using a small RCT comparing prolonged exposure vs cognitive behavioral conjoint therapy for PTSD, which had serious real-world challenges: 65% dropout in one arm, only 4 people at 12-month follow-up, and a common challenge of txs differing in session count and length
We expand Del Giudice & Gangestad's (2021) framework to trials and introduce "Type D" decisions, where a best-practice option is itself flawed (e.g., ITT when dropout is severe and differential). All options are flawed, so converging evidence across them is stronger than a single result
1/7 Clinical trials should preregister analyses. But even careful plans leave room for defensible alternatives, and problems often arise that create decisions no one anticipated. Instead of picking just one defensible option, what if you ran them all? New paper in @collabrapsychology.bsky.social
6/7 An interesting nuance: Sampling error in the SD and reliability terms were positively correlated, so they partially cancelled each other out. The full RCI was less variable than either component alone. But this offsetting was incomplete and couldn't rescue small samples.
1/7 The Reliable Change Index is supposed to be a common ruler for deciding whether a patient truly improved. But what if that ruler changes size from study to study? Recent paper in the International Journal of Social Research Methodology: doi.org/10.1080/13645579.2025.2585286