Just Emil Kirkegaard Things

Just Emil Kirkegaard Things

Interactions are still not important: not in medical data either

Analyzing NHANES data using improved methods still finds almost nothing

Emil O. W. Kirkegaard's avatar
Emil O. W. Kirkegaard
Oct 04, 2026
∙ Paid

Yesterday I wrote about a study I did of the OKCupid dataset. The point was to show that interaction effects are almost always tiny and you should adjust your effect size priors downwards, far far downwards. Naturally, some people objected to the dataset being used, which is amusingly a kind of interaction claim (that is, that interaction effects are weaker in dating site data than other data). However, this is all fine because there’s lots of other large scale datasets one can use for the same research design. So I decided to do it again and for this I chose the American NHANES. These large studies are done every 2 years and measure a large number of mostly biological and medical properties (e.g. blood work) that aren’t impacted by self-report issues or being on a dating site. To boost the sample size, I used data from the last 2 waves (from 2017–2020 and 2021–2023, total n = 14,608 adults). From these I selected 139 variables measured for at least 8,000 people; each model uses the people with data on all three of its variables (median n = 11,024).

The analytic strategy was about the same but expanded a bit:

  1. We pick 3 random variables, set the first as the outcome, and the other 2 as the predictors, thus getting the general model Y ~ X1 + X2. This we test against the interaction model Y ~ X1 * X2. The variables are continuous variables, occasionally quasi-continuous, so we don’t have to use anything other than OLS

  2. In variants of the above, we swap X2 for sex (binary), since sex interactions are probably those that are most common/largest.

  3. We use LOOCV to assess how much variance the interactions add. This is probably better than trusting the analytic approximations (adjusted R²).

Alright, so did this dataset change the conclusions? Not really. We find just about the same as before:

Or if you prefer a table:

So in a sizable number, even a majority, of models the variance explained % gain was negative, which is to say that there was so little signal of an interaction that model overfitting caused the gain to be negative. Some somewhat sizable interactions were found in about 1% of the models. But if you look more closely at them, you can see the problem is not an interaction but that the main effects are misspecified as linear when they are really nonlinear:

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Emil O. W. Kirkegaard · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture