Just Emil Kirkegaard Things

Just Emil Kirkegaard Things

Interactions are generally not important in social science

A large scale analysis

Emil O. W. Kirkegaard's avatar
Emil O. W. Kirkegaard
Oct 03, 2026
∙ Paid

Some years ago both me and Jonatan Pallesen looked into the generality of statistical interaction effects. These look like this:

In real life, however, interactions are mostly not a real thing. This is despite their sometimes great intuitive appeal and well-known examples (e.g. mating success/attractiveness as function of height and sex and their interaction). So I decided to hammer home this message so that there is something robust and convincing one can refer to to prove the point. For this purpose, I reached for my trusty revolver OKCupid dataset. It’s very large, and has an extreme variety of items. Yes, it is from a dating site, but every other well-known finding I’ve tried replicating in the dataset worked fine. The very fact that it works so well, as do other convenience samples, is evidence against the importance of interactions. So here’s what I did:

  1. Item-level data: pick a random ordinal variable as the outcome (Y), and 2 others as the predictors (X1, X2). Then fit the additive and interaction models: 1) Y ~ X1 + X2, vs. 2) Y ~ X1 * X2. Repeat 1000s of times. This approach needs ordinal regression since the outcomes are ordinal (all items have 2-4 answer options, which in the base case are ordinal or binary, sometimes nominal).

  2. Scale-level data: aggregate items into scales from my recent 10-dimension model using a mixture of item types (binary, ordinal, nominal), then repeat the same approach as above using scales instead of items.

  3. Scale-level data: aggregate items into scales that I had previously investigated in the cat-dog study, then repeat the same approach as above. The scales here are quite different from the ones above as they were handpicked to measure particular dimensions of interest instead of ‘falling out of’ the data.

  4. Variants of the above swapping X2 for sex, to specifically look at sex interactions, the most intuitively plausible subset.

First we have to think of what we care about. The most obvious is to look at the p-values. This was the focus of the prior studies on interactions in this dataset. It is, however, a mistake I think. The dataset is so large that many terms reach p<5% even with tiny effect sizes. What we really care about is not so much whether there is literally any signal larger than pure noise, but whether these signals are of practical concern. For this reason, here we will focus on changes in model explanatory power, that is, R² values (which can be deceptive but bear with me).

Alright, so what was found? It can be summarized in a few ways graphically:

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Emil O. W. Kirkegaard · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture