Interactions

Lecture 5

Emre Yucel

University of Texas at Austin

2026-10-01

Logistics

  • Team Case scores available (grading guide available in Canvas under Files)
  • What to do if your score was lower than expected?
  • AI practice tool — consider giving it a try if you are looking for extra practice

Motivation

Last week, we predicted evaluation score from beauty and gender, and the model forced the lines to be parallel:

What if we could build a more flexible model without forcing the lines to be parallel?

What is an interaction?

An interaction is an additional term in a regression model that allows the slope of one variable to depend on the value of another.

  • Without an interaction, a one-unit increase in beauty has the same predicted association with evaluations for both genders.
  • With an interaction, each gender can have its own beauty slope.

The question is: Does the relationship between beauty and evaluations differ by gender?

A categorical variable and a quantitative variable

Does beauty matter more for men, or for women?

  • We found that for the same level of attractiveness, male professors tend to get higher evaluation scores than female professors
  • But what if the effect of beauty depends on gender (is different for men vs women)?

Does beauty matter more for men, or for women?

Another way to think about it: what if non-parallel regression lines would fit the data better?

The algebra of interactions

The idea is to add a new variable that is itself the product of the two variables. Written in the same order R prints the output:

\widehat{\text{eval}} = \hat\beta_0 + \hat\beta_1(\text{beauty}) + \hat\beta_2(\text{male}) + \hat\beta_3(\text{beauty})(\text{male})

  • For female professors, male = 0, so the \hat\beta_2 and \hat\beta_3 terms become zero and drop out: \widehat{\text{eval}} = \hat\beta_0 + \hat\beta_1(\text{beauty})
  • For male professors, male = 1, so we get both a different intercept and a different slope for beauty: \widehat{\text{eval}} = (\hat\beta_0 + \hat\beta_2) + (\hat\beta_1+\hat\beta_3)(\text{beauty})

The interaction model

model1 <- lm(eval ~ beauty * gender, data=profs)
summary(model1)

Call:
lm(formula = eval ~ beauty * gender, data = profs)

Residuals:
     Min       1Q   Median       3Q      Max 
-1.83820 -0.37387  0.04551  0.39876  1.06764 

Coefficients:
                  Estimate Std. Error t value             Pr(>|t|)    
(Intercept)        3.89085    0.03878 100.337 < 0.0000000000000002 ***
beauty             0.08762    0.04706   1.862             0.063294 .  
gendermale         0.19510    0.05089   3.834             0.000144 ***
beauty:gendermale  0.11266    0.06398   1.761             0.078910 .  
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 0.5361 on 459 degrees of freedom
Multiple R-squared:  0.07256,   Adjusted R-squared:  0.0665 
F-statistic: 11.97 on 3 and 459 DF,  p-value: 0.000000147

Write out the fitted equations

Using the coefficient estimates from the model output, write out the fitted equation for:

  1. Female professors
  2. Male professors

Be ready to enter your answers in LC.

Fitted lines after adding the interaction

Women: \widehat{\text{eval}} = 3.891 + 0.088(\text{beauty})

Men: \widehat{\text{eval}} = 4.086 + 0.200(\text{beauty})

  • The estimated beauty slope is larger for men than for women
  • So the fitted gender gap grows with beauty
  • Whether we can trust that difference is a question for the confidence intervals

Main effects and interaction effects

  • In a model with an interaction term X_1X_2, you must also keep the main effects: the variables that are being interacted together. (R does this automatically when you use * in the lm command.)
  • The main effect of X_1 represents the predicted change in Y for a 1-unit change in X_1, holding X_2 constant at zero.

Main effects in our model

  • The main effect gendermale (0.20) is how much higher men are predicted to score than women, but only for an average-looking professor (beauty = 0).
  • The main effect beauty (0.09) is the predicted change in evaluation score for each additional beauty point, but only among women (gendermale = 0).

When zero is not meaningful

  • A main effect describes a slope or group difference when the other variable equals 0.
  • In this data set, beauty = 0 describes an average-looking professor, so 0 is a sensible reference point.
  • If 0 is outside the useful range, the main-effect coefficient may not answer a useful business question.

In that case, use the fitted equation to calculate a slope or prediction at an observed value of the other variable.

Coefficient confidence intervals

confint() gives a 95% confidence interval for each population coefficient.

confint(model1)
                   2.5 % 97.5 %
(Intercept)        3.815  3.967
beauty            -0.005  0.180
gendermale         0.095  0.295
beauty:gendermale -0.013  0.238
  • beauty: the slope for women, the reference group.
  • beauty:gendermale: the additional beauty effect for men.

Which row would you use to assess whether this additional effect differs from zero?

LC: do the slopes differ?

Quantity 95% CI
Women’s beauty slope (-0.005, 0.180)
Additional beauty effect for men (-0.013, 0.238)

Is there evidence at the 5% level that the beauty slopes differ for men and women?

A. Yes

B. No

Which interval supports your choice?

Interpreting the intervals

Quantity R coefficient
Women’s beauty slope beauty
Additional beauty effect for men beauty:gendermale
  • Among women, the 95% CI for the beauty slope is (-0.005, 0.180). It includes 0, so the data are inconclusive about a beauty-evaluation association for women.
  • The 95% CI for the additional beauty effect for men is (-0.013, 0.238). It includes 0, so this additional effect is not statistically significant at the 5% level.
  • The interaction row directly answers whether the two population slopes differ.

Can we build a CI for the men’s slope?

The estimated men’s slope is:

\hat\beta_{beauty}+\hat\beta_{beauty:male} = 0.088 + 0.113 = 0.200

But the original confint(model1) output has no row for this sum.

  • We cannot add standard errors or confidence-interval endpoints.
  • A direct calculation is beyond the scope of this course.

Releveling gives us the men’s coefficients

Make men the reference group. The fitted lines stay the same; only the coefficient meanings change.

profs_m <- profs %>% mutate(gender = relevel(factor(gender), "male"))
model1_men <- lm(eval ~ beauty * gender, data = profs_m)
summary(model1_men)

Call:
lm(formula = eval ~ beauty * gender, data = profs_m)

Residuals:
     Min       1Q   Median       3Q      Max 
-1.83820 -0.37387  0.04551  0.39876  1.06764 

Coefficients:
                    Estimate Std. Error t value             Pr(>|t|)    
(Intercept)          4.08595    0.03295 123.999 < 0.0000000000000002 ***
beauty               0.20027    0.04333   4.622           0.00000495 ***
genderfemale        -0.19510    0.05089  -3.834             0.000144 ***
beauty:genderfemale -0.11266    0.06398  -1.761             0.078910 .  
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 0.5361 on 459 degrees of freedom
Multiple R-squared:  0.07256,   Adjusted R-squared:  0.0665 
F-statistic: 11.97 on 3 and 459 DF,  p-value: 0.000000147

Releveling gives us the men’s CI

The beauty row now represents the men’s slope.

confint(model1_men)
                         2.5 %      97.5 %
(Intercept)          4.0211951  4.15070389
beauty               0.1151186  0.28543004
genderfemale        -0.2950979 -0.09509593
beauty:genderfemale -0.2383781  0.01306238

Slopes by gender

Slope estimate 95% CI
Women 0.088 (-0.005, 0.180)
Men 0.200 (0.115, 0.285)

Units: evaluation points per one-point increase in beauty.

The men’s interval excludes 0, but that does not show the men’s and women’s slopes differ. The interaction interval answers that question, and it includes 0.

Two categorical variables

Gender and tenure

  • Two simple models (eval ~ gender and eval ~ tenure) tell us that:
    • Male professors tend to get higher evaluation scores than female professors (\hat\beta_{\text{male}} = 0.168)
    • Professors with tenure tend to get lower evaluation scores than professors without tenure (\hat\beta_{\text{tenure}} = -0.173)
  • But what if the gender gap is different for professors with tenure vs professors without tenure?

Why won’t this model help?

model2_additive <- lm(eval ~ gender + tenure, data=profs)
summary(model2_additive)

Call:
lm(formula = eval ~ gender + tenure, data = profs)

Residuals:
     Min       1Q   Median       3Q      Max 
-1.93233 -0.35253  0.04747  0.44747  1.04747 

Coefficients:
            Estimate Std. Error t value             Pr(>|t|)    
(Intercept)  4.04167    0.05991  67.464 < 0.0000000000000002 ***
gendermale   0.17980    0.05136   3.501             0.000509 ***
tenureyes   -0.18914    0.06119  -3.091             0.002116 ** 
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 0.5442 on 460 degrees of freedom
Multiple R-squared:  0.04229,   Adjusted R-squared:  0.03813 
F-statistic: 10.16 on 2 and 460 DF,  p-value: 0.00004829

Why won’t this model help?

  • The additive model has only one gendermale coefficient (0.180)
  • That forces the gender gap to be the same with and without tenure
  • To let the gender gap differ by tenure status, we need gender * tenure

The interaction model

model2 <- lm(eval ~ gender * tenure, data=profs)
summary(model2)

Call:
lm(formula = eval ~ gender * tenure, data = profs)

Residuals:
     Min       1Q   Median       3Q      Max 
-1.89028 -0.36000  0.00972  0.40972  1.00972 

Coefficients:
                     Estimate Std. Error t value             Pr(>|t|)    
(Intercept)           3.86000    0.07585  50.890 < 0.0000000000000002 ***
gendermale            0.53615    0.10623   5.047          0.000000648 ***
tenureyes             0.05517    0.08796   0.627             0.530813    
gendermale:tenureyes -0.46105    0.12083  -3.816             0.000154 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 0.5363 on 459 degrees of freedom
Multiple R-squared:  0.07173,   Adjusted R-squared:  0.06567 
F-statistic: 11.82 on 3 and 459 DF,  p-value: 0.0000001795

Fill in the table

Copy down this table in your notes. Using the output on the previous slide, fill in the missing predictions. Write out the equation for each cell.

Female Male
Without tenure 3.86 ?
With tenure ? ?

Now interpret the results. What are your conclusions? (If you are having trouble interpreting the numbers, try graphing it by hand!)

Gender × tenure intervals

The reference group is female professors without tenure.

confint(model2)
                      2.5 % 97.5 %
(Intercept)           3.711  4.009
gendermale            0.327  0.745
tenureyes            -0.118  0.228
gendermale:tenureyes -0.699 -0.224
  • gendermale compares men with women without tenure.
  • tenureyes compares tenured with non-tenured professors among women.
  • gendermale:tenureyes compares the gender gap among tenured professors with the gender gap among professors without tenure.

Interpreting the interaction interval

  • The 95% CI for gendermale:tenureyes is (-0.699, -0.224).
  • The interval excludes 0, so there is evidence that the gender gap depends on tenure status.
  • We are 95% confident that the population gender gap (men minus women) is between 0.224 and 0.699 evaluation points smaller among tenured professors.

This interval concerns a difference between two gaps. It does not give the gender gap for either tenure group on its own.

LC: the gender gap among tenured professors

Using the original model2 output, can we determine whether the gender gap among tenured professors is statistically significant?

A. Yes, using the gendermale interval

B. Yes, using the gendermale:tenureyes interval

C. No, neither interval directly gives the gender gap among tenured professors

The gender gap among tenured professors

To get it, make “tenured” the reference group so that gendermale becomes the gap among tenured professors:

profs_t <- profs %>% mutate(tenure = relevel(factor(tenure), "yes"))
model2_tenured <- lm(eval ~ gender * tenure, data = profs_t)
confint(model2_tenured)

The gendermale row:

            2.5 % 97.5 %
gendermale -0.038  0.188
  • Without tenure: men score between 0.327 and 0.745 points higher (excludes 0)
  • With tenure: the 95% CI (-0.038, 0.188) includes 0, so we don’t have evidence of a gender gap among tenured professors

Two quantitative variables

The game

How the game works

  • A contestant chooses a case from 26 cases with hidden cash amounts.
  • Each round, the contestant opens other cases and learns which amounts are gone.
  • The banker offers to buy the chosen case. The contestant can accept or keep playing.

The remaining amounts and the round both change as the game progresses.

Our question: How does the banker’s offer relate to the average remaining at different rounds?

Data set

  • 77 rounds from 9 games of Deal or No Deal
  • Each row records a banker’s offer at one round of a game.
  • offer: the banker’s offer in dollars. average: the mean dollar value of the cases still in play. round: how far the game has progressed.

Might the banker offer a larger share of the remaining average later in the game?

First, predict the banker’s offer using the average remaining and the round, with no interaction.

model3 <- lm(offer ~ average + round, data=deal)
summary(model3)

Call:
lm(formula = offer ~ average + round, data = deal)

Residuals:
   Min     1Q Median     3Q    Max 
-48143 -19511  -1611  19314  58726 

Coefficients:
                Estimate   Std. Error t value            Pr(>|t|)    
(Intercept) -84382.33820   7293.84013  -11.57 <0.0000000000000002 ***
average          0.67797      0.03049   22.24 <0.0000000000000002 ***
round        13252.71669   1122.53557   11.81 <0.0000000000000002 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 24830 on 74 degrees of freedom
Multiple R-squared:  0.8971,    Adjusted R-squared:  0.8943 
F-statistic: 322.5 on 2 and 74 DF,  p-value: < 0.00000000000000022

Visualizing the model

Is this the right model?

The additive model assumes that an extra dollar in average changes the predicted offer by the same amount in every round.

  • As we look at each round, compare how quickly offers rise with the average remaining.
  • Would parallel lines capture the pattern, or do later rounds need steeper slopes?

An interaction would allow the slope of average to change as the game progresses.

Comparing early, middle, and late rounds

All panels use the same scales. These separate fitted lines help us assess whether one common slope is reasonable.

How does the slope change from Round 1 to Round 9?

The interaction model

model4 <- lm(offer ~ average * round, data=deal)
summary(model4)

Call:
lm(formula = offer ~ average * round, data = deal)

Residuals:
   Min     1Q Median     3Q    Max 
-37498  -5534  -1108   2167  44403 

Coefficients:
                  Estimate   Std. Error t value            Pr(>|t|)    
(Intercept)   13438.119288  7998.520955   1.680              0.0972 .  
average          -0.132290     0.060301  -2.194              0.0314 *  
round         -1204.893154  1193.578131  -1.009              0.3161    
average:round     0.123845     0.008885  13.939 <0.0000000000000002 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 13060 on 73 degrees of freedom
Multiple R-squared:  0.9719,    Adjusted R-squared:  0.9707 
F-statistic: 841.4 on 3 and 73 DF,  p-value: < 0.00000000000000022

Visualizing the model with an interaction

Write out the equations!

\widehat{\text{offer}} = 13438 + (-0.132)(\text{average}) + (-1205)(\text{round}) + 0.124(\text{average})(\text{round})

Group the terms by round:

\widehat{\text{offer}} = \underbrace{[13438 + (-1205)(\text{round})]}_{\text{intercept for this round}} + \underbrace{[-0.132 + 0.124(\text{round})]}_{\text{slope of average in this round}}(\text{average})

  • Round 1: \widehat{\text{offer}} = 12233 + (-0.008)(\text{average})
  • Round 2: \widehat{\text{offer}} = 11028 + (0.115)(\text{average})
  • Round 3: \widehat{\text{offer}} = 9823 + (0.239)(\text{average})
  • Round 4: \widehat{\text{offer}} = 8619 + (0.363)(\text{average})

What the interaction coefficient means

  • The slope of average increases by about $0.124 per dollar for each additional round. In friendlier units: each round, a $1,000 increase in the average remaining is worth about $124 more in the offer.
  • The impact of each additional round played on the banker’s offer increases by $0.124 for each additional dollar in the average amount remaining.

Offer-model confidence intervals

confint(model4)
                  2.5 %    97.5 %
(Intercept)   -2502.910 29379.149
average          -0.252    -0.012
round         -3583.691  1173.905
average:round     0.106     0.142
  • average is the slope of average at round = 0, outside the observed rounds.
  • round is the slope of round when average = 0 dollars.
  • average:round: we are 95% confident that each additional round increases the slope of average by between 0.106 and 0.142 dollars per dollar. The interval excludes 0.

LC: calculate the Round 7 slope

R coefficient Estimate
average -0.132
average:round 0.124

Calculate the estimated slope of average in Round 7.

Then interpret it: for a $1,000 increase in the average remaining, how much larger is the predicted offer?

Slopes at selected rounds

At Round r, the estimated slope of average is:

\hat\beta_{average}+r\hat\beta_{average:round} = -0.132 + 0.124r

Round Slope per $1 Change per $1,000
1 -0.008 $-8
5 0.487 $487
9 0.982 $982

These are point estimates. A confidence interval for one selected round would require combining coefficient uncertainty and covariance, which we will not calculate here.

Interpreting the selected slopes

  • In Round 1, a $1,000 increase in the average remaining changes the predicted offer very little.
  • In Round 5, it is associated with about a $487 larger predicted offer.
  • By Round 9, it is associated with about a $982 larger predicted offer, close to dollar-for-dollar.
  • The interaction confidence interval excludes 0, providing evidence that this slope changes across rounds.

LC: predict before we watch

A contestant reaches round 6 with $5, $5,000, $25,000, $200,000, and $750,000 still on the board.

Using the fitted interaction model:

  1. What offer does the model predict?
  2. What is the 95% prediction interval for one new offer?

Enter the predicted offer and the two interval endpoints in LC.

Prediction interval for one offer

  • Prediction interval: a range for one new offer is $99,607 to $152,237. Under the model assumptions, intervals constructed this way cover actual offers in about 95% of comparable new cases.
  • A rough approximation is the predicted offer \pm\ 2 \times \text{RSE}. Here, that gives ($99,794, $152,050).
  • It is usually better to use the prediction interval from R than a rough approximation.

Let’s see how well the model works to make a prediction!

When should you use interactions in a model?

From Perusall:

“I’m a bit confused on how you choose the variables to use to create an interaction. I understand the two regression equations are parallel, however, if you have 20+ variables within a data table, how would you choose which ones are most likely to have an interaction?”

When should you use interactions in a model?

  • Start with a substantive question: why might the relationship between a predictor and the outcome depend on another variable?
  • Examine the interaction estimate and its confidence interval. Is the difference large enough to matter, and how precisely can we estimate it?
  • Compare model fit as well. Adding a term cannot lower ordinary R^2, so look at adjusted R^2 and the residual standard error instead.

What interactions are not

Importantly, interactions are not about one X variable affecting another X variable (correlations between X variables)

  • Correlation question: Do male and female professors tend to have different beauty scores?
  • Interaction question: Is the beauty-evaluation slope different for male and female professors?

The interaction model addresses the second question. An observational association also does not, by itself, establish a causal effect.

What interactions are

Interactions let us model a situation where the relationship of one predictor variable and Y is different depending on the value of another X variable:

  • How much attractiveness matters for student evaluation scores depends on gender
  • How much gender matters for student evaluation scores depends on tenure
  • How much what you have left on the board matters depends on how far along you are in the game