Categorical predictors

STA 235 Lecture 4

Emre Yucel

University of Texas at Austin

2026-09-24

Categorical predictors with 2 categories

Evaluations Dataset

Evaluations of 463 instructors at UT Austin, 2000-2002, for each instructor:

  • eval: average student evaluation score of instructor
  • beauty: average beauty rating of instructor from six-student panel
  • gender: male or female
  • credits: single or multi-credit course
  • age: age of instructor
  • (and others)

What’s the impact of gender on student evaluations?

Incorporating gender into a regression model

  • Gender is a categorical variable (male or female in this data set) so we can’t use it as-is as a predictor
  • What we’ll do is recode gender into the quantitative “dummy” variable 1 = male, 0 = female
  • The ordering is arbitrary. R will choose the alphabetically first category as the 0, or baseline level
  • You will be able to tell which coefficient belongs to which level, i.e. gendermale means male=1, making female=0

model0 <- lm(eval ~ gender, data=profs)
summary(model0)

Call:
lm(formula = eval ~ gender, data = profs)

Residuals:
     Min       1Q   Median       3Q      Max 
-1.96903 -0.36903  0.03097  0.43097  0.99897 

Coefficients:
            Estimate Std. Error t value Pr(>|t|)    
(Intercept)  3.90103    0.03933   99.19  < 2e-16 ***
gendermale   0.16800    0.05169    3.25  0.00124 ** 
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 0.5492 on 461 degrees of freedom
Multiple R-squared:  0.0224,    Adjusted R-squared:  0.02028 
F-statistic: 10.56 on 1 and 461 DF,  p-value: 0.001239

Compare the fitted model to the means of the two groups:

\widehat{\text{eval}} = 3.901 + 0.168\times \text{male}

profs %>% 
  group_by(gender) %>%
  summarize(mean(eval))
# A tibble: 2 × 2
  gender `mean(eval)`
  <chr>         <dbl>
1 female         3.90
2 male           4.07

What’s going on here?

Plotting the Regression

Selecting a different reference category

You can also pick a reference category using relevel:

# Create a new data set where the gender variable has the "male" 
# category as the reference category
profs.releveled = profs %>% 
  mutate(gender=relevel(factor(gender), "male"))

# Fit the model with this new data set
lm(eval ~ gender, data=profs.releveled)

Call:
lm(formula = eval ~ gender, data = profs.releveled)

Coefficients:
 (Intercept)  genderfemale  
       4.069        -0.168  

Plotting the Releveled Regression

Incorporating gender into a multiple regression model

  • We can easily include categorical variables in multiple regression models too
  • We still need to think carefully about their interpretation, just like other multiple regression coefficients.

model1 <- lm(eval ~ beauty + gender, data=profs)
summary(model1)

Call:
lm(formula = eval ~ beauty + gender, data = profs)

Residuals:
     Min       1Q   Median       3Q      Max 
-1.87196 -0.36913  0.03493  0.39919  1.03237 

Coefficients:
            Estimate Std. Error t value Pr(>|t|)    
(Intercept)  3.88377    0.03866  100.47  < 2e-16 ***
beauty       0.14859    0.03195    4.65 4.34e-06 ***
gendermale   0.19781    0.05098    3.88  0.00012 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 0.5373 on 460 degrees of freedom
Multiple R-squared:  0.0663,    Adjusted R-squared:  0.06224 
F-statistic: 16.33 on 2 and 460 DF,  p-value: 1.407e-07

Gender Specific Equations

\widehat{\text{eval}} = \underbrace{[3.884 + 0.198 \times \text{male}]}_{\text{gender specific intercept}} + 0.149 \times \text{beauty}

For females we have \widehat{\text{eval}} = [3.884 + 0.198 \times 0] + 0.149 \times \text{beauty}

For males we have \widehat{\text{eval}} = [3.884 + 0.198 \times 1] + 0.149 \times \text{beauty}

Building a multiple regression model using gender

A multiple regression predicting evaluation score from beauty and gender effectively fits two parallel regression lines:

Categorical predictors with 3+ categories

Is there a generation gap?

  • The variable generation is either silent (born before 1945), boomer (born 1945-1964), or genx (born after 1965)
  • We already know that there is a difference between male and female professors, and that beauty matters too.
  • How would I build a model to determine whether generation matters above and beyond gender and beauty? (i.e., do professors of the same gender and beauty get different evaluations depending on their generation?)

Why not just convert to numerical variables?

How would things go wrong if we just recoded the generation variable as numeric (e.g., 0 = silent, 1 = boomer, 2 = gen X) and included that in the model?

Dummy Coding

Let’s arbitrarily pick boomers as a reference category:

Category genx silent
Boomers 0 0
Gen Xers 1 0
Silent Gens 0 1

R will do this automatically when you add a categorical variable with 3+ categories to a regression (it will arbitrarily pick a reference category)!

model2 <- lm(eval ~ beauty + gender + generation, data=profs)
summary(model2)

Call:
lm(formula = eval ~ beauty + gender + generation, data = profs)

Residuals:
     Min       1Q   Median       3Q      Max 
-1.91613 -0.36042  0.03609  0.42282  1.04398 

Coefficients:
                 Estimate Std. Error t value Pr(>|t|)    
(Intercept)       3.89727    0.04403  88.521  < 2e-16 ***
beauty            0.14021    0.03304   4.243 2.67e-05 ***
gendermale        0.22230    0.05345   4.159 3.82e-05 ***
generationgenx   -0.02831    0.06149  -0.460   0.6454    
generationsilent -0.16292    0.07992  -2.039   0.0421 *  
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 0.536 on 458 degrees of freedom
Multiple R-squared:  0.07477,   Adjusted R-squared:  0.06669 
F-statistic: 9.253 on 4 and 458 DF,  p-value: 3.386e-07


                 Estimate Std. Error t value Pr(>|t|)    
(Intercept)       3.89727    0.04403  88.521  < 2e-16 ***
beauty            0.14021    0.03304   4.243 2.67e-05 ***
gendermale        0.22230    0.05345   4.159 3.82e-05 ***
generationgenx   -0.02831    0.06149  -0.460   0.6454    
generationsilent -0.16292    0.07992  -2.039   0.0421 *  

All else equal (i.e., among professors of the same gender and beauty):

  • Gen X professors are predicted to get scores that are 0.03 points below those of boomers
  • Silent gen professors are predicted to get scores that are 0.16 points below those of boomers
  • Only the boomer/silent generation difference is statistically significant; gen X professors are not significantly different than boomers
  • In other words: age only seems to matter if you are really old

Age & gender specific regression lines

From each combination of age and gender, we have 6 different regression lines

Confidence intervals

confint(model2)
                       2.5 %       97.5 %
(Intercept)       3.81075083  3.983789776
beauty            0.07527196  0.205141913
gendermale        0.11725643  0.327335872
generationgenx   -0.14913816  0.092517802
generationsilent -0.31997630 -0.005870264

Among professors of the same gender and generation:

  • We are 95% confident that an additional beauty point raises evaluation score between 0.075 and 0.205 points
  • We have evidence of a statistically significant relationship between beauty and evaluation after adjusting for gender and generation

Confidence intervals

confint(model2)
                       2.5 %       97.5 %
(Intercept)       3.81075083  3.983789776
beauty            0.07527196  0.205141913
gendermale        0.11725643  0.327335872
generationgenx   -0.14913816  0.092517802
generationsilent -0.31997630 -0.005870264

Among professors of the same gender and beauty:

  • The difference between genx and boomer could be anywhere between -0.149 or 0.093 with 95% confidence; we can’t really conclude a relationship
  • The difference between silent and boomer could be anywhere between -0.32 or -0.006 with 95% confidence; even though this is statistically significant, it could be a rather weak effect

Predictions

While we could plug in values to our regression equation for prediction, it gets unwieldy rather quickly

avg_older_male <- data.frame(beauty=0, gender='male', generation='silent')
predict(model2, newdata=avg_older_male, interval='prediction')
       fit      lwr      upr
1 3.956643 2.893971 5.019315

Using R, we have the benefit of being able to get a prediction interval as well

  • There’s a 95% chance that an average looking, male, silent generation professor will have an evaluation score in this range
  • 95% of average looking, male, silent generation professors will receive evaluations between 2.89 and 5.02