# A tibble: 2 × 2
gender `mean(eval)`
<chr> <dbl>
1 female 3.90
2 male 4.07
What’s going on here?
Plotting the Regression
Selecting a different reference category
You can also pick a reference category using relevel:
# Create a new data set where the gender variable has the "male" # category as the reference categoryprofs.releveled = profs %>%mutate(gender=relevel(factor(gender), "male"))# Fit the model with this new data setlm(eval ~ gender, data=profs.releveled)
For females we have
\widehat{\text{eval}} = [3.884 + 0.198 \times 0] + 0.149 \times \text{beauty}
For males we have
\widehat{\text{eval}} = [3.884 + 0.198 \times 1] + 0.149 \times \text{beauty}
Building a multiple regression model using gender
A multiple regression predicting evaluation score from beauty and gender effectively fits two parallel regression lines:
Categorical predictors with 3+ categories
Is there a generation gap?
The variable generation is either silent (born before 1945), boomer (born 1945-1964), or genx (born after 1965)
We already know that there is a difference between male and female professors, and that beauty matters too.
How would I build a model to determine whether generation matters above and beyondgender and beauty? (i.e., do professors of the same gender and beauty get different evaluations depending on their generation?)
Why not just convert to numerical variables?
How would things go wrong if we just recoded the generation variable as numeric (e.g., 0 = silent, 1 = boomer, 2 = gen X) and included that in the model?
Dummy Coding
Let’s arbitrarily pick boomers as a reference category:
Category
genx
silent
Boomers
0
0
Gen Xers
1
0
Silent Gens
0
1
R will do this automatically when you add a categorical variable with 3+ categories to a regression (it will arbitrarily pick a reference category)!
The difference between genx and boomer could be anywhere between -0.149 or 0.093 with 95% confidence; we can’t really conclude a relationship
The difference between silent and boomer could be anywhere between -0.32 or -0.006 with 95% confidence; even though this is statistically significant, it could be a rather weak effect
Predictions
While we could plug in values to our regression equation for prediction, it gets unwieldy rather quickly