STAT 516 hw 3

Author

Karl Gregory

The code below imports into R a data set containing the self-reported heights (in feet and inches), lengths of index and pinky fingers (in millimeters), shoe size, and shoe size gender (“m”/“w”) of several statistics students. The code then removes an observation with shoe size gender given as “uk” and adds to the data set a column containing the heights in centimeters.

# import the data
link <- url("https://gregorkb.github.io/data/height_length_index_pinky_shoesize_combined.csv")
hl0 <- read.table(link,sep=",",header=T)

# remove a "uk" show size
rmv <- which(hl0$shmw == "uk")
hl <- hl0[-rmv,]

# convert ft and in. to centimeter heights
hl$height <- (hl$ft*12 + hl$in.)*2.54

# view the first few rows of the data frame
head(hl)
  ft in. ind pnk shoe shmw            class height
1  5   4  75  65  8.0    w STAT_516_sp_2026 162.56
2  5   4  70  56  7.0    w STAT_516_sp_2026 162.56
3  5   8  70  52 10.0    w STAT_516_sp_2026 172.72
4  5   5  68  62  7.5    w STAT_516_sp_2026 165.10
5  5   6  71  55  9.5    m STAT_516_sp_2026 167.64
6  6   4  78  65 13.0    m STAT_516_sp_2026 193.04

It is of interest to use the multiple linear regression model to predict the height of a person based on his or her index and pinky finger lengths, shoe size, and shoe size gender.

1.

Make a figure which shows scatterplots for all pairs of variables in the data set. Comment on which pairs of variables appear to be highly correlated.

plot(hl)

Height appears positively linearly related to both finger length 
measurements as well as to the shoe size. The heights 
appear to differ between shoe size gender. The index 
finger length appears highly positively correlated 
with the pinky finger length. Shoe size appears positively 
correlated with index and pinky finger lengths, and those wearing 
mens shoes appear to have longer index and pinky 
finger lengths. Lastly, the shoe sizes reported by those 
wearing mens shoes appear to be greater on average than 
those reported by those wearing womens shoes.

2.

Fit a multiple linear regression model for predicting height based on index and pinky finger length, shoe size, and shoe size gender. Then:

2.a

Report the estimated value of the regression coefficient for each covariate.

lm_all <- lm(height ~ ind + pnk + shoe + shmw, data = hl)
summary(lm_all)

Call:
lm(formula = height ~ ind + pnk + shoe + shmw, data = hl)

Residuals:
     Min       1Q   Median       3Q      Max 
-11.1911  -2.4660   0.2083   3.3303   9.9596 

Coefficients:
            Estimate Std. Error t value Pr(>|t|)    
(Intercept) 127.0611     7.6453  16.620  < 2e-16 ***
ind           0.1489     0.1602   0.929    0.356    
pnk           0.1562     0.1503   1.040    0.303    
shoe          3.0248     0.3969   7.622 1.64e-10 ***
shmww        -7.9098     1.3934  -5.677 3.73e-07 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 4.597 on 63 degrees of freedom
Multiple R-squared:  0.8358,    Adjusted R-squared:  0.8254 
F-statistic: 80.17 on 4 and 63 DF,  p-value: < 2.2e-16
The summary() function applied to the output of the lm() 
function gives the estimated coefficients in the first column of 
the coefficients table.
2.b

Give the value of the estimated standard error \(\widehat{\text{s.e.}}(\hat \beta_j) = \hat \sigma/\sqrt{\Omega_{jj}}\) for each of the covariates.

The summary() function applied to the output of the lm() 
function gives the estimated standard errors in the second 
column of the coefficients table.
2.c

Give the value of the test statistic \(T_{\operatorname{test}} = \hat \beta_j / (\hat \sigma/\sqrt{\Omega_{jj}})\) for each of the covariates.

The summary() function applied to the output of the lm() 
function gives the estimated standard errors in the third 
column of the coefficients table.
2.d

Give the p value for testing \(H_0\): \(\beta_j = 0\) versus \(H_1\): \(\beta_j \neq 0\) for each of the covariates.

The summary() function applied to the output of the lm() 
function gives the p values in the fourth 
column of the coefficients table.
2.e

Give an interpretation to the estimated coefficient \(\hat \beta_j\) for the shoe size covariate.

An increase in shoe size by one size corresponds to an 
increase in height of 3.024762cm, on average, 
with all other characteristics held fixed.
2.f

Give an interpretation to the estimated coefficient \(\hat \beta_j\) for the shoe size gender covariate.

Those wearing women's shoes are 7.909836cm shorter, 
on average, than those wearing men's shoes, when all 
other characteristics are held fixed.
2.g

Do the index and pinky finger lengths appear to be important predictors of height?

In this model, they do not appear to be 
statistically significant predictors of 
height, as the p values for testing 
whether their regression coefficients are 
equal to zero are not small.
2.h

Give an estimate of \(\sigma\), the standard deviation of the error term in the multiple linear regression model.

n <- nrow(hl)
p <- 4
sigmahat <- sqrt(sum(lm_all$residuals^2)/(n-(p+1)))
We obtain the estimate 4.597329.
2.i

Produce a normal quantile-quantile plot of the residuals as well as a residuals versus fitted values plot. Comment on whether you believe the assumptions of the multiple linear regression model to be satisfied.

plot(lm_all,which=1)

plot(lm_all,which=2)

The normal q-q plot does not show any alarming departures 
from normality. The residuals versus fitted values plot exhibits some 
fanning out from the left to the right, suggesting that the variance of 
the heights may not be constant over all covariate values. In 
particular, it appears that the variance in height is greater when the 
predicted height is greater (at greater values of the finger length and shoe 
size covariates.

3.

Fit a simple linear regression model using only the shoe size gender covariate. Then:

3.a

Give an interpretation of the estimated regression coefficient for the shoe size gender covariate.

lm_mw <- lm(height ~ shmw, data = hl)
summary(lm_mw)

Call:
lm(formula = height ~ shmw, data = hl)

Residuals:
     Min       1Q   Median       3Q      Max 
-17.2720  -6.0036  -0.2078   5.5880  23.3680 

Coefficients:
            Estimate Std. Error t value Pr(>|t|)    
(Intercept)  179.832      1.243 144.672  < 2e-16 ***
shmww        -16.348      1.784  -9.162 2.24e-13 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 7.354 on 66 degrees of freedom
Multiple R-squared:  0.5598,    Adjusted R-squared:  0.5532 
F-statistic: 83.94 on 1 and 66 DF,  p-value: 2.244e-13
Those wearing women's shoes are 16.34836cm shorter, 
on average, than those wearing men's shoes. We do not need to 
say 'with all other characteristics held fixed', because 
we have not included any other covariates in the model.
3.b

Why does this covariate appear to have a different effect when it is the sole covariate in the model?

When no other covariates are included, the shoe size 
gender covariate becomes equal to the difference between the 
average height of those wearing men's shoes and 
that of those wearing women's shoes, without regard for 
any other characteristics.

4.

Fit a multiple linear regression model using only the index and pinky finger lengths as predictors of height.

4.a

Does either covariate in this model appear to be significantly related to the height?

lm_fingers <- lm(height ~ ind + pnk, data = hl)
summary(lm_fingers)

Call:
lm(formula = height ~ ind + pnk, data = hl)

Residuals:
    Min      1Q  Median      3Q     Max 
-15.302  -5.802   0.048   5.703  17.190 

Coefficients:
            Estimate Std. Error t value Pr(>|t|)    
(Intercept)  88.9344    11.4224   7.786 7.01e-11 ***
ind           0.8865     0.2657   3.337   0.0014 ** 
pnk           0.3395     0.2678   1.268   0.2093    
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 8.248 on 65 degrees of freedom
Multiple R-squared:  0.4547,    Adjusted R-squared:  0.4379 
F-statistic:  27.1 on 2 and 65 DF,  p-value: 2.761e-09
The index finger length appears to be 
significantly related to height, but the pinky finger coefficient 
has a large p value, suggesting that this covariate does 
not contribute a significant information beyond 
that contributed by the index finger length.
4.b

What proportion of the total variation in heights does this model explain?

SSreg <- sum((lm_fingers$fitted.values - mean(hl$height))**2)
SStot <- sum((hl$height - mean(hl$height))**2)
Rsq <- SSreg/SStot
This is the coefficient of determination, 
appearing as 
'Multiple R-squared' in the 
summary output. The value is 0.4546846.

5.

Fit a multiple linear regression model using only the shoe size and shoe size gender covariates.

5.a

Does either covariate in this model appear to be significantly related to the height?

lm_shoes <- lm(height ~ shoe + shmw, data = hl)
summary(lm_shoes)

Call:
lm(formula = height ~ shoe + shmw, data = hl)

Residuals:
     Min       1Q   Median       3Q      Max 
-10.0529  -2.8899   0.4557   3.8680   8.5225 

Coefficients:
            Estimate Std. Error t value Pr(>|t|)    
(Intercept) 142.6470     3.9332  36.267  < 2e-16 ***
shoe          3.4892     0.3613   9.657 3.50e-14 ***
shmww        -8.6102     1.4035  -6.135 5.67e-08 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 4.749 on 65 degrees of freedom
Multiple R-squared:  0.8192,    Adjusted R-squared:  0.8137 
F-statistic: 147.3 on 2 and 65 DF,  p-value: < 2.2e-16
Both of these covariates appear to be 
significant predictors of height, as each has a 
small p value.
5.b

What proportion of the total variation in heights does this model explain?

SSreg <- sum((lm_shoes$fitted.values - mean(hl$height))**2)
SStot <- sum((hl$height - mean(hl$height))**2)
Rsq <- SSreg/SStot
This is the coefficient of determination, appearing as 
'Multiple R-squared' in the summary output. The value is 0.8192202.

6.

A forensic team analyzes a shoe print and a hand print, presumably left by the same person: The shoe print belongs to a size \(8\) women’s shoe and the index and pinky fingers measure \(70\)mm and \(60\)mm, respectively:

6.a

If the forensic team uses the models fitted above to make guesses about the height of the person who left the prints, will they be extrapolating beyond the range of the observed data? Explain your answer.

From careful study of the scatterplots 
produced in the first question, it does not 
appear that a person wearing a size 8 women's 
shoe and having index and pinky fingers 
measuring 70mm and 60mm, respectively, is an outlier 
beyond the range of observed data.
6.b

Give an interval such that the forensic team can be \(95\%\) certain it contains the average height of the population of all people wearing size \(8\) women’s shoes and having index and pinky fingers measuring \(70\)mm and \(60\)mm, respectively.

xnew <- data.frame(ind = 70, pnk = 60, shoe = 8, shmw = 'w')
ci <- predict(lm_all,newdata = xnew,int='conf')
The interval is (161.1451,165.1442).
6.c

Give an interval such that the forensic team can be \(95\%\) certain it contains the height of the person who left the prints.

pi <- predict(lm_all,newdata = xnew,int='pred')
The interval is (153.7426,172.5468).

7.

Suppose there is no shoe print, but only a hand print with index and pinky fingers measuring \(70\)mm and \(60\)mm, respectively:

7.a

Give an interval such that the forensic team can be \(95\%\) certain it contains the average height of the population of all people having index and pinky fingers measuring \(70\)mm and \(60\)mm, respectively.

xnew <- data.frame(ind = 70, pnk = 60)
ci <- predict(lm_fingers,newdata = xnew,int='conf')
The interval is (168.7724,173.9497).
7.b

Give an interval such that the forensic team can be \(95\%\) certain it contains the height of the person who left the hand print.

pi <- predict(lm_fingers,newdata = xnew,int='pred')
The interval is (154.6865,188.0356).

8.

Suppose there is no hand print, but only a shoe print belonging to a size 8 women’s shoe:

8.a

Give an interval such that the forensic team can be \(95\%\) certain it contains the average height of the population of all people wearing a size 8 women’s shoe.

xnew <- data.frame(shoe = 8, shmw = 'w')
ci <- predict(lm_shoes,newdata = xnew,int='conf')
The interval is (160.2693,163.6317).
8.b

Give an interval such that the forensic team can be \(95\%\) certain it contains the height of the person who left the shoe print.

pi <- predict(lm_shoes,newdata = xnew,int='pred')
The interval is (152.3183,171.5827).

9.

Answer the following based on careful study of the preceding model output and confidence and prediction intervals:

9.a

If a shoe print is found, does a hand print provide useful additional accuracy in guessing the height of the person leaving the prints?

The index and pinky finger lengths appear to 
contribute very little additional information if the 
shoe size and shoe size gender are known.  This is seen 
in two ways from the above output: The confidence and 
prediction intervals based on the complete information 
(with finger lengths) are scarcely narrower than those 
based only on shoe information.  Moreover, the model 
with all four covariates has a coefficient of 
determination scarcely higher than that of the model 
with only the shoe covariates. 
9.b

If a hand print is found, does a shoe print provide useful additional accuracy in guessing the height of the person leaving the prints?

The shoe print information does indeed allow 
the team to obtain a more accurate guess at the 
height of the person who left the prints; note 
how much narrower the confidence and prediction 
intervals became when the shoe size and gender 
information was included in the model.
9.c

If only a hand print is found, should the forensic team bother trying to use the index and pinky finger lengths to guess the height of the person who left it?

If no shoe information is available, the hand 
print information can still be useful.  Even 
though the confidence and prediction intervals 
are wide when only the hand print information is 
used, the intervals still allow the forensic team 
to narrow down the range of possible heights for 
the person leaving the print (the prediction 
interval does not cover the entire range of the 
observed data).