Several statistics students recorded their heights and the lengths of their index fingers. The measurements, in centimeters and millimeters, respectively, are read into R using the code below (copy and paste this into your .qmd file):
link <-url("https://gregorkb.github.io/data/height_length_index_pinky_shoesize_combined.csv")hl <-read.table(link,sep=",",header=T)head(hl)
ft in. ind pnk shoe shmw class
1 5 4 75 65 8.0 w STAT_516_sp_2026
2 5 4 70 56 7.0 w STAT_516_sp_2026
3 5 8 70 52 10.0 w STAT_516_sp_2026
4 5 5 68 62 7.5 w STAT_516_sp_2026
5 5 6 71 55 9.5 m STAT_516_sp_2026
6 6 4 78 65 13.0 m STAT_516_sp_2026
Y <- (hl$ft*12+ hl$in.)*2.54# heights in cmx <- hl$ind # index finger lengths in mmn <-length(Y)
It is of interest to use the simple linear regression model to predict the height of a person based on the length of his or her index finger.
1)
Make a scatterplot of the heights versus the index finger lengths with the least-squares line overlaid. Report the intercept and slope of the least-squares line as well as the value of Pearson’s correlation coefficient.
Make a normal quantile-quantile plot of the residuals as well as a residuals versus fitted values plot. Then carefully explain whether you believe the assumptions of the simple linear regression model are satisfied for these data.
plot(lm(Y~x),which =2)
plot(lm(Y~x),which =1)
The normal Q-Q plot indicates a small departure from normality; the
residuals versus fitted values plot shows fairly constant spread from
left to right, so the variance appears to be roughly constant. There is
perhaps some suggestion of a pattern in the points, which may indicate
non-linearity, but it is quite weak. It may be safe to assume that the
assumptions of the linear regression model are satisfied.
5)
Give a \(99\%\) confidence interval for the slope parameter \(\beta_1\).
99% confidence interval for slope: [0.7353599,1.578039]
6)
State whether you would reject \(H_0\): \(\beta_1 = 0\) versus \(H_1\): \(\beta_1 \neq 0\) at the \(\alpha = 0.05\) significance level.
We would reject H0: Note that a 95% confidence
interval for b1 would be narrower than the 99% CI, while being
centered at the same value. Since the 99% CI did not
contain 0, neither then will the 95% CI, implying that we would
reject H0 at the 0.05 significance level.
7)
Give an estimate of the mean height of persons with index finger length equal to \(72\) mm. In addition, give a 95% confidence interval for this mean height.
Estimated mean height of persons with index finger length
equal to 72 mm: 172.6428
95% confidence interval: [170.6545,174.6312]
8)
A hand print is found made by a hand with an index finger length of \(72\) mm. Give an interval which will contain with \(95\%\) probability the height of the person who made the print.
Do you think the above interval would be useful in identifying the person who made the print?
This interval contains most (91.30435%) of the heights of the
students in the class; if a similar percentage of
persons in the general population have heights in this interval,
then the interval does not narrow down very much the number
of persons who could possibly have made the hand print.
10)
Give the value of the coefficient of determination for these data. Interpret the value.
Coefficient of determination (R^2): 0.4415485
Index finger length is able to explain 44.15485 percent of
variation in height.
11)
Give the value of the test statistic \(\displaystyle T_{\operatorname{test}} = \frac{\hat \beta_1}{\hat \sigma / \sqrt{S_{xx}}}\) as well as the value of \(\displaystyle F_{\operatorname{test}} = \frac{\operatorname{MS}_{\operatorname{Reg}}}{\operatorname{MS}_{\operatorname{Error}}}\).
# T test statisticTtest <- b1 / (sigma_hat/sqrt(Sxx))# F test statisticSSerror <-sum((Y - Yhat)**2)MSerror <- SSerror/ (n-2)MSreg <- SSreg /1Ftest <- MSreg / MSerror # equal to Ttest**2 in SLR
Value of T test statistic: 7.278366
Value of F test statistic: 52.97462
12)
Give the p-value for testing \(H_0\): \(\beta = 0\) versus \(H_1\): \(\beta_1 \neq 0\) based on the value of the test statistic \(T_{\operatorname{test}}\). Interpret your answer.
pval <-2*(1-pt(abs(Ttest),n-2))
The p-value is the area to the right of
the test statistic under the pdf of the t distribution
with n-2 degrees of freedom. This is: 4.799214e-10.
There is strong evidence of a positive linear relationship
between length of index finger and height.
13)
Comment on whether there are any outliers in the data set. Show a plot to support your answer.
plot(lm(Y~x),which =4)
There are a few observations with somewhat large Cook's
distances, but they do not appear to be very extreme outliers.
14)
Convert the heights to inches (divide by \(2.54\)) and fit the regression model again. Report the coefficient of determination as well as the p-value for testing whether the slope coefficient is equal to zero. Compare these to the values you obtained previously.
Yin <- Y/2.54summary(lm(Yin ~ x))
Call:
lm(formula = Yin ~ x)
Residuals:
Min 1Q Median 3Q Max
-6.7020 -2.3358 -0.3265 2.4858 6.8520
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 35.18128 4.47928 7.854 4.41e-11 ***
x 0.45539 0.06257 7.278 4.80e-10 ***
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Residual standard error: 3.238 on 67 degrees of freedom
Multiple R-squared: 0.4415, Adjusted R-squared: 0.4332
F-statistic: 52.97 on 1 and 67 DF, p-value: 4.799e-10
The coefficient of determination (R^2)
and the p-value are no different. These do
not depend on the units in which the data
are recorded.