STA215 · TU past paper
Statistics II 2081 question paper
The complete TU 2081 exam paper for Statistics II (STA215), all 12 questions with solved model answers written to the mark scheme.
Tap a question to open its answer.
- 110 marksNumericalIntroduction of multiple linear regressionHideAnswer
Multiple Regression Analysis
Multiple Regression Analysis
Purpose of Multiple Regression Analysis
Multiple regression analysis is applied to:
- Establish a functional relationship between one dependent variable $Y$ and two or more independent variables $X_1, X_2, \ldots$
- Predict/estimate the value of the dependent variable from known independent variables
- Measure the separate effect of each independent variable while holding the others constant
- Identify which independent variables are significant contributors to the variation in $Y$
- Provide a more realistic model since real phenomena depend on several factors simultaneously
Given Data
Y X1 X2 64 6 2 70 6 1 85 9 0 50 5 8 60 6 2 72 7 1 75 8 5 55 6 9 80 8 1 70 6 1 $n = 10$
Part (i): Multiple Regression Equation
Model: $\hat{Y} = b_0 + b_1 X_1 + b_2 X_2$
Step 1: Summations
Y X1 X2 X1² X2² X1·X2 X1·Y X2·Y 64 6 2 36 4 12 384 128 70 6 1 36 1 6 420 70 85 9 0 81 0 0 765 0 50 5 8 25 64 40 250 400 60 6 2 36 4 12 360 120 72 7 1 49 1 7 504 72 75 8 5 64 25 40 600 375 55 6 9 36 81 54 330 495 80 8 1 64 1 8 640 80 70 6 1 36 1 6 420 70 681 67 30 463 182 185 4673 1810 $$\sum Y = 681,\ \sum X_1 = 67,\ \sum X_2 = 30,\ \sum X_1^2 = 463,$$ $$\sum X_2^2 = 182,\ \sum X_1 X_2 = 185,\ \sum X_1 Y = 4673,\ \sum X_2 Y = 1810$$
Step 2: Means
$$\bar{Y} = 68.1,\quad \bar{X}_1 = 6.7,\quad \bar{X}_2 = 3.0$$
Step 3: Normal Equations
$$681 = 10b_0 + 67b_1 + 30b_2 \quad (1)$$ $$4673 = 67b_0 + 463b_1 + 185b_2 \quad (2)$$ $$1810 = 30b_0 + 185b_1 + 182b_2 \quad (3)$$
Step 4: Solve
Eliminate $b_0$ using (1) and (2): multiply (1) by $6.7$: $$4562.7 = 67b_0 + 448.9b_1 + 201b_2 \quad (4)$$ $(2) - (4)$: $$110.3 = 14.1b_1 - 16b_2 \quad (5)$$
Eliminate $b_0$ using (1) and (3): multiply (1) by $3$: $$2043 = 30b_0 + 201b_1 + 90b_2 \quad (6)$$ $(3) - (6)$: $$-233 = -16b_1 + 92b_2 \quad (7)$$
From (5): $14.1b_1 - 16b_2 = 110.3$ From (7): $-16b_1 + 92b_2 = -233$
Solve. From (5): $b_1 = \dfrac{110.3 + 16b_2}{14.1}$.
Substitute in (7): $$-16\left(\frac{110.3 + 16b_2}{14.1}\right) + 92b_2 = -233$$
$$-16(110.3 + 16b_2) + 92(14.1)b_2 = -233(14.1)$$
$$-1764.8 - 256b_2 + 1297.2b_2 = -3285.3$$
$$1041.2b_2 = -1520.5$$
$$b_2 = -1.4603$$
Then: $$b_1 = \frac{110.3 + 16(-1.4603)}{14.1} = \frac{110.3 - 23.365}{14.1} = \frac{86.935}{14.1} = 6.1656$$
From (1): $$b_0 = \frac{681 - 67(6.1656) - 30(-1.4603)}{10}$$ $$= \frac{681 - 413.10 + 43.81}{10} = \frac{311.71}{10} = 31.171$$
Regression Equation
$$\boxed{\hat{Y} = 31.171 + 6.166,X_1 - 1.460,X_2}$$
Part (ii): Interpretation of $b_1$ and $b_2$
-
$b_1 = 6.166$: Holding the number of absent days ($X_2$) constant, a one-unit increase in the aptitude test score ($X_1$) raises the job-satisfaction score ($Y$) by about 6.17 units. The positive sign shows a direct relationship.
-
$b_2 = -1.460$: Holding the aptitude test score ($X_1$) constant, each additional day of absence ($X_2$) decreases the job-satisfaction score ($Y$) by about 1.46 units. The negative sign shows an inverse relationship.
Part (iii): Estimate for $X_1 = 9,\ X_2 = 6$
$$\hat{Y} = 31.171 + 6.166(9) - 1.460(6)$$ $$= 31.171 + 55.494 - 8.760$$ $$= \boxed{77.91}$$
The estimated job-satisfaction score is approximately 78.
- 210 marksNumericalCompletely Randomized DesignHideAnswer
One-Way ANOVA: Seminar Evaluation Scores by Manager Level
A management consulting company presents a 3-day seminar on project management to various clients. The seminar is basically the same each time it is given. However, sometimes it is presented to high-level managers, sometimes to mid-level managers, and sometimes to low-level managers. The seminar facilitators believe evaluations of the seminar may vary with the audience. Suppose the following data are some randomly selected evaluation scores from different levels of managers after they have attended the seminar. The ratings are on a scale from 1 to 100, with 100 being the highest score. Use a one-way ANOVA to determine whether there is a significant difference in the evaluation according to manager level. The following table gives the scores to the various clients due to different manager levels.
$$\begin{array}{|c|c|c|c|} \hline & \text{High Level} & \text{Mid Level} & \text{Low Level} \ \hline & 85 & 90 & 55 \ & 75 & 100 & 75 \ & 85 & 95 & 80 \ & 60 & 85 & 75 \ & 70 & 90 & \ & & 75 & \ \hline \end{array}$$
[10]
High Level Mid Level Low Level :-::-::-: 85 90 55 75 100 75 85 95 80 60 85 75 70 90 75 - $n1 = 5,\ n2 = 6,\ n3 = 4$ - Total $N = 15$, groups $k = 3$ - $\alpha = 0.05$ - $H0: \mu1 = \mu2 = \mu3$ (no significant difference) - $H1:$ at leas...
- 310 marksNumericalpaired sample t-testHideAnswer
T-Test for Difference Between Two Sample Means
Paired measurements of gripping strength for 8 left-handed writers: Person 1 2 3 4 5 6 7 8 -------------------------------- Left (L) 112 131 142 90 125 130 95 90 Right (R) 104 136 135 86 132 120 86 85 - $n = 8$ - $\alpha = 0.05$ - Claim:...
- 45 marksLatin Square DesignHideAnswer
What do you understand by design of experiments? Give the layout of a Latin Square Design. Explain why the number of treatments tested in a Latin Square Design should not be less than 3? [5]
Design of Experiments and Latin Square Design
Design of Experiments
Design of experiments refers to the systematic planning and arrangement of experimental conditions so that valid, unbiased, and reliable conclusions can be drawn from the experimental data with minimum experimental error.
According to R.A. Fisher, the design of experiments is based on three fundamental principles:
Three Principles of Design of Experiments
Principle Description Replication Execution of an experiment more than once. It provides a more reliable estimate and helps minimize experimental error (maximize accuracy). Randomization Treatments are allocated to various plots in a random manner, giving each treatment an equal chance of being assigned to any experimental unit. It reduces bias. Local Control Converts a heterogeneous experimental field into homogeneous blocks (row-wise or column-wise), thereby reducing experimental error.
Latin Square Design (LSD)
LSD is a design used for non-homogeneous experimental material. It applies local control simultaneously along both rows and columns, making it more efficient than RBD. The experimental layout is always square (equal number of rows and columns).
Mathematical Model
$$Y_{ijk} = \mu + \alpha_i + \beta_j + T_k + e_{ijk}$$
Where:
- $Y_{ijk}$ = observation in $i^{th}$ row and $j^{th}$ column receiving $k^{th}$ treatment
- $\mu$ = general mean effect
- $\alpha_i$ = effect due to $i^{th}$ row
- $\beta_j$ = effect due to $j^{th}$ column
- $T_k$ = effect due to $k^{th}$ treatment
- $e_{ijk}$ = random error due to chance
Layout of a Latin Square Design (4x4 Example)
In an LSD with 4 treatments (A, B, C, D), each treatment appears exactly once in each row and exactly once in each column:
$$ \begin{array}{|c|c|c|c|} \hline \textbf{Col 1} & \textbf{Col 2} & \textbf{Col 3} & \textbf{Col 4} \ \hline A & B & C & D \ \hline B & C & D & A \ \hline C & D & A & B \ \hline D & A & B & C \ \hline \end{array} $$
Here, 4 rows, 4 columns, and 4 treatments -- each appearing exactly once per row and per column.
Why the Number of Treatments Should NOT Be Less Than 3
In an LSD with $m$ treatments, the degrees of freedom for error is given by:
$$df_{error} = (m-1)(m-2)$$
If $m = 2$ (only 2 treatments):
$$df_{error} = (2-1)(2-2) = 1 \times 0 = 0$$
- With zero degrees of freedom for error, it is impossible to estimate the experimental error.
- Without an error estimate, no F-test can be performed, and no valid statistical conclusions can be drawn.
If $m = 3$ (minimum acceptable):
$$df_{error} = (3-1)(3-2) = 2 \times 1 = 2$$
- This gives at least 2 degrees of freedom for error, which allows estimation of experimental error and valid hypothesis testing.
Summary Table
Treatments ($m$) $df_{error} = (m-1)(m-2)$ Valid? 2 0 No -- error cannot be estimated 3 2 Yes -- minimum acceptable 4 6 Yes 5 12 Yes Conclusion: The number of treatments in an LSD must be at least 3 so that the degrees of freedom for error is greater than zero, enabling estimation of experimental error and valid F-tests in the ANOVA table.
- 55 marksNumericalTwo independent sample testHideAnswer
Chi-Square Test of Independence
Based on interviews of couples seeking divorce, a social worker compiles the following data related to the period of acquaintanceship before marriage and the duration of marriage. Perform a test to determine if the data substantiate an association between the duration of a marriage and the acquaintanceship prior to marriage. Use 5% level of significance.
Acquaintanceship before marriage Less than or equal to 5 years More than 5 years Total Below 0.5 year 15 7 22 0.5–1.5 years 26 22 48 Over 1.5 years 19 11 30 Total 60 40 100 [5]
Chi-Square Test for Association
Step 1: Given Data
Observed contingency table:
Acquaintanceship ≤ 5 years > 5 years Total Below 0.5 yr 15 7 22 0.5-1.5 yr 26 22 48 Over 1.5 yr 19 11 30 Total 60 40 100 Significance level $\alpha = 0.05$.
Step 2: Hypotheses
- $H_0$: No association between duration of marriage and acquaintanceship period (independent).
- $H_1$: There is an association.
Step 3: Expected Frequencies
$$E_{ij} = \frac{R_i \times C_j}{N}$$
Cell $E$ (Below 0.5, ≤5) $\frac{22 \times 60}{100} = 13.2$ (Below 0.5, >5) $\frac{22 \times 40}{100} = 8.8$ (0.5-1.5, ≤5) $\frac{48 \times 60}{100} = 28.8$ (0.5-1.5, >5) $\frac{48 \times 40}{100} = 19.2$ (Over 1.5, ≤5) $\frac{30 \times 60}{100} = 18.0$ (Over 1.5, >5) $\frac{30 \times 40}{100} = 12.0$ Step 4: Chi-Square Statistic
$$\chi^2 = \sum \frac{(O-E)^2}{E}$$
O E $(O-E)^2$ $(O-E)^2/E$ 15 13.2 3.24 0.24545 7 8.8 3.24 0.36818 26 28.8 7.84 0.27222 22 19.2 7.84 0.40833 19 18.0 1.00 0.05556 11 12.0 1.00 0.08333 $$\chi^2_{cal} = 0.24545 + 0.36818 + 0.27222 + 0.40833 + 0.05556 + 0.08333 = 1.4331$$
Step 5: Degrees of Freedom and Critical Value
$$df = (r-1)(c-1) = (3-1)(2-1) = 2$$
$$\chi^2_{0.05,,2} = 5.991$$
Step 6: Decision
$$\chi^2_{cal} = 1.433 < \chi^2_{tab} = 5.991$$
Fail to reject $H_0$.
Step 7: Conclusion
At the 5% level of significance, there is no significant association between the duration of marriage and the period of acquaintanceship before marriage. The data do not substantiate an association.
- 65 marksNumericalEstimationHideAnswer
Define confidence level in estimation. A quality control inspector collected a random sample of 400 tubes of toothpaste from the production line and found that 20 of the tubes had leaks from the tail end. Construct 96% confidence interval for the percentage of all the toothpaste tubes that had leakage and interpret the result. [5]
Confidence Level in Estimation and Confidence Interval for a Proportion
Definition: Confidence Level
The confidence level $(1-\alpha)$ is the probability that the interval estimate constructed from sample data will contain the true population parameter. A 96% confidence level means that if repeated random samples were taken and an interval built from each, about 96% of those intervals would enclose the true population proportion.
Given Data
Item Value Sample size $n = 400$ Leaking tubes $X = 20$ Confidence level $1-\alpha = 0.96$ Significance level $\alpha = 0.04$
Step 1: Sample Proportion
$$p = \frac{X}{n} = \frac{20}{400} = 0.05, \qquad q = 1 - p = 0.95$$
Step 2: Critical Value
$$\frac{\alpha}{2} = 0.02 \Rightarrow Z_{\alpha/2} = 2.054 \approx 2.05$$
Step 3: Standard Error
$$S.E.(p) = \sqrt{\frac{pq}{n}} = \sqrt{\frac{0.05 \times 0.95}{400}} = \sqrt{\frac{0.0475}{400}} = \sqrt{0.00011875} = 0.010897$$
Step 4: Confidence Interval
$$p \pm Z_{\alpha/2}\cdot S.E.(p)$$
Margin of error: $$E = 2.054 \times 0.010897 = 0.02238$$
Lower limit: $$0.05 - 0.02238 = 0.02762$$
Upper limit: $$0.05 + 0.02238 = 0.07238$$
Step 5: Express as Percentage
$$\boxed{2.76% \leq P \leq 7.24%}$$
(Using $Z = 2.05$: $2.77%$ to $7.24%$; essentially the same.)
Interpretation
We are 96% confident that the true percentage of all toothpaste tubes with leakage from the tail end lies between approximately 2.76% and 7.24%. If the sampling procedure were repeated many times, about 96% of the resulting intervals would contain the true population proportion of leaking tubes.
- 75 marksNumericalone sample tests for mean of normal populaHideAnswer
The mean drying time of a brand of spray paint is known to be 122 seconds. The research division of the company that produces this paint contemplates that adding a new chemical ingredient to the paint accelerate the drying process. To investigate this conjecture, the paint with the chemical additions is sprayed on 50 surfaces and the drying time is recorded. The mean and standard deviation of drying time computed from these recorded are found as 116 seconds and 16.8 seconds respectively. Does these data provide strong evidence that the mean drying time is reduced by the addition of the new chemical? Use 5% level of significance. Also find p-value. [5]
Hypothesis Test: Mean Drying Time of Spray Paint
Given Data
- Population (claimed) mean: $\mu_0 = 122$ seconds
- Sample size: $n = 50$
- Sample mean: $\bar{X} = 116$ seconds
- Sample standard deviation: $s = 16.8$ seconds
- Significance level: $\alpha = 0.05$
Step 1: Hypotheses
- $H_0: \mu = 122$ (no reduction)
- $H_1: \mu < 122$ (drying time reduced) - left-tailed test
Step 2: Test Statistic
Since $n = 50 > 30$, use the Z-test: $$Z = \frac{\bar{X} - \mu_0}{s/\sqrt{n}}$$
Step 3: Calculation
$$s/\sqrt{n} = \frac{16.8}{\sqrt{50}} = \frac{16.8}{7.0711} = 2.3759$$
$$Z = \frac{116 - 122}{2.3759} = \frac{-6}{2.3759} = -2.525$$
Step 4: Critical Value
Left-tailed at $\alpha = 0.05$: $Z_{tab} = -1.645$
Step 5: Decision
Since $Z_{cal} = -2.525 < -1.645$, we reject $H_0$.
Step 6: P-value
$$p = P(Z < -2.525) = 1 - P(Z < 2.525) = 1 - 0.9942 = 0.0058$$
Since $p = 0.0058 < \alpha = 0.05$, reject $H_0$.
Conclusion
The data provide strong evidence at the 5% level of significance that the new chemical additive reduces the mean drying time below 122 seconds. The small p-value ($\approx 0.0058$) confirms high significance.
- 85 marksNumericalMultiple and partial correlationHideAnswer
Explain the concept of multiple and partial correlation coefficients. Consider three variables X1, X2 and X3. If $r_{12}=0.40$, $r_{23}=0.50$ and $r_{13}=0.6$ find $R_{123}$ and $r_{23.1}$. [5]
Three variables $X1, X2, X3$ with zero-order correlation coefficients: $$r{12} = 0.40, \quad r{23} = 0.50, \quad r{13} = 0.60$$ Note on notation: The problem statement contains a typo. "$R{123}$" is standardly interpreted as the multiple...
- 95 marksNumericaltest for single proportionHideAnswer
Discuss the concept of level of significance in hypothesis testing. A manufacturer of laptop provides a particular model in one of three colors. Of the first 100 laptops sold, it is noted that 80 were the first color. Can you conclude that more than two third of all the customers have a preference for the first color? Use 5% level of significance. [5]
Level of Significance and Hypothesis Testing
Concept of Level of Significance
The level of significance (denoted α) is the probability of rejecting the null hypothesis $H_0$ when it is actually true. It is the maximum acceptable probability of committing a Type I error.
- Fixed in advance of the test (typically $1%$, $5%$, or $10%$).
- At $5%$ significance, we accept a $5%$ chance of wrongly rejecting a true $H_0$.
- It determines the critical (tabulated) value used to bound the acceptance/rejection regions.
- A smaller $\alpha$ makes the test more conservative (harder to reject $H_0$).
Numerical Solution
Given Data
Parameter Value Sample size $n$ $100$ Number preferring first color $x$ $80$ Hypothesized proportion $P$ $2/3 \approx 0.6667$ Level of significance $\alpha$ $0.05$ Sample proportion: $$\hat{p} = \frac{x}{n} = \frac{80}{100} = 0.80$$
Step 1: Hypotheses
- $H_0: P = \dfrac{2}{3}$ (two-thirds prefer the first color)
- $H_1: P > \dfrac{2}{3}$ (more than two-thirds prefer the first color)
This is a right-tailed test.
Step 2: Test Statistic
$$Z = \frac{\hat{p} - P}{\sqrt{\dfrac{P(1-P)}{n}}}$$
Numerator: $$\hat{p} - P = 0.80 - 0.6667 = 0.1333$$
Denominator: $$\sqrt{\frac{0.6667 \times 0.3333}{100}} = \sqrt{\frac{0.2222}{100}} = \sqrt{0.002222} = 0.04714$$
Therefore: $$Z = \frac{0.1333}{0.04714} = 2.83$$
Step 3: Critical Value
For a right-tailed test at $\alpha = 0.05$: $$Z_{tab} = 1.645$$
Step 4: Decision
$$Z_{cal} = 2.83 > Z_{tab} = 1.645$$
Since the calculated value falls in the rejection region, reject $H_0$.
Step 5: Conclusion
At the $5%$ level of significance, there is sufficient evidence to conclude that more than two-thirds of all customers prefer the first color of the laptop.
- 105 marksNumericalTest of significance of regressionHideAnswer
Multiple Regression Analysis Question
A multiple regression equation yields the following results:
Source Sum of square Degree of freedom Regression 740 2 Error 510 17 i) What is the total sample size?
ii) How many independent variables are being considered?
iii) Compute the coefficient of determination and interpret its value.
iv) Compute the standard error of estimate.
v) Test the hypothesis that the overall fit of the model is significant or not. Assume $\alpha=0.05$.
[5]
Source Sum of Squares (SS) Degrees of Freedom (df) --------------------------------------------------- Regression 740 2 Error 510 17 Total 1250 19 $\alpha = 0.05$ --- $$df{Total} = df{Reg} + df{Error} = 2 + 17 = 19$$ Since
- 115 marksNumericalM/M/1 systemHideAnswer
Explain briefly the queuing theory. Customers arrive at a one-man barber shop according to Poisson process with mean inter arrival time of 12 minutes. Customer spends an average of 10 minutes in the barber’s chair. What is the expected number of customers in the barber shop in the queue? [5]
Queuing Theory and Barber Shop Problem
Part 1: Brief Explanation of Queuing Theory
Queuing theory is the mathematical study of waiting lines. It models systems in which customers (or jobs) arrive, possibly wait in a queue, receive service from one or more servers, and then leave. Its aim is to predict system performance measures such as average waiting time, queue length, and server utilization so resources can be optimally designed.
Characteristics of a queuing system (Kendall notation A/B/c):
Element Description Arrival process Typically Poisson arrivals with rate $\lambda$ Service process Typically exponential service with rate $\mu$ Number of servers $c$ (here $c = 1$) Queue discipline Usually FCFS System capacity / population Usually infinite Utilization (traffic intensity): $\rho = \dfrac{\lambda}{c\mu}$; stability requires $\rho < 1$.
Part 2: Barber Shop Numerical Solution
Given data
- Mean inter-arrival time $= 12$ minutes
- Mean service time (in barber's chair) $= 10$ minutes
- One server (one-man shop) → M/M/1 model
Step 1: Rates
$$\lambda = \frac{1}{12}\text{ customers/min}, \qquad \mu = \frac{1}{10}\text{ customers/min}$$
Step 2: Traffic intensity
$$\rho = \frac{\lambda}{\mu} = \frac{1/12}{1/10} = \frac{10}{12} = \frac{5}{6} \approx 0.833 \quad (<1,\text{ stable})$$
Step 3: Expected number in queue $L_q$
$$L_q = \frac{\rho^2}{1-\rho} = \frac{(5/6)^2}{1 - 5/6} = \frac{25/36}{1/6} = \frac{25}{6} \approx 4.17 \text{ customers}$$
(Supporting) Expected number in system $L_s$
$$L_s = \frac{\rho}{1-\rho} = \frac{5/6}{1/6} = 5 \text{ customers}$$
Final Answer
$$\boxed{L_q = \frac{25}{6} \approx 4.17 \text{ customers waiting in the queue}}$$
The question specifically asks for customers in the queue, so $L_q \approx 4.17$ is the required answer ($L_s = 5$ is the number in the whole shop).
- 125 marksCentral Limit TheoremHideAnswer
Write short note on: a) Central limit theorem. b) Determination of required sample size to estimate population proportion. [5]
--- The Central Limit Theorem states that if a sufficiently large random sample of size n is drawn from any population (regardless of its shape) with population mean μ and population standard deviation σ, then the sampling distribution o...