STA215 · TU past paper
Statistics II 2078 question paper
The complete TU 2078 exam paper for Statistics II (STA215), all 12 questions with solved model answers written to the mark scheme.
Tap a question to open its answer.
- 110 marksNumericalCompletely Randomized DesignHideAnswer
Statistical Test for Computer Lifetime Across Brands
There are three brands of computers namely Dell, Lenovo, and HP. The following are the lifetime of 15 computers in years. Apply appropriate statistical tests to identify whether the average lifetime (in years) is significantly different across three brands of computers at a 5% level of significance. You can again tabulate the data initially in the required format for statistical analysis.
$$\begin{array}{|c|c|c|} \hline \text{Serial Number} & \text{Computer Brand} & \text{Lifetime in years} \ \hline 1 & \text{Dell} & 15 \ 2 & \text{Lenovo} & 10 \ 3 & \text{HP} & 9 \ 4 & \text{Dell} & 12 \ 5 & \text{Lenovo} & 6 \ 6 & \text{HP} & 7 \ 7 & \text{Dell} & 4 \ 8 & \text{Lenovo} & 8 \ 9 & \text{HP} & 13 \ 10 & \text{Dell} & 11 \ 11 & \text{HP} & 5 \ 12 & \text{Lenovo} & 7 \ 13 & \text{Dell} & 3 \ 14 & \text{HP} & 5 \ 15 & \text{Lenovo} & 4 \ \hline \end{array}$$
[10]
Dell Lenovo HP ------------------ 15 10 9 12 6 7 4 8 13 11 7 5 3 4 5 - $n1 = n2 = n3 = 5$, total $N = 15$, groups $k = 3$ - Significance level $\alpha = 0.05$ Comparing means of three independent groups on a continuous variable → One-Way...
- 210 markssampling distribution of mean and proportiHideAnswer
Explain the sample distribution of mean with reference to some numerical example. Illustrate the practical implications of the Central Limit Theorem (CLT) in inferential statistics.[10]
When all possible random samples of size n are drawn from a population of size N, and the mean of each sample is computed, the probability distribution formed by these sample means is called the sampling distribution of the sample mean. ...
- 310 marksNumericalIntroduction of multiple linear regressionHideAnswer
Multiple Linear Regression Analysis for Computer Vision Syndrome
ID CVS (Y) Age X₁ Work hrs X₂ ---------------------------------- 001 6 24 4 007 7 26 5 125 5 30 6 231 11 41 8 99 3 47 3 299 29.0 50 6 145 28 52 7 $n = 7$. CVS measured on scale 0 to 50. Dependent variable: Scale of CVS ($Y$) - it is the ...
- 45 marksNumericaltest for difference between two means and HideAnswer
The following are the details of working hours in the classroom per week of male and female faculty working in the area of Computer Science and Information Technology at Tribhuvan University. Apply independent t-test to examine the average working hour in the classroom per week is significantly different between male and female faculty, at 1% level of significance. State also null and alternative hypotheses appropriately.
Male Faculty Female Faculty Sample Size 60 30 Average working hours per week 12 9 The standard deviation of a working hour per week 4 3 [5]
Male Faculty Female Faculty --------- Sample size $n1 = 60$ $n2 = 30$ Mean working hours $\bar{X}1 = 12$ $\bar{X}2 = 9$ Standard deviation $S1 = 4$ $S2 = 3$ Level of significance: $\alpha = 0.01$ (two-tailed) --- Null Hypothesis
- 55 marksNumericalEstimationHideAnswer
A survey was conducted among 70 students studying B.Sc. CSIT in some colleges randomly. Among them, 50 students secured more than 80% marks in statistics. Compute 99% and 95% confidence intervals for the population proportion of students who secured more than 80% marks in subject statistics, and comment on the results. [5]
Confidence Interval for Population Proportion
STEP 1 - EXTRACT (Given Data)
Parameter Value Sample size $n$ 70 Students scoring > 80% ($X$) 50 Sample proportion $p = X/n$ $50/70 = 0.7143$ $q = 1 - p$ $0.2857$ $Z$ for 99% $2.576$ $Z$ for 95% $1.96$ All required data present.
STEP 2 - SOLVE
Standard Error
$$S.E.(p) = \sqrt{\frac{pq}{n}} = \sqrt{\frac{0.7143 \times 0.2857}{70}}$$
$$= \sqrt{\frac{0.2041}{70}} = \sqrt{0.002916} = 0.0540$$
99% Confidence Interval
$$CI = p \pm Z \cdot S.E. = 0.7143 \pm 2.576 \times 0.0540$$
Margin $= 2.576 \times 0.0540 = 0.1391$
$$0.7143 - 0.1391 = 0.5752, \quad 0.7143 + 0.1391 = 0.8534$$
$$\boxed{0.5752 \leq P \leq 0.8534}$$
95% Confidence Interval
$$CI = 0.7143 \pm 1.96 \times 0.0540$$
Margin $= 1.96 \times 0.0540 = 0.1058$
$$0.7143 - 0.1058 = 0.6085, \quad 0.7143 + 0.1058 = 0.8201$$
$$\boxed{0.6085 \leq P \leq 0.8201}$$
Summary Table
Confidence Level Z Margin Lower Upper 99% 2.576 0.1391 0.5752 0.8534 95% 1.96 0.1058 0.6085 0.8201 Comment
- We are 99% confident the true proportion lies between 57.52% and 85.34%.
- We are 95% confident the true proportion lies between 60.85% and 82.01%.
- The 99% interval is wider than the 95% interval, since higher confidence requires a larger margin. Thus higher confidence gives lower precision.
- 65 marksNumericaltest for difference between two means and HideAnswer
In location 1, there are 250 corona-positive cases out of 460 persons, and in location 2, 250 positive cases were reported out of 650 persons. Can it be concluded that the proportion of corona-positive cases is higher in location 1 compared to location 2? Test at a 10% level of significance. [5]
Location 1 Location 2 --------- Sample size $n1 = 460$ $n2 = 650$ Positive cases $x1 = 250$ $x2 = 250$ Level of significance: $\alpha = 0.10$ --- $$\hat{p}1 = \frac{250}{460} = 0.5435$$ $$\hat{p}2 = \frac{250}{650} = 0.3846$$ --- -
- 75 marksNumericalone sample tests for mean of normal populaHideAnswer
Hypothesis Testing for Average Age of CSIT Students
Step 1: Set up Hypotheses
- Null Hypothesis (H₀): $\mu = 22$ (The average age of enrolling students is 22 years)
- Alternative Hypothesis (H₁): $\mu < 22$ (The average age is less than 22 years)
This is a left-tailed test at 5% level of significance.
Step 2: Sample Data
| 20 | 19 | 22 | 23 | 19 | 20 | 20 | 21 | 22 | 20 | 19 | 20 |
Sample size: $n = 12$
Step 3: Calculate Sample Statistics
Sample mean: $$\bar{x} = \frac{20 + 19 + 22 + 23 + 19 + 20 + 20 + 21 + 22 + 20 + 19 + 20}{12} = \frac{245}{12} = 20.417$$
Sample standard deviation: $$s = \sqrt{\frac{\sum(x_i - \bar{x})^2}{n-1}} = \sqrt{\frac{26.917}{11}} = \sqrt{2.447} = 1.564$$
Step 4: Test Statistic
Since the population is normally distributed and population standard deviation is unknown, we use the t-test:
$$t = \frac{\bar{x} - \mu_0}{s/\sqrt{n}} = \frac{20.417 - 22}{1.564/\sqrt{12}} = \frac{-1.583}{0.451} = -3.511$$
Step 5: Critical Value
For a left-tailed test with $\alpha = 0.05$ and $df = n - 1 = 11$: $$t_{0.05, 11} = -1.796$$
Step 6: Decision
Since $t = -3.511 < -1.796$, we reject the null hypothesis.
Conclusion: At the 5% level of significance, there is sufficient evidence to support the researcher's doubt that the average age of CSIT enrolling students is less than 22 years. [5]
- Claimed population mean: $\mu0 = 22$ years - Sample data: 20, 19, 22, 23, 19, 20, 20, 21, 22, 20, 19, 20 - Sample size: $n = 12$ - Level of significance: $\alpha = 0.05$ - Population normal, population SD unknown → one-sample t-test - ...
- 85 marksNumericalTwo independent sample testHideAnswer
Apply the Mann-Whitney U test for examining the following knowledge score on IT among two groups of IT workers at a 5% level of significance.
Group A: 5 8 2 7 6 Group B: 9 12 4 6 [5]
Group A (n₁ = 5) 5 8 2 7 6 ------------------ Group B (n₂ = 4) 9 12 4 6 - H₀: No significant difference in knowledge scores between the two groups. - H₁: There is a significant difference between the two groups. - Significance level: α =...
- 95 marksNumericalTwo independent sample testHideAnswer
Chi-Square Test for Association Between Email Account Type and Hacking Status
A survey was conducted to see the association between hacking status of the email and the type of email account. The survey has reported the following cross tabulation. Do the information provide sufficient evidence to conclude that the type email account and the hacking status is associated? Use Chi-square test at 1% level of significance.
$$\begin{array}{|c|c|c|} \hline \text{Type of e-mail account} & \text{Hacking status Yes} & \text{Hacking status No} \ \hline \text{Yahoo} & 60 & 15 \ \text{Gmail} & 20 & 120 \ \hline \end{array}$$
[5]
Chi-Square Test for Association Between Email Account Type and Hacking Status
Step 1 - Given Data
Email Account Yes (Hacked) No (Not Hacked) Row Total Yahoo 60 15 75 Gmail 20 120 140 Column Total 80 135 215 Level of significance: $\alpha = 0.01$
Step 2 - Hypotheses
- $H_0$: Email account type and hacking status are independent (not associated)
- $H_1$: Email account type and hacking status are associated
Step 3 - Expected Frequencies
$$E_{ij} = \frac{(\text{Row Total}_i)(\text{Column Total}_j)}{\text{Grand Total}}$$
$$E_{11} = \frac{75 \times 80}{215} = 27.907$$ $$E_{12} = \frac{75 \times 135}{215} = 47.093$$ $$E_{21} = \frac{140 \times 80}{215} = 52.093$$ $$E_{22} = \frac{140 \times 135}{215} = 87.907$$
Step 4 - Test Statistic
$$\chi^2 = \sum \frac{(O - E)^2}{E}$$
O E O - E (O-E)² (O-E)²/E 60 27.907 32.093 1029.96 36.907 15 47.093 -32.093 1029.96 21.870 20 52.093 -32.093 1029.96 19.771 120 87.907 32.093 1029.96 11.717 $$\chi^2_{cal} = 36.907 + 21.870 + 19.771 + 11.717 = 90.27$$
(Small rounding differences; approximately $90.2$-$90.3$.)
Step 5 - Degrees of Freedom
$$df = (r-1)(c-1) = (2-1)(2-1) = 1$$
Step 6 - Critical Value
$$\chi^2_{0.01,,1} = 6.635$$
Step 7 - Decision
Since $\chi^2_{cal} = 90.27 > \chi^2_{tab} = 6.635$, we reject $H_0$.
Conclusion
At the 1% level of significance, there is sufficient evidence to conclude that the type of email account and the hacking status are associated. Yahoo accounts show a markedly higher hacking rate (60/75 = 80%) than Gmail accounts (20/140 ≈ 14%).
Note (optional): For a $2\times2$ table, Yates' continuity correction may be applied: $$\chi^2_{Yates} = \sum \frac{(|O-E| - 0.5)^2}{E} \approx 87.5$$ Either way, the conclusion is unchanged (reject $H_0$).
- 105 marksLatin Square DesignHideAnswer
State the mathematical model for Statistical analysis for m x m LSD for one observation per experimental unit. Also prepare a dummy ANOVA table for this. [5]
For a m × m Latin Square Design (LSD) with one observation per experimental unit, the linear statistical model is: $$X{ijk} = \mu + \alphai + \betaj + \gammak + \varepsilon{ijk}$$ where: Symbol Meaning ----------------- $X{ijk}$ Observat...
- 115 marksMarkov ProcessHideAnswer
Define the Markov chain and introduce its basic notations. Also, explain the characteristics of a Markov chain. [5]
A Markov chain is a stochastic process {X(t), t = 0, 1, 2, ...} that satisfies the Markov property (memoryless property), which states that the future state of the process depends only on the present state and not on the past states. Mat...
- 125 marksNeeds of applying non-parametric testsHideAnswer
Write short notes on the following: I. The rationale of using the non-parametric statistical test II. Estimation of minimum size for the given proportion [5]
--- Non-parametric statistical tests (also called distribution-free tests) are statistical methods that do not require assumptions about the underlying population distribution (e.g., normality). 1. No assumption about population distribu...