2078

STA215 · TU past paper

Statistics II 2078 question paper

The complete TU 2078 exam paper for Statistics II (STA215), all 12 questions with solved model answers written to the mark scheme.

Tap a question to open its answer.

  1. 110 marksNumericalCompletely Randomized DesignAnswer

    Statistical Test for Computer Lifetime Across Brands

    There are three brands of computers namely Dell, Lenovo, and HP. The following are the lifetime of 15 computers in years. Apply appropriate statistical tests to identify whether the average lifetime (in years) is significantly different across three brands of computers at a 5% level of significance. You can again tabulate the data initially in the required format for statistical analysis.

    $$\begin{array}{|c|c|c|} \hline \text{Serial Number} & \text{Computer Brand} & \text{Lifetime in years} \ \hline 1 & \text{Dell} & 15 \ 2 & \text{Lenovo} & 10 \ 3 & \text{HP} & 9 \ 4 & \text{Dell} & 12 \ 5 & \text{Lenovo} & 6 \ 6 & \text{HP} & 7 \ 7 & \text{Dell} & 4 \ 8 & \text{Lenovo} & 8 \ 9 & \text{HP} & 13 \ 10 & \text{Dell} & 11 \ 11 & \text{HP} & 5 \ 12 & \text{Lenovo} & 7 \ 13 & \text{Dell} & 3 \ 14 & \text{HP} & 5 \ 15 & \text{Lenovo} & 4 \ \hline \end{array}$$

    [10]

    Dell Lenovo HP ------------------ 15 10 9 12 6 7 4 8 13 11 7 5 3 4 5 - $n1 = n2 = n3 = 5$, total $N = 15$, groups $k = 3$ - Significance level $\alpha = 0.05$ Comparing means of three independent groups on a continuous variable → One-Way...

  2. 210 markssampling distribution of mean and proportiAnswer

    Explain the sample distribution of mean with reference to some numerical example. Illustrate the practical implications of the Central Limit Theorem (CLT) in inferential statistics.[10]

    When all possible random samples of size n are drawn from a population of size N, and the mean of each sample is computed, the probability distribution formed by these sample means is called the sampling distribution of the sample mean. ...

  3. 310 marksNumericalIntroduction of multiple linear regressionAnswer

    Multiple Linear Regression Analysis for Computer Vision Syndrome

    ID CVS (Y) Age X₁ Work hrs X₂ ---------------------------------- 001 6 24 4 007 7 26 5 125 5 30 6 231 11 41 8 99 3 47 3 299 29.0 50 6 145 28 52 7 $n = 7$. CVS measured on scale 0 to 50. Dependent variable: Scale of CVS ($Y$) - it is the ...

  4. 45 marksNumericaltest for difference between two means and Answer

    The following are the details of working hours in the classroom per week of male and female faculty working in the area of Computer Science and Information Technology at Tribhuvan University. Apply independent t-test to examine the average working hour in the classroom per week is significantly different between male and female faculty, at 1% level of significance. State also null and alternative hypotheses appropriately.

    Male FacultyFemale Faculty
    Sample Size6030
    Average working hours per week129
    The standard deviation of a working hour per week43

    [5]

    Male Faculty Female Faculty --------- Sample size $n1 = 60$ $n2 = 30$ Mean working hours $\bar{X}1 = 12$ $\bar{X}2 = 9$ Standard deviation $S1 = 4$ $S2 = 3$ Level of significance: $\alpha = 0.01$ (two-tailed) --- Null Hypothesis

  5. 55 marksNumericalEstimationAnswer

    A survey was conducted among 70 students studying B.Sc. CSIT in some colleges randomly. Among them, 50 students secured more than 80% marks in statistics. Compute 99% and 95% confidence intervals for the population proportion of students who secured more than 80% marks in subject statistics, and comment on the results. [5]

    Confidence Interval for Population Proportion

    STEP 1 - EXTRACT (Given Data)

    ParameterValue
    Sample size $n$70
    Students scoring > 80% ($X$)50
    Sample proportion $p = X/n$$50/70 = 0.7143$
    $q = 1 - p$$0.2857$
    $Z$ for 99%$2.576$
    $Z$ for 95%$1.96$

    All required data present.

    STEP 2 - SOLVE

    Standard Error

    $$S.E.(p) = \sqrt{\frac{pq}{n}} = \sqrt{\frac{0.7143 \times 0.2857}{70}}$$

    $$= \sqrt{\frac{0.2041}{70}} = \sqrt{0.002916} = 0.0540$$

    99% Confidence Interval

    $$CI = p \pm Z \cdot S.E. = 0.7143 \pm 2.576 \times 0.0540$$

    Margin $= 2.576 \times 0.0540 = 0.1391$

    $$0.7143 - 0.1391 = 0.5752, \quad 0.7143 + 0.1391 = 0.8534$$

    $$\boxed{0.5752 \leq P \leq 0.8534}$$

    95% Confidence Interval

    $$CI = 0.7143 \pm 1.96 \times 0.0540$$

    Margin $= 1.96 \times 0.0540 = 0.1058$

    $$0.7143 - 0.1058 = 0.6085, \quad 0.7143 + 0.1058 = 0.8201$$

    $$\boxed{0.6085 \leq P \leq 0.8201}$$

    Summary Table

    Confidence LevelZMarginLowerUpper
    99%2.5760.13910.57520.8534
    95%1.960.10580.60850.8201

    Comment

    • We are 99% confident the true proportion lies between 57.52% and 85.34%.
    • We are 95% confident the true proportion lies between 60.85% and 82.01%.
    • The 99% interval is wider than the 95% interval, since higher confidence requires a larger margin. Thus higher confidence gives lower precision.
  6. 65 marksNumericaltest for difference between two means and Answer

    In location 1, there are 250 corona-positive cases out of 460 persons, and in location 2, 250 positive cases were reported out of 650 persons. Can it be concluded that the proportion of corona-positive cases is higher in location 1 compared to location 2? Test at a 10% level of significance. [5]

    Location 1 Location 2 --------- Sample size $n1 = 460$ $n2 = 650$ Positive cases $x1 = 250$ $x2 = 250$ Level of significance: $\alpha = 0.10$ --- $$\hat{p}1 = \frac{250}{460} = 0.5435$$ $$\hat{p}2 = \frac{250}{650} = 0.3846$$ --- -

  7. 75 marksNumericalone sample tests for mean of normal populaAnswer

    Hypothesis Testing for Average Age of CSIT Students

    Step 1: Set up Hypotheses

    • Null Hypothesis (H₀): $\mu = 22$ (The average age of enrolling students is 22 years)
    • Alternative Hypothesis (H₁): $\mu < 22$ (The average age is less than 22 years)

    This is a left-tailed test at 5% level of significance.

    Step 2: Sample Data

    | 20 | 19 | 22 | 23 | 19 | 20 | 20 | 21 | 22 | 20 | 19 | 20 |

    Sample size: $n = 12$

    Step 3: Calculate Sample Statistics

    Sample mean: $$\bar{x} = \frac{20 + 19 + 22 + 23 + 19 + 20 + 20 + 21 + 22 + 20 + 19 + 20}{12} = \frac{245}{12} = 20.417$$

    Sample standard deviation: $$s = \sqrt{\frac{\sum(x_i - \bar{x})^2}{n-1}} = \sqrt{\frac{26.917}{11}} = \sqrt{2.447} = 1.564$$

    Step 4: Test Statistic

    Since the population is normally distributed and population standard deviation is unknown, we use the t-test:

    $$t = \frac{\bar{x} - \mu_0}{s/\sqrt{n}} = \frac{20.417 - 22}{1.564/\sqrt{12}} = \frac{-1.583}{0.451} = -3.511$$

    Step 5: Critical Value

    For a left-tailed test with $\alpha = 0.05$ and $df = n - 1 = 11$: $$t_{0.05, 11} = -1.796$$

    Step 6: Decision

    Since $t = -3.511 < -1.796$, we reject the null hypothesis.

    Conclusion: At the 5% level of significance, there is sufficient evidence to support the researcher's doubt that the average age of CSIT enrolling students is less than 22 years. [5]

    • Claimed population mean: $\mu0 = 22$ years - Sample data: 20, 19, 22, 23, 19, 20, 20, 21, 22, 20, 19, 20 - Sample size: $n = 12$ - Level of significance: $\alpha = 0.05$ - Population normal, population SD unknown → one-sample t-test - ...
  8. 85 marksNumericalTwo independent sample testAnswer

    Apply the Mann-Whitney U test for examining the following knowledge score on IT among two groups of IT workers at a 5% level of significance.

    Group A:58276
    Group B:91246

    [5]

    Group A (n₁ = 5) 5 8 2 7 6 ------------------ Group B (n₂ = 4) 9 12 4 6 - H₀: No significant difference in knowledge scores between the two groups. - H₁: There is a significant difference between the two groups. - Significance level: α =...

  9. 95 marksNumericalTwo independent sample testAnswer

    Chi-Square Test for Association Between Email Account Type and Hacking Status

    A survey was conducted to see the association between hacking status of the email and the type of email account. The survey has reported the following cross tabulation. Do the information provide sufficient evidence to conclude that the type email account and the hacking status is associated? Use Chi-square test at 1% level of significance.

    $$\begin{array}{|c|c|c|} \hline \text{Type of e-mail account} & \text{Hacking status Yes} & \text{Hacking status No} \ \hline \text{Yahoo} & 60 & 15 \ \text{Gmail} & 20 & 120 \ \hline \end{array}$$

    [5]

    Chi-Square Test for Association Between Email Account Type and Hacking Status

    Step 1 - Given Data

    Email AccountYes (Hacked)No (Not Hacked)Row Total
    Yahoo601575
    Gmail20120140
    Column Total80135215

    Level of significance: $\alpha = 0.01$

    Step 2 - Hypotheses

    • $H_0$: Email account type and hacking status are independent (not associated)
    • $H_1$: Email account type and hacking status are associated

    Step 3 - Expected Frequencies

    $$E_{ij} = \frac{(\text{Row Total}_i)(\text{Column Total}_j)}{\text{Grand Total}}$$

    $$E_{11} = \frac{75 \times 80}{215} = 27.907$$ $$E_{12} = \frac{75 \times 135}{215} = 47.093$$ $$E_{21} = \frac{140 \times 80}{215} = 52.093$$ $$E_{22} = \frac{140 \times 135}{215} = 87.907$$

    Step 4 - Test Statistic

    $$\chi^2 = \sum \frac{(O - E)^2}{E}$$

    OEO - E(O-E)²(O-E)²/E
    6027.90732.0931029.9636.907
    1547.093-32.0931029.9621.870
    2052.093-32.0931029.9619.771
    12087.90732.0931029.9611.717

    $$\chi^2_{cal} = 36.907 + 21.870 + 19.771 + 11.717 = 90.27$$

    (Small rounding differences; approximately $90.2$-$90.3$.)

    Step 5 - Degrees of Freedom

    $$df = (r-1)(c-1) = (2-1)(2-1) = 1$$

    Step 6 - Critical Value

    $$\chi^2_{0.01,,1} = 6.635$$

    Step 7 - Decision

    Since $\chi^2_{cal} = 90.27 > \chi^2_{tab} = 6.635$, we reject $H_0$.

    Conclusion

    At the 1% level of significance, there is sufficient evidence to conclude that the type of email account and the hacking status are associated. Yahoo accounts show a markedly higher hacking rate (60/75 = 80%) than Gmail accounts (20/140 ≈ 14%).

    Note (optional): For a $2\times2$ table, Yates' continuity correction may be applied: $$\chi^2_{Yates} = \sum \frac{(|O-E| - 0.5)^2}{E} \approx 87.5$$ Either way, the conclusion is unchanged (reject $H_0$).

  10. 105 marksLatin Square DesignAnswer

    State the mathematical model for Statistical analysis for m x m LSD for one observation per experimental unit. Also prepare a dummy ANOVA table for this. [5]

    For a m × m Latin Square Design (LSD) with one observation per experimental unit, the linear statistical model is: $$X{ijk} = \mu + \alphai + \betaj + \gammak + \varepsilon{ijk}$$ where: Symbol Meaning ----------------- $X{ijk}$ Observat...

  11. 115 marksMarkov ProcessAnswer

    Define the Markov chain and introduce its basic notations. Also, explain the characteristics of a Markov chain. [5]

    A Markov chain is a stochastic process {X(t), t = 0, 1, 2, ...} that satisfies the Markov property (memoryless property), which states that the future state of the process depends only on the present state and not on the past states. Mat...

  12. 125 marksNeeds of applying non-parametric testsAnswer

    Write short notes on the following: I. The rationale of using the non-parametric statistical test II. Estimation of minimum size for the given proportion [5]

    --- Non-parametric statistical tests (also called distribution-free tests) are statistical methods that do not require assumptions about the underlying population distribution (e.g., normality). 1. No assumption about population distribu...