2080

STA215 · TU past paper

Statistics II 2080 question paper

The complete TU 2080 exam paper for Statistics II (STA215), all 11 questions with solved model answers written to the mark scheme.

Tap a question to open its answer.

  1. 110 marksNumericalTest of significance of regressionAnswer

    Regression Analysis: Computer Program Efficiency

    Given regression equation: $$Y = 52.7 - 2.87X_1 + 0.85X_2$$

    Where:

    • Y = Processed requests per hour
    • X₁ = Data size (gigabytes)
    • X₂ = Number of tables
    • Total sum of squares (TSS) = 1452
    • Sum of squares due to regression (SSR) = 1143.3
    • Standard error of b₂ = 0.55
    • n = 7 observations

    Multiple Regression Analysis - Verified Model Answer

    Given Data

    • Regression equation: $\hat{Y} = 52.7 - 2.87X_1 + 0.85X_2$
    • Total Sum of Squares: $SST = 1452$
    • Sum of Squares due to Regression: $SSR = 1143.3$
    • Standard error of $b_2$: $S_{b_2} = 0.55$
    • Sample size: $n = 7$; number of predictors: $k = 2$
    • Significance level: $\alpha = 0.05$
    • Data table (Y, X1, X2) as given.

    a) Interpretation of $b_1$ and $b_2$

    $b_1 = -2.87$: Holding number of tables ($X_2$) constant, each additional gigabyte of data size decreases the processed requests by about 2.87 per hour on average.

    $b_2 = +0.85$: Holding data size ($X_1$) constant, each additional table used increases the processed requests by about 0.85 per hour on average.


    b) Significance of the Overall Model (F-test)

    $$SSE = SST - SSR = 1452 - 1143.3 = 308.7$$

    $$MSR = \frac{SSR}{k} = \frac{1143.3}{2} = 571.65$$

    $$MSE = \frac{SSE}{n-k-1} = \frac{308.7}{4} = 77.175$$

    $$F_{cal} = \frac{MSR}{MSE} = \frac{571.65}{77.175} = 7.41$$

    Critical value at $\alpha=0.05$, $df_1=2$, $df_2=4$: $$F_{tab}(2,4) = 6.94$$

    Since $F_{cal}=7.41 > F_{tab}=6.94$, reject $H_0$.

    Conclusion: The overall regression model is statistically significant at 0.05.


    c) Significance of $X_2$ (Number of Tables)

    Hypotheses: $H_0: \beta_2 = 0$ vs $H_1: \beta_2 \neq 0$

    $$t_{cal} = \frac{b_2}{S_{b_2}} = \frac{0.85}{0.55} = 1.545$$

    Critical value (two-tailed) at $\alpha=0.05$, $df = n-k-1 = 4$: $$t_{tab} = 2.776$$

    Since $|t_{cal}| = 1.545 < 2.776$, fail to reject $H_0$.

    Conclusion: There is no significant relationship between processed requests and number of tables at the 0.05 level.


    d) Percentage of Variation Explained ($R^2$)

    $$R^2 = \frac{SSR}{SST} = \frac{1143.3}{1452} = 0.7874$$

    $$R^2 \times 100 = 78.74%$$

    About 78.74% of the variation in processed requests is explained by data size and number of tables together.


    e) Standard Error of Estimate

    $$S_e = \sqrt{\frac{SSE}{n-k-1}} = \sqrt{\frac{308.7}{4}} = \sqrt{77.175} = 8.785$$

    $S_e \approx 8.79$ requests per hour.


    f) Estimate for $X_1 = 9$, $X_2 = 8$

    $$\hat{Y} = 52.7 - 2.87(9) + 0.85(8)$$ $$= 52.7 - 25.83 + 6.80 = 33.67$$

    Estimated processed requests $\approx 33.67$ per hour.


    Summary

    PartResult
    a$b_1=-2.87$, $b_2=+0.85$ (interpreted above)
    b$F=7.41 > 6.94$: model significant
    c$t=1.545 < 2.776$: $X_2$ not significant
    d$R^2 = 78.74%$
    e$S_e = 8.79$
    f$\hat{Y} = 33.67$
  2. 210 marksNumericalKruskal Wallis testAnswer

    Kruskal-Wallis Test for Propellant Burning Rates

    In an experiment to determine which of three different missile systems is preferable, the propellant burning rate is measured. The data after coding are given in the table. Use Kruskal-Wallis test (significance level of 0.01) to test the hypothesis that the propellant burning rates are same for three missile systems.

    1234567
    Missile system I22.316.722.719.318.5
    Missile system II23.419.517.520.816.019.9
    Missile system III18.419.517.818.019.622.817.1

    [10]

    Kruskal-Wallis Test for Missile System Propellant Burning Rates

    Step 1: Given Data

    System I (n₁=5)System II (n₂=6)System III (n₃=7)
    22.323.418.4
    16.719.519.5
    22.717.517.8
    19.320.818.0
    18.516.019.6
    19.922.8
    17.1

    Total $N = 5 + 6 + 7 = 18$

    Step 2: Hypotheses

    • $H_0$: The propellant burning rates are the same for all three missile systems.
    • $H_1$: At least one system differs.

    Step 3: Rank All 18 Observations (ascending)

    ValueSystemRank
    16.0II1
    16.7I2
    17.1III3
    17.5II4
    17.8III5
    18.0III6
    18.4III7
    18.5I8
    19.3I9
    19.5II10.5
    19.5III10.5
    19.6III12
    19.9II13
    20.8II14
    22.3I15
    22.7I16
    22.8III17
    23.4II18

    Tie at 19.5: average rank $= (10+11)/2 = 10.5$

    Step 4: Rank Sums

    System I: $R_1 = 2 + 8 + 9 + 15 + 16 = 50$

    System II: $R_2 = 1 + 4 + 10.5 + 13 + 14 + 18 = 60.5$

    System III: $R_3 = 3 + 5 + 6 + 7 + 10.5 + 12 + 17 = 60.5$

    Check: $50 + 60.5 + 60.5 = 171 = \dfrac{N(N+1)}{2} = \dfrac{18(19)}{2} = 171$ ✓

    Step 5: Test Statistic

    $$H = \frac{12}{N(N+1)} \sum \frac{R_i^2}{n_i} - 3(N+1)$$

    $$\frac{50^2}{5} = 500, \quad \frac{60.5^2}{6} = \frac{3660.25}{6} = 610.042, \quad \frac{60.5^2}{7} = \frac{3660.25}{7} = 522.893$$

    Sum $= 500 + 610.042 + 522.893 = 1632.935$

    $$H = \frac{12}{18 \times 19}(1632.935) - 3(19) = \frac{12}{342}(1632.935) - 57$$

    $$H = 0.035088 \times 1632.935 - 57 = 57.296 - 57 = 0.296 \approx 0.30$$

    Correction for ties (optional, one tie group of size 2):

    $$C = 1 - \frac{\sum(t^3 - t)}{N^3 - N} = 1 - \frac{2^3 - 2}{18^3 - 18} = 1 - \frac{6}{5814} = 1 - 0.001032 = 0.99897$$

    $$H_{corrected} = \frac{0.296}{0.99897} \approx 0.296$$

    The tie correction is negligible.

    Step 6: Critical Value

    $H$ follows $\chi^2$ with $df = k - 1 = 2$.

    $$\chi^2_{0.01, 2} = 9.210$$

    Step 7: Decision

    $$H = 0.30 < 9.210$$

    Fail to reject $H_0$.

    Conclusion

    At the 1% significance level, there is insufficient evidence to conclude that the propellant burning rates differ among the three missile systems. The burning rates may be considered the same.

  3. 310 marksLatin Square DesignAnswer

    What is Latin Square Design? Under what conditions can this be used? Give lay out and analysis of Latin Square Design.[10]

    Latin Square Design (LSD)

    Definition

    Latin Square Design (LSD) is a design of experiment used for non-homogeneous experimental material where local control is applied simultaneously in two directions (row-wise and column-wise), thereby converting non-homogeneous material into homogeneous. It is more efficient than Randomized Block Design (RBD) because it controls experimental error in two directions at the same time.

    The shape of LSD is always a square since it contains an equal number of rows, columns, and treatments. If there are m treatments, the layout is an m x m square.

    LSD is based on all three principles of experimental design: Randomization, Replication, and Local Control.


    Conditions for Use of LSD

    LSD can be used under the following conditions:

    1. The experimental material is non-homogeneous in two directions (rows and columns).
    2. The number of rows = number of columns = number of treatments (i.e., the layout must be a square).
    3. The number of treatments is neither too small (less than 4) nor too large (more than 8), as a very large square becomes difficult to manage.
    4. When it is desired to control variation in two directions simultaneously.
    5. When local control needs to be applied both row-wise and column-wise.

    Layout of Latin Square Design

    For a 4 x 4 LSD with treatments A, B, C, D, a typical layout is:

    Col 1Col 2Col 3Col 4
    Row 1ABCD
    Row 2BCDA
    Row 3CDAB
    Row 4DABC

    Key property: Each treatment appears exactly once in each row and exactly once in each column.


    Mathematical Model

    The mathematical model for LSD is:

    $$Y_{ijk} = \mu + \alpha_i + \beta_j + \tau_k + e_{ijk}$$

    Where:

    • $Y_{ijk}$ = observation in the $i$-th row and $j$-th column receiving the $k$-th treatment
    • $\mu$ = general mean effect
    • $\alpha_i$ = effect due to $i$-th row $(i = 1, 2, \ldots, m)$
    • $\beta_j$ = effect due to $j$-th column $(j = 1, 2, \ldots, m)$
    • $\tau_k$ = effect due to $k$-th treatment $(k = 1, 2, \ldots, m)$
    • $e_{ijk}$ = random error (chance variation), assumed $\sim N(0, \sigma^2)$

    Analysis of Latin Square Design

    Hypotheses

    • For Rows: $H_0$: There is no significant difference among rows vs $H_1$: There is significant difference among rows.
    • For Columns: $H_0$: There is no significant difference among columns vs $H_1$: There is significant difference among columns.
    • For Treatments: $H_0$: There is no significant difference among treatments vs $H_1$: There is significant difference among treatments.

    Notation

    Let:

    • $m$ = number of treatments (also number of rows and columns)
    • $N = m^2$ = total number of observations
    • $T$ = Grand total of all observations
    • $R_i$ = Total of $i$-th row
    • $C_j$ = Total of $j$-th column
    • $T_k$ = Total of $k$-th treatment

    Correction Factor and Sums of Squares

    Correction Factor (CF): $$CF = \frac{T^2}{N} = \frac{T^2}{m^2}$$

    Total Sum of Squares (TSS): $$TSS = \sum_{i}\sum_{j} Y_{ij}^2 - CF$$

    Sum of Squares for Rows (SSR): $$SSR = \frac{1}{m}\sum_{i=1}^{m} R_i^2 - CF$$

    Sum of Squares for Columns (SSC): $$SSC = \frac{1}{m}\sum_{j=1}^{m} C_j^2 - CF$$

    Sum of Squares for Treatments (SSTr): $$SSTr = \frac{1}{m}\sum_{k=1}^{m} T_k^2 - CF$$

    Error Sum of Squares (SSE): $$SSE = TSS - SSR - SSC - SSTr$$


    Degrees of Freedom

    SourceDegrees of Freedom (df)
    Rows$m - 1$
    Columns$m - 1$
    Treatments$m - 1$
    Error$(m-1)(m-2)$
    Total$m^2 - 1$

    Mean Squares

    $$MSR = \frac{SSR}{m-1}, \quad MSC = \frac{SSC}{m-1}, \quad MSTr = \frac{SSTr}{m-1}, \quad MSE = \frac{SSE}{(m-1)(m-2)}$$


    ANOVA Table for LSD

    Source of VariationSSdfMSF-ratio
    RowsSSR$m-1$MSR$F_R = \dfrac{MSR}{MSE}$
    ColumnsSSC$m-1$MSC$F_C = \dfrac{MSC}{MSE}$
    TreatmentsSSTr$m-1$MSTr$F_{Tr} = \dfrac{MSTr}{MSE}$
    ErrorSSE
  4. 45 marksNumericalDetermination of sample sizeAnswer

    What do you understand by estimation? If we want to determine average mechanical aptitude of a large group of workers, how large a random sample is needed to be able to assert with probability 0.95 that the sample mean will not differ from the true mean by more than 2.0 points? Assume that population standard deviation is 30. [5]

    Parameter Value ------------------ Confidence level $(1-\alpha)$ 0.95 Maximum allowable error $(E)$ 2.0 points Population standard deviation $(\sigma)$ 30 Critical value $Z{\alpha/2}$ at 95% 1.96 Estimation is the statistical procedure o...

  5. 55 marksNumericalTwo independent sample testAnswer

    Test of Independence: Opinion on Core Curriculum Change vs. Class Standing

    A random sample of students is asked their opinion on proposed core curriculum change. The results are as follows. Test the hypothesis that opinion on the change is independent of class standing. Use 0.01 significance level.

    $$\begin{array}{|c|c|c|} \hline \text{Class} & \text{Favoring} & \text{Opposing} \ \hline \text{Freshman} & 125 & 80 \ \text{Sophomore} & 60 & 140 \ \text{Junior} & 50 & 60 \ \text{Senior} & 40 & 55 \ \hline \end{array}$$

    [5]

    Observed frequencies: Class Favoring Opposing --------------------------- Freshman 125 80 Sophomore 60 140 Junior 50 60 Senior 40 55 Significance level: $\alpha = 0.01$ - $H0$: Opinion on the change is independent of class standing. -

  6. 65 marksNumericalCentral Limit TheoremAnswer

    Define Central limit theorem. The life of a certain brand of an electric bulb may be considered a random variable with mean 1350 hours and standard deviation 550 hours. Using central limit theorem, find the probability that the average life time of 100 bulbs exceeds 1440 hours. [5]

    Given data: - Population mean: $\mu = 1350$ hours - Population standard deviation: $\sigma = 550$ hours - Sample size: $n = 100$ bulbs - Value tested: $\bar{X} = 1440$ hours - Required: $P(\bar{X} 1440)$ All data present. The Central Lim...

  7. 75 marksNumericalMultiple and partial correlationAnswer

    Define multiple correlation. In a trivariate distribution X1, X2, and X3, the simple correlation coefficients are given as r12= 0.5, r23=0.6 and r13=0.7. Find i. partial correlation coefficient between X1 and X2 keeping X3 constant. ii. multiple correlation coefficient assuming X1 as dependent variable. [5]

    Multiple Correlation and Partial Correlation

    Definition of Multiple Correlation

    Multiple correlation is the correlation between one variable (the dependent variable) and the combined linear effect of two or more other variables (independent variables) taken together. For a trivariate distribution $X_1, X_2, X_3$, the multiple correlation coefficient $R_{1.23}$ measures the degree of association between $X_1$ and the joint effect of $X_2$ and $X_3$. Its square gives the proportion of variance in $X_1$ explained by $X_2$ and $X_3$.

    Given Data

    • $r_{12} = 0.5$
    • $r_{23} = 0.6$
    • $r_{13} = 0.7$

    Part (i): Partial Correlation $r_{12.3}$

    $$r_{12.3} = \frac{r_{12} - r_{13},r_{23}}{\sqrt{(1 - r_{13}^2)(1 - r_{23}^2)}}$$

    Numerator: $$0.5 - (0.7)(0.6) = 0.5 - 0.42 = 0.08$$

    Denominator: $$\sqrt{(1 - 0.49)(1 - 0.36)} = \sqrt{(0.51)(0.64)} = \sqrt{0.3264} = 0.5713$$

    Result: $$r_{12.3} = \frac{0.08}{0.5713} = 0.1400$$

    A weak positive partial correlation between $X_1$ and $X_2$ when $X_3$ is held constant.


    Part (ii): Multiple Correlation $R_{1.23}$

    $$R_{1.23} = \sqrt{\frac{r_{12}^2 + r_{13}^2 - 2,r_{12},r_{13},r_{23}}{1 - r_{23}^2}}$$

    Numerator: $$(0.5)^2 + (0.7)^2 - 2(0.5)(0.7)(0.6) = 0.25 + 0.49 - 0.42 = 0.32$$

    Denominator: $$1 - (0.6)^2 = 1 - 0.36 = 0.64$$

    Result: $$R_{1.23} = \sqrt{\frac{0.32}{0.64}} = \sqrt{0.5} = 0.7071$$

    Interpretation: $R_{1.23}^2 = 0.50$, so about 50% of the variation in $X_1$ is jointly explained by $X_2$ and $X_3$.


    Final Answers

    • $r_{12.3} = 0.1400$
    • $R_{1.23} = 0.7071$
  8. 85 marksNumericalANOVA tableAnswer

    What do you understand by Design of Experiment? Prepare one way analysis of variance table and carry out the test for the significance of difference in the average yields between different varieties of seed. Given: Total sum of squares = 258, Sum of square between varieties of seed = 50, Total number of observations = 20 [5]

    Design of Experiment (DOE) is the systematic planning, conducting, and statistical analysis of experiments so that valid and objective conclusions can be drawn about the effect of one or more factors on a response. It aims to control exp...

  9. 95 marksNumericaltest for single proportionAnswer

    Define type I and type II error in testing of hypothesis. It is claimed that Samsung and Huawei mobiles are equally popular in Kathmandu. A random sample of 600 people from Kathmandu showed 350 have Samsung mobile. Test the claim at 5% level of significance. [5]

    Type I and Type II Errors + Test of Hypothesis on Mobile Popularity

    Given Data

    • Sample size: $n = 600$
    • Number with Samsung: $X = 350$
    • Claim: equally popular, so $p_0 = 0.5$
    • Level of significance: $\alpha = 0.05$

    Part 1: Type I and Type II Error

    Type I Error ($\alpha$): Rejecting the null hypothesis $H_0$ when it is actually true. Its probability equals the level of significance $\alpha$.

    Type II Error ($\beta$): Accepting the null hypothesis $H_0$ when it is actually false.

    Decision$H_0$ True$H_0$ False
    Accept $H_0$Correct DecisionType II Error ($\beta$)
    Reject $H_0$Type I Error ($\alpha$)Correct Decision

    Part 2: Test of Hypothesis

    Step 1: Hypotheses

    $$H_0: p = 0.5 \quad \text{(equally popular)}$$ $$H_1: p \neq 0.5 \quad \text{(not equally popular)} \quad \text{(two-tailed test)}$$

    Step 2: Sample Proportion

    $$\hat{p} = \frac{X}{n} = \frac{350}{600} = 0.5833$$

    Step 3: Test Statistic (Z-test for proportion)

    $$Z = \frac{\hat{p} - p_0}{\sqrt{\dfrac{p_0 q_0}{n}}}$$

    where $q_0 = 1 - 0.5 = 0.5$.

    $$\text{S.E.} = \sqrt{\frac{0.5 \times 0.5}{600}} = \sqrt{\frac{0.25}{600}} = \sqrt{0.0004167} = 0.020412$$

    $$Z = \frac{0.5833 - 0.5}{0.020412} = \frac{0.0833}{0.020412} = 4.08$$

    On the continuity correction: applying it gives $Z_{cal} = 4.04$; without it (the standard approach at this level) $Z = 4.08$. The two are numerically equivalent for the decision. Using the raw count form:

    $$Z = \frac{X - np_0}{\sqrt{np_0 q_0}} = \frac{350 - 300}{\sqrt{150}} = \frac{50}{12.247} = 4.08$$

    Step 4: Critical Value

    At $\alpha = 0.05$, two-tailed:

    $$Z_{\alpha/2} = Z_{0.025} = 1.96$$

    Step 5: Decision

    $$|Z| = 4.08 > 1.96$$

    Since the calculated value exceeds the critical value, we reject $H_0$.

    Step 6: Conclusion

    At the 5% level of significance, there is sufficient evidence to reject the claim that Samsung and Huawei mobiles are equally popular in Kathmandu. The data indicates Samsung is significantly more popular.


    Final Result: $Z_{cal} = 4.08$ (without continuity correction) vs $Z_{tab} = 1.96$; reject $H_0$. With the continuity correction $Z_{cal} = 4.04$, and the decision and conclusion are identical.

  10. 105 marksNumericalCounting processAnswer

    Customers of certain Internet service provider connect to the internet at the average rate of 10 new connections per minute. Connections are modelle by binomial counting process. a. What frame length gives the probability 0.1 of an arrival during given frame? b. Find the mean and variance for the number of seconds between two consecutive connections. [5]

    • Arrival rate: $\lambdaA = 10$ connections per minute - Process model: Binomial counting process - Part (a): required frame probability $p = 0.1$ - Part (b): find mean and variance of inter-arrival time in seconds Convert rate to second...
  11. 115 marksParametric vs. non-parametric testAnswer

    Write short notes on any two: a. Difference between parametric and non-parametric test. b. Required assumptions for linear regression model. c. Stochastic process. [5]

    Short Notes on Any Two


    a. Difference Between Parametric and Non-Parametric Test

    BasisParametric TestNon-Parametric Test
    AssumptionRequires strict assumptions about population distribution (usually normality)Does not require assumptions about population distribution
    Data TypeWorks on interval or ratio scale dataWorks on nominal or ordinal scale data
    Population ParameterMakes inferences about population parameters (mean, variance)Makes inferences without reference to specific population parameters
    Sample SizeRequires relatively large sample sizeCan be used with small sample sizes
    PowerMore powerful when assumptions are metLess powerful compared to parametric tests
    Examplest-test, F-test, Z-testMann-Whitney U-test, Median test, Kruskal-Wallis test

    **Note ** The Mann-Whitney U-test is a non-parametric test used as an alternative to the T-test when T-test fails to satisfy the normality assumptions. This clearly illustrates the key distinction: parametric tests demand distributional assumptions, while non-parametric tests do not.


    b. Required Assumptions for Linear Regression Model

    For a linear regression model to produce valid and reliable results, the following assumptions must be satisfied:

    1. Linearity The relationship between the dependent variable (Y) and independent variable(s) (X) must be linear in nature.

    2. Independence of Errors The error terms (residuals) must be independent of each other. There should be no autocorrelation among residuals.

    3. Normality of Errors The error terms should follow a normal distribution with mean zero:

    $$E(\varepsilon_i) = 0$$

    4. Homoscedasticity (Constant Variance) The variance of error terms should be constant across all levels of the independent variable:

    $$Var(\varepsilon_i) = \sigma^2 \text{ (constant)}$$

    5. No Multicollinearity (for multiple regression) In multiple regression, the independent variables (X1, X2, ...) should not be highly correlated with each other. In multiple regression with variables X1 and X2, each variable is tested separately for its contribution, implying they should be independently contributing.

    6. No Measurement Error The independent variables are assumed to be measured without error.

    7. Significance of Regression Coefficients The regression coefficients (b1, b2, ...) should be significantly different from zero, tested using the t-test:

    $$t_{cal} = \frac{b_i}{S_{b_i}}$$

    **Example ** With n=20, b1=4, Sb1=1.2, the calculated t = 3.33 > t-tab = 2.110, confirming a significant linear relationship between Y and X1, validating the regression model assumption.