STA215 · TU past paper
Statistics II 2080 question paper
The complete TU 2080 exam paper for Statistics II (STA215), all 11 questions with solved model answers written to the mark scheme.
Tap a question to open its answer.
- 110 marksNumericalTest of significance of regressionHideAnswer
Regression Analysis: Computer Program Efficiency
Given regression equation: $$Y = 52.7 - 2.87X_1 + 0.85X_2$$
Where:
- Y = Processed requests per hour
- X₁ = Data size (gigabytes)
- X₂ = Number of tables
- Total sum of squares (TSS) = 1452
- Sum of squares due to regression (SSR) = 1143.3
- Standard error of b₂ = 0.55
- n = 7 observations
Multiple Regression Analysis - Verified Model Answer
Given Data
- Regression equation: $\hat{Y} = 52.7 - 2.87X_1 + 0.85X_2$
- Total Sum of Squares: $SST = 1452$
- Sum of Squares due to Regression: $SSR = 1143.3$
- Standard error of $b_2$: $S_{b_2} = 0.55$
- Sample size: $n = 7$; number of predictors: $k = 2$
- Significance level: $\alpha = 0.05$
- Data table (Y, X1, X2) as given.
a) Interpretation of $b_1$ and $b_2$
$b_1 = -2.87$: Holding number of tables ($X_2$) constant, each additional gigabyte of data size decreases the processed requests by about 2.87 per hour on average.
$b_2 = +0.85$: Holding data size ($X_1$) constant, each additional table used increases the processed requests by about 0.85 per hour on average.
b) Significance of the Overall Model (F-test)
$$SSE = SST - SSR = 1452 - 1143.3 = 308.7$$
$$MSR = \frac{SSR}{k} = \frac{1143.3}{2} = 571.65$$
$$MSE = \frac{SSE}{n-k-1} = \frac{308.7}{4} = 77.175$$
$$F_{cal} = \frac{MSR}{MSE} = \frac{571.65}{77.175} = 7.41$$
Critical value at $\alpha=0.05$, $df_1=2$, $df_2=4$: $$F_{tab}(2,4) = 6.94$$
Since $F_{cal}=7.41 > F_{tab}=6.94$, reject $H_0$.
Conclusion: The overall regression model is statistically significant at 0.05.
c) Significance of $X_2$ (Number of Tables)
Hypotheses: $H_0: \beta_2 = 0$ vs $H_1: \beta_2 \neq 0$
$$t_{cal} = \frac{b_2}{S_{b_2}} = \frac{0.85}{0.55} = 1.545$$
Critical value (two-tailed) at $\alpha=0.05$, $df = n-k-1 = 4$: $$t_{tab} = 2.776$$
Since $|t_{cal}| = 1.545 < 2.776$, fail to reject $H_0$.
Conclusion: There is no significant relationship between processed requests and number of tables at the 0.05 level.
d) Percentage of Variation Explained ($R^2$)
$$R^2 = \frac{SSR}{SST} = \frac{1143.3}{1452} = 0.7874$$
$$R^2 \times 100 = 78.74%$$
About 78.74% of the variation in processed requests is explained by data size and number of tables together.
e) Standard Error of Estimate
$$S_e = \sqrt{\frac{SSE}{n-k-1}} = \sqrt{\frac{308.7}{4}} = \sqrt{77.175} = 8.785$$
$S_e \approx 8.79$ requests per hour.
f) Estimate for $X_1 = 9$, $X_2 = 8$
$$\hat{Y} = 52.7 - 2.87(9) + 0.85(8)$$ $$= 52.7 - 25.83 + 6.80 = 33.67$$
Estimated processed requests $\approx 33.67$ per hour.
Summary
Part Result a $b_1=-2.87$, $b_2=+0.85$ (interpreted above) b $F=7.41 > 6.94$: model significant c $t=1.545 < 2.776$: $X_2$ not significant d $R^2 = 78.74%$ e $S_e = 8.79$ f $\hat{Y} = 33.67$ - 210 marksNumericalKruskal Wallis testHideAnswer
Kruskal-Wallis Test for Propellant Burning Rates
In an experiment to determine which of three different missile systems is preferable, the propellant burning rate is measured. The data after coding are given in the table. Use Kruskal-Wallis test (significance level of 0.01) to test the hypothesis that the propellant burning rates are same for three missile systems.
1 2 3 4 5 6 7 Missile system I 22.3 16.7 22.7 19.3 18.5 Missile system II 23.4 19.5 17.5 20.8 16.0 19.9 Missile system III 18.4 19.5 17.8 18.0 19.6 22.8 17.1 [10]
Kruskal-Wallis Test for Missile System Propellant Burning Rates
Step 1: Given Data
System I (n₁=5) System II (n₂=6) System III (n₃=7) 22.3 23.4 18.4 16.7 19.5 19.5 22.7 17.5 17.8 19.3 20.8 18.0 18.5 16.0 19.6 19.9 22.8 17.1 Total $N = 5 + 6 + 7 = 18$
Step 2: Hypotheses
- $H_0$: The propellant burning rates are the same for all three missile systems.
- $H_1$: At least one system differs.
Step 3: Rank All 18 Observations (ascending)
Value System Rank 16.0 II 1 16.7 I 2 17.1 III 3 17.5 II 4 17.8 III 5 18.0 III 6 18.4 III 7 18.5 I 8 19.3 I 9 19.5 II 10.5 19.5 III 10.5 19.6 III 12 19.9 II 13 20.8 II 14 22.3 I 15 22.7 I 16 22.8 III 17 23.4 II 18 Tie at 19.5: average rank $= (10+11)/2 = 10.5$
Step 4: Rank Sums
System I: $R_1 = 2 + 8 + 9 + 15 + 16 = 50$
System II: $R_2 = 1 + 4 + 10.5 + 13 + 14 + 18 = 60.5$
System III: $R_3 = 3 + 5 + 6 + 7 + 10.5 + 12 + 17 = 60.5$
Check: $50 + 60.5 + 60.5 = 171 = \dfrac{N(N+1)}{2} = \dfrac{18(19)}{2} = 171$ ✓
Step 5: Test Statistic
$$H = \frac{12}{N(N+1)} \sum \frac{R_i^2}{n_i} - 3(N+1)$$
$$\frac{50^2}{5} = 500, \quad \frac{60.5^2}{6} = \frac{3660.25}{6} = 610.042, \quad \frac{60.5^2}{7} = \frac{3660.25}{7} = 522.893$$
Sum $= 500 + 610.042 + 522.893 = 1632.935$
$$H = \frac{12}{18 \times 19}(1632.935) - 3(19) = \frac{12}{342}(1632.935) - 57$$
$$H = 0.035088 \times 1632.935 - 57 = 57.296 - 57 = 0.296 \approx 0.30$$
Correction for ties (optional, one tie group of size 2):
$$C = 1 - \frac{\sum(t^3 - t)}{N^3 - N} = 1 - \frac{2^3 - 2}{18^3 - 18} = 1 - \frac{6}{5814} = 1 - 0.001032 = 0.99897$$
$$H_{corrected} = \frac{0.296}{0.99897} \approx 0.296$$
The tie correction is negligible.
Step 6: Critical Value
$H$ follows $\chi^2$ with $df = k - 1 = 2$.
$$\chi^2_{0.01, 2} = 9.210$$
Step 7: Decision
$$H = 0.30 < 9.210$$
Fail to reject $H_0$.
Conclusion
At the 1% significance level, there is insufficient evidence to conclude that the propellant burning rates differ among the three missile systems. The burning rates may be considered the same.
- 310 marksLatin Square DesignHideAnswer
What is Latin Square Design? Under what conditions can this be used? Give lay out and analysis of Latin Square Design.[10]
Latin Square Design (LSD)
Definition
Latin Square Design (LSD) is a design of experiment used for non-homogeneous experimental material where local control is applied simultaneously in two directions (row-wise and column-wise), thereby converting non-homogeneous material into homogeneous. It is more efficient than Randomized Block Design (RBD) because it controls experimental error in two directions at the same time.
The shape of LSD is always a square since it contains an equal number of rows, columns, and treatments. If there are m treatments, the layout is an m x m square.
LSD is based on all three principles of experimental design: Randomization, Replication, and Local Control.
Conditions for Use of LSD
LSD can be used under the following conditions:
- The experimental material is non-homogeneous in two directions (rows and columns).
- The number of rows = number of columns = number of treatments (i.e., the layout must be a square).
- The number of treatments is neither too small (less than 4) nor too large (more than 8), as a very large square becomes difficult to manage.
- When it is desired to control variation in two directions simultaneously.
- When local control needs to be applied both row-wise and column-wise.
Layout of Latin Square Design
For a 4 x 4 LSD with treatments A, B, C, D, a typical layout is:
Col 1 Col 2 Col 3 Col 4 Row 1 A B C D Row 2 B C D A Row 3 C D A B Row 4 D A B C Key property: Each treatment appears exactly once in each row and exactly once in each column.
Mathematical Model
The mathematical model for LSD is:
$$Y_{ijk} = \mu + \alpha_i + \beta_j + \tau_k + e_{ijk}$$
Where:
- $Y_{ijk}$ = observation in the $i$-th row and $j$-th column receiving the $k$-th treatment
- $\mu$ = general mean effect
- $\alpha_i$ = effect due to $i$-th row $(i = 1, 2, \ldots, m)$
- $\beta_j$ = effect due to $j$-th column $(j = 1, 2, \ldots, m)$
- $\tau_k$ = effect due to $k$-th treatment $(k = 1, 2, \ldots, m)$
- $e_{ijk}$ = random error (chance variation), assumed $\sim N(0, \sigma^2)$
Analysis of Latin Square Design
Hypotheses
- For Rows: $H_0$: There is no significant difference among rows vs $H_1$: There is significant difference among rows.
- For Columns: $H_0$: There is no significant difference among columns vs $H_1$: There is significant difference among columns.
- For Treatments: $H_0$: There is no significant difference among treatments vs $H_1$: There is significant difference among treatments.
Notation
Let:
- $m$ = number of treatments (also number of rows and columns)
- $N = m^2$ = total number of observations
- $T$ = Grand total of all observations
- $R_i$ = Total of $i$-th row
- $C_j$ = Total of $j$-th column
- $T_k$ = Total of $k$-th treatment
Correction Factor and Sums of Squares
Correction Factor (CF): $$CF = \frac{T^2}{N} = \frac{T^2}{m^2}$$
Total Sum of Squares (TSS): $$TSS = \sum_{i}\sum_{j} Y_{ij}^2 - CF$$
Sum of Squares for Rows (SSR): $$SSR = \frac{1}{m}\sum_{i=1}^{m} R_i^2 - CF$$
Sum of Squares for Columns (SSC): $$SSC = \frac{1}{m}\sum_{j=1}^{m} C_j^2 - CF$$
Sum of Squares for Treatments (SSTr): $$SSTr = \frac{1}{m}\sum_{k=1}^{m} T_k^2 - CF$$
Error Sum of Squares (SSE): $$SSE = TSS - SSR - SSC - SSTr$$
Degrees of Freedom
Source Degrees of Freedom (df) Rows $m - 1$ Columns $m - 1$ Treatments $m - 1$ Error $(m-1)(m-2)$ Total $m^2 - 1$
Mean Squares
$$MSR = \frac{SSR}{m-1}, \quad MSC = \frac{SSC}{m-1}, \quad MSTr = \frac{SSTr}{m-1}, \quad MSE = \frac{SSE}{(m-1)(m-2)}$$
ANOVA Table for LSD
Source of Variation SS df MS F-ratio Rows SSR $m-1$ MSR $F_R = \dfrac{MSR}{MSE}$ Columns SSC $m-1$ MSC $F_C = \dfrac{MSC}{MSE}$ Treatments SSTr $m-1$ MSTr $F_{Tr} = \dfrac{MSTr}{MSE}$ Error SSE - 45 marksNumericalDetermination of sample sizeHideAnswer
What do you understand by estimation? If we want to determine average mechanical aptitude of a large group of workers, how large a random sample is needed to be able to assert with probability 0.95 that the sample mean will not differ from the true mean by more than 2.0 points? Assume that population standard deviation is 30. [5]
Parameter Value ------------------ Confidence level $(1-\alpha)$ 0.95 Maximum allowable error $(E)$ 2.0 points Population standard deviation $(\sigma)$ 30 Critical value $Z{\alpha/2}$ at 95% 1.96 Estimation is the statistical procedure o...
- 55 marksNumericalTwo independent sample testHideAnswer
Test of Independence: Opinion on Core Curriculum Change vs. Class Standing
A random sample of students is asked their opinion on proposed core curriculum change. The results are as follows. Test the hypothesis that opinion on the change is independent of class standing. Use 0.01 significance level.
$$\begin{array}{|c|c|c|} \hline \text{Class} & \text{Favoring} & \text{Opposing} \ \hline \text{Freshman} & 125 & 80 \ \text{Sophomore} & 60 & 140 \ \text{Junior} & 50 & 60 \ \text{Senior} & 40 & 55 \ \hline \end{array}$$
[5]
Observed frequencies: Class Favoring Opposing --------------------------- Freshman 125 80 Sophomore 60 140 Junior 50 60 Senior 40 55 Significance level: $\alpha = 0.01$ - $H0$: Opinion on the change is independent of class standing. -
- 65 marksNumericalCentral Limit TheoremHideAnswer
Define Central limit theorem. The life of a certain brand of an electric bulb may be considered a random variable with mean 1350 hours and standard deviation 550 hours. Using central limit theorem, find the probability that the average life time of 100 bulbs exceeds 1440 hours. [5]
Given data: - Population mean: $\mu = 1350$ hours - Population standard deviation: $\sigma = 550$ hours - Sample size: $n = 100$ bulbs - Value tested: $\bar{X} = 1440$ hours - Required: $P(\bar{X} 1440)$ All data present. The Central Lim...
- 75 marksNumericalMultiple and partial correlationHideAnswer
Define multiple correlation. In a trivariate distribution X1, X2, and X3, the simple correlation coefficients are given as r12= 0.5, r23=0.6 and r13=0.7. Find i. partial correlation coefficient between X1 and X2 keeping X3 constant. ii. multiple correlation coefficient assuming X1 as dependent variable. [5]
Multiple Correlation and Partial Correlation
Definition of Multiple Correlation
Multiple correlation is the correlation between one variable (the dependent variable) and the combined linear effect of two or more other variables (independent variables) taken together. For a trivariate distribution $X_1, X_2, X_3$, the multiple correlation coefficient $R_{1.23}$ measures the degree of association between $X_1$ and the joint effect of $X_2$ and $X_3$. Its square gives the proportion of variance in $X_1$ explained by $X_2$ and $X_3$.
Given Data
- $r_{12} = 0.5$
- $r_{23} = 0.6$
- $r_{13} = 0.7$
Part (i): Partial Correlation $r_{12.3}$
$$r_{12.3} = \frac{r_{12} - r_{13},r_{23}}{\sqrt{(1 - r_{13}^2)(1 - r_{23}^2)}}$$
Numerator: $$0.5 - (0.7)(0.6) = 0.5 - 0.42 = 0.08$$
Denominator: $$\sqrt{(1 - 0.49)(1 - 0.36)} = \sqrt{(0.51)(0.64)} = \sqrt{0.3264} = 0.5713$$
Result: $$r_{12.3} = \frac{0.08}{0.5713} = 0.1400$$
A weak positive partial correlation between $X_1$ and $X_2$ when $X_3$ is held constant.
Part (ii): Multiple Correlation $R_{1.23}$
$$R_{1.23} = \sqrt{\frac{r_{12}^2 + r_{13}^2 - 2,r_{12},r_{13},r_{23}}{1 - r_{23}^2}}$$
Numerator: $$(0.5)^2 + (0.7)^2 - 2(0.5)(0.7)(0.6) = 0.25 + 0.49 - 0.42 = 0.32$$
Denominator: $$1 - (0.6)^2 = 1 - 0.36 = 0.64$$
Result: $$R_{1.23} = \sqrt{\frac{0.32}{0.64}} = \sqrt{0.5} = 0.7071$$
Interpretation: $R_{1.23}^2 = 0.50$, so about 50% of the variation in $X_1$ is jointly explained by $X_2$ and $X_3$.
Final Answers
- $r_{12.3} = 0.1400$
- $R_{1.23} = 0.7071$
- 85 marksNumericalANOVA tableHideAnswer
What do you understand by Design of Experiment? Prepare one way analysis of variance table and carry out the test for the significance of difference in the average yields between different varieties of seed. Given: Total sum of squares = 258, Sum of square between varieties of seed = 50, Total number of observations = 20 [5]
Design of Experiment (DOE) is the systematic planning, conducting, and statistical analysis of experiments so that valid and objective conclusions can be drawn about the effect of one or more factors on a response. It aims to control exp...
- 95 marksNumericaltest for single proportionHideAnswer
Define type I and type II error in testing of hypothesis. It is claimed that Samsung and Huawei mobiles are equally popular in Kathmandu. A random sample of 600 people from Kathmandu showed 350 have Samsung mobile. Test the claim at 5% level of significance. [5]
Type I and Type II Errors + Test of Hypothesis on Mobile Popularity
Given Data
- Sample size: $n = 600$
- Number with Samsung: $X = 350$
- Claim: equally popular, so $p_0 = 0.5$
- Level of significance: $\alpha = 0.05$
Part 1: Type I and Type II Error
Type I Error ($\alpha$): Rejecting the null hypothesis $H_0$ when it is actually true. Its probability equals the level of significance $\alpha$.
Type II Error ($\beta$): Accepting the null hypothesis $H_0$ when it is actually false.
Decision $H_0$ True $H_0$ False Accept $H_0$ Correct Decision Type II Error ($\beta$) Reject $H_0$ Type I Error ($\alpha$) Correct Decision
Part 2: Test of Hypothesis
Step 1: Hypotheses
$$H_0: p = 0.5 \quad \text{(equally popular)}$$ $$H_1: p \neq 0.5 \quad \text{(not equally popular)} \quad \text{(two-tailed test)}$$
Step 2: Sample Proportion
$$\hat{p} = \frac{X}{n} = \frac{350}{600} = 0.5833$$
Step 3: Test Statistic (Z-test for proportion)
$$Z = \frac{\hat{p} - p_0}{\sqrt{\dfrac{p_0 q_0}{n}}}$$
where $q_0 = 1 - 0.5 = 0.5$.
$$\text{S.E.} = \sqrt{\frac{0.5 \times 0.5}{600}} = \sqrt{\frac{0.25}{600}} = \sqrt{0.0004167} = 0.020412$$
$$Z = \frac{0.5833 - 0.5}{0.020412} = \frac{0.0833}{0.020412} = 4.08$$
On the continuity correction: applying it gives $Z_{cal} = 4.04$; without it (the standard approach at this level) $Z = 4.08$. The two are numerically equivalent for the decision. Using the raw count form:
$$Z = \frac{X - np_0}{\sqrt{np_0 q_0}} = \frac{350 - 300}{\sqrt{150}} = \frac{50}{12.247} = 4.08$$
Step 4: Critical Value
At $\alpha = 0.05$, two-tailed:
$$Z_{\alpha/2} = Z_{0.025} = 1.96$$
Step 5: Decision
$$|Z| = 4.08 > 1.96$$
Since the calculated value exceeds the critical value, we reject $H_0$.
Step 6: Conclusion
At the 5% level of significance, there is sufficient evidence to reject the claim that Samsung and Huawei mobiles are equally popular in Kathmandu. The data indicates Samsung is significantly more popular.
Final Result: $Z_{cal} = 4.08$ (without continuity correction) vs $Z_{tab} = 1.96$; reject $H_0$. With the continuity correction $Z_{cal} = 4.04$, and the decision and conclusion are identical.
- 105 marksNumericalCounting processHideAnswer
Customers of certain Internet service provider connect to the internet at the average rate of 10 new connections per minute. Connections are modelle by binomial counting process. a. What frame length gives the probability 0.1 of an arrival during given frame? b. Find the mean and variance for the number of seconds between two consecutive connections. [5]
- Arrival rate: $\lambdaA = 10$ connections per minute - Process model: Binomial counting process - Part (a): required frame probability $p = 0.1$ - Part (b): find mean and variance of inter-arrival time in seconds Convert rate to second...
- 115 marksParametric vs. non-parametric testHideAnswer
Write short notes on any two: a. Difference between parametric and non-parametric test. b. Required assumptions for linear regression model. c. Stochastic process. [5]
Short Notes on Any Two
a. Difference Between Parametric and Non-Parametric Test
Basis Parametric Test Non-Parametric Test Assumption Requires strict assumptions about population distribution (usually normality) Does not require assumptions about population distribution Data Type Works on interval or ratio scale data Works on nominal or ordinal scale data Population Parameter Makes inferences about population parameters (mean, variance) Makes inferences without reference to specific population parameters Sample Size Requires relatively large sample size Can be used with small sample sizes Power More powerful when assumptions are met Less powerful compared to parametric tests Examples t-test, F-test, Z-test Mann-Whitney U-test, Median test, Kruskal-Wallis test **Note ** The Mann-Whitney U-test is a non-parametric test used as an alternative to the T-test when T-test fails to satisfy the normality assumptions. This clearly illustrates the key distinction: parametric tests demand distributional assumptions, while non-parametric tests do not.
b. Required Assumptions for Linear Regression Model
For a linear regression model to produce valid and reliable results, the following assumptions must be satisfied:
1. Linearity The relationship between the dependent variable (Y) and independent variable(s) (X) must be linear in nature.
2. Independence of Errors The error terms (residuals) must be independent of each other. There should be no autocorrelation among residuals.
3. Normality of Errors The error terms should follow a normal distribution with mean zero:
$$E(\varepsilon_i) = 0$$
4. Homoscedasticity (Constant Variance) The variance of error terms should be constant across all levels of the independent variable:
$$Var(\varepsilon_i) = \sigma^2 \text{ (constant)}$$
5. No Multicollinearity (for multiple regression) In multiple regression, the independent variables (X1, X2, ...) should not be highly correlated with each other. In multiple regression with variables X1 and X2, each variable is tested separately for its contribution, implying they should be independently contributing.
6. No Measurement Error The independent variables are assumed to be measured without error.
7. Significance of Regression Coefficients The regression coefficients (b1, b2, ...) should be significantly different from zero, tested using the t-test:
$$t_{cal} = \frac{b_i}{S_{b_i}}$$
**Example ** With n=20, b1=4, Sb1=1.2, the calculated t = 3.33 > t-tab = 2.110, confirming a significant linear relationship between Y and X1, validating the regression model assumption.