Important Questions

STA215 · Exam intelligence

Statistics II important questions

From 6 past TU papers: which questions keep coming back, how much they carry, and what is most likely to show up next. Every question links to a model answer.

Most likely in the next examStatistical

Ranked by how often a topic is asked, its marks weight, and whether it is due after skipping the 2081 paper. No guarantees; study the whole syllabus.

1asked 8xavg 5 marks · Two independent sample test
Answer

Chi-Square Test of Independence

Based on interviews of couples seeking divorce, a social worker compiles the following data related to the period of acquaintanceship before marriage and the duration of marriage. Perform a test to determine if the data substantiate an association between the duration of a marriage and the acquaintanceship prior to marriage. Use 5% level of significance.

Acquaintanceship before marriageLess than or equal to 5 yearsMore than 5 yearsTotal
Below 0.5 year15722
0.5–1.5 years262248
Over 1.5 years191130
Total6040100

[5]

Chi-Square Test for Association

Step 1: Given Data

Observed contingency table:

Acquaintanceship≤ 5 years> 5 yearsTotal
Below 0.5 yr15722
0.5-1.5 yr262248
Over 1.5 yr191130
Total6040100

Significance level $\alpha = 0.05$.

Step 2: Hypotheses

  • $H_0$: No association between duration of marriage and acquaintanceship period (independent).
  • $H_1$: There is an association.

Step 3: Expected Frequencies

$$E_{ij} = \frac{R_i \times C_j}{N}$$

Cell$E$
(Below 0.5, ≤5)$\frac{22 \times 60}{100} = 13.2$
(Below 0.5, >5)$\frac{22 \times 40}{100} = 8.8$
(0.5-1.5, ≤5)$\frac{48 \times 60}{100} = 28.8$
(0.5-1.5, >5)$\frac{48 \times 40}{100} = 19.2$
(Over 1.5, ≤5)$\frac{30 \times 60}{100} = 18.0$
(Over 1.5, >5)$\frac{30 \times 40}{100} = 12.0$

Step 4: Chi-Square Statistic

$$\chi^2 = \sum \frac{(O-E)^2}{E}$$

OE$(O-E)^2$$(O-E)^2/E$
1513.23.240.24545
78.83.240.36818
2628.87.840.27222
2219.27.840.40833
1918.01.000.05556
1112.01.000.08333

$$\chi^2_{cal} = 0.24545 + 0.36818 + 0.27222 + 0.40833 + 0.05556 + 0.08333 = 1.4331$$

Step 5: Degrees of Freedom and Critical Value

$$df = (r-1)(c-1) = (3-1)(2-1) = 2$$

$$\chi^2_{0.05,,2} = 5.991$$

Step 6: Decision

$$\chi^2_{cal} = 1.433 < \chi^2_{tab} = 5.991$$

Fail to reject $H_0$.

Step 7: Conclusion

At the 5% level of significance, there is no significant association between the duration of marriage and the period of acquaintanceship before marriage. The data do not substantiate an association.

2asked 6xavg 8 marks · Latin Square Design
Answer

What do you understand by design of experiments? Give the layout of a Latin Square Design. Explain why the number of treatments tested in a Latin Square Design should not be less than 3? [5]

Design of Experiments and Latin Square Design

Design of Experiments

Design of experiments refers to the systematic planning and arrangement of experimental conditions so that valid, unbiased, and reliable conclusions can be drawn from the experimental data with minimum experimental error.

According to R.A. Fisher, the design of experiments is based on three fundamental principles:

Three Principles of Design of Experiments

PrincipleDescription
ReplicationExecution of an experiment more than once. It provides a more reliable estimate and helps minimize experimental error (maximize accuracy).
RandomizationTreatments are allocated to various plots in a random manner, giving each treatment an equal chance of being assigned to any experimental unit. It reduces bias.
Local ControlConverts a heterogeneous experimental field into homogeneous blocks (row-wise or column-wise), thereby reducing experimental error.

Latin Square Design (LSD)

LSD is a design used for non-homogeneous experimental material. It applies local control simultaneously along both rows and columns, making it more efficient than RBD. The experimental layout is always square (equal number of rows and columns).

Mathematical Model

$$Y_{ijk} = \mu + \alpha_i + \beta_j + T_k + e_{ijk}$$

Where:

  • $Y_{ijk}$ = observation in $i^{th}$ row and $j^{th}$ column receiving $k^{th}$ treatment
  • $\mu$ = general mean effect
  • $\alpha_i$ = effect due to $i^{th}$ row
  • $\beta_j$ = effect due to $j^{th}$ column
  • $T_k$ = effect due to $k^{th}$ treatment
  • $e_{ijk}$ = random error due to chance

Layout of a Latin Square Design (4x4 Example)

In an LSD with 4 treatments (A, B, C, D), each treatment appears exactly once in each row and exactly once in each column:

$$ \begin{array}{|c|c|c|c|} \hline \textbf{Col 1} & \textbf{Col 2} & \textbf{Col 3} & \textbf{Col 4} \ \hline A & B & C & D \ \hline B & C & D & A \ \hline C & D & A & B \ \hline D & A & B & C \ \hline \end{array} $$

Here, 4 rows, 4 columns, and 4 treatments -- each appearing exactly once per row and per column.


Why the Number of Treatments Should NOT Be Less Than 3

In an LSD with $m$ treatments, the degrees of freedom for error is given by:

$$df_{error} = (m-1)(m-2)$$

If $m = 2$ (only 2 treatments):

$$df_{error} = (2-1)(2-2) = 1 \times 0 = 0$$

  • With zero degrees of freedom for error, it is impossible to estimate the experimental error.
  • Without an error estimate, no F-test can be performed, and no valid statistical conclusions can be drawn.

If $m = 3$ (minimum acceptable):

$$df_{error} = (3-1)(3-2) = 2 \times 1 = 2$$

  • This gives at least 2 degrees of freedom for error, which allows estimation of experimental error and valid hypothesis testing.

Summary Table

Treatments ($m$)$df_{error} = (m-1)(m-2)$Valid?
20No -- error cannot be estimated
32Yes -- minimum acceptable
46Yes
512Yes

Conclusion: The number of treatments in an LSD must be at least 3 so that the degrees of freedom for error is greater than zero, enabling estimation of experimental error and valid F-tests in the ANOVA table.

3asked 5xavg 10 marks · Introduction of multiple linear regression
Answer

Multiple Regression Analysis

Multiple Regression Analysis

Purpose of Multiple Regression Analysis

Multiple regression analysis is applied to:

  • Establish a functional relationship between one dependent variable $Y$ and two or more independent variables $X_1, X_2, \ldots$
  • Predict/estimate the value of the dependent variable from known independent variables
  • Measure the separate effect of each independent variable while holding the others constant
  • Identify which independent variables are significant contributors to the variation in $Y$
  • Provide a more realistic model since real phenomena depend on several factors simultaneously

Given Data

YX1X2
6462
7061
8590
5058
6062
7271
7585
5569
8081
7061

$n = 10$


Part (i): Multiple Regression Equation

Model: $\hat{Y} = b_0 + b_1 X_1 + b_2 X_2$

Step 1: Summations

YX1X2X1²X2²X1·X2X1·YX2·Y
646236412384128
7061361642070
859081007650
5058256440250400
606236412360120
7271491750472
7585642540600375
5569368154330495
8081641864080
7061361642070
681673046318218546731810

$$\sum Y = 681,\ \sum X_1 = 67,\ \sum X_2 = 30,\ \sum X_1^2 = 463,$$ $$\sum X_2^2 = 182,\ \sum X_1 X_2 = 185,\ \sum X_1 Y = 4673,\ \sum X_2 Y = 1810$$

Step 2: Means

$$\bar{Y} = 68.1,\quad \bar{X}_1 = 6.7,\quad \bar{X}_2 = 3.0$$

Step 3: Normal Equations

$$681 = 10b_0 + 67b_1 + 30b_2 \quad (1)$$ $$4673 = 67b_0 + 463b_1 + 185b_2 \quad (2)$$ $$1810 = 30b_0 + 185b_1 + 182b_2 \quad (3)$$

Step 4: Solve

Eliminate $b_0$ using (1) and (2): multiply (1) by $6.7$: $$4562.7 = 67b_0 + 448.9b_1 + 201b_2 \quad (4)$$ $(2) - (4)$: $$110.3 = 14.1b_1 - 16b_2 \quad (5)$$

Eliminate $b_0$ using (1) and (3): multiply (1) by $3$: $$2043 = 30b_0 + 201b_1 + 90b_2 \quad (6)$$ $(3) - (6)$: $$-233 = -16b_1 + 92b_2 \quad (7)$$

From (5): $14.1b_1 - 16b_2 = 110.3$ From (7): $-16b_1 + 92b_2 = -233$

Solve. From (5): $b_1 = \dfrac{110.3 + 16b_2}{14.1}$.

Substitute in (7): $$-16\left(\frac{110.3 + 16b_2}{14.1}\right) + 92b_2 = -233$$

$$-16(110.3 + 16b_2) + 92(14.1)b_2 = -233(14.1)$$

$$-1764.8 - 256b_2 + 1297.2b_2 = -3285.3$$

$$1041.2b_2 = -1520.5$$

$$b_2 = -1.4603$$

Then: $$b_1 = \frac{110.3 + 16(-1.4603)}{14.1} = \frac{110.3 - 23.365}{14.1} = \frac{86.935}{14.1} = 6.1656$$

From (1): $$b_0 = \frac{681 - 67(6.1656) - 30(-1.4603)}{10}$$ $$= \frac{681 - 413.10 + 43.81}{10} = \frac{311.71}{10} = 31.171$$

Regression Equation

$$\boxed{\hat{Y} = 31.171 + 6.166,X_1 - 1.460,X_2}$$


Part (ii): Interpretation of $b_1$ and $b_2$

  • $b_1 = 6.166$: Holding the number of absent days ($X_2$) constant, a one-unit increase in the aptitude test score ($X_1$) raises the job-satisfaction score ($Y$) by about 6.17 units. The positive sign shows a direct relationship.

  • $b_2 = -1.460$: Holding the aptitude test score ($X_1$) constant, each additional day of absence ($X_2$) decreases the job-satisfaction score ($Y$) by about 1.46 units. The negative sign shows an inverse relationship.


Part (iii): Estimate for $X_1 = 9,\ X_2 = 6$

$$\hat{Y} = 31.171 + 6.166(9) - 1.460(6)$$ $$= 31.171 + 55.494 - 8.760$$ $$= \boxed{77.91}$$

The estimated job-satisfaction score is approximately 78.

4asked 4xavg 5 marks · due (skipped 2081) · Markov Process
Answer

Define Markov chain and its characteristics. [5]

Markov Chain and Its Characteristics

Definition of Stochastic Process

A stochastic process is a family of random variables indexed by time as a parameter. It is denoted as {X(t), t ∈ T}, where:

  • T = index set (time parameter)
  • X(t) = random variable at time t
  • I = state space (the set of all possible values assumed by X(t))

The values assumed by the random variable X(t) are called states.


Definition of Markov Chain

A Markov Chain is a stochastic process {X(t), t = 0, 1, 2, ...} that satisfies the Markov property (memoryless property):

$$P[X(t+1) = j \mid X(t) = i,\ X(t-1) = i_{t-1},\ \ldots,\ X(0) = i_0] = P[X(t+1) = j \mid X(t) = i]$$

That is, the future state depends only on the present state, not on the past states.

The probability $P[X(t+1) = j \mid X(t) = i] = p_{ij}$ is called the transition probability from state $i$ to state $j$.


Characteristics of a Markov Chain

1. Markov Property (Memoryless Property)

The conditional probability of any future state depends only on the current state, not on the sequence of states that preceded it. This is the defining characteristic.

2. Transition Probability Matrix (TPM)

The one-step transition probabilities are arranged in a matrix $P = [p_{ij}]$, where:

  • $p_{ij} \geq 0$ for all $i, j$
  • $\sum_{j} p_{ij} = 1$ for each row $i$ (rows sum to 1)

Such a matrix is called a stochastic matrix.

3. Homogeneity (Time Homogeneous)

A Markov chain is said to be homogeneous if the transition probabilities do not depend on time $t$, i.e., $p_{ij}$ remains constant over time.

4. Chapman-Kolmogorov Equation

The $n$-step transition probability satisfies:

$$p_{ij}^{(n)} = \sum_{k} p_{ik}^{(r)} \cdot p_{kj}^{(n-r)}, \quad 0 < r < n$$

In matrix form: $P^{(n)} = P^n$

5. Steady State (Limiting) Distribution

As $n \to \infty$, the $n$-step transition probabilities converge to a steady state distribution $\pi$, such that:

$$\pi_j = \lim_{n \to \infty} P^n(x)$$

The steady state distribution satisfies: $$\pi P = \pi \quad \text{and} \quad \sum_{j} \pi_j = 1$$

6. Irreducibility

A Markov chain is irreducible if every state can be reached from every other state, i.e., all states communicate with each other.

7. Recurrence and Transience

  • A state $i$ is recurrent if the process returns to state $i$ with probability 1.
  • A state $i$ is transient if there is a positive probability of never returning to state $i$.

Summary Table

CharacteristicDescription
Markov PropertyFuture depends only on present
TPMNon-negative entries, rows sum to 1
HomogeneityTransition probs constant over time
Chapman-Kolmogorov$P^{(n)} = P^n$
Steady State$\pi P = \pi$, $\sum \pi_j = 1$
IrreducibilityAll states communicate
5asked 5xavg 5 marks · Estimation
Answer

Define confidence level in estimation. A quality control inspector collected a random sample of 400 tubes of toothpaste from the production line and found that 20 of the tubes had leaks from the tail end. Construct 96% confidence interval for the percentage of all the toothpaste tubes that had leakage and interpret the result. [5]

Confidence Level in Estimation and Confidence Interval for a Proportion

Definition: Confidence Level

The confidence level $(1-\alpha)$ is the probability that the interval estimate constructed from sample data will contain the true population parameter. A 96% confidence level means that if repeated random samples were taken and an interval built from each, about 96% of those intervals would enclose the true population proportion.


Given Data

ItemValue
Sample size$n = 400$
Leaking tubes$X = 20$
Confidence level$1-\alpha = 0.96$
Significance level$\alpha = 0.04$

Step 1: Sample Proportion

$$p = \frac{X}{n} = \frac{20}{400} = 0.05, \qquad q = 1 - p = 0.95$$

Step 2: Critical Value

$$\frac{\alpha}{2} = 0.02 \Rightarrow Z_{\alpha/2} = 2.054 \approx 2.05$$

Step 3: Standard Error

$$S.E.(p) = \sqrt{\frac{pq}{n}} = \sqrt{\frac{0.05 \times 0.95}{400}} = \sqrt{\frac{0.0475}{400}} = \sqrt{0.00011875} = 0.010897$$

Step 4: Confidence Interval

$$p \pm Z_{\alpha/2}\cdot S.E.(p)$$

Margin of error: $$E = 2.054 \times 0.010897 = 0.02238$$

Lower limit: $$0.05 - 0.02238 = 0.02762$$

Upper limit: $$0.05 + 0.02238 = 0.07238$$

Step 5: Express as Percentage

$$\boxed{2.76% \leq P \leq 7.24%}$$

(Using $Z = 2.05$: $2.77%$ to $7.24%$; essentially the same.)


Interpretation

We are 96% confident that the true percentage of all toothpaste tubes with leakage from the tail end lies between approximately 2.76% and 7.24%. If the sampling procedure were repeated many times, about 96% of the resulting intervals would contain the true population proportion of leaking tubes.

Most repeated questions

Topics asked at least twice, most-asked first.

asked 8xavg 5 marks · 2081, 2080, 2079, 2078, 2077...
Answer

Chi-Square Test of Independence

Based on interviews of couples seeking divorce, a social worker compiles the following data related to the period of acquaintanceship before marriage and the duration of marriage. Perform a test to determine if the data substantiate an association between the duration of a marriage and the acquaintanceship prior to marriage. Use 5% level of significance.

Acquaintanceship before marriageLess than or equal to 5 yearsMore than 5 yearsTotal
Below 0.5 year15722
0.5–1.5 years262248
Over 1.5 years191130
Total6040100

[5]

Chi-Square Test for Association

Step 1: Given Data

Observed contingency table:

Acquaintanceship≤ 5 years> 5 yearsTotal
Below 0.5 yr15722
0.5-1.5 yr262248
Over 1.5 yr191130
Total6040100

Significance level $\alpha = 0.05$.

Step 2: Hypotheses

  • $H_0$: No association between duration of marriage and acquaintanceship period (independent).
  • $H_1$: There is an association.

Step 3: Expected Frequencies

$$E_{ij} = \frac{R_i \times C_j}{N}$$

Cell$E$
(Below 0.5, ≤5)$\frac{22 \times 60}{100} = 13.2$
(Below 0.5, >5)$\frac{22 \times 40}{100} = 8.8$
(0.5-1.5, ≤5)$\frac{48 \times 60}{100} = 28.8$
(0.5-1.5, >5)$\frac{48 \times 40}{100} = 19.2$
(Over 1.5, ≤5)$\frac{30 \times 60}{100} = 18.0$
(Over 1.5, >5)$\frac{30 \times 40}{100} = 12.0$

Step 4: Chi-Square Statistic

$$\chi^2 = \sum \frac{(O-E)^2}{E}$$

OE$(O-E)^2$$(O-E)^2/E$
1513.23.240.24545
78.83.240.36818
2628.87.840.27222
2219.27.840.40833
1918.01.000.05556
1112.01.000.08333

$$\chi^2_{cal} = 0.24545 + 0.36818 + 0.27222 + 0.40833 + 0.05556 + 0.08333 = 1.4331$$

Step 5: Degrees of Freedom and Critical Value

$$df = (r-1)(c-1) = (3-1)(2-1) = 2$$

$$\chi^2_{0.05,,2} = 5.991$$

Step 6: Decision

$$\chi^2_{cal} = 1.433 < \chi^2_{tab} = 5.991$$

Fail to reject $H_0$.

Step 7: Conclusion

At the 5% level of significance, there is no significant association between the duration of marriage and the period of acquaintanceship before marriage. The data do not substantiate an association.

asked 6xavg 8 marks · 2081, 2080, 2078, 2077, 2075
Answer

What do you understand by design of experiments? Give the layout of a Latin Square Design. Explain why the number of treatments tested in a Latin Square Design should not be less than 3? [5]

Design of Experiments and Latin Square Design

Design of Experiments

Design of experiments refers to the systematic planning and arrangement of experimental conditions so that valid, unbiased, and reliable conclusions can be drawn from the experimental data with minimum experimental error.

According to R.A. Fisher, the design of experiments is based on three fundamental principles:

Three Principles of Design of Experiments

PrincipleDescription
ReplicationExecution of an experiment more than once. It provides a more reliable estimate and helps minimize experimental error (maximize accuracy).
RandomizationTreatments are allocated to various plots in a random manner, giving each treatment an equal chance of being assigned to any experimental unit. It reduces bias.
Local ControlConverts a heterogeneous experimental field into homogeneous blocks (row-wise or column-wise), thereby reducing experimental error.

Latin Square Design (LSD)

LSD is a design used for non-homogeneous experimental material. It applies local control simultaneously along both rows and columns, making it more efficient than RBD. The experimental layout is always square (equal number of rows and columns).

Mathematical Model

$$Y_{ijk} = \mu + \alpha_i + \beta_j + T_k + e_{ijk}$$

Where:

  • $Y_{ijk}$ = observation in $i^{th}$ row and $j^{th}$ column receiving $k^{th}$ treatment
  • $\mu$ = general mean effect
  • $\alpha_i$ = effect due to $i^{th}$ row
  • $\beta_j$ = effect due to $j^{th}$ column
  • $T_k$ = effect due to $k^{th}$ treatment
  • $e_{ijk}$ = random error due to chance

Layout of a Latin Square Design (4x4 Example)

In an LSD with 4 treatments (A, B, C, D), each treatment appears exactly once in each row and exactly once in each column:

$$ \begin{array}{|c|c|c|c|} \hline \textbf{Col 1} & \textbf{Col 2} & \textbf{Col 3} & \textbf{Col 4} \ \hline A & B & C & D \ \hline B & C & D & A \ \hline C & D & A & B \ \hline D & A & B & C \ \hline \end{array} $$

Here, 4 rows, 4 columns, and 4 treatments -- each appearing exactly once per row and per column.


Why the Number of Treatments Should NOT Be Less Than 3

In an LSD with $m$ treatments, the degrees of freedom for error is given by:

$$df_{error} = (m-1)(m-2)$$

If $m = 2$ (only 2 treatments):

$$df_{error} = (2-1)(2-2) = 1 \times 0 = 0$$

  • With zero degrees of freedom for error, it is impossible to estimate the experimental error.
  • Without an error estimate, no F-test can be performed, and no valid statistical conclusions can be drawn.

If $m = 3$ (minimum acceptable):

$$df_{error} = (3-1)(3-2) = 2 \times 1 = 2$$

  • This gives at least 2 degrees of freedom for error, which allows estimation of experimental error and valid hypothesis testing.

Summary Table

Treatments ($m$)$df_{error} = (m-1)(m-2)$Valid?
20No -- error cannot be estimated
32Yes -- minimum acceptable
46Yes
512Yes

Conclusion: The number of treatments in an LSD must be at least 3 so that the degrees of freedom for error is greater than zero, enabling estimation of experimental error and valid F-tests in the ANOVA table.

asked 5xavg 10 marks · 2081, 2079, 2078, 2077, 2075
Answer

Multiple Regression Analysis

Multiple Regression Analysis

Purpose of Multiple Regression Analysis

Multiple regression analysis is applied to:

  • Establish a functional relationship between one dependent variable $Y$ and two or more independent variables $X_1, X_2, \ldots$
  • Predict/estimate the value of the dependent variable from known independent variables
  • Measure the separate effect of each independent variable while holding the others constant
  • Identify which independent variables are significant contributors to the variation in $Y$
  • Provide a more realistic model since real phenomena depend on several factors simultaneously

Given Data

YX1X2
6462
7061
8590
5058
6062
7271
7585
5569
8081
7061

$n = 10$


Part (i): Multiple Regression Equation

Model: $\hat{Y} = b_0 + b_1 X_1 + b_2 X_2$

Step 1: Summations

YX1X2X1²X2²X1·X2X1·YX2·Y
646236412384128
7061361642070
859081007650
5058256440250400
606236412360120
7271491750472
7585642540600375
5569368154330495
8081641864080
7061361642070
681673046318218546731810

$$\sum Y = 681,\ \sum X_1 = 67,\ \sum X_2 = 30,\ \sum X_1^2 = 463,$$ $$\sum X_2^2 = 182,\ \sum X_1 X_2 = 185,\ \sum X_1 Y = 4673,\ \sum X_2 Y = 1810$$

Step 2: Means

$$\bar{Y} = 68.1,\quad \bar{X}_1 = 6.7,\quad \bar{X}_2 = 3.0$$

Step 3: Normal Equations

$$681 = 10b_0 + 67b_1 + 30b_2 \quad (1)$$ $$4673 = 67b_0 + 463b_1 + 185b_2 \quad (2)$$ $$1810 = 30b_0 + 185b_1 + 182b_2 \quad (3)$$

Step 4: Solve

Eliminate $b_0$ using (1) and (2): multiply (1) by $6.7$: $$4562.7 = 67b_0 + 448.9b_1 + 201b_2 \quad (4)$$ $(2) - (4)$: $$110.3 = 14.1b_1 - 16b_2 \quad (5)$$

Eliminate $b_0$ using (1) and (3): multiply (1) by $3$: $$2043 = 30b_0 + 201b_1 + 90b_2 \quad (6)$$ $(3) - (6)$: $$-233 = -16b_1 + 92b_2 \quad (7)$$

From (5): $14.1b_1 - 16b_2 = 110.3$ From (7): $-16b_1 + 92b_2 = -233$

Solve. From (5): $b_1 = \dfrac{110.3 + 16b_2}{14.1}$.

Substitute in (7): $$-16\left(\frac{110.3 + 16b_2}{14.1}\right) + 92b_2 = -233$$

$$-16(110.3 + 16b_2) + 92(14.1)b_2 = -233(14.1)$$

$$-1764.8 - 256b_2 + 1297.2b_2 = -3285.3$$

$$1041.2b_2 = -1520.5$$

$$b_2 = -1.4603$$

Then: $$b_1 = \frac{110.3 + 16(-1.4603)}{14.1} = \frac{110.3 - 23.365}{14.1} = \frac{86.935}{14.1} = 6.1656$$

From (1): $$b_0 = \frac{681 - 67(6.1656) - 30(-1.4603)}{10}$$ $$= \frac{681 - 413.10 + 43.81}{10} = \frac{311.71}{10} = 31.171$$

Regression Equation

$$\boxed{\hat{Y} = 31.171 + 6.166,X_1 - 1.460,X_2}$$


Part (ii): Interpretation of $b_1$ and $b_2$

  • $b_1 = 6.166$: Holding the number of absent days ($X_2$) constant, a one-unit increase in the aptitude test score ($X_1$) raises the job-satisfaction score ($Y$) by about 6.17 units. The positive sign shows a direct relationship.

  • $b_2 = -1.460$: Holding the aptitude test score ($X_1$) constant, each additional day of absence ($X_2$) decreases the job-satisfaction score ($Y$) by about 1.46 units. The negative sign shows an inverse relationship.


Part (iii): Estimate for $X_1 = 9,\ X_2 = 6$

$$\hat{Y} = 31.171 + 6.166(9) - 1.460(6)$$ $$= 31.171 + 55.494 - 8.760$$ $$= \boxed{77.91}$$

The estimated job-satisfaction score is approximately 78.

asked 5xavg 5 marks · 2081, 2079, 2078, 2077, 2075
Answer

Define confidence level in estimation. A quality control inspector collected a random sample of 400 tubes of toothpaste from the production line and found that 20 of the tubes had leaks from the tail end. Construct 96% confidence interval for the percentage of all the toothpaste tubes that had leakage and interpret the result. [5]

Confidence Level in Estimation and Confidence Interval for a Proportion

Definition: Confidence Level

The confidence level $(1-\alpha)$ is the probability that the interval estimate constructed from sample data will contain the true population parameter. A 96% confidence level means that if repeated random samples were taken and an interval built from each, about 96% of those intervals would enclose the true population proportion.


Given Data

ItemValue
Sample size$n = 400$
Leaking tubes$X = 20$
Confidence level$1-\alpha = 0.96$
Significance level$\alpha = 0.04$

Step 1: Sample Proportion

$$p = \frac{X}{n} = \frac{20}{400} = 0.05, \qquad q = 1 - p = 0.95$$

Step 2: Critical Value

$$\frac{\alpha}{2} = 0.02 \Rightarrow Z_{\alpha/2} = 2.054 \approx 2.05$$

Step 3: Standard Error

$$S.E.(p) = \sqrt{\frac{pq}{n}} = \sqrt{\frac{0.05 \times 0.95}{400}} = \sqrt{\frac{0.0475}{400}} = \sqrt{0.00011875} = 0.010897$$

Step 4: Confidence Interval

$$p \pm Z_{\alpha/2}\cdot S.E.(p)$$

Margin of error: $$E = 2.054 \times 0.010897 = 0.02238$$

Lower limit: $$0.05 - 0.02238 = 0.02762$$

Upper limit: $$0.05 + 0.02238 = 0.07238$$

Step 5: Express as Percentage

$$\boxed{2.76% \leq P \leq 7.24%}$$

(Using $Z = 2.05$: $2.77%$ to $7.24%$; essentially the same.)


Interpretation

We are 96% confident that the true percentage of all toothpaste tubes with leakage from the tail end lies between approximately 2.76% and 7.24%. If the sampling procedure were repeated many times, about 96% of the resulting intervals would contain the true population proportion of leaking tubes.

asked 4xavg 5 marks · 2079, 2078, 2077
Answer

Define Markov chain and its characteristics. [5]

Markov Chain and Its Characteristics

Definition of Stochastic Process

A stochastic process is a family of random variables indexed by time as a parameter. It is denoted as {X(t), t ∈ T}, where:

  • T = index set (time parameter)
  • X(t) = random variable at time t
  • I = state space (the set of all possible values assumed by X(t))

The values assumed by the random variable X(t) are called states.


Definition of Markov Chain

A Markov Chain is a stochastic process {X(t), t = 0, 1, 2, ...} that satisfies the Markov property (memoryless property):

$$P[X(t+1) = j \mid X(t) = i,\ X(t-1) = i_{t-1},\ \ldots,\ X(0) = i_0] = P[X(t+1) = j \mid X(t) = i]$$

That is, the future state depends only on the present state, not on the past states.

The probability $P[X(t+1) = j \mid X(t) = i] = p_{ij}$ is called the transition probability from state $i$ to state $j$.


Characteristics of a Markov Chain

1. Markov Property (Memoryless Property)

The conditional probability of any future state depends only on the current state, not on the sequence of states that preceded it. This is the defining characteristic.

2. Transition Probability Matrix (TPM)

The one-step transition probabilities are arranged in a matrix $P = [p_{ij}]$, where:

  • $p_{ij} \geq 0$ for all $i, j$
  • $\sum_{j} p_{ij} = 1$ for each row $i$ (rows sum to 1)

Such a matrix is called a stochastic matrix.

3. Homogeneity (Time Homogeneous)

A Markov chain is said to be homogeneous if the transition probabilities do not depend on time $t$, i.e., $p_{ij}$ remains constant over time.

4. Chapman-Kolmogorov Equation

The $n$-step transition probability satisfies:

$$p_{ij}^{(n)} = \sum_{k} p_{ik}^{(r)} \cdot p_{kj}^{(n-r)}, \quad 0 < r < n$$

In matrix form: $P^{(n)} = P^n$

5. Steady State (Limiting) Distribution

As $n \to \infty$, the $n$-step transition probabilities converge to a steady state distribution $\pi$, such that:

$$\pi_j = \lim_{n \to \infty} P^n(x)$$

The steady state distribution satisfies: $$\pi P = \pi \quad \text{and} \quad \sum_{j} \pi_j = 1$$

6. Irreducibility

A Markov chain is irreducible if every state can be reached from every other state, i.e., all states communicate with each other.

7. Recurrence and Transience

  • A state $i$ is recurrent if the process returns to state $i$ with probability 1.
  • A state $i$ is transient if there is a positive probability of never returning to state $i$.

Summary Table

CharacteristicDescription
Markov PropertyFuture depends only on present
TPMNon-negative entries, rows sum to 1
HomogeneityTransition probs constant over time
Chapman-Kolmogorov$P^{(n)} = P^n$
Steady State$\pi P = \pi$, $\sum \pi_j = 1$
IrreducibilityAll states communicate
asked 4xavg 6 marks · 2081, 2080, 2079, 2075
Answer

Multiple Regression Analysis Question

A multiple regression equation yields the following results:

SourceSum of squareDegree of freedom
Regression7402
Error51017

i) What is the total sample size?

ii) How many independent variables are being considered?

iii) Compute the coefficient of determination and interpret its value.

iv) Compute the standard error of estimate.

v) Test the hypothesis that the overall fit of the model is significant or not. Assume $\alpha=0.05$.

[5]

Source Sum of Squares (SS) Degrees of Freedom (df) --------------------------------------------------- Regression 740 2 Error 510 17 Total 1250 19 $\alpha = 0.05$ --- $$df{Total} = df{Reg} + df{Error} = 2 + 17 = 19$$ Since $df{Total} = n...

asked 3xavg 5 marks · 2080, 2077, 2075
Answer

What do you understand by estimation? If we want to determine average mechanical aptitude of a large group of workers, how large a random sample is needed to be able to assert with probability 0.95 that the sample mean will not differ from the true mean by more than 2.0 points? Assume that population standard deviation is 30. [5]

Parameter Value ------------------ Confidence level $(1-\alpha)$ 0.95 Maximum allowable error $(E)$ 2.0 points Population standard deviation $(\sigma)$ 30 Critical value $Z{\alpha/2}$ at 95% 1.96 Estimation is the statistical procedure o...

asked 3xavg 7 marks · 2078, 2075
Answer

The following are the details of working hours in the classroom per week of male and female faculty working in the area of Computer Science and Information Technology at Tribhuvan University. Apply independent t-test to examine the average working hour in the classroom per week is significantly different between male and female faculty, at 1% level of significance. State also null and alternative hypotheses appropriately.

Male FacultyFemale Faculty
Sample Size6030
Average working hours per week129
The standard deviation of a working hour per week43

[5]

Male Faculty Female Faculty --------- Sample size $n1 = 60$ $n2 = 30$ Mean working hours $\bar{X}1 = 12$ $\bar{X}2 = 9$ Standard deviation $S1 = 4$ $S2 = 3$ Level of significance: $\alpha = 0.01$ (two-tailed) --- Null Hypothesis $H0: \mu...

asked 3xavg 10 marks · 2081, 2079, 2078
Answer

One-Way ANOVA: Seminar Evaluation Scores by Manager Level

A management consulting company presents a 3-day seminar on project management to various clients. The seminar is basically the same each time it is given. However, sometimes it is presented to high-level managers, sometimes to mid-level managers, and sometimes to low-level managers. The seminar facilitators believe evaluations of the seminar may vary with the audience. Suppose the following data are some randomly selected evaluation scores from different levels of managers after they have attended the seminar. The ratings are on a scale from 1 to 100, with 100 being the highest score. Use a one-way ANOVA to determine whether there is a significant difference in the evaluation according to manager level. The following table gives the scores to the various clients due to different manager levels.

$$\begin{array}{|c|c|c|c|} \hline & \text{High Level} & \text{Mid Level} & \text{Low Level} \ \hline & 85 & 90 & 55 \ & 75 & 100 & 75 \ & 85 & 95 & 80 \ & 60 & 85 & 75 \ & 70 & 90 & \ & & 75 & \ \hline \end{array}$$

[10]

High Level Mid Level Low Level :-::-::-: 85 90 55 75 100 75 85 95 80 60 85 75 70 90 75 - $n1 = 5,\ n2 = 6,\ n3 = 4$ - Total $N = 15$, groups $k = 3$ - $\alpha = 0.05$ - $H0: \mu1 = \mu2 = \mu3$ (no significant difference) - $H1:$ at leas...

asked 3xavg 5 marks · 2081, 2078, 2077
Answer

The mean drying time of a brand of spray paint is known to be 122 seconds. The research division of the company that produces this paint contemplates that adding a new chemical ingredient to the paint accelerate the drying process. To investigate this conjecture, the paint with the chemical additions is sprayed on 50 surfaces and the drying time is recorded. The mean and standard deviation of drying time computed from these recorded are found as 116 seconds and 16.8 seconds respectively. Does these data provide strong evidence that the mean drying time is reduced by the addition of the new chemical? Use 5% level of significance. Also find p-value. [5]

Hypothesis Test: Mean Drying Time of Spray Paint

Given Data

  • Population (claimed) mean: $\mu_0 = 122$ seconds
  • Sample size: $n = 50$
  • Sample mean: $\bar{X} = 116$ seconds
  • Sample standard deviation: $s = 16.8$ seconds
  • Significance level: $\alpha = 0.05$

Step 1: Hypotheses

  • $H_0: \mu = 122$ (no reduction)
  • $H_1: \mu < 122$ (drying time reduced) - left-tailed test

Step 2: Test Statistic

Since $n = 50 > 30$, use the Z-test: $$Z = \frac{\bar{X} - \mu_0}{s/\sqrt{n}}$$

Step 3: Calculation

$$s/\sqrt{n} = \frac{16.8}{\sqrt{50}} = \frac{16.8}{7.0711} = 2.3759$$

$$Z = \frac{116 - 122}{2.3759} = \frac{-6}{2.3759} = -2.525$$

Step 4: Critical Value

Left-tailed at $\alpha = 0.05$: $Z_{tab} = -1.645$

Step 5: Decision

Since $Z_{cal} = -2.525 < -1.645$, we reject $H_0$.

Step 6: P-value

$$p = P(Z < -2.525) = 1 - P(Z < 2.525) = 1 - 0.9942 = 0.0058$$

Since $p = 0.0058 < \alpha = 0.05$, reject $H_0$.

Conclusion

The data provide strong evidence at the 5% level of significance that the new chemical additive reduces the mean drying time below 122 seconds. The small p-value ($\approx 0.0058$) confirms high significance.

asked 3xavg 5 marks · 2081, 2080, 2079
Answer

Explain the concept of multiple and partial correlation coefficients. Consider three variables X1, X2 and X3. If $r_{12}=0.40$, $r_{23}=0.50$ and $r_{13}=0.6$ find $R_{123}$ and $r_{23.1}$. [5]

Three variables $X1, X2, X3$ with zero-order correlation coefficients: $$r{12} = 0.40, \quad r{23} = 0.50, \quad r{13} = 0.60$$ Note on notation: The problem statement contains a typo. "$R{123}$" is standardly interpreted as the multiple...

asked 3xavg 5 marks · 2081, 2080, 2079
Answer

Discuss the concept of level of significance in hypothesis testing. A manufacturer of laptop provides a particular model in one of three colors. Of the first 100 laptops sold, it is noted that 80 were the first color. Can you conclude that more than two third of all the customers have a preference for the first color? Use 5% level of significance. [5]

Level of Significance and Hypothesis Testing

Concept of Level of Significance

The level of significance (denoted α) is the probability of rejecting the null hypothesis $H_0$ when it is actually true. It is the maximum acceptable probability of committing a Type I error.

  • Fixed in advance of the test (typically $1%$, $5%$, or $10%$).
  • At $5%$ significance, we accept a $5%$ chance of wrongly rejecting a true $H_0$.
  • It determines the critical (tabulated) value used to bound the acceptance/rejection regions.
  • A smaller $\alpha$ makes the test more conservative (harder to reject $H_0$).

Numerical Solution

Given Data

ParameterValue
Sample size $n$$100$
Number preferring first color $x$$80$
Hypothesized proportion $P$$2/3 \approx 0.6667$
Level of significance $\alpha$$0.05$

Sample proportion: $$\hat{p} = \frac{x}{n} = \frac{80}{100} = 0.80$$


Step 1: Hypotheses

  • $H_0: P = \dfrac{2}{3}$ (two-thirds prefer the first color)
  • $H_1: P > \dfrac{2}{3}$ (more than two-thirds prefer the first color)

This is a right-tailed test.


Step 2: Test Statistic

$$Z = \frac{\hat{p} - P}{\sqrt{\dfrac{P(1-P)}{n}}}$$

Numerator: $$\hat{p} - P = 0.80 - 0.6667 = 0.1333$$

Denominator: $$\sqrt{\frac{0.6667 \times 0.3333}{100}} = \sqrt{\frac{0.2222}{100}} = \sqrt{0.002222} = 0.04714$$

Therefore: $$Z = \frac{0.1333}{0.04714} = 2.83$$


Step 3: Critical Value

For a right-tailed test at $\alpha = 0.05$: $$Z_{tab} = 1.645$$


Step 4: Decision

$$Z_{cal} = 2.83 > Z_{tab} = 1.645$$

Since the calculated value falls in the rejection region, reject $H_0$.


Step 5: Conclusion

At the $5%$ level of significance, there is sufficient evidence to conclude that more than two-thirds of all customers prefer the first color of the laptop.

asked 2xavg 8 marks · 2080, 2075
Answer

Kruskal-Wallis Test for Propellant Burning Rates

In an experiment to determine which of three different missile systems is preferable, the propellant burning rate is measured. The data after coding are given in the table. Use Kruskal-Wallis test (significance level of 0.01) to test the hypothesis that the propellant burning rates are same for three missile systems.

1234567
Missile system I22.316.722.719.318.5
Missile system II23.419.517.520.816.019.9
Missile system III18.419.517.818.019.622.817.1

[10]

Kruskal-Wallis Test for Missile System Propellant Burning Rates

Step 1: Given Data

System I (n₁=5)System II (n₂=6)System III (n₃=7)
22.323.418.4
16.719.519.5
22.717.517.8
19.320.818.0
18.516.019.6
19.922.8
17.1

Total $N = 5 + 6 + 7 = 18$

Step 2: Hypotheses

  • $H_0$: The propellant burning rates are the same for all three missile systems.
  • $H_1$: At least one system differs.

Step 3: Rank All 18 Observations (ascending)

ValueSystemRank
16.0II1
16.7I2
17.1III3
17.5II4
17.8III5
18.0III6
18.4III7
18.5I8
19.3I9
19.5II10.5
19.5III10.5
19.6III12
19.9II13
20.8II14
22.3I15
22.7I16
22.8III17
23.4II18

Tie at 19.5: average rank $= (10+11)/2 = 10.5$

Step 4: Rank Sums

System I: $R_1 = 2 + 8 + 9 + 15 + 16 = 50$

System II: $R_2 = 1 + 4 + 10.5 + 13 + 14 + 18 = 60.5$

System III: $R_3 = 3 + 5 + 6 + 7 + 10.5 + 12 + 17 = 60.5$

Check: $50 + 60.5 + 60.5 = 171 = \dfrac{N(N+1)}{2} = \dfrac{18(19)}{2} = 171$ ✓

Step 5: Test Statistic

$$H = \frac{12}{N(N+1)} \sum \frac{R_i^2}{n_i} - 3(N+1)$$

$$\frac{50^2}{5} = 500, \quad \frac{60.5^2}{6} = \frac{3660.25}{6} = 610.042, \quad \frac{60.5^2}{7} = \frac{3660.25}{7} = 522.893$$

Sum $= 500 + 610.042 + 522.893 = 1632.935$

$$H = \frac{12}{18 \times 19}(1632.935) - 3(19) = \frac{12}{342}(1632.935) - 57$$

$$H = 0.035088 \times 1632.935 - 57 = 57.296 - 57 = 0.296 \approx 0.30$$

Correction for ties (optional, one tie group of size 2):

$$C = 1 - \frac{\sum(t^3 - t)}{N^3 - N} = 1 - \frac{2^3 - 2}{18^3 - 18} = 1 - \frac{6}{5814} = 1 - 0.001032 = 0.99897$$

$$H_{corrected} = \frac{0.296}{0.99897} \approx 0.296$$

The tie correction is negligible.

Step 6: Critical Value

$H$ follows $\chi^2$ with $df = k - 1 = 2$.

$$\chi^2_{0.01, 2} = 9.210$$

Step 7: Decision

$$H = 0.30 < 9.210$$

Fail to reject $H_0$.

Conclusion

At the 1% significance level, there is insufficient evidence to conclude that the propellant burning rates differ among the three missile systems. The burning rates may be considered the same.

asked 2xavg 5 marks · 2080, 2079
Answer

Write short notes on any two: a. Difference between parametric and non-parametric test. b. Required assumptions for linear regression model. c. Stochastic process. [5]

Short Notes on Any Two


a. Difference Between Parametric and Non-Parametric Test

BasisParametric TestNon-Parametric Test
AssumptionRequires strict assumptions about population distribution (usually normality)Does not require assumptions about population distribution
Data TypeWorks on interval or ratio scale dataWorks on nominal or ordinal scale data
Population ParameterMakes inferences about population parameters (mean, variance)Makes inferences without reference to specific population parameters
Sample SizeRequires relatively large sample sizeCan be used with small sample sizes
PowerMore powerful when assumptions are metLess powerful compared to parametric tests
Examplest-test, F-test, Z-testMann-Whitney U-test, Median test, Kruskal-Wallis test

**Note ** The Mann-Whitney U-test is a non-parametric test used as an alternative to the T-test when T-test fails to satisfy the normality assumptions. This clearly illustrates the key distinction: parametric tests demand distributional assumptions, while non-parametric tests do not.


b. Required Assumptions for Linear Regression Model

For a linear regression model to produce valid and reliable results, the following assumptions must be satisfied:

1. Linearity The relationship between the dependent variable (Y) and independent variable(s) (X) must be linear in nature.

2. Independence of Errors The error terms (residuals) must be independent of each other. There should be no autocorrelation among residuals.

3. Normality of Errors The error terms should follow a normal distribution with mean zero:

$$E(\varepsilon_i) = 0$$

4. Homoscedasticity (Constant Variance) The variance of error terms should be constant across all levels of the independent variable:

$$Var(\varepsilon_i) = \sigma^2 \text{ (constant)}$$

5. No Multicollinearity (for multiple regression) In multiple regression, the independent variables (X1, X2, ...) should not be highly correlated with each other. In multiple regression with variables X1 and X2, each variable is tested separately for its contribution, implying they should be independently contributing.

6. No Measurement Error The independent variables are assumed to be measured without error.

7. Significance of Regression Coefficients The regression coefficients (b1, b2, ...) should be significantly different from zero, tested using the t-test:

$$t_{cal} = \frac{b_i}{S_{b_i}}$$

**Example ** With n=20, b1=4, Sb1=1.2, the calculated t = 3.33 > t-tab = 2.110, confirming a significant linear relationship between Y and X1, validating the regression model assumption.


asked 2xavg 10 marks · 2078, 2077
Answer

Explain the sample distribution of mean with reference to some numerical example. Illustrate the practical implications of the Central Limit Theorem (CLT) in inferential statistics.[10]

When all possible random samples of size n are drawn from a population of size N, and the mean of each sample is computed, the probability distribution formed by these sample means is called the sampling distribution of the sample mean. ...

Study every one of these with model answers, flashcards, and MCQs.

Open STA215 study modes