2079

STA215 · TU past paper

Statistics II 2079 question paper

The complete TU 2079 exam paper for Statistics II (STA215), all 12 questions with solved model answers written to the mark scheme.

Tap a question to open its answer.

  1. 110 marksNumericalIntroduction of multiple linear regressionAnswer

    Multiple Regression Analysis: IRS Unpaid Taxes Estimation

    For the multiple regression model $Y = \beta0 + \beta1 X1 + \beta2 X2 + \epsilon$, the error term $\epsilon$ must satisfy: 1. Zero mean: $E(\epsiloni) = 0$ for all $i$. 2. Constant variance (homoscedasticity): $Var(\epsiloni) = \sigma^2$...

  2. 210 marksNumericalCompletely Randomized DesignAnswer

    Design of Experiments Question

    What do you understand by 'Design of an Experiment'? Physicians depend on laboratory test results when managing medical problems such as diabetes or epilepsy. In a uniformity test of glucose tolerance, three different laboratories were sent $n_t = 5$ identical blood samples from a person who had drunk 50 mg of glucose dissolved in water. The laboratory results are listed below: Do data indicate a difference in the average readings for three laboratories? Use $\alpha = 0.05$.

    Lab 1Lab 2Lab 3
    12.19.310.0
    11.711.110.5
    10.910.710.1
    10.210.911.0
    10.69.010.4

    [10]

    Design of an experiment is the systematic planning of an investigation so that valid, unbiased, and reliable conclusions can be drawn from collected data. It specifies how treatments are allocated to experimental units, how many replicat...

  3. 310 marksNumericalpaired sample t-testAnswer

    Hypothesis Testing: Type I and Type II Errors

    Define Type I and Type II error in testing of hypothesis. A psychologist wishes to verify that a certain drug increases the reaction time to given stimulus. The following reaction times (in tenth of seconds) were recorded before and after injection of the drug for each of four subjects. Test at 5% level of significance to determine whether the drug significantly increases the reaction time.

    Reaction TimeSubject 1Subject 2Subject 3Subject 4
    Before721212
    After1331813

    [10]

    Type I Error (α): Rejecting a null hypothesis $H0$ when it is actually true. Its probability equals the level of significance, denoted $\alpha$. Type II Error (β): Failing to reject (accepting) a null hypothesis $H0$ when it is actually ...

  4. 45 marksNumericalTest of significance of regressionAnswer

    ANOVA Analysis for Multiple Regression Model

    Source SS df MSS F --------------- Regression 12.62 2 ? ? Error 0.78 12 ? Total 13.40 14 - Number of independent variables: $k = 2$ - Total $df = n - 1 = 14 \Rightarrow n = 15$ --- MSS due to Regression: $$MSR = \frac{SSR}{dfR} = \frac{1...

  5. 55 marksParametric vs. non-parametric testAnswer

    What do you mean by non parametric test? Write down advantages of non parametric test over parametric tests? [5]

    A non-parametric test (also called a distribution-free test) is a statistical test that does not require any assumption about the population distribution from which the sample is drawn. Unlike parametric tests, these tests do not involve...

  6. 65 marksNumericalOne-sample testAnswer

    Bank of Nepal recorded the sex of first 30 customers who appeared last Monday with notation M M F M M F M F M F M F M F M F M F M F M F M F M F M F M F. At the 0.005 level of significance, test the randomness of this sequence. [5]

    Sequence (30 customers): M M F M M F M F M F M F M F M F M F M F M F M F M F M F M F - Total observations $n = 30$ - Level of significance $\alpha = 0.005$ --- Let me count carefully. Positions 1-30: 1:M 2:M 3:F 4:M 5:M 6:F 7:M 8:F 9:M 1...

  7. 75 marksNumericalTwo independent sample testAnswer

    A study showed the following results for the different age groups. At the 0.05 level of significance, is there evidence of a difference among the age groups with respect to use of mobile phones for accessing social networking?

    $$\begin{array}{|c|c|c|c|} \hline \text{Use mobile phones to access social networking?} & \text{18-34} & \text{35-64} & \text{65+} \ \hline \text{Yes} & 60 & 37 & 14 \ \text{No} & 40 & 63 & 86 \ \hline \end{array}$$

    [5]

    18-34 35-64 65+ Row Total --------------- Yes 60 37 14 111 No 40 63 86 189 Col Total 100 100 100 300 Level of significance: $\alpha = 0.05$ - $H0$: There is no difference among age groups regarding use of mobile phones for social network...

  8. 85 marksNumericaltest for single proportionAnswer

    It is claimed that Samsung and Redmi mobiles are equally popular in Kathmandu. A random sample of 500 people from Kathmandu showed 300 have Samsung mobile. Test the claim at 5% level of significance. [5]

    • Sample size: $n = 500$ - Samsung users: $X = 300$ - Level of significance: $\alpha = 0.05$ - Claimed proportion (equal popularity): $p0 = 0.5$ --- Sample proportion: $$\hat{p} = \frac{X}{n} = \frac{300}{500} = 0.60$$ - $H0: p = 0.5$ (S...
  9. 95 marksNumericalEstimationAnswer

    An effort to estimate the mean amount per customer for dinner at a major Atlanta restaurant, data were collected for a sample of 49 customers and sample mean is found at 24.80. Assume population standard deviation is 5. a. Compute standard error of mean. b. Find 95% confidence interval estimate for the population mean. [5]

    Confidence Interval Estimate for Population Mean

    STEP 1 - Given Data

    ParameterValue
    Sample size ($n$)49
    Sample mean ($\bar{x}$)24.80
    Population standard deviation ($\sigma$)5
    Confidence level95%

    Since $\sigma$ is known, we use the Z-distribution.


    STEP 2 - Solution

    Part (a): Standard Error of the Mean

    $$SE = \frac{\sigma}{\sqrt{n}} = \frac{5}{\sqrt{49}} = \frac{5}{7} = 0.7143$$

    $$\boxed{SE \approx 0.714}$$


    Part (b): 95% Confidence Interval

    Step 1: Z value

    For 95% confidence:

    • $\alpha = 0.05$, so $\alpha/2 = 0.025$
    • $Z_{0.025} = 1.96$

    Step 2: Margin of error

    $$E = Z_{\alpha/2} \cdot \frac{\sigma}{\sqrt{n}} = 1.96 \times 0.7143 = 1.4000$$

    Step 3: Confidence interval

    $$\text{Lower} = \bar{x} - E = 24.80 - 1.40 = 23.40$$

    $$\text{Upper} = \bar{x} + E = 24.80 + 1.40 = 26.20$$

    $$\boxed{23.40 \leq \mu \leq 26.20}$$


    Conclusion

    • Standard error of the mean $= 0.714$
    • 95% confidence interval for the population mean $$= $23.40$$ to $$$26.20$$

    We are 95% confident that the true mean dinner amount per customer lies between $23.40 and $26.20.

  10. 105 marksMarkov ProcessAnswer

    Define Markov chain and its characteristics. [5]

    Markov Chain and Its Characteristics

    Definition of Stochastic Process

    A stochastic process is a family of random variables indexed by time as a parameter. It is denoted as {X(t), t ∈ T}, where:

    • T = index set (time parameter)
    • X(t) = random variable at time t
    • I = state space (the set of all possible values assumed by X(t))

    The values assumed by the random variable X(t) are called states.


    Definition of Markov Chain

    A Markov Chain is a stochastic process {X(t), t = 0, 1, 2, ...} that satisfies the Markov property (memoryless property):

    $$P[X(t+1) = j \mid X(t) = i,\ X(t-1) = i_{t-1},\ \ldots,\ X(0) = i_0] = P[X(t+1) = j \mid X(t) = i]$$

    That is, the future state depends only on the present state, not on the past states.

    The probability $P[X(t+1) = j \mid X(t) = i] = p_{ij}$ is called the transition probability from state $i$ to state $j$.


    Characteristics of a Markov Chain

    1. Markov Property (Memoryless Property)

    The conditional probability of any future state depends only on the current state, not on the sequence of states that preceded it. This is the defining characteristic.

    2. Transition Probability Matrix (TPM)

    The one-step transition probabilities are arranged in a matrix $P = [p_{ij}]$, where:

    • $p_{ij} \geq 0$ for all $i, j$
    • $\sum_{j} p_{ij} = 1$ for each row $i$ (rows sum to 1)

    Such a matrix is called a stochastic matrix.

    3. Homogeneity (Time Homogeneous)

    A Markov chain is said to be homogeneous if the transition probabilities do not depend on time $t$, i.e., $p_{ij}$ remains constant over time.

    4. Chapman-Kolmogorov Equation

    The $n$-step transition probability satisfies:

    $$p_{ij}^{(n)} = \sum_{k} p_{ik}^{(r)} \cdot p_{kj}^{(n-r)}, \quad 0 < r < n$$

    In matrix form: $P^{(n)} = P^n$

    5. Steady State (Limiting) Distribution

    As $n \to \infty$, the $n$-step transition probabilities converge to a steady state distribution $\pi$, such that:

    $$\pi_j = \lim_{n \to \infty} P^n(x)$$

    The steady state distribution satisfies: $$\pi P = \pi \quad \text{and} \quad \sum_{j} \pi_j = 1$$

    6. Irreducibility

    A Markov chain is irreducible if every state can be reached from every other state, i.e., all states communicate with each other.

    7. Recurrence and Transience

    • A state $i$ is recurrent if the process returns to state $i$ with probability 1.
    • A state $i$ is transient if there is a positive probability of never returning to state $i$.

    Summary Table

    CharacteristicDescription
    Markov PropertyFuture depends only on present
    TPMNon-negative entries, rows sum to 1
    HomogeneityTransition probs constant over time
    Chapman-Kolmogorov$P^{(n)} = P^n$
    Steady State$\pi P = \pi$, $\sum \pi_j = 1$
    IrreducibilityAll states communicate
  11. 115 marksNumericalM/M/1 systemAnswer

    What are the basic concepts of queuing theory? In a super market, the average arrivals rate of customer is 10 per every 30 minutes following Poisson process. The average time taken by the cashier to list and calculate the customers purchase is 2.5 minutes following exponential distribution. What is the probability that queue length exceeds 6? What is the expected time spent by customer in the system? [5]

    • Arrival: 10 customers per 30 minutes (Poisson process) - Service time: 2.5 minutes per customer (exponential distribution) - Find: $P(\text{queue length} 6)$ and expected time in system $W$ --- Queuing theory is the mathematical study ...
  12. 125 marksMultiple and partial correlationAnswer

    Write short notes on the following. i. Partial and multiple correlation coefficient. ii. Properties of good estimator. [5]

    --- Definition: The partial correlation coefficient measures the degree of linear relationship between two variables while eliminating (holding constant) the effect of one or more other variables. For three variables X1, X2, and X3, the ...