STA169 · TU past paper
Statistics I 2076 question paper
The complete TU 2076 exam paper for Statistics I (STA169), all 13 questions with solved model answers written to the mark scheme.
Tap a question to open its answer.
- 110 marksNumericalMeasures of dispersionHideAnswer
What are the roles of measure of dispersion in descriptive statistics?
Following table gives the frequency distribution of thickness of computer chips (in nanometer) manufactured by two companies.
$$\begin{array}{|c|cccccc|}\hline \text{Thickness of computer chips} & 5 & 10 & 15 & 20 & 25 & 30 \ \hline \text{Number of chips} & & & & & & \ \text{Company A} & 10 & 15 & 24 & 20 & 18 & 13 \ \text{Company B} & 12 & 18 & 20 & 22 & 24 & 4 \ \hline \end{array}$$
[10]
Measures of Dispersion and Comparison of Two Companies
Part 1: Roles of Measures of Dispersion in Descriptive Statistics
A measure of dispersion quantifies the spread or scatter of observations around a central value. Its main roles are:
- Judging reliability of an average: Small dispersion means the average represents the data well; large dispersion weakens the average's reliability.
- Comparing consistency of two or more distributions (using relative measures such as CV), e.g. comparing two manufacturers.
- Quality control: Detects whether a process stays within acceptable variation limits.
- Foundation for advanced statistics: Standard deviation feeds into correlation, regression, skewness, kurtosis, and hypothesis testing.
- Homogeneity check: Low dispersion = homogeneous data; high dispersion = heterogeneous data.
- Unit-free comparison: The coefficient of variation allows comparison across datasets with different means/units.
Part 2: Comparison Using Coefficient of Variation
$$CV=\frac{\sigma}{\bar x}\times 100%,\qquad \text{lower CV } \Rightarrow \text{ more consistent.}$$
Given Data
- Thickness $x$: 5, 10, 15, 20, 25, 30 (nm)
- Company A frequencies: 10, 15, 24, 20, 18, 13
- Company B frequencies: 12, 18, 20, 22, 24, 4
Company A
$x$ $f$ $fx$ $fx^2$ 5 10 50 250 10 15 150 1500 15 24 360 5400 20 20 400 8000 25 18 450 11250 30 13 390 11700 Σ 100 1800 38100 $$\bar x_A=\frac{1800}{100}=18$$ $$\sigma_A^2=\frac{38100}{100}-18^2=381-324=57$$ $$\sigma_A=\sqrt{57}=7.5498$$ $$CV_A=\frac{7.5498}{18}\times100=41.94%$$
Company B
$x$ $f$ $fx$ $fx^2$ 5 12 60 300 10 18 180 1800 15 20 300 4500 20 22 440 8800 25 24 600 15000 30 4 120 3600 Σ 100 1700 34000 $$\bar x_B=\frac{1700}{100}=17$$ $$\sigma_B^2=\frac{34000}{100}-17^2=340-289=51$$ $$\sigma_B=\sqrt{51}=7.1414$$ $$CV_B=\frac{7.1414}{17}\times100=42.01%$$
Summary
Measure Company A Company B Mean 18 nm 17 nm SD 7.55 7.14 CV 41.94% 42.01% Conclusion: $CV_A (41.94%) < CV_B (42.01%)$, so Company A is marginally more consistent. The difference is very small, so the two production processes have almost identical relative variability.
- 210 marksNumericalRegression AnalysisHideAnswer
Question
A study was done to study the effect of ambient temperature on the electric power consumed by a chemical plant. Following table gives the data which are collected from an experimental pilot plant.
Temperature (°F) 27 45 72 58 31 60 34 74 Electric Power (BTU) 250 285 320 295 265 298 267 321 a. Identify which one is response variable, and fit a simple regression line, assuming that the relationship between them is linear.
b. Interpret the regression coefficient with reference to your problem.
c. Obtain coefficient of determination, and interpret this.
d. Based on the fitted model in (a), predict the power consumption for an ambient temperature of $65°F$. [10+0]
x = Temp (°F) 27 45 72 58 31 60 34 74 --------------------------- y = Power (BTU) 250 285 320 295 265 298 267 321 $n = 8$ x y x² xy y² ------------------ 27 250 729 6750 62500 45 285 2025 12825 81225 72 320 5184 23040 102400 58 295 3364 ...
- 310 marksContinuous distributionHideAnswer
a. Define Normal distribution. What are the main characteristics of a Normal distribution? b. What do you mean by probability density function? Write down its properties.[10]
--- A Normal Distribution is a continuous probability distribution of a random variable X with parameters μ (mean) and σ² (variance). Its probability density function is given by: $$f(x) = \frac{1}{\sigma\sqrt{2\pi}} \cdot e^{-\frac{1}{2...
- 45 marksNumericalMeasures of central tendencyHideAnswer
The following table gives the installation time (in minutes) for hardware on 50 different computers. If the average installation time is 30.2 minutes, find missing frequencies.
$$\begin{array}{|c|ccccc|c|}\hline \text{Installation Time} & 0-10 & 10-20 & 20-30 & 30-40 & 40-50 & \text{Total} \ \text{Number of computers} & 4 & - & 10 & - & 10 & 50 \ \hline \end{array}$$
[5]
Installation Time 0-10 10-20 20-30 30-40 40-50 Total :---::---::---::---::---::---::---: No. of computers 4 $f1$ 10 $f2$ 10 50 - Total frequency $N = 50$ - Mean $\bar{x} = 30.2$ minutes Missing frequencies: $f1$ (class 10-20) and $f2$ (c...
- 55 marksNumericalMeasures of central tendencyHideAnswer
Power Failure Duration Analysis
The length of power failure in minutes are recorded in the following table. Find $Q_3$, $D_2$ and $P_{40}$ and interpret the results.
Power failure time 22 23 24 25 26 27 28 Total Frequency 2 5 7 10 4 3 2 33 [5]
Power Failure Time (x) 22 23 24 25 26 27 28 ------------------------ Frequency (f) 2 5 7 10 4 3 2 Total $N = 33$. x f cf --------- 22 2 2 23 5 7 24 7 14 25 10 24 26 4 28 27 3 31 28 2 33 For a discrete series we use position
- 65 marksNumericalBayes theoremHideAnswer
A manufacturing company employs three analytical plans for the design and development of a particular product. For cost reasons, all three are used at varying times. In fact, plan 1, 2, and 3 are used for 30%, 20% and 50% of the products respectively. The defect rate in different procedures is as follows: $P(D/P_1) = 0.01$, $P(D/P_2) = 0.03$, $P(D/P_3) = 0.02$, where $P(D/P_j)$ is the probability of a defective product, given plan $j$. If a random product was observed and found to be defective, which plan was most likely used and thus responsible? [5]
Plan Prior $P(Pj)$ Defect Rate $P(D/Pj)$ ------------------ $P1$ 0.30 0.01 $P2$ 0.20 0.03 $P3$ 0.50 0.02 Find: posterior $P(Pj/D)$ for each plan; identify the largest. $$P(D) = (0.30)(0.01) + (0.20)(0.03) + (0.50)(0.02)$$ $$P(D) = 0.003 ...
- 75 marksNumericalMathematical expectation of a random variaHideAnswer
The random variable $X$ has the following probability distribution.
$$\begin{array}{|c|ccccccc|}\hline X & 0 & 1 & 2 & 3 & 4 & 5 & 6 \ P(X=x) & 0.03 & 0.15 & 0.4 & 0.2 & 0.1 & 0.07 & 0.05 \ \hline \end{array}$$
[5]
X 0 1 2 3 4 5 6 ------------------------ P(X=x) 0.03 0.15 0.40 0.20 0.10 0.07 0.05 The question statement does not specify what to compute. Standard interpretation for a 5-mark problem: verify validity, then find mean, variance, and stan...
- 85 marksNumericalJoint probability distribution of two randHideAnswer
If two random variables have the joint probability density function $$f(x,y) = \begin{cases} k(2x + 3y), & 0 \leq x \leq 1, 0 \leq y \leq 1 \ 0, & \text{otherwise} \end{cases}$$
find (i) constant $k$ (ii) conditional probability density function of $X$ (iii) Identify whether $X$ and $Y$ are independent. [5]
Joint PDF Problem: f(x,y) = k(2x + 3y)
Given Data
- Joint PDF: $f(x,y) = k(2x+3y)$ for $0 \le x \le 1$, $0 \le y \le 1$; zero otherwise.
(i) Finding the Constant k
The total probability must equal 1:
$$\int_0^1 \int_0^1 k(2x+3y), dx, dy = 1$$
Inner integral (over x):
$$\int_0^1 (2x+3y), dx = \left[x^2 + 3xy\right]_0^1 = 1 + 3y$$
Outer integral (over y):
$$k\int_0^1 (1+3y), dy = k\left[y + \tfrac{3y^2}{2}\right]_0^1 = k\left(1 + \tfrac{3}{2}\right) = \frac{5k}{2}$$
Setting equal to 1:
$$\frac{5k}{2} = 1 \implies \boxed{k = \frac{2}{5}}$$
(ii) Conditional PDF of X
The conditional PDF of $X$ given $Y=y$ is $f(x|y) = \dfrac{f(x,y)}{f_Y(y)}$.
Marginal PDF of Y:
$$f_Y(y) = \int_0^1 \frac{2}{5}(2x+3y), dx = \frac{2}{5}\left[x^2 + 3xy\right]_0^1 = \frac{2}{5}(1+3y), \quad 0 \le y \le 1$$
Conditional PDF:
$$f(x|y) = \frac{\frac{2}{5}(2x+3y)}{\frac{2}{5}(1+3y)} = \boxed{\frac{2x+3y}{1+3y}}, \quad 0 \le x \le 1$$
(iii) Independence of X and Y
Marginal PDF of X:
$$f_X(x) = \int_0^1 \frac{2}{5}(2x+3y), dy = \frac{2}{5}\left[2xy + \frac{3y^2}{2}\right]_0^1 = \frac{2}{5}\left(2x + \frac{3}{2}\right) = \frac{4x+3}{5}, \quad 0 \le x \le 1$$
Check factorization:
$$f_X(x)\cdot f_Y(y) = \frac{4x+3}{5}\cdot \frac{2(1+3y)}{5} = \frac{2(4x+3)(1+3y)}{25}$$
Compare with joint:
$$f(x,y) = \frac{2(2x+3y)}{5} = \frac{10(2x+3y)}{25}$$
Since
$$\frac{2(4x+3)(1+3y)}{25} \ne \frac{10(2x+3y)}{25},$$
the equality $f(x,y) = f_X(x)f_Y(y)$ fails.
Conclusion: X and Y are NOT independent.
- 95 marksNumericalDiscrete distributionsHideAnswer
A large chain retailer purchases a certain kind of electronic device from a manufacturer. The manufacturer indicates that the defective rate of the device is 15%. The inspector randomly picks 10 items from a shipment. What is the probability that there will be at least one defective item among these 10? [5]
- Defective rate (probability of success): $p = 0.15$ - Non-defective probability: $q = 1 - p = 0.85$ - Sample size: $n = 10$ - Required: $P(X \geq 1)$ --- Each device is defective or not, selections are independent, and $p$ is constant....
- 105 marksNumericalDiscrete distributionsHideAnswer
Message arrives at an electronic message center at random times, with an average of 9 messages per hour. a. What is the probability of receiving at least four messages during the next hour? b. What is the probability of receiving at most three messages during the next hour? [5]
- Average rate: 9 messages per hour - Time interval: next 1 hour - Model: Poisson distribution - Parameter: $\lambda = 9$ Poisson probability mass function: $$P(X = x) = \frac{e^{-\lambda}\lambda^x}{x!}, \quad \lambda = 9$$ With
- 115 marksNumericalSpearman's rank correlationHideAnswer
Question
Following data represent the preference of 10 students studying B.Sc(CSIT) towards two brands of computer namely Lenovo and Acer. Apply appropriate statistical tool to measure whether the brand preference is correlated. Also interpret your result.
$$\begin{array}{|c|c|c|c|c|c|c|c|c|c|c|}\hline \text{Computer} & \text{Student Preference} & & & & & & & & & \ \hline \text{Lenovo} & 5 & 2 & 9 & 8 & 1 & 10 & 3 & 4 & 6 & 7 \ \text{Acer} & 10 & 5 & 1 & 3 & 8 & 6 & 2 & 7 & 9 & 4 \ \hline \end{array}$$
[5]
Note: The question mentions DELL and HP in the text but the table gives Lenovo and Acer. I use the table data (the actual numbers). Ranks by 10 students, $n = 10$: Student Lenovo ($R1$) Acer ($R2$) --------- 1 5 10 2 2 5 3 9 1 4 8 3 5 1 ...
- 125 marksScales of measurementHideAnswer
What do your mean by measurement scale? Describe the different types of measurement scales used in statistics. [5]
The measurement consisting of counting the number of units or parts of units displayed by objects and phenomena is called a measurement scale. In other words, a measurement scale is a system or rule used to assign numbers or symbols to o...
- 135 marksTypes of samplingHideAnswer
What is sampling? Discuss various probability sampling techniques with merits and demerits. [5]
When one-by-one study of all units of a population is not possible due to factors like time, cost, manpower, resources, and destructive nature of study, we take a small representative part from the population for study. This small repres...