STA154 · TU past paper
Basic Statistics 2082 question paper
The complete TU 2082 exam paper for Basic Statistics (STA154), all 12 questions with solved model answers written to the mark scheme.
Tap a question to open its answer.
- 110 marksNumericalStandard normal distribution and Z-scoresHideAnswer
The distribution of monthly incomes of 5,000 employees of a certain industrial unit was found to be normally distributed with mean of Rs. 2,000 and a standard deviation of Rs. 200.(i) Estimate the range of incomes of the middle 60% employees.(ii) Estimate the lowest income of richest 10% employees.(iii) Estimate the highest income of poorest 10% employees.[10+0+0+0]
- Number of employees: $N = 5000$ - Distribution: Normal - Mean: $\mu = 2000$ - Standard deviation: $\sigma = 200$ Standardization: $Z = \dfrac{X - \mu}{\sigma}$, so $X = \mu + Z\sigma$. --- Middle 60% leaves 20% in each tail. - Lower cu...
- 210 marksNumericalComparison of consistency between datasetsHideAnswer
Build System Consistency Analysis
A software development team is tracking the build sizes (in MB) produced by two different automated build systems during nightly integrations over 10 days. Which build system is more consistent? Justify your answer.
$$\begin{array}{|c|cccccccccc|}\hline \text{Build System A} & 88 & 92 & 94 & 85 & 90 & 95 & 89 & 87 & 91 & 86 \ \hline \text{Build System B} & 130 & 135 & 132 & 140 & 128 & 133 & 137 & 131 & 129 & 138 \ \hline \end{array}$$
[10]
Model Answer: Build System Consistency Analysis
Given Data
Build System A (MB): 88, 92, 94, 85, 90, 95, 89, 87, 91, 86 Build System B (MB): 130, 135, 132, 140, 128, 133, 137, 131, 129, 138 $n = 10$ for each
Method
Since the two data sets have different means (different units/scale of build sizes), the correct measure of relative consistency is the Coefficient of Variation (CV). The lower the CV, the more consistent.
Step 1: Means
System A: $$\bar{x}_A = \frac{88+92+94+85+90+95+89+87+91+86}{10} = \frac{897}{10} = 89.7 \text{ MB}$$
System B: $$\bar{x}_B = \frac{130+135+132+140+128+133+137+131+129+138}{10} = \frac{1333}{10} = 133.3 \text{ MB}$$
Step 2: Variance and Standard Deviation
System A - squared deviations from 89.7:
$x$ $x-\bar{x}$ $(x-\bar{x})^2$ 88 -1.7 2.89 92 2.3 5.29 94 4.3 18.49 85 -4.7 22.09 90 0.3 0.09 95 5.3 28.09 89 -0.7 0.49 87 -2.7 7.29 91 1.3 1.69 86 -3.7 13.69 $$\sum (x-\bar{x})^2 = 100.10$$ $$\sigma_A^2 = \frac{100.10}{10} = 10.01, \qquad \sigma_A = \sqrt{10.01} \approx 3.164 \text{ MB}$$
System B - squared deviations from 133.3:
$x$ $x-\bar{x}$ $(x-\bar{x})^2$ 130 -3.3 10.89 135 1.7 2.89 132 -1.3 1.69 140 6.7 44.89 128 -5.3 28.09 133 -0.3 0.09 137 3.7 13.69 131 -2.3 5.29 129 -4.3 18.49 138 4.7 22.09 $$\sum (x-\bar{x})^2 = 148.10$$ $$\sigma_B^2 = \frac{148.10}{10} = 14.81, \qquad \sigma_B = \sqrt{14.81} \approx 3.849 \text{ MB}$$
Step 3: Coefficient of Variation
$$CV_A = \frac{\sigma_A}{\bar{x}_A}\times 100 = \frac{3.164}{89.7}\times 100 \approx 3.53%$$
$$CV_B = \frac{\sigma_B}{\bar{x}_B}\times 100 = \frac{3.849}{133.3}\times 100 \approx 2.89%$$
Conclusion
Consistency between two data sets with different means must be judged by the coefficient of variation, not the raw standard deviation.
- $CV_B \approx 2.89% < CV_A \approx 3.53%$
Therefore Build System B is more consistent, because relative to its own average build size its build sizes vary less. Although System A has a smaller absolute standard deviation, its builds are much smaller in magnitude, so a fair (relative) comparison shows System B fluctuates proportionally less around its mean.
Compare with the coefficient of variation, not the raw standard deviation: when the means differ substantially (89.7 against 133.3), the CV is the statistically correct basis, and it points to System B as the more consistent system.
- 310 marksNumericalKarl Pearson correlation coefficientHideAnswer
Question
A web development team records the number of hours spent on debugging (X) and the number of resolved issues (Y) across 10 sprints.
$$\begin{array}{|c|cccccccccc|}\hline X & 15 & 20 & 25 & 30 & 35 & 40 & 45 & 50 & 55 & 60 \ \hline Y & 3 & 6 & 8 & 11 & 12 & 14 & 16 & 19 & 21 & 22 \ \hline \end{array}$$
a) Calculate the Pearson correlation coefficient to assess the relationship between debugging hours and issues resolved.
b) Derive the regression equation of issues resolved on hours spent debugging.
c) Predict the number of issues resolved for 38 hours of debugging.
[10+0+0+0]
Model Answer: Correlation and Regression Analysis
Given Data
X (hours) 15 20 25 30 35 40 45 50 55 60 Y (issues) 3 6 8 11 12 14 16 19 21 22 $n = 10$
STEP 1: Compute the sums
X Y XY X² Y² 15 3 45 225 9 20 6 120 400 36 25 8 200 625 64 30 11 330 900 121 35 12 420 1225 144 40 14 560 1600 196 45 16 720 2025 256 50 19 950 2500 361 55 21 1155 3025 441 60 22 1320 3600 484 375 132 5820 16125 2112 $\sum X = 375,\ \sum Y = 132,\ \sum XY = 5820,\ \sum X^2 = 16125,\ \sum Y^2 = 2112$
a) Pearson Correlation Coefficient
$$r = \frac{n\sum XY - \sum X \sum Y}{\sqrt{[n\sum X^2 - (\sum X)^2][n\sum Y^2 - (\sum Y)^2]}}$$
Compute components:
- $n\sum XY = 10 \times 5820 = 58200$
- $\sum X \sum Y = 375 \times 132 = 49500$
- $n\sum X^2 = 161250,\quad (\sum X)^2 = 140625 \Rightarrow 20625$
- $n\sum Y^2 = 21120,\quad (\sum Y)^2 = 17424 \Rightarrow 3696$
$$r = \frac{58200 - 49500}{\sqrt{20625 \times 3696}} = \frac{8700}{\sqrt{76245000}}$$
$$\sqrt{76245000} = 8731.95$$
$$r = \frac{8700}{8731.95} = 0.9963$$
$r \approx 0.996$: very strong positive correlation.
b) Regression Equation of Y on X
$$b = \frac{n\sum XY - \sum X \sum Y}{n\sum X^2 - (\sum X)^2} = \frac{8700}{20625} = 0.42182$$
Means:
$$\bar{X} = \frac{375}{10} = 37.5,\qquad \bar{Y} = \frac{132}{10} = 13.2$$
Intercept:
$$a = \bar{Y} - b\bar{X} = 13.2 - (0.42182)(37.5) = 13.2 - 15.818 = -2.618$$
$$\boxed{\hat{Y} = -2.618 + 0.4218,X}$$
c) Prediction for X = 38 hours
$$\hat{Y} = -2.618 + 0.4218(38) = -2.618 + 16.029 = 13.41$$
Approximately 13 issues resolved (13.4 if retaining decimals).
Since $r = 0.996$ and $X = 38$ lies inside the observed range $[15,60]$, the prediction is reliable.
- 45 marksBar diagrams and Pareto diagramsHideAnswer
Differentiate Pareto chart and a bar diagram. [5]
Differentiate Pareto Chart and Bar Diagram
Definition and Purpose
Bar Diagram:
- A basic statistical graph that displays data using rectangular bars of equal width
- Bars are arranged side-by-side or in groups
- Used to compare quantities across different categories
- Primarily for simple data comparison and visualization
Pareto Chart:
- A specialized bar chart combined with a line graph (cumulative percentage curve)
- Based on the Pareto principle (80/20 rule)
- Bars are arranged in descending order of frequency/magnitude
- Used to identify the most significant factors contributing to a problem
Key Differences
Feature Bar Diagram Pareto Chart Arrangement Bars can be in any order (categorical or random) Bars arranged in descending order (highest to lowest) Components Only bars representing frequencies/values Bars + cumulative percentage line graph Purpose General comparison of data across categories Identify vital few factors from trivial many Application Simple data presentation and comparison Quality control, problem analysis, prioritization Principle No specific principle applied Based on Pareto principle (80% problems from 20% causes) Complexity Simple and straightforward More complex with dual axes Example Context
- Bar Diagram: Comparing sales of different products in a month
- Pareto Chart: Identifying which product defects account for 80% of quality issues
Conclusion
While a bar diagram is a general-purpose tool for data comparison, a Pareto chart is a specialized analytical tool designed to prioritize problems and identify the most impactful factors in quality management and process improvement.
- 55 marksMeasurement scales and typesHideAnswer
Fill the scale of measurement with the correct statistical test/measure.
Variable Measurement Scale Best Statistical Method Blood group Nominal Mode, Chi-square test Students' satisfaction (5-point Likert scale) Ordinal Median, Mann-Whitney U test, Spearman's rank correlation Annual income Ratio Mean, Standard deviation, t-test, Pearson correlation Age group (18–25, 26–35, 36–45, 46+) Ordinal Median, Mode, Chi-square test The lifetime of an electronic device Ratio Mean, Standard deviation, t-test, ANOVA [5]
Variable Measurement Scale Best Statistical Method ----------------------------------------------------- Blood group Nominal Mode, Chi-square test, Frequency distribution Students' satisfaction (5-point Likert scale) Ordinal Median, Mode...
- 65 marksNumericalBasic probability concepts and rulesHideAnswer
A piece of equipment will function only when all the components A, B, and C are working. The probability of A failing during one year is 0.15, that of B failing is 0.05, and that of C failing is 0.10. What is the probability that the equipment will not fail before the end of one year? [5]
Model Answer: Equipment Reliability
Step 1 - EXTRACT: Given Data
- Equipment functions only when all of A, B, C are working.
- $P(A \text{ fails}) = 0.15$
- $P(B \text{ fails}) = 0.05$
- $P(C \text{ fails}) = 0.10$
- Component failures assumed independent.
- Find: $P(\text{equipment does not fail in one year})$
Step 2 - SOLVE
Probability each component does NOT fail:
$$P(A \text{ works}) = 1 - 0.15 = 0.85$$ $$P(B \text{ works}) = 1 - 0.05 = 0.95$$ $$P(C \text{ works}) = 1 - 0.10 = 0.90$$
Condition: Equipment does not fail only when all three components work.
Using the multiplication rule for independent events:
$$P(\text{no failure}) = P(A) \cdot P(B) \cdot P(C) = 0.85 \times 0.95 \times 0.90$$
$$= 0.8075 \times 0.90 = 0.72675$$
Final Answer
$$\boxed{P(\text{equipment does not fail}) = 0.72675 \approx 0.7268 ;(72.68%)}$$
The value is 0.7268 (exactly 0.72675).
- 75 marksNumericalConditional probability and Bayes theoremHideAnswer
Three persons A, B, and C are being considered for appointment as Vice-Chancellor of a university, and whose chances of being selected are in the proportion 4:2:3 respectively. The probability that A, if selected, will introduce democratization is 0.3, and the corresponding probabilities for B and C are 0.5 and 0.8. What is the probability that democratization would be introduced? [5]
- Selection proportion for A : B : C = 4 : 2 : 3 - $P(D \mid A) = 0.3$ - $P(D \mid B) = 0.5$ - $P(D \mid C) = 0.8$ where $D$ = "democratization is introduced". Total parts $= 4 + 2 + 3 = 9$ $$P(A) = \frac{4}{9}, \quad P(B) = \frac{2}{9},...
- 85 marksNumericalRandom variables and probability functionsHideAnswer
A tech team records the number of bug reports closed per hour. The probability distribution is given below. Find the expected number of bug reports closed per hour and its variance.
$$\begin{array}{|c|ccccc|}\hline Y & 0 & 1 & 2 & 3 & 4 \ \hline P(Y) & 0.10 & 0.18 & 0.32 & 0.30 & 0.10 \ \hline \end{array}$$
[5]
Verified Model Answer: Expected Value and Variance
STEP 1 - Given Data
Probability distribution of $Y$ (bug reports closed per hour):
$Y$ 0 1 2 3 4 $P(Y)$ 0.10 0.18 0.32 0.30 0.10 Check: $\sum P(Y) = 0.10 + 0.18 + 0.32 + 0.30 + 0.10 = 1.00$ ✓ (valid distribution)
STEP 2 - Solve
Part 1: Expected Value $E(Y)$
$$E(Y) = \sum Y \cdot P(Y)$$
$Y$ $P(Y)$ $Y \cdot P(Y)$ 0 0.10 0.00 1 0.18 0.18 2 0.32 0.64 3 0.30 0.90 4 0.10 0.40 Total 1.00 2.12 $$E(Y) = 0 + 0.18 + 0.64 + 0.90 + 0.40 = 2.12 \text{ bug reports/hour}$$
Part 2: Variance $\text{Var}(Y)$
Compute $E(Y^2) = \sum Y^2 \cdot P(Y)$:
$Y$ $Y^2$ $P(Y)$ $Y^2 \cdot P(Y)$ 0 0 0.10 0.00 1 1 0.18 0.18 2 4 0.32 1.28 3 9 0.30 2.70 4 16 0.10 1.60 Total 1.00 5.76 $$E(Y^2) = 5.76$$
Apply the variance formula:
$$\text{Var}(Y) = E(Y^2) - [E(Y)]^2 = 5.76 - (2.12)^2$$
$$(2.12)^2 = 4.4944$$
$$\text{Var}(Y) = 5.76 - 4.4944 = 1.2656 \approx 1.27$$
Final Results
- Expected number of bug reports closed per hour: $E(Y) = 2.12$
- Variance: $\text{Var}(Y) = 1.2656 \approx 1.27$
- (Standard deviation: $\sigma = \sqrt{1.2656} \approx 1.125$)
- 95 marksNumericalConfidence interval for population proportHideAnswer
A sample survey of 400 customers shows that 350 are satisfied with ABC company providing internet service. Estimate the proportion of satisfied customers in the market with 95% and 99% confidence interval. [5]
Model Answer: Confidence Interval for Population Proportion
STEP 1 - Given Data
- Sample size: $n = 400$
- Number of satisfied customers: $x = 350$
- Confidence levels required: 95% and 99%
STEP 2 - Solution
Step 1: Sample Proportion
$$\hat{p} = \frac{x}{n} = \frac{350}{400} = 0.875$$ $$\hat{q} = 1 - \hat{p} = 0.125$$
Step 2: Standard Error
$$SE = \sqrt{\frac{\hat{p},\hat{q}}{n}} = \sqrt{\frac{0.875 \times 0.125}{400}} = \sqrt{\frac{0.109375}{400}}$$ $$SE = \sqrt{0.000273437} = 0.01654$$
Step 3: Critical Values
- 95% confidence: $Z = 1.96$
- 99% confidence: $Z = 2.576$
Step 4: Confidence Intervals
95% CI: $$\hat{p} \pm Z \cdot SE = 0.875 \pm 1.96 \times 0.01654$$ $$= 0.875 \pm 0.03242$$ $$= (0.8426,\ 0.9074)$$
99% CI: $$0.875 \pm 2.576 \times 0.01654$$ $$= 0.875 \pm 0.04261$$ $$= (0.8324,\ 0.9176)$$
Conclusion
- 95% CI: $(0.843,\ 0.907)$ → between 84.3% and 90.7% of customers are satisfied.
- 99% CI: $(0.832,\ 0.918)$ → between 83.2% and 91.8% of customers are satisfied.
The 99% interval is wider, reflecting higher confidence at the cost of precision.
- 105 marksNumericalBinomial distributionHideAnswer
It is observed that 80% of television viewers watch an entertainment channel. What is the probability that at least 80% of the viewers in a random sample of five watch an entertainment channel? [5]
Model Answer: Probability of At Least 80% Watching Entertainment Channel
STEP 1 - Given Data
- Population proportion watching entertainment channel: $p = 0.80$
- Probability of not watching: $q = 1 - p = 0.20$
- Sample size: $n = 5$
- Required: $P(\text{at least } 80% \text{ of sample watch})$
STEP 2 - Solve
Distribution: Let $X$ = number of viewers (out of 5) who watch the entertainment channel. Since each viewer independently watches with probability $p = 0.80$, $X$ follows a Binomial distribution:
$$P(X = k) = \binom{5}{k}(0.80)^k(0.20)^{5-k}$$
Interpret "at least 80% of 5": $$0.80 \times 5 = 4 \text{ viewers}$$
So we need $P(X \geq 4) = P(X = 4) + P(X = 5)$.
For $X = 4$: $$P(X = 4) = \binom{5}{4}(0.80)^4(0.20)^1 = 5 \times 0.4096 \times 0.20$$ $$= 5 \times 0.08192 = 0.4096$$
For $X = 5$: $$P(X = 5) = \binom{5}{5}(0.80)^5(0.20)^0 = 1 \times 0.32768 \times 1 = 0.32768$$
Sum: $$P(X \geq 4) = 0.4096 + 0.32768 = 0.73728$$
Final Answer
$$\boxed{P(X \geq 4) \approx 0.7373 \text{ or } 73.73%}$$
The probability that at least 80% of the viewers in a random sample of five watch the entertainment channel is approximately 0.737 (73.7%).
- 115 marksNumericalCentral moments and raw momentsHideAnswer
The first four moments about point 5 are 3, 10, 40, and 500. Compute the four central moments. [5]
Moments about the point $a = 5$: - $\mu'1 = 3$ - $\mu'2 = 10$ - $\mu'3 = 40$ - $\mu'4 = 500$ Required: The four central moments $\mu1, \mu2, \mu3, \mu4$. The central moments are obtained from the moments about an arbitrary point using:
- 125 marksCluster samplingHideAnswer
Write a short note on the following: (a) Cluster sampling (b) Primary data [5+0+0]
Model Answer: Cluster Sampling and Primary Data
(a) Cluster Sampling
Definition: Cluster sampling is a probability sampling technique in which the population is divided into groups or clusters, and a random sample of clusters is selected. All units within the selected clusters are then included in the sample.
Key Characteristics:
- The population is first divided into mutually exclusive and exhaustive clusters (groups)
- Clusters are typically geographic or natural groupings
- A random sample of clusters is chosen
- All elements within selected clusters are surveyed (or sometimes a subsample is taken)
Advantages:
- Cost-effective, especially when population is geographically dispersed
- Reduces travel and administrative costs
- Practical when a complete population list is unavailable
- Easier to implement than simple random sampling for large populations
Disadvantages:
- Higher sampling error compared to simple random sampling
- Clusters may be heterogeneous (not uniform within themselves)
- Less precise estimates if clusters are very different from each other
Example: To survey students' satisfaction in Tribhuvan University, divide the university into clusters (faculties/departments), randomly select some clusters, and survey all students in those selected clusters.
(b) Primary Data
Definition: Primary data is information collected directly from original sources for a specific research purpose. It is first-hand data gathered by the researcher through direct observation, surveys, interviews, or experiments.
Key Characteristics:
- Collected directly by the researcher or research team
- Specific to the research objective
- Original and first-hand in nature
- More accurate and reliable for the intended purpose
Methods of Collection:
- Questionnaires and surveys
- Interviews (structured, unstructured, semi-structured)
- Observation and experimentation
- Focus group discussions
- Case studies
Advantages:
- Highly relevant to research objectives
- Greater accuracy and control over data quality
- Researcher can verify authenticity
- Confidentiality can be maintained
Disadvantages:
- Time-consuming and expensive
- Requires trained personnel
- May have response bias or non-response issues
Example: Conducting a survey among BSc CSIT students to gather information about their programming skills is primary data collection.