2082

STA154 · TU past paper

Basic Statistics 2082 question paper

The complete TU 2082 exam paper for Basic Statistics (STA154), all 12 questions with solved model answers written to the mark scheme.

Tap a question to open its answer.

  1. 110 marksNumericalStandard normal distribution and Z-scoresAnswer

    The distribution of monthly incomes of 5,000 employees of a certain industrial unit was found to be normally distributed with mean of Rs. 2,000 and a standard deviation of Rs. 200.(i) Estimate the range of incomes of the middle 60% employees.(ii) Estimate the lowest income of richest 10% employees.(iii) Estimate the highest income of poorest 10% employees.[10+0+0+0]

    • Number of employees: $N = 5000$ - Distribution: Normal - Mean: $\mu = 2000$ - Standard deviation: $\sigma = 200$ Standardization: $Z = \dfrac{X - \mu}{\sigma}$, so $X = \mu + Z\sigma$. --- Middle 60% leaves 20% in each tail. - Lower cu...
  2. 210 marksNumericalComparison of consistency between datasetsAnswer

    Build System Consistency Analysis

    A software development team is tracking the build sizes (in MB) produced by two different automated build systems during nightly integrations over 10 days. Which build system is more consistent? Justify your answer.

    $$\begin{array}{|c|cccccccccc|}\hline \text{Build System A} & 88 & 92 & 94 & 85 & 90 & 95 & 89 & 87 & 91 & 86 \ \hline \text{Build System B} & 130 & 135 & 132 & 140 & 128 & 133 & 137 & 131 & 129 & 138 \ \hline \end{array}$$

    [10]

    Model Answer: Build System Consistency Analysis

    Given Data

    Build System A (MB): 88, 92, 94, 85, 90, 95, 89, 87, 91, 86 Build System B (MB): 130, 135, 132, 140, 128, 133, 137, 131, 129, 138 $n = 10$ for each

    Method

    Since the two data sets have different means (different units/scale of build sizes), the correct measure of relative consistency is the Coefficient of Variation (CV). The lower the CV, the more consistent.

    Step 1: Means

    System A: $$\bar{x}_A = \frac{88+92+94+85+90+95+89+87+91+86}{10} = \frac{897}{10} = 89.7 \text{ MB}$$

    System B: $$\bar{x}_B = \frac{130+135+132+140+128+133+137+131+129+138}{10} = \frac{1333}{10} = 133.3 \text{ MB}$$

    Step 2: Variance and Standard Deviation

    System A - squared deviations from 89.7:

    $x$$x-\bar{x}$$(x-\bar{x})^2$
    88-1.72.89
    922.35.29
    944.318.49
    85-4.722.09
    900.30.09
    955.328.09
    89-0.70.49
    87-2.77.29
    911.31.69
    86-3.713.69

    $$\sum (x-\bar{x})^2 = 100.10$$ $$\sigma_A^2 = \frac{100.10}{10} = 10.01, \qquad \sigma_A = \sqrt{10.01} \approx 3.164 \text{ MB}$$

    System B - squared deviations from 133.3:

    $x$$x-\bar{x}$$(x-\bar{x})^2$
    130-3.310.89
    1351.72.89
    132-1.31.69
    1406.744.89
    128-5.328.09
    133-0.30.09
    1373.713.69
    131-2.35.29
    129-4.318.49
    1384.722.09

    $$\sum (x-\bar{x})^2 = 148.10$$ $$\sigma_B^2 = \frac{148.10}{10} = 14.81, \qquad \sigma_B = \sqrt{14.81} \approx 3.849 \text{ MB}$$

    Step 3: Coefficient of Variation

    $$CV_A = \frac{\sigma_A}{\bar{x}_A}\times 100 = \frac{3.164}{89.7}\times 100 \approx 3.53%$$

    $$CV_B = \frac{\sigma_B}{\bar{x}_B}\times 100 = \frac{3.849}{133.3}\times 100 \approx 2.89%$$

    Conclusion

    Consistency between two data sets with different means must be judged by the coefficient of variation, not the raw standard deviation.

    • $CV_B \approx 2.89% < CV_A \approx 3.53%$

    Therefore Build System B is more consistent, because relative to its own average build size its build sizes vary less. Although System A has a smaller absolute standard deviation, its builds are much smaller in magnitude, so a fair (relative) comparison shows System B fluctuates proportionally less around its mean.

    Compare with the coefficient of variation, not the raw standard deviation: when the means differ substantially (89.7 against 133.3), the CV is the statistically correct basis, and it points to System B as the more consistent system.

  3. 310 marksNumericalKarl Pearson correlation coefficientAnswer

    Question

    A web development team records the number of hours spent on debugging (X) and the number of resolved issues (Y) across 10 sprints.

    $$\begin{array}{|c|cccccccccc|}\hline X & 15 & 20 & 25 & 30 & 35 & 40 & 45 & 50 & 55 & 60 \ \hline Y & 3 & 6 & 8 & 11 & 12 & 14 & 16 & 19 & 21 & 22 \ \hline \end{array}$$

    a) Calculate the Pearson correlation coefficient to assess the relationship between debugging hours and issues resolved.

    b) Derive the regression equation of issues resolved on hours spent debugging.

    c) Predict the number of issues resolved for 38 hours of debugging.

    [10+0+0+0]

    Model Answer: Correlation and Regression Analysis

    Given Data

    X (hours)15202530354045505560
    Y (issues)36811121416192122

    $n = 10$


    STEP 1: Compute the sums

    XYXYX²Y²
    153452259
    20612040036
    25820062564
    3011330900121
    35124201225144
    40145601600196
    45167202025256
    50199502500361
    552111553025441
    602213203600484
    3751325820161252112

    $\sum X = 375,\ \sum Y = 132,\ \sum XY = 5820,\ \sum X^2 = 16125,\ \sum Y^2 = 2112$


    a) Pearson Correlation Coefficient

    $$r = \frac{n\sum XY - \sum X \sum Y}{\sqrt{[n\sum X^2 - (\sum X)^2][n\sum Y^2 - (\sum Y)^2]}}$$

    Compute components:

    • $n\sum XY = 10 \times 5820 = 58200$
    • $\sum X \sum Y = 375 \times 132 = 49500$
    • $n\sum X^2 = 161250,\quad (\sum X)^2 = 140625 \Rightarrow 20625$
    • $n\sum Y^2 = 21120,\quad (\sum Y)^2 = 17424 \Rightarrow 3696$

    $$r = \frac{58200 - 49500}{\sqrt{20625 \times 3696}} = \frac{8700}{\sqrt{76245000}}$$

    $$\sqrt{76245000} = 8731.95$$

    $$r = \frac{8700}{8731.95} = 0.9963$$

    $r \approx 0.996$: very strong positive correlation.


    b) Regression Equation of Y on X

    $$b = \frac{n\sum XY - \sum X \sum Y}{n\sum X^2 - (\sum X)^2} = \frac{8700}{20625} = 0.42182$$

    Means:

    $$\bar{X} = \frac{375}{10} = 37.5,\qquad \bar{Y} = \frac{132}{10} = 13.2$$

    Intercept:

    $$a = \bar{Y} - b\bar{X} = 13.2 - (0.42182)(37.5) = 13.2 - 15.818 = -2.618$$

    $$\boxed{\hat{Y} = -2.618 + 0.4218,X}$$


    c) Prediction for X = 38 hours

    $$\hat{Y} = -2.618 + 0.4218(38) = -2.618 + 16.029 = 13.41$$

    Approximately 13 issues resolved (13.4 if retaining decimals).

    Since $r = 0.996$ and $X = 38$ lies inside the observed range $[15,60]$, the prediction is reliable.

  4. 45 marksBar diagrams and Pareto diagramsAnswer

    Differentiate Pareto chart and a bar diagram. [5]

    Differentiate Pareto Chart and Bar Diagram

    Definition and Purpose

    Bar Diagram:

    • A basic statistical graph that displays data using rectangular bars of equal width
    • Bars are arranged side-by-side or in groups
    • Used to compare quantities across different categories
    • Primarily for simple data comparison and visualization

    Pareto Chart:

    • A specialized bar chart combined with a line graph (cumulative percentage curve)
    • Based on the Pareto principle (80/20 rule)
    • Bars are arranged in descending order of frequency/magnitude
    • Used to identify the most significant factors contributing to a problem

    Key Differences

    FeatureBar DiagramPareto Chart
    ArrangementBars can be in any order (categorical or random)Bars arranged in descending order (highest to lowest)
    ComponentsOnly bars representing frequencies/valuesBars + cumulative percentage line graph
    PurposeGeneral comparison of data across categoriesIdentify vital few factors from trivial many
    ApplicationSimple data presentation and comparisonQuality control, problem analysis, prioritization
    PrincipleNo specific principle appliedBased on Pareto principle (80% problems from 20% causes)
    ComplexitySimple and straightforwardMore complex with dual axes

    Example Context

    • Bar Diagram: Comparing sales of different products in a month
    • Pareto Chart: Identifying which product defects account for 80% of quality issues

    Conclusion

    While a bar diagram is a general-purpose tool for data comparison, a Pareto chart is a specialized analytical tool designed to prioritize problems and identify the most impactful factors in quality management and process improvement.

  5. 55 marksMeasurement scales and typesAnswer

    Fill the scale of measurement with the correct statistical test/measure.

    VariableMeasurement ScaleBest Statistical Method
    Blood groupNominalMode, Chi-square test
    Students' satisfaction (5-point Likert scale)OrdinalMedian, Mann-Whitney U test, Spearman's rank correlation
    Annual incomeRatioMean, Standard deviation, t-test, Pearson correlation
    Age group (18–25, 26–35, 36–45, 46+)OrdinalMedian, Mode, Chi-square test
    The lifetime of an electronic deviceRatioMean, Standard deviation, t-test, ANOVA

    [5]

    Variable Measurement Scale Best Statistical Method ----------------------------------------------------- Blood group Nominal Mode, Chi-square test, Frequency distribution Students' satisfaction (5-point Likert scale) Ordinal Median, Mode...

  6. 65 marksNumericalBasic probability concepts and rulesAnswer

    A piece of equipment will function only when all the components A, B, and C are working. The probability of A failing during one year is 0.15, that of B failing is 0.05, and that of C failing is 0.10. What is the probability that the equipment will not fail before the end of one year? [5]

    Model Answer: Equipment Reliability

    Step 1 - EXTRACT: Given Data

    • Equipment functions only when all of A, B, C are working.
    • $P(A \text{ fails}) = 0.15$
    • $P(B \text{ fails}) = 0.05$
    • $P(C \text{ fails}) = 0.10$
    • Component failures assumed independent.
    • Find: $P(\text{equipment does not fail in one year})$

    Step 2 - SOLVE

    Probability each component does NOT fail:

    $$P(A \text{ works}) = 1 - 0.15 = 0.85$$ $$P(B \text{ works}) = 1 - 0.05 = 0.95$$ $$P(C \text{ works}) = 1 - 0.10 = 0.90$$

    Condition: Equipment does not fail only when all three components work.

    Using the multiplication rule for independent events:

    $$P(\text{no failure}) = P(A) \cdot P(B) \cdot P(C) = 0.85 \times 0.95 \times 0.90$$

    $$= 0.8075 \times 0.90 = 0.72675$$

    Final Answer

    $$\boxed{P(\text{equipment does not fail}) = 0.72675 \approx 0.7268 ;(72.68%)}$$

    The value is 0.7268 (exactly 0.72675).

  7. 75 marksNumericalConditional probability and Bayes theoremAnswer

    Three persons A, B, and C are being considered for appointment as Vice-Chancellor of a university, and whose chances of being selected are in the proportion 4:2:3 respectively. The probability that A, if selected, will introduce democratization is 0.3, and the corresponding probabilities for B and C are 0.5 and 0.8. What is the probability that democratization would be introduced? [5]

    • Selection proportion for A : B : C = 4 : 2 : 3 - $P(D \mid A) = 0.3$ - $P(D \mid B) = 0.5$ - $P(D \mid C) = 0.8$ where $D$ = "democratization is introduced". Total parts $= 4 + 2 + 3 = 9$ $$P(A) = \frac{4}{9}, \quad P(B) = \frac{2}{9},...
  8. 85 marksNumericalRandom variables and probability functionsAnswer

    A tech team records the number of bug reports closed per hour. The probability distribution is given below. Find the expected number of bug reports closed per hour and its variance.

    $$\begin{array}{|c|ccccc|}\hline Y & 0 & 1 & 2 & 3 & 4 \ \hline P(Y) & 0.10 & 0.18 & 0.32 & 0.30 & 0.10 \ \hline \end{array}$$

    [5]

    Verified Model Answer: Expected Value and Variance

    STEP 1 - Given Data

    Probability distribution of $Y$ (bug reports closed per hour):

    $Y$01234
    $P(Y)$0.100.180.320.300.10

    Check: $\sum P(Y) = 0.10 + 0.18 + 0.32 + 0.30 + 0.10 = 1.00$ ✓ (valid distribution)

    STEP 2 - Solve

    Part 1: Expected Value $E(Y)$

    $$E(Y) = \sum Y \cdot P(Y)$$

    $Y$$P(Y)$$Y \cdot P(Y)$
    00.100.00
    10.180.18
    20.320.64
    30.300.90
    40.100.40
    Total1.002.12

    $$E(Y) = 0 + 0.18 + 0.64 + 0.90 + 0.40 = 2.12 \text{ bug reports/hour}$$

    Part 2: Variance $\text{Var}(Y)$

    Compute $E(Y^2) = \sum Y^2 \cdot P(Y)$:

    $Y$$Y^2$$P(Y)$$Y^2 \cdot P(Y)$
    000.100.00
    110.180.18
    240.321.28
    390.302.70
    4160.101.60
    Total1.005.76

    $$E(Y^2) = 5.76$$

    Apply the variance formula:

    $$\text{Var}(Y) = E(Y^2) - [E(Y)]^2 = 5.76 - (2.12)^2$$

    $$(2.12)^2 = 4.4944$$

    $$\text{Var}(Y) = 5.76 - 4.4944 = 1.2656 \approx 1.27$$

    Final Results

    • Expected number of bug reports closed per hour: $E(Y) = 2.12$
    • Variance: $\text{Var}(Y) = 1.2656 \approx 1.27$
    • (Standard deviation: $\sigma = \sqrt{1.2656} \approx 1.125$)
  9. 95 marksNumericalConfidence interval for population proportAnswer

    A sample survey of 400 customers shows that 350 are satisfied with ABC company providing internet service. Estimate the proportion of satisfied customers in the market with 95% and 99% confidence interval. [5]

    Model Answer: Confidence Interval for Population Proportion

    STEP 1 - Given Data

    • Sample size: $n = 400$
    • Number of satisfied customers: $x = 350$
    • Confidence levels required: 95% and 99%

    STEP 2 - Solution

    Step 1: Sample Proportion

    $$\hat{p} = \frac{x}{n} = \frac{350}{400} = 0.875$$ $$\hat{q} = 1 - \hat{p} = 0.125$$

    Step 2: Standard Error

    $$SE = \sqrt{\frac{\hat{p},\hat{q}}{n}} = \sqrt{\frac{0.875 \times 0.125}{400}} = \sqrt{\frac{0.109375}{400}}$$ $$SE = \sqrt{0.000273437} = 0.01654$$

    Step 3: Critical Values

    • 95% confidence: $Z = 1.96$
    • 99% confidence: $Z = 2.576$

    Step 4: Confidence Intervals

    95% CI: $$\hat{p} \pm Z \cdot SE = 0.875 \pm 1.96 \times 0.01654$$ $$= 0.875 \pm 0.03242$$ $$= (0.8426,\ 0.9074)$$

    99% CI: $$0.875 \pm 2.576 \times 0.01654$$ $$= 0.875 \pm 0.04261$$ $$= (0.8324,\ 0.9176)$$

    Conclusion

    • 95% CI: $(0.843,\ 0.907)$ → between 84.3% and 90.7% of customers are satisfied.
    • 99% CI: $(0.832,\ 0.918)$ → between 83.2% and 91.8% of customers are satisfied.

    The 99% interval is wider, reflecting higher confidence at the cost of precision.

  10. 105 marksNumericalBinomial distributionAnswer

    It is observed that 80% of television viewers watch an entertainment channel. What is the probability that at least 80% of the viewers in a random sample of five watch an entertainment channel? [5]

    Model Answer: Probability of At Least 80% Watching Entertainment Channel

    STEP 1 - Given Data

    • Population proportion watching entertainment channel: $p = 0.80$
    • Probability of not watching: $q = 1 - p = 0.20$
    • Sample size: $n = 5$
    • Required: $P(\text{at least } 80% \text{ of sample watch})$

    STEP 2 - Solve

    Distribution: Let $X$ = number of viewers (out of 5) who watch the entertainment channel. Since each viewer independently watches with probability $p = 0.80$, $X$ follows a Binomial distribution:

    $$P(X = k) = \binom{5}{k}(0.80)^k(0.20)^{5-k}$$

    Interpret "at least 80% of 5": $$0.80 \times 5 = 4 \text{ viewers}$$

    So we need $P(X \geq 4) = P(X = 4) + P(X = 5)$.

    For $X = 4$: $$P(X = 4) = \binom{5}{4}(0.80)^4(0.20)^1 = 5 \times 0.4096 \times 0.20$$ $$= 5 \times 0.08192 = 0.4096$$

    For $X = 5$: $$P(X = 5) = \binom{5}{5}(0.80)^5(0.20)^0 = 1 \times 0.32768 \times 1 = 0.32768$$

    Sum: $$P(X \geq 4) = 0.4096 + 0.32768 = 0.73728$$

    Final Answer

    $$\boxed{P(X \geq 4) \approx 0.7373 \text{ or } 73.73%}$$

    The probability that at least 80% of the viewers in a random sample of five watch the entertainment channel is approximately 0.737 (73.7%).

  11. 115 marksNumericalCentral moments and raw momentsAnswer

    The first four moments about point 5 are 3, 10, 40, and 500. Compute the four central moments. [5]

    Moments about the point $a = 5$: - $\mu'1 = 3$ - $\mu'2 = 10$ - $\mu'3 = 40$ - $\mu'4 = 500$ Required: The four central moments $\mu1, \mu2, \mu3, \mu4$. The central moments are obtained from the moments about an arbitrary point using:

  12. 125 marksCluster samplingAnswer

    Write a short note on the following: (a) Cluster sampling (b) Primary data [5+0+0]

    Model Answer: Cluster Sampling and Primary Data

    (a) Cluster Sampling

    Definition: Cluster sampling is a probability sampling technique in which the population is divided into groups or clusters, and a random sample of clusters is selected. All units within the selected clusters are then included in the sample.

    Key Characteristics:

    • The population is first divided into mutually exclusive and exhaustive clusters (groups)
    • Clusters are typically geographic or natural groupings
    • A random sample of clusters is chosen
    • All elements within selected clusters are surveyed (or sometimes a subsample is taken)

    Advantages:

    • Cost-effective, especially when population is geographically dispersed
    • Reduces travel and administrative costs
    • Practical when a complete population list is unavailable
    • Easier to implement than simple random sampling for large populations

    Disadvantages:

    • Higher sampling error compared to simple random sampling
    • Clusters may be heterogeneous (not uniform within themselves)
    • Less precise estimates if clusters are very different from each other

    Example: To survey students' satisfaction in Tribhuvan University, divide the university into clusters (faculties/departments), randomly select some clusters, and survey all students in those selected clusters.


    (b) Primary Data

    Definition: Primary data is information collected directly from original sources for a specific research purpose. It is first-hand data gathered by the researcher through direct observation, surveys, interviews, or experiments.

    Key Characteristics:

    • Collected directly by the researcher or research team
    • Specific to the research objective
    • Original and first-hand in nature
    • More accurate and reliable for the intended purpose

    Methods of Collection:

    • Questionnaires and surveys
    • Interviews (structured, unstructured, semi-structured)
    • Observation and experimentation
    • Focus group discussions
    • Case studies

    Advantages:

    • Highly relevant to research objectives
    • Greater accuracy and control over data quality
    • Researcher can verify authenticity
    • Confidentiality can be maintained

    Disadvantages:

    • Time-consuming and expensive
    • Requires trained personnel
    • May have response bias or non-response issues

    Example: Conducting a survey among BSc CSIT students to gather information about their programming skills is primary data collection.