2080.1

STA169 · TU past paper

Statistics I 2080.1 question paper

The complete TU 2080.1 exam paper for Statistics I (STA169), all 13 questions with solved model answers written to the mark scheme.

Tap a question to open its answer.

  1. 110 marksNumericalMeasures of central tendencyAnswer

    Define statistics and discuss its importance in the field of computational sciences. The following are the numbers of minutes that a person had to wait for the bus to work on 20 working days: 15, 10, 2, 17, 5, 8, 3, 10, 2, 9, 5, 9, 13, 1, 10, 12, 5, 10, 8, 4. Compute mean, median, mode, standard, variance and coefficient of variation.[10]

    Statistics: Definition, Importance, and Computations

    Part 1: Definition of Statistics (2 marks)

    Statistics is the branch of science that deals with the collection, organization, presentation, analysis, and interpretation of numerical data in order to draw valid conclusions and make rational decisions under conditions of uncertainty.

    Part 2: Importance of Statistics in Computational Sciences (3 marks)

    1. Fast computation: Computers perform statistical operations (mean, variance, ANOVA, regression) rapidly and reliably, removing tedious hand calculation.
    2. Handling large/multivariate data: Complex methods like multivariate analysis, clustering, and linear programming become feasible.
    3. Basis of modern techniques: Data mining, machine learning, pattern recognition, bootstrap methods, and image analysis are inherently statistical and computational.
    4. Research and reliability: Enables reproducible, high-speed experimental research and simulation.
    5. Practical CS/IT problem solving: Reliability testing (e.g., expected lifetime of hardware), performance benchmarking, and predictive modeling all depend on statistics.

    Part 3: Statistical Computations (5 marks)

    Given data (n = 20)

    $$15, 10, 2, 17, 5, 8, 3, 10, 2, 9, 5, 9, 13, 1, 10, 12, 5, 10, 8, 4$$

    Sorted: $$1, 2, 2, 3, 4, 5, 5, 5, 8, 8, 9, 9, 10, 10, 10, 10, 12, 13, 15, 17$$

    Mean

    $$\sum X = 158$$ $$\bar{X} = \frac{158}{20} = 7.9 \text{ minutes}$$

    Median (n even → average of 10th and 11th values)

    10th = 8, 11th = 9 $$\text{Median} = \frac{8+9}{2} = 8.5 \text{ minutes}$$

    Mode

    Value 10 occurs 4 times (highest frequency). $$\text{Mode} = 10 \text{ minutes}$$

    Variance and Standard Deviation

    $$\sum X^2 = 1626$$

    $$\sigma^2 = \frac{\sum X^2}{n} - \bar{X}^2 = \frac{1626}{20} - (7.9)^2$$ $$= 81.3 - 62.41 = 18.89$$

    $$\sigma = \sqrt{18.89} = 4.3463 \approx 4.35 \text{ minutes}$$

    (Note: if the sample formula with $n-1$ is used: $$s^2 = \frac{\sum X^2 - n\bar{X}^2}{n-1} = \frac{1626 - 1248.2}{19} = \frac{377.8}{19} = 19.8842$$ $$s = 4.4592 \text{ minutes})$$

    Coefficient of Variation

    $$CV = \frac{\sigma}{\bar{X}} \times 100 = \frac{4.3463}{7.9} \times 100 = 55.02%$$

    (Using sample SD: $CV = \frac{4.4592}{7.9}\times100 = 56.45%$)

    Final Summary

    MeasureValue
    Mean7.9 min
    Median8.5 min
    Mode10 min
    Variance (population)18.89
    Standard Deviation (population)4.35 min
    Coefficient of Variation55.02%
  2. 210 marksNumericalContinuous distributionAnswer

    Define normal distribution. What are the main characteristics of normal distribution? Extruded plastic rods are automatically cut into length 5 inches. Actual length are normally distributed about a mean of 5 inches and their standard deviation is 0.05 inches. (i) What proportion of rods exceed tolerance limits of 4.9 inches to 5.1 inches? (ii) Proportion of rods having tolerance rod which is greater than 6.5 inches.[10]

    Normal Distribution: Definition, Characteristics, and Numerical

    Definition

    A normal distribution is a continuous probability distribution, symmetrical and bell-shaped about its mean, defined by two parameters, mean $\mu$ and standard deviation $\sigma$. Its probability density function is:

    $$f(x) = \frac{1}{\sigma\sqrt{2\pi}} e^{-\frac{1}{2}\left(\frac{x-\mu}{\sigma}\right)^2}, \quad -\infty < x < \infty$$

    Main Characteristics

    1. It is a continuous distribution with parameters $\mu$ (mean) and $\sigma^2$ (variance).
    2. The curve is bell-shaped and symmetrical about the mean $\mu$.
    3. Mean = Median = Mode (all coincide at the centre).
    4. Maximum ordinate occurs at $x = \mu$, equal to $\frac{1}{\sigma\sqrt{2\pi}}$.
    5. Skewness $\beta_1 = 0$ (perfectly symmetric).
    6. Kurtosis $\beta_2 = 3$ (mesokurtic).
    7. Empirical area rule:
      • $P(\mu-\sigma < X < \mu+\sigma) = 0.6826$
      • $P(\mu-2\sigma < X < \mu+2\sigma) = 0.9544$
      • $P(\mu-3\sigma < X < \mu+3\sigma) = 0.9974$
    8. Total area under the curve = 1.
    9. Asymptotic: the curve never touches the X-axis.

    Given Data

    • Mean $\mu = 5$ inches
    • Standard deviation $\sigma = 0.05$ inches
    • $X \sim N(5, 0.05^2)$
    • Standard normal variate: $Z = \dfrac{X - 5}{0.05}$

    Part (i): Proportion exceeding tolerance limits 4.9 to 5.1 inches

    Z-scores:

    $$Z_1 = \frac{4.9 - 5}{0.05} = -2, \qquad Z_2 = \frac{5.1 - 5}{0.05} = +2$$

    Proportion within limits:

    $$P(4.9 < X < 5.1) = P(-2 < Z < 2) = 2 \times P(0 < Z < 2) = 2 \times 0.4772 = 0.9544$$

    Proportion exceeding limits:

    $$P(X < 4.9 \text{ or } X > 5.1) = 1 - 0.9544 = \boxed{0.0456 \ (4.56%)}$$

    Part (ii): Proportion having length greater than 6.5 inches

    Z-score:

    $$Z = \frac{6.5 - 5}{0.05} = \frac{1.5}{0.05} = 30$$

    Since $Z = 30$ lies far beyond the standard normal table (which effectively ends near $Z \approx 4$):

    $$P(X > 6.5) = P(Z > 30) \approx \boxed{0}$$

    A rod exceeding 6.5 inches is 30 standard deviations above the mean, so the proportion is practically zero.

    Summary

    PartConditionZ-value(s)Proportion
    (i)X < 4.9 or X > 5.1$\pm 2$0.0456 (4.56%)
    (ii)X > 6.530$\approx 0$
  3. 310 marksNumericalKarl Pearson's coefficient of correlationAnswer

    Difference between Correlation and Regression Analysis

    Raw materials are used in the production of synthesis and fiber are stored in a place which has no humidity control. Measurement of relative humidity in the storage place and the moisture content of a sample of raw materials (both in percentage) on 10 days yields the following results.

    $$\begin{array}{c|cccccccccc} \text{Humidity % (X)} & 48 & 55 & 30 & 44 & 36 & 31 & 62 & 48 & 42 & 50 \ \text{Moisture content % (Y)} & 11 & 13 & 10 & 12 & 9 & 7 & 16 & 11 & 9 & 14 \ \end{array}$$

    i. Compute correlation coefficient between humidity and moisture content and interpret the result.

    ii. Find the regression equation of moisture content on humidity.

    iii. Estimate the moisture content if humidity is 45%.

    iv. Interpret the value of regression coefficient.

    [10]

    Difference Between Correlation and Regression Analysis

    BasisCorrelationRegression
    MeaningMeasures degree and direction of linear relationship between two variablesEstablishes functional relationship to predict one variable from another
    PurposeFind strength of associationEstimate/predict dependent variable
    VariablesBoth treated equallyOne independent (X), one dependent (Y)
    ResultA single value $r \in [-1, +1]$An equation (regression line)
    Symmetry$r(X,Y)=r(Y,X)$Regression of Y on X $\neq$ X on Y
    Cause-EffectDoes not imply causationAssumes a dependence relationship

    Given Data

    $n = 10$

    X48553044363162484250
    Y11131012971611914

    Preliminary Calculations

    XYX²Y²XY
    48112304121528
    55133025169715
    3010900100300
    44121936144528
    369129681324
    31796149217
    62163844256992
    48112304121528
    429176481378
    50142500196700
    4461122083413185210

    $$\bar{X}=\frac{446}{10}=44.6,\qquad \bar{Y}=\frac{112}{10}=11.2$$

    $$\Sigma X^2-n\bar X^2 = 20834-10(44.6)^2=20834-19891.6=942.4$$ $$\Sigma Y^2-n\bar Y^2 = 1318-10(11.2)^2=1318-1254.4=63.6$$ $$\Sigma XY-n\bar X\bar Y = 5210-10(44.6)(11.2)=5210-4995.2=214.8$$


    Part (i): Correlation Coefficient

    $$r=\frac{\Sigma XY-n\bar X\bar Y}{\sqrt{(\Sigma X^2-n\bar X^2)(\Sigma Y^2-n\bar Y^2)}}=\frac{214.8}{\sqrt{942.4\times63.6}}$$

    $$r=\frac{214.8}{\sqrt{59916.64}}=\frac{214.8}{244.78}\approx \boxed{0.877}$$

    Interpretation: $r=0.877$ shows a strong positive linear relationship. As humidity increases, moisture content also increases.


    Part (ii): Regression Equation of Y on X

    $$b_{YX}=\frac{\Sigma XY-n\bar X\bar Y}{\Sigma X^2-n\bar X^2}=\frac{214.8}{942.4}=0.2279\approx 0.228$$

    $$a=\bar Y-b\bar X=11.2-0.228(44.6)=11.2-10.169=1.031$$

    $$\boxed{\hat Y = 1.03 + 0.228X}$$


    Part (iii): Estimate Y when X = 45

    $$\hat Y = 1.03 + 0.228(45)=1.03+10.26=\boxed{11.29%}$$

    When humidity is 45%, estimated moisture content ≈ 11.29%.


    Part (iv): Interpretation of Regression Coefficient

    $b=0.228$ means that for every 1 percentage point increase in relative humidity, the moisture content increases on average by about 0.228 percentage points. The positive sign confirms a direct relationship: higher storage humidity leads to higher moisture in the raw materials.

  4. 45 marksTypes of samplingAnswer

    What is sampling? Explain the main purpose of sampling. Describe briefly stratified sampling. [5]

    Sampling: Definition, Purpose, and Stratified Sampling


    1. What is Sampling?

    Sampling is the process of selecting a subset (called a sample) of units from a larger group (called the population) in order to draw conclusions or make inferences about the entire population. Instead of studying every unit in the population (census), sampling allows us to study only a representative portion, saving time, cost, and effort.

    The difference between the value of the sample statistic obtained from a sample and the value of the corresponding population parameter is called the sampling error.


    2. Main Purpose of Sampling

    The main purposes of sampling are:

    • Economy: Studying the entire population is costly; sampling reduces expenditure of money, time, and manpower.
    • Feasibility: In many cases, examining every unit is practically impossible (e.g., testing every light bulb for lifespan).
    • Speed: Results can be obtained more quickly from a sample than from a complete census.
    • Accuracy: A carefully conducted sample survey can sometimes be more accurate than a full census, since more attention can be given to each selected unit.
    • Defining survey objectives: The first step in sampling is to clearly define the objective of the survey, commensurate with available resources in terms of money, manpower, and time.

    3. Stratified Random Sampling

    Definition

    When the units in the population are not similar in nature, the population is first divided into sub-groups called strata, and then a simple random sample is drawn from each stratum in proportion to its size. This technique is known as Stratified Random Sampling.

    Conditions for Stratification

    The strata must satisfy the following conditions:

    • The strata should be non-overlapping and together should comprise the whole population.
    • The strata should be as homogeneous within groups (similar units inside each stratum) and heterogeneous between groups (different units across strata) as possible.

    Purpose of Stratification

    According to the notes, the purposes of stratification are:

    1. To make the sample more representative of the population.
    2. For greater accuracy in the results.
    3. For administrative convenience in conducting the survey.

    Example

    Suppose we want to study the income level of people in a city. The population can be stratified into groups such as:

    • Low income group
    • Middle income group
    • High income group

    A random sample is then drawn from each group separately, ensuring all income levels are properly represented in the final sample.

    Advantage over Simple Random Sampling

    For a given level of precision, stratified random sampling usually requires a smaller sample size as compared to simple random sampling, making it more efficient.


    Summary Table:

    FeatureStratified Sampling
    Population divided intoStrata (sub-groups)
    Within strataHomogeneous
    Between strataHeterogeneous
    Sampling from each stratumSimple Random Sampling
    Main benefitMore representative and accurate
  5. 55 marksNumericalMeasures of dispersionAnswer

    Measurement of Dispersion and Consistency Analysis

    Measurement of dispersion refers to statistical measures describing the spread or variability of data values around a central value. It indicates how much individual observations deviate from the average. Common measures: Range, Mean Dev...

  6. 65 marksNumericalMeasures of skewnessAnswer

    Define skewness and kurtosis. The first four moments about mean are 0, 14.75, 39.75 and 152.31. Compute skewness and kurtosis and interpret the results. [5]

    Skewness measures the lack of symmetry in a frequency distribution. It indicates the direction and degree of departure from symmetry. A perfectly symmetrical distribution has skewness = 0. Kurtosis measures the degree of peakedness or fl...

  7. 75 marksNumericalMeasures of central tendencyAnswer

    Partition Values

    What are partition values?

    Partition values are values that divide a distribution into equal parts. Common partition values include quartiles (dividing into 4 parts), deciles (dividing into 10 parts), and percentiles (dividing into 100 parts).

    From the following distribution of scores of 200 students of a college, compute:

    Scores30-4040-5050-6060-7070-8080-90
    Number of students145060452011

    i. The minimum scores obtained by top 10% students.

    The top 10% corresponds to the 90th percentile. We need to find the score below which 90% of students fall.

    Position of 90th percentile = $\frac{90}{100} \times 200 = 180$

    Cumulative frequencies:

    • 30-40: 14
    • 40-50: 64
    • 50-60: 124
    • 60-70: 169
    • 70-80: 189
    • 80-90: 200

    The 180th student falls in the 70-80 class.

    Using the formula: $P_{90} = L + \frac{\frac{90N}{100} - CF}{f} \times h$

    $$P_{90} = 70 + \frac{180 - 169}{20} \times 10 = 70 + \frac{11}{20} \times 10 = 70 + 5.5 = 75.5$$

    The minimum score obtained by top 10% students is 75.5

    ii. The range of middle 60% students.

    The middle 60% corresponds to the 20th to 80th percentiles.

    For 20th percentile: Position = $\frac{20}{100} \times 200 = 40$

    The 40th student falls in the 40-50 class.

    $$P_{20} = 40 + \frac{40 - 14}{50} \times 10 = 40 + \frac{26}{50} \times 10 = 40 + 5.2 = 45.2$$

    For 80th percentile: Position = $\frac{80}{100} \times 200 = 160$

    The 160th student falls in the 60-70 class.

    $$P_{80} = 60 + \frac{160 - 124}{45} \times 10 = 60 + \frac{36}{45} \times 10 = 60 + 8 = 68$$

    The range of middle 60% students is 45.2 to 68

    [5]

    Partition Values and Computation

    Definition

    Partition values are values that divide a distribution (arranged in order) into a number of equal parts. Common ones:

    • Quartiles ($Q$): divide data into 4 equal parts
    • Deciles ($D$): divide data into 10 equal parts
    • Percentiles ($P$): divide data into 100 equal parts

    Given Data (Cumulative Frequency Table)

    Scores$f$$cf$
    30-401414
    40-505064
    50-6060124
    60-7045169
    70-8020189
    80-9011200

    $N = 200$, class width $h = 10$.


    Part (i): Minimum Score of Top 10% Students

    Top 10% lie above $P_{90}$, so the minimum score of top 10% = $P_{90}$.

    $$\frac{90N}{100} = \frac{90 \times 200}{100} = 180\text{th item}$$

    Since $169 < 180 \le 189$, the $P_{90}$ class is 70-80.

    • $L = 70,\ cf = 169,\ f = 20,\ h = 10$

    $$P_{90} = 70 + \frac{180 - 169}{20} \times 10 = 70 + \frac{11}{20}\times 10 = 70 + 5.5 = \boxed{75.5}$$

    Minimum score of top 10% students = 75.5


    Part (ii): Range of Middle 60% Students

    Middle 60% lies between $P_{20}$ and $P_{80}$.

    Calculation of $P_{20}$:

    $$\frac{20 \times 200}{100} = 40\text{th item}$$

    Since $14 < 40 \le 64$, $P_{20}$ class is 40-50.

    $$P_{20} = 40 + \frac{40 - 14}{50}\times 10 = 40 + \frac{26}{50}\times 10 = 40 + 5.2 = 45.2$$

    Calculation of $P_{80}$:

    $$\frac{80 \times 200}{100} = 160\text{th item}$$

    Since $124 < 160 \le 169$, $P_{80}$ class is 60-70.

    $$P_{80} = 60 + \frac{160 - 124}{45}\times 10 = 60 + \frac{36}{45}\times 10 = 60 + 8 = 68$$

    Range of middle 60%:

    $$P_{80} - P_{20} = 68 - 45.2 = \boxed{22.8}$$

    Range of the middle 60% of students = 22.8 marks

  8. 85 marksNumericalLaws of probabilityAnswer

    Define mutually exclusive events and independent events in probability. A problem of mathematics is given to three students, A, B and C whose chances of solving the problem are in ratio 2 : 3 : 5. Find the probability that (i) all of them solve the problem (ii) none of them solve the problem (iii) the problem will be solved. [5]

    • Chances (probabilities) of solving the problem for A, B, C are in ratio $2 : 3 : 5$. - Total ratio parts $= 2 + 3 + 5 = 10$. Individual probabilities: $$P(A) = \frac{2}{10} = \frac{1}{5}, \quad P(B) = \frac{3}{10}, \quad P(C) = \frac{5...
  9. 95 marksNumericalBayes theoremAnswer

    Define Baye's theorem. Store A, B and C have 100, 75 and 50 employees and, respectively 70, 60 and 50 percent of these are women. Registration are equally likely among all employees regardless of sex. One employee resigns, and this is woman. What is the probability that she works in store B? [5]

    Bayes' Theorem - Definition and Application

    STEP 1: Given Data

    StoreEmployees% WomenWomen Count
    A10070%70
    B7560%45
    C5050%25
    Total225140

    Resignations equally likely among all employees. Observed event: the resigning employee is a woman. Find $P(\text{Store B} \mid \text{Woman})$.

    Definition of Bayes' Theorem

    If $B_1, B_2, \dots, B_n$ are mutually exclusive and exhaustive events with $P(B_i) > 0$, and $A$ is any event with $P(A) > 0$, then:

    $$P(B_i \mid A) = \frac{P(B_i), P(A \mid B_i)}{\sum_{j=1}^{n} P(B_j), P(A \mid B_j)}$$

    • $P(B_i)$ = prior probability
    • $P(A \mid B_i)$ = likelihood
    • $P(B_i \mid A)$ = posterior probability

    STEP 2: Solution

    Define Events

    • $S_A, S_B, S_C$ = employee works in Store A, B, C
    • $W$ = resigning employee is a woman

    Prior Probabilities (equal likelihood over 225 employees)

    $$P(S_A) = \frac{100}{225}, \quad P(S_B) = \frac{75}{225}, \quad P(S_C) = \frac{50}{225}$$

    Likelihoods

    $$P(W \mid S_A) = 0.70, \quad P(W \mid S_B) = 0.60, \quad P(W \mid S_C) = 0.50$$

    Total Probability of a Woman Resigning

    $$P(W) = \frac{100}{225}(0.70) + \frac{75}{225}(0.60) + \frac{50}{225}(0.50)$$

    $$P(W) = \frac{70 + 45 + 25}{225} = \frac{140}{225}$$

    Apply Bayes' Theorem

    $$P(S_B \mid W) = \frac{P(S_B), P(W \mid S_B)}{P(W)} = \frac{\dfrac{75}{225}(0.60)}{\dfrac{140}{225}} = \frac{45}{140}$$

    $$\boxed{P(S_B \mid W) = \frac{45}{140} = \frac{9}{28} \approx 0.3214 \approx 32.14%}$$

    Conclusion

    The probability that the resigning woman works in Store B is approximately 0.3214 (32.14%).

  10. 105 marksNumericalDiscrete distributionsAnswer

    Fitting a Binomial Distribution

    Under what condition binomial probability distribution? Five unbiased coins are tossed 100 times and the following results were obtained. Fit the binomial distribution.

    No of heads012345
    Frequency5243522104

    [5]

    The binomial probability distribution applies under the following conditions: - Each trial results in only two mutually exclusive outcomes (success/failure). - The number of trials n is fixed and finite. - The trials are independent of o...

  11. 115 marksNumericalDiscrete distributionsAnswer

    Define poisson probability distribution. Cars arrive at a petrol station at an average rate of 3 per minute. Assuming that the cars arrive at random, find the probability that (i) no car arrives during a particular minute. (ii) at least one car arrive during a particular minute. (iii) four cars arrives in any 2 minutes. [5]

    Poisson Probability Distribution

    Definition

    The Poisson distribution is a discrete probability distribution that describes the number of events occurring in a fixed interval of time (or space), given that these events occur with a known constant average rate $\lambda$ and independently of each other. It is a limiting case of the binomial distribution when $n \to \infty$, $p \to 0$, while $np = \lambda$ stays finite.

    Probability mass function:

    $$P(X = x) = \frac{e^{-\lambda},\lambda^x}{x!}, \quad x = 0, 1, 2, 3, \ldots$$

    where $\lambda$ is the mean number of occurrences, and $\text{mean} = \text{variance} = \lambda$.


    Given Data

    • Average arrival rate = $3$ cars per minute
    • $\lambda = 3$ (per minute)

    (i) No car arrives during a particular minute

    $x = 0$, $\lambda = 3$:

    $$P(X = 0) = \frac{e^{-3},3^0}{0!} = e^{-3} \approx 0.0498$$


    (ii) At least one car arrives during a particular minute

    $$P(X \geq 1) = 1 - P(X = 0) = 1 - e^{-3}$$

    $$P(X \geq 1) = 1 - 0.0498 \approx 0.9502$$


    (iii) Four cars arrive in any 2 minutes

    For a 2-minute interval, the mean scales:

    $$\lambda' = 3 \times 2 = 6$$

    With $x = 4$:

    $$P(X = 4) = \frac{e^{-6},6^4}{4!} = \frac{e^{-6}\times 1296}{24}$$

    $$= \frac{0.00247875 \times 1296}{24} = \frac{3.2124}{24} \approx 0.1339$$


    Summary

    PartResult
    (i) $P(X=0)$$e^{-3} \approx 0.0498$
    (ii) $P(X\ge 1)$$1 - e^{-3} \approx 0.9502$
    (iii) $P(X=4),\ \lambda'=6$$\approx 0.1339$
  12. 125 marksNumericalJoint probability distribution of two randAnswer

    Let X and Y be two continuous random variable having joint pdf $f(x, y) = c(x^2 + y^2)$, $0 < x < 1$, $0 < y < 1$, $= 0$ otherwise. Determine (a) the value of c. (b) $P(x < 0.5, y > 0.5)$. [5]

    $$f(x, y) = c(x^2 + y^2), \quad 0 < x < 1,\ 0 < y < 1$$ $$f(x, y) = 0 \quad \text{otherwise}$$ Required: (a) value of $c$; (b) $P(X < 0.5,\ Y 0.5)$. --- Normalization condition: $$\int{0}^{1}\int{0}^{1} c(x^2 + y^2), dx, dy = 1$$ Inner...

  13. 135 marksProbability distribution of a random variaAnswer

    Write notes on any two: a. Random variable and probability distribution b. Five number summary c. Sampling error. [5]

    --- Definition: The five number summary is a set of five descriptive statistical values that together provide a concise summary of a dataset. These five values are: Symbol Meaning ----------------- Xs Smallest value (minimum) Q1 First qu...