2080

STA154 · TU past paper

Basic Statistics 2080 question paper

The complete TU 2080 exam paper for Basic Statistics (STA154), all 12 questions with solved model answers written to the mark scheme.

Tap a question to open its answer.

  1. 110 marksNumericalSimple linear regression modelAnswer

    Regression Analysis: Blood Pressure and Age

    Regression is a statistical technique used to model and analyze the relationship between a dependent variable and one or more independent variables. It enables estimation or prediction of the dependent variable's value from known values ...

  2. 210 marksNumericalNormal distribution properties and applicaAnswer

    Under which situation Normal distribution will be used. The annual salaries of employees in a large company are approximately normally distributed with a mean of $50,000 and a standard deviation of $20,000. (a) What is the probability of people earns less than $40,000? (b) What is the probability of people earns between $45,000 and $65,000? (c) What is the probability of people earns more than $70,000?[10]

    • Distribution: Normal (approximately) - Mean $\mu = 50{,}000$ - Standard deviation $\sigma = 20{,}000$ - Required: - (a) $P(X < 40{,}000)$ - (b) $P(45{,}000 < X < 65{,}000)$ - (c) $P(X 70{,}000)$ --- The normal distribution is applied w...
  3. 310 marksNumericalDescriptive versus inferential statisticsAnswer

    Distinguish between Descriptive and Inferential Statistics

    Descriptive Statistics involves organizing, summarizing, and presenting data in a meaningful way using measures like mean, median, mode, standard deviation, and visual representations. It describes the characteristics of the data itself.

    Inferential Statistics involves using sample data to make conclusions, predictions, or inferences about a larger population. It uses probability theory and hypothesis testing to draw generalizations beyond the observed data.

    Model Answer: Descriptive vs Inferential Statistics and Consistency Analysis

    Part 1: Distinction between Descriptive and Inferential Statistics

    AspectDescriptive StatisticsInferential Statistics
    DefinitionMethods used to organize, summarize, and describe the features of a datasetMethods used to draw conclusions/generalizations about a population from a sample
    PurposePresent and describe data in a meaningful formTest hypotheses and make predictions beyond the data
    ScopeLimited to the data actually collectedExtends findings from sample to population
    Tools/MethodsMean, median, mode, range, variance, standard deviation, tables, graphsEstimation, confidence intervals, hypothesis testing, regression, probability
    UncertaintyNo probability/uncertainty involvedInvolves probability and margin of error
    ExampleAverage marks of 10 studentsPredicting all students' performance from these 10

    Part 2: Consistency using Coefficient of Variation (CV)

    Consistency is measured by the Coefficient of Variation. The subject with the lower CV is more consistent.

    $$CV = \frac{\sigma}{\bar{x}} \times 100%$$

    Given data

    • Basic Statistics: 53, 55, 52, 32, 30, 60, 47, 46, 35, 58
    • Computer Science: 57, 45, 24, 31, 25, 84, 43, 80, 32, 72
    • n = 10 each

    Basic Statistics

    Mean: $$\bar{x}_1 = \frac{468}{10} = 46.8$$

    Squared deviations:

    $x$$x-\bar{x}$$(x-\bar{x})^2$
    536.238.44
    558.267.24
    525.227.04
    32-14.8219.04
    30-16.8282.24
    6013.2174.24
    470.20.04
    46-0.80.64
    35-11.8139.24
    5811.2125.44
    Σ = 1073.60

    $$\sigma_1 = \sqrt{\frac{1073.6}{10}} = \sqrt{107.36} = 10.36$$

    $$CV_1 = \frac{10.36}{46.8}\times100 = 22.14%$$


    Computer Science

    Mean: $$\bar{x}_2 = \frac{493}{10} = 49.3$$

    Squared deviations:

    $x$$x-\bar{x}$$(x-\bar{x})^2$
    577.759.29
    45-4.318.49
    24-25.3640.09
    31-18.3334.89
    25-24.3590.49
    8434.71204.09
    43-6.339.69
    8030.7942.49
    32-17.3299.29
    7222.7515.29
    Σ = 4644.10

    Check on the sum of squared deviations:

    $$59.29+18.49+640.09+334.89+590.49+1204.09+39.69+942.49+299.29+515.29 = 4644.10$$

    $$\sigma_2 = \sqrt{\frac{4644.1}{10}} = \sqrt{464.41} = 21.55$$

    $$CV_2 = \frac{21.55}{49.3}\times100 = 43.71%$$


    Conclusion

    $$CV_1 = 22.14% < CV_2 = 43.71%$$

    Since Basic Statistics has the lower coefficient of variation, Basic Statistics is the more consistent subject. The marks in Basic Statistics are more uniformly clustered about the mean than those in Computer Science.

  4. 45 marksDefinition and role of statistics in ITAnswer

    Describe the role of Statistics in information technology. [5]

    Statistics plays a fundamental and multifaceted role in modern Information Technology, serving as a bridge between raw data and actionable insights. It provides the mathematical and analytical foundation for decision-making in IT systems...

  5. 55 marksNumericalBar diagrams and Pareto diagramsAnswer

    The following table shows the mode of transport used by 400 students of a school. Represent the following information on the bar diagram and Pareto diagram.

    Mode of transportBusBicycleOn footBy car
    No. of students2001008020

    [5]

    Mode of transport Bus Bicycle On foot By car --------------- No. of students 200 100 80 20 Total students $= 200 + 100 + 80 + 20 = 400$ --- A bar diagram represents each category by a rectangular bar whose height is proportional to its f...

  6. 65 marksNumericalKarl Pearson correlation coefficientAnswer

    Properties of Correlation and Analysis

    1. Range: The correlation coefficient always lies between $-1$ and $+1$, i.e., $-1 \le r \le +1$. 2. Independent of change of origin and scale: Correlation is unaffected by adding/subtracting a constant or multiplying/dividing by a posit...
  7. 75 marksNumericalRandom variables and probability functionsAnswer

    Define random variable. A random variable $X$ has following probability function. Find, (i) the value of $C$, (ii) $P(X < 3)$, $P(X \geq 1)$, (iii) mean and variance.

    X012345
    P(X)0.1C0.22C0.3C

    [5]

    A random variable is a real-valued function that assigns a numerical value to each outcome in the sample space of a random experiment. Random variables are classified as discrete (countable values) or continuous (values over an interval)...

  8. 85 marksNumericalBasic probability concepts and rulesAnswer

    The probability that an integrated circuit chip will have defective etching is 0.12, the probability that it will have crack defect is 0.29, and the probability that it has both defects is 0.07. What is the probability that a newly manufactured chip will have either an etching or crack defect. [5]

    • P(Etching defect) = P(E) = 0.12 - P(Crack defect) = P(C) = 0.29 - P(Both defects) = P(E ∩ C) = 0.07 - Required: P(E ∪ C) Addition Rule of Probability: $$P(E \cup C) = P(E) + P(C) - P(E \cap C)$$ Substituting: $$P(E \cup C) = 0.12 + 0.2...
  9. 95 marksNumericalBinomial distributionAnswer

    Define binomial distribution. Fit a binomial distribution on the following data:

    x01234
    f286246104

    [5]

    Binomial Distribution - Definition and Fitting

    STEP 1 - EXTRACT: Given Data

    $x$01234
    $f$286246104

    Maximum value of $x = 4 \Rightarrow n = 4$.

    Definition

    A binomial distribution is a discrete probability distribution giving the number of successes in $n$ independent trials, each with two outcomes (success/failure) and constant success probability $p$.

    $$P(X = x) = \binom{n}{x} p^x q^{n-x}, \quad q = 1-p, \quad x = 0,1,2,\dots,n$$

    STEP 2 - SOLVE

    Step 1: Mean

    $x$$f$$xf$
    0280
    16262
    24692
    31030
    4416
    TotalN = 150Σxf = 200

    $$\bar{x} = \frac{\sum xf}{N} = \frac{200}{150} = 1.3333$$

    Step 2: Parameters

    $$np = \bar{x} \Rightarrow p = \frac{\bar{x}}{n} = \frac{1.3333}{4} = \frac{1}{3} = 0.3333$$ $$q = 1 - p = \frac{2}{3} = 0.6667$$

    Step 3: Theoretical frequencies

    $$f_x = N \binom{4}{x} p^x q^{4-x} = 150 \binom{4}{x}\left(\tfrac{1}{3}\right)^x \left(\tfrac{2}{3}\right)^{4-x}$$

    • $f_0 = 150 \cdot 1 \cdot \frac{16}{81} = 150 \times 0.19753 = 29.63 \approx 30$
    • $f_1 = 150 \cdot 4 \cdot \frac{1}{3}\cdot\frac{8}{27} = 150 \times 0.39506 = 59.26 \approx 59$
    • $f_2 = 150 \cdot 6 \cdot \frac{1}{9}\cdot\frac{4}{9} = 150 \times 0.29630 = 44.44 \approx 44$
    • $f_3 = 150 \cdot 4 \cdot \frac{1}{27}\cdot\frac{2}{3} = 150 \times 0.09877 = 14.81 \approx 15$
    • $f_4 = 150 \cdot 1 \cdot \frac{1}{81} = 150 \times 0.01235 = 1.85 \approx 2$

    Check: $30 + 59 + 44 + 15 + 2 = 150$ ✓

    Fitted Distribution:

    $x$01234
    Observed $f$286246104
    Expected $f'$305944152

    Parameters: $n = 4$, $p = \tfrac{1}{3}$.

  10. 105 marksNumericalCoefficient of skewnessAnswer

    If the first four moments about mean are 0, 2.8, -2 and 24.5 respectively. Compute coefficient of skewness and kurtosis and comment upon result. [5]

    • First moment about mean: $\mu1 = 0$ - Second moment about mean: $\mu2 = 2.8$ - Third moment about mean: $\mu3 = -2$ - Fourth moment about mean: $\mu4 = 24.5$ Using the beta coefficient: $$\beta1 = \frac{\mu3^2}{\mu2^3} = \frac{(-2)^2}{...
  11. 115 marksNumericalConfidence interval for population meanAnswer

    The fuel consumption of a new model of cars is being tested. In one trial, 50 cars chosen at random were driven under the identical conditions and the distances, x km, covered on 1 liter of petrol were recorded. The results gave the following totals: Σ x = 525, Σ x² = 5625. Calculate the 99% confidence interval for the mean petrol consumption, in km per liter. Interpret the result. [5]

    • Sample size: $n = 50$ - $\Sigma x = 525$ - $\Sigma x^2 = 5625$ - Confidence level: 99% $$\bar{x} = \frac{\Sigma x}{n} = \frac{525}{50} = 10.5 \text{ km/liter}$$ Using the sample estimate of variance (with $n$ divisor, common at this le...
  12. 125 marksBox and whisker plotsAnswer

    Write short note on the following: a) Use of Box and whisker plot. b) Parameter and Statistic. [5]

    Model Answer: Box and Whisker Plot, Parameter and Statistic

    a) Use of Box and Whisker Plot

    A box and whisker plot (or boxplot) is a graphical method used to display the distribution and spread of numerical data. Its main uses include:

    • Visual Summary of Data: Displays five-number summary - minimum, Q1 (first quartile), median (Q2), Q3 (third quartile), and maximum values in a single diagram.

    • Identifying Outliers: Points that fall beyond 1.5 × IQR (Interquartile Range) from Q1 or Q3 are plotted separately as outliers, making them easily identifiable.

    • Comparing Distributions: Multiple boxplots can be drawn side-by-side to compare the central tendency, spread, and skewness of different datasets or groups.

    • Assessing Skewness: The position of the median line within the box and the length of whiskers indicate whether data is symmetrically distributed or skewed.

    • Understanding Data Spread: The box width (IQR) shows where the middle 50% of data lies, while whiskers extend to show the range of typical values.

    Example: In quality control, boxplots help compare product measurements across different batches to identify inconsistencies.


    b) Parameter and Statistic

    ParameterStatistic
    A numerical value that describes a characteristic of a populationA numerical value that describes a characteristic of a sample
    Fixed and constant for a given populationVaries from sample to sample
    Usually unknown and estimated from sample dataCalculated from observed sample data
    Denoted by Greek letters (μ, σ, ρ)Denoted by Roman letters (x̄, s, r)
    Example: Population mean μ, population standard deviation σExample: Sample mean x̄, sample standard deviation s

    Key Relationship: Statistics are used as estimators of population parameters. For instance, the sample mean (x̄) is used to estimate the population mean (μ).