STA154 · TU past paper
Basic Statistics 2080 question paper
The complete TU 2080 exam paper for Basic Statistics (STA154), all 12 questions with solved model answers written to the mark scheme.
Tap a question to open its answer.
- 110 marksNumericalSimple linear regression modelHideAnswer
Regression Analysis: Blood Pressure and Age
Regression is a statistical technique used to model and analyze the relationship between a dependent variable and one or more independent variables. It enables estimation or prediction of the dependent variable's value from known values ...
- 210 marksNumericalNormal distribution properties and applicaHideAnswer
Under which situation Normal distribution will be used. The annual salaries of employees in a large company are approximately normally distributed with a mean of $50,000 and a standard deviation of $20,000. (a) What is the probability of people earns less than $40,000? (b) What is the probability of people earns between $45,000 and $65,000? (c) What is the probability of people earns more than $70,000?[10]
- Distribution: Normal (approximately) - Mean $\mu = 50{,}000$ - Standard deviation $\sigma = 20{,}000$ - Required: - (a) $P(X < 40{,}000)$ - (b) $P(45{,}000 < X < 65{,}000)$ - (c) $P(X 70{,}000)$ --- The normal distribution is applied w...
- 310 marksNumericalDescriptive versus inferential statisticsHideAnswer
Distinguish between Descriptive and Inferential Statistics
Descriptive Statistics involves organizing, summarizing, and presenting data in a meaningful way using measures like mean, median, mode, standard deviation, and visual representations. It describes the characteristics of the data itself.
Inferential Statistics involves using sample data to make conclusions, predictions, or inferences about a larger population. It uses probability theory and hypothesis testing to draw generalizations beyond the observed data.
Model Answer: Descriptive vs Inferential Statistics and Consistency Analysis
Part 1: Distinction between Descriptive and Inferential Statistics
Aspect Descriptive Statistics Inferential Statistics Definition Methods used to organize, summarize, and describe the features of a dataset Methods used to draw conclusions/generalizations about a population from a sample Purpose Present and describe data in a meaningful form Test hypotheses and make predictions beyond the data Scope Limited to the data actually collected Extends findings from sample to population Tools/Methods Mean, median, mode, range, variance, standard deviation, tables, graphs Estimation, confidence intervals, hypothesis testing, regression, probability Uncertainty No probability/uncertainty involved Involves probability and margin of error Example Average marks of 10 students Predicting all students' performance from these 10
Part 2: Consistency using Coefficient of Variation (CV)
Consistency is measured by the Coefficient of Variation. The subject with the lower CV is more consistent.
$$CV = \frac{\sigma}{\bar{x}} \times 100%$$
Given data
- Basic Statistics: 53, 55, 52, 32, 30, 60, 47, 46, 35, 58
- Computer Science: 57, 45, 24, 31, 25, 84, 43, 80, 32, 72
- n = 10 each
Basic Statistics
Mean: $$\bar{x}_1 = \frac{468}{10} = 46.8$$
Squared deviations:
$x$ $x-\bar{x}$ $(x-\bar{x})^2$ 53 6.2 38.44 55 8.2 67.24 52 5.2 27.04 32 -14.8 219.04 30 -16.8 282.24 60 13.2 174.24 47 0.2 0.04 46 -0.8 0.64 35 -11.8 139.24 58 11.2 125.44 Σ = 1073.60 $$\sigma_1 = \sqrt{\frac{1073.6}{10}} = \sqrt{107.36} = 10.36$$
$$CV_1 = \frac{10.36}{46.8}\times100 = 22.14%$$
Computer Science
Mean: $$\bar{x}_2 = \frac{493}{10} = 49.3$$
Squared deviations:
$x$ $x-\bar{x}$ $(x-\bar{x})^2$ 57 7.7 59.29 45 -4.3 18.49 24 -25.3 640.09 31 -18.3 334.89 25 -24.3 590.49 84 34.7 1204.09 43 -6.3 39.69 80 30.7 942.49 32 -17.3 299.29 72 22.7 515.29 Σ = 4644.10 Check on the sum of squared deviations:
$$59.29+18.49+640.09+334.89+590.49+1204.09+39.69+942.49+299.29+515.29 = 4644.10$$
$$\sigma_2 = \sqrt{\frac{4644.1}{10}} = \sqrt{464.41} = 21.55$$
$$CV_2 = \frac{21.55}{49.3}\times100 = 43.71%$$
Conclusion
$$CV_1 = 22.14% < CV_2 = 43.71%$$
Since Basic Statistics has the lower coefficient of variation, Basic Statistics is the more consistent subject. The marks in Basic Statistics are more uniformly clustered about the mean than those in Computer Science.
- 45 marksDefinition and role of statistics in ITHideAnswer
Describe the role of Statistics in information technology. [5]
Statistics plays a fundamental and multifaceted role in modern Information Technology, serving as a bridge between raw data and actionable insights. It provides the mathematical and analytical foundation for decision-making in IT systems...
- 55 marksNumericalBar diagrams and Pareto diagramsHideAnswer
The following table shows the mode of transport used by 400 students of a school. Represent the following information on the bar diagram and Pareto diagram.
Mode of transport Bus Bicycle On foot By car No. of students 200 100 80 20 [5]
Mode of transport Bus Bicycle On foot By car --------------- No. of students 200 100 80 20 Total students $= 200 + 100 + 80 + 20 = 400$ --- A bar diagram represents each category by a rectangular bar whose height is proportional to its f...
- 65 marksNumericalKarl Pearson correlation coefficientHideAnswer
Properties of Correlation and Analysis
- Range: The correlation coefficient always lies between $-1$ and $+1$, i.e., $-1 \le r \le +1$. 2. Independent of change of origin and scale: Correlation is unaffected by adding/subtracting a constant or multiplying/dividing by a posit...
- 75 marksNumericalRandom variables and probability functionsHideAnswer
Define random variable. A random variable $X$ has following probability function. Find, (i) the value of $C$, (ii) $P(X < 3)$, $P(X \geq 1)$, (iii) mean and variance.
X 0 1 2 3 4 5 P(X) 0.1 C 0.2 2C 0.3 C [5]
A random variable is a real-valued function that assigns a numerical value to each outcome in the sample space of a random experiment. Random variables are classified as discrete (countable values) or continuous (values over an interval)...
- 85 marksNumericalBasic probability concepts and rulesHideAnswer
The probability that an integrated circuit chip will have defective etching is 0.12, the probability that it will have crack defect is 0.29, and the probability that it has both defects is 0.07. What is the probability that a newly manufactured chip will have either an etching or crack defect. [5]
- P(Etching defect) = P(E) = 0.12 - P(Crack defect) = P(C) = 0.29 - P(Both defects) = P(E ∩ C) = 0.07 - Required: P(E ∪ C) Addition Rule of Probability: $$P(E \cup C) = P(E) + P(C) - P(E \cap C)$$ Substituting: $$P(E \cup C) = 0.12 + 0.2...
- 95 marksNumericalBinomial distributionHideAnswer
Define binomial distribution. Fit a binomial distribution on the following data:
x 0 1 2 3 4 f 28 62 46 10 4 [5]
Binomial Distribution - Definition and Fitting
STEP 1 - EXTRACT: Given Data
$x$ 0 1 2 3 4 $f$ 28 62 46 10 4 Maximum value of $x = 4 \Rightarrow n = 4$.
Definition
A binomial distribution is a discrete probability distribution giving the number of successes in $n$ independent trials, each with two outcomes (success/failure) and constant success probability $p$.
$$P(X = x) = \binom{n}{x} p^x q^{n-x}, \quad q = 1-p, \quad x = 0,1,2,\dots,n$$
STEP 2 - SOLVE
Step 1: Mean
$x$ $f$ $xf$ 0 28 0 1 62 62 2 46 92 3 10 30 4 4 16 Total N = 150 Σxf = 200 $$\bar{x} = \frac{\sum xf}{N} = \frac{200}{150} = 1.3333$$
Step 2: Parameters
$$np = \bar{x} \Rightarrow p = \frac{\bar{x}}{n} = \frac{1.3333}{4} = \frac{1}{3} = 0.3333$$ $$q = 1 - p = \frac{2}{3} = 0.6667$$
Step 3: Theoretical frequencies
$$f_x = N \binom{4}{x} p^x q^{4-x} = 150 \binom{4}{x}\left(\tfrac{1}{3}\right)^x \left(\tfrac{2}{3}\right)^{4-x}$$
- $f_0 = 150 \cdot 1 \cdot \frac{16}{81} = 150 \times 0.19753 = 29.63 \approx 30$
- $f_1 = 150 \cdot 4 \cdot \frac{1}{3}\cdot\frac{8}{27} = 150 \times 0.39506 = 59.26 \approx 59$
- $f_2 = 150 \cdot 6 \cdot \frac{1}{9}\cdot\frac{4}{9} = 150 \times 0.29630 = 44.44 \approx 44$
- $f_3 = 150 \cdot 4 \cdot \frac{1}{27}\cdot\frac{2}{3} = 150 \times 0.09877 = 14.81 \approx 15$
- $f_4 = 150 \cdot 1 \cdot \frac{1}{81} = 150 \times 0.01235 = 1.85 \approx 2$
Check: $30 + 59 + 44 + 15 + 2 = 150$ ✓
Fitted Distribution:
$x$ 0 1 2 3 4 Observed $f$ 28 62 46 10 4 Expected $f'$ 30 59 44 15 2 Parameters: $n = 4$, $p = \tfrac{1}{3}$.
- 105 marksNumericalCoefficient of skewnessHideAnswer
If the first four moments about mean are 0, 2.8, -2 and 24.5 respectively. Compute coefficient of skewness and kurtosis and comment upon result. [5]
- First moment about mean: $\mu1 = 0$ - Second moment about mean: $\mu2 = 2.8$ - Third moment about mean: $\mu3 = -2$ - Fourth moment about mean: $\mu4 = 24.5$ Using the beta coefficient: $$\beta1 = \frac{\mu3^2}{\mu2^3} = \frac{(-2)^2}{...
- 115 marksNumericalConfidence interval for population meanHideAnswer
The fuel consumption of a new model of cars is being tested. In one trial, 50 cars chosen at random were driven under the identical conditions and the distances, x km, covered on 1 liter of petrol were recorded. The results gave the following totals: Σ x = 525, Σ x² = 5625. Calculate the 99% confidence interval for the mean petrol consumption, in km per liter. Interpret the result. [5]
- Sample size: $n = 50$ - $\Sigma x = 525$ - $\Sigma x^2 = 5625$ - Confidence level: 99% $$\bar{x} = \frac{\Sigma x}{n} = \frac{525}{50} = 10.5 \text{ km/liter}$$ Using the sample estimate of variance (with $n$ divisor, common at this le...
- 125 marksBox and whisker plotsHideAnswer
Write short note on the following: a) Use of Box and whisker plot. b) Parameter and Statistic. [5]
Model Answer: Box and Whisker Plot, Parameter and Statistic
a) Use of Box and Whisker Plot
A box and whisker plot (or boxplot) is a graphical method used to display the distribution and spread of numerical data. Its main uses include:
-
Visual Summary of Data: Displays five-number summary - minimum, Q1 (first quartile), median (Q2), Q3 (third quartile), and maximum values in a single diagram.
-
Identifying Outliers: Points that fall beyond 1.5 × IQR (Interquartile Range) from Q1 or Q3 are plotted separately as outliers, making them easily identifiable.
-
Comparing Distributions: Multiple boxplots can be drawn side-by-side to compare the central tendency, spread, and skewness of different datasets or groups.
-
Assessing Skewness: The position of the median line within the box and the length of whiskers indicate whether data is symmetrically distributed or skewed.
-
Understanding Data Spread: The box width (IQR) shows where the middle 50% of data lies, while whiskers extend to show the range of typical values.
Example: In quality control, boxplots help compare product measurements across different batches to identify inconsistencies.
b) Parameter and Statistic
Parameter Statistic A numerical value that describes a characteristic of a population A numerical value that describes a characteristic of a sample Fixed and constant for a given population Varies from sample to sample Usually unknown and estimated from sample data Calculated from observed sample data Denoted by Greek letters (μ, σ, ρ) Denoted by Roman letters (x̄, s, r) Example: Population mean μ, population standard deviation σ Example: Sample mean x̄, sample standard deviation s Key Relationship: Statistics are used as estimators of population parameters. For instance, the sample mean (x̄) is used to estimate the population mean (μ).
-