STA169 · Exam intelligence
Statistics I important questions
From 6 past TU papers: which questions keep coming back, how much they carry, and what is most likely to show up next. Every question links to a model answer.
Most likely in the next examStatistical
Ranked by how often a topic is asked, its marks weight, and whether it is due after skipping the 2080.1 paper. No guarantees; study the whole syllabus.
1asked 11xavg 6 marks · Measures of central tendencyAnswerHideDefine statistics and discuss its importance in the field of computational sciences. The following are the numbers of minutes that a person had to wait for the bus to work on 20 working days: 15, 10, 2, 17, 5, 8, 3, 10, 2, 9, 5, 9, 13, 1, 10, 12, 5, 10, 8, 4. Compute mean, median, mode, standard, variance and coefficient of variation.[10]
Define statistics and discuss its importance in the field of computational sciences. The following are the numbers of minutes that a person had to wait for the bus to work on 20 working days: 15, 10, 2, 17, 5, 8, 3, 10, 2, 9, 5, 9, 13, 1, 10, 12, 5, 10, 8, 4. Compute mean, median, mode, standard, variance and coefficient of variation.[10]
Statistics: Definition, Importance, and Computations
Part 1: Definition of Statistics (2 marks)
Statistics is the branch of science that deals with the collection, organization, presentation, analysis, and interpretation of numerical data in order to draw valid conclusions and make rational decisions under conditions of uncertainty.
Part 2: Importance of Statistics in Computational Sciences (3 marks)
- Fast computation: Computers perform statistical operations (mean, variance, ANOVA, regression) rapidly and reliably, removing tedious hand calculation.
- Handling large/multivariate data: Complex methods like multivariate analysis, clustering, and linear programming become feasible.
- Basis of modern techniques: Data mining, machine learning, pattern recognition, bootstrap methods, and image analysis are inherently statistical and computational.
- Research and reliability: Enables reproducible, high-speed experimental research and simulation.
- Practical CS/IT problem solving: Reliability testing (e.g., expected lifetime of hardware), performance benchmarking, and predictive modeling all depend on statistics.
Part 3: Statistical Computations (5 marks)
Given data (n = 20)
$$15, 10, 2, 17, 5, 8, 3, 10, 2, 9, 5, 9, 13, 1, 10, 12, 5, 10, 8, 4$$
Sorted: $$1, 2, 2, 3, 4, 5, 5, 5, 8, 8, 9, 9, 10, 10, 10, 10, 12, 13, 15, 17$$
Mean
$$\sum X = 158$$ $$\bar{X} = \frac{158}{20} = 7.9 \text{ minutes}$$
Median (n even → average of 10th and 11th values)
10th = 8, 11th = 9 $$\text{Median} = \frac{8+9}{2} = 8.5 \text{ minutes}$$
Mode
Value 10 occurs 4 times (highest frequency). $$\text{Mode} = 10 \text{ minutes}$$
Variance and Standard Deviation
$$\sum X^2 = 1626$$
$$\sigma^2 = \frac{\sum X^2}{n} - \bar{X}^2 = \frac{1626}{20} - (7.9)^2$$ $$= 81.3 - 62.41 = 18.89$$
$$\sigma = \sqrt{18.89} = 4.3463 \approx 4.35 \text{ minutes}$$
(Note: if the sample formula with $n-1$ is used: $$s^2 = \frac{\sum X^2 - n\bar{X}^2}{n-1} = \frac{1626 - 1248.2}{19} = \frac{377.8}{19} = 19.8842$$ $$s = 4.4592 \text{ minutes})$$
Coefficient of Variation
$$CV = \frac{\sigma}{\bar{X}} \times 100 = \frac{4.3463}{7.9} \times 100 = 55.02%$$
(Using sample SD: $CV = \frac{4.4592}{7.9}\times100 = 56.45%$)
Final Summary
| Measure | Value |
|---|---|
| Mean | 7.9 min |
| Median | 8.5 min |
| Mode | 10 min |
| Variance (population) | 18.89 |
| Standard Deviation (population) | 4.35 min |
| Coefficient of Variation | 55.02% |
2asked 11xavg 6 marks · Discrete distributionsAnswerHideFitting a Binomial Distribution
Under what condition binomial probability distribution? Five unbiased coins are tossed 100 times and the following results were obtained. Fit the binomial distribution.
No of heads 0 1 2 3 4 5 Frequency 5 24 35 22 10 4
[5]
Fitting a Binomial Distribution
Under what condition binomial probability distribution? Five unbiased coins are tossed 100 times and the following results were obtained. Fit the binomial distribution.
| No of heads | 0 | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|---|
| Frequency | 5 | 24 | 35 | 22 | 10 | 4 |
[5]
The binomial probability distribution applies under the following conditions: - Each trial results in only two mutually exclusive outcomes (success/failure). - The number of trials n is fixed and finite. - The trials are independent of o...
3asked 7xavg 6 marks · Joint probability distribution of two random variablesAnswerHideLet X and Y be two continuous random variable having joint pdf $f(x, y) = c(x^2 + y^2)$, $0 < x < 1$, $0 < y < 1$, $= 0$ otherwise. Determine
(a) the value of c.
(b) $P(x < 0.5, y > 0.5)$. [5]
Let X and Y be two continuous random variable having joint pdf $f(x, y) = c(x^2 + y^2)$, $0 < x < 1$, $0 < y < 1$, $= 0$ otherwise. Determine
(a) the value of c.
(b) $P(x < 0.5, y > 0.5)$. [5]
$$f(x, y) = c(x^2 + y^2), \quad 0 < x < 1,\ 0 < y < 1$$ $$f(x, y) = 0 \quad \text{otherwise}$$ Required: (a) value of $c$; (b) $P(X < 0.5,\ Y 0.5)$. --- Normalization condition: $$\int{0}^{1}\int{0}^{1} c(x^2 + y^2), dx, dy = 1$$ Inner...
4asked 6xavg 8 marks · Measures of dispersionAnswerHideMeasurement of Dispersion and Consistency Analysis
Measurement of Dispersion and Consistency Analysis
Measurement of dispersion refers to statistical measures describing the spread or variability of data values around a central value. It indicates how much individual observations deviate from the average. Common measures: Range, Mean Dev...
5asked 6xavg 7 marks · Continuous distributionAnswerHideDefine normal distribution. What are the main characteristics of normal distribution? Extruded plastic rods are automatically cut into length 5 inches. Actual length are normally distributed about a mean of 5 inches and their standard deviation is 0.05 inches.
(i) What proportion of rods exceed tolerance limits of 4.9 inches to 5.1 inches?
(ii) Proportion of rods having tolerance rod which is greater than 6.5 inches.[10]
Define normal distribution. What are the main characteristics of normal distribution? Extruded plastic rods are automatically cut into length 5 inches. Actual length are normally distributed about a mean of 5 inches and their standard deviation is 0.05 inches.
(i) What proportion of rods exceed tolerance limits of 4.9 inches to 5.1 inches?
(ii) Proportion of rods having tolerance rod which is greater than 6.5 inches.[10]
Normal Distribution: Definition, Characteristics, and Numerical
Definition
A normal distribution is a continuous probability distribution, symmetrical and bell-shaped about its mean, defined by two parameters, mean $\mu$ and standard deviation $\sigma$. Its probability density function is:
$$f(x) = \frac{1}{\sigma\sqrt{2\pi}} e^{-\frac{1}{2}\left(\frac{x-\mu}{\sigma}\right)^2}, \quad -\infty < x < \infty$$
Main Characteristics
- It is a continuous distribution with parameters $\mu$ (mean) and $\sigma^2$ (variance).
- The curve is bell-shaped and symmetrical about the mean $\mu$.
- Mean = Median = Mode (all coincide at the centre).
- Maximum ordinate occurs at $x = \mu$, equal to $\frac{1}{\sigma\sqrt{2\pi}}$.
- Skewness $\beta_1 = 0$ (perfectly symmetric).
- Kurtosis $\beta_2 = 3$ (mesokurtic).
- Empirical area rule:
- $P(\mu-\sigma < X < \mu+\sigma) = 0.6826$
- $P(\mu-2\sigma < X < \mu+2\sigma) = 0.9544$
- $P(\mu-3\sigma < X < \mu+3\sigma) = 0.9974$
- Total area under the curve = 1.
- Asymptotic: the curve never touches the X-axis.
Given Data
- Mean $\mu = 5$ inches
- Standard deviation $\sigma = 0.05$ inches
- $X \sim N(5, 0.05^2)$
- Standard normal variate: $Z = \dfrac{X - 5}{0.05}$
Part (i): Proportion exceeding tolerance limits 4.9 to 5.1 inches
Z-scores:
$$Z_1 = \frac{4.9 - 5}{0.05} = -2, \qquad Z_2 = \frac{5.1 - 5}{0.05} = +2$$
Proportion within limits:
$$P(4.9 < X < 5.1) = P(-2 < Z < 2) = 2 \times P(0 < Z < 2) = 2 \times 0.4772 = 0.9544$$
Proportion exceeding limits:
$$P(X < 4.9 \text{ or } X > 5.1) = 1 - 0.9544 = \boxed{0.0456 \ (4.56%)}$$
Part (ii): Proportion having length greater than 6.5 inches
Z-score:
$$Z = \frac{6.5 - 5}{0.05} = \frac{1.5}{0.05} = 30$$
Since $Z = 30$ lies far beyond the standard normal table (which effectively ends near $Z \approx 4$):
$$P(X > 6.5) = P(Z > 30) \approx \boxed{0}$$
A rod exceeding 6.5 inches is 30 standard deviations above the mean, so the proportion is practically zero.
Summary
| Part | Condition | Z-value(s) | Proportion |
|---|---|---|---|
| (i) | X < 4.9 or X > 5.1 | $\pm 2$ | 0.0456 (4.56%) |
| (ii) | X > 6.5 | 30 | $\approx 0$ |
Most repeated questions
Topics asked at least twice, most-asked first.
asked 11xavg 6 marks · 2080.1, 2080, 2079, 2078, 2076...AnswerHideDefine statistics and discuss its importance in the field of computational sciences. The following are the numbers of minutes that a person had to wait for the bus to work on 20 working days: 15, 10, 2, 17, 5, 8, 3, 10, 2, 9, 5, 9, 13, 1, 10, 12, 5, 10, 8, 4. Compute mean, median, mode, standard, variance and coefficient of variation.[10]
Define statistics and discuss its importance in the field of computational sciences. The following are the numbers of minutes that a person had to wait for the bus to work on 20 working days: 15, 10, 2, 17, 5, 8, 3, 10, 2, 9, 5, 9, 13, 1, 10, 12, 5, 10, 8, 4. Compute mean, median, mode, standard, variance and coefficient of variation.[10]
Statistics: Definition, Importance, and Computations
Part 1: Definition of Statistics (2 marks)
Statistics is the branch of science that deals with the collection, organization, presentation, analysis, and interpretation of numerical data in order to draw valid conclusions and make rational decisions under conditions of uncertainty.
Part 2: Importance of Statistics in Computational Sciences (3 marks)
- Fast computation: Computers perform statistical operations (mean, variance, ANOVA, regression) rapidly and reliably, removing tedious hand calculation.
- Handling large/multivariate data: Complex methods like multivariate analysis, clustering, and linear programming become feasible.
- Basis of modern techniques: Data mining, machine learning, pattern recognition, bootstrap methods, and image analysis are inherently statistical and computational.
- Research and reliability: Enables reproducible, high-speed experimental research and simulation.
- Practical CS/IT problem solving: Reliability testing (e.g., expected lifetime of hardware), performance benchmarking, and predictive modeling all depend on statistics.
Part 3: Statistical Computations (5 marks)
Given data (n = 20)
$$15, 10, 2, 17, 5, 8, 3, 10, 2, 9, 5, 9, 13, 1, 10, 12, 5, 10, 8, 4$$
Sorted: $$1, 2, 2, 3, 4, 5, 5, 5, 8, 8, 9, 9, 10, 10, 10, 10, 12, 13, 15, 17$$
Mean
$$\sum X = 158$$ $$\bar{X} = \frac{158}{20} = 7.9 \text{ minutes}$$
Median (n even → average of 10th and 11th values)
10th = 8, 11th = 9 $$\text{Median} = \frac{8+9}{2} = 8.5 \text{ minutes}$$
Mode
Value 10 occurs 4 times (highest frequency). $$\text{Mode} = 10 \text{ minutes}$$
Variance and Standard Deviation
$$\sum X^2 = 1626$$
$$\sigma^2 = \frac{\sum X^2}{n} - \bar{X}^2 = \frac{1626}{20} - (7.9)^2$$ $$= 81.3 - 62.41 = 18.89$$
$$\sigma = \sqrt{18.89} = 4.3463 \approx 4.35 \text{ minutes}$$
(Note: if the sample formula with $n-1$ is used: $$s^2 = \frac{\sum X^2 - n\bar{X}^2}{n-1} = \frac{1626 - 1248.2}{19} = \frac{377.8}{19} = 19.8842$$ $$s = 4.4592 \text{ minutes})$$
Coefficient of Variation
$$CV = \frac{\sigma}{\bar{X}} \times 100 = \frac{4.3463}{7.9} \times 100 = 55.02%$$
(Using sample SD: $CV = \frac{4.4592}{7.9}\times100 = 56.45%$)
Final Summary
| Measure | Value |
|---|---|
| Mean | 7.9 min |
| Median | 8.5 min |
| Mode | 10 min |
| Variance (population) | 18.89 |
| Standard Deviation (population) | 4.35 min |
| Coefficient of Variation | 55.02% |
asked 11xavg 6 marks · 2080.1, 2080, 2079, 2078, 2076...AnswerHideFitting a Binomial Distribution
Under what condition binomial probability distribution? Five unbiased coins are tossed 100 times and the following results were obtained. Fit the binomial distribution.
No of heads 0 1 2 3 4 5 Frequency 5 24 35 22 10 4
[5]
Fitting a Binomial Distribution
Under what condition binomial probability distribution? Five unbiased coins are tossed 100 times and the following results were obtained. Fit the binomial distribution.
| No of heads | 0 | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|---|
| Frequency | 5 | 24 | 35 | 22 | 10 | 4 |
[5]
The binomial probability distribution applies under the following conditions: - Each trial results in only two mutually exclusive outcomes (success/failure). - The number of trials n is fixed and finite. - The trials are independent of o...
asked 7xavg 6 marks · 2080.1, 2080, 2078, 2076, 2075AnswerHideLet X and Y be two continuous random variable having joint pdf $f(x, y) = c(x^2 + y^2)$, $0 < x < 1$, $0 < y < 1$, $= 0$ otherwise. Determine
(a) the value of c.
(b) $P(x < 0.5, y > 0.5)$. [5]
Let X and Y be two continuous random variable having joint pdf $f(x, y) = c(x^2 + y^2)$, $0 < x < 1$, $0 < y < 1$, $= 0$ otherwise. Determine
(a) the value of c.
(b) $P(x < 0.5, y > 0.5)$. [5]
$$f(x, y) = c(x^2 + y^2), \quad 0 < x < 1,\ 0 < y < 1$$ $$f(x, y) = 0 \quad \text{otherwise}$$ Required: (a) value of $c$; (b) $P(X < 0.5,\ Y 0.5)$. --- Normalization condition: $$\int{0}^{1}\int{0}^{1} c(x^2 + y^2), dx, dy = 1$$ Inner...
asked 6xavg 8 marks · 2080.1, 2079, 2078, 2076, 2075AnswerHideMeasurement of Dispersion and Consistency Analysis
Measurement of Dispersion and Consistency Analysis
Measurement of dispersion refers to statistical measures describing the spread or variability of data values around a central value. It indicates how much individual observations deviate from the average. Common measures: Range, Mean Dev...
asked 6xavg 7 marks · 2080.1, 2080, 2079, 2078, 2076...AnswerHideDefine normal distribution. What are the main characteristics of normal distribution? Extruded plastic rods are automatically cut into length 5 inches. Actual length are normally distributed about a mean of 5 inches and their standard deviation is 0.05 inches.
(i) What proportion of rods exceed tolerance limits of 4.9 inches to 5.1 inches?
(ii) Proportion of rods having tolerance rod which is greater than 6.5 inches.[10]
Define normal distribution. What are the main characteristics of normal distribution? Extruded plastic rods are automatically cut into length 5 inches. Actual length are normally distributed about a mean of 5 inches and their standard deviation is 0.05 inches.
(i) What proportion of rods exceed tolerance limits of 4.9 inches to 5.1 inches?
(ii) Proportion of rods having tolerance rod which is greater than 6.5 inches.[10]
Normal Distribution: Definition, Characteristics, and Numerical
Definition
A normal distribution is a continuous probability distribution, symmetrical and bell-shaped about its mean, defined by two parameters, mean $\mu$ and standard deviation $\sigma$. Its probability density function is:
$$f(x) = \frac{1}{\sigma\sqrt{2\pi}} e^{-\frac{1}{2}\left(\frac{x-\mu}{\sigma}\right)^2}, \quad -\infty < x < \infty$$
Main Characteristics
- It is a continuous distribution with parameters $\mu$ (mean) and $\sigma^2$ (variance).
- The curve is bell-shaped and symmetrical about the mean $\mu$.
- Mean = Median = Mode (all coincide at the centre).
- Maximum ordinate occurs at $x = \mu$, equal to $\frac{1}{\sigma\sqrt{2\pi}}$.
- Skewness $\beta_1 = 0$ (perfectly symmetric).
- Kurtosis $\beta_2 = 3$ (mesokurtic).
- Empirical area rule:
- $P(\mu-\sigma < X < \mu+\sigma) = 0.6826$
- $P(\mu-2\sigma < X < \mu+2\sigma) = 0.9544$
- $P(\mu-3\sigma < X < \mu+3\sigma) = 0.9974$
- Total area under the curve = 1.
- Asymptotic: the curve never touches the X-axis.
Given Data
- Mean $\mu = 5$ inches
- Standard deviation $\sigma = 0.05$ inches
- $X \sim N(5, 0.05^2)$
- Standard normal variate: $Z = \dfrac{X - 5}{0.05}$
Part (i): Proportion exceeding tolerance limits 4.9 to 5.1 inches
Z-scores:
$$Z_1 = \frac{4.9 - 5}{0.05} = -2, \qquad Z_2 = \frac{5.1 - 5}{0.05} = +2$$
Proportion within limits:
$$P(4.9 < X < 5.1) = P(-2 < Z < 2) = 2 \times P(0 < Z < 2) = 2 \times 0.4772 = 0.9544$$
Proportion exceeding limits:
$$P(X < 4.9 \text{ or } X > 5.1) = 1 - 0.9544 = \boxed{0.0456 \ (4.56%)}$$
Part (ii): Proportion having length greater than 6.5 inches
Z-score:
$$Z = \frac{6.5 - 5}{0.05} = \frac{1.5}{0.05} = 30$$
Since $Z = 30$ lies far beyond the standard normal table (which effectively ends near $Z \approx 4$):
$$P(X > 6.5) = P(Z > 30) \approx \boxed{0}$$
A rod exceeding 6.5 inches is 30 standard deviations above the mean, so the proportion is practically zero.
Summary
| Part | Condition | Z-value(s) | Proportion |
|---|---|---|---|
| (i) | X < 4.9 or X > 5.1 | $\pm 2$ | 0.0456 (4.56%) |
| (ii) | X > 6.5 | 30 | $\approx 0$ |
asked 6xavg 5 marks · 2080.1, 2080, 2079, 2078, 2076...AnswerHideWhat is sampling? Explain the main purpose of sampling. Describe briefly stratified sampling. [5]
What is sampling? Explain the main purpose of sampling. Describe briefly stratified sampling. [5]
Sampling: Definition, Purpose, and Stratified Sampling
1. What is Sampling?
Sampling is the process of selecting a subset (called a sample) of units from a larger group (called the population) in order to draw conclusions or make inferences about the entire population. Instead of studying every unit in the population (census), sampling allows us to study only a representative portion, saving time, cost, and effort.
The difference between the value of the sample statistic obtained from a sample and the value of the corresponding population parameter is called the sampling error.
2. Main Purpose of Sampling
The main purposes of sampling are:
- Economy: Studying the entire population is costly; sampling reduces expenditure of money, time, and manpower.
- Feasibility: In many cases, examining every unit is practically impossible (e.g., testing every light bulb for lifespan).
- Speed: Results can be obtained more quickly from a sample than from a complete census.
- Accuracy: A carefully conducted sample survey can sometimes be more accurate than a full census, since more attention can be given to each selected unit.
- Defining survey objectives: The first step in sampling is to clearly define the objective of the survey, commensurate with available resources in terms of money, manpower, and time.
3. Stratified Random Sampling
Definition
When the units in the population are not similar in nature, the population is first divided into sub-groups called strata, and then a simple random sample is drawn from each stratum in proportion to its size. This technique is known as Stratified Random Sampling.
Conditions for Stratification
The strata must satisfy the following conditions:
- The strata should be non-overlapping and together should comprise the whole population.
- The strata should be as homogeneous within groups (similar units inside each stratum) and heterogeneous between groups (different units across strata) as possible.
Purpose of Stratification
According to the notes, the purposes of stratification are:
- To make the sample more representative of the population.
- For greater accuracy in the results.
- For administrative convenience in conducting the survey.
Example
Suppose we want to study the income level of people in a city. The population can be stratified into groups such as:
- Low income group
- Middle income group
- High income group
A random sample is then drawn from each group separately, ensuring all income levels are properly represented in the final sample.
Advantage over Simple Random Sampling
For a given level of precision, stratified random sampling usually requires a smaller sample size as compared to simple random sampling, making it more efficient.
Summary Table:
| Feature | Stratified Sampling |
|---|---|
| Population divided into | Strata (sub-groups) |
| Within strata | Homogeneous |
| Between strata | Heterogeneous |
| Sampling from each stratum | Simple Random Sampling |
| Main benefit | More representative and accurate |
asked 4xavg 5 marks · 2080.1, 2080, 2079, 2076AnswerHideDefine Baye's theorem. Store A, B and C have 100, 75 and 50 employees and, respectively 70, 60 and 50 percent of these are women. Registration are equally likely among all employees regardless of sex. One employee resigns, and this is woman. What is the probability that she works in store B? [5]
Define Baye's theorem. Store A, B and C have 100, 75 and 50 employees and, respectively 70, 60 and 50 percent of these are women. Registration are equally likely among all employees regardless of sex. One employee resigns, and this is woman. What is the probability that she works in store B? [5]
Bayes' Theorem - Definition and Application
STEP 1: Given Data
| Store | Employees | % Women | Women Count |
|---|---|---|---|
| A | 100 | 70% | 70 |
| B | 75 | 60% | 45 |
| C | 50 | 50% | 25 |
| Total | 225 | 140 |
Resignations equally likely among all employees. Observed event: the resigning employee is a woman. Find $P(\text{Store B} \mid \text{Woman})$.
Definition of Bayes' Theorem
If $B_1, B_2, \dots, B_n$ are mutually exclusive and exhaustive events with $P(B_i) > 0$, and $A$ is any event with $P(A) > 0$, then:
$$P(B_i \mid A) = \frac{P(B_i), P(A \mid B_i)}{\sum_{j=1}^{n} P(B_j), P(A \mid B_j)}$$
- $P(B_i)$ = prior probability
- $P(A \mid B_i)$ = likelihood
- $P(B_i \mid A)$ = posterior probability
STEP 2: Solution
Define Events
- $S_A, S_B, S_C$ = employee works in Store A, B, C
- $W$ = resigning employee is a woman
Prior Probabilities (equal likelihood over 225 employees)
$$P(S_A) = \frac{100}{225}, \quad P(S_B) = \frac{75}{225}, \quad P(S_C) = \frac{50}{225}$$
Likelihoods
$$P(W \mid S_A) = 0.70, \quad P(W \mid S_B) = 0.60, \quad P(W \mid S_C) = 0.50$$
Total Probability of a Woman Resigning
$$P(W) = \frac{100}{225}(0.70) + \frac{75}{225}(0.60) + \frac{50}{225}(0.50)$$
$$P(W) = \frac{70 + 45 + 25}{225} = \frac{140}{225}$$
Apply Bayes' Theorem
$$P(S_B \mid W) = \frac{P(S_B), P(W \mid S_B)}{P(W)} = \frac{\dfrac{75}{225}(0.60)}{\dfrac{140}{225}} = \frac{45}{140}$$
$$\boxed{P(S_B \mid W) = \frac{45}{140} = \frac{9}{28} \approx 0.3214 \approx 32.14%}$$
Conclusion
The probability that the resigning woman works in Store B is approximately 0.3214 (32.14%).
asked 3xavg 5 marks · 2080, 2079, 2076AnswerHideDefine random variable. If X is the number of points rolled with a balanced die, find E(X) and variance of X. Also find the expected value of random variable g(X) = 2X^2 + 1. [5]
Define random variable. If X is the number of points rolled with a balanced die, find E(X) and variance of X. Also find the expected value of random variable g(X) = 2X^2 + 1. [5]
Random Variable: Definition, E(X), Variance, and E(g(X))
Step 1 - Extract (Given Data)
- Experiment: rolling a balanced (fair) die
- Random variable $X$ = number of points shown
- Possible values: $x = 1, 2, 3, 4, 5, 6$
- $P(X = x) = \frac{1}{6}$ for each value (balanced die)
- Function required: $g(X) = 2X^2 + 1$
All data present and sufficient.
Step 2 - Solve
Definition of Random Variable
A random variable is a real-valued function that assigns a numerical value to each outcome in the sample space of a random experiment. It may be discrete (countable values) or continuous (values over an interval).
Probability Distribution
$$P(X = x) = \frac{1}{6}, \quad x = 1, 2, 3, 4, 5, 6$$
Calculation of E(X)
$$E(X) = \sum x \cdot P(X=x) = \frac{1+2+3+4+5+6}{6} = \frac{21}{6} = \frac{7}{2} = 3.5$$
Calculation of Variance
$$E(X^2) = \frac{1^2+2^2+3^2+4^2+5^2+6^2}{6} = \frac{1+4+9+16+25+36}{6} = \frac{91}{6}$$
$$\text{Var}(X) = E(X^2) - [E(X)]^2 = \frac{91}{6} - \left(\frac{7}{2}\right)^2 = \frac{91}{6} - \frac{49}{4}$$
$$= \frac{182}{12} - \frac{147}{12} = \frac{35}{12} \approx 2.917$$
Expected Value of g(X) = 2X² + 1
$$E(2X^2 + 1) = 2E(X^2) + 1 = 2 \cdot \frac{91}{6} + 1 = \frac{182}{6} + \frac{6}{6} = \frac{188}{6} = \frac{94}{3} \approx 31.33$$
Summary
| Quantity | Value |
|---|---|
| $E(X)$ | $\frac{7}{2} = 3.5$ |
| $\text{Var}(X)$ | $\frac{35}{12} \approx 2.917$ |
| $E(2X^2+1)$ | $\frac{94}{3} \approx 31.33$ |
asked 3xavg 10 marks · 2078, 2076, 2075AnswerHideWrite the properties of correlation coefficient. The time it takes to transmit a file always depends on the file size. Suppose you transmitted 30 files, with the average size of 126 Kbytes and the standard deviation of 35 Kbytes. The average transmitted time was 0.04 seconds with the standard deviation 0.01 seconds. The correlation coefficient between the time and size was 0.86. Based on these data, fit a linear regression model and predict the time it take to transmit a 400Kbyte file.[10]
Write the properties of correlation coefficient. The time it takes to transmit a file always depends on the file size. Suppose you transmitted 30 files, with the average size of 126 Kbytes and the standard deviation of 35 Kbytes. The average transmitted time was 0.04 seconds with the standard deviation 0.01 seconds. The correlation coefficient between the time and size was 0.86. Based on these data, fit a linear regression model and predict the time it take to transmit a 400Kbyte file.[10]
Let $X$ = file size (Kbytes), $Y$ = transmission time (seconds). Parameter X (size) Y (time) ------------------------------- Number of files $n = 30$ Mean $\bar{X} = 126$ $\bar{Y} = 0.04$ Std. deviation $SX = 35$ $SY = 0.01$ Correlation ...
asked 3xavg 5 marks · 2078, 2076, 2075AnswerHideCalculate Spearman's rank correlation coefficient for the following ranks given by three judges in a music contest. Indicate which pair of judges has the nearest approaches to music.
$$\begin{array}{c|cccccccccc} \text{1}^{\text{st}}\text{ Judge} & 2 & 1 & 4 & 6 & 5 & 8 & 9 & 10 & 7 & 3 \ \text{2}^{\text{nd}}\text{ Judge} & 4 & 3 & 2 & 5 & 1 & 6 & 8 & 9 & 10 & 7 \ \text{3}^{\text{rd}}\text{ Judge} & 5 & 8 & 4 & 7 & 10 & 2 & 1 & 6 & 9 & 3 \ \end{array}$$
[5]
Calculate Spearman's rank correlation coefficient for the following ranks given by three judges in a music contest. Indicate which pair of judges has the nearest approaches to music.
$$\begin{array}{c|cccccccccc} \text{1}^{\text{st}}\text{ Judge} & 2 & 1 & 4 & 6 & 5 & 8 & 9 & 10 & 7 & 3 \ \text{2}^{\text{nd}}\text{ Judge} & 4 & 3 & 2 & 5 & 1 & 6 & 8 & 9 & 10 & 7 \ \text{3}^{\text{rd}}\text{ Judge} & 5 & 8 & 4 & 7 & 10 & 2 & 1 & 6 & 9 & 3 \ \end{array}$$
[5]
$n = 10$ competitors. Competitor 1st Judge ($R1$) 2nd Judge ($R2$) 3rd Judge ($R3$) :---::---::---::---: 1 2 4 5 2 1 3 8 3 4 2 4 4 6 5 7 5 5 1 10 6 8 6 2 7 9 8 1 8 10 9 6 9 7 10 9 10 3 7 3 Formula: $$rs = 1 - \frac{6\sum d^2}{n(n^2 - 1)}...
asked 3xavg 10 marks · 2080.1, 2080, 2079AnswerHideDifference between Correlation and Regression Analysis
Raw materials are used in the production of synthesis and fiber are stored in a place which has no humidity control. Measurement of relative humidity in the storage place and the moisture content of a sample of raw materials (both in percentage) on 10 days yields the following results.
$$\begin{array}{c|cccccccccc} \text{Humidity % (X)} & 48 & 55 & 30 & 44 & 36 & 31 & 62 & 48 & 42 & 50 \ \text{Moisture content % (Y)} & 11 & 13 & 10 & 12 & 9 & 7 & 16 & 11 & 9 & 14 \ \end{array}$$
i. Compute correlation coefficient between humidity and moisture content and interpret the result.
ii. Find the regression equation of moisture content on humidity.
iii. Estimate the moisture content if humidity is 45%.
iv. Interpret the value of regression coefficient.
[10]
Difference between Correlation and Regression Analysis
Raw materials are used in the production of synthesis and fiber are stored in a place which has no humidity control. Measurement of relative humidity in the storage place and the moisture content of a sample of raw materials (both in percentage) on 10 days yields the following results.
$$\begin{array}{c|cccccccccc} \text{Humidity % (X)} & 48 & 55 & 30 & 44 & 36 & 31 & 62 & 48 & 42 & 50 \ \text{Moisture content % (Y)} & 11 & 13 & 10 & 12 & 9 & 7 & 16 & 11 & 9 & 14 \ \end{array}$$
i. Compute correlation coefficient between humidity and moisture content and interpret the result.
ii. Find the regression equation of moisture content on humidity.
iii. Estimate the moisture content if humidity is 45%.
iv. Interpret the value of regression coefficient.
[10]
Difference Between Correlation and Regression Analysis
| Basis | Correlation | Regression |
|---|---|---|
| Meaning | Measures degree and direction of linear relationship between two variables | Establishes functional relationship to predict one variable from another |
| Purpose | Find strength of association | Estimate/predict dependent variable |
| Variables | Both treated equally | One independent (X), one dependent (Y) |
| Result | A single value $r \in [-1, +1]$ | An equation (regression line) |
| Symmetry | $r(X,Y)=r(Y,X)$ | Regression of Y on X $\neq$ X on Y |
| Cause-Effect | Does not imply causation | Assumes a dependence relationship |
Given Data
$n = 10$
| X | 48 | 55 | 30 | 44 | 36 | 31 | 62 | 48 | 42 | 50 |
|---|---|---|---|---|---|---|---|---|---|---|
| Y | 11 | 13 | 10 | 12 | 9 | 7 | 16 | 11 | 9 | 14 |
Preliminary Calculations
| X | Y | X² | Y² | XY |
|---|---|---|---|---|
| 48 | 11 | 2304 | 121 | 528 |
| 55 | 13 | 3025 | 169 | 715 |
| 30 | 10 | 900 | 100 | 300 |
| 44 | 12 | 1936 | 144 | 528 |
| 36 | 9 | 1296 | 81 | 324 |
| 31 | 7 | 961 | 49 | 217 |
| 62 | 16 | 3844 | 256 | 992 |
| 48 | 11 | 2304 | 121 | 528 |
| 42 | 9 | 1764 | 81 | 378 |
| 50 | 14 | 2500 | 196 | 700 |
| 446 | 112 | 20834 | 1318 | 5210 |
$$\bar{X}=\frac{446}{10}=44.6,\qquad \bar{Y}=\frac{112}{10}=11.2$$
$$\Sigma X^2-n\bar X^2 = 20834-10(44.6)^2=20834-19891.6=942.4$$ $$\Sigma Y^2-n\bar Y^2 = 1318-10(11.2)^2=1318-1254.4=63.6$$ $$\Sigma XY-n\bar X\bar Y = 5210-10(44.6)(11.2)=5210-4995.2=214.8$$
Part (i): Correlation Coefficient
$$r=\frac{\Sigma XY-n\bar X\bar Y}{\sqrt{(\Sigma X^2-n\bar X^2)(\Sigma Y^2-n\bar Y^2)}}=\frac{214.8}{\sqrt{942.4\times63.6}}$$
$$r=\frac{214.8}{\sqrt{59916.64}}=\frac{214.8}{244.78}\approx \boxed{0.877}$$
Interpretation: $r=0.877$ shows a strong positive linear relationship. As humidity increases, moisture content also increases.
Part (ii): Regression Equation of Y on X
$$b_{YX}=\frac{\Sigma XY-n\bar X\bar Y}{\Sigma X^2-n\bar X^2}=\frac{214.8}{942.4}=0.2279\approx 0.228$$
$$a=\bar Y-b\bar X=11.2-0.228(44.6)=11.2-10.169=1.031$$
$$\boxed{\hat Y = 1.03 + 0.228X}$$
Part (iii): Estimate Y when X = 45
$$\hat Y = 1.03 + 0.228(45)=1.03+10.26=\boxed{11.29%}$$
When humidity is 45%, estimated moisture content ≈ 11.29%.
Part (iv): Interpretation of Regression Coefficient
$b=0.228$ means that for every 1 percentage point increase in relative humidity, the moisture content increases on average by about 0.228 percentage points. The positive sign confirms a direct relationship: higher storage humidity leads to higher moisture in the raw materials.
asked 2xavg 5 marks · 2080, 2075AnswerHideDifferentiate between primary data and secondary data. What are the sources of secondary data? [5]
Differentiate between primary data and secondary data. What are the sources of secondary data? [5]
Primary Data: Primary data are those data which are collected by the investigator himself/herself for the first time. They are original in nature and collected for a specific purpose of inquiry. Example: Data collected by CBS (Central Bu...
asked 2xavg 5 marks · 2080, 2079AnswerHideWhat aspect of summary measures of data can be explained by the measures of skewness? Kelvin Hota is the national sales manager for National Text Books. He has a sales staff of 10 who visit college professors all over the United States. Each Sunday morning he requires his sales staffs to send him a report. Listed below are the number of visits last week. Compute the five number summary. 25, 6, 10, 13, 15, 2, 18, 5, 20, 30 [5]
What aspect of summary measures of data can be explained by the measures of skewness? Kelvin Hota is the national sales manager for National Text Books. He has a sales staff of 10 who visit college professors all over the United States. Each Sunday morning he requires his sales staffs to send him a report. Listed below are the number of visits last week. Compute the five number summary. 25, 6, 10, 13, 15, 2, 18, 5, 20, 30 [5]
Summary measures describe several aspects of data: central tendency (mean, median, mode), dispersion (range, variance), and shape. Measures of skewness explain the shape aspect, specifically the lack of symmetry of a distribution. Skewne...
asked 2xavg 5 marks · 2080, 2078AnswerHideThe first four moments of a distribution about x=2 are 1,2.5,5.5 and 16. Calculate the first four moments about the mean. Test the skewness and kurtosis. Interpret the results. [5]
The first four moments of a distribution about x=2 are 1,2.5,5.5 and 16. Calculate the first four moments about the mean. Test the skewness and kurtosis. Interpret the results. [5]
Moments about $x = 2$ (arbitrary origin $a = 2$): $$\mu1' = 1, \quad \mu2' = 2.5, \quad \mu3' = 5.5, \quad \mu4' = 16$$ --- First: $$\mu1 = 0$$ Second: $$\mu2 = \mu2' - (\mu1')^2 = 2.5 - 1 = 1.5$$ Third: $$\mu3 = \mu3' - 3\mu2'\mu1' + 2(...
asked 2xavg 5 marks · 2080, 2079AnswerHideDefine independent and mutually exclusive events. Three groups of children 2 boys and 2 girls, 3 boys and 1 girl, 1 boy and 3 girls respectively. One child is selected at random from each group. Find the probability of selecting one boy and two girls. [5]
Define independent and mutually exclusive events. Three groups of children 2 boys and 2 girls, 3 boys and 1 girl, 1 boy and 3 girls respectively. One child is selected at random from each group. Find the probability of selecting one boy and two girls. [5]
- Group 1: 2 boys, 2 girls (total 4) - Group 2: 3 boys, 1 girl (total 4) - Group 3: 1 boy, 3 girls (total 4) - One child selected at random from each group. - Required: $P(\text{1 boy and 2 girls})$ Mutually Exclusive Events: Two events ...
Study every one of these with model answers, flashcards, and MCQs.
Open STA169 study modes