Important Questions

STA169 · Exam intelligence

Statistics I important questions

From 6 past TU papers: which questions keep coming back, how much they carry, and what is most likely to show up next. Every question links to a model answer.

Most likely in the next examStatistical

Ranked by how often a topic is asked, its marks weight, and whether it is due after skipping the 2080.1 paper. No guarantees; study the whole syllabus.

1asked 11xavg 6 marks · Measures of central tendency
Answer

Define statistics and discuss its importance in the field of computational sciences. The following are the numbers of minutes that a person had to wait for the bus to work on 20 working days: 15, 10, 2, 17, 5, 8, 3, 10, 2, 9, 5, 9, 13, 1, 10, 12, 5, 10, 8, 4. Compute mean, median, mode, standard, variance and coefficient of variation.[10]

Statistics: Definition, Importance, and Computations

Part 1: Definition of Statistics (2 marks)

Statistics is the branch of science that deals with the collection, organization, presentation, analysis, and interpretation of numerical data in order to draw valid conclusions and make rational decisions under conditions of uncertainty.

Part 2: Importance of Statistics in Computational Sciences (3 marks)

  1. Fast computation: Computers perform statistical operations (mean, variance, ANOVA, regression) rapidly and reliably, removing tedious hand calculation.
  2. Handling large/multivariate data: Complex methods like multivariate analysis, clustering, and linear programming become feasible.
  3. Basis of modern techniques: Data mining, machine learning, pattern recognition, bootstrap methods, and image analysis are inherently statistical and computational.
  4. Research and reliability: Enables reproducible, high-speed experimental research and simulation.
  5. Practical CS/IT problem solving: Reliability testing (e.g., expected lifetime of hardware), performance benchmarking, and predictive modeling all depend on statistics.

Part 3: Statistical Computations (5 marks)

Given data (n = 20)

$$15, 10, 2, 17, 5, 8, 3, 10, 2, 9, 5, 9, 13, 1, 10, 12, 5, 10, 8, 4$$

Sorted: $$1, 2, 2, 3, 4, 5, 5, 5, 8, 8, 9, 9, 10, 10, 10, 10, 12, 13, 15, 17$$

Mean

$$\sum X = 158$$ $$\bar{X} = \frac{158}{20} = 7.9 \text{ minutes}$$

Median (n even → average of 10th and 11th values)

10th = 8, 11th = 9 $$\text{Median} = \frac{8+9}{2} = 8.5 \text{ minutes}$$

Mode

Value 10 occurs 4 times (highest frequency). $$\text{Mode} = 10 \text{ minutes}$$

Variance and Standard Deviation

$$\sum X^2 = 1626$$

$$\sigma^2 = \frac{\sum X^2}{n} - \bar{X}^2 = \frac{1626}{20} - (7.9)^2$$ $$= 81.3 - 62.41 = 18.89$$

$$\sigma = \sqrt{18.89} = 4.3463 \approx 4.35 \text{ minutes}$$

(Note: if the sample formula with $n-1$ is used: $$s^2 = \frac{\sum X^2 - n\bar{X}^2}{n-1} = \frac{1626 - 1248.2}{19} = \frac{377.8}{19} = 19.8842$$ $$s = 4.4592 \text{ minutes})$$

Coefficient of Variation

$$CV = \frac{\sigma}{\bar{X}} \times 100 = \frac{4.3463}{7.9} \times 100 = 55.02%$$

(Using sample SD: $CV = \frac{4.4592}{7.9}\times100 = 56.45%$)

Final Summary

MeasureValue
Mean7.9 min
Median8.5 min
Mode10 min
Variance (population)18.89
Standard Deviation (population)4.35 min
Coefficient of Variation55.02%
2asked 11xavg 6 marks · Discrete distributions
Answer

Fitting a Binomial Distribution

Under what condition binomial probability distribution? Five unbiased coins are tossed 100 times and the following results were obtained. Fit the binomial distribution.

No of heads012345
Frequency5243522104

[5]

The binomial probability distribution applies under the following conditions: - Each trial results in only two mutually exclusive outcomes (success/failure). - The number of trials n is fixed and finite. - The trials are independent of o...

3asked 7xavg 6 marks · Joint probability distribution of two random variables
Answer

Let X and Y be two continuous random variable having joint pdf $f(x, y) = c(x^2 + y^2)$, $0 < x < 1$, $0 < y < 1$, $= 0$ otherwise. Determine

(a) the value of c.

(b) $P(x < 0.5, y > 0.5)$. [5]

$$f(x, y) = c(x^2 + y^2), \quad 0 < x < 1,\ 0 < y < 1$$ $$f(x, y) = 0 \quad \text{otherwise}$$ Required: (a) value of $c$; (b) $P(X < 0.5,\ Y 0.5)$. --- Normalization condition: $$\int{0}^{1}\int{0}^{1} c(x^2 + y^2), dx, dy = 1$$ Inner...

4asked 6xavg 8 marks · Measures of dispersion
Answer

Measurement of Dispersion and Consistency Analysis

Measurement of dispersion refers to statistical measures describing the spread or variability of data values around a central value. It indicates how much individual observations deviate from the average. Common measures: Range, Mean Dev...

5asked 6xavg 7 marks · Continuous distribution
Answer

Define normal distribution. What are the main characteristics of normal distribution? Extruded plastic rods are automatically cut into length 5 inches. Actual length are normally distributed about a mean of 5 inches and their standard deviation is 0.05 inches.

(i) What proportion of rods exceed tolerance limits of 4.9 inches to 5.1 inches?

(ii) Proportion of rods having tolerance rod which is greater than 6.5 inches.[10]

Normal Distribution: Definition, Characteristics, and Numerical

Definition

A normal distribution is a continuous probability distribution, symmetrical and bell-shaped about its mean, defined by two parameters, mean $\mu$ and standard deviation $\sigma$. Its probability density function is:

$$f(x) = \frac{1}{\sigma\sqrt{2\pi}} e^{-\frac{1}{2}\left(\frac{x-\mu}{\sigma}\right)^2}, \quad -\infty < x < \infty$$

Main Characteristics

  1. It is a continuous distribution with parameters $\mu$ (mean) and $\sigma^2$ (variance).
  2. The curve is bell-shaped and symmetrical about the mean $\mu$.
  3. Mean = Median = Mode (all coincide at the centre).
  4. Maximum ordinate occurs at $x = \mu$, equal to $\frac{1}{\sigma\sqrt{2\pi}}$.
  5. Skewness $\beta_1 = 0$ (perfectly symmetric).
  6. Kurtosis $\beta_2 = 3$ (mesokurtic).
  7. Empirical area rule:
    • $P(\mu-\sigma < X < \mu+\sigma) = 0.6826$
    • $P(\mu-2\sigma < X < \mu+2\sigma) = 0.9544$
    • $P(\mu-3\sigma < X < \mu+3\sigma) = 0.9974$
  8. Total area under the curve = 1.
  9. Asymptotic: the curve never touches the X-axis.

Given Data

  • Mean $\mu = 5$ inches
  • Standard deviation $\sigma = 0.05$ inches
  • $X \sim N(5, 0.05^2)$
  • Standard normal variate: $Z = \dfrac{X - 5}{0.05}$

Part (i): Proportion exceeding tolerance limits 4.9 to 5.1 inches

Z-scores:

$$Z_1 = \frac{4.9 - 5}{0.05} = -2, \qquad Z_2 = \frac{5.1 - 5}{0.05} = +2$$

Proportion within limits:

$$P(4.9 < X < 5.1) = P(-2 < Z < 2) = 2 \times P(0 < Z < 2) = 2 \times 0.4772 = 0.9544$$

Proportion exceeding limits:

$$P(X < 4.9 \text{ or } X > 5.1) = 1 - 0.9544 = \boxed{0.0456 \ (4.56%)}$$

Part (ii): Proportion having length greater than 6.5 inches

Z-score:

$$Z = \frac{6.5 - 5}{0.05} = \frac{1.5}{0.05} = 30$$

Since $Z = 30$ lies far beyond the standard normal table (which effectively ends near $Z \approx 4$):

$$P(X > 6.5) = P(Z > 30) \approx \boxed{0}$$

A rod exceeding 6.5 inches is 30 standard deviations above the mean, so the proportion is practically zero.

Summary

PartConditionZ-value(s)Proportion
(i)X < 4.9 or X > 5.1$\pm 2$0.0456 (4.56%)
(ii)X > 6.530$\approx 0$

Most repeated questions

Topics asked at least twice, most-asked first.

asked 11xavg 6 marks · 2080.1, 2080, 2079, 2078, 2076...
Answer

Define statistics and discuss its importance in the field of computational sciences. The following are the numbers of minutes that a person had to wait for the bus to work on 20 working days: 15, 10, 2, 17, 5, 8, 3, 10, 2, 9, 5, 9, 13, 1, 10, 12, 5, 10, 8, 4. Compute mean, median, mode, standard, variance and coefficient of variation.[10]

Statistics: Definition, Importance, and Computations

Part 1: Definition of Statistics (2 marks)

Statistics is the branch of science that deals with the collection, organization, presentation, analysis, and interpretation of numerical data in order to draw valid conclusions and make rational decisions under conditions of uncertainty.

Part 2: Importance of Statistics in Computational Sciences (3 marks)

  1. Fast computation: Computers perform statistical operations (mean, variance, ANOVA, regression) rapidly and reliably, removing tedious hand calculation.
  2. Handling large/multivariate data: Complex methods like multivariate analysis, clustering, and linear programming become feasible.
  3. Basis of modern techniques: Data mining, machine learning, pattern recognition, bootstrap methods, and image analysis are inherently statistical and computational.
  4. Research and reliability: Enables reproducible, high-speed experimental research and simulation.
  5. Practical CS/IT problem solving: Reliability testing (e.g., expected lifetime of hardware), performance benchmarking, and predictive modeling all depend on statistics.

Part 3: Statistical Computations (5 marks)

Given data (n = 20)

$$15, 10, 2, 17, 5, 8, 3, 10, 2, 9, 5, 9, 13, 1, 10, 12, 5, 10, 8, 4$$

Sorted: $$1, 2, 2, 3, 4, 5, 5, 5, 8, 8, 9, 9, 10, 10, 10, 10, 12, 13, 15, 17$$

Mean

$$\sum X = 158$$ $$\bar{X} = \frac{158}{20} = 7.9 \text{ minutes}$$

Median (n even → average of 10th and 11th values)

10th = 8, 11th = 9 $$\text{Median} = \frac{8+9}{2} = 8.5 \text{ minutes}$$

Mode

Value 10 occurs 4 times (highest frequency). $$\text{Mode} = 10 \text{ minutes}$$

Variance and Standard Deviation

$$\sum X^2 = 1626$$

$$\sigma^2 = \frac{\sum X^2}{n} - \bar{X}^2 = \frac{1626}{20} - (7.9)^2$$ $$= 81.3 - 62.41 = 18.89$$

$$\sigma = \sqrt{18.89} = 4.3463 \approx 4.35 \text{ minutes}$$

(Note: if the sample formula with $n-1$ is used: $$s^2 = \frac{\sum X^2 - n\bar{X}^2}{n-1} = \frac{1626 - 1248.2}{19} = \frac{377.8}{19} = 19.8842$$ $$s = 4.4592 \text{ minutes})$$

Coefficient of Variation

$$CV = \frac{\sigma}{\bar{X}} \times 100 = \frac{4.3463}{7.9} \times 100 = 55.02%$$

(Using sample SD: $CV = \frac{4.4592}{7.9}\times100 = 56.45%$)

Final Summary

MeasureValue
Mean7.9 min
Median8.5 min
Mode10 min
Variance (population)18.89
Standard Deviation (population)4.35 min
Coefficient of Variation55.02%
asked 11xavg 6 marks · 2080.1, 2080, 2079, 2078, 2076...
Answer

Fitting a Binomial Distribution

Under what condition binomial probability distribution? Five unbiased coins are tossed 100 times and the following results were obtained. Fit the binomial distribution.

No of heads012345
Frequency5243522104

[5]

The binomial probability distribution applies under the following conditions: - Each trial results in only two mutually exclusive outcomes (success/failure). - The number of trials n is fixed and finite. - The trials are independent of o...

asked 7xavg 6 marks · 2080.1, 2080, 2078, 2076, 2075
Answer

Let X and Y be two continuous random variable having joint pdf $f(x, y) = c(x^2 + y^2)$, $0 < x < 1$, $0 < y < 1$, $= 0$ otherwise. Determine

(a) the value of c.

(b) $P(x < 0.5, y > 0.5)$. [5]

$$f(x, y) = c(x^2 + y^2), \quad 0 < x < 1,\ 0 < y < 1$$ $$f(x, y) = 0 \quad \text{otherwise}$$ Required: (a) value of $c$; (b) $P(X < 0.5,\ Y 0.5)$. --- Normalization condition: $$\int{0}^{1}\int{0}^{1} c(x^2 + y^2), dx, dy = 1$$ Inner...

asked 6xavg 8 marks · 2080.1, 2079, 2078, 2076, 2075
Answer

Measurement of Dispersion and Consistency Analysis

Measurement of dispersion refers to statistical measures describing the spread or variability of data values around a central value. It indicates how much individual observations deviate from the average. Common measures: Range, Mean Dev...

asked 6xavg 7 marks · 2080.1, 2080, 2079, 2078, 2076...
Answer

Define normal distribution. What are the main characteristics of normal distribution? Extruded plastic rods are automatically cut into length 5 inches. Actual length are normally distributed about a mean of 5 inches and their standard deviation is 0.05 inches.

(i) What proportion of rods exceed tolerance limits of 4.9 inches to 5.1 inches?

(ii) Proportion of rods having tolerance rod which is greater than 6.5 inches.[10]

Normal Distribution: Definition, Characteristics, and Numerical

Definition

A normal distribution is a continuous probability distribution, symmetrical and bell-shaped about its mean, defined by two parameters, mean $\mu$ and standard deviation $\sigma$. Its probability density function is:

$$f(x) = \frac{1}{\sigma\sqrt{2\pi}} e^{-\frac{1}{2}\left(\frac{x-\mu}{\sigma}\right)^2}, \quad -\infty < x < \infty$$

Main Characteristics

  1. It is a continuous distribution with parameters $\mu$ (mean) and $\sigma^2$ (variance).
  2. The curve is bell-shaped and symmetrical about the mean $\mu$.
  3. Mean = Median = Mode (all coincide at the centre).
  4. Maximum ordinate occurs at $x = \mu$, equal to $\frac{1}{\sigma\sqrt{2\pi}}$.
  5. Skewness $\beta_1 = 0$ (perfectly symmetric).
  6. Kurtosis $\beta_2 = 3$ (mesokurtic).
  7. Empirical area rule:
    • $P(\mu-\sigma < X < \mu+\sigma) = 0.6826$
    • $P(\mu-2\sigma < X < \mu+2\sigma) = 0.9544$
    • $P(\mu-3\sigma < X < \mu+3\sigma) = 0.9974$
  8. Total area under the curve = 1.
  9. Asymptotic: the curve never touches the X-axis.

Given Data

  • Mean $\mu = 5$ inches
  • Standard deviation $\sigma = 0.05$ inches
  • $X \sim N(5, 0.05^2)$
  • Standard normal variate: $Z = \dfrac{X - 5}{0.05}$

Part (i): Proportion exceeding tolerance limits 4.9 to 5.1 inches

Z-scores:

$$Z_1 = \frac{4.9 - 5}{0.05} = -2, \qquad Z_2 = \frac{5.1 - 5}{0.05} = +2$$

Proportion within limits:

$$P(4.9 < X < 5.1) = P(-2 < Z < 2) = 2 \times P(0 < Z < 2) = 2 \times 0.4772 = 0.9544$$

Proportion exceeding limits:

$$P(X < 4.9 \text{ or } X > 5.1) = 1 - 0.9544 = \boxed{0.0456 \ (4.56%)}$$

Part (ii): Proportion having length greater than 6.5 inches

Z-score:

$$Z = \frac{6.5 - 5}{0.05} = \frac{1.5}{0.05} = 30$$

Since $Z = 30$ lies far beyond the standard normal table (which effectively ends near $Z \approx 4$):

$$P(X > 6.5) = P(Z > 30) \approx \boxed{0}$$

A rod exceeding 6.5 inches is 30 standard deviations above the mean, so the proportion is practically zero.

Summary

PartConditionZ-value(s)Proportion
(i)X < 4.9 or X > 5.1$\pm 2$0.0456 (4.56%)
(ii)X > 6.530$\approx 0$
asked 6xavg 5 marks · 2080.1, 2080, 2079, 2078, 2076...
Answer

What is sampling? Explain the main purpose of sampling. Describe briefly stratified sampling. [5]

Sampling: Definition, Purpose, and Stratified Sampling


1. What is Sampling?

Sampling is the process of selecting a subset (called a sample) of units from a larger group (called the population) in order to draw conclusions or make inferences about the entire population. Instead of studying every unit in the population (census), sampling allows us to study only a representative portion, saving time, cost, and effort.

The difference between the value of the sample statistic obtained from a sample and the value of the corresponding population parameter is called the sampling error.


2. Main Purpose of Sampling

The main purposes of sampling are:

  • Economy: Studying the entire population is costly; sampling reduces expenditure of money, time, and manpower.
  • Feasibility: In many cases, examining every unit is practically impossible (e.g., testing every light bulb for lifespan).
  • Speed: Results can be obtained more quickly from a sample than from a complete census.
  • Accuracy: A carefully conducted sample survey can sometimes be more accurate than a full census, since more attention can be given to each selected unit.
  • Defining survey objectives: The first step in sampling is to clearly define the objective of the survey, commensurate with available resources in terms of money, manpower, and time.

3. Stratified Random Sampling

Definition

When the units in the population are not similar in nature, the population is first divided into sub-groups called strata, and then a simple random sample is drawn from each stratum in proportion to its size. This technique is known as Stratified Random Sampling.

Conditions for Stratification

The strata must satisfy the following conditions:

  • The strata should be non-overlapping and together should comprise the whole population.
  • The strata should be as homogeneous within groups (similar units inside each stratum) and heterogeneous between groups (different units across strata) as possible.

Purpose of Stratification

According to the notes, the purposes of stratification are:

  1. To make the sample more representative of the population.
  2. For greater accuracy in the results.
  3. For administrative convenience in conducting the survey.

Example

Suppose we want to study the income level of people in a city. The population can be stratified into groups such as:

  • Low income group
  • Middle income group
  • High income group

A random sample is then drawn from each group separately, ensuring all income levels are properly represented in the final sample.

Advantage over Simple Random Sampling

For a given level of precision, stratified random sampling usually requires a smaller sample size as compared to simple random sampling, making it more efficient.


Summary Table:

FeatureStratified Sampling
Population divided intoStrata (sub-groups)
Within strataHomogeneous
Between strataHeterogeneous
Sampling from each stratumSimple Random Sampling
Main benefitMore representative and accurate
asked 4xavg 5 marks · 2080.1, 2080, 2079, 2076
Answer

Define Baye's theorem. Store A, B and C have 100, 75 and 50 employees and, respectively 70, 60 and 50 percent of these are women. Registration are equally likely among all employees regardless of sex. One employee resigns, and this is woman. What is the probability that she works in store B? [5]

Bayes' Theorem - Definition and Application

STEP 1: Given Data

StoreEmployees% WomenWomen Count
A10070%70
B7560%45
C5050%25
Total225140

Resignations equally likely among all employees. Observed event: the resigning employee is a woman. Find $P(\text{Store B} \mid \text{Woman})$.

Definition of Bayes' Theorem

If $B_1, B_2, \dots, B_n$ are mutually exclusive and exhaustive events with $P(B_i) > 0$, and $A$ is any event with $P(A) > 0$, then:

$$P(B_i \mid A) = \frac{P(B_i), P(A \mid B_i)}{\sum_{j=1}^{n} P(B_j), P(A \mid B_j)}$$

  • $P(B_i)$ = prior probability
  • $P(A \mid B_i)$ = likelihood
  • $P(B_i \mid A)$ = posterior probability

STEP 2: Solution

Define Events

  • $S_A, S_B, S_C$ = employee works in Store A, B, C
  • $W$ = resigning employee is a woman

Prior Probabilities (equal likelihood over 225 employees)

$$P(S_A) = \frac{100}{225}, \quad P(S_B) = \frac{75}{225}, \quad P(S_C) = \frac{50}{225}$$

Likelihoods

$$P(W \mid S_A) = 0.70, \quad P(W \mid S_B) = 0.60, \quad P(W \mid S_C) = 0.50$$

Total Probability of a Woman Resigning

$$P(W) = \frac{100}{225}(0.70) + \frac{75}{225}(0.60) + \frac{50}{225}(0.50)$$

$$P(W) = \frac{70 + 45 + 25}{225} = \frac{140}{225}$$

Apply Bayes' Theorem

$$P(S_B \mid W) = \frac{P(S_B), P(W \mid S_B)}{P(W)} = \frac{\dfrac{75}{225}(0.60)}{\dfrac{140}{225}} = \frac{45}{140}$$

$$\boxed{P(S_B \mid W) = \frac{45}{140} = \frac{9}{28} \approx 0.3214 \approx 32.14%}$$

Conclusion

The probability that the resigning woman works in Store B is approximately 0.3214 (32.14%).

asked 3xavg 5 marks · 2080, 2079, 2076
Answer

Define random variable. If X is the number of points rolled with a balanced die, find E(X) and variance of X. Also find the expected value of random variable g(X) = 2X^2 + 1. [5]

Random Variable: Definition, E(X), Variance, and E(g(X))

Step 1 - Extract (Given Data)

  • Experiment: rolling a balanced (fair) die
  • Random variable $X$ = number of points shown
  • Possible values: $x = 1, 2, 3, 4, 5, 6$
  • $P(X = x) = \frac{1}{6}$ for each value (balanced die)
  • Function required: $g(X) = 2X^2 + 1$

All data present and sufficient.

Step 2 - Solve

Definition of Random Variable

A random variable is a real-valued function that assigns a numerical value to each outcome in the sample space of a random experiment. It may be discrete (countable values) or continuous (values over an interval).

Probability Distribution

$$P(X = x) = \frac{1}{6}, \quad x = 1, 2, 3, 4, 5, 6$$

Calculation of E(X)

$$E(X) = \sum x \cdot P(X=x) = \frac{1+2+3+4+5+6}{6} = \frac{21}{6} = \frac{7}{2} = 3.5$$

Calculation of Variance

$$E(X^2) = \frac{1^2+2^2+3^2+4^2+5^2+6^2}{6} = \frac{1+4+9+16+25+36}{6} = \frac{91}{6}$$

$$\text{Var}(X) = E(X^2) - [E(X)]^2 = \frac{91}{6} - \left(\frac{7}{2}\right)^2 = \frac{91}{6} - \frac{49}{4}$$

$$= \frac{182}{12} - \frac{147}{12} = \frac{35}{12} \approx 2.917$$

Expected Value of g(X) = 2X² + 1

$$E(2X^2 + 1) = 2E(X^2) + 1 = 2 \cdot \frac{91}{6} + 1 = \frac{182}{6} + \frac{6}{6} = \frac{188}{6} = \frac{94}{3} \approx 31.33$$

Summary

QuantityValue
$E(X)$$\frac{7}{2} = 3.5$
$\text{Var}(X)$$\frac{35}{12} \approx 2.917$
$E(2X^2+1)$$\frac{94}{3} \approx 31.33$
asked 3xavg 10 marks · 2078, 2076, 2075
Answer

Write the properties of correlation coefficient. The time it takes to transmit a file always depends on the file size. Suppose you transmitted 30 files, with the average size of 126 Kbytes and the standard deviation of 35 Kbytes. The average transmitted time was 0.04 seconds with the standard deviation 0.01 seconds. The correlation coefficient between the time and size was 0.86. Based on these data, fit a linear regression model and predict the time it take to transmit a 400Kbyte file.[10]

Let $X$ = file size (Kbytes), $Y$ = transmission time (seconds). Parameter X (size) Y (time) ------------------------------- Number of files $n = 30$ Mean $\bar{X} = 126$ $\bar{Y} = 0.04$ Std. deviation $SX = 35$ $SY = 0.01$ Correlation ...

asked 3xavg 5 marks · 2078, 2076, 2075
Answer

Calculate Spearman's rank correlation coefficient for the following ranks given by three judges in a music contest. Indicate which pair of judges has the nearest approaches to music.

$$\begin{array}{c|cccccccccc} \text{1}^{\text{st}}\text{ Judge} & 2 & 1 & 4 & 6 & 5 & 8 & 9 & 10 & 7 & 3 \ \text{2}^{\text{nd}}\text{ Judge} & 4 & 3 & 2 & 5 & 1 & 6 & 8 & 9 & 10 & 7 \ \text{3}^{\text{rd}}\text{ Judge} & 5 & 8 & 4 & 7 & 10 & 2 & 1 & 6 & 9 & 3 \ \end{array}$$

[5]

$n = 10$ competitors. Competitor 1st Judge ($R1$) 2nd Judge ($R2$) 3rd Judge ($R3$) :---::---::---::---: 1 2 4 5 2 1 3 8 3 4 2 4 4 6 5 7 5 5 1 10 6 8 6 2 7 9 8 1 8 10 9 6 9 7 10 9 10 3 7 3 Formula: $$rs = 1 - \frac{6\sum d^2}{n(n^2 - 1)}...

asked 3xavg 10 marks · 2080.1, 2080, 2079
Answer

Difference between Correlation and Regression Analysis

Raw materials are used in the production of synthesis and fiber are stored in a place which has no humidity control. Measurement of relative humidity in the storage place and the moisture content of a sample of raw materials (both in percentage) on 10 days yields the following results.

$$\begin{array}{c|cccccccccc} \text{Humidity % (X)} & 48 & 55 & 30 & 44 & 36 & 31 & 62 & 48 & 42 & 50 \ \text{Moisture content % (Y)} & 11 & 13 & 10 & 12 & 9 & 7 & 16 & 11 & 9 & 14 \ \end{array}$$

i. Compute correlation coefficient between humidity and moisture content and interpret the result.

ii. Find the regression equation of moisture content on humidity.

iii. Estimate the moisture content if humidity is 45%.

iv. Interpret the value of regression coefficient.

[10]

Difference Between Correlation and Regression Analysis

BasisCorrelationRegression
MeaningMeasures degree and direction of linear relationship between two variablesEstablishes functional relationship to predict one variable from another
PurposeFind strength of associationEstimate/predict dependent variable
VariablesBoth treated equallyOne independent (X), one dependent (Y)
ResultA single value $r \in [-1, +1]$An equation (regression line)
Symmetry$r(X,Y)=r(Y,X)$Regression of Y on X $\neq$ X on Y
Cause-EffectDoes not imply causationAssumes a dependence relationship

Given Data

$n = 10$

X48553044363162484250
Y11131012971611914

Preliminary Calculations

XYX²Y²XY
48112304121528
55133025169715
3010900100300
44121936144528
369129681324
31796149217
62163844256992
48112304121528
429176481378
50142500196700
4461122083413185210

$$\bar{X}=\frac{446}{10}=44.6,\qquad \bar{Y}=\frac{112}{10}=11.2$$

$$\Sigma X^2-n\bar X^2 = 20834-10(44.6)^2=20834-19891.6=942.4$$ $$\Sigma Y^2-n\bar Y^2 = 1318-10(11.2)^2=1318-1254.4=63.6$$ $$\Sigma XY-n\bar X\bar Y = 5210-10(44.6)(11.2)=5210-4995.2=214.8$$


Part (i): Correlation Coefficient

$$r=\frac{\Sigma XY-n\bar X\bar Y}{\sqrt{(\Sigma X^2-n\bar X^2)(\Sigma Y^2-n\bar Y^2)}}=\frac{214.8}{\sqrt{942.4\times63.6}}$$

$$r=\frac{214.8}{\sqrt{59916.64}}=\frac{214.8}{244.78}\approx \boxed{0.877}$$

Interpretation: $r=0.877$ shows a strong positive linear relationship. As humidity increases, moisture content also increases.


Part (ii): Regression Equation of Y on X

$$b_{YX}=\frac{\Sigma XY-n\bar X\bar Y}{\Sigma X^2-n\bar X^2}=\frac{214.8}{942.4}=0.2279\approx 0.228$$

$$a=\bar Y-b\bar X=11.2-0.228(44.6)=11.2-10.169=1.031$$

$$\boxed{\hat Y = 1.03 + 0.228X}$$


Part (iii): Estimate Y when X = 45

$$\hat Y = 1.03 + 0.228(45)=1.03+10.26=\boxed{11.29%}$$

When humidity is 45%, estimated moisture content ≈ 11.29%.


Part (iv): Interpretation of Regression Coefficient

$b=0.228$ means that for every 1 percentage point increase in relative humidity, the moisture content increases on average by about 0.228 percentage points. The positive sign confirms a direct relationship: higher storage humidity leads to higher moisture in the raw materials.

asked 2xavg 5 marks · 2080, 2075
Answer

Differentiate between primary data and secondary data. What are the sources of secondary data? [5]

Primary Data: Primary data are those data which are collected by the investigator himself/herself for the first time. They are original in nature and collected for a specific purpose of inquiry. Example: Data collected by CBS (Central Bu...

asked 2xavg 5 marks · 2080, 2079
Answer

What aspect of summary measures of data can be explained by the measures of skewness? Kelvin Hota is the national sales manager for National Text Books. He has a sales staff of 10 who visit college professors all over the United States. Each Sunday morning he requires his sales staffs to send him a report. Listed below are the number of visits last week. Compute the five number summary. 25, 6, 10, 13, 15, 2, 18, 5, 20, 30 [5]

Summary measures describe several aspects of data: central tendency (mean, median, mode), dispersion (range, variance), and shape. Measures of skewness explain the shape aspect, specifically the lack of symmetry of a distribution. Skewne...

asked 2xavg 5 marks · 2080, 2078
Answer

The first four moments of a distribution about x=2 are 1,2.5,5.5 and 16. Calculate the first four moments about the mean. Test the skewness and kurtosis. Interpret the results. [5]

Moments about $x = 2$ (arbitrary origin $a = 2$): $$\mu1' = 1, \quad \mu2' = 2.5, \quad \mu3' = 5.5, \quad \mu4' = 16$$ --- First: $$\mu1 = 0$$ Second: $$\mu2 = \mu2' - (\mu1')^2 = 2.5 - 1 = 1.5$$ Third: $$\mu3 = \mu3' - 3\mu2'\mu1' + 2(...

asked 2xavg 5 marks · 2080, 2079
Answer

Define independent and mutually exclusive events. Three groups of children 2 boys and 2 girls, 3 boys and 1 girl, 1 boy and 3 girls respectively. One child is selected at random from each group. Find the probability of selecting one boy and two girls. [5]

  • Group 1: 2 boys, 2 girls (total 4) - Group 2: 3 boys, 1 girl (total 4) - Group 3: 1 boy, 3 girls (total 4) - One child selected at random from each group. - Required: $P(\text{1 boy and 2 girls})$ Mutually Exclusive Events: Two events ...

Study every one of these with model answers, flashcards, and MCQs.

Open STA169 study modes