STA169 · TU past paper
Statistics I 2075 question paper
The complete TU 2075 exam paper for Statistics I (STA169), all 13 questions with solved model answers written to the mark scheme.
Tap a question to open its answer.
- 110 marksNumericalMeasures of dispersionHideAnswer
Distinction Between Absolute and Relative Measures of Dispersion
Absolute Measure of Dispersion: An absolute measure expresses the dispersion in the same units as the original data. Examples include range, variance, standard deviation, and mean deviation. These measures are useful when comparing datasets with similar means and units.
Relative Measure of Dispersion: A relative measure is a dimensionless quantity expressed as a ratio or percentage of a central tendency measure (usually the mean). Examples include coefficient of variation, coefficient of range, and coefficient of quartile deviation. These measures are useful for comparing the consistency of datasets with different means or different units.
Computer Consistency Analysis
Time (in seconds) 0-2 2-4 4-6 6-8 8-10 10-12 A 5 16 13 7 5 4 B 2 7 12 19 9 1 For Company A:
Mean: $\bar{x}_A = \frac{1 \times 5 + 3 \times 16 + 5 \times 13 + 7 \times 7 + 9 \times 5 + 11 \times 4}{50} = \frac{5 + 48 + 65 + 49 + 45 + 44}{50} = \frac{256}{50} = 5.12$ seconds
Variance: $\sigma_A^2 = \frac{\sum f(x - \bar{x})^2}{n} = \frac{5(1-5.12)^2 + 16(3-5.12)^2 + 13(5-5.12)^2 + 7(7-5.12)^2 + 5(9-5.12)^2 + 4(11-5.12)^2}{50} = 6.8896$
Standard Deviation: $\sigma_A = \sqrt{6.8896} = 2.624$ seconds
Coefficient of Variation: $CV_A = \frac{\sigma_A}{\bar{x}_A} \times 100 = \frac{2.624}{5.12} \times 100 = 51.25%$
For Company B:
Mean: $\bar{x}_B = \frac{1 \times 2 + 3 \times 7 + 5 \times 12 + 7 \times 19 + 9 \times 9 + 11 \times 1}{50} = \frac{2 + 21 + 60 + 133 + 81 + 11}{50} = \frac{308}{50} = 6.16$ seconds
Variance: $\sigma_B^2 = \frac{2(1-6.16)^2 + 7(3-6.16)^2 + 12(5-6.16)^2 + 19(7-6.16)^2 + 9(9-6.16)^2 + 1(11-6.16)^2}{50} = 5.8704$
Standard Deviation: $\sigma_B = \sqrt{5.8704} = 2.423$ seconds
Coefficient of Variation: $CV_B = \frac{\sigma_B}{\bar{x}_B} \times 100 = \frac{2.423}{6.16} \times 100 = 39.37%$
Conclusion: Since $CV_B (39.37%) < CV_A (51.25%)$, Company B's computers are more consistent as they have a lower coefficient of variation, indicating less relative variability in execution time.
Basis Absolute Measure Relative Measure ----------------------------------------- Definition Expresses dispersion in the same units as the original data Expresses dispersion as a pure number (ratio/percentage), free of units Unit Has uni...
- 210 marksNumericalRegression AnalysisHideAnswer
Regression Analysis: Normal Stress and Shear Resistance
There is a data inconsistency in the problem statement. The header row lists 9 pairs, but the intro table string lists 10 values for each variable. I will use the 9 clearly-tabulated pairs from the bracketed array: X (Normal stress) 26 2...
- 310 marksDiscrete distributionsHideAnswer
What do you understand by binomial distribution? What are its main features? What do you mean by marginal probability distribution? Write down its properties.[10]
--- A binomial distribution is a discrete probability distribution that describes the number of successes in n independent trials, where each trial has only two possible outcomes: success (with probability p) and failure (with probabilit...
- 45 marksNumericalMeasures of central tendencyHideAnswer
Measurement of computer chip's thickness (in nanometers) is recorded below. Find the mode of thickness of computer chips and interpret the result.
Thickness of chips in n.m. 34-39 39-44 44-49 49-54 54-59 Total No. of computers 3 11 16 25 5 60 [5]
Thickness (n.m.) No. of computers (f) -------------------------------------- 34 - 39 3 39 - 44 11 44 - 49 16 49 - 54 25 54 - 59 5 Total 60 Class width $h = 5$, continuous class intervals. --- The modal class is the class with the highest...
- 55 marksNumericalMeasures of central tendencyHideAnswer
Calculate Q3, D6 and P80 from the following data and interpret the results.
Respiratory rate 10 15 20 25 30 35 40 45 50 No. of Person 8 12 36 25 28 18 9 12 6 [5]
Discrete (ungrouped) frequency distribution: Respiratory rate (x) 10 15 20 25 30 35 40 45 50 ------------------------------ No. of persons (f) 8 12 36 25 28 18 9 12 6 x f cf --------- 10 8 8 15 12 20 20 36 56 25 25 81 30 28 109 35 18 127...
- 65 marksNumericalJoint probability distribution of two randHideAnswer
Random Variable and Bivariate Probability Distribution
Definition of a Random Variable: A random variable is a function that assigns numerical values to the outcomes of a random experiment. It can be discrete (taking countable values) or continuous (taking any value in an interval).
A random variable is a real-valued function that assigns a numerical value to each outcome in the sample space of a random experiment. A random variable is discrete if it takes finitely many or countably many values, each with a probabil...
- 75 marksNumericalJoint probability distribution of two randHideAnswer
If two random variables have the joint probability density function
$$f(x,y) = \begin{cases} k e^{-(x+y)}, & 0 < x < \infty, 0 < y < \infty \ 0, & \text{otherwise} \end{cases}$$
Find (i) constant $k$ (ii) conditional probability density function of $X$ given $Y$ (iii) Var$(3X + 2Y)$. [5]
Joint PDF: $$f(x,y) = \begin{cases} k,e^{-(x+y)}, & 0<x<\infty,\ 0<y<\infty \ 0, & \text{otherwise} \end{cases}$$ Find: (i) constant $k$, (ii) conditional PDF of $X$ given $Y$, (iii) $\text{Var}(3X+2Y)$. --- $$\int0^\infty\int0^\infty ...
- 85 marksNumericalContinuous distributionHideAnswer
A certain machine makes electrical resistors having a mean resistance of 40 ohms and standard deviation of 2 ohms. Assuming that the resistance follows a normal distribution.(i) What percentage of resistors will have a resistance exceeding 43 ohms?(ii) What percentage of resistors will have a resistance between 30 ohms to 45 ohms? [5]
- Mean resistance: $\mu = 40$ ohms - Standard deviation: $\sigma = 2$ ohms - Distribution: Normal Standardization formula: $$Z = \frac{X - \mu}{\sigma}$$ --- Z-score for X = 43: $$Z = \frac{43 - 40}{2} = \frac{3}{2} = 1.5$$ Probability: ...
- 95 marksNumericalSpearman's rank correlationHideAnswer
Spearman's Rank Correlation Coefficient
As part of study of the psychological correlates of success in athletes, the following measurements are obtained from members of Nepal national football team. Calculate Spearman's rank correlation coefficient.
$$ \begin{array}{c|cccccccc} \text{Anger} & 6 & 7 & 5 & 21 & 13 & 5 & 13 & 14 \ \text{Vigor} & 30 & 23 & 29 & 22 & 19 & 19 & 28 & 19 \end{array} $$
[5]
Player Anger Vigor ---------------------- 1 6 30 2 7 23 3 5 29 4 21 22 5 13 19 6 5 19 7 13 28 8 14 19 $n = 8$ Anger: values sorted: 5, 5, 6, 7, 13, 13, 14, 21 - Two 5's occupy positions 1, 2 → rank $\frac{1+2}{2} = 1.5$ - 6 → rank 3 - 7 ...
- 105 marksNumericalMeasures of kurtosisHideAnswer
Compute percentile coefficient of kurtosis from the following data and interpret the result.
Hourly wages (Rs) 23-27 28-32 33-37 38-42 43-47 48-52 Number of workers 22 16 9 4 3 1 [5]
Class (Wages) Boundaries f cf ------------ 23-27 22.5-27.5 22 22 28-32 27.5-32.5 16 38 33-37 32.5-37.5 9 47 38-42 37.5-42.5 4 51 43-47 42.5-47.5 3 54 48-52 47.5-52.5 1 55 Total N = 55 Class width $h = 5$. $$k = \frac{QD}{P{90} - P{10}}, ...
- 115 marksNumericalDiscrete distributionsHideAnswer
Write the properties of Poisson distribution. Fit a Poisson distribution and find the expected frequencies.
$$\begin{array}{c|cccccccc} \text{X} & 0 & 1 & 2 & 3 & 4 & 5 & 6 & 7 \ \text{Y} & 71 & 112 & 117 & 57 & 27 & 11 & 3 & 1 \end{array}$$
[5]
Frequency distribution: X 0 1 2 3 4 5 6 7 --------------------------- Y (f) 71 112 117 57 27 11 3 1 All values readable. Nothing missing. --- Let $X \sim P(\lambda)$. 1. PMF:
- 125 marksTypes of DataHideAnswer
Define primary data and secondary data and explain the difference between them. [5]
Primary Data and Secondary Data
Definition of Data
Data is the raw information which is useful for any kind of research or inquiry.
Primary Data
Primary data are those data which are collected by the investigator himself/herself for the first time for a specific purpose.
In other words, primary data are first-hand or original data collected directly from the source.
Example: The Central Bureau of Statistics (CBS) in Nepal conducts a population census every 10 years. The data collected by CBS itself is primary data for CBS.
Secondary Data
Secondary data are those data which are collected by one person/organization and used by another person/organization for their own purpose.
In other words, secondary data are second-hand data.
Example: If a researcher uses the population data published by CBS for their own study, that data becomes secondary data for the researcher.
Differences Between Primary Data and Secondary Data
Basis Primary Data Secondary Data Originality Original in nature; collected by the investigator himself/herself Not original; collected by one and used by another Cost Collection is expensive and time-consuming Less expensive than primary data Method of Collection Collected through direct interview, indirect interview, mailed questionnaire, telephone, etc. Collected from published (e.g., WHO, CBS, Nepal Rastra Bank) and unpublished sources Purpose Collected as per the specific requirement of the investigator May have been collected for a different objective Investigator's Bias Influenced by investigator's bias Not influenced by investigator's bias Reliability Generally more reliable as the investigator controls the process Reliability depends on the original collector
Summary
In short, primary data are fresh, first-time collected data that are more accurate but costly, while secondary data are already available data that are economical but may not perfectly fit the current research objective. The choice between them depends on the nature, scope, and budget of the study.
- 135 marksTypes of samplingHideAnswer
What do you mean by sampling? Explain non probability sampling with merits and demerits. [5]
When one-by-one study of all units of a population is not possible due to factors like time, cost, manpower, resources, and destructive nature of study, we take a small representative part from the population for study. This small repres...