2080

STA169 · TU past paper

Statistics I 2080 question paper

The complete TU 2080 exam paper for Statistics I (STA169), all 13 questions with solved model answers written to the mark scheme.

Tap a question to open its answer.

  1. 110 marksNumericalMeasures of central tendencyAnswer

    Measures of Central Tendency and Dispersion in Descriptive Statistics

    Measures of Central Tendency and Dispersion - Verified Model Answer

    STEP 1 - EXTRACT (Given Data)

    Grouped frequency distribution, $n = 50$:

    Mass (lbs)1-22-33-44-55-66-7
    Frequency $f$81015962

    $\sum f = 8+10+15+9+6+2 = 50$ ✓

    Required: mean, standard deviation, variance, coefficient of variation, plus conceptual explanation.


    STEP 2 - SOLVE

    Part 1: Concepts

    Measures of central tendency describe the center or typical value of a dataset: mean, median, mode.

    Measures of dispersion describe the spread or variability of data about the central value.

    Different measures of dispersion:

    • Range (max − min)
    • Quartile deviation / semi-interquartile range = $(Q_3 - Q_1)/2$
    • Mean deviation (average absolute deviation)
    • Standard deviation and variance
    • Coefficient of variation (relative measure)

    Part 2: Computation

    Midpoints and products:

    Class$x$$f$$fx$$fx^2$
    1-21.5812.018.00
    2-32.51025.062.50
    3-43.51552.5183.75
    4-54.5940.5182.25
    5-65.5633.0181.50
    6-76.5213.084.50
    Total50176.0712.50

    $$\sum f = 50,\quad \sum fx = 176.0,\quad \sum fx^2 = 712.50$$

    Mean: $$\bar{x} = \frac{\sum fx}{\sum f} = \frac{176.0}{50} = 3.52 \text{ lbs}$$

    Variance and Standard Deviation (population form, standard for grouped data):

    $$\sigma^2 = \frac{\sum fx^2}{n} - \bar{x}^2 = \frac{712.50}{50} - (3.52)^2$$ $$= 14.25 - 12.3904 = 1.8596 \approx 1.86 \text{ lbs}^2$$

    $$\sigma = \sqrt{1.8596} = 1.3637 \approx 1.36 \text{ lbs}$$

    (Note: the sample form with $n-1$ gives $s^2 = 92.98/49 = 1.8975 \approx 1.90$ and $s = 1.38$. Both are acceptable; the population form is the conventional choice for descriptive grouped-data problems and is the one reported above.)

    Coefficient of Variation: $$CV = \frac{\sigma}{\bar{x}} \times 100 = \frac{1.3637}{3.52} \times 100 = 38.74%$$


    Summary Table

    MeasureValue (population)Value (sample, n−1)
    Mean $\bar{x}$3.52 lbs3.52 lbs
    Variance1.86 lbs²1.90 lbs²
    Standard Deviation1.36 lbs1.38 lbs
    Coefficient of Variation38.74%39.20%

    Part 4: Interpretation

    • Mean = 3.52 lbs: The average mass is about 3.52 lbs, centered in the modal 3-4 class (highest frequency = 15).
    • Standard Deviation ≈ 1.36 lbs: Individual masses deviate from the mean by about 1.36 lbs on average, indicating a moderate spread.
    • Variance ≈ 1.86 lbs²: The squared spread measure; SD is more interpretable being in original units.
    • CV ≈ 38.7%: A relatively high CV, showing the data are fairly variable relative to the mean; since CV > 30% the distribution is not highly homogeneous.

    The only judgement call here is the use of the sample ($n-1$) rather than the population ($n$) divisor, which yields slightly different SD, variance and CV values; both are defensible.

  2. 210 marksNumericalKarl Pearson's coefficient of correlationAnswer

    Correlation and Regression Analysis

    Correlation and Regression Analysis

    Given Data

    $n = 8$

    Number of weeks (x)453971045
    Speed gain (y)851204819216423474110

    Meaning of Correlation Types

    i. Positive Correlation: Two variables move in the same direction. When one increases, the other also increases (and vice versa). Example: height and weight.

    ii. Negative Correlation: Two variables move in opposite directions. When one increases, the other decreases. Example: price and demand.

    iii. Perfect Correlation: A change in one variable produces an exactly proportional change in the other, so all points lie on a straight line. Perfect positive gives $r = +1$; perfect negative gives $r = -1$.


    Calculation Table

    xyx²y²xy
    485167225340
    51202514400600
    34892304144
    919281368641728
    716449268961148
    10234100547562340
    474165476296
    51102512100550
    Σx=47Σy=1027Σx²=321Σy²=160021Σxy=7146

    Verification of sums: $\Sigma x = 47$, $\Sigma y = 1027$, $\Sigma x^2 = 321$, $\Sigma xy = 7146$, and $\Sigma y^2 = 160021$. All confirmed.


    Part (a): Correlation Coefficient

    $$r = \frac{n\Sigma xy - \Sigma x \Sigma y}{\sqrt{[n\Sigma x^2 - (\Sigma x)^2][n\Sigma y^2 - (\Sigma y)^2]}}$$

    Numerator: $$8(7146) - 47(1027) = 57168 - 48269 = 8899$$

    First bracket: $$8(321) - 47^2 = 2568 - 2209 = 359$$

    Second bracket: $$8(160021) - 1027^2 = 1280168 - 1054729 = 225439$$

    $$r = \frac{8899}{\sqrt{359 \times 225439}} = \frac{8899}{\sqrt{80932601}} = \frac{8899}{8996.26}$$

    $$\boxed{r \approx 0.989}$$

    Interpretation: Since $r = 0.989$ is very close to $+1$, there is a very high positive correlation between weeks in the program and speed gain. As weeks increase, speed gain increases strongly.


    Part (b): Regression Equation of Speed Gain (y) on Weeks (x)

    $$\hat{y} = a + bx$$

    $$b = \frac{n\Sigma xy - \Sigma x \Sigma y}{n\Sigma x^2 - (\Sigma x)^2} = \frac{8899}{359} = 24.79$$

    $$\bar{x} = \frac{47}{8} = 5.875, \quad \bar{y} = \frac{1027}{8} = 128.375$$

    $$a = \bar{y} - b\bar{x} = 128.375 - 24.79(5.875) = 128.375 - 145.64 = -17.26$$

    $$\boxed{\hat{y} = -17.26 + 24.79x}$$


    Part (c): Estimate for 6 Weeks and Slope Interpretation

    $$\hat{y} = -17.26 + 24.79(6) = -17.26 + 148.74 = 131.48$$

    $$\boxed{\hat{y} \approx 131.48 \text{ units of speed gain}}$$

    Slope interpretation: $b = 24.79$ means that for each additional week in the program, the reading speed gain increases by about 24.79 units on average.

  3. 310 marksNumericalJoint probability distribution of two randAnswer

    The joint density function of two continuous random variables X and Y is f(x, y) = kxy 0 < x < 4, 1 < y < 5= 0 otherwisea. Find the value of constant k. b. Find P(x > 3, y < 2) c. Find P(1 < x < 2, 2 < y < 3)[10]

    Joint Probability Density Function - Verified Solution

    STEP 1 - Given Data

    $$f(x, y) = kxy, \quad 0 < x < 4, \quad 1 < y < 5$$ $$f(x, y) = 0, \quad \text{otherwise}$$

    Required:

    • (a) constant $k$
    • (b) $P(X > 3, Y < 2)$
    • (c) $P(1 < X < 2, 2 < Y < 3)$

    STEP 2 - Solve

    Part (a): Value of k

    Normalization condition:

    $$\int_{1}^{5} \int_{0}^{4} kxy , dx , dy = 1$$

    Inner integral (over $x$):

    $$\int_{0}^{4} x , dx = \left[\frac{x^2}{2}\right]_0^4 = \frac{16}{2} = 8$$

    So:

    $$k \int_{1}^{5} 8y , dy = 8k \left[\frac{y^2}{2}\right]_1^5 = 8k \cdot \frac{25 - 1}{2} = 8k \cdot 12 = 96k$$

    Setting equal to 1:

    $$96k = 1 \implies \boxed{k = \frac{1}{96}}$$


    Part (b): P(X > 3, Y < 2)

    The region intersects the support: $3 < x < 4$ and $1 < y < 2$.

    $$P = \frac{1}{96} \int_{1}^{2} \int_{3}^{4} xy , dx , dy$$

    Inner integral:

    $$\int_{3}^{4} x , dx = \left[\frac{x^2}{2}\right]_3^4 = \frac{16 - 9}{2} = \frac{7}{2}$$

    Then:

    $$P = \frac{1}{96} \cdot \frac{7}{2} \int_{1}^{2} y , dy = \frac{7}{192} \left[\frac{y^2}{2}\right]_1^2 = \frac{7}{192} \cdot \frac{4-1}{2} = \frac{7}{192} \cdot \frac{3}{2} = \frac{21}{384}$$

    $$\boxed{P(X > 3, Y < 2) = \frac{7}{128} \approx 0.0547}$$


    Part (c): P(1 < X < 2, 2 < Y < 3)

    Both limits lie within the support.

    $$P = \frac{1}{96} \int_{2}^{3} \int_{1}^{2} xy , dx , dy$$

    Inner integral:

    $$\int_{1}^{2} x , dx = \left[\frac{x^2}{2}\right]_1^2 = \frac{4 - 1}{2} = \frac{3}{2}$$

    Then:

    $$P = \frac{1}{96} \cdot \frac{3}{2} \int_{2}^{3} y , dy = \frac{3}{192} \left[\frac{y^2}{2}\right]_2^3 = \frac{1}{64} \cdot \frac{9-4}{2} = \frac{1}{64} \cdot \frac{5}{2}$$

    $$\boxed{P(1 < X < 2, 2 < Y < 3) = \frac{5}{128} \approx 0.0391}$$


    Summary

    PartResult
    (a) $k$$\dfrac{1}{96}$
    (b) $P(X>3, Y<2)$$\dfrac{7}{128} \approx 0.0547$
    (c) $P(1<X<2, 2<Y<3)$$\dfrac{5}{128} \approx 0.0391$
  4. 45 marksTypes of DataAnswer

    Differentiate between primary data and secondary data. What are the sources of secondary data? [5]

    Primary Data: Primary data are those data which are collected by the investigator himself/herself for the first time. They are original in nature and collected for a specific purpose of inquiry. Example: Data collected by CBS (Central Bu...

  5. 55 marksTypes of samplingAnswer

    What is sampling? Define simple random sampling and stratified random sampling with some relevant examples. [5]

    Sampling, Simple Random Sampling, and Stratified Random Sampling

    Sampling

    When one-by-one study of all units of a population is not possible due to factors like time, cost, manpower, resources, and destructive nature of study, we take a small representative part from the population for study. This small representative part selected for study from the population is called a sample, and the process of selecting a sample from a population is called sampling.

    Example: A pathologist takes a syringe of blood as a sample to find out a disease.


    Simple Random Sampling

    Simple random sampling is the most common and simplest method of sampling in which each sample unit is selected from a population with equal probability. Every unit in the population has a fixed and equal chance of being selected in the sample.

    Example: Suppose a teacher wants to select 5 students from a class of 50. Each student is assigned a number from 1 to 50, and 5 numbers are drawn randomly (using a lottery or random number table). Every student has an equal probability (5/50 = 1/10) of being selected.

    Limitations of Simple Random Sampling

    • Requires an up-to-date sampling frame.
    • When sampling units are widely spread geographically, the cost of collecting data may be high in terms of time and money.
    • For a given precision, it usually requires a larger sample size compared to stratified random sampling.

    Stratified Random Sampling

    When units in the population are not similar in nature, the population is first divided into subgroups called strata before the sample is drawn. Then a simple random sample is drawn from each stratum in proportion to its size. This method is called stratified random sampling.

    Rules for Stratification

    • The strata should be non-overlapping and should together comprise the whole population.
    • Strata should be as homogeneous within groups and heterogeneous between groups as possible.

    Purposes of Stratification

    • To make the sample more representative.
    • For greater accuracy.
    • For administrative convenience.

    Example: Suppose a researcher wants to study the income level of 1000 people in a city divided into three income groups:

    StratumGroupPopulation SizeSample (10%)
    1Low income50050
    2Middle income30030
    3High income20020

    A simple random sample is drawn from each stratum proportionally, ensuring all income groups are represented in the final sample of 100.


    Key Difference

    FeatureSimple Random SamplingStratified Random Sampling
    Population natureHomogeneousHeterogeneous
    DivisionNo divisionDivided into strata
    AccuracyRelatively lowerHigher
    RepresentationMay miss subgroupsAll subgroups represented
  6. 65 marksNumericalfive number summaryAnswer

    What aspect of summary measures of data can be explained by the measures of skewness? Kelvin Hota is the national sales manager for National Text Books. He has a sales staff of 10 who visit college professors all over the United States. Each Sunday morning he requires his sales staffs to send him a report. Listed below are the number of visits last week. Compute the five number summary. 25, 6, 10, 13, 15, 2, 18, 5, 20, 30 [5]

    Summary measures describe several aspects of data: central tendency (mean, median, mode), dispersion (range, variance), and shape. Measures of skewness explain the shape aspect, specifically the lack of symmetry of a distribution. Skewne...

  7. 75 marksNumericalMeasures of central tendencyAnswer

    What are the requisites for good average?

    From the following distribution of marks of 200 students of a college:

    $$\begin{array}{c|ccccccc} \text{Marks} & 30\text{-}40 & 40\text{-}50 & 50\text{-}60 & 60\text{-}70 & 70\text{-}80 & 80\text{-}90 \ \hline \text{No. of students} & 14 & 50 & 60 & 45 & 20 & 11 \ \end{array}$$

    Compute:

    i. The minimum marks obtained by top 10% students.

    ii. Modal marks.

    [5]

    Requisites for a Good Average & Distribution Analysis

    Requisites for a Good Average

    1. Rigidly defined - It should have a clear, unambiguous definition.
    2. Based on all observations - It should use every value in the data.
    3. Easy to understand and compute - Simple to calculate and interpret.
    4. Capable of further algebraic treatment - Suitable for further mathematical work.
    5. Least affected by sampling fluctuations - Stable across samples.
    6. Not unduly affected by extreme values - Robust against abnormal observations.

    Given Data

    Marksfcf
    30-401414
    40-505064
    50-6060124
    60-7045169
    70-8020189
    80-9011200
    Total200

    Part (i): Minimum Marks of Top 10% Students

    Top 10% means the highest 10% of scores, i.e. the marks exceeded by 90% of students. This is the 90th percentile $P_{90}$.

    $$\frac{90N}{100} = \frac{90 \times 200}{100} = 180$$

    Locate 180 in the cf column: cf = 169 (up to 60-70), cf = 189 (up to 70-80). So the 180th value lies in class 70-80.

    • $L = 70$, $cf = 169$, $f = 20$, $h = 10$

    $$P_{90} = L + \frac{\frac{90N}{100} - cf}{f} \times h = 70 + \frac{180 - 169}{20} \times 10$$

    $$P_{90} = 70 + \frac{11}{20} \times 10 = 70 + 5.5 = 75.5$$

    $$\boxed{P_{90} = 75.5 \text{ marks}}$$

    The minimum marks obtained by the top 10% of students is 75.5.


    Part (ii): Modal Marks

    Highest frequency = 60, so modal class is 50-60.

    • $L = 50$, $f_1 = 60$, $f_0 = 50$, $f_2 = 45$, $h = 10$

    $$\text{Mode} = L + \frac{f_1 - f_0}{2f_1 - f_0 - f_2} \times h$$

    $$= 50 + \frac{60 - 50}{2(60) - 50 - 45} \times 10 = 50 + \frac{10}{120 - 95} \times 10$$

    $$= 50 + \frac{10}{25} \times 10 = 50 + 4 = 54$$

    $$\boxed{\text{Mode} = 54 \text{ marks}}$$

    The modal marks of the distribution is 54.

  8. 85 marksNumericalMomentsAnswer

    The first four moments of a distribution about x=2 are 1,2.5,5.5 and 16. Calculate the first four moments about the mean. Test the skewness and kurtosis. Interpret the results. [5]

    Moments about $x = 2$ (arbitrary origin $a = 2$): $$\mu1' = 1, \quad \mu2' = 2.5, \quad \mu3' = 5.5, \quad \mu4' = 16$$ --- First: $$\mu1 = 0$$ Second: $$\mu2 = \mu2' - (\mu1')^2 = 2.5 - 1 = 1.5$$ Third: $$\mu3 = \mu3' - 3\mu2'\mu1' + 2(...

  9. 95 marksNumericalConcepts of probabilityAnswer

    Define independent and mutually exclusive events. Three groups of children 2 boys and 2 girls, 3 boys and 1 girl, 1 boy and 3 girls respectively. One child is selected at random from each group. Find the probability of selecting one boy and two girls. [5]

    • Group 1: 2 boys, 2 girls (total 4) - Group 2: 3 boys, 1 girl (total 4) - Group 3: 1 boy, 3 girls (total 4) - One child selected at random from each group. - Required: $P(\text{1 boy and 2 girls})$ Mutually Exclusive Events: Two events ...
  10. 105 marksNumericalBayes theoremAnswer

    What is conditional probability? Three roads A, B and C lead away from a jail. A prisoner escaping from the jail selects a road at random. If road A is selected, the probability of escaping is 1/10. Similarly for road B it is 1/8 and for road C it is 1/5. i. What is the probability that the prisoner will succeed in escaping? ii. If the prisoner has succeeded in escaping, what is the probability that he had chosen the road A? [5]

    Conditional Probability and Prisoner Escape Problem

    Step 1 - Given Data

    • Roads: A, B, C selected at random, so $P(A) = P(B) = P(C) = \dfrac{1}{3}$
    • $P(E \mid A) = \dfrac{1}{10}$
    • $P(E \mid B) = \dfrac{1}{8}$
    • $P(E \mid C) = \dfrac{1}{5}$

    where $E$ = event of escaping.

    Definition: Conditional Probability

    Conditional probability is the probability of an event occurring given that another event has already occurred. For events $A$ and $B$:

    $$P(A \mid B) = \frac{P(A \cap B)}{P(B)}, \quad P(B) > 0$$

    Step 2 - Solve

    Part (i): Probability of Escape

    By the Theorem of Total Probability:

    $$P(E) = P(A)P(E \mid A) + P(B)P(E \mid B) + P(C)P(E \mid C)$$

    $$P(E) = \frac{1}{3}\cdot\frac{1}{10} + \frac{1}{3}\cdot\frac{1}{8} + \frac{1}{3}\cdot\frac{1}{5}$$

    $$P(E) = \frac{1}{3}\left(\frac{1}{10} + \frac{1}{8} + \frac{1}{5}\right)$$

    LCM of $10, 8, 5 = 40$:

    $$\frac{1}{10} + \frac{1}{8} + \frac{1}{5} = \frac{4 + 5 + 8}{40} = \frac{17}{40}$$

    $$P(E) = \frac{1}{3}\cdot\frac{17}{40} = \frac{17}{120} \approx 0.1417$$

    Part (ii): Probability Road A Chosen Given Escape

    By Bayes' Theorem:

    $$P(A \mid E) = \frac{P(A)P(E \mid A)}{P(E)} = \frac{\frac{1}{3}\cdot\frac{1}{10}}{\frac{17}{120}} = \frac{\frac{1}{30}}{\frac{17}{120}}$$

    $$P(A \mid E) = \frac{1}{30}\times\frac{120}{17} = \frac{120}{510} = \frac{4}{17} \approx 0.2353$$

    Summary

    Road$P(\cdot)$$P(E\mid\cdot)$Product
    A1/31/104/120
    B1/31/85/120
    C1/31/58/120
    Total17/120
    • $P(\text{escape}) = \dfrac{17}{120} \approx 0.142$
    • $P(A \mid \text{escape}) = \dfrac{4}{17} \approx 0.235$
  11. 115 marksNumericalDiscrete distributionsAnswer

    Define binomial distribution. Under what conditions is the binomial distribution appropriate? The probability of a novice archer hitting the target with any shot is 0.3. Given that the archer shoots six arrows, find the probability that the target is hit at least once. [5]

    Binomial Distribution

    STEP 1 - Given data

    • Probability of success (hit) per shot: $p = 0.3$
    • Probability of failure (miss): $q = 1 - 0.3 = 0.7$
    • Number of independent trials (arrows): $n = 6$
    • Required: $P(\text{at least one hit}) = P(X \ge 1)$

    All required data present.

    STEP 2 - Solution

    Definition

    A binomial distribution is a discrete probability distribution giving the probability of obtaining exactly $x$ successes in $n$ independent trials, where each trial results in only one of two outcomes (success or failure) with a constant success probability $p$.

    Its probability mass function is:

    $$P(X = x) = \binom{n}{x} p^x q^{,n-x}, \quad x = 0, 1, 2, \ldots, n$$

    where $q = 1 - p$. It is a biparametric distribution with parameters $n$ and $p$.

    Mean $= np$, Variance $= npq$.

    Conditions for applicability

    1. The number of trials $n$ is fixed and finite.
    2. Each trial has only two mutually exclusive outcomes (success/failure).
    3. The trials are independent.
    4. The probability of success $p$ is constant across all trials.

    Numerical part

    Here each shot is a trial with two outcomes (hit/miss), $n = 6$ fixed, independent shots, and constant $p = 0.3$: so the binomial model applies.

    $$X \sim B(n = 6,\ p = 0.3)$$

    Using the complement:

    $$P(X \ge 1) = 1 - P(X = 0)$$

    $$P(X = 0) = \binom{6}{0}(0.3)^0 (0.7)^6 = 1 \cdot 1 \cdot (0.7)^6$$

    $$(0.7)^6 = 0.117649$$

    Therefore:

    $$P(X \ge 1) = 1 - 0.117649 = 0.882351$$

    $$\boxed{P(X \ge 1) \approx 0.8824}$$

    Interpretation: There is about an 88.24% chance the archer hits the target at least once in six shots.

  12. 125 marksNumericalContinuous distributionAnswer

    State features of normal distribution. In a photographic process, the developing time of prints as a random variable having normal distribution with mean of 18.25 seconds with standard deviation 0.34 seconds. Find the probability that at least 17.64 seconds to develop one of the prints. [5]

    Normal Distribution: Features and Probability

    Features of Normal Distribution

    1. Bell-shaped and symmetric about the mean $\mu$; the two halves are mirror images.
    2. Mean = Median = Mode, all located at the center $\mu$.
    3. Completely defined by two parameters: mean $\mu$ and standard deviation $\sigma$.
    4. Total area under the curve equals $1$ (total probability = 1).
    5. Asymptotic: the curve approaches but never touches the x-axis, ranging from $-\infty$ to $+\infty$.
    6. It follows the empirical rule: about 68% of values lie within $\mu \pm \sigma$, 95% within $\mu \pm 2\sigma$, and 99.7% within $\mu \pm 3\sigma$.
    7. PDF:

    $$f(x) = \frac{1}{\sigma\sqrt{2\pi}} e^{-\frac{1}{2}\left(\frac{x-\mu}{\sigma}\right)^2}, \quad -\infty < x < \infty$$


    Numerical Solution

    Given data:

    • Mean $\mu = 18.25$ s
    • Standard deviation $\sigma = 0.34$ s
    • Required: $P(X \geq 17.64)$

    Step 1: Standardize

    $$Z = \frac{X - \mu}{\sigma} = \frac{17.64 - 18.25}{0.34} = \frac{-0.61}{0.34} = -1.79$$

    Step 2: Express probability

    $$P(X \geq 17.64) = P(Z \geq -1.79)$$

    Step 3: Apply symmetry

    $$P(Z \geq -1.79) = P(Z \leq 1.79) = 0.5 + P(0 \leq Z \leq 1.79)$$

    Step 4: Z-table value

    $$P(0 \leq Z \leq 1.79) = 0.4633$$

    Step 5: Final result

    $$P(X \geq 17.64) = 0.5 + 0.4633 = \boxed{0.9633}$$

    Interpretation: There is approximately a 96.33% probability that a print takes at least 17.64 seconds to develop.

  13. 135 marksNumericalMathematical expectation of a random variaAnswer

    Define random variable. If X is the number of points rolled with a balanced die, find E(X) and variance of X. Also find the expected value of random variable g(X) = 2X^2 + 1. [5]

    Random Variable: Definition, E(X), Variance, and E(g(X))

    Step 1 - Extract (Given Data)

    • Experiment: rolling a balanced (fair) die
    • Random variable $X$ = number of points shown
    • Possible values: $x = 1, 2, 3, 4, 5, 6$
    • $P(X = x) = \frac{1}{6}$ for each value (balanced die)
    • Function required: $g(X) = 2X^2 + 1$

    All data present and sufficient.

    Step 2 - Solve

    Definition of Random Variable

    A random variable is a real-valued function that assigns a numerical value to each outcome in the sample space of a random experiment. It may be discrete (countable values) or continuous (values over an interval).

    Probability Distribution

    $$P(X = x) = \frac{1}{6}, \quad x = 1, 2, 3, 4, 5, 6$$

    Calculation of E(X)

    $$E(X) = \sum x \cdot P(X=x) = \frac{1+2+3+4+5+6}{6} = \frac{21}{6} = \frac{7}{2} = 3.5$$

    Calculation of Variance

    $$E(X^2) = \frac{1^2+2^2+3^2+4^2+5^2+6^2}{6} = \frac{1+4+9+16+25+36}{6} = \frac{91}{6}$$

    $$\text{Var}(X) = E(X^2) - [E(X)]^2 = \frac{91}{6} - \left(\frac{7}{2}\right)^2 = \frac{91}{6} - \frac{49}{4}$$

    $$= \frac{182}{12} - \frac{147}{12} = \frac{35}{12} \approx 2.917$$

    Expected Value of g(X) = 2X² + 1

    $$E(2X^2 + 1) = 2E(X^2) + 1 = 2 \cdot \frac{91}{6} + 1 = \frac{182}{6} + \frac{6}{6} = \frac{188}{6} = \frac{94}{3} \approx 31.33$$

    Summary

    QuantityValue
    $E(X)$$\frac{7}{2} = 3.5$
    $\text{Var}(X)$$\frac{35}{12} \approx 2.917$
    $E(2X^2+1)$$\frac{94}{3} \approx 31.33$