Mathematics · Chapter 5
Study notes aligned to the official NEB syllabus.
A measure of central tendency (mean, median, mode) tells us the single value around which a data set clusters, but it says nothing about how tightly or loosely the values are spread. Two classes can share the same average marks yet be completely different: one where almost everyone scores near the mean, and one where marks are scattered from very low to very high. Statistics in this chapter studies that scatter through measures of dispersion, refines the picture with skewness, and then turns to probability, the numerical measure of uncertainty. Learn the formulas exactly, keep the columns of your working table neat, and always state units where they apply.
Throughout, $N = \sum f$ is the total frequency (total number of items), $\bar{x}$ is the arithmetic mean, and for grouped data $x$ is the mid-value of a class.
Dispersion is the extent to which the individual items of a data set differ from a central value and from one another. The measures fall into two families.
A good measure of dispersion should be rigidly defined, easy to understand, based on all the observations, suitable for further algebraic treatment, and not unduly affected by extreme values or sampling fluctuation.
The range is the crudest measure, just the gap between the largest and smallest values:
$$ \begin{aligned} R &= L - S \ \qquad \text{coefficient of range} &= \frac{L - S}{L + S}. \end{aligned} $$
The quartile deviation (semi-interquartile range) uses the first and third quartiles $Q_1$ and $Q_3$, so it ignores the extreme quarter on each end:
$$ \begin{aligned} Q.D. &= \frac{Q_3 - Q_1}{2} \ \qquad \text{coefficient of Q.D.} &= \frac{Q_3 - Q_1}{Q_3 + Q_1}. \end{aligned} $$
The mean deviation is the average of the absolute deviations of the items from a central value $A$ (usually the mean or median):
$$ \begin{aligned} M.D. &= \frac{\sum f,|x - A|}{N} \ \qquad \text{coefficient of M.D.} &= \frac{M.D.}{A}. \end{aligned} $$
The modulus is essential, because without it the signed deviations from the mean always add to zero. Mean deviation uses every item but the absolute value makes it awkward for further algebra, which is exactly why the standard deviation is preferred.
The standard deviation is the most important and most widely used measure of dispersion. It is defined as the positive square root of the mean of the squared deviations of the items from their arithmetic mean, and is denoted by $\sigma$ (sigma). Squaring removes the sign problem cleanly and keeps the measure open to algebra.
Individual series (ungrouped data $x_1, x_2, \dots, x_n$):
$$ \begin{aligned} \sigma &= \sqrt{\frac{\sum (x - \bar{x})^2}{n}} \ &= \sqrt{\frac{\sum x^2}{n} - \left(\frac{\sum x}{n}\right)^2}. \end{aligned} $$
Discrete series (values $x$ with frequencies $f$, $N = \sum f$):
$$ \begin{aligned} \sigma &= \sqrt{\frac{\sum f (x - \bar{x})^2}{N}} \ &= \sqrt{\frac{\sum f x^2}{N} - \left(\frac{\sum f x}{N}\right)^2}. \end{aligned} $$
Continuous series (classes with mid-values $x$ and frequencies $f$): the same formula as the discrete case, with $x$ the class mid-value. When the numbers are large it is easier to use the step-deviation (shortcut) method with an assumed mean $A$ and common class width $h$, putting $d = \dfrac{x - A}{h}$:
$$\sigma = h\sqrt{\frac{\sum f d^2}{N} - \left(\frac{\sum f d}{N}\right)^2}.$$
The second (root-mean-square) form of each formula is almost always faster in the exam than computing every $(x - \bar{x})^2$ by hand.
A measure of central tendency (mean, median, mode) tells us the single value around which a data set clusters, but it says nothing about how tightly or loosely the values are spread. Two classes can share the same average marks yet be completely different: one where almost everyone scores near the mean, and one where marks are scattered from very low to very high. Statistics in this chapter studies that scatter through measures of dispersion, refines the picture with skewness, and then turns to probability, the numerical measure of uncertainty. Learn the formulas exactly, keep the columns of your working table neat, and always state units where they apply.
Throughout, is the total frequency (total number of items), is the arithmetic mean, and for grouped data is the mid-value of a class.
Dispersion is the extent to which the individual items of a data set differ from a central value and from one another. The measures fall into two families.
A good measure of dispersion should be rigidly defined, easy to understand, based on all the observations, suitable for further algebraic treatment, and not unduly affected by extreme values or sampling fluctuation.
The range is the crudest measure, just the gap between the largest and smallest values:
The quartile deviation (semi-interquartile range) uses the first and third quartiles and , so it ignores the extreme quarter on each end:
The mean deviation is the average of the absolute deviations of the items from a central value (usually the mean or median):
The modulus is essential, because without it the signed deviations from the mean always add to zero. Mean deviation uses every item but the absolute value makes it awkward for further algebra, which is exactly why the standard deviation is preferred.
The standard deviation is the most important and most widely used measure of dispersion. It is defined as the positive square root of the mean of the squared deviations of the items from their arithmetic mean, and is denoted by (sigma). Squaring removes the sign problem cleanly and keeps the measure open to algebra.
Individual series (ungrouped data ):
Discrete series (values with frequencies , ):
Continuous series (classes with mid-values and frequencies ): the same formula as the discrete case, with the class mid-value. When the numbers are large it is easier to use the step-deviation (shortcut) method with an assumed mean and common class width , putting :
The second (root-mean-square) form of each formula is almost always faster in the exam than computing every by hand.