Kindly check out the below video to learn about the concept of mean. Watch the video till the end to complete the task.
 
 
The mean of the ungrouped frequency distribution can be determined using the formula:
\(\overline X = \frac{f_1 x_1 + f_2 x_2 + ... + f_n x_n}{f_1 + f_2 + ... + f_n}\) \(= \frac{\sum_{i=1}^{n} f_i x_i}{\sum_{i=1}^{n} f_i}\)
Example:
The height(in \(cm\)) of \(20\) students in a classroom are:
 
Height
\(x_i\)
Number of students
\(f_i\)
\(130\) \(1\)
\(135\) \(2\)
\(140\) \(1\)
\(155\) \(2\)
\(163\) \(1\)
\(165\) \(3\)
\(177\) \(2\)
\(189\) \(2\)
\(196\) \(2\)
\(100\) \(4\)
 
Find the mean height of the \(20\) students.
 
Solution:
 
To find the value of \(f_ix_i\), multiply the value of \(x\) and \(f\) of each entry.
 
Consider for the mark \(130\). That is, \(130 \times 1 = 130\)
 
Similarly, for the mark \(135\), we have \(135 \times 2 = 270\) and so on.
 
Tabulating these values, we get:
 
Marks
\(x_i\)
Frequency
\(f_i\)
\(f_ix_i\)
\(130\) \(1\) \(130\)
\(135\) \(2\) \(270\)
\(140\) \(1\) \(140\)
\(155\) \(2\) \(310\)
\(163\) \(1\) \(163\)
\(165\) \(3\) \(495\)
\(177\) \(2\) \(354\)
\(189\) \(2\) \(378\)
\(196\) \(2\) \(392\)
\(100\) \(4\) \(400\)
Total \(\sum f_i = 20\) \(\sum f_ix_i = 3032\)
 
Substituting the known values in the above formula, we get:
 
Mean \(\overline X = \frac{3032}{20}\) \(= 151.6\)
 
Therefore, the mean of the given data is \(151.6\).
Mean - Grouped frequency distribution
In most situations, we usually consider a very large amount of data(like population census) for a purposeful study. In such cases, we may find it difficult to write the data in ungrouped data.
 
Hence, to simplify our work, we need to convert the given ungrouped data into grouped data.
 
Let us convert the given height(in \(cm\)) of the students in the classroom into grouped frequency data.
 
\(100\), \(165\), \(189\), \(155\), \(140\), \(196\), \(165\), \(135\), \(135\), \(100\), \(100\), \(100\), \(155\), \(165\), \(196\), \(189\), \(177\), \(163\), \(177\), \(130\).
 
Consider the frequency distribution table.
 
Height (in cm) \(100 - 120\) \(120 - 140\) \(140 - 160\) \(160 - 180\) \(180 - 200\)
Students \(4\) \(3\) \(3\) \(6\) \(4\)
 
The above frequency table shows that the data are grouped in class intervals.
 
Consider the interval \(140 - 160\). There are \(3\) students in the heights between \(140 - 160\) metres. In grouped frequency, individual observations are not available. Thus, we need to determine the value that indicates the particular interval. This value is called a midpoint or class mark. The midpoint can be determined using the formula:
 
Midpoint \(= \frac{UCL + LCL}{2}\)
 
Where \(UCL\) is the upper class limit and \(LCL\) is the lower class limit.
Example:
Consider the interval \(140 - 160\). Let us find the midpoint of this interval.
 
Here, \(UCL = 140\) and \(LCL = 160\)
 
Midpoint of \(140 - 160\) is \(\frac{140 + 160}{2} =\) \(\frac{300}{2}\) \(= 150\)
 
Therefore, the midpoint of the interval \(140 - 160\) is \(150\).
The mean of a grouped frequency distribution can be determined using any one of the following methods.
  • Direct method
  • Assumed mean method
  • Step deviation method
Direct method:
The formula for finding the arithmetic mean using the direct method is given by:
 
\(\overline X = \frac{\sum f_ix_i}{\sum f_i}\)
 
Where \(i\) varies from \(1\) to \(n\), \(x_i\) is the midpoint of the class interval and \(f_i\) is the frequency.
 
Steps:
 
1. Calculate the midpoint of the class interval and name it as \(x_i\).
 
2. Multiply the midpoints\(x_i\) with the frequency\(f_i\) of each class interval and name it as \(f_ix_i\).
 
3. Find the values \(\sum f_ix_i\) and \(\sum f_i\).
 
4. Divide \(\sum f_ix_i\) by \(\sum f_i\) to determine the mean of the data.
Example:
The following frequency distribution table shows that the number of trees based on the height in metres. Find the average height of the trees.
 
Height (in \(m\)) \(30 - 40\) \(40 - 50\) \(50 - 60\) \(60 - 70\) \(70 - 80\)
Number of trees \(124\) \(156\) \(200\) \(10\) \(10\)
 
Solution:
 
Let us form a frequency distribution table.
 
Height
(in \(m\))
Number of trees
(\(f_i\))
Midpoint
(\(x_i\))
\(f_ix_i\)
\(30 - 40\) \(124\) \(35\) \(4340\)
\(40 - 50\) \(156\) \(45\) \(7020\)
\(50 - 60\) \(200\) \(55\) \(11000\)
\(60 - 70\) \(10\) \(65\) \(650\)
\(70 - 80\) \(10\) \(75\) \(750\)
Total \(\sum f_i = 500\)   \(\sum f_ix_i = 23760\)
 
Mean \(\overline X = \frac{23760}{500}\) \(= 47.52\)
 
Therefore, the average height of the trees is \(47.52\).
 
Assumed mean method:
Consider if the data is very large and finding the products of the observations and then adding them becomes tedious and may result in errors. Let us use the assumed mean method to find the mean of grouped frequency to avoid such complications.
 
Steps:
 
1. Calculate the midpoint of the class interval and name it as \(x_i\).
 
2. From the data of \(x_i\), choose any value(preferably in the middle) as the assumed mean(\(a\)).
 
3. Determine the deviation \(d=x-a\) for each of the classes.
 
4. Multiply the deviation and frequency of each class interval and name it \(f_id_i\).
 
5. Find the values \(\sum f_id_i\) and \(\sum f_i\).
 
6. Calculate the mean by applying the formula \(\overline X = a + \frac{\sum f_id_i}{\sum f_i}\)
Example:
Find the mean of the following frequency distribution:
 
Class interval \(10 - 20\) \(20 - 30\) \(30 - 40\) \(40 - 50\)
\(50 - 60\)
\(60 - 70\) \(70 - 80\)
Frequency \(23\) \(15\) \(10\) \(28\) \(5\) \(7\) \(11\)
 
Solution:
 
Class interval
Frequency
(\(f_i\))
Midpoint
(\(x_i\))
Deviation
\(d_i = x - 45\)
\(f_id_i\)
\(10 - 20\) \(23\) \(15\) \(-30\) \(-690\)
\(20 - 30\) \(15\) \(25\) \(-20\) \(-300\)
\(30 - 40\) \(10\) \(35\) \(-10\) \(-100\)
\(40 - 50\) \(28\) \(45\) \(0\) \(0\)
\(50 - 60\) \(5\) \(55\) \(50\) \(250\)
\(60 - 70\) \(8\) \(65\) \(20\) \(160\)
\(70 - 80\) \(11\) \(75\) \(30\) \(330\)
Total \(\sum f_i = 100\)     \(\sum f_id_i = -350\)
We know that the mean of grouped data using the assumed mean method can be determined using the formula \(\overline X = a + \frac{\sum f_id_i}{\sum f_i}\)
Substituting the known values in the above formula, we get:
 
\(\overline X = 45 + (\frac{-350}{100})\)
 
\(\overline X = 45 - 3.5\)
 
\(\overline X = 41.5\)
 
Therefore, the mean of the given data is \(41.5\).
 
Step deviation method:
Let us consider the steps for finding the mean of grouped data using the step deviation method.
 
Steps:
 
1. Calculate the midpoint of the class interval and name it as \(x_i\).
 
2. From the data of \(x_i\), choose any value(preferably in the middle) as the assumed mean(\(a\)).
 
3. Determine the deviation (\(d = x_i - a\)) for each class.
 
4. Determine the deviation (\(u = \frac{x_i - a}{h}\) where \(h\) is the class size) for each class.
 
5. Multiply the frequency and \(u_i\) of each class interval and name it as \(f_iu_i\).
 
6. Calculate the mean by applying the formula \(\overline X = a + \left[\frac{\sum fd}{\sum f} \times h \right]\).
Example:
Find the mean of the following frequency distribution:
 
Class interval \(30 - 40\) \(40 - 50\) \(50 - 60\) \(60 - 70\) \(70 - 80\)
Frequency \(124\) \(156\) \(200\) \(10\) \(10\)
 
Solution:
 
Let the assumed mean be \(a = 55\) and class width \(h = 10\).
 
Class interval
Frequency
\(f_i\)
Midpoint
\(x_i\)
Deviation
\(d_i = x_i - 55\)
\(u_i = \frac{x - a}{h}\)
\(f_iu_i\)
\(30 - 40\) \(124\) \(35\) \(-20\) \(-2\) \(-248\)
\(40 - 50\) \(156\) \(45\) \(-10\) \(-1\) \(-156\)
\(50 - 60\) \(200\) \(55\) \(0\) \(0\) \(0\)
\(60 - 70\) \(10\) \(65\) \(10\) \(1\) \(10\)
\(70 - 80\) \(10\) \(75\) \(20\) \(2\) \(20\)
Total \(\sum f_i = 500\)       \(\sum f_iu_i = -374\)
We know that the mean of the grouped frequency distribution using the step deviation method can be determined using the formula, \(\overline X = a + \left[\frac{\sum f_iu_i}{\sum f_i} \times h \right]\).
Substituting the known values in the above formula, we have:
 
\(\overline X = 55 + \left[\frac{-374}{500} \times 10 \right]\)
 
\(\overline X = 55 - 7.48\)
 
\(\overline X = 47.52\)
 
Therefore, the mean of the given data is \(47.52\).