01 · Explore
Introduction to Grouped Data
In previous classes, we studied ungrouped data and basic pictorial representations. In Class 10, we extend these concepts to grouped data, which is essential for handling large real-world datasets.
Statistics involves the collection, organization, and analysis of data. While ungrouped data lists every individual observation, grouped data organizes these observations into class intervals. This condensation makes it easier to identify patterns in large datasets, such as the marks of all students in a city or the heights of a large population.
To analyze grouped data, we focus on three measures of central tendency: the mean (average), the mode (most frequent value), and the median (middle-most value). Because grouping data involves some loss of individual detail, the values we calculate for grouped data are often estimates rather than the exact values obtained from raw data.
A key concept in grouped data is the 'class mark' or mid-point. We assume that the frequency of each class interval is centered around its mid-point. The class mark is calculated as the average of the upper and lower limits of the class: Class mark = (Upper class limit + Lower class limit) / 2.
Pause & Try
Think it through first. Writing and checking your answer is free with an account.
Question
What is the class mark for the interval 25 - 40?
Sign in to see the answer
Write your own answer and compare it with ours. It’s free.
Sign inNew here? Sign up freeNCERT reference: chapter PDF pages 1, 3.
02 · Explore
Mean of Grouped Data: Direct Method
The mean is the sum of all observations divided by the total number of observations. For grouped data, we use class marks to represent each interval.
In the Direct Method, we multiply each class mark (xi) by its corresponding frequency (fi) to find the product fi*xi. The sum of these products (Σfixi) is then divided by the total frequency (Σfi).
The formula is expressed as: x̄ = (Σfixi) / (Σfi). This method is most suitable when the numerical values of the class marks and frequencies are small and easy to multiply.
It is important to note that the mean calculated from grouped data may differ slightly from the mean of the original ungrouped data. This is because the grouping process assumes all values in an interval are exactly at the mid-point, which is an approximation.
Calculating Mean via Direct Method
x̄ = 1860 / 30 = 62
Given class intervals and frequencies: 10-25 (f=2, xi=17.5), 25-40 (f=3, xi=32.5), 40-55 (f=7, xi=47.5), 55-70 (f=6, xi=62.5), 70-85 (f=6, xi=77.5), 85-100 (f=6, xi=92.5). Sum of fi = 30. Sum of fixi = (2*17.5) + (3*32.5) + (7*47.5) + (6*62.5) + (6*77.5) + (6*92.5) = 35 + 97.5 + 332.5 + 375 + 465 + 555 = 1860. Dividing 1860 by 30 gives a mean of 62.
Pause & Try
Think it through first. Writing and checking your answer is free with an account.
Question
When is the Direct Method considered appropriate?
Sign in to see the answer
Write your own answer and compare it with ours. It’s free.
Sign inNew here? Sign up freeNCERT reference: chapter PDF pages 2, 3, 4.
03 · Explore
Assumed Mean Method
When class marks are large, the Assumed Mean Method reduces the size of the numbers we work with by shifting the origin.
In this method, we choose one of the class marks (usually the middle one) as the 'assumed mean', denoted by 'a'. We then calculate the deviation (di) for each class mark: di = xi - a.
Instead of multiplying frequencies by the large xi values, we multiply them by the smaller deviations (di). The mean is then found by adding the average of these deviations back to the assumed mean.
The formula is: x̄ = a + (Σfidi) / (Σfi). This method significantly simplifies calculations when dealing with large numbers, as the products fidi are much smaller than fixi.
| Class Mark (xi) | Assumed Mean (a) | Deviation (di = xi - a) |
|---|---|---|
| 17.5 | 47.5 | -30 |
| 32.5 | 47.5 | -15 |
| 47.5 | 47.5 | 0 |
| 62.5 | 47.5 | 15 |
Pause & Try
Think it through first. Writing and checking your answer is free with an account.
Question
If the assumed mean 'a' is 50 and the average of deviations (Σfidi/Σfi) is -2.5, what is the actual mean?
Sign in to see the answer
Write your own answer and compare it with ours. It’s free.
Sign inNew here? Sign up freeNCERT reference: chapter PDF pages 4, 5, 6.
04 · Explore
Step-Deviation Method
The Step-Deviation Method is the most advanced technique for calculating the mean, further simplifying the Assumed Mean Method by dividing deviations by the class size.
If all deviations (di) have a common factor, usually the class size 'h', we can divide each deviation by 'h' to get even smaller values: ui = (xi - a) / h.
We then calculate the mean of these step-deviations (ū = Σfiui / Σfi). To find the actual mean, we multiply this average by 'h' and add it to the assumed mean 'a'.
The formula is: x̄ = a + h * (Σfiui / Σfi). This method is highly efficient for datasets with equal class sizes and large numerical values, as it reduces the multiplication to the smallest possible integers.
- 1
Identify class marks (xi) and choose an assumed mean (a).
- 2
Determine the class size (h).
- 3
Calculate ui = (xi - a) / h for each class.
- 4
Multiply each frequency (fi) by its corresponding ui to get fiui.
- 5
Find the sum Σfiui and the total frequency Σfi.
- 6
Apply the formula: x̄ = a + h * (Σfiui / Σfi).
The mean calculated by the step-deviation method will always match the mean calculated by the direct or assumed mean methods.
Pause & Try
Think it through first. Writing and checking your answer is free with an account.
Question
Does the choice of assumed mean 'a' affect the final result of the mean?
Sign in to see the answer
Write your own answer and compare it with ours. It’s free.
Sign inNew here? Sign up freeNCERT reference: chapter PDF pages 6, 7, 9.
05 · Explore
Mode of Grouped Data
The mode is the value that occurs most frequently. In grouped data, we first identify the modal class.
Unlike ungrouped data where the mode is simply the most frequent observation, in grouped data, we can only identify the 'modal class'—the class interval with the highest frequency. The actual mode is a value within this class.
To calculate the estimated mode, we use a formula that considers the frequencies of the classes immediately before and after the modal class. This accounts for the distribution of data around the peak.
The formula is: Mode = l + [(f1 - f0) / (2f1 - f0 - f2)] * h, where 'l' is the lower limit of the modal class, 'f1' is its frequency, 'f0' is the frequency of the preceding class, 'f2' is the frequency of the succeeding class, and 'h' is the class size.
Identifying the Mode
- 1
Find Max Frequency
Look at the frequency column and find the highest value.
- 2
Identify Modal Class
The class interval corresponding to the maximum frequency is the modal class.
- 3
Apply Formula
Use the lower limit, class size, and surrounding frequencies in the mode formula.
The process of determining the mode for a grouped frequency distribution.
Estimate a mode inside its class
Mode = 20 + [(5 − 2)/(2 × 5 − 2 − 3)] × 10 = 26
For continuous equal-width classes 10–20, 20–30 and 30–40 with frequencies 2, 5 and 3, the modal class is 20–30. Set l = 20, f₁ = 5, f₀ = 2, f₂ = 3 and h = 10. The denominator is 10 − 2 − 3 = 5, so mode = 20 + (3/5) × 10 = 26. The modal class is an interval; the estimated mode is a value. This estimate does not recover the exact most frequent raw observation.
Pause & Try
Think it through first. Writing and checking your answer is free with an account.
Question
In a distribution, if the class 30-40 has the highest frequency of 15, and the classes 20-30 and 40-50 have frequencies 7 and 10 respectively, identify f1, f0, and f2.
Sign in to see the answer
Write your own answer and compare it with ours. It’s free.
Sign inNew here? Sign up freeNCERT reference: chapter PDF pages 13, 14, 15.
