STA301 — Midterm Summary (Lectures 1–22)
📘 Lecture 1 — What is Statistics?
📖 Overview: This introductory lecture defines statistics as a science for drawing conclusions from sample data and explores its nature, meanings, and applications across diverse fields. It establishes the foundational framework of descriptive statistics, probability, and inferential statistics while introducing key concepts about data, variables, measurement scales, and measurement errors.
🗂️ Topics Covered
The lecture covers the definition and nature of statistics as a discipline including descriptive statistics, probability, and inferential statistics. It examines the multiple meanings of the word "statistics" (plural vs. singular usage), characteristics of statistical science, how statistics works, and its importance across various fields. The lecture then introduces fundamental concepts including the meaning of data, observations and variables (quantitative/qualitative, discrete/continuous), measurement scales (nominal, ordinal, interval, ratio), and errors of measurement including biased and random errors.
📝 Lecture Summary
What is Statistics?
Statistics is defined as the science which enables us to draw conclusions about various phenomena based on real data collected on a sample basis. It serves as a tool for data-based research and is also known as Quantitative Analysis. Statistics has extensive applications in virtually every discipline from Agriculture to Zoology (A to Z!), including Anthropology, Astronomy, Biology, Economics, Engineering, Environment, Geology, Genetics, Medicine, Physics, Psychology, Sociology, and many others. Any scientific inquiry where conclusions and decisions are based on real-life data requires statistical techniques. Currently, in developed countries, there is an active movement for Statistical Literacy.
The Nature of This Discipline
The discipline of statistics is structured around three main branches:
Descriptive Statistics — Methods for summarizing and describing data
Probability — The mathematical foundation for dealing with uncertainty
Inferential Statistics — Methods for drawing conclusions about populations from sample data
Meanings of 'Statistics'
The word "Statistics" comes from the Latin word status, meaning a political state, and originally meant information useful to the state (e.g., population sizes and armed forces). The word has now acquired three different meanings:
💡 Why this matters: Understanding which meaning is being used is critical for correctly interpreting statistical information.
First meaning (plural usage): Statistics refers to "numerical facts systematically arranged" — a set of numerical data in respective fields (e.g., statistics of prices, road accidents, crimes, births, educational institutions). Most people use the word data instead in this context.
Second meaning (singular usage): Statistics is defined as a discipline that includes procedures and techniques used to collect, process, and analyze numerical data to make inferences and research decisions in the face of uncertainty. Uncertainty does not imply ignorance but refers to the incompleteness and instability of available data. In this sense, statistics is characterized as a science and is mathematical in character.
Third meaning (plural usage): Statistics are numerical quantities calculated from sample observations; a single quantity so collected is called a statistic. For example, the mean of a sample is a statistic.
Characteristics of the Science of Statistics
The important characteristics of statistics as a discipline are:
- Deals with aggregates or large groups of data, not individual observations
- Deals with aggregates of observations of the same kind rather than isolated figures
- Deals with variability that obscures underlying patterns — no two objects are exactly alike
- Deals with uncertainties — every process involves deficiencies or chance variation, hence probability is needed
- Deals with numerically describable characteristics — either by counts or measurements
- Deals with aggregates subject to random causes (e.g., heights of persons are subject to race, ancestry, age, diet, habits, climate)
- Statistical laws are valid on the average or in the long run — there is no guarantee a law will hold in all cases
- Statistical results might be misleading if insufficient care is exercised or if handled by someone not well-versed in statistics
The Way in Which Statistics Works
The main functions of statistics:
- Assists in summarizing larger sets of data in easily understandable form
- Assists in efficient design of laboratory and field experiments as well as surveys
- Assists in sound and effective planning in any field of inquiry
- Assists in drawing general conclusions and making predictions
Importance of Statistics in Various Fields
Statistical techniques are powerful tools for analyzing numerical data used in almost every branch of learning:
- Modern administrators lean on statistical data for factual decision-making
- Politicians use statistics to lend support and credence to their arguments
- Businessmen, industrialists, and research workers employ statistical methods
- Banks, Insurance companies, and Government have statistics departments
- Social scientists use statistical methods in various areas of socio-economic life — it is said that "a social scientist without an adequate understanding of statistics, is often like the blind man groping in a dark room for a black cat that is not there"
The Meaning of Data
The word "data" is Latin for "those that are given" (singular: datum). Data may be thought of as the results of observation. Data are collected in many aspects of everyday life: statements given to a police officer or physician, correct/incorrect answers on an exam, athletic event results (time to complete a marathon, errors in baseball), and scientific inquiry (positions of artifacts/fossils, interactions between animal colony members, spectral composition of starlight).
Observations and Variables
In statistics, an observation means any sort of numerical recording of information — a physical measurement (height, weight), a classification (heads or tails), or an answer (yes or no).
A variable is a characteristic that varies with an individual or object (e.g., age varies from person to person). A variable can assume a number of values. The given set of all possible values from which the variable takes a value is called its Domain. If the domain contains only one value, the variable is referred to as a constant.
Quantitative and Qualitative Variables
Variables are classified as quantitative or qualitative according to the form of the characteristic:
Quantitative variable: A characteristic that can be expressed numerically (e.g., age, weight, income, number of children)
Qualitative variable: A characteristic that is non-numerical (e.g., education, sex, eye-colour, quality, intelligence, poverty, satisfaction). A qualitative characteristic is also called an attribute. Individuals or objects with such characteristics can be counted or enumerated after being assigned to mutually exclusive classes or categories.
Discrete and Continuous Variables
A quantitative variable may be classified as discrete or continuous:
Discrete variable: Can take only a discrete set of integers or whole numbers — values are taken by jumps or breaks. Represents count data (e.g., number of persons in a family, number of rooms in a house, number of deaths in an accident, income of an individual)
Continuous variable: Can take any value — fractional or integral — within a given interval; its domain is an interval with all possible values without gaps. Represents measurement data (e.g., age of a person, height of a plant, weight of a commodity, temperature at a place)
A variable is generally denoted by a symbol such as X or Y, and Xᵢ or Xⱼ represents the i-th or j-th value of the variable.
Measurement Scales
By measurement, we mean assigning numbers to observations or objects, and scaling is a process of measuring. The four scales of measurements are:
Nominal Scale: Classification or grouping of observations into mutually exclusive qualitative categories or classes. Numbers may be used to identify categories but carry no numerical significance and there is no particular order (e.g., male/female classified as 1 and 2; rainfall classified as heavy=1, moderate=2, light=3)
Ordinal or Ranking Scale: Includes the characteristics of a nominal scale and additionally has the property of ordering or ranking of measurements. The only relation between pairs of categories is "greater than" or "more preferred" (e.g., student performance rated as excellent, good, fair, poor; ranks indicated by 1, 2, 3, 4)
Interval Scale: A measurement scale possessing a constant interval size (distance) but not a true zero point. Arithmetic operations of addition and subtraction are meaningful, but ratios are not (e.g., temperature in Celsius or Fahrenheit — the difference between 20°C and 30°C equals the difference between 5°C and 15°C, but 40°C is not twice as hot as 20°C)
Ratio Scale: A special kind of interval scale where the scale of measurement has a true zero point as its origin. Used to measure weight, volume, distance, money, etc. The key difference from interval scale is that the zero point is meaningful for ratio scale.
Errors of Measurement
A continuous variable can never be measured with perfect fineness due to habits, practices, methods, and instruments. Measurements are recorded correct to the nearest unit with limited accuracy. Actual or true values are assumed to exist.
🔑 Definition — Error of measurement: The departure from the true value. If observed value is x and true value is x + ε, then the difference (x + ε) – x = ε is the error.
🔑 Definition — Absolute error: The error ε involving the unit of measurement of x
📐 Formula: Relative error = ε/(x + ε) — independent of units of measurement
📐 Formula: Percentage error = Relative error × 100
📌 Example: If a student's weight is recorded as 60 kg (correct to nearest kilogram), the true weight lies between 59.5 kg and 60.5 kg. If recorded as 60.00 kg, the true weight lies between 59.995 kg and 60.005 kg.
An error has both magnitude and direction, and the word "error" in statistics does not mean mistake (which is a chance inaccuracy).
Biased and Random Errors
🔑 Definition — Biased error: When the observed value is consistently and constantly higher or lower than the true value. Also called cumulative or systematic errors.
Biased errors arise from personal limitations of the observer, imperfection in instruments, or other conditions controlling measurements. They are not revealed by repeating measurements and are cumulative in nature — the greater the number of measurements, the greater the magnitude of error. They are more troublesome.
🔑 Definition — Unbiased error: When deviations (excesses and defects) from the true value tend to occur equally often. Also called random errors or accidental errors.
Unbiased errors are revealed when measurements are repeated and tend to cancel out in the long run. They are compensating errors.
⭐ Key Takeaways
Statistics is a science that deals with variability, uncertainty, and aggregates of data — it has three distinct meanings depending on context (plural data, singular discipline, or sample quantities). Understanding the difference between descriptive statistics, probability, and inferential statistics provides the foundational framework for the entire course. Variables must be correctly classified as quantitative/qualitative and discrete/continuous to apply appropriate analytical methods, and measurement scales (nominal, ordinal, interval, ratio) determine what mathematical operations are permissible. Finally, measurement always involves error — both biased (systematic, cumulative) and random (compensating) — which must be understood for proper interpretation of statistical results.
🧠 Quick Revision Questions
- What are the three distinct meanings of the word "statistics" and how can you tell which meaning is being used in a given context?
- How does a discrete variable differ from a continuous variable? Give two examples of each.
- What are the four measurement scales and what distinguishes an interval scale from a ratio scale?
- What is the difference between a biased error and a random error in measurement, and why are biased errors considered more troublesome?
- Explain the difference between absolute error and relative error. If a true weight is 50.3 kg and the measured weight is 50 kg, what are the absolute, relative, and percentage errors?
📘 Lecture 2 — Steps involved in a Statistical Research-Project
📖 Overview: This lecture covers the complete process of conducting a statistical research project, from defining objectives to collecting data and drawing samples. It explains the crucial distinction between primary and secondary data, various data collection methods, the concept of sampling versus census, and the mechanics of simple random sampling using random number tables—foundational knowledge for any statistical inquiry.
🗂️ Topics Covered
The lecture begins with the steps involved in any statistical research: topic significance, objectives, and methodology for data collection (source, sampling, instrument). It then covers collection of data, differentiating between primary and secondary data, and details five methods for collecting primary data: direct personal investigation, indirect investigation, questionnaires, enumerators, and local sources. Sources of secondary data are listed. The concept of population (finite, infinite, hypothetical) is introduced, along with sampled vs. target populations and the sampling frame. Advantages of sampling and the distinction between sampling and non-sampling errors are explained. Finally, non-random sampling (quota sampling) and random sampling (with a focus on simple random sampling) are discussed, concluding with an example of selecting a sample using random numbers.
📝 Lecture Summary
Steps involved in a Statistical Research-Project
A statistical research project begins with defining the topic and significance of the study, followed by stating the objective clearly—exactly what you are trying to find out. The methodology for data-collection includes identifying the source of your data (the statistical population), the sampling methodology, and the instrument for collecting data. The most important part of statistical work is often the collection of data. Data can be collected by a CENSUS (complete enumeration of the whole field), which is often too costly and time-consuming, or by a SAMPLE (partial enumeration), which saves much time and money.
Collection of Data
PRIMARY DATA are data that have been originally collected (raw data) and have not undergone any sort of statistical treatment. SECONDARY DATA are data that have undergone any sort of treatment by statistical methods at least once—collected, classified, tabulated, or presented for a certain purpose.
Five methods are employed to collect primary data:
Direct Personal Investigation: The investigator collects information personally from the individuals concerned. The information is generally accurate and complete, but this method can be very costly and time-consuming for vast areas and is subject to the personal bias of the investigator.
Indirect Investigation: Third parties or witnesses are interviewed when direct sources do not exist or informants hesitate to respond. This method is useful for complex information or reluctant informants.
Collection through Questionnaires: A questionnaire is an inquiry form with pertinent questions sent by mail. This method is cheap and expeditious for extensive inquiries, but the main difficulty is non-response—respondents may not fill or return the questionnaire, or may return it incomplete. Despite drawbacks, it is the standard method for routine business and administrative inquiries. Questions should be few, brief, simple, clearly worded, and not offensive.
Collection through Enumerators: Trained enumerators assist informants in making entries correctly. This method gives the most reliable information and is considered the best method for large-scale governmental inquiries, though it is generally too costly for private individuals.
Collection through Local Sources: Agents or local correspondents collect and send information using their own judgment. This method is cheap and expeditious but gives only estimates.
Secondary data may be obtained from official sources (e.g., Statistical Division), semi-official sources (e.g., State Bank of Pakistan), trade associations, technical journals, and research organizations.
Population
A statistical population is the collection of every member of a group possessing the same basic and defined characteristic, but varying in amount or quality from one member to another.
Examples:
- Finite population: IQ's of all children in a school.
- Infinite population: Barometric pressure (indefinitely large number of points on earth's surface).
- Hypothetical population: The aggregate of all conceivable ways a specified event can happen, e.g., all possible outcomes from the throw of a die.
🔑 Definition — Sampled population: The population from which a sample is actually chosen. 🔑 Definition — Target population: The population about which information is sought.
💡 Why this matters: A clear distinction between sampled and target populations is critical for the validity of a study's conclusions.
SAMPLING FRAME: A sampling frame is a complete list of all the elements in the population. It should be free from defects: it should not contain inaccurate elements, should be complete, free from duplication, and up-to-date.
Advantages of Sampling
- Savings in time and money. Although cost per unit is greater in sampling, total cost is less because the sample is smaller.
- More detailed information may be obtained from each sample unit.
- Possibility of follow-up for queries and omissions.
- Sampling is the only feasible possibility where tests to destruction are undertaken or where the population is effectively infinite.
Sampling & Non-Sampling Errors
SAMPLING ERROR is the difference between the estimate derived from the sample (the statistic) and the true population value (the parameter). For example, sampling error = ( \bar{X} - \mu ). It arises because a sample cannot exactly represent the population, even if drawn correctly.
NON-SAMPLING ERROR are errors not attributable to sampling but arising in data collection, even during a complete count. Main sources include: defects in the sampling frame, faulty reporting due to personal preferences, negligence of investigators, and non-response to mail questionnaires.
Non-response can be:
- Partial non-response: The respondent refuses to answer some questions.
- Total non-response: The respondent refuses to answer any questions.
Non-sampling errors can be reduced by following up non-response, proper training of investigators, and correct manipulation of collected information. Providing information about the survey's purpose and sending a pre-paid, addressed envelope can stimulate response. A well-defined cut-off date should be established for data collection.
Nonrandom Sampling
Nonrandom sampling (also called purposive sampling) involves drawing population units into the sample using personal judgment.
QUOTA SAMPLING: In this type, selection is not dictated by chance; the interviewer is restricted by quota controls but chooses actual sample units. For example, an interviewer may be told to interview ten married women between thirty and forty years of age living in town X.
Advantages of Quota Sampling:
- No need to construct a frame.
- Very quick investigation.
- Cost reduction.
Disadvantages of Quota Sampling:
- Subjective method (choice between objectivity and convenience).
- Cannot evaluate sampling error objectively (no probability basis).
- Bias creeps in as the interviewer is free to select individuals within quotas.
- Falsification of returns is more dangerous.
Random Sampling
By random sampling, we mean sampling done by adopting the lottery method. The theory of statistical sampling rests on the assumption of random selection.
Types of Random Sampling:
- Simple Random Sampling
- Stratified Random Sampling
- Systematic Sampling
- Cluster Sampling
- Multi-stage Sampling
Simple Random Sampling
In simple random sampling, the chance of any one element of the parent population being included in the sample is the same as for any other element. Haphazard selection is not equivalent to simple random sampling; humans are poor random selectors.
A much more convenient alternative to the lottery method is the use of RANDOM NUMBERS TABLES. A random number table is a page full of digits from 0 to 9 printed in a totally random manner. Computers can also generate random numbers.
📌 Example: Selecting a sample of 10 students from a population of 1000 college students to compare sample mean age with population mean age.
First, allocate sampling numbers to each student using cumulative frequencies:
| Age (X) | No. of Students (f) | Cumulative Frequency (cf) | Sampling Numbers |
|---|---|---|---|
| 13 | 6 | 6 | 000 – 005 |
| 14 | 61 | 67 | 006 – 066 |
| 15 | 270 | 337 | 067 – 336 |
| 16 | 491 | 828 | 337 – 827 |
| 17 | 153 | 981 | 828 – 980 |
| 18 | 15 | 996 | 981 – 995 |
| 19 | 4 | 1000 | 996 – 999 |
The first student gets 000, the 1000th student gets 999 (three-digit numbers for simplicity).
Select 10 random numbers from the table (e.g., by closing eyes and pointing). Suppose the selected random numbers are: 041, 103, 374, 171, 508, 652, 880, 066, 715, 471. The corresponding ages are found by checking which class each number falls into:
- 041 falls in the 14-year class (cf 6 to 67) → age 14
- 103 falls in the 15-year class (cf 67 to 337) → age 15
- 374 falls in the 16-year class (cf 337 to 828) → age 16
- and so on.
The sample ages are: 14, 15, 16, 15, 16, 16, 17, 15, 16, 16.
📐 Formula for Population Mean: ( \mu = \frac{\sum fx}{\sum f} = \frac{15785}{1000} = 15.785 ) years 📐 Formula for Sample Mean: ( \bar{X} = \frac{\sum X}{n} = \frac{156}{10} = 15.6 ) years
Sampling Error = ( \bar{X} - \mu = 15.6 - 15.785 = -0.185 ) years
💡 Why this matters: The small sampling error demonstrates that random sampling produces a sample that is a good representative of the population.
Other Types of Random Sampling
- Stratified sampling: Used if the population is heterogeneous.
- Systematic sampling: Practically more convenient than simple random sampling.
- Cluster sampling: Used when sampling units exist in natural clusters.
- Multi-stage sampling
All these are forms of PROBABILITY sampling—each sampling unit has a known (but not necessarily equal) probability of being selected. Because of this, precision and reliability of estimates can be calculated OBJECTIVELY.
⭐ Key Takeaways
The lecture emphasizes that the most critical part of statistical research is proper data collection, distinguishing between primary (raw, untreated) and secondary (previously treated) data. Sampling is often preferred over a complete census due to cost and time savings, with the key advantage being that a properly drawn random sample can closely represent the population. The fundamental distinction between sampling error (unavoidable difference between sample statistic and population parameter) and non-sampling error (avoidable errors from faulty framing, non-response, or bias) is essential. Simple random sampling, using tools like random number tables, provides the theoretical foundation for all probability sampling methods by ensuring each element has an equal chance of selection.
🧠 Quick Revision Questions
- What is the key difference between primary data and secondary data?
- List three methods for collecting primary data and briefly state one advantage and one disadvantage of each.
- What is the difference between a sampling error and a non-sampling error? Provide one example of each.
- Explain the steps involved in drawing a simple random sample from a finite population using a random number table.
- What are two major disadvantages of quota sampling compared to simple random sampling?
📘 Lecture 3 — Tabulation, Bar Charts & Pie Charts
📖 Overview: This lecture introduces techniques for summarizing and describing qualitative data, focusing on univariate and bivariate frequency tables and their graphical representations. It covers pie charts, simple bar charts, component bar charts, and multiple bar charts, explaining how to construct and interpret each for effective data communication.
🗂️ Topics Covered
The lecture begins with a tree diagram overview of data summarization techniques for qualitative and quantitative data. It then covers univariate frequency tables for qualitative data using an example of student schooling medium, followed by pie chart construction. Simple bar charts are introduced with a turnover example, then bivariate frequency tables for two qualitative variables, component bar charts for displaying totals and their components, and finally multiple bar charts for comparing different datasets like imports and exports.
📝 Lecture Summary
Tabulation
The lecture begins by introducing the concept of tabulation for summarizing data. The first step in handling qualitative data is to count occurrences of each category, creating a frequency table. Using the example of 1200 first-year students surveyed about their schooling medium (Urdu or English):
- 719 students came from Urdu medium schools
- 481 students came from English medium schools
The technical term for these counts is frequency — meaning "how frequently something happens."
💡 Why this matters: Raw frequencies alone are not as useful as proportions or percentages. Dividing cell frequencies by total frequency and multiplying by 100 gives percentages:
| Medium of Institution | f | % |
|---|---|---|
| Urdu | 719 | 59.9% ≈ 60% |
| English | 481 | 40.1% ≈ 40% |
| Total | 1200 | 100% |
This is an example of a univariate frequency table pertaining to qualitative data.
Simple Bar Chart
A simple bar chart consists of horizontal or vertical bars of equal width and lengths proportional to values they represent. The basis of comparison is one-dimensional, so bar widths have no mathematical significance but are used for visual appeal.
Example: Turnover of a company for 5 years:
| Years | 1965 | 1966 | 1967 | 1968 | 1969 |
|---|---|---|---|---|---|
| Turnover (Rupees) | 35,000 | 42,000 | 43,500 | 48,000 | 48,500 |
To construct: Take years along x-axis, construct a scale for turnover along y-axis, and draw vertical bars of equal width and different heights according to turnover figures.
📌 Rule: When values do not relate to time, they should be arranged in ascending or descending order before charting.
Pie Chart
A pie chart consists of a circle divided into two or more parts according to the number of distinct categories in the data.
🔑 Definition — Pie Chart Angle Calculation: To determine where to cut the circle, divide the cell frequency by the total frequency and multiply by 360°.
📐 Formula: Angle = (Cell Frequency / Total Frequency) × 360°
Example: For the schooling medium data:
| Medium of Institution | f | Angle |
|---|---|---|
| Urdu | 719 | 215.7° |
| English | 481 | 144.3° |
| Total | 1200 | 360° |
Bivariate Frequency Table
When analyzing two variables simultaneously, we create a bivariate frequency table. Using the student example with both medium of schooling and sex of student:
- The boxhead is the top row of the table
- The stub is the first column of the table
Four categories are counted:
- Male student from Urdu medium
- Female student from Urdu medium
- Male student from English medium
- Female student from English medium
Resulting table:
| Medium | Male | Female | Total |
|---|---|---|---|
| Urdu | 202 | 517 | 719 |
| English | 350 | 131 | 481 |
| Total | 552 | 648 | 1200 |
This is an example of a bivariate frequency table pertaining to two qualitative variables.
Component Bar Chart
A component bar chart (also called subdivided bar chart) is used when information is available about totals and their components.
In this chart, each bar is divided into parts. Using the student data:
- Lower part of each bar represents students from English medium schools
- Upper part of each bar represents students from Urdu medium schools
The advantage is being able to ascertain the situation of both variables at a glance — comparing male vs. female totals while also comparing English medium proportions within each sex group.
Multiple Bar Chart
A multiple bar chart consists of a set of grouped bars, with lengths proportionate to variable values, each shaded or colored differently for identification.
Example: Imports and Exports of Pakistan (1970-71 to 1974-75):
| Years | Imports (Crores Rs.) | Exports (Crores Rs.) |
|---|---|---|
| 1970-71 | 370 | 200 |
| 1971-72 | 350 | 337 |
| 1972-73 | 840 | 855 |
| 1973-74 | 1438 | 1016 |
| 1974-75 | 2092 | 1029 |
This is a good device for comparing two different kinds of information. If additional variables like production were available, three bars could be grouped together.
Key Difference Between Component and Multiple Bar Charts:
- Component bar chart: Used when information is available about totals and their components (parts add up to the total)
- Multiple bar chart: Used when values do not add up to a total (e.g., imports and exports are separate measurements, not components of a single total)
⭐ Key Takeaways
The most critical concepts from this lecture are understanding the distinction between univariate and bivariate frequency tables for qualitative data and knowing which graphical representation to use for each situation. Pie charts display proportions of a whole by dividing 360° according to frequencies, while simple bar charts show one variable with bar heights proportional to values. Component bar charts are specifically for displaying totals broken into components that add up to the whole, whereas multiple bar charts compare separate variables that do not sum to a total. Proper tabulation is the foundation — always convert raw frequencies to percentages for meaningful interpretation.
🧠 Quick Revision Questions
- How do you calculate the angle for each sector in a pie chart, and what must the sum of all angles equal?
- What is the key difference between a component bar chart and a multiple bar chart in terms of the relationship between the values being displayed?
- In a bivariate frequency table, what are the terms for the top row and first column respectively?
- Why is converting raw frequencies to percentages more useful than reporting only the frequencies?
- When constructing a simple bar chart for non-time-series data, what ordering rule should be followed?
📘 Lecture 4 — Frequency Distribution of a Continuous Variable
📖 Overview: This lecture covers the construction of frequency distributions specifically for continuous variables, using the example of EPA mileage ratings for 30 cars. It explains the step-by-step process of creating classes with appropriate boundaries, and introduces key graphical representations: histograms, frequency polygons, and frequency curves. Understanding these techniques is essential for visualizing and analyzing continuous data.
🗂️ Topics Covered
The lecture begins with the steps for constructing a frequency distribution for a continuous variable, including identifying the range, determining class intervals and limits, and forming classes. It then explains class boundaries to avoid tallying ambiguity. Next, it covers relative and percentage frequency distributions for comparing datasets. Finally, it describes three graphical methods: the histogram using adjacent rectangles on class boundaries, the frequency polygon using midpoints and including zero-frequency classes to close the figure, and the frequency curve as a smoothed version of the polygon.
📝 Lecture Summary
CONSTRUCTION OF A FREQUENCY DISTRIBUTION
The construction of a frequency distribution for a continuous variable involves a systematic series of steps. First, identify the smallest and largest values in the dataset. Next, compute the range, which is the difference between the largest and smallest values. Then, decide on the number of classes, usually between 10 and 20 for large datasets, but fewer for smaller sets. Divide the range by the number of classes to get the approximate class interval width, rounding it for convenience. Finally, determine the lower and upper class limits, starting from a number slightly less than the smallest value.
🔑 Definition — Range: The difference between the largest value and the smallest value in a dataset.
📐 Formula: Range = Xm – X0 → The highest value minus the lowest value.
📌 Example: In the EPA mileage data, the smallest value is 30.1 and the largest is 44.9. Range = 44.9 – 30.1 = 14.8. The chosen number of classes is 5, giving a class interval of h = 14.8 / 5 = 2.96, which is rounded to 3. The lowest class lower limit is set at 30.0, and successive classes are: 30.0-32.9, 33.0-35.9, 36.0-38.9, 39.0-41.9, and 42.0-44.9. The tally yields frequencies of 2, 4, 14, 8, and 2, respectively.
CLASS BOUNDARIES
Class boundaries are the true limits of a class, designed to prevent ambiguity when an observation falls exactly on a class limit. They are typically taken to one more decimal place than the original data. The difference between the upper and lower class boundary of any class is equal to the class interval h.
🔑 Definition — Class Boundaries: The true class limits of a class, ensuring no observation falls exactly on a boundary.
📌 Example: The class 30.0 – 32.9 has class boundaries of 29.95 – 32.95. The class interval is still h = 32.95 – 29.95 = 3. The boundaries for all classes are: 29.95-32.95, 32.95-35.95, 35.95-38.95, 38.95-41.95, and 41.95-44.95.
💡 Why this matters: Class boundaries eliminate the difficulty of tallying values that equal a class limit, ensuring each observation falls into a unique class.
RELATIVE AND PERCENTAGE FREQUENCY DISTRIBUTIONS
A relative frequency distribution is obtained by dividing each class frequency by the total number of observations. Multiplying each relative frequency by 100 gives the percentage frequency distribution. This transformation allows for comparison between datasets of different sizes.
🔑 Definition — Relative Frequency: The proportion of observations in a class, obtained by dividing the class frequency by the total number of observations.
📐 Formula: Relative Frequency = f / n → Class frequency divided by total observations.
📌 Example: For the class 36.0-38.9, the frequency is 14 and the total is 30. The relative frequency is 14/30 ≈ 0.467, and the percentage frequency is 0.467 * 100 = 46.7%. Comparing two car models, Model A has 6.7% of cars in the 42.0-44.9 group, while Model B has 16% in the same group, indicating a different performance distribution.
HISTOGRAM
A histogram is a graphical representation of a continuous frequency distribution. It consists of a set of adjacent rectangles. The bases of the rectangles are marked off by class boundaries along the X-axis, and their heights are proportional to the frequencies of the respective classes.
🔑 Definition — Histogram: A graph of adjacent rectangles representing a frequency distribution, with bases on class boundaries and heights proportional to class frequencies.
📌 Example: For the EPA mileage data, a histogram is drawn with the X-axis representing miles per gallon from 29.95 to 44.95 using class boundaries. The Y-axis shows the number of cars (frequency). The first rectangle (29.95-32.95) has a height of 2, the second (32.95-35.95) has a height of 4, the third (35.95-38.95) has a height of 14, the fourth (38.95-41.95) has a height of 8, and the fifth (41.95-44.95) has a height of 2.
💡 Why this matters: A histogram gives an immediate visual indication of the overall pattern of the frequency distribution, showing where data are concentrated.
FREQUENCY POLYGON
A frequency polygon is constructed by plotting class frequencies against the mid-points of the classes. These points are then connected by straight line segments. To create a closed figure, an additional class is added at the beginning and end of the distribution, each with a frequency of zero.
🔑 Definition — Frequency Polygon: A graph obtained by plotting class frequencies against class mid-points and connecting the points with straight lines, typically closed by adding zero-frequency classes.
📐 Formula: Mid-point of a class = (Lower Class Boundary + Upper Class Boundary) / 2
📌 Example: For the EPA data, the mid-point of the first class (29.95-32.95) is (29.95+32.95)/2 = 31.45. The complete table includes two extra classes with zero frequency: 26.95-29.95 (mid-point 28.45, frequency 0) and 44.95-47.95 (mid-point 46.45, frequency 0). Points are plotted at (28.45, 0), (31.45, 2), (34.45, 4), (37.45, 14), (40.45, 8), (43.45, 2), and (46.45, 0), then connected with straight lines to form a closed polygon.
FREQUENCY CURVE
A frequency curve is a smooth curve that approximates the shape of the frequency polygon. It is obtained by smoothing out the jagged edges of the polygon, providing a more generalized and often more realistic representation of the underlying distribution.
🔑 Definition — Frequency Curve: A smoothed version of the frequency polygon, representing the general pattern of the data.
📌 Example: Starting from the jagged frequency polygon of the EPA mileage data, a smooth curved line is drawn that passes through the general shape of the polygon, eliminating the sharp angles at each plotted point.
⭐ Key Takeaways
The construction of a frequency distribution for a continuous variable requires calculating the range, selecting an appropriate number of classes, and determining class limits that are convenient. Class boundaries, taken to one more decimal than the data, prevent ambiguity in tallying. Relative and percentage frequency distributions are crucial for comparing datasets of different sizes. Histograms use adjacent rectangles on class boundaries to display frequencies, while frequency polygons use class mid-points and must be closed with zero-frequency classes to be a true polygon. A frequency curve provides a smoothed, generalized picture of the data’s distribution pattern.
🧠 Quick Revision Questions
- What are the first two steps in constructing a frequency distribution for continuous data?
- Why are class boundaries used instead of class limits when constructing a histogram?
- How do you calculate the relative frequency of a class?
- What is the key difference between a histogram and a frequency polygon?
- Why is it necessary to add two extra classes with zero frequency when constructing a frequency polygon?
📘 Lecture 5 — FREQUENCY POLYGON, FREQUENCY CURVE, AND CUMULATIVE FREQUENCY DISTRIBUTION
📖 Overview: This lecture continues from the previous one, covering various types of frequency curves encountered in practice, including symmetrical, skewed, and U-shaped distributions. It also introduces the cumulative frequency distribution and the cumulative frequency polygon (ogive) for continuous variables, which are essential tools for understanding data patterns and analyzing real-world phenomena.
🗂️ Topics Covered
The lecture begins with the frequency polygon and its smoothed version, the frequency curve, explaining how histograms approximate smooth curves with smaller class intervals. It then discusses various types of frequency curves: symmetrical, moderately skewed (positive and negative), extremely skewed (J-shaped and reverse J-shaped), and U-shaped distributions, noting that moderately skewed distributions are most common in natural and social phenomena. The lecture also covers discrete frequency distributions with an example of a positively skewed discrete distribution. Finally, it introduces cumulative frequency distributions for continuous variables, focusing on the “less than” type, and the construction of cumulative frequency polygons or ogives, including how to close them on both sides.
📝 Lecture Summary
FREQUENCY POLYGON:
A frequency polygon is obtained by plotting class frequencies against the mid-points of the classes and connecting the points with straight line segments. For example, from the EPA mileage ratings data:
| Class Boundaries | Mid-Point (X) | Frequency (f) |
|---|---|---|
| 26.95 – 29.95 | 28.45 | |
| 29.95 – 32.95 | 31.45 | 2 |
| 32.95 – 35.95 | 34.45 | 4 |
| 35.95 – 38.95 | 37.45 | 14 |
| 38.95 – 41.95 | 40.45 | 8 |
| 41.95 – 44.95 | 43.45 | 2 |
| 44.95 – 47.95 | 46.45 |
The resulting polygon shows the distribution of mileage.
When the frequency polygon is smoothed, we obtain a frequency curve. In the example, the dotted line represents the frequency curve. It does not need to touch all points; it is drawn by free-hand to display the overall pattern of the distribution. The frequency curve is a theoretical concept. If the class interval of a histogram is made very small and the number of classes very large, the rectangles become narrow, and the histogram approaches a smooth curve. Despite being theoretical, frequency curves are useful in analyzing real-world problems because close approximations to theoretical curves often occur in practice.
VARIOUS TYPES OF FREQUENCY CURVES
- The symmetrical frequency curve: If a vertical mirror is placed in the centre, the left side is a mirror image of the right side.
- The moderately skewed frequency curve: Positively skewed has a longer right tail than left; negatively skewed has a longer left tail than right.
- The extremely skewed frequency curve: An extremely negatively skewed curve occurs when the maximum frequency is at the end of the frequency table. For example, death rates of adult males:
| Age Group | No. of deaths per thousand |
|---|---|
| 20 – 29 | 2.1 |
| 30 – 39 | 4.3 |
| 40 – 49 | 5.7 |
| 50 – 59 | 8.9 |
| 60 – 69 | 12.4 |
| 70 – 79 | 16.7 |
This results in a J-shaped distribution. The extremely positively skewed distribution is the reverse J-shaped distribution.
- The U-shaped distribution: A less common type, e.g., death rates for all age groups.
💡 Why this matters: The moderately skewed frequency distribution is the MOST frequently encountered, occurring in thousands of natural and social phenomena like weights, heights, and marks of children.
Various types of discrete frequency distributions also exist, e.g., a positively skewed discrete distribution shown in the lecture.
CUMULATIVE FREQUENCY DISTRIBUTION
As with discrete variables, starting from the first frequency and adding successively gives cumulative frequencies.
From the EPA example:
| Class Boundaries | Frequency | Cumulative Frequency |
|---|---|---|
| 29.95 – 32.95 | 2 | 2 |
| 32.95 – 35.95 | 4 | 2+4 = 6 |
| 35.95 – 38.95 | 14 | 6+14 = 20 |
| 38.95 – 41.95 | 8 | 20+8 = 28 |
| 41.95 – 44.95 | 2 | 28+2 = 30 |
Each cumulative frequency represents the total frequency from the lower class boundary of the lowest class to the UPPER class boundary of that class. For example, the number of cars with mileage less than 35.95 is 6, and less than 41.95 is 28. This is called a “less than” type cumulative frequency distribution.
CUMULATIVE FREQUENCY POLYGON or OGIVE
A “less than” type ogive is obtained by plotting upper class boundaries on the X-axis and cumulative frequencies on the Y-axis, then joining the points with straight line segments. The ogive touches the X-axis on the left by adding a class with zero frequency at the beginning:
| Class Boundaries | Frequency | Cumulative Frequency |
|---|---|---|
| 26.95 – 29.95 | 0 | 0 |
| 29.95 – 32.95 | 2 | 2 |
| 32.95 – 35.95 | 4 | 6 |
| 35.95 – 38.95 | 14 | 20 |
| 38.95 – 41.95 | 8 | 28 |
| 41.95 – 44.95 | 2 | 30 |
To close on the right-hand side, connect the last point to the X-axis with a vertical line.
EXAMPLE:
For 40 pizza products, cost of a slice (S Cost) in dollars, with minimum 0.52 and maximum 1.90.
Range = 1.90 – 0.52 = 1.38
Desired number of classes = 8
Class interval h = 1.38 / 8 = 0.1725 ≈ 0.20
Lower limit of first class = 0.51
Class limits and boundaries:
| Class Limits | Class Boundaries |
|---|---|
| 0.51 – 0.70 | 0.505 – 0.705 |
| 0.71 – 0.90 | 0.705 – 0.905 |
| 0.91 – 1.10 | 0.905 – 1.105 |
| 1.11 – 1.30 | 1.105 – 1.305 |
| 1.31 – 1.50 | 1.305 – 1.505 |
| 1.51 – 1.70 | 1.505 – 1.705 |
| 1.71 – 1.90 | 1.705 – 1.905 |
Tallying data into classes yields a frequency distribution, and constructing its histogram helps determine symmetry or skewness.
⭐ Key Takeaways
A frequency polygon connects mid-point frequencies, and smoothing it yields a theoretical frequency curve that shows overall distribution patterns. The most common real-world distributions are moderately skewed, while symmetrical, extremely skewed (J-shaped), and U-shaped curves also occur. Cumulative frequency distributions (less than type) help count observations below class boundaries, and their graph (ogive) is constructed by plotting upper class boundaries against cumulative frequencies, closed by adding a zero-frequency class. Understanding these curve shapes is critical for analyzing and summarizing continuous and discrete data in statistics.
🧠 Quick Revision Questions
- How is a frequency polygon constructed, and what does it represent?
- What is the difference between a positively skewed and a negatively skewed frequency curve?
- Give an example of a real-world phenomenon that yields a J-shaped distribution.
- How is a "less than" type cumulative frequency distribution interpreted?
- How do you close an ogive on both the left and right sides?
📘 Lecture 6 — Stem-and-Leaf Display and Mode
📖 Overview: This lecture introduces the stem-and-leaf display, a technique invented by John Tukey (1977) that preserves individual data values while sorting and displaying data. It then transitions to measures of central tendency, focusing on the mode — the most frequently occurring value — and explains how to compute it for raw data, discrete frequency distributions, and grouped (continuous) data.
🗂️ Topics Covered
The lecture covers the construction and interpretation of stem-and-leaf displays, converting them into frequency distributions and histograms, and noting their shape similarity. It then introduces the concept of central tendency as a method for describing variable data, distinguishes between measures of central tendency and dispersion, and provides a detailed explanation of the mode, including its calculation for raw data (using a dot plot), discrete frequency distributions, and continuous grouped data using a specific formula.
📝 Lecture Summary
Stem-and-Leaf Display
A stem-and-leaf display is a technique for simultaneously sorting and displaying data, overcoming the disadvantage of a frequency table where the identity of individual observations is lost in the grouping process. Each number in the data set is divided into two parts: a stem, which is the leading digit(s) used for sorting, and a leaf, which is the trailing digit(s) shown in the display. A vertical line separates the leaf from the stem. For example, the number 243 can be split as "24 | 3" or "2 | 43". It is common practice to arrange the trailing digits in each row from smallest to highest, which also produces an ordered array of the data.
🔑 Definition — Stem: The leading digit(s) of each number in a data set, used for sorting in a stem-and-leaf display. 🔑 Definition — Leaf: The trailing digit(s) of each number in a data set, shown in the display of a stem-and-leaf display.
📌 Example (Ages): The ages of 30 patients are: 48, 31, 54, 37, 18, 64, 61, 43, 40, 71, 51, 12, 52, 65, 53, 42, 39, 62, 74, 48, 29, 67, 30, 49, 68, 35, 57, 26, 27, 58. Using the first digit as the stem and the second as the leaf, and ordering the leaves, we get the display:
Stem (Leading Digit) | Leaf (Trailing Digit)
1 | 2 8
2 | 6 7 9
3 | 0 1 5 7 9
4 | 0 2 3 8 8 9
5 | 1 2 3 4 7 8
6 | 1 2 4 5 7 8
7 | 1 4
This display can be easily converted into a frequency distribution (e.g., class 10-19 has frequency 2, class 20-29 has frequency 3, etc.). If this frequency distribution is converted into a histogram and rotated 90 degrees, its shape is exactly like the shape of the stem-and-leaf display.
📌 Example (Death Rates): Construct a stem-and-leaf display for mean annual death rates per thousand: 7.5, 8.2, 7.2, 8.9, etc. Using the decimal part as the leaf and the rest of the digits as the stem, we get an ordered display:
Stem | Leaf
3 | 9
4 | 6 6
5 | 0 4 5
6 | 0 0 2 2 5 5 6 6 6 8 9
7 | 1 3 3 3 4 4 5 6 7 7 8 8 8 9
8 | 1 1 2 4 6 6 7 7 8 8 8 9 9
9 | 0 1 2 3 3 3 3 4 4 4 6 7 7 9 9
10 | 0 1 1 2 3 3 4 4 6 6 7 8 8 9 9
11 | 1 4 4 6 6 9
12 | 0 1 4 5 6 8 8
14 | 0
💡 Why this matters: The stem-and-leaf display provides a quick visual summary of a dataset's distribution while retaining the original data values, which is a key advantage over histograms.
Description of Variable Data and Measures of Central Tendency
In any statistical enquiry, a concise numerical description of variable data is needed. Measures of central tendency (averages) enable us to measure the central tendency of variable data, indicating the location or general position of the distribution on the X-axis. Measures of dispersion enable us to measure its variability. An average is a single value representing a set of data, more or less central to the values which cluster around it. The most common types are the arithmetic mean, geometric mean, harmonic mean, median, and mode. The mathematical averages (arithmetic, geometric, harmonic) indicate the magnitude of values; the median indicates the middle position; the mode indicates the most frequent value.
The Mode
The mode is defined as that value which occurs most frequently in a set of data, indicating the most common result.
📌 Example (Raw Data): Marks of eight students: 2, 7, 9, 5, 8, 9, 10, 9. The most common mark is 9. Hence, Mode = 9.
Mode in Case of Raw Data (Continuous Variable)
For raw data of a continuous variable not grouped into a frequency distribution, the mode is obtained by counting the number of times each value occurs. A dot plot can be used to visualize this, where dots are placed above one another for repeating values, forming piles.
📌 Example (R&D Percentages): Data on percentages of revenues spent on R&D by 49 companies. The dot plot shows value 6.9 occurs 3 times, while all others occur once or twice. Hence, the modal value is 6.9. The dot plot also shows almost all percentages are between 6% and 12%, with most between 7% and 9%.
Mode in Case of a Discrete Frequency Distribution
In a discrete frequency distribution, the mode is the value with the highest frequency.
📌 Example (Airline Passengers): Number of passengers on 50 flights of a 40-seater plane:
| No. of Passengers (X) | No. of Flights (f) |
|---|---|
| 28 | 1 |
| 33 | 1 |
| 34 | 2 |
| 35 | 3 |
| 36 | 5 |
| 37 | 7 |
| 38 | 10 |
| 39 | 13 |
| 40 | 8 |
| Total | 50 |
| Highest frequency (fm) = 13, occurs against X value 39. Hence, Mode = 39. The company should be satisfied that a 40-seater is the correct size. |
Mode in Case of the Frequency Distribution of a Continuous Variable
For grouped data, the modal class is the one with the highest frequency. To find the exact mode within this class, the following formula is used:
📐 Formula: Mode (X̂) = l + [ (fm - f1) / {(fm - f1) + (fm - f2)} ] × h
Where:
- l = lower class boundary of the modal class
- fm = frequency of the modal class
- f1 = frequency of the class preceding the modal class
- f2 = frequency of the class following the modal class
- h = length of class interval of the modal class
📌 Example (EPA Mileage Ratings):
| Mileage Rating | Class Boundaries | No. of Cars |
|---|---|---|
| 30.0 – 32.9 | 29.95 – 32.95 | 2 |
| 33.0 – 35.9 | 32.95 – 35.95 | 4 = f1 |
| 36.0 – 38.9 | 35.95 – 38.95 | 14 = fm |
| 39.0 – 41.9 | 38.95 – 41.95 | 8 = f2 |
| 42.0 – 44.9 | 41.95 – 44.95 | 2 |
The third class is the modal class. Using the formula: X̂ = 35.95 + [ (14 - 4) / {(14 - 4) + (14 - 8)} ] × 3 X̂ = 35.95 + [10 / (10 + 6)] × 3 X̂ = 35.95 + (10/16) × 3 X̂ = 35.95 + 1.875 X̂ = 37.825
Locating this value on the X-axis of the frequency curve shows the mode exists in the middle of the data values, confirming it is a measure of central tendency.
⭐ Key Takeaways
The stem-and-leaf display is a powerful tool for exploratory data analysis because it sorts data and shows its distribution without losing individual values. Converting this display into a frequency table and histogram reveals that the shape of the distribution is preserved. The mode is a straightforward measure of central tendency representing the most frequent value, but its calculation method depends on data type: by counting for raw data, by identifying the highest frequency for discrete distributions, and via a specific formula for continuous grouped data. The formula for the mode in grouped data interpolates within the modal class using the frequencies of the classes before and after it. Understanding the mode's calculation is essential for describing the most typical or common outcome in a dataset.
🧠 Quick Revision Questions
- What is the main advantage of a stem-and-leaf display over a frequency distribution?
- In the stem-and-leaf display for the age data (12 to 74), how many observations (leaves) are in the row for stem "5"?
- How is the modal class identified in a frequency distribution of a continuous variable?
- In the formula for the mode of grouped data (Mode = l + [ (fm - f1) / {(fm - f1) + (fm - f2)} ] × h), what does "f1" represent?
- For the EPA mileage ratings example, if the frequency of the modal class (fm) was 14, the class before it (f1) was 4, and the class after it (f2) was 8, what is the value of the term (fm - f2)?
📘 Lecture 7 — Measures of Central Tendency (Mode, Arithmetic Mean, Median)
📖 Overview: This lecture continues the discussion of measures of central tendency, focusing on the mode, its desirable properties, and when it is most useful. It then introduces the arithmetic mean as the most widely used average, covering its calculation for both ungrouped and grouped data, along with the concept of grouping error. Finally, the lecture discusses the weighted arithmetic mean and introduces the median as a measure of central tendency, particularly useful when data contains extreme values.
🗂️ Topics Covered
The lecture begins by detailing the desirable properties and appropriate uses of the mode, including the concept of a bi-modal distribution. It then formally introduces the arithmetic mean, demonstrating its calculation for ungrouped data and grouped frequency distributions using class-marks, while also explaining the concept of grouping error. The lecture covers the desirable properties and limitations of the arithmetic mean before moving to the weighted mean to handle data with unequal importance. The session concludes with the definition and calculation of the median for raw data and discrete frequency distributions, highlighting its value when extreme values distort the arithmetic mean.
📝 Lecture Summary
DESIRABLE PROPERTIES OF THE MODE
The mode is a valuable measure of central tendency, particularly easy to understand and ascertain for discrete frequency distributions. A key advantage is that it is not affected by a few very high or low values. The mode is most useful in practical business contexts, such as inventory management, where knowing the most common size or value is critical. The manager of a men's clothing store, for instance, needs to know the modal hat size to stock in greatest quantity.
🔑 Definition — Mode: The value that occurs most frequently in a data set.
💡 Why this matters: The mode is crucial for business decisions like stock control, where the most frequent item is the most important to have in large supply.
It is important to note that sometimes a frequency distribution contains two modes, in which case it is called a bi-modal distribution.
THE ARITHMETIC MEAN
The arithmetic mean is the statistician’s term for the average and is the most widely used average. It is the value that is numerically MOST representative of the whole series. Its formal definition is: “The arithmetic mean or simply the mean is a value obtained by dividing the sum of all the observations by their number.”
For ungrouped data, the formula is:
📐 Formula: X̄ = ΣX / n → The mean (X̄) equals the sum of all observations (ΣX) divided by the number of observations (n).
📌 Example: A news agent's receipts for a week are: Monday £9.90, Tuesday £7.75, Wednesday £19.50, Thursday £32.75, Friday £63.75, Saturday £75.50, Sunday £50.70. The total is £259.85. The mean sales per day are £259.85 / 7 = £37.12.
For grouped data in a frequency distribution, the observations in each class are assumed to be identical with the class-mark (midpoint, X). The formula becomes:
📐 Formula: X̄ = ΣfX / Σf or X̄ = ΣfX / n → The mean equals the sum of (frequency * midpoint) divided by the total frequency.
📌 Example: For the EPA mileage ratings of 30 cars, the midpoints (X) are computed (e.g., 31.45, 34.45), multiplied by their frequencies (f), and summed (ΣfX = 1135.5). The mean is 1135.5 / 30 = 37.85 miles per gallon.
GROUPING ERROR
Grouping error refers to the error introduced by the assumption that all values in a class are equal to the class midpoint. In reality, this is rarely true, so the mean from a frequency distribution is an approximation. However, this error is usually small for the arithmetic mean. In the EPA example, the true mean from raw data was 37.82, while the grouped data gave 37.85, a very slight difference.
DESIRABLE PROPERTIES OF THE ARITHMETIC MEAN:
- Best understood average in statistics.
- Relatively easy to calculate.
- Takes into account every value in the series.
LIMITATION: A few very high or very low values can drag the arithmetic mean towards them, making it unrepresentative.
📌 Example: The number of floors in a city’s buildings: 5, 4, 3, 4, 5, 4, 3, 4, 5, 20, 5, 6, 32, 8, 27. The mean is 9, even though 12 out of 15 buildings have 6 floors or less. The three skyscrapers disproportionately affect the mean.
WEIGHTED MEAN
The weighted mean is used when data values cannot be regarded as having equal weightage. Each value (X_i) is assigned a weight (W_i) according to a suitable criterion.
📐 Formula: X̄_w = ΣW_i X_i / ΣW_i → The weighted mean is the sum of (weight * value) divided by the sum of the weights.
📌 Example: In a high school with 100 Freshmen (15% absent), 80 Sophomores (5% absent), 70 Juniors (10% absent), and 50 Seniors (2% absent), the simple average is incorrectly 8%. Using the number of students as weights gives the correct answer: (100*15 + 80*5 + 70*10 + 50*2) / (100+80+70+50) = 2700/300 = 9%.
MEDIAN
The median is the middle value of a series when the variable values are placed in order of magnitude. It is defined as a “value which divides a set of data into two halves, one half comprising of observations greater than and the other half smaller than it.” The median is a better average than the mean when data contains extreme values.
For ungrouped data with an odd number of observations, it is the middle value. 📌 Example (Odd n): For the building floors data sorted as 3, 3, 4, 4, 4, 4, 5, 5, 5, 5, 6, 8, 20, 27, 32, the median is 5. This is much more representative than the mean of 9.
For ungrouped data with an even number of observations, the median is the arithmetic mean of the two middle values. 📌 Example (Even n): For passengers sorted as 5, 14, 18, 23, 34, 47, the median is (18+23)/2 = 20.5.
For a discrete frequency distribution, the median is found by locating the value corresponding to the (n+1)/2 th item in the cumulative frequency distribution. 📌 Example: In a school with 45 classes, the median class-size is the 23rd class. The cumulative frequency column shows 20 classes have up to 28 pupils and 28 classes have up to 29 pupils. Therefore, the 23rd class is the one with 29 pupils per class.
⭐ Key Takeaways
The mode is the most frequent value, useful for inventory but can be absent or not unique (bi-modal). The arithmetic mean is the most common average, calculated as the sum divided by the number of values, but it is sensitive to extreme outliers. The weighted mean is essential when data points have different levels of importance. The median, as the middle value, is a more robust measure for skewed data or when outliers are present. For grouped data, these measures are approximations due to grouping error, though this error is usually small for the mean.
🧠 Quick Revision Questions
- What is the formal definition of the arithmetic mean, and what is the symbol used to represent it?
- Explain the concept of "grouping error" and why the mean calculated from a frequency distribution is an approximation.
- In what type of situation is the median a more appropriate measure of central tendency than the arithmetic mean? Provide a brief example.
- Describe the difference between calculating the median for a raw data set with an odd number of observations versus one with an even number of observations.
- When is it necessary to use the weighted arithmetic mean instead of the simple arithmetic mean?
📘 Lecture 8 — Median, Empirical Relation, and Quantiles
📖 Overview: This lecture explains how to compute the median for a continuous frequency distribution and open-ended distributions. It introduces the empirical relation between the mean, median, and mode, and then extends the concept of partitioning data into quartiles, deciles, and percentiles. Understanding these measures is essential for describing data distribution and relative standings.
🗂️ Topics Covered
This lecture covers the median formula for continuous frequency distributions and its interpretation, handling open-ended frequency distributions with median, the empirical relation between mean, median, and mode for skewed distributions, the concept of quantiles including quartiles, deciles, and percentiles with their formulas, and the graphic location of quantiles using an ogive.
📝 Lecture Summary
Median in case of a frequency distribution of a continuous variable
The median for a grouped frequency distribution is found using the formula: $$\widetilde{X} = l + \frac{h}{f}\left(\frac{n}{2} - c\right)$$ where l = lower class boundary of the median class, h = class interval size, f = frequency of median class, n = total frequency, and c = cumulative frequency of class preceding median class. The median class is the first class with cumulative frequency exceeding n/2.
📐 Formula: $$\widetilde{X} = l + \frac{h}{f}\left(\frac{n}{2} - c\right)$$ → The median equals the lower boundary plus a proportional part of the class interval based on how far into the class the middle observation lies.
📌 Example: For EPA mileage ratings with n=30, n/2=15. Median class is 36.0-38.9 (boundaries 35.95-38.95). Here l=35.95, h=3, f=14, c=6. Calculating: $$\widetilde{X} = 35.95 + \frac{3}{14}(15-6) = 35.95 + 1.93 = 37.88$$. This means half of cars have mileage less than or equal to 37.88 mpg, half above.
Median in case of an open-ended frequency distribution
An open-ended frequency distribution has classes like "less than 2000" or "5000 and above" without defined boundaries. The median is especially valuable here because unless the median falls in an open-ended class, we don't need to estimate boundaries. If the median falls in an intermediate class, the first or last open-ended class is not involved in computation.
Empirical relation between the mean, median and the mode
The word 'empirical' means 'based on observation.' For a unimodal curve of moderate skewness, the median lies between the mean and the mode. The relationship is: Median - Mode = 2(Mean - Median) and equivalently Mean - Mode = 3(Mean - Median).
📐 Formula: Mode = 3 Median – 2 Mean → In moderately skewed distributions, if you know any two measures, you can estimate the third. 💡 Why this matters: This helps verify consistency of calculated measures or estimate missing ones.
This relation does NOT hold for J-shaped or extremely skewed distributions.
QUARTILES
Quartiles divide the total area under the frequency curve into four equal parts. The formulas follow the same pattern as median but with different numerators:
📌 First Quartile (Q₁): $$Q_1 = l + \frac{h}{f}\left(\frac{n}{4} - c\right)$$ — Divides lower 25% of data 📌 Second Quartile (Q₂): $$Q_2 = l + \frac{h}{f}\left(\frac{2n}{4} - c\right) = l + \frac{h}{f}\left(\frac{n}{2} - c\right)$$ — Same as median 📌 Third Quartile (Q₃): $$Q_3 = l + \frac{h}{f}\left(\frac{3n}{4} - c\right)$$ — Divides upper 25% of data
DECILES & PERCENTILES
Deciles divide the distribution into 10 equal parts, and percentiles divide into 100 equal parts.
📐 Formula for first decile: $$D_1 = l + \frac{h}{f}\left(\frac{n}{10} - c\right)$$ 📐 Formula for first percentile: $$P_1 = l + \frac{h}{f}\left(\frac{n}{100} - c\right)$$
The 5th decile (D₅) equals the median, the 25th percentile (P₂₅) equals Q₁, the 75th percentile (P₇₅) equals Q₃, and the 40th percentile (P₄₀) equals the 4th decile (D₄).
All these measures — median, quartiles, deciles, and percentiles — are collectively called quantiles (or fractiles).
📌 Example: If Company A's yearly sales are at the 90th percentile, this means 90% of all companies have sales less than Company A's, and only 10% have sales exceeding Company A's. Percentile rankings are of practical value only for large data sets.
Graphic location of quantiles
The ogive (cumulative frequency polygon) allows convenient graphical determination of quantiles. To find the median graphically:
- Compute n/2 (here 30/2 = 15)
- Locate 15 on the y-axis
- Draw a horizontal line from y=15 to the ogive curve
- From the intersection point, drop a vertical line to the x-axis
- The x-value where the vertical line touches is the median
Similarly, for Q₁ draw horizontal line at n/4, and for Q₃ draw at 3n/4. For deciles, draw at n/10, 2n/10, 3n/10, etc. For percentiles, draw at n/100, 2n/100, 3n/100, etc.
⭐ Key Takeaways
The median formula for grouped data is the foundation for computing all quantiles, with the only change being the numerator (n/2 for median, n/4 for Q₁, etc.). The empirical relation Mode = 3Median – 2Mean applies only to moderately skewed unimodal distributions, not to symmetrical or extremely skewed ones. The second quartile, 5th decile, and 50th percentile are all equivalent to the median. Percentile ranking is particularly useful for describing relative standing in large datasets. The ogive provides a quick graphical method to locate any quantile by drawing horizontal lines at appropriate cumulative frequency values.
🧠 Quick Revision Questions
- What are the symbols l, h, f, n, and c in the median formula for grouped data, and what does each represent?
- Why is the median preferred over the mean for open-ended frequency distributions?
- If the mean of a moderately skewed distribution is 50 and the median is 48, what is the approximate mode?
- Which decile is equivalent to the median, and which percentile is equivalent to the first quartile?
- How would you find the 60th percentile graphically using an ogive?
📘 Lecture 9 — Geometric Mean, Harmonic Mean, Relation Between the Arithmetic, Geometric and Harmonic Means, Some Other Measures of Central Tendency
📖 Overview: This lecture introduces two additional measures of central tendency beyond the arithmetic mean: the geometric mean and the harmonic mean. It explains their formulas for both raw and grouped data, provides detailed examples of their application, and clarifies the specific situations where each measure is most appropriate. The lecture also establishes the mathematical relationship between the arithmetic, geometric, and harmonic means and briefly discusses other measures like the mid-range and mid-quartile range.
🗂️ Topics Covered
This lecture covers the geometric mean, its definition, formula, and computation using logarithms for both raw data and grouped frequency distributions, with detailed examples including the EPA mileage data and a firm's turnover growth. It then covers the harmonic mean, its definition and formula for raw and grouped data, illustrated by a car speed example to show its correct application for averaging rates. The lecture ends with the relation between the arithmetic, geometric, and harmonic means, and introduces other measures of central tendency: the mid-range and mid-quartile range.
📝 Lecture Summary
Geometric Mean
The geometric mean, G, of a set of n positive values X₁, X₂,..., Xₙ is defined as the positive nth root of their product. The formula is G = ⁿ√(X₁ × X₂ × ... × Xₙ) where Xᵢ > 0. When n is large, computing the geometric mean becomes laborious because you have to extract the nth root of the product of all values. The calculation is simplified by using logarithms. Taking logarithms to base 10, the formula becomes log G = (1/n)[log X₁ + log X₂ + ... + log Xₙ] = Σlog X / n. Therefore, G = antilog [Σlog X / n].
🔑 Definition — Geometric Mean (G): The positive nth root of the product of n positive values. It is used when averaging relative changes or rates of growth. 📐 Formula: G = antilog [Σlog X / n] (for raw data) → This means the geometric mean is the antilogarithm of the average of the logarithms of the data values. 📌 Example: Find the geometric mean of the numbers: 45, 32, 37, 46, 39, 36, 41, 48, 36. Solution: The product is ⁹√(45×32×37×46×39×36×41×48×36), which is cumbersome to compute directly. Using logarithms, the sum of log X is 14.3870. Then, log G = 14.3870 / 9 = 1.5986. Hence, G = antilog 1.5986 = 39.68.
Geometric Mean for Grouped Data
In case of a frequency distribution having k classes with midpoints X₁, X₂,..., Xₖ and corresponding frequencies f₁, f₂,..., fₖ (such that Σfᵢ = n), the geometric mean is given by G = ⁿ√(X₁^f₁ × X₂^f₂ × ... × Xₖ^fₖ). Each value of X must be multiplied by itself f times, making this quite difficult. In terms of logarithms, the formula becomes log G = (1/n)[f₁log X₁ + f₂log X₂ + ... + fₖlog Xₖ] = Σf log X / n. Hence, G = antilog [Σf log X / n]. 📌 Example: For the EPA mileage ratings data:
| Mileage Rating | No. of Cars (f) | Class-mark (X) | log X | f log X |
|---|---|---|---|---|
| 30.0 - 32.9 | 2 | 31.45 | 1.4976 | 2.9952 |
| 33.0 - 35.9 | 4 | 34.45 | 1.5372 | 6.1488 |
| 36.0 - 38.9 | 14 | 37.45 | 1.5735 | 22.0290 |
| 39.0 - 41.9 | 8 | 40.45 | 1.6069 | 12.8552 |
| 42.0 - 44.9 | 2 | 43.45 | 1.6380 | 3.2760 |
| Total | 30 | 47.3042 | ||
| G = antilog (47.3042 / 30) = antilog 1.5768 = 37.74 miles per gallon. |
💡 Why this matters: The geometric mean is used when relative changes in a variable are to be averaged, such as rates of growth over time. Using the arithmetic mean can exaggerate the average rate. 📌 Example: A firm’s turnover increased over four years, with yearly percentages of 125%, 200%, 150%, and 140% compared to the previous year. The arithmetic mean of these percentages is (125+200+150+140)/4 = 153.75%. Applying this rate to the initial £2,000 gives an incorrect final turnover of £11,176, whereas the actual was £10,500. The geometric mean is ⁴√(125×200×150×140) = 151.37%. Applying this rate gives the correct final turnover of £10,500. This shows the average annual increase was 51.37%, not 53.75%.
Harmonic Mean
The harmonic mean is defined as the reciprocal of the arithmetic mean of the reciprocals of the values. In case of raw data: H.M. = n / Σ(1/X). In case of grouped data: H.M. = n / [Σf(1/X)], where X represents the midpoints of the classes.
🔑 Definition — Harmonic Mean (H.M.): The reciprocal of the arithmetic mean of the reciprocals of the data values. It is the appropriate average when values are given as x per y, where x is constant and y is variable. 📐 Formula: H.M. = n / Σ(1/X) (for raw data) → This means the harmonic mean is the total number of observations divided by the sum of the reciprocals of the observations. 📌 Example: A car travels 100 miles in 10 intervals of 10 miles each, at speeds of 30, 35, 40, 40, 45, 40, 50, 55, 55, and 30 mph. The arithmetic mean of the speeds is 42 mph, which is incorrect. The correct average speed is total distance (100 miles) divided by total time (2.4881 hours), which equals 40.2 mph. The harmonic mean correctly computes this: n = 10, Σ(1/X) = 0.2488, so H.M. = 10 / 0.2488 = 40.2 mph.
Rules for Using Different Means
- When values are given as x per y where x is constant and y is variable, the Harmonic Mean is the appropriate average to use.
- When values are given as x per y where y is constant and x is variable, the Arithmetic Mean is the appropriate average to use.
- When relative changes in some variable quantity are to be averaged, the Geometric Mean is the appropriate average to use.
📌 Example: If 10 students have marks out of 20 (13, 11, 9, 9, 6, 5, 19, 17, 12, 9), the average marks by the arithmetic mean is 11. This is correct because all marks are "marks per 20", where the denominator (y=20) is constant, and the numerator (marks, x) is variable.
Relation Between Arithmetic, Geometric and Harmonic Means
For any set of positive data, the following inequality holds: Arithmetic Mean ≥ Geometric Mean ≥ Harmonic Mean.
Some Other Measures of Central Tendency
- Mid-Range: For n observations with x₀ and xₘ as the smallest and largest observations, the mid-range is (x₀ + xₘ) / 2.
- Mid-Quartile Range: For n observations with Q₁ and Q₃ as the first and third quartiles, the mid-quartile range is (Q₁ + Q₃) / 2. This is also known as the mid-hinge.
A measure of central tendency is a single number that represents a whole set of data, often called an "average". The choice of which average to use depends on the nature of the data and the specific question being asked.
⭐ Key Takeaways
The geometric mean (antilog of the average of logs) is essential for averaging rates of growth or relative changes, as using the arithmetic mean in such cases leads to an overestimation of the average rate. The harmonic mean (reciprocal of the average of reciprocals) is the correct average for data expressed as a ratio where the numerator is constant and the denominator varies, such as calculating average speed when distances are equal. A fundamental inequality states that for any set of positive data, the arithmetic mean is always greater than or equal to the geometric mean, which is always greater than or equal to the harmonic mean. The mid-range (average of the minimum and maximum values) and the mid-quartile range (average of the first and third quartiles) are additional, though less commonly used, measures of central tendency. The key skill is to correctly identify which measure—arithmetic, geometric, or harmonic mean—is appropriate for a given problem based on the nature of the data.
🧠 Quick Revision Questions
- What is the formula for the geometric mean of raw data using logarithms?
- Why does the arithmetic mean give an incorrect "average" rate of increase for a firm's turnover over several years, and which measure should be used instead?
- In the car speed example, why did the arithmetic mean of the speeds (42 mph) give an incorrect average speed, and how did the harmonic mean provide the correct answer?
- State the mathematical relationship between the arithmetic mean, geometric mean, and harmonic mean for a set of positive values.
- Define the mid-range and the mid-quartile range.
📘 Lecture 10 — Dispersion: Absolute and Relative Measures, Range, Coefficient of Dispersion, Quartile Deviation, Coefficient of Quartile Deviation
📖 Overview: This lecture introduces the concept of dispersion as a measure of the variability or spread in a dataset. It explains why knowledge of the average alone is insufficient, and differentiates between absolute and relative measures. The first two specific measures—Range and Quartile Deviation—are defined, computed, and compared, along with their associated coefficients for relative comparison.
🗂️ Topics Covered
The lecture begins by establishing the need for measures of dispersion using examples of datasets with the same mean but different spreads. It then classifies measures into absolute and relative types. The Range and its associated Coefficient of Dispersion are explained, followed by the Quartile Deviation (Semi-Interquartile Range) and its relative measure, the Coefficient of Quartile Deviation. Finally, the lecture introduces the rationale for the upcoming measures of Mean Deviation and Standard Deviation.
📝 Lecture Summary
Concept of Dispersion
Just as variable series differ in their location (average), they also differ in the amount of variability or scatter they exhibit. Two datasets can have the same arithmetic mean but be entirely different in how their values are distributed. Therefore, a measure of dispersion is needed to accompany the relevant measure of central tendency.
🔑 Definition — Dispersion: The amount of variability, scatter, or spread exhibited by the values in a dataset. It distinguishes datasets that share the same average but have different distributions. 📌 Example: A first-year class (ages 17, 18, 19) shows consistent ages, while an evening class (ages 18 to 58) shows great variability. Both groups have ages, but their dispersion is very different. 💡 Why this matters: Knowing the average alone can be misleading. Dispersion provides a fuller description of the data.
Absolute Versus Relative Measures of Dispersion
There are two types of dispersion measurements: absolute and relative.
An absolute measure of dispersion measures dispersion in the same units (or the square of units) as the data.
A relative measure of dispersion is expressed as a ratio, coefficient, or percentage and is independent of units. It is useful for comparing datasets of different natures (e.g., comparing dispersion of heights in meters with weights in kilograms).
This lecture discusses four measures: the Range, the Quartile Deviation, the Mean Deviation, and the Standard Deviation.
Range
The range is the difference between the two extreme values of a dataset.
📐 Formula: R = Xm – X0, where Xm = highest value and X0 = lowest value.
It is a simple measure of mental arithmetic but gives no idea of the distribution of observations between the ends. For grouped data, the range is the difference between the upper boundary of the highest class and the lower boundary of the lowest class. It is appropriately used in quality control charts, daily temperatures, and stock prices. Disadvantages:
- It ignores all information from intermediate observations.
- It can give a misleading picture of the spread.
Viewpoint of Range: The range can be seen as twice the arithmetic mean of the deviations of the smallest and largest values from the mid-range. As such, it is associated with the mid-range as the measure of central tendency.
Coefficient of Dispersion
The range is an absolute measure; its relative measure is the Coefficient of Dispersion (C.D.).
📐 Formula: Coefficient of Dispersion = (Range) / (Mid-Range) / 2 = (Xm – X0) / (Xm + X0)
This is a dimensionless (pure) number used for comparing dispersion across different datasets. 📌 Example: If C.D. for one dataset is 0.6 and for another is 0.4, the first dataset has greater dispersion.
Quartile Deviation
The Quartile Deviation (Q.D.) is defined as half the difference between the third quartile (Q3) and the first quartile (Q1). It is also known as the Semi-Interquartile Range.
📐 Formula: Q.D. = (Q3 – Q1) / 2
It is not an extremely satisfactory measure as it only uses two values. An attractive feature is that the range Median ± Q.D. contains approximately 50% of the data.
Viewpoint of Quartile Deviation: The Q.D. can be seen as the arithmetic mean of the deviations of the first and third quartiles from the Median. As such, it is associated with the median and should be used when the median is the chosen average. 📌 Example: For Company X: Q1=60, Q3=270 → Q.D. = (270-60)/2 = 105 shares (high scatter). For Company Y: Q1=165, Q3=210 → Q.D. = (210-165)/2 = 22.5 shares (concentrated around median). 💡 Why this matters: A larger Q.D. indicates greater scatter. Q.D. is superior to the Range as it is not affected by extreme values.
Coefficient of Quartile Deviation
The quartile deviation is an absolute measure; its relative measure is the Coefficient of Quartile Deviation.
📐 Formula: Coefficient of Quartile Deviation = Q.D. / Mid-Quartile Range = (Q3 – Q1) / (Q3 + Q1)
This is a pure number used for comparing variation in different datasets.
Introduction to Mean Deviation and Standard Deviation
Unlike the Range and Quartile Deviation, the Mean Deviation and Standard Deviation are based on each and every data value.
To measure dispersion around the arithmetic mean, the natural idea is to compute distances from the mean. However, the sum of deviations from the mean is always zero.
The solution is to ignore the signs of these deviations. Summing the absolute deviations gives a non-zero total. Averaging these absolute differences yields the Mean Deviation.
🔑 Definition — Mean Deviation (M.D.): The arithmetic mean of the absolute deviations of observations from their mean. Its full name is Mean Absolute Deviation. 📐 Formula: M.D. = Σ |d| / n, where |d| = absolute deviation from the mean.
This concept will be discussed in detail in the next lecture, followed by the Standard Deviation.
⭐ Key Takeaways
The most critical concepts from this lecture are: (1) Dispersion measures the spread of data and is essential alongside central tendency for a complete description. (2) Absolute measures are in the original units; relative measures (coefficients) are unitless and allow comparison between different datasets. (3) The Range is the simplest but most limited measure, ignoring all intermediate data; it is associated with the mid-range. (4) The Quartile Deviation (Semi-Interquartile Range) is a more robust measure that focuses on the middle 50% of data and is associated with the median. (5) Both the Range and Quartile Deviation have relative coefficients (Coefficient of Dispersion and Coefficient of Quartile Deviation, respectively) for comparative analysis.
🧠 Quick Revision Questions
- Why is it necessary to use a measure of dispersion alongside a measure of central tendency? Give a specific example from the lecture.
- What is the fundamental difference between an absolute measure and a relative measure of dispersion?
- Write the formula for the Range. What are its two main disadvantages?
- Calculate the Quartile Deviation for a dataset where Q1 = 25 and Q3 = 75. What does this value represent?
- Why can't the simple sum of deviations from the mean (Σ (X – X̄)) be used as a measure of dispersion? What solution does the Mean Deviation propose?
📘 Lecture 11 — Mean Deviation, Standard Deviation and Variance, Coefficient of Variation
📖 Overview: This lecture introduces two important measures of dispersion that involve all data values: the mean deviation and the standard deviation. It explains how to compute these for raw data and frequency distributions, and introduces the coefficient of variation as a relative measure for comparing variability across different datasets.
🗂️ Topics Covered
The lecture covers the mean deviation as an absolute measure of dispersion based on absolute deviations from the mean, followed by the standard deviation and variance which square the deviations to overcome mathematical limitations. It then discusses the coefficient of variation as a relative measure useful for comparing variability between different datasets or variables.
📝 Lecture Summary
Mean Deviation
The mean deviation is a measure of dispersion that involves each data value. Since the sum of deviations from the arithmetic mean is always zero, we take the absolute values of these deviations. The mean deviation for raw data is the average of these absolute deviations from the mean. For grouped data (frequency distribution), the formula accounts for frequencies.
🔑 Definition — Mean Deviation: The average of the absolute deviations of observations from their mean. 📐 Formula: M.D. = Σ|dᵢ| / n (for raw data); M.D. = Σfᵢ|xᵢ - x̄| / n (for grouped data) 📌 Example: For fatalities data (4,6,2,0,3,5,8), mean = 4. Absolute deviations: 0,2,2,4,1,1,4. Sum = 14. M.D. = 14/7 = 2 fatalities. 💡 Why this matters: The mean deviation provides a quick and simple measure using all data, but ignoring signs creates mathematical limitations for further analysis.
The coefficient of mean deviation is the relative form obtained by dividing the mean deviation by the average used (mean or median). The median is preferred for datasets with extreme values. 📐 Formula: Coefficient of M.D. = M.D./Mean or M.D./Median
Standard Deviation and Variance
The variance overcomes the sign problem by squaring the deviations from the mean. The standard deviation is the positive square root of the variance, bringing the measure back to the original units. A short cut formula simplifies computation using Σx and Σx².
🔑 Definition — Variance: The average of the squared deviations of values from their mean. 🔑 Definition — Standard Deviation: The positive square root of the variance. 📐 Formula: Variance = Σ(x - x̄)²/n; Standard Deviation S = √[Σ(x - x̄)²/n] 📐 Short Cut Formula: S = √[(Σx²/n) - (Σx/n)²] 📌 Example: For fatalities data, Σ(x-x̄)² = 42. Variance = 42/7 = 6 squared fatalities. Using short cut: Σx=28, Σx²=154. S = √[(154/7) - (28/7)²] = √(22-16) = √6 = 2.45 fatalities.
For grouped data, the standard deviation formula includes frequencies: 📐 Formula: S = √[Σf(x - x̄)²/n] = √[(Σfx²/n) - (Σfx/n)²] 📌 Example: For light bulb data (life in hundreds of hours): Σfx=2437.5, Σfx²=78781.25, n=100. Mean=24.375. S = √[(78781.25/100) - (2437.5/100)²] = √(787.8125-594.14) = √193.67 = 13.9 hundred hours = 1390 hours.
Coefficient of Variation
The coefficient of variation (CV) is the most important relative measure of dispersion, expressed as a percentage. It is used to compare variability between different variables or datasets with very different means.
🔑 Definition — Coefficient of Variation: A relative measure of dispersion that expresses standard deviation as a percentage of the mean. 📐 Formula: C.V. = (S/x̄) × 100 📌 Example 1 — Comparing earnings variability: Country 1: mean=$19.50, S=$4, CV=20.5%. Country 2: mean=Rs.75, S=Rs.28, CV=37.3%. Country 2 has greater variability. 📌 Example 2 — Comparing crop yields: Untreated land: mean=35 bushels, S=10, CV=28.57%. Treated land: mean=58 bushels, S=10, CV=17.24%. The treated land actually has lower relative variability despite the same absolute standard deviation.
⭐ Key Takeaways
The mean deviation is the average absolute distance from the mean, but ignoring signs makes it unsuitable for further mathematical work. The standard deviation and variance overcome this by squaring deviations, with the variance in squared units and the standard deviation in original units. The short cut formulas using Σx and Σx² simplify computation significantly. The coefficient of variation is essential when comparing variability between different datasets or variables, as it normalizes standard deviation by the mean to show relative dispersion. For light bulb data with open-ended classes, midpoint approximation is necessary and standard deviation was 1390 hours.
🧠 Quick Revision Questions
- Why must we use absolute deviations for mean deviation instead of actual deviations from the mean?
- What is the mathematical advantage of squaring deviations in standard deviation over taking absolute values in mean deviation?
- Calculate the standard deviation (using short cut formula) for the dataset: 10, 12, 8, 15, 10.
- Why does the coefficient of variation show that variability has "decreased" for treated wheat land even though standard deviation remains 10 bushels?
- In which type of frequency distribution situation would we prefer mean deviation from median over mean deviation from mean?
📘 Lecture 12 — Chebychev's Inequality, The Empirical Rule, The Five-Number Summary
📖 Overview: This lecture explores how the standard deviation provides a measure of variability for a single data set. It introduces Chebychev's Theorem, which applies to any data set regardless of distribution shape, and the Empirical Rule, which applies to mound-shaped symmetric distributions. The lecture also covers the Five-Number Summary, a tool for exploratory data analysis that helps determine the shape of a distribution.
🗂️ Topics Covered
This lecture covers three major topics: Chebychev's Inequality, which gives the minimum proportion of data falling within k standard deviations of the mean for any data set; the Empirical Rule, which provides approximate percentages for mound-shaped symmetric distributions; and the Five-Number Summary, consisting of the minimum, first quartile, median, third quartile, and maximum, which helps identify the shape of a distribution. The lecture includes detailed examples and comparisons of theoretical versus actual results.
📝 Lecture Summary
Chebychev's Theorem
Chebychev's Theorem, named after Russian mathematician P.L. Chebychev (1821-1894), applies to any set of data, regardless of the shape of its frequency distribution. For any number k greater than 1, at least 1 – 1/k² of the data-values fall within k standard deviations of the mean, i.e., within the interval (X̄ – kS, X̄ + kS). This means at least 3/4 (75%) of data-values will fall within 2 standard deviations of the mean, and at least 8/9 (89%) will fall within 3 standard deviations of the mean. Because k must be greater than 1, no useful information is provided about the fraction of measurements within 1 standard deviation of the mean.
The theorem tells us that at least 75% of values fall within 2 standard deviations, at least 89% within 3 standard deviations, and at least 94% (15/16) within 4 standard deviations. For a set of data with mean 150 and standard deviation 25, at least 75% of values lie between 100 and 200, at least 89% lie between 75 and 225, and at least 96% lie between 25 and 275. For another set with the same mean but smaller standard deviation of 10, these intervals are narrower: at least 75% lie between 130 and 170, at least 89% lie between 120 and 180, and at least 96% lie between 100 and 200.
💡 Why this matters: Chebychev's Theorem provides a guarantee that works for any distribution, unlike the Empirical Rule which only works for mound-shaped distributions. However, a limitation is that it provides weak information — for many random variables, the actual probability within 2 standard deviations is far greater than 75%.
🔑 Definition — Chebychev's Theorem: Given a set of n observations x₁, x₂, x₃... xₙ on variable X, the probability is at least (1 – 1/k²) that X will take on a value within k standard deviations of the mean (where k > 1). 📐 Formula: At least 1 – 1/k² of data falls within (X̄ – kS, X̄ + kS) for k > 1 → Plain-English meaning: For any data set, we can guarantee a minimum percentage of data falls within a certain distance from the mean — the larger the distance (k), the larger the percentage. 📌 Example: Data with mean = 150, S = 25, k = 2: At least 1 – 1/4 = 75% of values lie between 150 – 50 = 100 and 150 + 50 = 200.
The Empirical Rule
The Empirical Rule is a rule of thumb that applies to data sets with frequency distributions that are mound-shaped and symmetric. According to this rule: approximately 68% of measurements will fall within 1 standard deviation of the mean, i.e., within (X̄ – S, X̄ + S); approximately 95% will fall within 2 standard deviations, i.e., within (X̄ – 2S, X̄ + 2S); and approximately 100% (practically all) will fall within 3 standard deviations, i.e., within (X̄ – 3S, X̄ + 3S).
Example: For 50 companies' percentages of revenues spent on R&D, with mean = 8.49 and standard deviation = 1.98: Within 1 standard deviation (6.51, 10.47): 34 of 50 measurements (68%); within 2 standard deviations (4.53, 12.45): 47 of 50 measurements (94%); within 3 standard deviations (2.55, 14.43): all 50 measurements (100%). Even though the data distribution is skewed to the right, the percentages are remarkably close to the theoretical values (68%, 95%, 100%). Unless the distribution is extremely skewed, the mound-shaped approximations will be reasonably accurate.
📐 Rule: For mound-shaped symmetric distributions: ≈68% within X̄ ± S; ≈95% within X̄ ± 2S; ≈100% within X̄ ± 3S 📌 Example: R&D spending data, X̄ = 8.49, S = 1.98: X̄ ± S = (6.51, 10.47) contains 68% of data.
The Five-Number Summary
A five-number summary consists of X₀ (minimum), Q₁ (first quartile), Median (X̃), Q₃ (third quartile), and Xₘ (maximum). It provides a good idea about the shape of the distribution without actually drawing its graph.
For perfectly symmetrical distributions: the distance from Q₁ to the median equals the distance from the median to Q₃; the distance from X₀ to Q₁ equals the distance from Q₃ to Xₘ; and the median, mid-quartile range, and midrange are all equal (also equal to the arithmetic mean).
For right-skewed (positively skewed) distributions: the distance from Q₃ to Xₘ greatly exceeds the distance from X₀ to Q₁; and the median < mid-quartile range < midrange.
For left-skewed (negatively skewed) distributions: the distance from X₀ to Q₁ greatly exceeds the distance from Q₃ to Xₘ; and midrange < mid-quartile range < median.
Example: Annual costs (in $000) at 10 Big Ten universities: 13.0, 14.3, 14.9, 15.2, 15.2, 15.4, 15.6, 16.4, 17.0, 23.1. The five-number summary is: X₀ = 13.0, Q₁ = 14.9, Median = 15.3, Q₃ = 16.4, Xₘ = 23.1. The distance from Q₃ to Xₘ (6.7) greatly exceeds the distance from X₀ to Q₁ (1.9), and median (15.3) < mid-quartile range (15.65) < midrange (18.05), both indicating the distribution is positively skewed.
🔑 Definition — Five-Number Summary: A set of five descriptive statistics (X₀, Q₁, Median, Q₃, Xₘ) that summarizes a data set and helps determine the shape of its distribution. 📌 Example: Big Ten university costs: X₀=13.0, Q₁=14.9, Median=15.3, Q₃=16.4, Xₘ=23.1 → positively skewed distribution.
⭐ Key Takeaways
Chebychev's Theorem applies to ANY data set regardless of shape, guaranteeing at least 75% of data within 2 standard deviations and 89% within 3 standard deviations of the mean, but it cannot be used for 1 standard deviation. The Empirical Rule applies only to mound-shaped symmetric (or near-symmetric) distributions, giving approximate percentages of 68%, 95%, and 100% within 1, 2, and 3 standard deviations respectively. The Five-Number Summary (minimum, Q1, median, Q3, maximum) is an effective tool for determining distribution shape without graphing: for right-skewed data, the right tail (Q3 to max) is longer than the left tail (min to Q1), and median < mid-quartile range < midrange. Chebychev's Theorem provides weaker information than the Empirical Rule but has the advantage of universal applicability.
🧠 Quick Revision Questions
- What is the minimum proportion of data that must fall within 2 standard deviations of the mean according to Chebychev's Theorem?
- Why does Chebychev's Theorem give no information about data within 1 standard deviation of the mean?
- What are the three approximate percentages stated by the Empirical Rule for mound-shaped symmetric distributions?
- In a five-number summary for a right-skewed distribution, which distances should be compared and what inequality should hold?
- Given a five-number summary of X₀=5, Q₁=12, Median=18, Q₃=22, Xₘ=25, is the distribution symmetric, right-skewed, or left-skewed?
📘 Lecture 13 — Box and Whisker Plot & Pearson’s Coefficient of Skewness
📖 Overview: This lecture introduces two key methods for analyzing the shape of a distribution. First, the Box and Whisker Plot is presented as a graphical tool based on the Five-Number Summary to visualize spread, concentration, and skewness. Second, Pearson’s Coefficient of Skewness is introduced as a numerical measure to quantify the asymmetry of a distribution, which is crucial for understanding data patterns that mean and standard deviation alone cannot reveal.
🗂️ Topics Covered
The lecture begins by reviewing the Five-Number Summary and its properties for symmetrical and skewed distributions, illustrated with an example. It then details the step-by-step construction of a Box and Whisker Plot using a downtime dataset and interprets it for skewness. Finally, it introduces Pearson’s Coefficient of Skewness, including its formula and calculation using an example of asthma onset ages to differentiate between symmetrical and skewed distributions.
📝 Lecture Summary
FIVE-NUMBER SUMMARY
A Five-Number Summary consists of X₀, Q₁, Median, Q₃, and Xₘ. It provides insight into the shape of the distribution. For a perfectly symmetrical distribution: (1) the distance from Q₁ to the median equals the distance from the median to Q₃; (2) the distance from X₀ to Q₁ equals the distance from Q₃ to Xₘ; and (3) the median, mid-quartile range, and midrange are all equal, and also equal to the arithmetic mean.
For a right-skewed (positively-skewed) distribution, the distance from Q₃ to Xₘ greatly exceeds the distance from X₀ to Q₁. Also, the median is less than the mid-quartile range, which is less than the midrange. For a left-skewed distribution, the opposite is true: the distance from X₀ to Q₁ exceeds the distance from Q₃ to Xₘ, and the midrange is less than the mid-quartile range, which is less than the median.
🔑 Definition — Five-Number Summary: A set of five descriptive statistics (X₀, Q₁, Median, Q₃, Xₘ) that provides a summarized view of the location, spread, and shape of a dataset. 🔑 Definition — Mid-Quartile Range: The midpoint between the first and third quartiles, computed as (Q₁ + Q₃)/2. 🔑 Definition — Midrange: The midpoint between the smallest and largest values in a dataset, computed as (X₀ + Xₘ)/2.
📌 Example: For the Big Ten Universities annual cost data, the ordered array is: 13.0, 14.3, 14.9, 15.2, 15.2, 15.4, 15.6, 16.4, 17.0, 23.1. The five-number summary is: X₀ = 13.0, Q₁ = 14.9, Median = 15.3, Q₃ = 16.4, Xₘ = 23.1. The distance from Q₃ to Xₘ (23.1 - 16.4 = 6.7) greatly exceeds the distance from X₀ to Q₁ (14.9 - 13.0 = 1.9). The median (15.3) < mid-quartile range (15.65) < midrange (18.05). This indicates the data is right-skewed.
BOX AND WHISKER PLOT
A Box and Whisker Plot is a graphical representation of the data through its five-number summary. It is constructed by representing the variable on a horizontal axis, drawing a box from Q₁ to Q₃, dividing the box with a vertical line at the median, and extending "whiskers" from the left end of the box to X₀ and from the right end of the box to Xₘ.
📌 Example: For the downtime data of 30 machines (X₀ = 1, Q₁ = 4, Median = 5, Q₃ = 8.25, Xₘ = 13), the plot shows 50% of measurements are between 4 and 8.25. The median line is closer to the left end of the box, indicating the data is skewed to the right. The left whisker is shorter than the right whisker, confirming this skewness.
🔑 Definition — Box and Whisker Plot: A graphical display that uses a box to represent the middle 50% of the data (the interquartile range) and whiskers to represent the extreme values, allowing for a quick visual assessment of spread, concentration, and symmetry. 💡 Why this matters: The Box and Whisker Plot allows for a quick visual assessment of the shape of a distribution without needing to draw a frequency polygon or curve.
PEARSON’S COEFFICIENT OF SKEWNESS
Pearson’s Coefficient of Skewness is a numerical measure of the degree of asymmetry in a distribution. It is defined as: (mean - mode) / standard deviation. Since the mode is often difficult to determine, a modified formula using the empirical relation between mean, median, and mode is used. The general formula for Pearson’s Coefficient of Skewness is: 3(mean - median) / standard deviation.
For a symmetrical distribution, the coefficient is zero. For a distribution skewed to the right, the answer is positive (mean > median). For one skewed to the left, the answer is negative (mean < median).
🔑 Formula: Pearson’s Coefficient of Skewness = 3(mean – median) / standard deviation → This quantifies the skewness by scaling the difference between the mean and median by three and dividing by the standard deviation. 💡 Why this matters: This coefficient provides a single number that describes the shape of the distribution, making it possible to compare skewness across different datasets, even when their means and standard deviations are identical.
📌 Example: For the ages of onset of nervous asthma in children, both manual and non-manual worker children have the same mean (8.50 years) and standard deviation (3.61 years). However, the median for manual worker children is 8.50, and for non-manual worker children is 9.16. The Pearson’s coefficient for manual workers is 3(8.50 – 8.50)/3.61 = 0, indicating a symmetrical distribution. For non-manual workers, it is 3(8.50 – 9.16)/3.61 = –0.55, indicating a negatively skewed (left-skewed) distribution.
⭐ Key Takeaways
The Five-Number Summary (X₀, Q₁, Median, Q₃, Xₘ) is a powerful tool for determining the shape of a distribution by comparing distances between quartiles and extremes. A Box and Whisker Plot provides a visual representation of this summary, where the position of the median line and the lengths of the whiskers indicate skewness: a median closer to the left and a longer right whisker suggest right-skewness. Pearson’s Coefficient of Skewness offers a numerical value to measure asymmetry, where a positive value indicates right-skewness, a negative value indicates left-skewness, and zero indicates symmetry. This coefficient is especially useful when two distributions have the same mean and standard deviation but different shapes, as it reveals the direction and degree of skewness.
🧠 Quick Revision Questions
- What are the five components of a Five-Number Summary, and how can their relative distances indicate the shape of a distribution?
- Describe the step-by-step process for constructing a Box and Whisker Plot.
- In a Box and Whisker Plot for a right-skewed distribution, where is the median line located, and which whisker is longer?
- What is the formula for Pearson’s Coefficient of Skewness (modified version), and what does a positive value indicate about the distribution?
- In the example of the asthma onset ages, the two groups had the same mean and standard deviation but different coefficients of skewness. What does this demonstrate about the importance of measuring skewness?
📘 Lecture 14 — Bowley’s coefficient of skewness, The Concept of Kurtosis, Percentile Coefficient of Kurtosis, Moments & Moment Ratios, Sheppard’s Corrections, The Role of Moments in Describing Frequency Distributions
📖 Overview: This lecture extends the analysis of skewness by introducing Bowley’s coefficient, a measure that relies solely on quartiles and the median, making it ideal when the mean and standard deviation are not used. It then introduces the concept of kurtosis, which measures the peakedness of a distribution, and defines three types of curves: leptokurtic, mesokurtic, and platykurtic. Finally, it establishes the concept of moments, which are powerful mathematical tools for describing the shape of any frequency distribution, leading to moment ratios for skewness and kurtosis.
🗂️ Topics Covered
This lecture begins with Bowley's coefficient of skewness, a quartile-based measure, and applies it to the asthma onset data from the previous lecture. It then defines kurtosis and the percentile coefficient of kurtosis. The core of the lecture is an introduction to moments: raw moments about the mean and an arbitrary origin. Sheppard's corrections for grouped data are presented, followed by the relationships between moments about the mean and an arbitrary origin. Finally, the moment ratios b1 and b2 are introduced as key tools for measuring skewness and kurtosis, linking the first four moments to describing the center, dispersion, and shape of a distribution.
📝 Lecture Summary
Bowley’s coefficient of skewness
Bowley’s coefficient of skewness uses quartiles and the median, offering an advantage when the mean or standard deviation is not required. In an asymmetrical distribution, the quartiles are not equidistant from the median. For a positively skewed distribution, Q3 is farther from the median than Q1, making the quantity Q1 + Q3 - 2 median > 0. The relative measure is obtained by dividing this by the inter-quartile range, so the coefficient is a pure number between 0 and ±1.
🔑 Definition — Bowley’s coefficient of skewness: A relative measure of skewness based on quartiles.
📐 Formula: (Q1 + Q3 - 2 * median) / (Q3 - Q1) → Measures the asymmetry of data by comparing the distances of Q1 and Q3 from the median, relative to the total spread of the middle 50% of the data.
📌 Example: For children of manual workers: Q1=6.00, Q3=11.00, median=8.50. Calculation: (11.00 + 6.00 - 28.50) / (11.00 - 6.00) = (17.00 - 17.00)/5.00 = 0. This indicates a symmetrical distribution. For children of non-manual workers: Q1=5.50, Q3=10.83, median=9.16. Calculation: (10.83 + 5.50 - 29.16) / (10.83 - 5.50) = (16.33 - 18.32)/5.33 = (-1.99)/5.33 = -0.37. This indicates a negatively skewed distribution.
The Concept of Kurtosis
Kurtosis, introduced by Karl Pearson, measures the degree of peakedness or flatness of a unimodal frequency curve. A curve with a high, sharp peak is leptokurtic. A curve with a flat top is platykurtic. The normal curve, which is neither very peaked nor flat, is called mesokurtic and serves as the basis for comparison.
Percentile Coefficient of Kurtosis
This measure of kurtosis is based on quartiles and percentiles. It compares the spread of the middle 50% of the data to the spread between the 10th and 90th percentiles.
🔑 Definition — Percentile Coefficient of Kurtosis: A measure of the peakedness or flatness of a distribution.
📐 Formula: K = Q.D. / (P90 - P10) → where Q.D. is the quartile deviation. For a normal (mesokurtic) distribution, K = 0.263. For a leptokurtic distribution, K < 0.263. For a platykurtic distribution, K > 0.263.
Moments
A moment designates the power to which deviations are raised before averaging them. The r-th sample moment about the mean (central moment) is m_r = (1/n) * Σ(x_i - x̄)^r. For raw data, m_1 = 0 and m_2 equals the sample variance. Moments can also be taken about an arbitrary origin (m'_r) or about zero (m'_r = (1/n) * Σ x_i^r).
🔑 Definition — First moment about the mean (m₁): (1/n) * Σ(x - x̄) = 0.
🔑 Definition — Second moment about the mean (m₂): (1/n) * Σ(x - x̄)² = sample variance.
🔑 Definition — First moment about zero: (1/n) * Σ x = arithmetic mean.
📌 Example: For the marks 32, 36, 36, 37, 39, 41, 45, 46, 48, the mean x̄ = 40. Calculations for (x - x̄) are in the table. m₁ = Σ(x - x̄) / n = 0/9 = 0. m₂ = 232/9 = 25.78 marks². m₃ = 186/9 = 20.67 marks³. m₄ = 10708/9 = 1189.78 marks⁴.
Sheppard’s Corrections
When calculating moments from a grouped frequency distribution, an error is introduced by assuming all values in a class are at the midpoint. W.F. Sheppard introduced corrections, assuming the distribution is continuous and tails off to zero. The corrections are:
m₂ (corrected) = m₂ (uncorrected) - h²/12m₃ (corrected) = m₃ (uncorrected)m₄ (corrected) = m₄ (uncorrected) - (h²/2) * m₂ (uncorrected) + (7h⁴/240)These are not applicable to highly skewed distributions or those with unequal class intervals.
The Role of Moments in Describing Frequency Distributions
The first four moments about the mean play a key role in describing a distribution, allowing comparison with the normal distribution through moment ratios. The moment ratios, b₁ and b₂, are pure numbers independent of units.
🔑 Definition — Moment Ratios: b₁ = (m₃)² / (m₂)³ and b₂ = m₄ / (m₂)².
b₁measures skewness. For a symmetric distribution,m₃ = 0andb₁ = 0. Sinceb₁involves a square, it does not indicate direction; the sign ofm₃does.b₂measures kurtosis. For a normal distribution,b₂ = 3. For a leptokurtic distribution,b₂ > 3. For a platykurtic distribution,b₂ < 3.
💡 Why this matters: This shows that the first moment (zero) gives the mean, the second gives the variance, the third gives skewness, and the fourth gives kurtosis, making all four essential for a complete description of a frequency distribution's shape.
⭐ Key Takeaways
Bowley’s coefficient of skewness is a powerful alternative to Pearson’s coefficient, especially useful when only quartiles and medians are available, and its sign directly indicates the direction of skewness. Kurtosis is a crucial concept for describing the peakedness of a distribution, with the mesokurtic normal curve (K=0.263, b₂=3) serving as the reference against which leptokurtic (more peaked) and platykurtic (flatter) distributions are compared. Moments are the fundamental building blocks for describing a distribution, where the first four moments correspond to location (mean), dispersion (variance), skewness (m₃), and kurtosis (m₄). The moment ratios b₁ and b₂ are the ultimate, unit-less tools for quantifying skewness and kurtosis, respectively. Sheppard’s corrections are essential for obtaining accurate moment estimates from grouped frequency data, though they have specific applicability conditions.
🧠 Quick Revision Questions
- What is the formula for Bowley’s coefficient of skewness, and what does a positive value indicate about a distribution?
- Define leptokurtic, mesokurtic, and platykurtic curves. How does the percentile coefficient of kurtosis (K) distinguish between them?
- Write the general formula for the
r-thsample moment about the mean (m_r) for raw data. What doesm₂equal? - State Sheppard’s correction for the second moment about the mean (
m₂). Under what conditions is this correction applicable? - What are the moment ratios
b₁andb₂? What do they measure, and what are their values for a perfectly normal distribution?
📘 Lecture 15 — Simple Linear Regression, Standard Error of Estimate, and Correlation
📖 Overview: This lecture introduces the foundational concepts of bivariate analysis, focusing on how to model the relationship between two variables using simple linear regression. It covers the method of least squares for fitting the best line, the standard error of estimate as a measure of prediction reliability, and Pearson's correlation coefficient to quantify the strength of a linear relationship.
🗂️ Topics Covered
The lecture begins by defining bivariate data and illustrating it with a scatter diagram using an example of drug percentage and reaction time. It then explains the simple linear regression model, the equation of a straight line, and the principle of least squares to find the line of best fit via normal equations. The discussion moves to the standard error of estimate as a measure of scatter around the regression line, and concludes with Pearson’s product-moment correlation coefficient, its interpretation, and a detailed example.
📝 Lecture Summary
Simple Linear Regression
In many real-world situations, we are interested in the relationship between two or more variables. For example, the yield of a crop depends on factors like soil fertility, fertilizer, and rainfall. We begin with a bivariate example from a pharmaceutical company. A drug is administered to five subjects, and we record the percentage of drug in the bloodstream (independent variable X) and the reaction time in milliseconds (dependent variable Y).
| Subject | Percentage (X) | Reaction Time (Y) |
|---|---|---|
| A | 1 | 1 |
| B | 2 | 1 |
| C | 3 | 2 |
| D | 4 | 2 |
| E | 5 | 4 |
The first step is to draw a scatter diagram, a graph of X against Y. In this example, an upward trend is visible: as X increases, Y also increases. The points do not all fall on a straight line, but an overall linear pattern is apparent. The lecture notes that in behavioral and social sciences, perfect linear relationships are rare, and only a general linear tendency is expected. The reasons for this include uncontrolled factors, such as firm efficiency or market share, that cause variation even for the same X value. The goal is to superimpose a general linear relationship on this pattern to remove the effect of outside factors.
The equation of a straight line is given as: [ Y = a + bX ] where:
- (Y) is the dependent variable.
- (X) is the independent variable.
- (a) is the Y-intercept (the value of Y when X = 0).
- (b) is the slope of the line (the change in Y for a one-unit change in X).
🔑 Definition — Regression Line: The line of best fit obtained by the method of least squares. It is the line that minimizes the sum of the squares of the vertical deviations between the data points and the line.
📌 Example: Many lines can be drawn through a scatter diagram, but we need the "best" one. The method of least squares finds a unique line. The normal equations, which are solved simultaneously to find (a) and (b), are:
- (\sum Y = na + b\sum X)
- (\sum XY = a\sum X + b\sum X^2) For the drug example, we compute the necessary sums:
- (\sum X = 15), (\sum Y = 10), (\sum X^2 = 55), (\sum XY = 37), (n = 5) The normal equations become:
- (10 = 5a + 15b)
- (37 = 15a + 55b) Solving these gives (b = 0.7) and (a = -0.1). The regression line is: [ \hat{Y} = -0.1 + 0.7X ] This line can be used to estimate Y for a given X, but not for extrapolation (predicting outside the range of the data).
Standard Error of Estimate
The observed values do not all fall exactly on the regression line. The standard error of estimate measures the degree of scatter about this line.
🔑 Definition — Standard Error of Estimate ((S_{y.x})): The standard deviation of the observed Y values around the predicted values from the regression line. It measures the reliability of predictions.
📐 Formula: [ S_{y.x} = \sqrt{\frac{\sum Y^2 - a\sum Y - b\sum XY}{n-2}} ]
- (Y) is the observed value.
- (n) is the sample size.
📌 Interpretation: (S_{y.x}) lies between 0 and (S_y) (standard deviation of Y).
- (S_{y.x} = 0): All points lie on the line (perfect relationship).
- (S_{y.x} = S_y): No linear relationship.
- The closer (S_{y.x}) is to 0, the more reliable the regression line is for prediction.
📌 Example: For the drug example, with (a = -0.1), (b = 0.7), (\sum Y^2 = 26), (\sum Y = 10), (\sum XY = 37), and (n = 5): [ S_{y.x} = \sqrt{\frac{26 - (-0.1)(10) - (0.7)(37)}{5-2}} = \sqrt{\frac{26 + 1 - 25.9}{3}} = \sqrt{\frac{1.1}{3}} = \sqrt{0.3667} = 0.61 ] The standard deviation of Y ((S_y)) is 1.10. Since (S_{y.x} = 0.61) is not very small compared to (S_y = 1.10), the regression line is probably not very reliable for prediction.
Correlation
Correlation measures the strength or degree of relationship between two random variables. The Pearson’s product-moment coefficient of correlation ((r)) is a numerical measure of the strength of the linear relationship.
🔑 Definition — Pearson’s Correlation Coefficient ((r)): A pure number that measures the strength and direction of the linear relationship between two variables. It always lies between -1 and 1.
📐 Short-cut Formula: [ r = \frac{\sum XY - \frac{(\sum X)(\sum Y)}{n}}{\sqrt{[\sum X^2 - \frac{(\sum X)^2}{n}][\sum Y^2 - \frac{(\sum Y)^2}{n}]}} ]
📌 Interpretation:
- (0 < r < 1): Positive correlation (as X increases, Y tends to increase). The closer to 1, the stronger the positive linear relationship.
- (r = -1): Perfect negative linear correlation (all points on a downward-sloping line).
- (-1 < r < 0): Negative correlation (as X increases, Y tends to decrease). The closer to -1, the stronger the negative linear relationship.
- (r = 0): No linear correlation (X and Y are uncorrelated).
📌 Example: A principal wants to know if there is a correlation between grades in Mathematics (X) and Statistics (Y) for a sample of 9 students.
| X | Y | X² | Y² | XY |
|---|---|---|---|---|
| 5 | 11 | 25 | 121 | 55 |
| 12 | 16 | 144 | 256 | 192 |
| 14 | 15 | 196 | 225 | 210 |
| 16 | 20 | 256 | 400 | 320 |
| 18 | 17 | 324 | 289 | 306 |
| 21 | 19 | 441 | 361 | 399 |
| 22 | 25 | 484 | 625 | 550 |
| 23 | 24 | 529 | 576 | 552 |
| 25 | 21 | 625 | 441 | 525 |
| 156 | 168 | 3024 | 3294 | 3109 |
[ r = \frac{3109 - \frac{(156)(168)}{9}}{\sqrt{[3024 - \frac{(156)^2}{9}][3294 - \frac{(168)^2}{9}]}} = \frac{3109 - 2912}{\sqrt{[3024 - 2704][3294 - 3136]}} = \frac{197}{\sqrt{320 \times 158}} = \frac{197}{\sqrt{50560}} = \frac{197}{224.86} = 0.88 ] There is a strong positive linear correlation between marks in Mathematics and Statistics.
⭐ Key Takeaways
- Simple linear regression uses the method of least squares to find the line that minimizes the sum of squared vertical deviations, defined by (Y = a + bX), where (a) and (b) are solved from the normal equations.
- The standard error of estimate ((S_{y.x})) quantifies the scatter of data points around the regression line; a smaller value indicates more reliable predictions.
- Pearson’s correlation coefficient ((r)) measures the strength and direction of a linear relationship, ranging from -1 (perfect negative) to +1 (perfect positive), with 0 indicating no linear relationship.
- Regression lines should not be used for extrapolation (predicting outside the data range), and the regression of Y on X is not the same as the regression of X on Y.
- The lecture concludes the Descriptive Statistics portion of the course, emphasizing regression and correlation as foundations for more advanced topics like multiple regression.
🧠 Quick Revision Questions
- What is the primary purpose of drawing a scatter diagram in bivariate analysis?
- Write down the two normal equations used to find the least squares regression line (Y = a + bX).
- What does the standard error of estimate ((S_{y.x})) measure, and what does a value close to zero indicate?
- If Pearson’s correlation coefficient (r = -0.95), what does this tell you about the relationship between the two variables?
- Why is it considered unwise to use a regression line for extrapolation?
📘 Lecture 16 — Set Theory, Counting Rules: The Rule of Multiplication
📖 Overview: This lecture introduces the fundamental concepts of set theory, which is essential for understanding probability. It covers definitions of sets, subsets, operations on sets (union, intersection, difference, complementation), and the algebra of sets. The lecture also presents the basic counting rule known as the Rule of Multiplication, which is crucial for determining the total number of outcomes in compound experiments and for computing probabilities.
🗂️ Topics Covered
The lecture covers set theory definitions including empty, finite, and infinite sets; subsets and proper subsets; the universal set; Venn diagrams; operations on sets (union, intersection, difference, complementation); the algebra of sets including commutative, associative, distributive, idempotent, identity, complementation, and De Morgan’s laws; partition of sets; power set; Cartesian product of sets; tree diagrams; and the Rule of Multiplication for counting outcomes in compound experiments.
📝 Lecture Summary
Set Theory
A set is any well-defined collection or list of distinct objects. The term well-defined means any object must be classified as either belonging or not belonging to the set, and distinct implies each object appears only once. The objects in a set are called members or elements. Sets are denoted by capital letters (A, B, C), while elements are represented by small letters and enclosed in parentheses. The number of elements in set A is written as n(A). If x is an element of A, we write x ∈ A; if not, x ∉ A.
🔑 Definition — Empty Set: A set that has no elements is called an empty or null set and is denoted by the symbol φ. (Note: {0} is not an empty set as it contains the element 0.)
A set containing only one element is called a unit set or a singleton set. A set may be specified in two ways: (1) The "Roster" method lists all elements, e.g., A = {1, 3, 5, 7, 9, 11}; (2) The "Rule" method (Set Builder) states a rule that determines membership, e.g., A = {x | x is an odd number and x < 12} (the vertical line is read as "such that"). The repetition or order of elements does not change the nature of the set. A set is finite when it contains a finite number of elements; otherwise, it is infinite. The empty set is regarded as a finite set.
📌 Example of Finite Sets:
- A = {1, 2, 3, ..., 99, 100}
- B = {x | x is a month of the year}
- C = {x | x is a printing mistake in a book}
- D = {x | x is a living citizen of Pakistan}
📌 Example of Infinite Sets:
- A = {x | x is an even integer}
- B = {x | x is a real number between 0 and 1 inclusive}
- C = {x | x is a point on a line}
- D = {x | x is a sentence in the English language}
Subsets
A subset is a set that consists of some elements of another set. If B is a subset of A, then every member of set B is also a member of set A. This is written as B ⊆ A or A ⊇ B, read as "B is contained in A" or "A contains B." Any set is always regarded as a subset of itself, and the empty set φ is considered a subset of every set. Two sets A and B are equal or identical if and only if they contain exactly the same elements, i.e., A = B if and only if A ⊆ B and B ⊆ A.
🔑 Definition — Proper Subset: If a set B contains some but not all elements of another set A, while A contains each element of B (i.e., B ⊆ A and B ≠ A), then B is a proper subset of A.
Universal Set: The original set of which all sets we discuss are subsets is called the universal set (or the space) and is generally denoted by S or Ω. It contains all possible elements under consideration. A set S with n elements will produce 2^n subsets, including S and φ.
📌 Example: Consider set A = {1, 2, 3}. All possible subsets are: φ, {1}, {2}, {3}, {1,2}, {1,3}, {2,3}, and {1,2,3}. Hence, there are 2^3 = 8 subsets.
Venn Diagram
A Venn diagram is a diagram that represents sets by circular regions, parts of circular regions, or their complements with respect to a rectangle representing the space S. Named after English logician John Venn (1834-1923), these diagrams are used to represent sets and subsets pictorially and to verify relationships among sets and subsets. A simple Venn diagram shows two sets A and B within the universal set S; if they do not overlap, they are called disjoint sets.
Operations on Sets
Let sets A and B be subsets of some universal set S. These sets may be combined and operated on in various ways to form new sets.
Union of Sets: The union or sum of two sets A and B, denoted by A ∪ B (read as "A union B"), is the set of all elements that belong to at least one of the sets A and B: A ∪ B = {x | x ∈ A or x ∈ B}.
📌 Example: Let A = {1, 2, 3, 4} and B = {3, 4, 5, 6}. Then A ∪ B = {1, 2, 3, 4, 5, 6}.
Intersection of Sets: The intersection of two sets A and B, denoted by A ∩ B (read as "A intersection B"), is the set of all elements that belong to both A and B: A ∩ B = {x | x ∈ A and x ∈ B}.
📌 Example: Let A = {1, 2, 3, 4} and B = {3, 4, 5, 6}. Then A ∩ B = {3, 4}.
Disjoint Sets: Two sets A and B are disjoint or mutually exclusive or non-overlapping when they have no elements in common, i.e., when their intersection is an empty set: A ∩ B = φ. Two sets are conjoint when they have at least one element in common.
Set Difference: The difference of two sets A and B, denoted by A – B or by A – (A ∩ B), is the set of all elements of A which do not belong to B: A – B = {x | x ∈ A and x ∉ B}. In general, A – B ≠ B – A. Note that A – B and B are disjoint sets. If A and B are disjoint, then A – B coincides with set A.
Complementation: The particular difference S – A (the set of all elements of S which do not belong to A) is called the complement of A and is denoted by à or A^c: à = {x | x ∈ S and x ∉ A}. The complement of S is the empty set φ. It should be noted that A – B and A ∩ B̃ (where B̃ is the complement of set B) are the same set.
Algebra of Sets
The algebra of sets provides laws used to solve many problems in probability calculations. Let A, B, and C be any subsets of the universal set S.
-
Commutative laws:
- A ∪ B = B ∪ A
- A ∩ B = B ∩ A
-
Associative laws:
- (A ∪ B) ∪ C = A ∪ (B ∪ C)
- (A ∩ B) ∩ C = A ∩ (B ∩ C)
-
Distributive laws:
- A ∩ (B ∪ C) = (A ∩ B) ∪ (A ∩ C)
- A ∪ (B ∩ C) = (A ∪ B) ∩ (A ∪ C)
-
Idempotent laws:
- A ∪ A = A
- A ∩ A = A
-
Identity laws:
- A ∪ S = S
- A ∩ S = A
- A ∪ φ = A
- A ∩ φ = φ
-
Complementation laws:
- A ∪ Ã = S
- A ∩ Ã = φ
- (Ã) = A
- Š = φ
- φ̃ = S
-
De Morgan’s laws:
- (A ∪ B) = Ã ∩ B̃
- (A ∩ B) = Ã ∪ B̃
Partition of Sets
A partition of a set S is a sub-division of the set into non-empty subsets that are disjoint and exhaustive, i.e., their union is the set S itself. This implies:
- (i) A_i ∩ A_j = φ, where i ≠ j
- (ii) A_1 ∪ A_2 ∪ ... ∪ A_n = S
The subsets in a partition are called cells.
📌 Example: Let S = {a, b, c, d, e}. Then {a, b} and {c, d, e} is a partition of S as each element belongs to exactly one cell.
Class of Sets and Power Set
A set of sets is called a class. The class of ALL subsets of a set A is called the Power Set of A and is denoted by P(A).
📌 Example: If A = {H, T}, then P(A) = {φ, {H}, {T}, {H, T}}.
Cartesian Product of Sets
The Cartesian product of sets A and B, denoted by A × B (read as "A cross B"), is a set that contains all ordered pairs (x, y) where x belongs to A and y belongs to B: A × B = {(x, y) | x ∈ A and y ∈ B}. This is named after French mathematician René Descartes (1596-1605). The product of a set A by itself is denoted by A². This concept may be extended to any finite number of sets.
📌 Example: Let A = {H, T} and B = {1, 2, 3, 4, 5, 6}. Then A × B = {(H,1); (H,2); (H,3); (H,4); (H,5); (H,6); (T,1); (T,2); (T,3); (T,4); (T,5); (T,6)} — 12 (2 × 6) ordered pairs. These twelve elements make up the universal set S when a coin and a die are tossed together. In general, A × B ≠ B × A.
Tree Diagram
A tree diagram is a useful device for enumerating all possible outcomes of two or more sequential events. The "tree" is constructed from left to right, and the possible outcomes are represented by the individual paths or branches of the tree. The tree diagram conveniently finds the Cartesian product A × B.
Rule of Multiplication
The Rule of Multiplication states: If a compound experiment consists of two experiments where the first experiment has exactly m distinct outcomes, and corresponding to each outcome of the first experiment there can be n distinct outcomes of the second experiment, then the compound experiment has exactly m × n outcomes.
📌 Example: The compound experiment of tossing a coin and throwing a die consists of two experiments. The coin has 2 outcomes (H, T), and the die has 6 outcomes (1, 2, 3, 4, 5, 6). The total number of distinct outcomes is 2 × 6 = 12.
💡 Why this matters: The tree diagram and Cartesian product both verify that the total possible outcomes = 2 × 6 = 12.
The rule of multiplication can be readily extended to compound experiments consisting of any number of experiments performed in a given sequence. This rule can also be called the Multiple Choice Rule.
📌 Example: A restaurant offers 3 types of soups, 4 types of sandwiches, and 2 types of desserts. A customer can order 3 × 4 × 2 = 24 different meals.
📌 Example (Combination Lock): Suppose a combination lock has 8 rings. In how many ways can the lock be adjusted? Solution: Each of the 8 rings can have any of the 10 digits 0 to 9. Hence, the total number of ways is 10 × 10 × 10 × 10 × 10 × 10 × 10 × 10 = 10^8 = 100,000,000 (one hundred million).
⭐ Key Takeaways
The most critical concepts from this lecture are: (1) The fundamental definitions of sets, subsets, empty set, universal set, and the different ways to specify sets (Roster and Rule methods). (2) The four basic set operations—union, intersection, difference, and complementation—along with a solid understanding of Venn diagrams to visualize these operations. (3) The seven algebraic laws (commutative, associative, distributive, idempotent, identity, complementation, and De Morgan's laws) which are essential tools for simplifying probability problems. (4) The concept of Cartesian product and the Rule of Multiplication, which provide the foundation for counting the total number of outcomes in compound experiments. (5) The use of tree diagrams as a practical method for enumerating all possible outcomes when sequential events occur.
🧠 Quick Revision Questions
- What is the difference between an empty set φ and a unit set {0}?
- If set A has 4 elements, how many total subsets does A have, and how many of those are proper subsets?
- Using the sets A = {1, 2, 3, 4, 5} and B = {4, 5, 6, 7}, find: a) A ∪ B, b) A ∩ B, c) A – B, d) Verify De Morgan's law for (A ∪ B).
- A student can choose from 5 different shirts, 3 different pants, and 4 different ties. How many different outfits (shirt, pants, and tie) can the student create?
- Explain why the Cartesian product A × B is generally not equal to B × A, and give a simple example with sets of numbers to illustrate this.
📘 Lecture 17 — COUNTING RULES
📖 Overview: This lecture covers the fundamental counting rules used to calculate probabilities, specifically permutations and combinations. It then introduces the foundational concepts of probability theory, including random experiments, sample spaces, and types of events. Understanding these concepts is crucial for correctly computing probabilities in a wide range of statistical problems.
🗂️ Topics Covered
The lecture reviews counting rules (multiple choice, permutations, combinations), then defines factorial notation. It explains the rule of permutations for distinct and non-distinct objects, followed by the rule of combinations and its properties as binomial coefficients. The lecture then transitions to probability concepts, defining random experiments, sample spaces, events (simple, compound, complementary), and the key event relationships: mutually exclusive, exhaustive, and equally likely events.
📝 Lecture Summary
COUNTING RULES
As discussed in the last lecture, there are certain rules that facilitate the calculations of probabilities in certain situations. They are known as counting rules and include concepts of; Multiple Choice, Permutations, and Combinations. The rule of multiplication was already discussed in the last lecture. Let us now consider the rule of permutations.
RULE OF PERMUTATION
A permutation is any ordered subset from a set of n distinct objects. The number of permutations of r objects, selected in a definite order from n distinct objects is denoted by the symbol ( _nP_r ) and is given by: [ _nP_r = n(n-1)(n-2)...(n-r+1) = \frac{n!}{(n-r)!} ]
FACTORIALS A factorial is the product of all positive integers less than or equal to a given positive integer. For example, ( 7! = 7 \times 6 \times 5 \times 4 \times 3 \times 2 \times 1 ). Also, we define ( 0! = 1 ).
📐 Formula: ( _nP_r = \frac{n!}{(n-r)!} ) → The number of ways to arrange 'r' items selected from a set of 'n' distinct items, where order matters.
📌 Example: A club consists of four members. How many ways are there of selecting three officers: president, secretary and treasurer? It is evident that the order, in which 3 officers are to be chosen, is of significance. Thus there are 4 choices for the first office, 3 choices for the second office, and 2 choices for the third office. Hence the total number of ways is ( 4 \times 3 \times 2 = 24 ). The same result is obtained by applying the rule of permutations: [ _4P_3 = \frac{4!}{(4-3)!} = \frac{4 \times 3 \times 2 \times 1}{1!} = 24 ] Let the four members be, A, B, C and D. A tree diagram provides an organized way of listing the possible arrangements.
In the formula of ( _nP_r ), if we put r = n, we obtain: ( _nP_n = n! ). I.e. the total number of permutations of n distinct objects, taking all n at a time, is equal to ( n! ).
📌 Example: Suppose that there are three persons A, B & D, and that they wish to have a photograph taken. The total number of ways in which they can be seated on three chairs is ( _3P_3 = 3! = 6 ). These are: ABD, ADB, BAD, BDA, DAB, DBA.
The above discussion pertained to the case when all the objects under consideration are distinct objects. If some of the objects are not distinct, the formula of permutations modifies as given below: The number of permutations of n objects, selected all at a time, when n objects consist of ( n_1 ) of one kind, ( n_2 ) of a second kind, ..., ( n_k ) of a kth kind, is: [ P = \frac{n!}{n_1! n_2! ... n_k!} \quad (\text{where } \sum n_i = n) ]
📐 Formula: ( P = \frac{n!}{n_1! n_2! ... n_k!} ) → The number of distinct arrangements of 'n' total objects where some objects are identical to each other.
📌 Example: How many different (meaningless) words can be formed from the word ‘committee’? In this example: n = 9 (total letters), ( n_1 = 1 ) (c), ( n_2 = 1 ) (o), ( n_3 = 2 ) (m’s), ( n_4 = 1 ) (i), ( n_5 = 2 ) (t’s), and ( n_6 = 2 ) (e’s). Hence, the total number of (meaningless) words is: [ P = \frac{9!}{1! 1! 2! 1! 2! 2!} = \frac{9 \times 8 \times 7 \times 6 \times 5 \times 4 \times 3 \times 2 \times 1}{1 \times 1 \times 2 \times 1 \times 1 \times 2 \times 1 \times 2 \times 1} = 45360 ]
RULE OF COMBINATION
A combination is any subset of r objects, selected without regard to their order, from a set of n distinct objects. The total number of such combinations is denoted by the symbol ( _nC_r ) or ( \binom{n}{r} ) and is given by: [ \binom{n}{r} = \frac{n!}{r! (n-r)!} ] It should be noted that ( _nP_r = r! \binom{n}{r} ). In other words, every combination of r objects (out of n objects) generates ( r! ) permutations.
🔑 Definition — Combination: A selection of r objects from a set of n distinct objects where the order of selection does not matter.
📐 Formula: ( \binom{n}{r} = \frac{n!}{r!(n-r)!} ) → The number of ways to choose 'r' items from a set of 'n' distinct items, where order does not matter.
📌 Example: Suppose we have a group of three persons, A, B, & C. If we wish to select a group of two persons out of these three, the three possible groups are {A, B}, {A, C} and {B, C}. In other words, the total number of combinations of size two out of this set of size three is 3. Now, suppose that our interest lies in forming a committee of two persons, one of whom is to be the president and the other the secretary. The six possible committees (permutations) are: (A, B), (B, A), (A, C), (C, A), (B, C) & (C, B). The point to note is that each of the three combinations generates 2 = 2! permutations.
The quantity ( \binom{n}{r} ) or ( nC_r ) is also called a binomial coefficient because of its appearance in the binomial expansion of: [ (a+b)^n = \sum{r=0}^n \binom{n}{r} a^{n-r} b^r ] The binomial coefficient has two important properties: (i) ( \binom{n}{r} = \binom{n}{n-r} ), and (ii) ( \binom{n}{n-r} + \binom{n}{r} = \binom{n+1}{r} ) Also, it should be noted that ( \binom{n}{0} = 1 = \binom{n}{n} ) and ( \binom{n}{1} = n = \binom{n}{n-1} ).
📌 Example: A three-person committee is to be formed out of a group of ten persons. In how many ways can this be done? Since the order is unimportant, it is a problem involving combinations. Thus the desired number is: [ \binom{10}{3} = \frac{10!}{3! (10-3)!} = \frac{10!}{3! 7!} = \frac{10 \times 9 \times 8 \times 7!}{3 \times 2 \times 1 \times 7!} = 120 ]
📌 Example: In how many ways can a person draw a hand of 5 cards from a well-shuffled ordinary deck of 52 cards? The total number of ways is given by: [ \binom{52}{5} = \frac{52 \times 51 \times 50 \times 49 \times 48}{5 \times 4 \times 3 \times 2 \times 1} = 2,598,960 ]
RANDOM EXPERIMENT
Having reviewed the counting rules, let us now begin the discussion of concepts that lead to the formal definitions of probability. The term experiment means a planned activity or process whose results yield a set of data. A single performance of an experiment is called a trial. The result obtained from an experiment or a trial is called an outcome.
🔑 Definition — Random Experiment: An experiment which produces different results even though it is repeated a large number of times under essentially similar conditions is called a Random Experiment. The tossing of a fair coin, the throwing of a balanced die, drawing of a card from a well-shuffled deck of 52 playing cards, selecting a sample, etc. are examples of random experiments.
PROPERTIES OF A RANDOM EXPERIMENT A random experiment has three properties:
- The experiment can be repeated, practically or theoretically, any number of times.
- The experiment always has two or more possible outcomes.
- The outcome of each repetition is unpredictable, i.e. it has some degree of uncertainty.
📌 Example: Interviewing a person to find out whether or not he or she is a smoker is an example of a random experiment. This process can be repeated a large number of times, there are at least two possible replies (‘I am a smoker’ and ‘I am not a smoker’), and the answer is not known in advance.
SAMPLE SPACE
🔑 Definition — Sample Space: A set consisting of all possible outcomes that can result from a random experiment (real or conceptual) is called the sample space for the experiment and is denoted by the letter S. Each possible outcome is a member of the sample space and is called a sample point in that space.
📌 Example-1: The experiment of tossing a coin results in either of the two possible outcomes: a head (H) or a tail (T). The sample space is S = {H, T}.
📌 Example-2: The sample space for tossing two coins once (or tossing a coin twice) will contain four possible outcomes: S = {HH, HT, TH, TT}. S is the Cartesian product A × A, where A = {H, T}.
📌 Example-3: The sample space S for the random experiment of throwing two six-sided dice can be described by the Cartesian product A × A, where A = {1, 2, 3, 4, 5, 6}. S contains 36 outcomes or sample points.
EVENTS
🔑 Definition — Event: Any subset of a sample space S of a random experiment is called an event. An event is an individual outcome or any number of outcomes (sample points) of a random experiment.
SIMPLE & COMPOUND EVENTS An event that contains exactly one sample point is defined as a simple event. A compound event contains more than one sample point and is produced by the union of simple events.
📌 Example: The occurrence of a 6 when a die is thrown is a simple event, while the occurrence of a sum of 10 with a pair of dice is a compound event, as it can be decomposed into three simple events (4, 6), (5, 5) and (6, 4).
OCCURRENCE OF AN EVENT An event A is said to occur if and only if the outcome of the experiment corresponds to some element of A.
📌 Example: Suppose we toss a die, and we are interested in the occurrence of an even number. If ANY of the three numbers ‘2’, ‘4’ or ‘6’ occurs, we say that the event of our interest has occurred. The event A is represented by the set {2, 4, 6}.
COMPLEMENTARY EVENT The event “not-A” is denoted by ( \bar{A} ) or ( A^c ) and called the negation (or complementary event) of A.
📌 Example: If we toss a coin once, then the complement of “heads” is “tails”. If we toss a coin four times, then the complement of “at least one head” is “no heads”.
A sample space consisting of n sample points can produce ( 2^n ) different subsets (or simple and compound events). The subset that is the sample space itself is called the certain or sure event. The empty set ( \phi ) is called the impossible event.
MUTUALLY EXCLUSIVE EVENTS
🔑 Definition — Mutually Exclusive Events: Two events A and B of a single experiment are said to be mutually exclusive or disjoint if and only if they cannot both occur at the same time, i.e. they have no points in common.
📌 Example-1: When we toss a coin, we get either a head or a tail, but not both at the same time. The two events 'head' and 'tail' are therefore mutually exclusive.
📌 Example-2: When a die is rolled, the events ‘even number’ and ‘odd number’ are mutually exclusive. Three or more events originating from the same experiment are mutually exclusive if pairwise they are mutually exclusive. If the two events can occur at the same time, they are not mutually exclusive. For example, if we draw a card from an ordinary deck of 52 playing cards, it can be both a king and a diamond. Therefore, kings and diamonds are not mutually exclusive.
EXHAUSTIVE EVENTS
🔑 Definition — Exhaustive Events: Events are said to be collectively exhaustive when the union of mutually exclusive events is equal to the entire sample space S.
📌 Examples: In the coin-tossing experiment, ‘head’ and ‘tail’ are collectively exhaustive events. In the die-tossing experiment, ‘even number’ and ‘odd number’ are collectively exhaustive events.
PARTITION OF THE SAMPLE SPACE A group of mutually exclusive and exhaustive events belonging to a sample space is called a partition of the sample space. With reference to any sample space S, events A and ( \bar{A} ) form a partition as they are mutually exclusive and their union is the entire sample space.
EQUALLY LIKELY EVENTS
🔑 Definition — Equally Likely Events: Two events A and B are said to be equally likely when one event is as likely to occur as the other. In other words, each event should occur in equal number in repeated trials.
📌 Example: When a fair coin is tossed, the head is as likely to appear as the tail, and the proportion of times each side is expected to appear is 1/2.
📌 Example: If a card is drawn out of a deck of well-shuffled cards, each card is equally likely to be drawn, and the probability that any card will be drawn is 1/52.
💡 Why this matters: The concepts of mutually exclusive, exhaustive, and equally likely events are the fundamental building blocks for the classical definition of probability, which is the ratio of favorable outcomes to the total number of equally likely outcomes. A proper understanding of these event types is essential for correctly setting up and solving probability problems.
⭐ Key Takeaways
The critical concepts from this lecture are the distinction between permutations (ordered arrangements) and combinations (unordered selections), and the formulas to calculate their respective numbers. You must also be able to define a random experiment, its sample space, and the different types of events, including simple, compound, mutually exclusive, exhaustive, and equally likely events. A firm grasp of these definitions and formulas is essential for understanding how probabilities are formally defined and calculated in subsequent lectures.
🧠 Quick Revision Questions
- What is the formula for the number of permutations of r objects chosen from n distinct objects, and how does it differ from the formula for combinations?
- How does the permutation formula change when some of the objects are not distinct?
- Define a random experiment and state its three essential properties.
- What is the difference between a simple event and a compound event? Give an example of each from the experiment of rolling a pair of dice.
- Explain the relationship between mutually exclusive and exhaustive events. What is a partition of a sample space?
📘 Lecture 18 — Definitions of Probability
📖 Overview: This lecture introduces the fundamental definitions of probability, beginning with essential concepts like mutually exclusive, exhaustive, and equally likely events. It then explores both subjective and objective approaches to probability, with a detailed focus on the classical (a priori) definition and the relative frequency definition, highlighting their applications and limitations.
🗂️ Topics Covered
The lecture begins with a review of mutually exclusive events, exhaustive events, equally likely events, and the partition of the sample space. It then introduces the subjective or personalistic approach to probability. The objective approach is divided into two parts: the classical definition of probability, with several examples and its shortcomings, followed by the relative frequency definition of probability.
📝 Lecture Summary
Mutually Exclusive Events
Two events A and B of a single experiment are said to be mutually exclusive or disjoint if they cannot both occur at the same time, meaning they have no points in common. For example, when tossing a coin, getting a head and getting a tail are mutually exclusive events. When rolling a die, the events ‘even number’ and ‘odd number’ are mutually exclusive. Three or more events are mutually exclusive if they are pairwise mutually exclusive. If two events can occur at the same time, they are not mutually exclusive; for instance, drawing a card from a deck could result in both a king and a diamond.
🔑 Definition — Mutually Exclusive Events: Events that cannot occur at the same time; they have no sample points in common.
📌 Example: In a single coin toss, the events "Head" and "Tail" are mutually exclusive because you cannot get both simultaneously.
Exhaustive Events
Events are said to be collectively exhaustive when the union of mutually exclusive events is equal to the entire sample space S. For instance, in the coin-tossing experiment, ‘head’ and ‘tail’ are collectively exhaustive events. In the die-tossing experiment, ‘even number’ and ‘odd number’ are collectively exhaustive events.
🔑 Definition — Exhaustive Events: A set of events whose union equals the entire sample space.
Partition of the Sample Space
A group of mutually exclusive and exhaustive events belonging to a sample space is called a partition of the sample space. With reference to any sample space S, events A and its complement (Ā) form a partition as they are mutually exclusive and their union is the entire sample space.
🔑 Definition — Partition of the Sample Space: A set of events that are both mutually exclusive and collectively exhaustive.
Equally Likely Events
Two events A and B are said to be equally likely when one event is as likely to occur as the other. In other words, each event should occur in equal number in repeated trials. For example, when a fair coin is tossed, a head is as likely to appear as a tail, with a proportion of 1/2. If a card is drawn from a well-shuffled deck, each card is equally likely to be drawn, with a proportion of 1/52.
🔑 Definition — Equally Likely Events: Events that have the same chance of occurring in a random experiment.
Subjective or Personalistic Probability
The subjective or personalistic probability is a measure of the strength of a person’s belief regarding the occurrence of an event A. Probability in this sense is purely subjective and is based on whatever evidence is available to the individual. Its disadvantage is that two or more persons faced with the same evidence may arrive at different probabilities. For example, in a trial, two judges might find the accused guilty while a third finds the evidence insufficient, illustrating the subjective nature of this probability.
🔑 Definition — Subjective Probability: A measure of an individual's personal belief in the likelihood of an event, which can vary from person to person.
1. The Classical or ‘A Priori’ Definition of Probability
If a random experiment can produce n mutually exclusive and equally likely outcomes, and if m out of these outcomes are considered favorable to the occurrence of a certain event A, then the probability of the event A, denoted by P(A), is defined as the ratio m/n.
📐 Formula: P(A) = m/n = Number of favorable outcomes / Total number of possible outcomes
This definition was formulated by the French mathematician P.S. Laplace and can be conveniently used in experiments where the total number of possible outcomes and the number of outcomes favorable to an event can be determined.
📌 Example-1: If a card is drawn from an ordinary deck of 52 playing cards, the probability that it is a red card is:
- Let A be the event that the card is red. Number of favorable outcomes = 26 (13 diamonds + 13 hearts).
- P(A) = 26/52 = 1/2. The probability that the card is a 10:
- Let B be the event that the card is a 10. Number of favorable outcomes = 4.
- P(B) = 4/52 = 1/13.
📌 Example-2: A fair coin is tossed three times. The sample space S = {HHH, HHT, HTH, THH, HTT, THT, TTH, TTT}, so n(S) = 8. Let A be the event of at least one head. Then A = {HHH, HHT, HTH, THH, HTT, THT, TTH}, so n(A) = 7.
- P(A) = 7/8.
📌 Example-3: Four items are taken at random from a box of 12 items (3 faulty, 9 good). The box is rejected if more than 1 item is faulty. Find the probability that the box is accepted. The total sample points are 12C4 = 495. The box is accepted if 0 or 1 faulty items are chosen.
- n(A) = (3C0 × 9C4) + (3C1 × 9C3) = 126 + 252 = 378.
- P(A) = 378/495 = 0.76 or 76%.
Shortcomings of the Classical Definition
- It involves circular reasoning because the term "equally likely" really means "equally probable," thus defining probability using concepts that presume a prior knowledge of the meaning of probability.
- This definition becomes vague when the possible outcomes are infinite or uncountable.
- This definition is not applicable when the assumption of equally likely does not hold, which occurs in numerous real-world situations.
The Relative Frequency Definition of Probability
The essence of this definition is that if an experiment is repeated a large number of times under identical conditions, and if the event of our interest occurs a certain number of times, then the proportion in which this event occurs is regarded as the probability of that event. For example, the proportion of students who obtain a first division in a large exam can be regarded as the probability of obtaining a first division.
🔑 Definition — Relative Frequency Probability: The probability of an event is the proportion of times it occurs in a large number of repeated trials under identical conditions.
💡 Why this matters: The relative frequency definition overcomes the limitations of the classical definition by being applicable to situations where outcomes are not equally likely or are infinite.
⭐ Key Takeaways
This lecture introduces the foundational concepts for probability theory, including mutually exclusive, exhaustive, and equally likely events, which are essential for defining a sample space's partition. The classical definition, P(A) = m/n, is useful only when all outcomes are equally likely and finite, but it suffers from circular reasoning and inapplicability to many real-world scenarios. The relative frequency definition provides a practical alternative by defining probability as the proportion of occurrences in many trials, making it more widely applicable. The subjective probability approach acknowledges individual belief differences but lacks objectivity. Understanding these definitions and their limitations is crucial for selecting the correct approach in statistical problems.
🧠 Quick Revision Questions
- What is meant by mutually exclusive events? Give an example.
- What is a partition of the sample space?
- State the classical definition of probability and write its formula.
- Explain one major shortcoming of the classical definition of probability.
- How does the relative frequency definition of probability differ from the classical definition?
📘 Lecture 19 — Relative Frequency Definition of Probability • Axiomatic Definition of Probability • Laws of Probability • Rule of Complementation • Addition Theorem
📖 Overview: This lecture presents multiple definitions of probability, including the relative frequency (a posteriori) and axiomatic definitions. It also introduces fundamental laws of probability, specifically the rule of complementation and the addition theorem, which are essential for solving probability problems.
🗂️ Topics Covered
The lecture covers the relative frequency definition of probability with empirical examples from coin-tossing and birth statistics, the axiomatic definition of probability based on Kolmogorov's three axioms, and two key laws of probability: the rule of complementation and the addition theorem for non-mutually exclusive events.
📝 Lecture Summary
The Relative Frequency Definition of Probability ('A Posteriori' Definition of Probability)
If a random experiment is repeated a large number of times, say n times, under identical conditions and if an event A is observed to occur m times, then the probability of the event A is defined as the limit of the relative frequency m/n as n tends to infinity. Symbolically, P(A) = Lim (m/n) as n → ∞. The definition assumes that as n increases indefinitely, the ratio m/n tends to become stable at the numerical value P(A). As its name suggests, the relative frequency definition relates to the relative frequency with which an event occurs in the long run. This definition is very useful in practical situations where the classical definition cannot be applied because various possible outcomes are NOT equally likely. This type of probability is also called empirical probability as it is based on EMPIRICAL evidence (observational data), and can also be called STATISTICAL PROBABILITY.
🔑 Definition — Relative Frequency Definition of Probability: If a random experiment is repeated n times and event A occurs m times, then P(A) = limit of m/n as n approaches infinity. 📐 Formula: P(A) = lim (m/n) as n → ∞ → The probability of an event is the stable long-run proportion of times it occurs. 📌 Example — Coin tossing by Kerrich in 1946: A coin was tossed 10,000 times, yielding 5067 heads and 4933 tails. The proportion of heads fluctuated widely at first (e.g., around 0.6 after 10 tosses) but settled down near 0.5 as the number of tosses increased to 10,000. This hypothetical limiting value of 0.5 is the statistical probability of heads.
📌 Example — Male birth proportions in England (1956): For major regions (each based on ~100,000 births), proportions of male births ranged between 0.512 and 0.517 (countrywide: 0.514). For rural districts of Dorset (each based on ~200 births), proportions fluctuated widely from 0.38 to 0.59. The larger sample size produces greater constancy, and the hypothetical limiting value is the statistical probability of a male birth (approximately 0.514).
💡 Why this matters: The relative frequency definition is the foundation of mathematical statistics and is used when classical equally-likely assumptions do not hold. It relies on empirical data collected from repeated experiments.
The Axiomatic Definition of Probability
This definition, introduced in 1933 by the Russian mathematician Andrei N. Kolmogorov, is based on a set of AXIOMS. Let S be a sample space with sample points E₁, E₂, ... Eᵢ, ... Eₙ. To each sample point, we assign a real number, denoted P(Eᵢ) and called the probability of Eᵢ, that must satisfy the following basic axioms:
Axiom 1: For any event Eᵢ, 0 ≤ P(Eᵢ) ≤ 1.
Axiom 2: P(S) = 1 for the sure event S.
Axiom 3: If A and B are mutually exclusive events (subsets of S), then P(A ∪ B) = P(A) + P(B).
According to the axiomatic theory, probability is defined as a non-negative real number attached to each sample point Eᵢ such that the sum of all such numbers must equal ONE. The assignment of probabilities may be based on past evidence (empirical probability) or on underlying conditions ensuring equally likely outcomes (classical definition).
🔑 Definition — Axiomatic Probability: Probability assignment to sample points satisfying three axioms: non-negativity (0 ≤ P(Eᵢ) ≤ 1), certainty (P(S) = 1), and additivity for mutually exclusive events (P(A∪B) = P(A) + P(B)).
📌 Example — Births in England and Wales (1956): The table shows 716,740 total births. Relative frequencies (treated as probabilities): Male liveborn = 0.5021, Male stillborn = 0.0120, Female liveborn = 0.4750, Female stillborn = 0.0109. For compound events: P(Male birth M) = P(A or B) = P(A) + P(B) = 0.5021 + 0.0120 = 0.5141. P(Stillbirth S) = P(B or D) = P(B) + P(D) = 0.0120 + 0.0109 = 0.0229.
💡 Why this matters: The axiomatic definition provides a formal mathematical framework for probability that unifies both classical and empirical approaches.
Law of Complementation
If Ā is the complement of an event A relative to the sample space S, then P(Ā) = 1 - P(A). Hence the probability of the complement of an event is equal to one minus the probability of the event. Complementary probabilities are very useful for solving questions like "What is the probability that, in tossing two fair dice, at least one even number will appear?"
🔑 Definition — Rule of Complementation: P(Ā) = 1 - P(A), where Ā is the complement of event A in sample space S. 📐 Formula: P(Ā) = 1 - P(A) → The probability an event does NOT occur is 1 minus the probability it does occur. 📌 Example — A coin is tossed 4 times in succession: The sample space has 2⁴ = 16 equally likely sample points. Let A be the event "at least one head occurs." Its complement Ā is "no head" = {TTTT}. P(Ā) = 1/16. Therefore, P(A) = 1 - P(Ā) = 1 - 1/16 = 15/16.
Addition Law
If A and B are any two events defined in a sample space S, then P(A ∪ B) = P(A) + P(B) - P(A ∩ B). In words: "If two events A and B are not mutually exclusive, then the probability that at least one of them occurs is given by the sum of the separate probabilities of events A and B minus the probability of the joint event A ∩ B."
🔑 Definition — Addition Theorem (General Addition Law): P(A ∪ B) = P(A) + P(B) - P(A ∩ B), for any two events A and B. 📐 Formula: P(A ∪ B) = P(A) + P(B) - P(A ∩ B) → The probability of A or B occurring equals the sum of individual probabilities minus the probability of both occurring together. 📌 Example — From the birth data: P(Male) = 0.5141, P(Stillbirth) = 0.0229, P(Male and Stillbirth) = 0.0120. Then P(Male or Stillbirth) = 0.5141 + 0.0229 - 0.0120 = 0.5250.
⭐ Key Takeaways
The most critical points from this lecture are: (1) The relative frequency definition defines probability as the limit of m/n as n approaches infinity, making it useful for empirical situations where outcomes are not equally likely. (2) The axiomatic definition, based on Kolmogorov's three axioms, provides a formal mathematical foundation for all probability. (3) The rule of complementation P(Ā) = 1 - P(A) is extremely useful when the complement of an event is easier to compute, as shown in the coin-tossing example where P(at least one head) = 15/16. (4) The addition theorem P(A∪B) = P(A) + P(B) - P(A∩B) accounts for double-counting when events are not mutually exclusive. (5) Inductive (subjective) probability is non-quantifiable and cannot be used in mathematical statistics, which relies on statistical probability.
🧠 Quick Revision Questions
-
State the relative frequency definition of probability and explain why it is called "a posteriori" probability.
-
What are Kolmogorov's three axioms of probability? Explain each axiom in simple terms.
-
Using the rule of complementation, find the probability of getting at least one head when a coin is tossed 4 times. Show your work.
-
For the addition theorem, why must we subtract P(A∩B) when computing P(A∪B)? Provide a simple intuitive explanation.
-
From the birth data in England and Wales (1956), if P(Male) = 0.5141, P(Stillbirth) = 0.0229, and P(Male and Stillbirth) = 0.0120, compute P(Male or Stillbirth) using the addition theorem.
📘 Lecture 20 — Application of Addition Theorem, Conditional Probability, Multiplication Theorem
📖 Overview: This lecture extends the foundations of probability by discussing the Addition Theorem in detail, including its application to both mutually exclusive and non-mutually exclusive events. It then introduces the crucial concept of conditional probability, which leads directly to the Multiplication Theorem for finding the probability of joint events. These theorems are essential for solving a wide range of practical probability problems.
🗂️ Topics Covered
The lecture begins by revisiting the Addition Law and demonstrating its application through examples involving cards and dice. It then covers corollaries for mutually exclusive events and uses a horse race example to illustrate probability for non-equally likely events. The concept of Conditional Probability is introduced, showing how additional information reduces the sample space. Finally, the Multiplication Theorem of probability is derived and applied to a problem involving defective and good items.
📝 Lecture Summary
ADDITION LAW
The Addition Law (General Addition Theorem) states that for any two events A and B defined in a sample space S, the probability that at least one of them occurs is given by: P(A ∪ B) = P(A) + P(B) – P(A ∩ B) . In words, it is the sum of the separate probabilities minus the probability of their joint occurrence. This law applies when events are not mutually exclusive.
🔑 Definition — Addition Law: The rule that gives the probability of the union of two events, accounting for any overlap between them.
📐 Formula: P(A ∪ B) = P(A) + P(B) – P(A ∩ B) → The probability of either A or B (or both) is the sum of their individual probabilities minus the probability of them both happening.
📌 Example: Selecting one card from a deck of 52. Let A be the event "card is a club" and B be "card is a face card". Find P(A ∪ B) (a club or a face card or both).
- Step 1: Identify probabilities. P(A) = 13/52 (13 clubs). P(B) = 12/52 (12 face cards). P(A ∩ B) = 3/52 (3 face cards are clubs).
- Step 2: Apply the formula. P(A ∪ B) = 13/52 + 12/52 - 3/52 = 22/52.
COROLLARY-1
If A and B are mutually exclusive events (they cannot occur simultaneously, meaning A ∩ B = ∅), then the Addition Law simplifies. Since P(A ∩ B) = 0, the formula becomes: P(A ∪ B) = P(A) + P(B) .
📌 Example: Tossing a pair of dice and being interested in the probability of getting a total of 5 (event A) or a total of 11 (event B). These events are mutually exclusive.
- Step 1: Find individual probabilities. P(A) = 4/36 (outcomes: (1,4), (2,3), (3,2), (4,1)). P(B) = 2/36 (outcomes: (5,6), (6,5)).
- Step 2: Apply formula for mutually exclusive events. P(A ∪ B) = 4/36 + 2/36 = 6/36 = 1/6 ≈ 16.67%.
COROLLARY-2
If A₁, A₂, ..., Aₖ are k mutually exclusive events, then the probability that one of them occurs is the sum of the probabilities of the separate events: P(A₁ ∪ A₂ ∪ ... ∪ Aₖ) = P(A₁) + P(A₂) + ... + P(Aₖ) .
📌 Example (Race Problem): Three horses A, B, C are in a race. A is twice as likely to win as B, and B is twice as likely to win as C. What is P(A or B wins)? The events are not equally likely.
- Step 1: Define probabilities relative to C. Let P(C) = p. Then P(B) = 2p and P(A) = 4p.
- Step 2: Use the fact that A, B, C are mutually exclusive and collectively exhaustive (sum = 1). So, p + 2p + 4p = 1 → 7p = 1 → p = 1/7.
- Step 3: Find specific probabilities. P(A) = 4/7, P(B) = 2/7, P(C) = 1/7.
- Step 4: Since A and B are mutually exclusive, P(A ∪ B) = P(A) + P(B) = 4/7 + 2/7 = 6/7.
💡 Why this matters: This example shows how to solve problems when the events are not equally likely but their relationships are known.
CONDITIONAL PROBABILITY
Conditional probability is the probability of an event being calculated after some additional information is received, which reduces the original sample space. The conditional probability of event A given that event B has occurred (and P(B) > 0) is defined as: P(A|B) = P(A ∩ B) / P(B) . Similarly, P(B|A) = P(A ∩ B) / P(A) if P(A) > 0. P(A|B) satisfies all basic axioms of probability.
🔑 Definition — Conditional Probability (P(A|B)): The probability that event A occurs given that event B has already occurred. It is valid only if P(B) > 0.
📐 Formula: P(A|B) = P(A ∩ B) / P(B) (where P(B) > 0) → The probability of A given B is the probability of both A and B happening divided by the probability of B.
📌 Example 1: Tossing a fair die. What is the probability of getting a 6 (event A) given that the die shows an even number (event B)?
- Step 1: The reduced sample space is B = {2, 4, 6}. There is one favorable outcome for A (6).
- Step 2: In the reduced sample space, P(A|B) = 1/3.
📌 Example 2: School records show 75% of students are from two-parent homes (B), and 20% are low-achievers from two-parent homes (A ∩ B). Find P(low achiever | two-parent home) = P(A|B).
- Step 1: Identify probabilities. P(B) = 0.75, P(A ∩ B) = 0.20.
- Step 2: Apply formula. P(A|B) = P(A ∩ B) / P(B) = 0.20 / 0.75 = 0.27.
MULTIPLICATION THEOREM OF PROBABILITY
The Multiplication Theorem is derived from the definition of conditional probability. By rearranging the formula P(A|B) = P(A ∩ B) / P(B), we get P(A ∩ B) = P(B) * P(A|B) . Similarly, P(A ∩ B) = P(A) * P(B|A) . This is the general rule of multiplication for any two events.
🔑 Definition — Multiplication Law: The rule that the probability that two events A and B will both occur is equal to the probability of one event multiplied by the conditional probability of the other given that the first has occurred.
📐 Formulas: P(A ∩ B) = P(A) * P(B|A) (if P(A) > 0) ; P(A ∩ B) = P(B) * P(A|B) (if P(B) > 0) → The probability of both A and B is the probability of the first event times the probability of the second event after the first has happened.
📌 Example: A box has 15 items (11 good, 4 defective). Two items are selected without replacement. Find P(first is good (A) AND second is defective (B)).
- Step 1: Find P(A). P(A) = 11/15.
- Step 2: Given A occurred, 14 items remain (10 good, 4 defective). Find P(B|A). P(B|A) = 4/14.
- Step 3: Apply the multiplication theorem. P(A ∩ B) = P(A) * P(B|A) = (11/15) * (4/14) = 44/210 = 0.16.
💡 Why this matters: The distinction between the two theorems is key: use the addition theorem for the probability of "either A or B" and the multiplication theorem for the probability of "both A and B."
⭐ Key Takeaways
The most critical concepts from this lecture are the formulas and distinctions between the addition and multiplication theorems. You must memorize the general addition law P(A∪B) = P(A) + P(B) – P(A∩B) and know when to simplify it to P(A) + P(B) for mutually exclusive events. The definition of conditional probability, P(A|B) = P(A∩B) / P(B), is fundamental, as it leads directly to the multiplication law P(A∩B) = P(A) * P(B|A). Finally, remember that the addition theorem answers "or" questions, while the multiplication theorem answers "and" questions, and both are applied based on the specific conditions and sample space given.
🧠 Quick Revision Questions
- What is the general formula for the Addition Law of probability, and when can it be simplified?
- Define conditional probability P(A|B) in terms of a joint and a marginal probability. What condition is required for it to be valid?
- Write down the two forms of the general Multiplication Theorem.
- In the "Three horses" example, why couldn't the probabilities be set directly to 1/3 each? How were they derived instead?
- A bag contains 5 red and 3 blue marbles. Two marbles are drawn without replacement. What is the probability the first is red and the second is blue?
📘 Lecture 21 — Independent and Dependent Events
📖 Overview: This lecture introduces the critical distinction between independent and dependent events in probability theory. It demonstrates how to determine independence through the Multiplication Theorem and illustrates the practical application of marginal probability using real-world birth statistics.
🗂️ Topics Covered
The lecture begins with an extended example applying both Addition and Multiplication Theorems of probability, then formally defines independent events and the special case of the Multiplication Theorem. It provides a detailed example with dice to demonstrate independence, examines a real birth statistics table to determine dependence between sex and stillbirth status, and concludes with the definition and computation of marginal probabilities from joint probability tables.
📝 Lecture Summary
Example Illustrating Addition and Multiplication Theorems
A bag contains 10 white and 3 black balls. Another bag contains 3 white and 5 black balls. Two balls are transferred from first bag and placed in the second, and then one ball is taken from the latter. We need the probability that it is a white ball.
Let A represent the event that 2 balls are drawn from the first bag and transferred to the second bag. A can occur in three mutually exclusive ways:
- A₁ = 2 white balls transferred
- A₂ = 1 white ball and 1 black ball transferred
- A₃ = 2 black balls transferred
P(A₁) = C(10,2)/C(13,2) = 45/78 P(A₂) = C(10,1)×C(3,1)/C(13,2) = 30/78 P(A₃) = C(3,2)/C(13,2) = 3/78
After transferring 2 balls, the second bag contains: i) If 2 white balls transferred: 5 white, 5 black → P(W/A₁) = 5/10 ii) If 1 white and 1 black transferred: 4 white, 6 black → P(W/A₂) = 4/10 iii) If 2 black balls transferred: 3 white, 7 black → P(W/A₃) = 3/10
Let W represent the event that a white ball is drawn from the second bag. P(W) = P(A₁∩W) + P(A₂∩W) + P(A₃∩W) P(A₁∩W) = P(A₁)P(W/A₁) = 45/78 × 5/10 = 15/52 P(A₂∩W) = P(A₂)P(W/A₂) = 30/78 × 4/10 = 2/13 P(A₃∩W) = P(A₃)P(W/A₃) = 3/78 × 3/10 = 3/260
P(W) = 15/52 + 2/13 + 3/260 = 59/130 = 0.45
Independent Events
Two events A and B in the same sample space S are defined to be independent (or statistically independent) if the probability that one event occurs is not affected by whether the other event has or has not occurred. That is: P(A/B) = P(A) and P(B/A) = P(B)
🔑 Definition — Independent Events: Two events A and B are independent if and only if P(A∩B) = P(A)P(B). This is known as the special case of the Multiplication Theorem of Probability.
Rationale
According to the multiplication theorem of probability: P(A∩B) = P(A)P(B/A). Putting P(B/A) = P(B), we obtain P(A∩B) = P(A)P(B).
The events A and B are defined to be dependent if P(A∩B) ≠ P(A) × P(B). This means the occurrence of one event in some way affects the probability of the other.
💡 Why this matters: Two events that are independent can never be mutually exclusive.
Example: Demonstrating Independence with Dice
Two fair dice, one red and one green, are thrown. Let A denote the event that the red die shows an even number. Let B denote the event that the green die shows a 5 or a 6.
The sample space S has 36 outcomes. P(A) = 3/6, P(B) = 2/6. The joint event A∩B contains 6 outcomes: (2,5), (4,5), (6,5), (2,6), (4,6), (6,6). P(A∩B) = 6/36.
P(A)P(B) = 3/6 × 2/6 = 6/36 = P(A∩B)
Therefore, events A and B are independent.
Example: Sex and Stillbirth Dependence
Table of births in England and Wales in 1956 (proportions):
| Liveborn | Stillborn | Total | |
|---|---|---|---|
| Male | .5021 | .0120 | .5141 |
| Female | .4750 | .0109 | .4859 |
| Total | .9771 | .0229 | 1.0000 |
Let M = male birth, S = stillbirth.
Conditional probability of stillbirth given male birth = P(S/M) = 8609/368490 = 0.0234 Conditional probability of stillbirth given female birth = 7796/348258 = 0.0224 Overall (unconditional) proportion of stillbirths = 16405/716740 = 0.0229
The conditional probability of stillbirth among boys is slightly higher than the overall proportion, while among girls it is slightly lower. Therefore, sex and stillbirth are statistically dependent — the sex of a baby affects its chance of being stillborn.
Marginal Probability
The probabilities that appear in the margins of a table are known as Marginal Probabilities.
🔑 Definition — Marginal Probability: A probability obtained by summing joint probabilities in a row or column of a contingency table.
From the table above:
- P(male birth) = 0.5141 (appears in the margin)
- P(female birth) = 0.4859
- P(live birth) = 0.9771
- P(stillbirth) = 0.0229
P(male birth) = P(male live-born) + P(male stillborn) = 0.5021 + 0.0120 = 0.5141
This follows the Addition Theorem of Probability for mutually exclusive events — joint probabilities in any row add up to yield the corresponding marginal probability.
📐 Formula: P(marginal event) = Σ joint probabilities in that row (or column)
Example: Conditional Probability Calculation
P(stillbirth/male birth) = P(male birth and stillbirth)/P(male birth) = 0.0120/0.5141 = 0.0233
⭐ Key Takeaways
The Multiplication Theorem for independent events states that P(A∩B) = P(A)P(B), which is the essential test for independence. Two events are independent if the occurrence of one does not affect the probability of the other; they are dependent otherwise. Importantly, independent events can never be mutually exclusive. Marginal probabilities are found by summing joint probabilities across rows or columns of a contingency table, and they represent unconditional probabilities of individual events. Conditional probabilities can be compared to marginal probabilities to determine independence — if P(A/B) ≠ P(A), the events are dependent.
🧠 Quick Revision Questions
- What is the mathematical condition that must hold for two events A and B to be considered independent?
- In the bag transfer example, what was the final probability of drawing a white ball from the second bag, and why was the Multiplication Theorem needed?
- Why can two independent events never be mutually exclusive? Explain the logical reasoning.
- In the birth statistics example, how did comparing conditional probabilities to the overall proportion reveal dependence between sex and stillbirth?
- How is a marginal probability computed from a table of joint probabilities, and which probability theorem justifies this computation?
📘 Lecture 22 — Bayes’ Theorem, Discrete Random Variable, Discrete Probability Distribution, Graphical Representation, Mean, Standard Deviation, Coefficient of Variation, Distribution Function
📖 Overview: This lecture introduces Bayes’ Theorem, a powerful extension of conditional probability for calculating “reverse” probabilities when events form a partition. It then transitions to the foundational concepts of discrete random variables, discrete probability distributions, their graphical representation, and key numerical measures (mean, standard deviation, coefficient of variation), concluding with the cumulative distribution function. These concepts are essential for modeling and analyzing random phenomena in statistics.
🗂️ Topics Covered
The lecture begins with Bayes’ Theorem for k mutually exclusive and exhaustive events, illustrated by a detailed example on car emission testing. It then defines a random variable, specifically a discrete random variable, and introduces the discrete probability distribution with its two properties. The graphical representation of a discrete probability distribution is shown via a line chart. The computation of the mean (expected value), standard deviation, and coefficient of variation for a discrete probability distribution is explained mathematically and with a worked example on flower petals. Another example on the sum of two dice illustrates finding probabilities from a distribution. Finally, the concept of the distribution function (cumulative distribution function) is defined and applied to the dice example.
📝 Lecture Summary
Bayes’ Theorem
This theorem deals with conditional probabilities in an interesting way. If events A1, A2… Ak form a partition of a sample space S (meaning they are mutually exclusive and exhaustive), and if B is any other event of S such that it can occur ONLY IF ONE OF THE Ai OCCURS, then for any i, the probability that Ai occurred given that B occurred is given by the formula below. In simpler terms, if you know that event B happened, but B could only be caused by one of several prior events (A1, A2, … Ak), Bayes’ Theorem tells you the probability that a particular prior event (Ai) was the cause.
🔑 Definition — Partition: A set of events A1, A2... Ak that are mutually exclusive (no two can occur at the same time) and exhaustive (their union is the whole sample space S).
📐 Formula (General): P(Ai | B) = [P(Ai).P(B|Ai)] / [ Σ (from i=1 to k) P(Ai).P(B|Ai) ] 📐 Formula (for k=2): P(A1 | B) = [P(A1).P(B|A1)] / [ P(A1).P(B|A1) + P(A2).P(B|A2) ] → Plain-English Meaning: The probability of a specific cause (Ai) after observing an effect (B) is the likelihood of that cause producing the effect, divided by the total likelihood of the effect being produced by any cause.
📌 Example: In a developed country, 25% of all cars emit excessive pollutants. When tested, 99% of cars that emit excessive pollutants fail, but 17% of cars that do not emit excessive pollutants also fail. What is the probability that a car that fails the test actually emits excessive pollutants?
- Step 1: Define Events. A1: Car emits excessive pollutants. A2: Car does not emit excessive pollutants. B: Car fails the test.
- Step 2: Identify Given Probabilities. P(A1) = 0.25, P(A2) = 1 - 0.25 = 0.75, P(B|A1) = 0.99, P(B|A2) = 0.17.
- Step 3: Apply Bayes’ Theorem (k=2). P(A1|B) = (0.25 * 0.99) / (0.25 * 0.99 + 0.75 * 0.17) = 0.2475 / (0.2475 + 0.1275) = 0.2475 / 0.3750
- Step 4: Result. P(A1|B) = 0.66. There is a 66% probability that a car which fails the test actually emits excessive pollutants. 💡 Why this matters: This example shows a classic “false positive” scenario. The test is very good at catching polluting cars (99% fail rate), but it also has a high rate of failing non-polluting cars (17%). Bayes’ Theorem corrects our intuition and gives the true probability, which is much lower than 99%.
Random Variable
Such a numerical quantity whose value is determined by the outcome of a random experiment is called a random variable. For example, if we toss three dice together and let X denote the number of heads, the random variable X consists of the values 0, 1, 2, and 3. In this example, X is a discrete random variable.
Discrete Probability Distribution
A probability distribution for a discrete random variable is a table or formula that lists all possible values the variable can take and their corresponding probabilities.
🔑 Properties of a Discrete Probability Distribution:
- 0 ≤ P(Xi) ≤ 1 for each value Xi.
- Σ P(Xi) = 1 (The sum of all probabilities must equal 1).
📌 Example: A biologist studying a flower species observes the number of petals on 1000 flowers.
- Given Frequency Distribution: [Table shows number of petals (3-9) and their frequencies (50, 100, 200, 300, 250, 75, 25)]. Because 1000 is a large number, the proportions f/Σf are regarded as probabilities.
- Discrete Probability Distribution: [Table shows the same petal values (x1=3 to x7=9) and their probabilities (0.05, 0.10, 0.20, 0.30, 0.25, 0.075, 0.025)]. The sum of probabilities is 1. This is a valid discrete probability distribution.
Graphical Representation of a Discrete Probability Distribution
This distribution can be represented in the form of a line chart. [The chart shows a line chart of the petal example, with the x-axis for “No. of Petals (x)” and the y-axis for “Probability P(x)”, showing a roughly symmetric distribution centered around 6]. This graph clearly shows that every discrete probability distribution has a central point and a spread, similar to a frequency distribution.
Mean, Standard Deviation and Coefficient of Variation of a Discrete Probability Distribution
Mean (Expected Value, μ or E(X)) The mean of a discrete probability distribution is given by the formula below. It is the long-run average value of the random variable.
📐 Formula: μ = E(X) = Σ [X * P(X)] → Plain-English Meaning: Multiply each value of X by its probability and sum up all those products.
📌 Example (Flower Petals): [Table shows x, P(x), and xP(x) columns]. μ = Σ xP(x) = 0.15 + 0.40 + 1.00 + 1.80 + 1.75 + 0.60 + 0.225 = 5.925. Interpretation: Considering a very large number of flowers, we would expect the average number of petals per flower to be 5.925 (approximately 6 petals). This is why the mean is technically called the expected value.
Standard Deviation (σ or S.D.(X)) The standard deviation measures the spread or dispersion of the probability distribution.
📐 Formula: σ = S.D.(X) = √ [ Σ (X²P(X)) – (Σ XP(X))² ] → Plain-English Meaning: This is the square root of the difference between the average of the squares and the square of the average.
📌 Example (Flower Petals): [Table shows x, P(x), xP(x), and x²P(x) columns].
- Σ xP(x) = 5.925
- Σ x²P(x) = 0.45 + 1.60 + 5.00 + 10.80 + 12.25 + 4.80 + 2.025 = 36.925
- σ = √ [36.925 – (5.925)²] = √ [36.925 – 35.106] = √(1.819) ≈ 1.3 [The graph of the petal distribution is shown again, with arrows indicating μ = 5.925 and σ = 1.3].
Coefficient of Variation (C.V.) 📐 Formula: C.V. = (σ / μ) * 100 → Plain-English Meaning: A relative measure of dispersion, expressed as a percentage. 📌 Example (Flower Petals): C.V. = (1.3 / 5.925) * 100 = 21.9%
📌 Example: Sum of Dots on Two Fair Dice.
- Part a: The sample space has 36 equally likely outcomes. The random variable X (sum of dots) can be 2, 3, …, 12. [The probability distribution table is constructed, e.g., P(X=2) = 1/36, P(X=3) = 2/36, P(X=7) = 6/36, P(X=12) = 1/36].
- Part b(i): Find P(a sum > 8) = P(X=9) + P(X=10) + P(X=11) + P(X=12) = 4/36 + 3/36 + 2/36 + 1/36 = 10/36.
- Part b(ii): Find P(5 < X ≤ 10) = P(X=6) + P(X=7) + P(X=8) + P(X=9) + P(X=10) = 5/36 + 6/36 + 5/36 + 4/36 + 3/36 = 23/36.
Distribution Function of a Discrete Random Variable
The distribution function (d.f.) or cumulative distribution function (cdf) of a random variable X, denoted by F(x), is defined by F(x) = P(X ≤ x). It gives the probability of the event that X takes a value LESS THAN OR EQUAL TO a specified value x.
📌 Example (Sum of Two Dice): [The table now has a third row for F(xi), which is the cumulative sum of f(xi)].
- F(5) = P(X ≤ 5) = 1/36 + 2/36 + 3/36 + 4/36 = 10/36. This means the probability of obtaining a sum of five or less is 10/36.
⭐ Key Takeaways
Bayes’ Theorem is a crucial method for revising probabilities when new information (event B) becomes available, and it requires that the prior events (Ai) form a partition. A discrete random variable takes a countable number of values, and its probability distribution must have all probabilities between 0 and 1 and sum to 1. The mean (or expected value) of a distribution is its center of gravity, computed as Σ xP(x), while the standard deviation, computed from Σ x²P(x), measures its spread. The cumulative distribution function, F(x) = P(X ≤ x), is a powerful tool for finding probabilities of ranges of values.
🧠 Quick Revision Questions
- What are the two defining properties of a discrete probability distribution?
- In Bayes’ Theorem, what does it mean for events A1, A2, … Ak to form a “partition” of the sample space?
- Calculate the expected value (mean) of the sum of dots when two fair dice are thrown.
- What is the formula for the standard deviation of a discrete probability distribution?
- For the sum of two dice, explain the difference between f(7), which is P(X=7), and F(7), which is P(X ≤ 7).