MKT611 — Final Term Summary (Lectures 23–45)
📘 Lecture 24 — Fieldwork
📖 Overview: This lecture covers the critical process of fieldwork in marketing research, emphasizing that poor management of this stage can neutralize all previous research efforts. It details the four main aspects of fieldwork—time schedule, budget, personnel, and performance measurement—and provides comprehensive guidance on selecting, training, and managing the field force to ensure high-quality data collection.
🗂️ Topics Covered
The lecture begins by defining fieldwork and its importance in data collection, then outlines the four main aspects: time schedule, budget, personnel, and performance measurement. It delves into details of creating realistic time schedules and budgets, then extensively covers personnel management including selection criteria, training requirements, probing techniques, and proper procedures for recording replies and terminating interviews.
📝 Lecture Summary
Fieldwork
Fieldwork is an important step in the research process equated to data collection and related activities such as contacting respondents, administering instruments, recording data, and returning data to a central location for processing. If it is poorly managed, all previous steps will be neutralized. Collection of data may be by telephonic, face-to-face, or mail methods, and the field design must be organized accordingly, including activities of supervision, monitoring, mailing, and observing all field operations. If fieldwork is not properly organized, it will result into many kinds of errors. The four main aspects of fieldwork are: time schedule, budget, personnel, and performance measurement.
💡 Why this matters: Proper fieldwork management is the foundation upon which the validity of all research findings rests; errors at this stage cannot be corrected later.
Time Schedule
When preparing a time schedule of activities, we must be realistic and allocate a reasonable amount of time to data collection within the total time budget. To make a time schedule, the planner should specify the beginning and ending of the project, sequence activities within the time frame, determine days needed to complete various activities, and prepare a Gantt chart.
Budget
Budget includes assignment of cost to various activities identified in the time schedule. Both the time schedule and budget are prepared together. The budget shows a detailed breakdown of the cost of all activities and includes some amount for unseen contingencies. It is reviewed and approved by some individuals in the organization. Various cost items include: wages and salaries, telephone expenses, material and supplies, salaries of field supervisor/s, reproduction expenses like photocopy, and miscellaneous.
Personnel
Personnel means human resources for the fieldwork, collectively called the field force. Management of personnel requires selection, training, supervising, authenticating, and evaluation of field workers. For selection, we have to develop Job Description and Job Specification. The characteristics of the field force should include: being healthy and having stamina to undertake strenuous work, having a pleasant appearance, being communicative with listening and speaking skills, being educated up to BA, and preferably having some experience. Inexperienced personnel result in excessive coding error, misreporting, and larger refusal rate.
Training of Field Force
After personnel have been recruited, they should be trained in the skills of data collection. Training is critical for fieldworkers and should be given in person at some central location. It should cover: initial contact, appointments, and opening remarks to seek permission and cooperation; training about observation; asking questions, which brings high dividends in eliminating potential bias; being thoroughly familiar with the questionnaire; observing the order of the questions; using exact wording; slow reading to help respondents understand the question; and repeating questions if not understood. Probing should be done carefully.
Probing Questions
Probing questions are used to motivate, clarify, enlarge, or explain the answers. Techniques include: repeating the question in the same words; repeating the reply verbatim and confirming; silent probe by unexpected pause and looks that should not be embarrassing; reassuring and boosting the respondent that there is no right or wrong, just what it means to him/her; and eliciting clarification by saying "I don't quite understand", "please tell me more", "what do you mean", "anything else", "any other reason?"
Recording Replies
Recording replies seems simple but mistakes are common. Unstructured questions must be recorded verbatim. Guidelines for recording include: record identification data like name, date, interviewer, and project in the beginning; record during the interview and not afterward; record in respondents' own words; do not summarize; include everything; include all probe comments; and repeat the response so that respondent listens and verifies.
Terminating Interview
Guidelines for closing the interview include: don't close before all information is gathered; leave the respondent with good feelings; and express your appreciation for his time and cooperation. In telephonic interviews: verify the telephone number once again; be courteous and polite; put a smile in your voice; speak clearly; don't sound bored; and close the interview properly.
⭐ Key Takeaways
The four critical aspects of fieldwork management are time schedule, budget, personnel, and performance measurement, and these must be carefully coordinated to avoid errors that can invalidate the entire research project. Proper selection of field force members requires specific characteristics including health, appearance, communication skills, education, and experience, as inexperienced personnel lead to excessive errors and refusal rates. Training must cover initial contact procedures, question administration with exact wording and proper order, and careful probing techniques. Interviewers must record replies verbatim during the interview using respondents' own words without summarizing, and must terminate interviews professionally by ensuring all information is gathered and leaving respondents with positive feelings.
🧠 Quick Revision Questions
- What are the four main aspects of fieldwork that must be managed to avoid research errors?
- What specific characteristics should a field force member possess according to the job specification?
- List at least five techniques that can be used for probing questions during interviews.
- What are the key guidelines for properly recording replies during data collection?
- How should an interviewer properly terminate an interview, and what special considerations apply to telephonic interviews?
📘 Lecture 25 — Data Analysis
📖 Overview: This lecture covers the critical processes of performance measurement of field force, data preparation, and data editing in marketing research. It emphasizes the importance of controlling non-sampling errors that occur during fieldwork, distinguishing between errors committed by fieldworkers and respondents, and providing practical methods to minimize these errors.
🗂️ Topics Covered
The lecture begins with performance measurement of the field force, including authentication and evaluation methods. It then transitions into data preparation and the principles of data analysis, followed by data editing procedures to fix common problems. A major portion is dedicated to non-sampling errors, categorized into fieldworker errors (intentional and unintentional) and respondent errors (intentional and unintentional), concluding with detailed strategies for controlling each type of error.
📝 Lecture Summary
Performance Measurement of the Field Force
Performance measurement and management are two important aspects of fieldwork. The supervisor should ensure that procedures are being followed; if not, provide additional training to the field workers. Guidelines include examining and editing collected questionnaires, ensuring the sampling plan is followed strictly (not based on convenience and accessibility), and controlling cheating and fake answers by calling respondents on the phone.
Authentication & Evaluation A good method to authenticate that the data were collected genuinely is to call 10 to 15% of respondents to validate. Ask their basic demographic data and cross-check. Also ask respondents about the length, quality, and their reaction to the interviewer. To evaluate the performance of each field worker, check: contact rate, response rates, refusal rate, cost incurred and time spent, quality of interviewing (including precision, ability to probe, ability to ask sensitive questions, interpersonal skills, and termination of interview), and quality of data (legibility, non-response, whether instructions were followed, and if answers were completely recorded).
Data Preparation
We should understand that data is recorded measures of phenomenon, but information is a body of facts in a format suitable for decision making. The purpose of research is to provide information. Data analysis is the process whereby data is converted into information. Understanding the principles of data analysis is important because it: leads the researcher to develop insight into information and data, helps avoid erroneous judgment and conclusions, and helps interpret the analysis of others.
💡 Why this matters: Knowledge and power of data analysis can constructively influence research design, but it cannot rescue or compensate for a study that was not well conceived. If research hypotheses were non-viable, sampling was inadequate, or fieldwork was sloppy, data analysis will not provide any remedy.
Data Editing
Although data analysis techniques are unique for each study, all studies require data editing. Data editing means identifying omissions, ambiguities, and errors in the responses and taking necessary corrective actions. The responsibility for these errors could lie with the interviewer, supervisor, or data analyst.
Problems identified during editing include: incorrect instructions by the interviewer; omissions and ambiguities (e.g., two boxes checked in an MCQ); inconsistencies (e.g., not married but having children); lack of cooperation (always agreeing); and ineligible respondents (e.g., age under 18). Alternatives for data editing include: approaching the respondent again for clarification, devising a new category, creating a category of "no answer" or "missing" , or not including a question or the whole questionnaire in the analysis.
Non Sampling Errors
All errors in a survey, except those related to the sampling plan and sampling size, are non-sampling errors. These include: non-response errors, data gathering errors, data handling errors, data analysis errors, interpretation errors, ambiguity in problem definition, and inappropriate wording of questions. The greatest potential for these errors is at the data collection stage or fieldwork. Such errors may occur due to the fieldworker or the respondents.
Fieldworker Errors These errors are committed by the person who administers the questionnaire or takes the interview. They may be intentional (committed deliberately) or unintentional (occur without willful intent).
- Intentional Fieldworker Errors: Cheating (e.g., intentionally falsifying responses if paid per completed interview), contacting a convenient interviewee instead of a genuine sample, completing fewer questionnaires than agreed, leading the respondent to a particular answer through wording or body language, and re-wording the question.
- Unintentional Fieldworker Errors: These arise from three sources: personal characteristics (accent, gender), misunderstanding of instructions about scales or recording responses, a gap between the educational levels of the research designer and the fieldworker, and miscellaneous sources like fatigue or monotony.
Respondent Errors These errors are committed by the person who fills the questionnaire or responds to an interview. They are also either intentional or unintentional.
- Respondents' Intentional Errors: Respondents willfully misrepresent themselves by telling a lie (due to privacy or embarrassment regarding income, age, etc.) or by not giving a response in whole or part (due to a busy schedule or privacy).
- Respondents' Unintentional Errors: Respondents believe an invalid response is true. Examples include answering without understanding (e.g., income with or without taxes), checking two answers instead of one, guessing due to low recall, attention loss, distraction, and respondent fatigue.
How to Control Field Errors
Precautions should be taken to minimize such errors, but they cannot be eliminated entirely. Several methods are suggested for controlling field errors.
Control of Intentional Fieldworker Errors:
- Supervision: Oversee the work, use central telephone monitoring to spot cheating, and reprimand the interviewer if needed.
- Validation: Validate the work by re-contacting about 10% of respondents, sometimes re-administering the instrument for comparison.
Control of Unintentional Fieldworker Errors:
- Selection and Training: Careful selection and training can best care for such errors.
- Orientation Session: Meetings of field workers with the supervisor to give instructions.
- Role Playing (Rehearsals): The supervisor or another person plays the role of respondent to check interviewing skills.
Control of Intentional Respondent Errors:
- Anonymity: Assure confidentiality of the respondent's name.
- Incentives: Give or promise gifts like t-shirts, ballpoints, cash, or diaries.
- Validation Checks: Confirm the response by cross-checking with other ways.
- Third Person Techniques: Design questions so they are indirect, in the name of a friend or neighbor.
Control of Unintentional Respondent Errors:
- Well Drafted Questionnaire: Clear instructions and examples remove misunderstandings.
- Reversal of Scale on Endpoints: Place negative adjectives on different sides (sometimes right, sometimes left) to compel thoughtful answers.
- Prompters: Use written or oral statements like "we are almost finished" to keep respondents on task and reduce fatigue.
💡 Why this matters: Good design of the questionnaire and technology can reduce errors of data collection in general.
⭐ Key Takeaways
Data analysis converts raw data into decision-making information, but it cannot fix a poorly designed study. All field measurement errors, from both fieldworkers and respondents, are non-sampling errors which must be actively controlled. Intentional errors like cheating can be managed through supervision and validation, while unintentional errors require better training, clear instructions, and well-drafted questionnaires. Authenticating 10-15% of respondents is a critical step to ensure data validity and to measure field force performance.
🧠 Quick Revision Questions
- What is the difference between "data" and "information" as defined in the lecture?
- What are the five alternatives for data editing when a problem is found in a questionnaire?
- List the three methods for controlling unintentional fieldworker errors.
- What is the "third person technique" and which type of error is it designed to control?
- Why is it important to reverse scale endpoints on a questionnaire?
📘 Lecture 26 — Data Analysis
📖 Overview: This lecture introduces the process of data analysis in marketing research, emphasizing that it is as important as other research phases. It covers the distinction between descriptive and inferential statistics, and explains basic data analysis techniques including frequency distribution, measures of central tendency (mean, median, mode), and their appropriate applications based on data type and measurement levels.
🗂️ Topics Covered
The lecture begins by positioning data analysis within the broader marketing research process, then categorizes statistical techniques into descriptive and inferential statistics. It explains how variable count and measurement levels influence technique selection. The main focus is on basic data analysis procedures: frequency distribution with examples, arithmetic mean calculation for interval data (both raw and grouped), median as a measure for ordinal data with odd/even observation calculations, and mode as the most frequent observation.
📝 Lecture Summary
Data Analysis Overview
Data analysis proceeds after data have been collected and checked for legibility, completeness, consistency, accuracy, and response classification. It is one of many activities in marketing research but not the most important aspect; rather, it is as important as other aspects such as good problem definition, appropriate study design, proper sampling procedures, well-designed and tested instruments of data collection, and well-monitored fieldwork. In data analysis, the researcher needs adequate knowledge of analysis techniques and must select techniques aligned with research objectives.
Categories of Statistical Techniques
Statistical techniques fall into two categories: descriptive statistics and inferential statistics. Descriptive statistics describe the characteristics of the sample under study, while inferential statistics help the researcher make inferences about the population from which the sample was drawn. Selection of statistical techniques also depends on how many variables are to be analyzed at the same time: uni-variable, bi-variable, and multi-variable analysis. Another relevant question is what level of measurement is available: nominal, ordinal, or interval.
Basic Data Analysis Procedures
Three basic data analysis procedures discussed are frequency distribution, cross-tabulation, and hypothesis testing.
Frequency Distribution
A frequency distribution simply reports the number of responses that each question received. When a market researcher is finding an answer to questions about a single variable, they start with frequency distribution. Absolute frequency and relative frequency (percentage) may be the subject of frequency distribution.
🔑 Definition — Frequency Distribution: Reports the number of responses each question received.
📌 Example 1 — Distribution of customers of brand X of the 1300cc car province wise (n=780):
- Islamabad: 034 (04.4%)
- Punjab: 305 (39.1%)
- Sindh: 230 (29.5%)
- NWFP: 148 (19.0%)
- Baluchistan: 063 (08.0%)
- Total: 780 (100%)
📌 Example 2 — Factors important in purchase of Refrigerator (n=300):
- Size: 183 (61%)
- Color, appearance: 90 (30%)
- Price: 57 (19%)
- Less consumption of electricity: 90 (30%)
- Size of freezer: 54 (18%)
- Self defrosting: 60 (20%)
- Ice maker, Water dispenser: 84 (28%)
- Brand name: 51 (17%)
📌 Example 3 — Monthly income of respondents (n=200):
- 0-5000: 15 (7.5%)
- 5000-8000: 22 (11.0%)
- 8000-11000: 42 (21.0%)
- 11000-14000: 60 (30.0%)
- 14000-17000: 33 (16.5%)
- 17000-20000: 28 (14.0%)
- Total: 200 (100%)
Mean
Mean is a measure of central tendency. There are different kinds of means, e.g., arithmetic mean, geometric mean, harmonic mean. In marketing research, the arithmetic mean (arithmetic average) is most widely used. It is the ratio between the sum of observations and number of observations. Mean is the sum of values divided by the sample size or number of observations.
📐 Formula: Mean = Sum of values / Sample size (number of observations)
📌 Example — Raw data (n=10): Family # Monthly Income ('000 rupees')
- 1: 41
- 2: 37
- 3: 28
- 4: 51
- 5: 50
- 6: 36
- 7: 30
- 8: 41
- 9: 47
- 10: 29 Sum: 390 Mean Income = 390/10 = 39 (or rupees 39,000)
If the data are interval grouped into categories and classes, mean is calculated by taking the midpoints of the categories. The midpoint of a category is multiplied by the frequency in the category, summing up the multiplied values and dividing by the sum of frequencies.
📌 Example — Grouped data (n=200): Monthly Income (Rs.) | # of Respondents (f) | Midpoint (x) | fx 0-5000 | 15 | 2500 | 37,500 5000-8000 | 22 | 6500 | 143,000 8000-11000 | 42 | 9500 | 399,000 11000-14000 | 60 | 12,500 | 750,000 14000-17000 | 33 | 15,500 | 511,500 17000-20000 | 28 | 18,500 | 518,000 Total | 200 | | 2,359,000 Mean income = 2,359,000 / 200 = Rs. 11,795
Median (for Ordinal Data)
In ordinal data, median is the measure of central tendency. Median is the midpoint of the frequencies. As per this measure, 50% of the observations are above median value whereas 50% are below.
🔑 Definition — Median: The midpoint of frequencies such that 50% of observations are above and 50% are below.
If the number of observations is odd, the middle observation is the median.
📌 Example — Odd number of observations: Sales values by a salesman in a month: Rs. 125,000, Rs. 175,000, Rs. 325,000, Rs. 77,000, Rs. 180,000. Arrange in ascending order: Rs. 77,000 Rs. 125,000 Rs. 175,000 Rs. 180,000 Rs. 325,000 The median (middle value) is Rs. 175,000.
If the observations are even numbers, the median would be the arithmetic mean of the middlemost two values.
📌 Example — Even number of observations: Sales order values in a month: Rs. 25,000, Rs. 178,000, Rs. 510,000, Rs. 450,000, Rs. 275,000, Rs. 370,000. Arrange in descending order: Rs. 510,000 Rs. 450,000 Rs. 370,000 Rs. 275,000 Rs. 250,000 Rs. 178,000 Middlemost two values: Rs. 370,000 and Rs. 275,000 Median = (370,000 + 275,000) / 2 = 645,000 / 2 = Rs. 322,500
Mode
Mode is still another measure of central tendency. This is the observation (x) where the frequency is maximum.
🔑 Definition — Mode: The observation where the frequency is maximum.
📌 Example — Sales returns by model: Model # | 1 | 2 | 3 | 4 | 5 | 6 | 7
of returns | 6 | 12 | 8 | 2 | 1 | 7 | 2
The mode of sales return is Model #2 because the number of returns was 12 (maximum frequency).
⭐ Key Takeaways
Data analysis is as important as other marketing research phases and requires selecting appropriate techniques based on research objectives, number of variables, and measurement levels. The key distinction is between descriptive statistics (describing sample characteristics) and inferential statistics (making population inferences). For single-variable analysis, frequency distribution reports response counts and percentages. The arithmetic mean is the most widely used measure of central tendency for interval data, calculated as sum of values divided by number of observations; for grouped data, midpoints are used. Median is appropriate for ordinal data and is the midpoint of frequencies, calculated differently for odd (middle value) versus even (average of two middle values) numbers of observations. Mode identifies the most frequent observation in a dataset.
🧠 Quick Revision Questions
- What are the two main categories of statistical techniques in marketing research, and how do they differ?
- How do you calculate the mean for grouped interval data, and what is the formula?
- What is the difference between calculating median for an odd versus even number of observations?
- In the example of factors important in refrigerator purchase, what was the most frequently cited attribute and its percentage?
- What measure of central tendency is most appropriate for ordinal data, and why?
📘 Lecture 27 — Measure of Dispersion
📖 Overview: This lecture moves beyond measures of central tendency to explore measures of dispersion, which describe the variation or spread in data. Understanding dispersion is crucial for market researchers because two datasets can have identical means but very different distributions, affecting business decisions like distribution planning. The lecture covers range, variance, standard deviation, and coefficient of variation, including how to compute and interpret them.
🗂️ Topics Covered
The lecture begins by explaining why measures of dispersion are needed, using an example of two household groups with identical mean tea spending but different spreads. It defines Range as the simplest measure, then explains Variance as the mean of squared deviations from the mean, and Standard Deviation as its square root. A detailed example using ABC Real Estate shows how to compute sample variance and standard deviation step-by-step. The Rules of Standard Deviation for normal distributions are presented, followed by a bank waiting-time example that calculates all measures and evaluates a manager's statement. Finally, the Coefficient of Variation (CV) is introduced as a relative measure of dispersion for comparing datasets with different magnitudes, illustrated with a stock investment example.
📝 Lecture Summary
Why Measure Dispersion?
Measures of location (mean, mode, median) describe central tendency, but market research also needs to know the variation or scatter among values. Two variables can have identical central tendencies but very different spread. For example, spending on tea per annum for two groups of households: Group A (3000, 1500, 2100, 1800, 2800) and Group B (3600, 972, 1319, 2700, 2609) both have a mean of Rs. 2240, but their variation is not the same. This information is important for a tea manufacturer making a distribution plan.
Range
The range is the difference between the largest and smallest values in the data. In the previous example, the range of Group A is Rs. 3000 – 1500 = 1500, while the range of Group B is Rs. 3600 – 972 = 2628.
Variance and Standard Deviation
The distance between the mean and an observed value is the deviation from the mean. When we square these deviations and find the mean of these squared deviations, this is called variance. When data are scattered to a great extent, variance is large; if data are clustered around the mean, variance is small.
🔑 Definition — Sample Variance (s²): It is found with the formula ( s^2 = \frac{\sum (x - \bar{x})^2}{n - 1} ), where the sum of square deviations for the value from the mean ( \bar{x} ) is divided by ( n - 1 ).
🔑 Definition — Standard Deviation of Sample (s): It is the square root of the sample variance. The formula is ( s = \sqrt{\frac{\sum (x - \bar{x})^2}{n - 1}} ).
📌 Example — ABC Real Estate in Karachi wants to know how long it takes to sell listed homes. The director took a sample of 10 homes and recorded weeks to sell: 21, 6, 9, 23, 1, 10, 8, 11, 5, 7.
| Home | Period x | (x - (\bar{x})) | (x - (\bar{x}))² |
|---|---|---|---|
| 1 | 21 | 10.9 | 118.81 |
| 2 | 6 | -4.1 | 16.81 |
| 3 | 9 | -1.1 | 1.21 |
| 4 | 23 | 12.9 | 166.41 |
| 5 | 1 | -9.1 | 82.81 |
| 6 | 10 | -0.1 | 0.01 |
| 7 | 8 | -2.1 | 4.41 |
| 8 | 11 | 0.9 | 0.81 |
| 9 | 5 | -5.1 | 26.01 |
| 10 | 7 | -3.1 | 9.61 |
| n=10 | (\sum x=101) (\bar{x}=101/10=10.1) | (\sum = 426.9) |
( s = \sqrt{\frac{426.9}{10 - 1}} = \sqrt{\frac{426.9}{9}} = \sqrt{47.433} \approx 6.89 ) weeks.
Rules of Standard Deviation
For a normal distribution: 68% of the population falls within ±1 standard deviation. 95% of the population falls between ±2 standard deviations. 99.7% (virtually all) of the population falls between ±3 standard deviations.
📌 Example — A bank branch records waiting time (in minutes) during lunch hour for 12 customers: 4, 5, 3, 5, 6, 2, 7, 2, 4, 3, 2.
Solution:
| Customer # | Waiting time x | (x - (\bar{x})) | (x - (\bar{x}))² |
|---|---|---|---|
| 1 | 4 | 0 | 0 |
| 2 | 5 | 1 | 1 |
| 3 | 3 | -1 | 1 |
| 4 | 5 | 1 | 1 |
| 5 | 5 | 1 | 1 |
| 6 | 6 | 2 | 4 |
| 7 | 2 | -2 | 4 |
| 8 | 7 | 3 | 9 |
| 9 | 2 | -2 | 4 |
| 10 | 4 | 0 | 0 |
| 11 | 3 | -1 | 1 |
| 12 | 2 | -2 | 4 |
| n=12 | (\sum x=48) (\bar{x}=48/12=4) | (\sum = 1.82) |
Mean = 4 minutes; Median: arranging in descending order (7,6,5,5,5,4,4,3,3,2,2,2), median = (4+4)/2 = 4; Mode = 5; Variance (s²) = 3.312; s = 1.82. The manager's statement "almost certainly not more than six minutes" is correct because the range of ± one standard deviation is (4-1.82) to (4+1.82) = 2.18 to 5.82 minutes. Thus, majority of people wait not more than 5.82 minutes, which supports the manager's claim.
💡 Why this matters: The standard deviation rule allows managers to make probabilistic statements about expected customer experience, such as waiting times, which is valuable for service quality management.
Coefficient of Variation
The coefficient of variation (CV) is the ratio of the standard deviation to the mean, expressed as a percentage: ( CV = \frac{s}{\bar{x}} \times 100 ). CV is expressed in percentage and is useful when the variable is measured on a ratio scale. It is a useful measure of relative dispersion when means are positive, allowing comparison of sets of numbers with different magnitudes.
📌 Example — The standard deviation of closing prices of two shares X and Y were Rs. 5 and Rs. 50, and mean closing prices during a week were Rs. 10 and Rs. 1000 respectively.
If we look only at standard deviation, we might invest in share X because it has less volatility. But using CV:
( CV_x = \frac{5}{10} \times 100 = 50% )
( CV_y = \frac{50}{1000} \times 100 = 5% )
Now we change the decision in favor of Y, as fluctuation in Y's prices is much lesser than share X. The coefficient of variation is a good measure for comparing riskiness in this case.
⭐ Key Takeaways
The most critical concepts from this lecture are: (1) Range is the simplest measure of dispersion but only uses two extremes, while Variance and Standard Deviation measure the average squared deviation from the mean and its square root. (2) The standard deviation for a sample is computed by dividing the sum of squared deviations by ( n-1 ), not ( n ). (3) For normally distributed data, approximately 68%, 95%, and 99.7% of values fall within ±1, ±2, and ±3 standard deviations of the mean, respectively. (4) Coefficient of Variation is essential when comparing variability between datasets with different units or vastly different means (e.g., comparing stock price risk). (5) Businesses use dispersion measures to assess risk, consistency, and to make data-driven decisions beyond just averages.
🧠 Quick Revision Questions
- What is the difference between variance and standard deviation?
- In the ABC Real Estate example, what was the mean and standard deviation of the time to sell homes?
- According to the rules of standard deviation, what percentage of the population falls within ±2 standard deviations of the mean?
- Why is the coefficient of variation preferred over the standard deviation when comparing the riskiness of two stocks with very different average prices?
- In the bank waiting time example, how did the manager justify his statement that customers would "almost certainly" wait not more than six minutes?
📘 Lecture 28 — Charts and Graphs
📖 Overview: This lecture introduces the use of visual representations in marketing research data analysis, focusing on the three most common types of charts: pie charts, line charts, and bar charts. It matters because charts and graphs allow researchers to present complex data in an easily digestible and interpretable format at a glance.
🗂️ Topics Covered
The lecture begins by stating that data can be shown by charts and graphs in addition to text and tables. It then covers three main types of charts: the Pie Chart, which is used for representing categorical data as parts of a whole; the Line Chart (or Graph), which is used to show dynamic changes over time; and the Bar Chart, which uses bars to compare the magnitudes of different categories. Variations of the bar chart, such as using pictures instead of bars and the grouped bar chart, are also discussed.
📝 Lecture Summary
Charts and Graphs
Data can be shown by charts and graphs in addition to the text and tables. Although the text and tables are useful for explaining and interpretation, charts and graphic illustration may add to the value of information because they provide information to the reader in a glance. Due to availability of computer software, it is now very easy and quick to make these charts and graphs. There are many kinds of charts that can be used in data analysis, but we will mention only three types which are more common: Pie chart, Line chart (graph), Bar chart.
Pie Chart
A pie chart is probably the most familiar chart used in representing quantitative data. A pie chart is simply a circle divided into sections, with each of the sections representing a portion of the total. As the sections are presented as part of the whole, pie chart is particularly more useful in showing relative size. Pie chart is prepared for categorical data.
🔑 Definition — Pie Chart: A circle divided into sections, with each section representing a portion of the total. 📐 Formula: The pie is based on 360 degrees, and each slice’s angle is calculated based on its percentage of the whole. 📌 Example: A pie chart showing the share of market of different brands of mobile phone sets, where Brand A has 35%, Brand B has 31%, Brand C has 20%, and Brand D has 14%.
Pie chart is based on the fact that the circle has 360 degrees. The pie is divided into different slices according to percentage in the category. Different colors can be used. To show emphasis on some category of data, the relevant portion of the pie can be lifted (it called exploding the pie). A pie chart can have many sections or slices, but it is recommended that no more than six sections should be generated in a pie.
📌 Example: In order to determine the significant source of business, WR hotel examined the check-in cards of 1000 customers randomly. They found the following break-up of data: Individual travelers 230, Tour groups 125, Business travelers 378, Government officials 143, Others 124. The solution involves calculating relative frequencies: Individual travelers (23.0%), Tour groups (12.5%), Business travelers (37.8%), Government officials (14.3%), Others (12.4%). 💡 Why this matters: This allows the hotel to see at a glance that Business travelers are their most significant customer source.
Bar Chart
In a bar chart, each category is depicted by a bar, vertical or horizontal. The length or height of the bar shows the frequency or percentage of observations falling into a category. The lengths or heights of different bars allow the user to compare the magnitudes of different categories easily.
📌 Example: The expansion of a bank in terms of opening of new branches from 2001 to 2007 was as follows: 2001 – 15, 2002 – 28, 2003 – 14, 2004 – 21, 2005 – 19, 2006 – 23, 2007 – 25. This data can be represented on a bar chart with years on the X-axis and the number of branches on the Y-axis.
Variation in Bar Chart
The bar chart has variations. Pictures can be used instead of bars, e.g., people for population, pictures of cars for automobile production, and piles of 1000 rupees note for sales. Another variation of bar chart used frequently is the grouped bar chart where more than one category can be captured and can be compared side by side in different colours.
📌 Example: A survey of retail sales of split air conditioners of 3 brands S, T, Y from 2000 to 2007, revealed the following values in million of rupees: Year 2000: S=4.3, T=3.2, Y=5.2 Year 2001: S=4.8, T=4.2, Y=6 Year 2002: S=6, T=5, Y=6.2 Year 2003: S=5.8, T=3.1, Y=6.8 Year 2004: S=7.2, T=3.7, Y=7.5 Year 2005: S=4.6, T=4.1, Y=6 Year 2006: S=6, T=5.1, Y=7 Year 2007: S=7.5, T=5.4, Y=8.5
A line chart with three different lines for S, T, and Y can be used, or a grouped bar chart with three bars side-by-side for each year. 💡 Why this matters: These variations allow for more engaging and comparative visualizations of data.
Line Chart or Graph
The pie chart is a one-scale chart. It is best used for static comparison, that is, the phenomenon at one time. The bar chart is having two dimensions, one of which usually is the time. It shows dynamic relation of the changes with time such as time series fluctuations. In the chart, the X-axis represents time and the Y-axis, the values of the variables. More than one variable can be plotted on the same graph but each variable is represented by different lines in different colors or form (dashes or dots) with explanations in the legend at the bottom of the graph.
📌 Example: Using the same data for brands S, T, and Y from the bar chart variation example, a line chart shows the trend of sales for each brand over the years 2000-2007.
⭐ Key Takeaways
For exam success, a student must remember that pie charts are best for showing parts of a whole (relative size) with categorical data, and it is recommended to have no more than six slices. Line charts are specifically used to show trends and dynamic changes over time, with time on the X-axis. Bar charts are for comparing magnitudes across categories, and they have variations like using pictures or creating grouped bar charts to compare multiple variables side-by-side. Finally, computer software makes the creation of these charts quick and easy, but the analyst must choose the correct chart type for the data and the message.
🧠 Quick Revision Questions
- What is the primary purpose of a pie chart, and what type of data is it best suited for?
- When would you choose a line chart over a bar chart to represent data?
- What is a "grouped bar chart" and what advantage does it offer over a standard bar chart?
- How is a pie chart's "exploding" feature used to improve data presentation?
- What is the recommended maximum number of sections (slices) a pie chart should have, and why?
📘 Lecture 29 — Hypothesis Testing
📖 Overview: This lecture introduces hypothesis testing as a key component of data analysis in marketing research. It covers the formulation of null and alternative hypotheses, the distinction between one-tailed and two-tailed tests, and the fundamental steps involved in conducting a statistical hypothesis test.
🗂️ Topics Covered
This lecture begins by defining a hypothesis as an educated guess about relationships between variables, which must be empirically tested and formally stated. It then explains the null and alternative forms of hypothesis, followed by the concepts of one-tailed and two-tailed tests with specific examples related to marketing metrics like defective parts and sales returns. Finally, it outlines the step-by-step process of hypothesis testing and mentions the selection of appropriate statistical techniques based on research objectives.
📝 Lecture Summary
Data analysis
Hypothesis is an educated guess, a tentative statement about the relationship between two or more variables. It is to be empirically tested and be stated before the marketing project begins. Hypotheses must be formally stated, as they are focal points for researchers in marketing. Hypotheses may be in operational (general) terms or null and alternative forms.
Null and Alternative Form of Hypothesis
Testing of hypothesis usually begins with stating the hypothesis in a null and alternative form. For example, we might want to see whether mean age of a class of consumer is 30 years. 🔑 Definition — Null hypothesis (H₀): A statement that there is no effect or no difference, serving as the default assumption to be tested. It is written as Ho: μ = 30. 🔑 Definition — Alternative hypothesis (H₁): A statement that contradicts the null hypothesis, indicating the presence of an effect or difference. It can be written as H₁: μ ≠ 30 or H₁: μ > 30. 📐 Formula: Ho: μ = 30 and H₁: μ ≠ 30 → The null states the mean is exactly 30, while the alternative states the mean is not equal to 30. 📌 Example: For a consumer class mean age, the null hypothesis is Ho: μ = 30, and the two-tailed alternative is H₁: μ ≠ 30.
One tailed and two tailed test
The previous example of alternative hypothesis is two-tailed as we will reject the null hypothesis if the mean age was lesser or greater than 30. 🔑 Definition — One-tailed test: A test where the alternative hypothesis specifies a direction (greater than or less than a value). Another alternative hypothesis in this situation could be that the mean age is greater than 30. In this case we would write null and alternate form of hypothesis as: Ho: μ = 30 H₁: μ > 30 Here the hypothesis is one-tailed as we have a specific direction in mind for the alternative hypothesis. We can also phrase the null hypothesis to cover a range of values. For example, Ho: μ ≤ 30 Which implies an alternative hypothesis H₁: μ > 30 Here again one-tailed test would apply. The researcher should be careful to phrase the alternative hypothesis in a way as to accept the alternative hypothesis that is of real interest if null hypothesis is rejected.
One-Tailed or Two-Tailed Test
- Defective parts are more than 2% → One-tailed
- Sales returns are less than 4% p.m. → One-tailed
- Within one per cent of the mean → Two-tailed
- Between 5% and 6% → Two-tailed We use one and two-tailed concept when we look critical values in the probability tables. We should have this concept clear so that we can look into the appropriate table.
Steps in Hypothesis Testing
Following steps are usually followed in hypothesis testing:
- Formulate a null and alternative hypothesis
- Specify the significance level
- Select the appropriate statistical technique according to the nature and type of data collected
- Perform the statistical test applying the technique above
- Look for the value of test statistics (critical value) in the relevant standard normal table on the confidence level as specified in step #2 above
- Compare the value of statistics as calculated in step #4 with critical value and accept or reject the hypothesis
- At the end researcher draws conclusion.
Different Statistical Tests
Which statistical technique or test to select for our analysis or hypothesis testing depends on our objectives of research project. Objectives may be translated into research questions and/or research hypotheses. There are many statistical tests and techniques which are used in data analysis in marketing research. Some of them are described in the next lectures in detail. 💡 Why this matters: Choosing the wrong statistical test can lead to incorrect conclusions, making it essential to align the test with the research objectives and data type.
⭐ Key Takeaways
Hypothesis testing begins with formally stating a null hypothesis (H₀) and an alternative hypothesis (H₁), which must be empirically tested. The key distinction between one-tailed and two-tailed tests depends on whether the alternative hypothesis specifies a direction (e.g., greater than) or simply a difference. The process involves seven steps: formulating hypotheses, setting a significance level, selecting and performing a statistical test, comparing the calculated test statistic with a critical value, and drawing a conclusion. The choice between one-tailed and two-tailed tests is critical when determining critical values from probability tables.
🧠 Quick Revision Questions
- What is the difference between a null hypothesis and an alternative hypothesis?
- When would you use a one-tailed test instead of a two-tailed test?
- What are the seven steps in hypothesis testing?
- In the example "Sales returns are less than 4% p.m.", is this a one-tailed or two-tailed test?
- Why must a hypothesis be stated before the marketing research project begins?
📘 Lecture 30 — Chi-Square Test
📖 Overview: This lecture introduces the Chi-Square test, a statistical method used to determine whether there is a systematic association between two variables under investigation. It explains how cross-tabulation works and demonstrates the application of Chi-Square statistics through both univariate and bivariate data examples.
🗂️ Topics Covered
The lecture covers the Chi-Square test as a test of association, explaining null and alternative hypotheses, observed and expected frequencies, and cross tabulation. It then examines univariate data with two detailed examples — one analyzing color preferences for plastic products and another evaluating TV picture quality among different brands — showing step-by-step calculations and hypothesis testing procedures.
📝 Lecture Summary
Chi –Square Test
It is a test of association. It determines whether there is a systematic association in two variables under investigation. The null hypothesis would be that there is no association (or dependence) between the variables and accordingly an alternative hypothesis will be framed. In this test, observed and expected frequencies are computed and cross tabulation is done.
Cross Tabulation
Cross tabulation shows two or more variables at a time. It merges responses with regard to two or more variables in the same table. Frequencies of one variable are cross tabulated and distributed according to other categories.
📌 Example: Gender of students who visited Lahore zoo:
| Boys | Girls | Total | |
|---|---|---|---|
| Frequent visitors | 45 | 20 | 65 |
| Less frequent visitors | 23 | 47 | 70 |
| Total | 68 | 67 | 135 |
How Many Variables?
We may have one or two variables to test through Chi-Square. Thus the data are called univariate and bivariate respectively. Cross tabulation is done accordingly.
Univariate Data – (Example # 01)
A plastic company sells its products in three primary colors: yellow, blue & red. The marketing manager thinks that customers do not have any color preference. Manager set up a test where 240 purchasers were provided equal opportunity to buy any of the color. Results: 40 bought blue, 120 bought red and 80 yellow.
🔑 Hypothesis:
- H₀: There is no color preference of purchaser of plastic among three colors
- H₁: There is a color preference
📐 Formula: χ² = Σ [(fo - fe)² / fe] → Sum of (observed frequency minus expected frequency squared divided by expected frequency)
📌 Example Calculation:
- Yellow: fo=80, fe=80, (fo-fe)=0, (fo-fe)²=0, value=0
- Blue: fo=40, fe=80, (fo-fe)=-40, (fo-fe)²=1600, value=20
- Red: fo=120, fe=80, (fo-fe)=40, (fo-fe)²=1600, value=20
- Total χ² = 40
Degrees of freedom: V = k-1 = 3-1 = 2 Critical value at 0.05 level: 5.99
Since 40 > 5.99, H₀ is rejected and H₁ accepted. Red is most popular color, Blue is least.
Univariate Data (Example # 02)
A research company wanted to find out the best quality of picture in various brands of TV. 6 color TV sets were placed with brand names covered. 1302 people gave their opinion.
| TV brand | Number labeling as best picture |
|---|---|
| M | 675 |
| N | 280 |
| O | 81 |
| P | 119 |
| Q | 123 |
| R | 22 |
Test: Null hypothesis at α = 0.05 that all brands have same picture quality.
📌 Analysis: A contingency table is prepared showing expected values. Each expected value = 1302/6 = 217
📐 Critical value: With v = k-1 = 6-1 = 5 degrees of freedom at α = 0.05, the critical χ² value = 11.07
📌 Result: Calculated χ² = 1331.39, which is greater than critical value 11.07. Hence null hypothesis is rejected.
💡 Why this matters: When the calculated Chi-Square value far exceeds the critical value, it strongly indicates that the observed differences are statistically significant and not due to random chance.
⭐ Key Takeaways
The Chi-Square test determines whether there is a systematic association between variables by comparing observed frequencies with expected frequencies under the null hypothesis of no association. Cross tabulation organizes data by merging responses across two or more variables simultaneously. For univariate data, degrees of freedom are calculated as k-1 where k is the number of categories. A calculated Chi-Square value exceeding the critical value leads to rejection of the null hypothesis, indicating significant differences exist among categories. Expected frequencies are computed by dividing the total number of observations by the number of categories when testing for equal distribution.
🧠 Quick Revision Questions
- What is the null hypothesis in a Chi-Square test of association?
- How is cross tabulation defined and what purpose does it serve?
- In Example #01 (color preference), why was the expected frequency 80 for all three colors?
- What is the formula for degrees of freedom in univariate Chi-Square analysis?
- Why was the null hypothesis rejected in Example #02 (TV picture quality) when the calculated value was 1331.39 and the critical value was 11.07?
📘 Lecture 31 — Bivariate Chi-Square
📖 Overview: This lecture introduces the Chi-Square test of independence for analyzing relationships between two categorical variables. Using real hotel guest satisfaction data and education-confidence examples, it demonstrates how to compute expected frequencies, test statistics, and interpret results to determine whether variables are related or independent.
🗂️ Topics Covered
The lecture covers the Bivariate Chi-Square test for contingency tables, including formulation of null and alternative hypotheses, computation of observed and expected frequencies, degrees of freedom determination, calculation of the chi-square test statistic, comparison with critical values, and interpretation of results through two worked examples involving hotel guest reasons for not returning and education-confidence relationships.
📝 Lecture Summary
Data analysis
The lecture presents a survey on hotel guest satisfaction where respondents from three hotels (A, B, and C) indicated they were not likely to return and provided their primary reason. The reasons included room rent (78 guests), location (70 guests), room space (42 guests), and other reasons (39 guests). Across hotels, 109 guests were from Hotel A, 42 from Hotel B, and 78 from Hotel C who were not planning to return.
The null hypothesis (H₀) states there is no relationship between the primary reason for not returning and the hotels. The alternative hypothesis (H₁) states there is a relationship between these two variables.
Observed Frequencies
The contingency table presents the joint tallies of sampled guests with respect to primary reason for not returning and hotel property:
| Reason for not Returning | Hotel A | Hotel B | Hotel C | Total |
|---|---|---|---|---|
| Room rent | 40 | 14 | 24 | 78 |
| Location | 32 | 12 | 26 | 70 |
| Room space | 19 | 8 | 15 | 42 |
| Other | 18 | 8 | 13 | 39 |
| Total | 109 | 42 | 78 | 229 |
To test the null hypothesis of independence, the chi-square test statistic is computed using the formula:
📐 Formula: χ² = Σ (f₀ - fₑ)² / fₑ → Where f₀ = observed frequency in each cell, fₑ = expected frequency in each cell, and the sum is taken over all cells
Computed the Expected Frequencies
The expected frequency in a cell is the product of its row total and column total divided by the overall sample size:
🔑 Definition — Expected frequency (fₑ): fₑ = (row total × column total) / n where row total = sum of all frequencies in a row, column total = sum of all frequencies in a column, n = overall sample size
| Reason for not Returning | Hotel A | Hotel B | Hotel C | Total |
|---|---|---|---|---|
| Room rent | 37.12 | 14.30 | 26.56 | 78 |
| Location | 33.31 | 12.83 | 23.84 | 70 |
| Room space | 19.99 | 7.7 | 14.30 | 42 |
| Other | 18.56 | 7.15 | 13.28 | 39 |
| Total | 109 | 42 | 78 | 229 |
📐 Formula: Degrees of freedom = (r - 1)(c - 1) → For an r×c contingency table, where r = number of rows and c = number of columns
The chi-square test statistic is computed as shown:
| CELL | f₀ | fₑ | (f₀-fₑ) | (f₀-fₑ)² | (f₀-fₑ)²/fₑ |
|---|---|---|---|---|---|
| Room rent/Hotel A | 40 | 37.12 | 2.88 | 8.29 | 0.22 |
| Room rent/Hotel B | 14 | 14.30 | -0.3 | 0.09 | 0.006 |
| Room rent/Hotel C | 24 | 26.56 | -2.56 | 6.55 | 0.25 |
| Location/Hotel A | 32 | 33.31 | -1.31 | 1.71 | 0.05 |
| Location/Hotel B | 12 | 12.83 | -0.83 | 0.688 | 0.05 |
| Location/Hotel C | 26 | 23.84 | 2.16 | 4.66 | 0.19 |
| Room space/Hotel A | 19 | 19.99 | -0.99 | 0.98 | 0.04 |
| Room space/Hotel B | 8 | 7.7 | 0.3 | 0.09 | 0.01 |
| Room space/Hotel C | 15 | 14.30 | 0.7 | 0.49 | 0.03 |
| Other/Hotel A | 18 | 18.56 | -0.56 | 0.31 | 0.01 |
| Other/Hotel B | 8 | 7.15 | 0.85 | 0.72 | 0.100 |
| Other/Hotel C | 13 | 13.28 | -0.28 | 0.07 | 0.005 |
| Total | 0.98 |
Using a level of significance of α = 0.05, the computed test statistic χ² = 0.98 is less than 12.59, the upper-tail critical value from the chi-square distribution with (4-1)(3-1) = 6 degrees of freedom. Therefore, the null hypothesis of independence is accepted.
💡 Why this matters: When χ² is less than the critical value, we conclude there is no significant relationship between the variables — in this case, the primary reason for not returning does not differ significantly across hotels.
Another Example of Chi-Square
The second example examines the relationship between education level and confidence in television programs:
| Education | Under Metric | Upto B.A. | M.A. & Above | Total |
|---|---|---|---|---|
| A great extent | 95 | 57 | 39 | 191 |
| Some extent | 272 | 274 | 214 | 760 |
| Very little | 140 | 163 | 148 | 451 |
| Total | 507 | 494 | 401 | 1402 |
The chi-square test statistic computation:
| Cell | f₀ | fₑ | Contribution (f₀-fₑ)²/fₑ |
|---|---|---|---|
| Cell₁,₁ | 95 | 69.1 | 9.71 |
| Cell₁,₂ | 57 | 67.3 | 1.58 |
| Cell₁,₃ | 39 | 54.6 | 4.46 |
| Cell₂,₁ | 272 | 274.8 | 0.03 |
| Cell₂,₂ | 274 | 267.8 | 0.14 |
| Cell₂,₃ | 214 | 217.4 | 0.05 |
| Cell₃,₁ | 140 | 163.1 | 3.27 |
| Cell₃,₂ | 163 | 158.9 | 0.11 |
| Cell₃,₃ | 148 | 129.0 | 2.80 |
| χ² | 22.15 |
Reject or Accept Null Hypothesis
Given a significance level of 0.01 and 4 degrees of freedom [df = (3-1)(3-1) = 4], the critical value of χ² from the chi-square distribution table is 15.09.
The hypotheses are:
- H₀: The variables are independent (education and confidence have no relationship)
- Hₐ: The variables are not independent (there is a relationship)
The decision rule is:
- If χ² ≤ 15.09, accept H₀
- If χ² ≥ 15.09, reject H₀
Since the computed test statistic χ² = 22.15 is greater than the critical value 15.09 (22.15 > 15.09), the null hypothesis is rejected. Therefore, the two variables are not independent — there is a significant relationship between education and confidence in television programs. Insights into the nature of this relationship can be obtained by investigating individual cell contributions to the chi-square statistic.
💡 Why this matters: Individual cell contributions show which combinations of categories contribute most to the relationship. Here, Cell₁,₁ (Under Metric with "A great extent") has the highest contribution (9.71), suggesting this group differs most from what would be expected under independence.
⭐ Key Takeaways
The Bivariate Chi-Square test determines whether two categorical variables are independent or related by comparing observed frequencies with expected frequencies under independence. The expected frequency for each cell is calculated as (row total × column total) ÷ overall sample size, and degrees of freedom are (r-1)(c-1). The test statistic χ² = Σ(f₀-fₑ)²/fₑ follows a chi-square distribution, and the null hypothesis of independence is rejected when the computed χ² exceeds the critical value at the chosen significance level. In the hotel example, χ² = 0.98 was less than the critical value of 12.59, so independence was accepted, while in the education-confidence example, χ² = 22.15 exceeded 15.09, indicating a significant relationship. Individual cell contributions help identify which specific category combinations drive the relationship.
🧠 Quick Revision Questions
- What is the formula for calculating expected frequencies in a chi-square test of independence?
- How many degrees of freedom are there for a 4×3 contingency table?
- In the hotel guest satisfaction example, was the null hypothesis accepted or rejected, and why?
- What does it mean when an individual cell's contribution to the chi-square statistic is large?
- For the education-confidence example, what were the null and alternative hypotheses, and what was the conclusion at α = 0.01?
📘 Lecture 32 — Correlation and Regression
📖 Overview: This lecture explores the measurement of relationships between variables in marketing research. It explains how to quantify associations using correlation, differentiating between types like non-monotonic, monotonic, and linear relationships. The primary focus is on the Product Moment (Pearson) correlation coefficient and Spearman's Rank Order correlation, providing formulas, examples, and interpretation guidelines essential for analyzing interval and ordinal data.
🗂️ Topics Covered
This lecture begins by discussing different types of relationships between variables, specifically non-monotonic and monotonic relationships. It then introduces linear correlation and the coefficient of correlation (r) as a measure of association for interval/ratio data, including its strength and direction. The Product Moment (Pearson) Correlation coefficient is defined with its formula and demonstrated through two detailed examples. Finally, the lecture covers Rank Order Correlation, specifically Spearman's rank order correlation, for ordinal data, with examples.
📝 Lecture Summary
Correlation and Regression
Many marketing research projects study relationships, such as between sales revenue and advertising expenses, market share and number of salespeople, or sales and distance to competing stores. Association between categories from nominal data was discussed with chi-square. Association between two interval-scaled data is measured through correlation. Correlations are various types and used for various purposes such as nonmonotonic, monotonic, linear, curvilinear, rank order, etc.
Non-monotonic Relationship
A non-monotonic relationship is one where the presence or absence of one variable is systematically associated with the presence or absence of another variable, but there is no discernable direction of relationship. However, the relationship exists. For example, in a fast food restaurant, experience tells us that morning customers buy tea but afternoon customers typically purchase soft drink. This relationship is in no way exclusive. There is no guarantee that a morning customer would always order tea and an afternoon customer always order soft drink. However, a relationship does exist.
Monotonic Relationship
A monotonic relationship is one where a researcher can assign only a general direction of association between two variables. A monotonic relationship can be either increasing or decreasing. If one variable increases and the other also increases, it is called a monotonic increasing relationship. On the other hand, a monotonic decreasing relationship would be when one variable increases but the other decreases. Monotonic means that the direction of the relationship is described. However, the exact amount of change in one variable as the other variable changes cannot be indicated.
LINEAR CORRELATION
Correlation or linear correlation is a measure of the nature and degree of association or co-variation between two variables. These variables are measured by interval or ratio scales. Co-variation is defined as the amount of change in one variable systematically associated with a change in another variable. The Coefficient of Correlation (r) is the index number that communicates the strength of association between two variables, say X and Y. Strength is indicated by the absolute size of the correlation.
🔑 Definition — Coefficient of Correlation (r): The index number that communicates the strength and direction of association between two variables.
Here is a rule of thumb about the strength of correlation (after it is proved that correlation is statistically significant):
| Range of Coefficient r | Strength of Association |
|---|---|
| 0.00 to ±0.20 | None |
| ±0.21 to ±0.40 | Very weak |
| ±0.41 to ±0.80 | Moderate |
| ±0.81 to ±1.00 | Strong |
Correlation coefficients that are closer to zero indicate that there is no systematic association between the two variables, and the ones which are closer to ±1.0 show that there is some systematic association between two variables. Direction of association is indicated by the sign (+ or -) of the size of the coefficient. If the sign is positive, it means there is a positive direction of the co-variation between the two variables, that is, both variables change in the same direction. In this case if variable X increases, Y would also increase. On the other hand, a negative sign in correlation indicates a negative co-variation, which means if one variable increases the other would decrease and vice versa.
📌 Example: If we find that number of years of education (variable X) and hours in reading newspaper (variable Y) have a correlation of 0.87, it means they are positively correlated. It will be interpreted like this: people with more education spend more time reading newspapers. But if we find a correlation of -0.81 between smoking cigarettes and education, it means more educated people smoke less.
Product Moment Correlation
Product Moment correlation is the statistic widely used in determining the size or strength of association between two interval or ratio scaled variables. It is also known as the Pearson Correlation Coefficient as it was originally proposed by Karl Pearson. Sometimes it is also called bivariate correlation, simple correlation or correlation coefficient.
📐 Formula: $$r_{xy} = \frac{\sum{(x - \bar{x})(y - \bar{y})}}{\sqrt{\sum{(x - \bar{x})^2} \cdot \sum{(y - \bar{y})^2}}}$$
Where X and Y are the two variables, $\bar{x}$ and $\bar{y}$ are their respective means and $r_{xy}$ is the correlation between X and Y.
Example 1: A company wants to find out if the number of salespersons in a sales territory has some relationship with the sales revenue. From a random sample of 10 territories, the following data were obtained:
| Territory # | # of Sales Persons (X) | Sales in 000 Rupees (Y) |
|---|---|---|
| 1 | 5 | 261 |
| 2 | 7 | 288 |
| 3 | 6 | 381 |
| 4 | 9 | 412 |
| 5 | 12 | 440 |
| 6 | 8 | 317 |
| 7 | 11 | 567 |
| 8 | 16 | 572 |
| 9 | 13 | 428 |
| 10 | 7 | 317 |
Solution: First, calculate the means: $\bar{x} = 9.4$, $\bar{y} = 398$.
| x | y | $(x - \bar{x})$ | $(x - \bar{x})^2$ | $(y - \bar{y})$ | $(y - \bar{y})^2$ | $(x - \bar{x})(y - \bar{y})$ |
|---|---|---|---|---|---|---|
| 5 | 261 | -4.4 | 19.36 | -127 | 16129 | 558.8 |
| 7 | 288 | -2.4 | 5.76 | -100 | 10000 | 240 |
| 6 | 381 | -3.4 | 11.56 | -7 | 49 | 23.8 |
| 9 | 412 | -.4 | .16 | 16 | 576 | -9.6 |
| 12 | 440 | 2.6 | 6.76 | 52 | 2704 | 135.2 |
| 8 | 317 | -1.4 | 1.96 | -71 | 5041 | 99.4 |
| 11 | 567 | 1.6 | 2.56 | 179 | 32041 | 286.4 |
| 16 | 569 | 6.6 | 43.56 | 181 | 32761 | 1194.6 |
| 13 | 428 | 3.6 | 12.96 | 40 | 1600 | 144 |
| 7 | 317 | -2.4 | 5.76 | -71 | 5041 | 170.4 |
| Sum | 110.4 | 104942 | 2843 |
$$r_{xy} = \frac{2843}{\sqrt{110.4 \times 104942}} = \frac{2843}{\sqrt{11585670.4}} = \frac{2843}{3403.76} = .84$$
There is a moderate positive correlation between the two variables.
Example II: A researcher is interested to find out whether the high price of large refrigerators compensates in terms of saving energy. He collected data on price and annual cost of consumption of electricity in terms of rupees from a random sample of 8 brands of refrigerators. The data are presented below.
| Refrigerator Brand | Price Rs. 000 (X) | Annual cost of Consumption on electricity Rs. 000 (Y) |
|---|---|---|
| A | 85 | 4.8 |
| B | 76 | 5.4 |
| C | 90 | 5.8 |
| D | 87 | 6.6 |
| E | 110 | 7.7 |
| F | 80 | 6.6 |
| G | 65 | 7.0 |
| H | 75 | 8.1 |
Find the coefficient of correlation r. What do you say about it?
$$r_{xy} = \frac{10.01}{10.49} = .96$$
There is a strong positive correlation between the two variables.
Rank Order Correlation
We use Rank Order Correlation coefficient in data analysis to measure the monotonic relationship between two variables measured on an ordinal scale. That is when the data are ordered into ranks.
🔑 Definition — Spearman's Rank Order Correlation (rs): A measure that indicates the direction and degree of association between two sets of rankings.
📐 Formula: $$r_s = 1 - \frac{6(\sum d^2)}{n(n^2 - 1)}$$
Where:
- $r_s$ = Spearman’s rank order correlation
- d = difference of ranks in the paired ranking
- n = number of items ranked
Example: Ranking of various toothpaste brands according to consumer perception on their decay prevention and whitening ability.
| Toothpaste Brand | Decay Prevention Rating | Whitening Rating |
|---|---|---|
| ... | ... | ... |
The computation of Spearman rank order correlation in this example is as follows: $$r_s = 1 - \frac{6(\sum d^2)}{n(n^2 - 1)} = 1 - \frac{6 \times 156}{8(8^2 - 1)} = 1 - \frac{936}{8(63)} = 1 - \frac{936}{504} = 1 - 1.86 = -.86$$
A Spearman rank order correlation of -.86 shows a fairly strong relationship between the two rank orders i.e., decay prevention and whitening ability for the toothpaste brands under study. A negative sign indicates that the relationship is negative, which means if the ranks are higher on decay prevention, these are lower on whitening ability. These are the perceived rankings by the respondent.
Example: Spearman’s Rank Order Correlation: Two judges were asked to rate six brands of tissue papers. Their ratings are given below. Find out the Spearman’s rank order correlation between the ratings assigned by two judges.
| Tissue Brand | Judge I Rank | Judge II Rank |
|---|---|---|
| G | 2 | 6 |
| H | 3 | 3 |
| I | 6 | 5 |
| J | 1 | 4 |
| K | 5 | 2 |
| L | 4 | 1 |
The Spearman rank correlation coefficient is –.26, which is very weak and it is in the opposite direction.
💡 Why this matters: Understanding correlation helps marketers quantify the strength and direction of relationships between variables, such as advertising spend and sales. This allows for more informed decision-making and resource allocation. The choice of correlation method depends on the data's scale of measurement (interval vs. ordinal).
⭐ Key Takeaways
The core of this lecture is understanding that correlation measures the strength and direction of a linear relationship between two variables. The Pearson correlation coefficient (r) is for interval/ratio data, ranging from -1 to +1, where its absolute value indicates strength (with ±0.81 being strong) and its sign indicates direction. Spearman's rank order correlation (rs) is the non-parametric equivalent used for ordinal (ranked) data. It is crucial to remember that correlation does not imply causation; a strong correlation only indicates an association, not that one variable causes the other to change.
🧠 Quick Revision Questions
- What is the fundamental difference between a monotonic and a non-monotonic relationship?
- For an interval-scaled variable, which correlation coefficient would you use to measure its linear association with another interval-scaled variable?
- Interpret a Pearson correlation coefficient (r) of -0.85 between advertising spend and customer complaints. What does this tell you about the relationship?
- You have two sets of rankings for five brands. What formula would you use to measure the association between these rankings?
- According to the lecture's rule of thumb, what is the strength of association for a correlation coefficient of +0.35?
📘 Lecture 33 — Testing the difference between the Means
📖 Overview: This lecture addresses how marketing researchers test hypotheses about differences between means, such as comparing a sample mean to a population mean or comparing two independent sample means. It explains the appropriate use of the z-test and t-test based on sample size and knowledge of population standard deviation, and provides worked examples to illustrate the procedure.
🗂️ Topics Covered
The lecture begins by establishing the context for testing differences between means, distinguishing when to use chi-square (nominal data), Spearman’s correlation (ordinal data), and z/t-tests (interval data). It then explains the conditions for choosing between the z-test and t-test based on sample size and population standard deviation. The method for calculating standard error of the mean difference for one mean and two means is presented. Finally, two complete examples are given: a one-sample z-test for a shoe company’s new design decision, and a two-sample z-test comparing weekly consumption of two cola beverages.
📝 Lecture Summary
Testing the difference between the Means
Marketing research may face the problem of testing hypotheses relating to the difference between means. Such means may be the sample mean and population mean or two independent means.
Chi-square is one procedure used to test the significance of difference of data derived from nominal data. For ordinal data, we may use Spearman’s correlation. We use different statistics to test the difference between means based on interval data. These statistical tests are classified by whether one or two samples are involved: one mean from a sample may be compared with a mean hypothesized to exist in the population, or two means generated from independent samples.
The z-test and the t-test are appropriate to test the difference between the means. The choice is made based on the researcher’s knowledge about the standard deviation of the population and the sample size. It is appropriate to use these tests in the following situations:
- If sample size is more than 30 & population standard deviation is unknown, use z-test.
- If sample size is less than 30 and population standard deviation is unknown, use t-test.
How to Calculate Z or T?
In order to calculate z or t statistics, we need to have calculated the population mean and standard deviation of one or more samples being tested. Then we calculate the standard error of the mean difference. In the case of one mean, the standard error of the mean is:
🔑 Definition — Standard Error of the Mean (Sx): The standard deviation of the sampling distribution of the sample mean, measuring how much the sample mean is expected to vary from the population mean. 📐 Formula: Sx = S / √n → (Sample standard deviation divided by the square root of the sample size)
In the case of two means, the standard error will be: 📐 Formula: Sx₁-x₂ = √[(S₁²/n₁) + (S₂²/n₂)] → (Square root of the sum of the variance of sample 1 divided by its size plus the variance of sample 2 divided by its size)
The z or t is calculated by the following formulas.
For One Mean: 📐 Formula: z (or t) = (x̄ - μ) / Sx Where:
- μ = population mean (hypothesized)
- x̄ = sample mean
- Sx = standard error of the mean difference
Example I: One-Sample Z-Test
A shoe company is investigating the desirability of adding a new design of shoes to its shop. The company has decided that it will add this design only if it sells 100 pairs per week in each store. It was put on 40 different shops randomly and data on their sale was calculated. The data revealed an average sale of 106 pairs per week per store. The standard deviation was 13.8. Should the company introduce the new design?
Solution:
- Hypotheses:
- H₀: μ ≤ 100 (Null hypothesis: mean sales are less than or equal to 100)
- H₁: μ > 100 (Alternative hypothesis: mean sales are greater than 100)
- Given:
- Sample mean (x̄) = 106
- Sample standard deviation (S) = 13.8
- Sample size (n) = 40
- Population mean (μ) = 100
- Calculate Standard Error:
- Sx = S / √n = 13.8 / √40 = 13.8 / 6.32 = 2.18
- Choose test: Since sample size > 30 and population standard deviation is unknown, use z-test.
- Calculate z-statistic:
- z = (x̄ - μ) / Sx = (106 - 100) / 2.18 = 6 / 2.18 = 2.75
- Decision: The critical value of z at the 0.05 confidence level for a one-tailed test is 1.645 (the lecture text says 1.96, which is for a two-tailed test; 1.645 is the standard for a one-tailed test at 0.05). Since the calculated z (2.75) exceeds the critical value (1.645), the null hypothesis is rejected.
- Conclusion: The company should go ahead and introduce the new shoes.
📌 Example: The shoe company found a sample mean of 106 pairs, which was 2.75 standard errors above the hypothesized mean of 100. This is a statistically significant difference, justifying the decision to add the new design.
Example of Z-Test for Two Independent Means
Let us take the example of a z-test for two means. Suppose two independent samples of 50 each of the two cola beverages A and B yield the following data on average weekly per household consumption in Gulberg area of Lahore.
- Cola beverage A: mean consumption = 6.3 liters, standard deviation = 2.1
- Cola beverage B: mean consumption = 5.7 liters, standard deviation = 1.3
Solution:
- Hypotheses:
- H₀: μA = μB (No significant difference between the means)
- H₁: μA ≠ μB (Significant difference between the means)
- Given:
- x̄A = 6.3, SA = 2.1, nA = 50
- x̄B = 5.7, SB = 1.3, nB = 50
- Calculate Standard Error of the Difference between Two Means:
- Sx₁-x₂ = √[(S₁²/n₁) + (S₂²/n₂)] = √[(2.1² / 50) + (1.3² / 50)] = √[(4.41 / 50) + (1.69 / 50)]
- Sx₁-x₂ = √[0.0882 + 0.0338] = √0.122 = 0.349
- Choose test: Since both samples are more than 30, we select z-test.
- Calculate z-statistic:
- z = (x̄A - x̄B) / Sx₁-x₂ = (6.3 - 5.7) / 0.349 = 0.6 / 0.349 = 1.72
- Decision: This is a two-tailed test. The critical value of z at the 0.05 significance level is 1.96. Since the calculated value of z (1.72) is less than the critical value (1.96), we fail to reject the null hypothesis.
- Conclusion: There is no significant difference between the average weekly consumption per household of the two brands.
📌 Example: The difference in average consumption between Cola A (6.3 liters) and Cola B (5.7 liters) was 0.6 liters. The calculated z-statistic of 1.72 was less than the critical value of 1.96, so this observed difference is not statistically significant at the 0.05 level.
💡 Why this matters: These tests are fundamental for making data-driven marketing decisions. The shoe company example shows how to use a one-sample test to decide on product launches, while the cola example shows how to use a two-sample test to compare brand performance, without relying on assumptions.
⭐ Key Takeaways
The choice between a z-test and t-test depends on sample size and knowledge of population standard deviation; use z-test for samples over 30 when the population standard deviation is unknown, and t-test for samples under 30 in the same situation. The standard error of the mean is calculated as S/√n for one sample and √[(S₁²/n₁)+(S₂²/n₂)] for two independent samples. In the one-sample shoe example, a calculated z of 2.75 exceeded the critical value, leading to the rejection of the null hypothesis and a recommendation to launch the new design. In the two-sample cola example, a calculated z of 1.72 was less than the critical value of 1.96, leading to the conclusion that there is no significant difference in weekly consumption between the two brands. The z-statistic is the difference between sample and hypothesized means (or difference between two sample means) divided by the standard error of that difference.
🧠 Quick Revision Questions
- What are the conditions for choosing a z-test over a t-test when testing the difference between means?
- How do you calculate the standard error of the mean for a single sample and for the difference between two independent samples?
- In the shoe company example, what was the null hypothesis, and why was the calculated z-statistic (2.75) sufficient to reject it?
- In the two-sample cola example, why was a two-tailed test used, and what does a calculated z of 1.72 compared to a critical value of 1.96 lead you to conclude?
- Explain why the standard error is a crucial component in calculating both the z-statistic and the t-statistic.
📘 Lecture 34 — Testing the Difference between Means
📖 Overview: This lecture introduces the t-test for hypothesis testing when sample sizes are small (less than 30) and population variance is unknown. It covers both the t-test for one mean and for two independent means, demonstrating how to calculate test statistics, determine degrees of freedom, and interpret results using real business examples.
🗂️ Topics Covered
The lecture begins by reviewing when to use the t-test versus the z-test and explains how to calculate degrees of freedom. It then walks through a complete example of the t-test for one mean involving customer exchanges at a departmental store. Finally, it presents a detailed example of the t-test for two independent means comparing sales of car polish in plastic versus metal containers across 10 stores each.
📝 Lecture Summary
t-Test
You may recall that we use t-test if the sample is less than 30 and population variance is not known. z and t are calculated in the same manner with the same formula but in reading the critical value of t from the table of t distribution, we need degree of freedom which is calculated as:
🔑 Definition — Degree of freedom (df): A statistical parameter that determines the specific t-distribution to use when finding critical values.
📐 Formula: df = (n₁ + n₂ - 2) → The total sample size minus 2, used when comparing two independent samples.
t-Test for One Mean
Suppose a departmental store manager believes that an average number of customers who exchange merchandise each day is not more than 20. The store records the number of exchanges each day for 26 days it was open for a given month. The researcher calculates a sample mean equal to 22 and standard deviation equal to 5. Do you think manager’s assumption is correct?
Solution
x̄ = 22
n = 26
s = 5
Hypotheses:
H₀: μ ≤ 20 (Manager's assumption - average exchanges not more than 20)
H₁: μ > 20 (Alternate - average exchanges are more than 20)
Standard error of the mean:
Sₓ = s/√n = 5/√26 = 5/5.1 = 0.98
Calculated t-value:
t = (x̄ - μ)/Sₓ = (22 - 20)/0.98 = 2.04
Degree of freedom: df = n - 1 = 26 - 1 = 25
Critical value at 0.05 level of significance at 25 df = 1.70 (from t-distribution table for one-tailed test)
As the calculated t (2.04) is greater than 1.70, H₀ is rejected.
Conclusion: Average number of customers who exchange merchandise each day in the store is more than 20. Manager's assumption is incorrect.
t-Test for Two Independent Means
A manufacturer of car polish has recently developed a new polish. The company is considering two different containers for the polish, one metal and one plastic. The company will make final decision after test marketing. The company introduced metal and plastic containers on 10 independent random samples of 10 stores each.
Results:
| Store # | Plastic container | Metal container | Store # | Plastic container | Metal container |
|---|---|---|---|---|---|
| 1 | 416 | 340 | 6 | 358 | 358 |
| 2 | 327 | 385 | 7 | 400 | 352 |
| 3 | 370 | 380 | 8 | 394 | 390 |
| 4 | 380 | 376 | 9 | 390 | 360 |
| 5 | 400 | 382 | 10 | 381 | 385 |
Which container should they introduce?
Solution
Hypotheses:
H₀: μ₁ = μ₂ (No difference between sales of plastic and metal containers)
H₁: μ₁ ≠ μ₂ (There is a difference in sales)
Descriptive statistics (calculated from data):
Mean of plastic container (x̄₁) = 381.6
Mean of metallic container (x̄₂) = 370.8
Standard deviation of plastic container (s₁) = 25.29
Standard deviation of metallic container (s₂) = 16.97
💡 Why this matters: In calculating standard deviation for interval data, we derive the sum of squared deviation by (n-1) to get an unbiased estimate of the population variance.
Calculated t-value: The t-statistic for two independent means uses the formula that incorporates the standard deviations and sample sizes of both groups.
Critical value for a two-tailed test at (n₁ + n₂ - 2) = 10 + 10 - 2 = 18 df at 0.05 level of significance is 2.10.
Since the calculated t-value is less than the critical value of 2.10, the null hypothesis is accepted.
Conclusion: Company can introduce any container. Customers do not have special preference for either container.
⭐ Key Takeaways
The t-test is essential when working with small samples (n < 30) and unknown population variance. For one-mean tests, degrees of freedom equal n-1, while for two independent means, df equals n₁ + n₂ - 2. The critical value is read from the t-distribution table based on the chosen significance level (typically 0.05) and the direction of the test (one-tailed or two-tailed). A calculated t-value exceeding the critical value leads to rejection of the null hypothesis. In business contexts, t-tests help managers make data-driven decisions about assumptions and product strategies, as demonstrated by the container preference example.
🧠 Quick Revision Questions
- What conditions must be met to use a t-test instead of a z-test?
- Calculate the degrees of freedom for a one-mean t-test with a sample of 26 observations.
- In the customer exchange example, why was the null hypothesis rejected when t = 2.04 and the critical value was 1.70?
- For two independent samples of 10 stores each, what is the degrees of freedom for the t-test?
- What was the final business decision in the car polish container example, and why?
📘 Lecture 35 — Analysis of Variance
📖 Overview: This lecture introduces Analysis of Variance (ANOVA), a statistical technique used to test differences among more than two means. It explains when to use ANOVA instead of t or z tests, and provides a detailed procedure for conducting one-way ANOVA with a full worked example from marketing research.
🗂️ Topics Covered
The lecture covers the definition and purpose of ANOVA, the distinction between one-way and n-way ANOVA, the key statistical components (SS total, SS between, SS within, mean squares, and F-ratio), and the step-by-step procedure for conducting one-way ANOVA. A complete example involving test marketing of shampoo at three different price levels is worked through to demonstrate the calculations and hypothesis testing process.
📝 Lecture Summary
Analysis of Variance (ANOVA)
ANOVA is used when a market researcher wants to examine differences among more than two means. While t or z tests handle comparisons between two independent means, ANOVA is the appropriate technique for multiple group comparisons. Although traditionally used for experimental data, ANOVA can also analyze survey or observational data.
In ANOVA, we have one or more independent variables which must be non-metric or categorical, and a dependent variable which is metric (measured through interval or ratio scale). Independent variables that are categorical are also called factors. A particular level of factors or independent variables is called a treatment. If one factor (at different levels) is the treatment, one-way analysis of variance is used. If more than one factor are the treatment, n-way analysis of variance is used. If the set of independent variables contains both categorical and metric variables, analysis of covariance (ANCOVA) is used.
One-way Analysis of Variance
One-way ANOVA is used when the researcher wants to examine differences in the mean values of the dependent variable for several levels of a single independent variable or factor. Examples include:
- Are the attitudes of various channels of distribution (wholesalers, retailers, agents) different towards the company's distribution policies?
- Do different regions differ in sales?
- Are the results of various test markets at different price levels really different?
- Are brands evaluated differently by different groups exposed to ads?
The key statistics related to one-way analysis of variance include:
- Mean Square
- Sum of Squares Between (SS between) — also called variation between groups (VB)
- Sum of Squares Total (SS total) — also called total variation (TV)
- Sum of Squares Within (SS within) — also called variation within groups (VW)
- F ratios
Conducting One-Way ANOVA
The procedure for conducting one-way analysis of variance involves the following steps:
A. Identifying the variables: The researcher first identifies the independent variable with all its levels and the dependent variable. The independent variable is generally denoted by X and the dependent variable by Y. In one-way ANOVA, there is one categorical variable having more than two categories, say the number of categories is c. If each category has n observations, the total sample size will be n × c. The researcher then prepares the null and alternate hypotheses.
B. Measure and decompose the variation: Find the total variation in Y and separate it into variation between groups and variation within groups. This is expressed by the equation:
🔑 Definition — SS total = SS between + SS within
SS total is computed by squaring the deviation of each score from the grand mean and summing these squares.
SS within is the variability observed within each group, calculated by squaring the deviation of each score from its group mean and summing these scores. 📐 Formula: SS within = Sum c (x – x̄_group)²
SS between is the variability of the group means about the grand mean, computed by squaring the deviation of each group mean from the grand mean, multiplying by n, and summing them up. 📐 Formula: SS between = Sum [n (x̄_group – grand mean)²]
After calculating SS total, SS between, and SS within, we compute the variance or mean square by dividing the various sum of squares by the appropriate degrees of freedom.
To get Mean Square Between Groups (MS between), SS between is divided by categories minus one (c-1): 📐 Formula: MS between = SS between / (c-1)
To obtain Mean Square Within Groups (MS within), SS within is divided by cn-c degrees of freedom: 📐 Formula: MS within = SS within / (cn - c)
Finally, the F-ratio is found by dividing MS between by MS within: 📐 Formula: F = MS between / MS within
The F-ratio is then checked against the critical value in the relevant F-table for comparison. At the end, the null hypothesis is accepted or rejected, and the conclusion is drawn.
Example
A company wants to launch a new shampoo but is unsure which price will bring more sales. They ran a test in the market before launching the product to decide the final price. The company chose four separate areas, and within each area the product was sold at three different prices in different markets. There were 12 test markets total.
Data from the Test Markets (Unit Sales in 000)
| Market Area | Regular Price (Rs. 325) | Reduced Price (Rs. 315) | Discount (Coupon) |
|---|---|---|---|
| M, N, O | 11 | 12 | 13 |
| P, Q, R | 11 | 14 | 12 |
| S, T, U | 9 | 12 | 9 |
| W, X, Y | 8 | 13 | 10 |
| Mean | 9.75 | 12.75 | 11 |
Grand Mean = 11.17
Solution
Step 1: Calculate SS total SS total = (11-11.17)² + (11-11.17)² + (9-11.17)² + (8-11.17)² + (12-11.17)² + (14-11.17)² + (12-11.17)² + (13-11.17)² + (13-11.17)² + (12-11.17)² + (9-11.17)² + (10-11.17)² = 37.67
Step 2: Calculate SS within SS within = (11-9.75)² + (11-9.75)² + (9-9.75)² + (8-9.75)² + (12-12.75)² + (14-12.75)² + (12-12.75)² + (13-12.75)² + (13-11)² + (12-11)² + (9-11)² + (10-11)² = 19.56
Step 3: Calculate SS between SS between = 4(9.75-11.17)² + 4(12.75-11.17)² + 4(11-11.17)² = 18.17
Step 4: Calculate Mean Squares MS between = 18.17 / (c-1) = 18.17 / 2 = 9.08 MS within = 19.57 / (cn-c) = 19.57 / 9 = 2.17
Step 5: Calculate F-ratio F = 9.08 / 2.17 = 4.178
Step 6: Compare with critical value The F-ratio is 4.178. Comparing with the F-table at (c-1) = 2 degrees of freedom in the numerator and (cn-c) = 9 degrees of freedom in the denominator, the critical value is 4.26. Since 4.178 is less than 4.26, we accept the null hypothesis.
💡 Why this matters: The calculated F-value being less than the table value means the observed differences in sales across price treatments are not statistically significant.
Conclusion: It appears that there is no real difference in sales produced by different prices. All price treatments produce almost the same sales volume.
⭐ Key Takeaways
ANOVA is the correct statistical technique when comparing means across more than two groups, replacing multiple t-tests. The total variation in the dependent variable (SS total) is decomposed into variation between groups (SS between) and variation within groups (SS within). The F-ratio compares the mean square between groups to the mean square within groups, and this value is compared against a critical value from the F-distribution table using appropriate degrees of freedom. In the shampoo pricing example, the calculated F-value of 4.178 did not exceed the critical value of 4.26, leading to acceptance of the null hypothesis that different prices produce equivalent sales. A non-significant ANOVA result indicates that group means are not statistically different from each other.
🧠 Quick Revision Questions
- When should a researcher use ANOVA instead of a t-test or z-test?
- What are the three types of sum of squares in one-way ANOVA, and what does each represent?
- How is the F-ratio calculated in one-way ANOVA, and what do the numerator and denominator represent?
- In the shampoo pricing example, why was the null hypothesis accepted even though the group means (9.75, 12.75, and 11) appeared different?
- What are the degrees of freedom for the numerator and denominator when testing one factor with 4 categories and 5 observations per category?
📘 Lecture 36 — ANOVA and Post Hoc Analysis
📖 Overview: This lecture teaches how to use Analysis of Variance (ANOVA) to test whether there are significant differences among three or more group means, using a marketing example of product location affecting sales. It then covers post hoc analysis to identify which specific groups differ after finding a significant ANOVA result. This matters because marketing managers need to pinpoint which store locations (or other factors) truly perform differently.
🗂️ Topics Covered
The lecture begins with a one-way ANOVA example testing if product location (front, middle, back) affects toy sales, computing SS total, SS within, SS between, MS between, MS within, and the F-ratio. It interprets the F-test result by comparing it to the critical value, then introduces multiple comparisons and post hoc analysis. Post hoc steps include determining the number of comparisons, computing mean differences, obtaining the critical range using the studentized range distribution (Q), and comparing each pair to identify significant differences.
📝 Lecture Summary
ANOVA and Post Hoc Analysis
A marketing manager of a toy company wants to find out whether product location in the store (front, middle, and back) affects the sales of toys. Six stores are randomly selected for each location. Price, store size and display area for the product is constant for all stores. After one month of the experimental research period, the sales volume of each location in the stores was noted. The task is to test at the 0.05 level of significance whether there is evidence of significant difference of sales volume among various locations, and determine which location appears to be different significantly in average sales.
Null Hypothesis: Mean X₁ = Mean X₂ = Mean X₃
Alternate Hypothesis: All means are not equal.
🔑 Definition — Grand Mean: The overall mean of all data points combined across all groups.
🔑 Definition — SS Total (Total Sum of Squares): The sum of squared deviations of each individual score from the grand mean. It measures total variation in the data.
📐 Formula for SS Total: SS Total = Σ(Xᵢ – X̄_grand)²
📌 Example — Computing SS Total:
Grand Mean = 39.61
SS Total = (86 - 39.61)² + (72 - 39.61)² + (54 - 39.61)² + (62 - 39.61)² + (50 - 39.61)² + (40 - 39.61)² + (46 - 39.61)² + (60 - 39.61)² + (40 - 39.61)² + (28 - 39.61)² + (22 - 39.61)² + (24 - 39.61)² + (33 - 39.61)² + (24 - 39.61)² + (20 - 39.61)² + (14 - 39.61)² + (18 - 39.61)² + (16 - 39.61)²
= 7406.28
🔑 Definition — SS Within (Within-Groups Sum of Squares): The sum of squared deviations of each score from its own group mean. It measures variation within groups (error or unexplained variation).
📐 Formula for SS Within: SS Within = Σ(Xᵢ – X̄_group)²
📌 Example — Computing SS Within:
Group means: Front = 60.67, Middle = 37.3, Back = 20.83
SS Within = (86 - 60.67)² + (72 - 60.67)² + (54 - 60.67)² + (62 - 60.67)² + (50 - 60.67)² + (40 - 60.67)² + (46 - 37.3)² + (60 - 37.3)² + (40 - 37.3)² + (28 - 37.3)² + (22 - 37.3)² + (28 - 37.3)² + (33 - 20.83)² + (24 - 20.83)² + (20 - 20.83)² + (14 - 20.83)² + (18 - 20.83)² + (16 - 20.83)²
= 2599.51
🔑 Definition — SS Between (Between-Groups Sum of Squares): The sum of squared deviations of each group mean from the grand mean, weighted by sample size. It measures variation due to the treatment/group effect.
📐 Formula for SS Between: SS Between = Σ nⱼ (X̄ⱼ – X̄_grand)²
📌 Example — Computing SS Between:
Each group has n = 6
SS Between = 6(60.67 – 39.61)² + 6(37.3 – 39.61)² + 6(20.83 – 39.61)²
= 4809.29
🔑 Definition — MS Between (Mean Square Between Groups): Average between-group variation. Calculated by dividing SS Between by its degrees of freedom (c – 1).
📐 Formula: MS Between = SS Between / (c – 1)
📌 Example: MS Between = 4809.29 / 2 = 2404.64
🔑 Definition — MS Within (Mean Square Within Groups): Average within-group variation. Calculated by dividing SS Within by its degrees of freedom (cn – c).
📐 Formula: MS Within = SS Within / (cn – c)
📌 Example: MS Within = 2599.51 / 15 = 173.30
🔑 Definition — F-ratio: The test statistic for ANOVA. Ratio of MS Between to MS Within. A large F suggests group means are not all equal.
📐 Formula: F = MS Between / MS Within
📌 Example: F = 2404.64 / 173.30 = 13.88
Now we find the critical value of F at the 0.05 level of significance. We need degrees of freedom in numerator which is c – 1 = 2, and degrees of freedom in denominator which is cn – c = 15. The value of F at these degrees of freedom at 0.05 is 3.68, which is lesser than the calculated value (13.88). Hence Null Hypothesis is rejected. It means that there is evidence of significant difference in average sale among the various product locations in the store.
💡 Why this matters: When we look at the table, we find that sales at the front location is different (greater) from the middle and from the back location.
Multiple Comparisons
Once the differences in the means of groups have been established, it becomes important to determine means of which particular groups are significantly different from each other. This is done with a procedure called post hoc comparison. It is called post hoc because other hypotheses are formulated after the data have been inspected under one-way ANOVA.
Post hoc Analysis
- Determine the number of comparisons by the formula c(c-1)/2.
- Compute differences between various means.
- Obtain critical range for this procedure.
- Compare each of the c(c-1)/2 pair of means with the critical range. (Note that critical range would remain same if the sample size in all pairs is equal. In case it is different, then critical range will be calculated for each pair of means.)
- If the absolute difference of a pair of means is greater than the critical range, the difference between that pair of means is significant; otherwise not.
Now let us go back to our previous example of location of toys in the store and perform the post hoc analysis.
Step 1: Number of comparisons
Possible number of comparisons = c(c-1)/2 = 3(3-1)/2 = 3
Step 2: Absolute mean differences
- X̄₁ – X̄₂ = 60.67 – 37.3 = 23.37
- X̄₁ – X̄₃ = 60.67 – 20.83 = 39.84
- X̄₂ – X̄₃ = 37.3 – 20.83 = 16.47
Step 3: Obtain critical range
Obtain only one critical range because the three groups have the same sample size. The equation to obtain critical range is:
📐 Formula — Critical Range: Critical Range = Q * √(MSW / n)
Where Q is the upper-tail critical value from a studentized range distribution table with c degrees of freedom in the numerator and cn – c degrees of freedom in the denominator.
The degree of freedom in numerator is 3 and denominator is 18 – 3 = 15. When we look up the value of Qᵤ in the table against these degrees of freedom, it is 3.67.
Taking figures of MSW and n from our previous example: Critical Range = 3.67 * √(173.30 / 6) = 19.63
Step 4: Compare differences to critical range
- 23.37 > 19.63 → Significant
- 39.84 > 19.63 → Significant
- 16.47 < 19.63 → Not Significant
Step 5: Conclusions
- X̄₁ – X̄₂ = is SIGNIFICANT
- X̄₁ – X̄₃ = is SIGNIFICANT
- X̄₂ – X̄₃ = is NOT SIGNIFICANT
⭐ Key Takeaways
The key to understanding ANOVA is recognizing that it partitions total variation into between-group and within-group components, with the F-ratio testing whether group differences are larger than random error. In this example, the front location had significantly higher average sales than both middle and back, while middle and back were not significantly different from each other. Post hoc analysis using the studentized range (Q) provides a critical range to compare pairs, ensuring that multiple comparisons do not inflate Type I error. Always remember that ANOVA only tells you if at least one group differs; post hoc tests identify exactly which groups differ.
🧠 Quick Revision Questions
- What are the three sums of squares computed in one-way ANOVA, and what does each measure?
- How do you calculate the F-ratio, and what does a large F value indicate about the null hypothesis?
- Why is post hoc analysis necessary after a significant ANOVA result, and what is the formula to determine the number of comparisons?
- In the toy store example, which pairs of locations showed significant sales differences, and why was one pair not significant?
- What is the studentized range statistic Q used for in post hoc analysis, and what degrees of freedom are needed to find its critical value?
📘 Lecture 37 — Regression Analysis
📖 Overview: This lecture introduces regression analysis, a statistical method used for prediction. It distinguishes regression from correlation and explains how to build a simple linear regression model to predict a dependent variable based on one independent variable. The lecture provides step-by-step examples for calculating regression coefficients using the least squares method.
🗂️ Topics Covered
This lecture first defines regression analysis and contrasts it with correlation. It then describes the types of regression models, including simple and multiple regression. The core of the lecture focuses on the simple linear regression model, the least squares method for calculating the regression coefficients (intercept and slope), and concludes with two detailed examples demonstrating how to find the regression equation and use it for prediction.
📝 Lecture Summary
Regression and Correlation
Regression analysis is used to analyze the associative relationship between a metric dependent variable and one or more independent variables. It shows whether a relationship exists between the criterion and predictor variable and whether the independent variable explains a significant variation in the dependent variable. Its goal is to predict the value of a dependent variable (criterion/response variable) based on the value of an independent variable (predictor/explanatory variable).
Correlation, in contrast, measures the strength and direction of association between two numerical variables. These variables may not necessarily have a dependent-independent relationship. The objective is not to use one variable to predict the other but to measure the strength of covariation between them.
🔑 Definition — Regression: The dependence of a variable on one or more other variables, primarily used for prediction.
Types of Regression Models
The relationship between the dependent and independent variables can take many forms. The simplest relationship is the straight line or linear relationship with only one predictor variable. This is called a simple regression model, straight-line regression, or bivariate regression model. If there are more than one independent variable, it is called a multiple regression model.
Simple Regression Model
Construction of a regression model starts with identifying the dependent and independent variables. The simple linear regression model is: 📐 Formula: Y = a + bx → The value of the dependent variable (Y) is predicted to be the intercept (a) plus the slope (b) times the value of the independent variable (X). Where ‘a’ is the intercept and ‘b’ is the slope. Slope means the amount of change in Y if there is a change of one unit in X. The best fit on a scatter diagram is found by minimizing the differences between Y (actual values) and ŷ (the predicted value). 📐 Formula: ŷ = a + bx → The equation for predicted values of Y.
A mathematical technique called the least squares method is used to determine the regression parameters ‘a’ and ‘b’. This method results in the minimum sum of squared differences between the actual value of Y and the predicted value of ŷ. To apply this method, calculate the following: x̄ (mean of X), ȳ (mean of Y), n (sample size), ∑X (sum of X), ∑Y (sum of Y), ∑X² (sum of squared X values), and ∑XY (sum of product of X & Y). 📐 Formula: b = ∑(x - x̄)(y - ȳ) / ∑(x - x̄)² = (∑xy - n x̄ ȳ) / (∑x² - n x̄²) → The formula for the slope. 📐 Formula: a = ȳ - b x̄ → The formula for the intercept.
📌 Example 1: A manager wishes to predict delivery time based on the number of crates delivered. A sample of 15 customers was taken.
- Given: ∑X = 2432, ∑Y = 710, ∑X² = 486360, ∑XY = 128722, n = 15, x̄ = 162.13, ȳ = 47.33
- Step 1: Calculate b. b = (128722 - 15 * (162.13 * 47.33)) / (486360 - 15 * (162.13)²) b = (128722 - 115104) / (486360 - 394292) b = 13618 / 92068 = 0.1479
- Step 2: Calculate a. a = 47.33 - 0.1479 * 162.13 a = 47.33 - 23.98 = 23.35
- Step 3: State the regression equation: ŷ = 23.35 + 0.1479X
- Step 4: Interpret a and b. The intercept 'a' (23.35) is the estimated delivery time when no crates are delivered. The slope 'b' (0.1479) means that for each additional crate delivered, the delivery time is expected to increase by 0.1479 minutes.
- Step 5: Predict the delivery time for 150 crates. ŷ = 23.35 + 0.1479 * 150 = 23.35 + 22.185 = 45.535 minutes.
📌 Example 2: A company wants to check the sale of washing machines as a function of expenditures on R&D for 14 years.
- Given: ∑X = 156831834 (Note: this is ∑X²), ∑XY = 301298822, n = 14, x̄ = 2921.28, ȳ = 5826.93
- Step 1: Calculate b. b = (301298822 - 14 * (2921.28 * 5826.93)) / (156831834 - 14 * (2921.28)²) b = (301298822 - 238309317) / (156831834 - 119474276) b = 62989505 / 37357558 = 1.686
- Step 2: Calculate a. a = 5826.93 - 1.686 * 2921.28 a = 5826.93 - 4925.28 = 901.6
- Step 3: State the regression equation: ŷ = 901.6 + 1.686X
- Step 4: Predict sales if R&D expenditures are Rs.50000 (note: R&D data is in Rs. 00, so X=500). ŷ = 901.6 + 1.686 * 500 ŷ = 901.6 + 843 = 1744.6 thousand rupees or Rs. 1,744,600.
💡 Why this matters: Understanding the slope and intercept allows you to quantify the relationship between variables and make data-driven predictions.
⭐ Key Takeaways
Regression is a predictive tool that models the relationship between a dependent and one or more independent variables, while correlation only measures the strength of association between two variables. The simple linear regression model (ŷ = a + bx) is used for prediction with one independent variable. The least squares method minimizes the sum of squared errors to find the best fitting line, with the slope (b) showing the change in Y for a one-unit change in X. Always interpret the coefficients in the context of the problem and ensure the independent variable's units match the prediction scenario.
🧠 Quick Revision Questions
- What is the primary goal of regression analysis?
- How does regression differ from correlation?
- In the simple linear regression equation ŷ = a + bx, what does the coefficient ‘b’ represent?
- What is the purpose of the least squares method?
- An estimated regression equation is ŷ = 5 + 2x. Predict y when x = 10.
📘 Lecture 38 — Report Writing
📖 Overview: This lecture covers the final and critical step in the marketing research process — writing the research report. It explains why the report is the only tangible output of a research project and provides detailed guidelines for creating an effective, professional report that communicates findings clearly to management for decision-making.
🗂️ Topics Covered
The lecture begins by defining the research report as the only tangible output of a marketing research project and its importance as a historical record. It then discusses the importance of the research report, emphasizing that it is the culminating activity and the only part the client sees. Next, it provides comprehensive guidelines for writing the report, including considering the reader, avoiding technical jargon, logical organization, using headings, being objective, and following all principles of good communication, along with ensuring a professional look. Finally, it outlines the standard format of a research report, covering front matter, body parts, and end matter.
📝 Lecture Summary
Research Report
The report represents the efforts of the research team. If it is poorly written with lots of errors, the quality of the whole research may become suspicious. On the other hand, if all aspects of a research report are done well, the credibility of the researcher will be high.
Marketing research is conducted to assist marketing management in taking decisions and reducing risk. It should be remembered that the research report is the only tangible output of a marketing research project. At the same time, it is documentary evidence that the research was conducted and becomes a historical record of the organization. As such, due attention must be given to the preparation of the report.
Importance of Research Report
The final step in the research process is the preparation of a research report. This is the culminating activity in the research project and can be the most important part of the research process. The research report is the only part of the research project that the client will actually see.
By definition, the purpose of marketing research is to provide information that facilitates decision making by management. Unless this information is properly communicated, even the most carefully designed and well-executed research project has a value equal to naught. There is an iron law of marketing research that “people would rather live with a problem that they cannot solve than accept a solution they cannot understand.” It simply means that the main criterion to evaluate the research report is how well it communicates the findings of research to the reader.
Writing the Report—Guidelines
Here are some guidelines to follow in the preparation of the research report.
Consider your Reader Before you start writing, carefully consider your audience who will read your report, usually marketing managers. You should paint a picture of your audience in your mind. Consider how much information he/she already has, how much detail you should provide him/her to take a decision, how much technical knowledge the reader has, and to what extent you can use technical terms.
Avoid Technical Jargon It is generally recommended to avoid technical jargon while writing a research report. The reason is simple; you may have more than one reader, and while the principal reader or some of the audience may know the meaning of the technical terms, some readers may not. Thus, it is better to use descriptive explanations. If it is necessary to use technical terms, define these terms for your reader, preferably in a glossary or appendix.
Logical Organization The report should be structured logically so that it is easy to follow. Logical structure should be visible, especially in the body of the report. This makes the parts of the report coherent and enhances clarity.
Use Headings & Subheadings The report may be divided into headings and subheadings for the topics and subtopics respectively. A topic may be the main idea of each section. The headings and subheadings are signals or signposts on a map. Topics may be in the form of a single word, phrase, sentence, or question; whatever fits the purpose of your report. However, there should be consistency in the format, font type, and font size in different levels of headings and subheadings throughout the report.
Be Objective Research is an objective and systematic method of collecting, analyzing, and interpreting data to assist marketing managers in their decision. This objectivity should be maintained when communicating the results as well. The report should accurately present the details (methodology, data analysis, results, and conclusion) without regard to the expectations of the management or client. Factual results should be presented in the report no matter whether these results are seen favorably or unfavorably by the client or user of the research.
Follow all Principles of Good Communication The report is meant to communicate the research findings. You must follow all principles of good communication while writing a research report. Some of these are:
- Use simple language. If a simple word is available instead of a hard one, use it (e.g., use instead of utilize).
- Use strong action verbs (e.g., investigate instead of “performing an investigation”).
- Generally write in active instead of passive voice (e.g., “Asghar wrote a report” instead of “The report was written by Asghar”).
- Eliminate unnecessary words. Write ‘now’ instead of ‘at this point of time’.
- Add graphs and charts to enhance understanding. However, brevity should not sacrifice completeness.
- Observe all seven Cs of communication: completeness, conciseness, coherence, clarity, correctness, courtesy, and consideration.
Professional Look The final report should give a professional look. It should be produced with good quality paper, typing, margins, headings, subheadings, and binding. The professional appearance of the report speaks about the professional work that has been carried out by the researcher.
Format of the Report
The format of the research report may vary with the purpose of the research, the researcher’s style, or the user’s instructions. If the client or user of the research wants the research report in a specific format, follow his/her instructions. Unless there are specific instructions from the organization, the following elements may be included to develop a format for the research report.
Front matter or Prefatory parts
- Title page
- Letter of transmittal
- Letter of authorization
- Table of content
- Executive summary
Body or Textual Parts
- Introduction
- Background of the problem
- Statement of the problem
- Research objectives
- Research Design and Methodology
- Type of research design
- Data collection from secondary sources
- Primary data collection
- Instrument of data collection
- Sampling techniques
- Fieldwork
- Data Analysis
- Results
- Limitations
- Conclusions and Recommendations
- References
End matter or Supplementary parts
- Appendices
⭐ Key Takeaways
The research report is the only tangible output of a marketing research project and serves as both a decision-making tool for management and a historical record for the organization. The most important criterion for evaluating a report is how well it communicates findings to the reader, following the principle that people prefer a problem they cannot solve over a solution they cannot understand. Key writing guidelines include considering the audience, avoiding technical jargon without definitions, maintaining objectivity by presenting factual results regardless of client expectations, and following all principles of good communication (including the seven Cs). Finally, a professional report must be logically organized with headings and subheadings, have a professional appearance, and follow a standard format consisting of front matter, body parts, and end matter.
🧠 Quick Revision Questions
- Why is the research report considered the most important part of the research process?
- What is the "iron law of marketing research" mentioned in the lecture, and what does it mean?
- List three specific guidelines for writing a research report, as discussed in the lecture.
- What are the three main parts of a research report's format, and give two examples of what each part contains?
- What should a researcher do if it is necessary to use technical jargon in a research report?
📘 Lecture 39 — Components of Research Report
📖 Overview: This lecture outlines the standard components of a formal research report, from front matter to appendices. Understanding these components is critical because a well-structured report ensures clarity, professionalism, and effective communication of research findings to decision-makers.
🗂️ Topics Covered
This lecture covers all major parts of a formal research report, beginning with prefatory parts like the title page, transmittal letter, letter of authorization, table of contents, and executive summary. It then details the body of the report, including the introduction, research design and methodology, data presentation, results, conclusions and recommendations, and limitations. Finally, it discusses end matter (appendices) and provides general formatting guidelines.
📝 Lecture Summary
Prefatory Parts or Front Matter
The prefatory parts or front matter are the initial sections of a research report that prepare the reader for the main content.
Title Page: The title page includes the title of the report, name and address of the researcher and the organization conducting the research, the name of the client, and the date the research report is being submitted. The title should be concise, clear, and crisp, indicating the nature of the project.
Transmittal Letter: This part of a formal research report is usually developed at the end. The transmittal letter introduces the research report to the recipient of the report. It can draw attention to particular project characteristics, contractual obligations, or noteworthy conclusions, and as such can generate interest in the subject matter of the report.
Letter of Authorization: A letter of authorization is a letter which was written by the client or person who wanted the research to be done to the researcher, authorizing him to start research for the writer. A copy of the letter is enough to be included in the report.
Table of Contents: The table of contents lists major report topics and subtopics (sections, chapters, appendices, etc.) and their beginning page numbers. It varies from being very detailed to consisting of general topic headings only. Usually, the major headings and subheadings are included in the table of contents. The table of contents is followed by a list of tables, list of graphs, list of appendixes, and list of exhibits, if any, if the report is lengthy.
Executive Summary: The executive summary is an important part of the report, as this is the only portion of the report that executives often read. The executive summary should be written after the rest of the report has been completed. An executive summary is a mini report within the report. It is not simply a brief of the report, but it is a distillation of the research project outlining the methodology, major findings, and conclusions. An executive summary is a bottom-line report created for decision makers who have no time or desire to go into the project’s technical details.
🔑 Definition — Executive Summary: A mini report within the report that distills the methodology, major findings, and conclusions for decision-makers.
Body—Textual Parts
Introduction: To begin with the body of the report, an introduction to the research background, discussions with the client and possibly the industry expert to find the direction of doing research to solve this management problem. This part provides a clear statement of the research problem and objectives/questions/hypotheses of the research. After reading this section, one can understand the reason and rationale for conducting this study.
Research Design and Methodology: A complete understanding and evaluation of a project depends on the research methodology used. Methodology — ranging from sampling frame and procedures, to mode of data collection, to research instrument, to techniques of data analysis employed — should be described adequately. Who collected data and how fieldwork was organized and monitored to ensure the quality of data collection is explained in this section. Secondary data collection methods and sources are also discussed in this section.
Data Presentation: This section may contain several chapters or subsections showing data analysis in the form of tables, description, graphs, etc. Data analysis is quantitative, qualitative, or both. This part of the body is critical for the research project as the ultimate results and conclusions of the report are based on data analysis.
Results:
- As stated earlier, results may comprise several chapters or sub-sections. Most of the time, the results are presented both at the aggregate and the subgroup level, for example, market segment, market area, wholesale, retail level.
- Tables and graphs may highlight the results with the main findings discussed in the text.
- The results should be organized in a coherent and logical way.
The standard formats that are used to arrange data in the tables are:
- Alphabetically.
- Chronologically.
- Geographically.
- According to size.
- According to interest of the reader.
- According to tradition.
- According to importance.
Conclusions and Recommendations: This part of the report gives the main findings, conclusions, and recommendations for the organization. A summary of the statistical findings is not enough. The researcher needs to discuss the results in light of the management problem being addressed to arrive at major conclusions. Based on the results, the researcher may give some suggestions to the decision-makers. Sometimes marketing researchers are not asked to recommend anything but confine themselves to giving their findings and conclusions.
Limitations: After the conclusions and recommendations component of the research report, limitations of the research are mentioned. Limitations may originate due to time, budget, or some other organizational constraints. No research project is without shortcomings. A researcher is ethically and professionally bound to fully disclose the shortcomings or setbacks of the research that may have an impact on its validity, reliability, or predictability. Sometimes limitations are given before the conclusions and recommendation section so that the readers should know the limitations of findings and conclusions.
💡 Why this matters: Disclosing limitations adds credibility to the research by showing professional honesty and helps decision-makers assess the reliability and applicability of findings.
Guidelines for Visuals: In the research report, visuals may be used for enhancing the understanding of the report. Mostly these include tables, charts and graphs, maps, figures, and flow charts. The guidelines for preparing and using these visuals are:
- Make visual aids simple and convenient to understand.
- Primary objective of including such visuals should be to augment the clarity and understanding of the content of the report.
End Matter or Supplementary Parts
This is the section of the report which contains 'too material'. It includes any material that the researcher thinks should also be included in the report to aid the understanding of the reader. It is in the form of appendixes which are labeled as Appendix A, B, C, etc. and contain the headings as well. These appendixes may range from a blank copy of questionnaire to price lists, tables, diagrams, statistical illustrations, photographs, etc.
General Guidelines
- Type or print on one side only of heavy, white, unrolled paper.
- Paper size: 8½ X 11 inches.
- Double-space the entire paper.
- Left justify text only.
- Leave a minimum one-inch margin on the sides, top, and bottom of each page.
- Number pages consecutively in the top right corner, beginning with the title page.
- Just before the page number, use a shortened form of the title as a header.
- Font size 12-point.
- Times Roman or Courier are acceptable typefaces.
- Only black toner.
- Indent paragraphs 5-7 spaces.
- No more than 27 lines of text per page.
⭐ Key Takeaways
A formal research report is structured into three main parts: prefatory parts (title page, transmittal letter, letter of authorization, table of contents, and executive summary), body (introduction, methodology, data presentation, results, conclusions/recommendations, and limitations), and end matter (appendices). The executive summary is critical because it is often the only part executives read and must distill the entire project's methodology, findings, and conclusions. The results section must be organized logically (alphabetically, chronologically, geographically, etc.) and presented at aggregate and subgroup levels. Disclosing limitations is an ethical and professional obligation that affects the perceived validity of the research. Finally, visuals should be simple and used only to enhance clarity and understanding of the content.
🧠 Quick Revision Questions
- What are the five standard components of the prefatory parts (front matter) of a research report?
- Why is the executive summary considered a "mini report within the report," and when should it be written?
- List at least three of the seven standard formats for arranging data in tables within the results section.
- What is the ethical and professional obligation regarding the "limitations" section of a research report?
- According to the general guidelines, what are the specifications for font size, typeface, and maximum lines of text per page?
📘 Lecture 40 — Citation of References in Report
📖 Overview: This lecture explains the importance of proper citation in research reports to avoid plagiarism and strengthen literature reviews. It provides a comprehensive guide to the APA style of citation, covering various source types such as books, journals, electronic sources, and multimedia. Detailed formatting rules for reference lists and quotations are also included to ensure academic integrity and consistency.
🗂️ Topics Covered
The lecture begins by introducing various style manuals available for research writing, with a focus on the APA style as the most commonly used in marketing research. It then delves into specific citation formats for books (with 1-2 authors, 3-5 authors, edited books, chapters), journal articles (with 2 authors, 3-6 authors), newspaper articles, magazine articles, and electronic sources (books, documents, journal articles, abstracts). Finally, it covers formatting rules for reference lists and the correct way to make short and long quotations.
📝 Lecture Summary
Citation of References in Report
Understanding citation protects the researcher from the offence of plagiarism. It helps in compiling Literature Review. Both instructors and students must be vigilant about citations.
Various Style Manuals available in Market
There are different manuals available to research writers to learn and use citations in the research report. Some of them include APA – American Psychological Association, MLA – Modern Language Association, Chicago Style – Chicago Manual of Style, Turabian Style – based on Chicago Style, Harvard Referencing System, ASA – American Sociological Association, CBE – Council of Biology Editors, and APSA – American Political Science Association. Mostly, the first manual listed (APA) is used in marketing research projects.
APA Style Guide
How various sources are cited in APA style is explained below with examples.
Book with 1 to 2 authors For End Notes, the format is: Author, A. A., & Author, B. B. (Year). Title of work. Location: Publisher. For In-Text citations, the format is: (Author & Author, Year, p. Page Number). When material taken is on more than two pages, include the page range (e.g., p. 11-14).
🔑 Definition — End Notes: Full bibliographic details of a source listed at the end of a research paper, usually on a separate page. 📐 Example: In-Text: (John & Spencer, 2004, p. 11-14)
Book with 3 to 5 authors For the first citation in-text, list all authors: (Author, Author, & Author, Year). For subsequent citations, use the first author's last name followed by "et al.": (Author et al., Year).
🔑 Definition — et al.: A Latin abbreviation meaning "and others," used in citations to refer to multiple authors after the first one is listed. 📌 Example: First citation: (Parkinson, Butcher, & Greenwood, 2001). Subsequent citations: (Parkinson et al., 2001).
Edited book For End Notes, include "(Eds.)" after the editors' names. The format is: Editor, A. A., & Editor, B. B. (Eds.). (Year). Title of work. Location: Publisher. 📌 Example: Gibbs, J.T., & Huang, L.N. (Eds.). (1991). Selling in the children market. San Francisco: Jossey-Bass.
Chapter from a book For End Notes, the format includes the chapter author, the year, the chapter title, the book editors, the book title, page range, location, and publisher. 📌 Example: Masaro, D. (1992). Broadening the domain of the fuzzy logical model of perception. In H.L. Pick, Jr., P. van den Broek, & D.C. Knill (Eds.), Cognition: Conceptual and methodological issues (pp. 51-84). Washington, DC: American Psychological Association.
Journal article with 2 authors For End Notes, the format is: Author, A. A., & Author, B. B. (Year). Title of article. Title of Periodical, Volume Number(Issue Number), Pages. 📌 Example: James, R., & Cramer, S. (2003). The hiring process in organizations. Consulting Psychology Journal: Practice and Research, 45(2), 10-36.
Journal article with 3 to 6 authors Similar to books with 3-5 authors, the first citation lists all authors, and subsequent citations use "et al." 📌 Example: First citation: (Kendall, Stark, & Adam, 1990). Subsequent citations: (Kendall et al., 1990).
Newspaper article with no author For End Notes, the entry begins with the article title. For In-Text, use a shortened version of the title in double quotation marks and the year. 📌 Example: In-Text: ("New Medicine," 1999).
Magazine Article For End Notes, include the specific date of the magazine issue. 📌 Example: Steiner, M.I. (2006, October 9). Measuring the mind. Science, 262, 113-114.
Electronic book retrieved from database For End Notes, specify the edition (if applicable) and state that it was retrieved from a database. 📌 Example: Naraynswamy, R. M. (2008). Fundamentals of social research (5th ed.). Retrieved from STAT! Ref database.
Document on university program or department Web site For End Notes, include the retrieval date and the full URL of the document. 📌 Example: Trapp, Y. U. (2005). Multiple intelligences: The learning process in our students. Retrieved July 1, 2006, from Yale University, Yale-New Haven Teachers Institute Web site: http://www.yale.edu/ynhti/curriculum/units/2001/6/01.06.10.x.html
Electronic journal article with 1 to 2 authors, retrieved from database For End Notes, include a DOI (Digital Object Identifier) if available. 📌 Example: Shoemaker J. I and Bradman S. W: A model for the study of celebrity preference in youth. Journal of Consumer Behaviour, 20(4), 580-588. doi:10.1177/0269881105058776.
🔑 Definition — DOI (Digital Object Identifier): A unique alphanumeric string assigned to a digital object, such as a journal article, to provide a persistent link to its location on the internet.
Electronic journal article, 3-5 authors, retrieved from database, without DOI For End Notes, state that it was retrieved from a database. 📌 Example: Tang, P., Yuan, W., & Tseng, H. (2005). Clinical follow-up study on diabetes patients participating in a health management plan. Journal of Nursing Research, 13(4), 253-261. Retrieved from CINAHL database.
Electronic journal article with 1 to 2 authors, freely available, without DOI For End Notes, provide a direct URL to the article. 📌 Example: Munch, T. J., & Barrete, N. S. (2001). Emotional intelligence, self-esteem and parental love. E-Journal of Psychology and consumer behavior, 2(2), 38-48. Retrieved from http://ojs.lib.swin.edu.au/index.php/ejap/article/view/71/100
Abstract For End Notes, if you are citing an abstract of a work, indicate "[Abstract]" after the journal information, followed by the abstract's own source. 📌 Example: Mehjabeen, A., & Fatima, N. (2002). Effects of consumer perceptions in supermarket organization development. Journal of Social Psychology, 21, 96-111. [Abstract] Psychological Abstracts, 2002, 68, Abstract No. 1122.
Video Tape For End Notes, include the producer, the year, the title, and the medium in brackets. 📌 Example: National Institute of Medical Sciences. (2008). Drug abuse [videotape]. Islamabad.
Formatting
Reference List Order
- Place the list of references cited at the end of the paper, starting on a new page.
- Begin each entry flush with the left margin, but indent subsequent lines five to seven spaces (hanging indent).
- Double space both within and between entries.
- Italicize the title of books, magazines, etc.
- Arrange sources alphabetically beginning with the author’s last name.
- If an author has more than one source, arrange entries by year, earliest first.
- When an author appears both as a sole author and as the first author of a group, list the one author entries first.
- If no author is given, begin the entry with the title and alphabetize without counting "a," "an," or "the." Do not underline, italicize, or use quote marks for titles used instead of an author name.
Capitalization in Reference List
- Capitalize only the first word of the title, the first word after a colon or dash, and proper nouns in titles of books, articles, etc.
- Capitalize all major words and all words of four letters or more in periodical titles.
How to Make a Quotations When fewer than 40 words, put prose quotations in running text and put quote marks around quoted material. The author’s last name, publication year, and page number(s) of the quote must appear in the text. 📌 Example 1: Herman (1996) states that a traumatic response frequently entails a “delayed, uncontrolled repetitive appearance of hallucinations and other intrusive phenomena” (p. 11). 📌 Example 2: A traumatic response frequently entails a “delayed, uncontrolled repetitive appearance of hallucinations and other intrusive phenomena” (Herman, 1996, p. 11).
Long Quotations When 40 words or more, present the quotation in block form. Indent it 5-7 spaces from the left margin and omit the quotation marks. If the quotation has internal paragraphs, indent the internal paragraphs a further 5-7 spaces. Double space the block quote. Cite the source after the end punctuation of the quote. 📌 Example: Meile (1993) found the following: The “placebo effect,” which had been verified in previous studies, disappeared when behaviors were studied in this manner. Furthermore, the behaviors were never exhibited again, even when real drugs were administered. Earlier studies were clearly premature in attributing the results to a placebo effect. (p. 276)
💡 Why this matters: Mastering these formatting rules ensures that your research report is professional, credible, and easy for readers to navigate, directly impacting your academic standing and the clarity of your work.
⭐ Key Takeaways
The most critical rule is to always cite sources to avoid plagiarism, with APA being the primary style for marketing research. For in-text citations, remember the "et al." rule: for sources with 3-5 authors, list all authors on first citation and use "et al." thereafter. The reference list must be organized alphabetically by author's last name, formatted with a hanging indent, and double-spaced. For quotations, the 40-word rule is critical: enclose shorter quotes in quotation marks within the text, while longer quotes must be in a separate, indented block without quotation marks. Finally, pay attention to detail for different source types—books, journals, electronic sources—as each has a specific structure for the end note and in-text citation.
🧠 Quick Revision Questions
- What is the primary purpose of citing references in a research report?
- According to APA style, how do you cite a book with four authors in the text for the first time versus subsequent times?
- What is the correct format for the end note of a journal article with two authors retrieved from a database that has a DOI?
- Describe the APA formatting rules for a block quotation (40 words or more).
- If a newspaper article has no author, how do you format the in-text citation and the end note entry?
📘 Lecture 41 — Presentation of Reports
📖 Overview: This lecture covers the process of orally presenting research results to clients or management. It emphasizes that oral presentations are often the basis for first impressions about research quality, providing guidelines for preparation, use of visual aids, body language, and effective delivery before, during, and after the presentation.
🗂️ Topics Covered
The lecture discusses the importance and strategy of oral presentations in research, preparation techniques including rehearsal and outlines, the role and various types of visual aids (transparencies, charts, videotapes, PowerPoint), the critical nature of body language and personal mannerisms during delivery, and provides detailed guidelines for effective presentation delivery before, during, and after the event.
📝 Lecture Summary
Oral Presentations
Sometimes it is desirable or mandatory for a researcher to present project results orally to the client or management. One strategy consultants follow is to initially distribute a written research report, then follow it with an oral presentation. This allows the audience to become familiar with the project before discussion. Many clients form their first impressions about the quality of the research project based on the oral presentation. The oral presentation serves as an executive overview; no attempt should be made to communicate all technical details. Decision makers are typically interested in the gist of everything, not technical details.
Preparing a Presentation
Key to a successful presentation is preparation. The presenter should be prepared to answer any question about the research process or results. Extensive rehearsals are recommended. It is also desirable for the presenter to prepare a detailed outline for assistance.
Prepare Visual Aids for Presentation
Prepare visual aids as they greatly enhance oral communication. Visual aids provide a framework for discussion. Numerical data are better understood in visual rather than verbal form, which is why tables and graphs are used. Significant points can easily be emphasized; complex ideas that are difficult to communicate otherwise can be illustrated with diagrams or pictures. Visual aids also provide variety to the presentation and help refer back to critical points for discussion. Various kinds include transparencies, charts, handouts, slides, videotapes, films, and samples. Videotapes are particularly effective for presenting proceedings of focus groups and dynamic fieldwork. PowerPoint and other software are easily available for making visuals.
💡 Why this matters: Visual aids transform abstract data into comprehensible formats, but must not dominate the presentation.
Use of Visual Aids
Visual aids should not dominate the presentation. The researcher should remain the center of attention for the audience. No one should depend on visual aids to the extent that the presentation stops if the equipment fails (e.g., electricity breaks down with no generator).
Body Language in Presentation
'What to say' is important, but 'how to say' is more important. 'How to say' distinguishes a mediocre communicator from an outstanding one. Body language must be effectively used. Body gestures clarify verbal communication. The personal mannerisms of a researcher can either help or hinder an oral presentation. A researcher with good oral communication skills and no offensive mannerisms is more effective. Avoid distracting the audience by fidgeting with key rings, pens, or other objects; remove everything from pockets except notes.
Guidelines for Preparing and Delivering Effective Presentation
Before the Presentation:
- Write an outline of the presentation
- Prepare necessary visual aids
- Check all equipment
- Have a contingency plan in case visual aid equipment fails
- Analyze your audience in terms of their reaction to research findings (agree, hostile, or indifferent). It is better to begin with ideas the audience would most likely agree with
- Practice the presentation many times; have someone witness and comment on how to improve effectiveness
During Presentation:
- Start with an overview, then go into details
- Face the audience at all times and maintain eye contact
- Talk to the audience rather than reading excessively from a script or screen
- Use visual aids effectively – they should aid the presenter
- Avoid distracting mannerisms including unnecessary movement; ensure movements have purpose
- Be concerned with your voice – not too soft, loud, fast, slow, or monotonous. Use pauses to allow the audience time to digest material
- Involve the audience
After Presentation: After completing the presentation, ask the audience if they have questions. The question-answer session is an interesting and important part of the presentation. This often concludes the talk, but audience can be permitted to ask questions during the presentation. Pause and make sure the question is understood; then give as compact a response as possible. The presenter should anticipate questions beforehand and take questions seriously.
Question – Answer Session
During the question-answer period the presenter should:
- Concentrate on the question
- Pause and repeat the question – this allows time to think about the answer
- Don't fake an answer. If you don't know the answer, say so
- Answer questions concisely but support answers with as much evidence as possible
⭐ Key Takeaways
The oral presentation is often the basis for a client's first impression of research quality and should serve as an executive overview rather than a technical deep dive. Extensive preparation, including rehearsals, audience analysis, and contingency plans for equipment failure, is critical for success. Visual aids like graphs, tables, and videotapes enhance understanding but must not dominate the presentation, with the researcher remaining the center of attention. Effective body language, avoiding distracting mannerisms, and focusing on how you say things distinguishes an outstanding communicator. The question-answer session is vital and should be handled by pausing, repeating questions, answering concisely with evidence, and never faking an answer.
🧠 Quick Revision Questions
- What is the recommended strategy for combining written and oral research reports?
- Why should an oral presentation function as an executive overview rather than communicating all technical details?
- List four ways in which visual aids enhance oral communication in research presentations.
- What is the most important contingency plan a presenter should have for visual aid equipment?
- During the question-answer session, why should the presenter pause and repeat the question before answering?
📘 Lecture 42 — Demand Forecasting
📖 Overview: This lecture explores the critical process of estimating market and sales potential for new or existing products. It details why accurate forecasting is essential for making informed decisions in marketing, production, finance, and human resources, and categorizes the various methods used to generate these forecasts.
🗂️ Topics Covered
The lecture begins by explaining the importance and accuracy of sales forecasts. It then provides a comprehensive classification of forecasting methods into Qualitative and Quantitative types. For qualitative methods, it details the Jury of Executive Judgment, Sales Force Estimates, Survey of Customer Intentions, and the Delphi Approach. For quantitative methods, it covers Time-Series Analysis (with extrapolation) and Causal Models, which include Leading Indicators and Regression Models.
📝 Lecture Summary
Demand Forecasting
Frequently, marketing researchers are requested to estimate the current market and sales potential for a new or existing product. This information is essential to configure sales territories, assign sales quotas, determine the number of salespersons needed and their compensation level, set appropriate advertising and sales promotion budgets, find new prospect accounts, drop slow products, and make new product decisions. Sales or demand potential for new or established products can be estimated.
Importance of Forecasting
The forecasting of sales or demand is a critical input to marketing decisions and making decisions in other functional areas like production, finance, and human resources. Poor forecasting would result in excessive inventory, inefficient sales expenses, heavy discounts, lost sales, inefficient scheduling of production, and poor planning for cash flow and capital investments. We should understand that forecasting provides the basis of almost all planning and control. If the forecasts are unreliable, it is most difficult to make the right tactical or strategic decision.
Accuracy of Sales Forecasts
Sales forecasting comprises numerical estimates and these are just estimates and are never absolutely correct; that is, the numerical estimates always differ from actual sales results. This can be established only after the sales have been recorded. As such, there is no direct measure of forecasting accuracy before the forecasting period. Therefore, the tactics to be closer to accurate forecasting are that we, as good researchers, should:
- Choose systematic and objective procedures and employ them adequately; and
- Select valid data sources that yield information on time and in adequate detail.
Methods of Forecasting
There is a variety of approaches that can be used for forecasting. These are classified as Qualitative and Quantitative. Quantitative methods may further be sub-divided into Time Series Extrapolation and Causal Models.
List of Forecasting Methods
- Qualitative Methods:
- Jury of executive judgment
- Sales force estimates
- Survey of customer intentions
- Delphi
- Time-Series Extrapolation:
- Trend projection
- Moving average
- Seasonal and cyclical index
- Causal Models:
- Leading indicators
- Regression models
Qualitative Methods
These methods are based on the subjective judgments of various individuals in the situation, although these individuals may have access to quantitative information about the past to aid their estimates. These individuals may get an opportunity to revise and refine their estimates, but still the estimates are subjective.
Jury of executive judgment
This method involves combining the judgment of a group of managers on the issue of forecast. A variety of concerned and informed managers representing such functional areas as marketing, sales, operations, manufacturing, purchasing, accounting, and finance are invited, combined in one place and asked to give their sales estimates for the next specified period. The estimates are consolidated and may be averaged or a range is determined. This method is widely used in forecasting, but it is mostly used to estimate the potential of consumer products and sales of service companies.
- Advantages: It is fast and efficient; it is quite timely as the forecast is generated by executives with the most current information; the forecast is based on the collective knowledge and experience of the managers.
- Disadvantages: The main disadvantage is the subjectivity of the executives.
Sales force Estimates
This method is based on the judgments of the sales force which is actually working in the field. Each salesperson is asked to give their estimate of sales in their territory for the next period. All estimates are added, and this gives a total of sales potential in all territories for the next period. Then these estimates are fine-tuned by the sales supervisors and estimates are finalized.
- Advantages: The forecasts from the sales force are drawn on complete, sensitive, and current knowledge of the customer and market, making these estimates very close to the actual.
- Disadvantages: Individual salespeople can be naturally optimistic or pessimistic. A serious bias occurs when the forecast is linked to the performance measure of the salesperson, as they may intentionally underestimate the potential of their territory so that fewer quotas are assigned and they can easily achieve them. This method is mostly used in industrial organizations.
Survey of customer intentions
In this method, customers are requested to make their own forecasts about how much of this product they intend to buy and use in the next period. The sales forecast, in turn, is worked out on the basis of their buying intentions. The sampling frame is usually the existing customer or client list, as the bulk of sales or demand usually comes from existing customers. The right person in the customer organization must be contacted. The survey of customer buying intentions works best when the number of customers, or at least the major customers, is small. As such, the maximum use of this technique is made in industrial organizations. Compared to Jury of Executive Judgment or Sales force Estimate methods, this method is more expensive and time-consuming.
Delphi Approach
The Delphi Approach is an extension of the jury of executive judgment method to refine the forecasting process. In this approach, group members are asked to make individual judgments about the forecast. Then these judgments are compiled and the whole package is returned to each member, so that they can compare their own estimate with those of the other members. The names are obscured, and codes are given instead so that the personalities or positions of some members in the group do not bias the opinion of other members. The members are asked to revise their estimates in the light of others’ judgments, and if they differ from others, state the reason why they believe their estimates are correct. They return the package to the coordinator who is conducting this session and serves as a clearing house. The coordinator forwards the revised estimates with comments of each member to other members. This process is repeated three or four times, and the group usually reaches the final forecast of sales.
Quantitative Methods
Time-Series Analysis
Time series analysis is simply the extrapolation of historical data into the next period. Statistical formulas are used to extrapolate the data into the future. Three factors are prerequisite for time-series extrapolation:
- Data must exist in time series.
- Environmental change influencing the time series can make the extrapolation err, with little ability to forecast “turning points”.
- Detection of patterns or trends in the past data must be possible.
Causal Models
Causal Models involve statistical techniques that relate historical sales data to the economic factors or forces that become the cause to increase or decrease the sales. These methods are indeed the most sophisticated sales forecasting tools. They prove to be very correct when relevant historical data on major forces causing changes in sales are available. Two methods are mostly used in the Causal Models: Leading Indicators and Regression.
Leading Indicators
This approach involves the identification of leading indicators which become the cause of the variation in sales of a good or service. These factors can move the sales of a particular good or service up or down. For example:
- Urbanization may lead to new housing.
- New housing leads to major appliance sales.
- Number of births leads to the sale of infant-related goods and services.
Regression Models
We have already studied the simple Regression model in which independent variable/s are identified and their values are input into the model to forecast the sale for a particular period or year.
- Simple regression model formula: Sales Forecast = Y = a + bX
- Multiple Regression Model formula: Sales Forecast = Y = a + b1X1 + b2X2 + b3X3
⭐ Key Takeaways
A student must remember that demand forecasting is the basis for critical decisions across all business functions. The methods are classified into qualitative (subjective judgments like Jury of Executive Opinion and Delphi) and quantitative (data-driven like Time-Series and Causal Models). Each method has specific advantages and disadvantages, with considerations for cost, time, objectivity, and context (e.g., consumer vs. industrial goods). Crucially, forecasts are only estimates and are never absolutely correct, though systematic procedures and valid data improve accuracy. For the exam, be prepared to explain each method, its use case, and its pros and cons, and recall the basic formula for a regression model.
🧠 Quick Revision Questions
- What are the four main methods of qualitative forecasting?
- Explain a key disadvantage of the Sales Force Estimates method, particularly when performance is linked to the forecast.
- How does the Delphi Approach improve upon the simple Jury of Executive Judgment method?
- What is the fundamental difference between Time-Series Analysis and Causal Models?
- Write the formula for a multiple regression model used in sales forecasting.
📘 Lecture 43 — New Product Research
📖 Overview: This lecture explores the critical role of marketing research in reducing uncertainty during new product development. It focuses on three key stages where research is essential: idea generation, concept development and testing, and test marketing, detailing specific techniques and methodologies for each.
🗂️ Topics Covered
The lecture outlines the stages of new product development where marketing research is most valuable. It first covers Idea Generation, examining methods like focus groups and benefit structure analysis to generate product ideas. Then, it details Concept Development and Testing, explaining how ideas are refined into testable concepts and the methodologies for evaluating them. Finally, it discusses Test Marketing as both a predictive and managerial tool, including formulas for sales estimation and potential problems.
📝 Lecture Summary
Innovation and New Product Development
Innovation and the development of new products are critical to the life of almost all business firms as they must adapt to a changing environment. Some uncertainty is associated with new products because, by definition, they contain aspects with which the organization is unfamiliar. Therefore, a good proportion of marketing research is directed toward reducing the risk involved in introducing new products. A marketing manager needs the support and confirmation from market researchers at various stages of new product development.
Stages in New Product Development
The known stages in the development of a new product are: idea generation, idea screening, concept development and testing, business analysis, product development and laboratory testing, test marketing or field testing, and commercialization. Marketing research may not be needed in all stages but is definitely required in stages 1, 3, and 6. The specific techniques used in different stages are different.
Idea Generation
The objective of idea generation research is to come up with completely new ideas for products, new attributes for current products, or new uses for current products. Ideas may come from various sources like salespersons, dealers, maintenance people, and customer service personnel, all of whom have direct contact with customers. A market researcher can rely on the opinions of such personnel or plan to accompany these people to listen to and refine their ideas.
Focus Groups are extensively used in idea generation. As many focus groups as time permits are used primarily as brainstorming sessions, with the objective being to generate as many ideas as possible without being critical. In these interviews, a technique called benefit structure analysis can be used, where product users identify the benefits they desire and the extent to which the current product delivers those benefits. This identifies benefits that current products are not delivering, hence the need for a new or innovated product. The researcher should seek new dimensions of consumer perception about established products. In focus groups, social and environmental trends can also be analyzed; for example, a trend of using natural foods might suggest that biscuits filled with fruit could be a good option.
Perceptual Maps are prepared by the researcher, who positions various products or brands in the market along dimensions that users perceive as critical and use for evaluation. A perceptual map can suggest gaps where new products might fit.
Concept Development and Evaluation
Concept development and testing is another stage where research is required. A concept is a fully developed and elaborated idea, which is different from a raw idea. First, the researcher translates the idea into a perceivable concept. For example, a product concept is formed that includes major attributes of the new product, relative advantage over current products, tentative price, packaging, advertising approach, and a suggested name. As there is no tangible product to test at this point, the concept must be developed and defined well enough to be clearly communicable, possibly as a verbal description or with three-dimensional models. Questions are then asked of respondents to test the concept.
Methodology of Concept Testing As the nature of most concept testing is exploratory, focus group interviews are the most frequently used technique. Usually, the discussion centers around testing one concept, but a paired comparison technique is used when the objective is to test alternate concepts. In this method, each respondent tests a set of product concepts two at a time and states which of the two is preferred. However, in a paired-comparison test, respondents may select one product over another using a very trivial attribute.
Concept Test Group The respondents for concept testing normally include people from the target segments. The aim of concept testing is to determine if a viable market exists, so no potential segment should be ignored. Generally, the concept is exposed to respondents for testing through a personal mode, in a central location like a shopping mall or a facility available with the researcher.
Objectives of Concept Testing The main purpose of concept testing is to help refine the product features, determine how it should be positioned, and suggest something about different components of the marketing mix. Such a test provides an overall indication of attitudes, interest, and likelihood of purchase by the target segment. The objectives are:
- To get a first-hand reaction of potential consumers’ views of the product idea.
- To select the most promising concepts for further development.
- To get an initial evaluation of the commercialization of the newly developed product.
Concept Testing Questions Since concept testing requires diagnostic information, questions can be posed to help determine:
- Whether the respondents comprehend the product or not.
- How do the respondents perceive the attributes of the new product?
- What are the possible advantages and disadvantages of the intended product?
- Segments/situations in which the product can be used and how frequently.
- What alternative concepts would be preferred?
- Does the concept have a crucial flaw?
Test Marketing or Field Testing
In concept testing, we test only an imaginary product, but in test marketing, we test an actual tangible product after it has been developed. Test marketing is a controlled experiment, done in a limited but carefully selected part of the marketplace, where the aim is to predict the sales or profit consequences of one or more proposed marketing actions. Here, the focus of testing is the acceptability of the newly developed product. Test marketing has two objectives: prediction of sales and managerial control.
Test Marketing as a Managerial Control Tool
- We can gain experience in physically handling the product—shelf life, breakage, storage, shipping—and identify costly mistakes to avoid them on a national basis.
- We can learn the difficulties of gaining distribution, producing a new commercial, and making our price hold at retail. This experience would be used later in our national rollout.
Test Marketing as a Predictive Research Tool To find out the potential sales of a new product, we select a test area, run the test for a desired period, and predict sales for the country as a whole using two methods:
-
Buying Income Method The sales of the test product are expanded by the ratio of the test area’s buying income to the buying income of the country. 📐 Formula: Country’s Sales Estimate = (Total Country’s income / Test area income) × Test area sales
-
The Share-of-Market Method The sales of the test brand are worked out by relating them to sales of the entire product category in the test area. 📐 Formula: Country’s Sales Estimate = (Test area sales of the new brand / Test area sales of the whole product category) × Country’s sales of the whole product category
Problems of Test Marketing
- Salespersons in the selected area are stimulated beyond normal activity.
- Special introductory offers and promotions are often made to the trade and consumers, which may not be available at the test scale for a national rollout.
- Competitors can attempt to destroy judgment ability by increasing their efforts in the test cities out of proportion with their national efforts.
- Measurement accuracy can yield ambiguous data; auditing store sales can be inaccurate due to poor store records.
- Competitors may use your test market to learn of your activities and monitor your results.
⭐ Key Takeaways
The most critical points from this lecture are the three key stages in new product development that require marketing research: idea generation, concept development and testing, and test marketing. For idea generation, focus groups and benefit structure analysis are crucial for uncovering unmet needs and market gaps. Concept development and testing uses focus groups and paired comparison methods to refine a product idea and assess its market viability before significant investment. Finally, test marketing is a controlled experiment that serves both as a managerial control tool for learning about product handling and distribution, and as a predictive tool using the Buying Income Method or Share-of-Market Method to forecast national sales. However, test marketing has several potential problems, such as competitive interference and data accuracy issues, that must be considered.
🧠 Quick Revision Questions
- What is the main difference between an "idea" and a "concept" in the context of new product development?
- Describe the "benefit structure analysis" approach used in focus groups for idea generation.
- What are the two primary objectives of test marketing?
- Explain the "Share-of-Market Method" formula for estimating a new product's national sales from test market results.
- List three potential problems that can compromise the validity of a test marketing study.
📘 Lecture 44 — Advertising Research
📖 Overview: This lecture explores the major applications of marketing research in advertising and promotion, focusing on two core areas: media research and message effectiveness (copy testing). It explains how companies measure media audiences, test ad effectiveness through various procedures, and evaluate whether an advertisement achieves recognition, recall, persuasion, and sales impact.
🗂️ Topics Covered
The lecture covers media research, including media vehicle distribution and audience measurement for newspapers, television, and radio. It then examines copy testing procedures such as consumer jury, physiological methods (eye camera, GSR, tachistoscope, brain-wave analysis), inquiry tests, on-the-air tests, trailer tests, and sales tests. Finally, it discusses criteria for good advertisements, the role of focus groups in advertising research, and sample questions used in ad testing.
📝 Lecture Summary
Media Research
Media research is a critical topic within marketing research that helps companies select the most effective media for their advertising plans. Choices must be made between various media types (television vs. radio vs. newspapers) and even between specific newspapers, television channels, or programs within a channel. Media research attempts to answer several key questions:
- Media Vehicle Distribution: How many television sets, radio sets, magazines, or newspapers carry the advertisement?
- Media Audience: How many people actually watch, listen to, or read the media in which the ad appears? Media audience is larger than vehicle distribution because more than one person may be exposed to a single vehicle.
- Exposure to Advertisement: People may be exposed to a medium but may not notice a specific advertisement. Research finds how many people were actually exposed and noticed the ad, which is less than the media audience.
- Advertising Perception: Among those who noticed the ad, how many correctly perceived and comprehended its message?
- Sales Response: How many of those exposed, who noticed and comprehended the ad, actually purchased the product in response?
Media Vehicle Distribution
Data on media vehicle distribution is readily available from several sources and is generally considered accurate. Advertising research frequently uses these sources:
- Audit Bureau of Circulation (ABC) for newspapers
- Newspapers and magazines' own reports
- Data about radio and television set sales from markets
- Government agency surveys about television users
For broadcast media, the measurement of vehicle distribution is less important compared to media audience.
Media Audiences
The media audience is defined as the number of people actually exposed to the vehicle at least once.
Newspapers Readers: A newspaper reader is one who claims to have read at least part of the newspaper in question on a given day. Newspapers sometimes collect readers' data, and advertisers rely on these data.
Television Viewers: Television viewership can be found using several methods:
- Diary Method: Household viewers record the names of shows they watched and mail the diary back to the researcher. The researcher can break down audience estimates by age, sex, and geographical area.
- Audi meter: This device is connected electronically to a computer and records what channel the television is tuned to and whether anything was being watched. The meter is placed out of view to avoid self-consciousness. However, it does not indicate how many people are watching.
- People meter: This solves the audimeter's disadvantage by allowing each family member to "log on" and "log off" their viewing time.
- Coincident telephone recall method: A sample of households is telephoned and asked what show is being watched at that time, if any, and asked to identify the sponsor or product being advertised.
- Personal interview recall method: A sample of respondents is interviewed at home shortly after the program of interest, usually during prime time shows.
Radio Audience: No formal or syndicated sources are available for radio, so the advertising researcher may rely on information collected by the media itself.
Copy Testing
Copy testing refers to testing the effectiveness of all aspects of an advertisement (color, graphics, pictures, action, etc.). It involves exposing an audience to the advertisement and observing their response. Copy testing is done at different stages of development: as a written concept, a set of drawings, an animated version, or a finished advertisement. It may be tested before or after running on media, called pretests and posttests. The final version test is ideal because it represents what people will actually see.
Ad Testing Procedures
Consumer Jury: In this procedure, 50 to 100 consumers from the target audience are interviewed either individually or in small groups.
Physiological Methods: These methods use devices to record physiological movements of viewers and draw inferences for research. They measure physiological arousal that is normally uncontrollable by the respondent, such as skin resistance, heart beat, facial expressions, muscle movement, and voice pitch.
- Eye camera: Tracks eye movement as it watches an ad to determine which sections caught and held attention. In print ads, it reveals where the eye focused, what the reader "returned to" for reexamination, and what point was "fixed on."
- Galvanic skin response (GSR): Measures response to skin by attaching electrodes from a recording device to respondents when they are exposed to ads.
- Tachistoscope: Measures the rate at which an ad conveys information or recognition.
- Brain-wave analysis: The audience is exposed to the advertisement, and attention, interest, or emotional reaction is assessed through wave analysis. Higher wave amplitude indicates more brain activity at that point. This takes place in a laboratory setting.
Inquiry Tests: These measure ad effectiveness based on consumer inquiries that result directly from the ad placed mostly in newspapers or magazines. It provides a direct measure of response with no interview, reducing costs and artificial reactions.
On-the-Air Tests: Over 100 respondents are contacted by telephone in big cities who claim they watched a particular television show the night before. They are asked what they remember about the specific ad, the sales points in the ad, and whether they had any favorable attitude toward the ad. This method is mostly used in television copy testing.
Trailer Tests: Respondents are chosen from shopping malls and taken to a trailer or room in the mall. They are shown several ads with or without surrounding programming and asked questions including recall tests.
Sales Tests: In standard advertising tracking, respondents matching the target market profile are interviewed personally or by telephone to measure their levels of awareness, attitudes about the ad and brand, and recent purchases after being exposed to the ad.
Criteria for Good Advertisement
The criteria for a good advertisement message are:
- Advertisement recognition: Recognition is a necessary condition for effective advertising. If the advertisement cannot pass this minimal test, it probably will not be effective.
- Recall of its contents
- How does it persuade?
- Impact on purchase behavior (i.e., purchase)
Focus Group in Advertising Research
Focus-group research is widely used in the development of an advertising campaign. Focus groups are mainly used to generate ideas for advertisements and to test reactions to rough executions. Opinions about advertisement concepts and actual advertisements are sought, including audience impressions about what the ad was, what ideas were presented, and interest in those ideas. The goal is to detect potent misperceptions as well.
Sample Questions in Ad Research
Typical questions include:
- Do you remember seeing this ad on TV? (Yes / No / Not sure-I may have)
- What did you see in it?
- How much interested are you in what this ad is trying to show you? (Very interested / to some extent / not interested)
- How does it make you feel about the product? (It's a good product / It's Ok / It's bad / Not sure)
- Please check whether this commercial was: (Appealing / Clever / Confusing / Convincing / Dull / Effective / Interesting / Irritating)
💡 Why this matters: Advertising research directly impacts budget allocation, creative strategy, and campaign effectiveness. Knowing which media reach the right audience and which messages generate recall, persuasion, and sales allows companies to maximize return on their advertising investment.
⭐ Key Takeaways
The most critical concepts from this lecture are the distinction between media vehicle distribution and media audience, the various methods for measuring television viewership (diary, audimeter, people meter, coincident telephone recall, personal interview recall), and the different copy testing procedures (consumer jury, physiological methods, inquiry tests, on-the-air tests, trailer tests, sales tests). A good advertisement must achieve recognition, recall, persuasion, and impact on purchase behavior. Focus groups are essential for generating ad ideas and testing rough executions, and the criteria for ad effectiveness provide a framework for evaluating whether an advertisement passes minimal tests for success.
🧠 Quick Revision Questions
- What are the five key questions that media research attempts to answer about an advertisement?
- Explain the difference between an audimeter and a people meter in television audience measurement.
- What are the four criteria for a good advertisement message according to the lecture?
- Describe how an eye camera is used in physiological testing of advertisements.
- In on-the-air tests, how are respondents contacted and what are they asked about the advertisement?
📘 Lecture 45 — International Marketing Research
📖 Overview: This lecture explores the complexities and unique challenges of conducting marketing research in international markets. It highlights why research is essential for avoiding costly mistakes and making informed decisions in diverse global environments, and examines specific issues related to research design, data collection methods, and cultural equivalence.
🗂️ Topics Covered
The lecture begins by establishing the need for international marketing research due to increasing global complexity and management's lack of familiarity with foreign markets. It then delves into the inherent complexity and diversity of the international environment, covering differences in consumer behavior, infrastructure, and regulations. The discussion moves to specific information needs based on a firm's market experience, followed by a deep dive into critical issues like establishing comparability and equivalence, and various survey methods adapted for global contexts. The topics of questionnaire translation and key variable importance indicators for market assessment are also covered, concluding with practical considerations for conducting international research and a recap of the entire course.
📝 Lecture Summary
Need for International Marketing Research
As the international environment becomes more complex and management of many domestic firms lacks familiarity with foreign markets, it becomes all the more important to undertake research prior to making international marketing decisions and marketing strategy. This is equally important for decisions relating to initial market entry, product positioning, marketing mix, or subsequent expansion decisions. Research will save us from costly mistakes in marketing and loss of valuable opportunities in international markets.
Complexity of International Marketing Research
Although marketing research follows the same six steps as domestic research—Understanding the management dilemma, Defining research problem and developing research objectives, Formulating research design, Collecting data/Fieldwork, Analyzing data, and Writing research report and oral presentation—it is more complex. Marketing on global scales poses problems that are inherently more complex than those encountered in a firm’s domestic market. Operations take place on a much broader scale and scope, often involving a range of different types of activities and management systems including licensing, strategic alliances and joint ventures.
International marketing entails operation in a variety of diverse environmental contexts. These range from the mature industrialized markets of Europe, the US and Japan, the unstable but blossoming markets of Latin America, the politically uncertain markets of the Middle East or Russia, and the volatile markets of South East Asia to the emerging African markets. International markets are also characterized by rapid rates of change in the technological, economic, social and political forces that shape their development. Change is rapid and all pervasive, but as well as unpredictable, altering the nature of opportunities and threats in international markets. Research aids in assessing where the best opportunities lie, where and how to enter new markets.
Diversity of International Environment
Diversity occurs particularly in relation to consumer tastes, preferences and behavior, and to a lesser extent, business-to-business markets. The banking system, the structure of distribution adds a further level of complexity to strategy development and implementation. This, in turn, is further compounded by government regulation of business operations, product formulation and packaging, advertising, promotion, pricing as well as trade barriers such as tariffs, import quotas, etc. Level of literacy also varies from country to country. While levels of literacy in industrialized countries are typically 99%, it is important to remember that is far from the case in other countries.
Information Needs
Information needs vary depending on the firm’s experience and degree of involvement in international markets. In the initial phase of entry into international markets, information is needed to assess opportunities and risks in different countries. Key decisions require specific information:
- Which markets and target segments will be entered?
- Which mode of entry and operation should be adopted for specific target markets?
- What should be the timing for entry?
- How marketing resources must be allocated between different levels of marketing management (product/product line level, customer level and market segment/ country market level)?
- How to establish a control system to monitor performance in the target market?
Issues in International Marketing Research
Complexity of Research Design and Difficulties in Establishing Comparability and Equivalence are two predominant issues. The relevant respondent may differ from country to country. For example, in European countries, children play an important role in decisions related to the purchase of chocolate or cereals, while in other countries which are less child oriented, the mother may be the relevant decision maker. Equally, the role of women is enhancing in financial and insurance decisions in some societies, but in Arab society, this is rarely the case. Secondary data such as data on motor vehicle registrations may not provide equivalent data between many countries.
Survey Methods
Telephone Interviewing and CATI
Telephone interviewing is the dominant mode of questionnaire administration. However, telephone (land line) penetration is still not complete in rural areas. In developing countries, only a few households have telephones. Telephone incidence is low in Africa. India is a predominantly rural society where the penetration of telephones is less than 10% of households in the villages. With the decline of costs for international telephone calls, multi country studies can be conducted from a single location. This greatly reduces the time and costs. Cell phone penetration is high. Computer-assisted telephone interviewing (CATI) facilities are well developed in the United States and Canada and in some European countries, such as Germany.
Mall Intercept
In North America, many marketing research organizations have permanent facilities in malls, equipped with interviewing rooms, kitchens, observation areas, and other devices.
Mailed Questionnaires
Because of low cost, mail interviews continue to be used in most developed countries where literacy is high and the postal system is well developed. In Africa, Asia, and South America, however, the use of mail surveys and mail panels is low because of illiteracy and the large proportion of the population living in rural areas.
Electronic Surveys
In the United States and Canada, the use of e-mail and the internet is growing by leaps and bounds. Use of these methods for conducting survey is growing not only with business and institutional respondents, but also with households.
Questionnaire Translation
The questions may have to be translated for administration in different cultures. Direct translation, in which a bilingual translator translates the questionnaire directly from a base language to the respondent’s language, is frequently used. Procedures such as back translation and parallel translation have been suggested to avoid errors. In back translation, the questionnaire is translated from the base language by a bilingual speaker whose native language is the language into which the questionnaire is being translated. This version is then retranslated back into the original language. Translation errors can then be identified. An alternative procedure is parallel translation, where a committee of translators translate the questionnaire simultaneously and the translations are compared to decide on the final version.
Variable Importance Indicators
The lecture outlines several key variable categories for assessing international markets:
- Economic: Measure of Economic Wealth, Macro level Indicator of Market Potential, etc. Indicators include GNP, GNP per capita, Population, Inflation, Unemployment Rate, Interest Rates.
- Political: Measure of Political Stability and Political Risk, Government’s Attitude towards business, etc. Indicators include Type of Government, Expert ratings of political stability.
- Legal: Measure of legal risk, protectionism, marketing mix strategies, etc. Indicators include Import-Export laws, Tariffs, Non-tariff barriers, taxes, copyright laws.
- Socio-Cultural: Measure of High/Low Context Cultures, Attitude of people, Differences in lifestyles. Indicators include Religion, language, literacy, values, work ethics, role of family, gender roles.
- Infrastructural: Measure of technological advancement, available media and their relative influence. Indicators include Energy Costs, Extent of computerization, No. of Telephones, Fax Machines, Presence of Mass media.
Personnel
International research needs a real commitment in terms of personnel resources, such as sampling experts, telephone interviewing experts, and executives with appropriate skills.
International Marketing Research in Practice
There are many major research agencies with multinational operations that provide the benefit of coordinating the project from the home country and assuring the clients of comparability. These agencies also ensure that they have local staff in all of these countries that are familiar with the local culture and traditions and will be in a position to provide better insight about the market. Quotations from two or three international research organizations are obtained. Price is not necessarily the deciding factor. A cheaper quote may mean less rigorous procedures.
💡 Why this matters: The practical execution of international research requires balancing centralized coordination for comparability with local expertise for cultural insight, and price should not be the primary criterion for selecting a research agency.
Recap of the Course
Now that you have completed the course, you should be able to:
- Understand the management dilemma and the decision making situation confronting the marketing manager.
- Discuss and finalize the research problem.
- Write the marketing research objectives.
- Review the related literature and develop research questions and or research hypotheses.
- Prepare a research design.
- Determine sample size and select the sample using an appropriate sampling method.
- Develop the data collection instrument appropriate for your research project.
- Collect data and monitor the field work.
- Analyze data using appropriate statistical techniques.
- Write a professional research report and give oral presentation if required by the management/client.
⭐ Key Takeaways
International marketing research follows the same six steps as domestic research but is far more complex due to operating in diverse and rapidly changing environmental contexts. The critical issues for a student to remember are the need to establish comparability and equivalence across different countries—this affects everything from identifying the relevant respondent (e.g., children vs. mother for purchase decisions) to choosing survey methods (e.g., mail surveys are ineffective where literacy is low). Questionnaire translation is a major challenge, and using back translation or parallel translation is essential to avoid errors. Finally, when assessing international markets, one must systematically consider economic, political, legal, socio-cultural, and infrastructural variable indicators, and be aware that local staff familiar with the culture is crucial for effective practice.
🧠 Quick Revision Questions
- What are the six steps of the marketing research process, and why is applying them internationally more complex than domestically?
- Explain the two predominant issues in international marketing research concerning respondent selection and data comparability. Provide an example for each.
- Why might mail questionnaires be a poor choice for conducting survey research in many developing countries, and what alternative survey methods might be more appropriate?
- Describe the difference between "back translation" and "parallel translation" for questionnaire translation. Which procedure is more collaborative?
- List the five categories of variable importance indicators for assessing international markets and provide one example indicator for each category.