PSY631 — Midterm Summary (Lectures 1–22)
📘 Lecture 1 — Psychological Testing and Measurement (PSY-631)
📖 Overview: This introductory lecture establishes the foundational context for psychological testing by directing students to key professional resources and information sources. It emphasizes the importance of consulting authoritative bodies like the American Psychological Association (APA), test manuals, and internet-based repositories for understanding and evaluating tests.
🗂️ Topics Covered
The lecture covers the identification of key APA divisions relevant to psychological testing, such as Division 5 (Evaluation, Measurement, and Statistics) and Division 14 (Society for Industrial and Organizational Psychology). It also outlines essential sources of test information, including test manuals, catalogues, and internet sites.
📝 Lecture Summary
Some Sources of Information on Tests
This section lists the primary channels through which students and professionals can access reliable information about psychological tests. It begins by directing attention to the American Psychological Association (APA) and its specific divisions that focus on testing and measurement.
🔑 Definition — APA Divisions: Specialized subgroups within the American Psychological Association that focus on particular areas of psychology. For testing, Division 5 (Evaluation, Measurement, and Statistics) is the primary relevant division.
📐 Formula: N/A
📌 Example: The lecture provides two specific links for accessing APA division information:
- For Division 5, Evaluation, Measurement, and Statistics: http//www.apa.org/about/division/div5.html
- For Division 14, Society for Industrial and Organizational Psychology: http//www.apa.org/about/division/div14.html
💡 Why this matters: Knowing which APA division to consult allows a psychologist to find peer-reviewed, professionally endorsed information about test development, validity, and proper usage.
The section continues by listing additional sources of information:
- Test manuals and catalogues: These are official documents provided by test publishers. A test manual contains the technical details of a test’s development, its reliability, validity, norms, and instructions for administration and scoring. Catalogues provide an overview of available tests from a particular publisher.
- Internet sites: The internet serves as a vast repository of information, allowing access to test reviews, publisher websites, professional forums, and databases for research on specific tests.
🔑 Definition — Test manual: The official, comprehensive document that accompanies a standardized test, detailing its construction, psychometric properties, and administration procedures.
📐 Formula: N/A
📌 Example: A test manual for the Wechsler Adult Intelligence Scale (WAIS) would include chapters on its theoretical foundation, standardization sample (e.g., age ranges, demographics), reliability coefficients (e.g., split-half, test-retest), validity evidence (e.g., content, construct, criterion), and step-by-step administration and scoring rules.
💡 Why this matters: Relying solely on internet summaries or second-hand descriptions of a test can lead to misuse. Accessing the primary test manual ensures the test is administered and interpreted according to the developer's validated procedures.
⭐ Key Takeaways
A student must remember that the primary professional source for information on psychological tests is the APA, specifically Division 5 for evaluation and measurement. Test manuals are the definitive, technical source for understanding a test’s psychometric properties, while catalogues help identify available tests. Internet sites are a useful supplementary source, but must be used critically. Knowing these information sources is the first step toward responsible and competent test selection, use, and interpretation.
🧠 Quick Revision Questions
- Which APA division is most directly concerned with evaluation, measurement, and statistics?
- What is the primary difference between a test manual and a test catalogue?
- How can a student find a list of all tests published by a specific company?
- Why is it considered essential to read the test manual before administering a psychological test?
- Name two specific types of professional information that can be found on internet sites related to psychological testing.
📘 Lecture 2 — Historical Background of Psychological Testing
📖 Overview: This lecture traces the historical roots of psychological testing from ancient China through the 19th and early 20th centuries in Europe and America. It explains how early civil service examinations, Darwinian theory, and pioneering psychologists like Galton, Cattell, and Binet laid the foundation for modern mental measurement. Understanding this history is essential because it reveals the origins of key concepts such as individual differences, standardized testing, and intelligence measurement.
🗂️ Topics Covered
The lecture begins with the origins of mental testing in ancient China, describing civil service examinations from 2200 BC through the Ming Dynasty. It then covers the spread of Chinese testing methods to Europe and America. The summary of key developments includes Darwin’s influence on studying individual differences, contributions of German experimental psychologists, and the pioneering work of Galton (hereditary genius, sensorimotor tests, correlation), Cattell (first use of "mental test"), and Binet (first effective intelligence test with Simon). The lecture concludes with other key pioneers in testing.
📝 Lecture Summary
Historical Background of Psychological Testing
Major developments in psychological testing took place in the West, especially the U.S., but the earliest roots are in the Orient. China was the first country to develop and use tests in the formal sense of mental measurement. The Greek philosophers Plato and Aristotle proposed ideas about individual differences around 300-400 BC, but the Chinese had a system of mental measurement even in 2200 BC.
More than 4000 years ago, the Chinese developed a civil service testing program. Under Chinese emperors, oral examinations were held every third year for work evaluations and promotions. During the Han Dynasty (206 BCE to 220 CE), test batteries (multiple measures) were common, covering topics like revenue, civil law, military affairs, agriculture, and geography.
The system broadened during the Chan dynasty (beginning 1115 BCE), covering archery, horsemanship, music, writing, arithmetic, civil law, agriculture, revenue, military affairs, geography, and social rites. By the Ming Dynasty (1368-1644 CE), a national multistage testing program existed, with tests at local and regional centers using special testing booths. Successful candidates at local level went to provincial capitals for extensive essay exams, then to the nation’s capital for a final round. Only those who passed the third set were eligible for public office. This civil service system prevailed until 1905.
The Chinese pattern was copied by other nations. In 1832, the English East India Company copied the Chinese model for selecting employees for overseas duty. The German and French governments later adopted it. In 1883, the U.S. Civil Service Commission was established, developing competitive exams for government jobs. This marked a period when the testing movement grew rapidly in the Western world.
By the sixteenth century, European society had become more advanced and capitalistic, with greater recognition of individuality. Major developments occurred during the Renaissance, with a rebirth of individualism. By the early nineteenth century, human knowledge was gathered through observation, but physical scientists were developing more precise instruments. The most significant shift came after Charles Darwin wrote On the Origins of Species in 1859. Darwin proposed that humans descended from apes as a result of chance variation, and that species are selected by nature based on adaptability and survival value (the "survival of the fittest").
Darwin generated scientific interest in the study of individual differences, noting that slight differences in offspring provide materials for natural selection. Around the same time, Gustav Fechner, Wilhelm Wundt, Hermann Ebbinghaus, and other German experimental psychologists showed that psychological phenomena could be expressed in quantitative, rational terms. U.S. psychologists also studied and reported on individual differences, and American experts began developing standardized measures of scholastic achievement. In France, psychiatrists and psychologists studied mental disorders, influencing the development of clinical assessment techniques.
The most significant names that initiated and contributed to the study of individual differences and test development in the 19th century included Sir Francis Galton, James McKean Cattell, and Alfred Binet.
🔑 Definition — Individual differences: The variations among people in their psychological traits, abilities, and behaviors, which became a central focus for scientific study after Darwin’s work.
📌 Example: Darwin wrote that “the many slight differences which appear in the offspring from the same parents... may be called individual differences... [they] afford materials for natural selection to act on.”
Sir Francis Galton
Galton was a cousin of Charles Darwin, born into a family of geniuses with an IQ over 200. He was a geographer, meteorologist, tropical explorer, "founder of differential psychology," inventor of fingerprint identification, pioneer of statistical correlation and regression, and a proponent of hereditarianism and eugenics. He gave the concept of "hereditary genius" in his 1869 book Hereditary Genius, arguing that gifted individuals tend to come from families with other gifted individuals. He analyzed biographical dictionaries and encyclopedias, concluding that talent in science, professions, and arts ran in families.
He attempted to measure human traits quantitatively to determine the distribution of heredity, using a word association test and tests of mental imagery. Galton argued it would be "quite practicable to produce a highly gifted race of men by judicious marriages during several consecutive generations." He defined eugenics as the study of agencies under social control that may improve or repair racial qualities of future generations, either physically or mentally. He stated, "What Nature does blindly, slowly, and ruthlessly, man may do providently, quickly, and kindly," and "Intelligence must be bred, not trained." These arguments appealed to many and were taken to extremes, supporting apartheid policies, sterilization programs, and withholding basic human rights from minority groups.
Galton was interested in the hereditary basis of intelligence and techniques for measuring abilities. He constructed sensorimotor tests and devised methods for investigating individual differences in abilities and temperament. Using these tests, he collected measurements of over 9000 people aged 5 to 80. Among his methodological contributions was the technique of "co-relations" (correlation), which remains popular for analyzing test scores.
🔑 Definition — Eugenics: The study of the agencies under social control that may improve or repair the racial qualities of future generations, either physically or mentally.
🔑 Definition — Correlation (co-relations): A statistical technique developed by Galton for analyzing relationships between test scores, showing how variables are associated.
📌 Example: Galton collected measurements from over 9000 people using sensorimotor tests to study individual differences in abilities and temperament.
James McKeen Cattell
James McKeen Cattell was an American psychologist who gave more importance to mental processes. He was the first ever to use the term "mental test" for devices used to measure intelligence. He developed tasks aimed to measure reaction time, word association, keenness of vision, and weight discrimination. These tests proved to be a failure because they were not comprehensive and complex enough to measure intelligence. Cattell joined Galton in his methods and tests, trying to relate scores on mental tests of reaction time and sensory discrimination to school marks.
🔑 Definition — Mental test: A term first used by James McKeen Cattell to refer to devices used to measure intelligence, such as reaction time and sensory discrimination tasks.
📌 Example: Cattell developed tasks measuring reaction time, word association, keenness of vision, and weight discrimination, but these tests failed because they were not comprehensive enough to predict intelligence.
Alfred Binet and Theodore Simon
It remained for the Frenchman Alfred Binet to construct the first mental test that proved to be an effective predictor of scholastic achievement. Binet, along with Theodore Simon, developed the first formal measure of intelligence in 1905 in France. The test or scale was developed to assist the education ministry in identifying "dull" students in the Paris school system so they could be provided remedial aid.
The main idea was that intelligence could be measured in terms of a child's performance. The test could identify more intelligent children within a particular age group and differentiate intelligent children from less intelligent ones. The testing procedure involved: (1) Binet developed a number of tasks; (2) he took groups of students categorized as 'dull' or 'bright' by teachers; (3) tasks were presented, and those completed by 'bright' students were retained; (4) retained tasks were considered indicative of intelligence; (5) with further work, dull or bright children could be identified with reference to their age.
The original Binet-Simon scale was revised multiple times. American psychologist Lewis Terman gave the first Stanford revision in 1916, covering American standards from age 3 to adulthood. Further revisions were made in 1937 and 1960. The Stanford-Binet remains one of the most widely used tests today. The 1905 test was an individually administered test of 30 problems arranged in ascending difficulty, emphasizing the ability to judge, understand, and reason. The 1908 revision contained subtests grouped by age levels from 3 to 13 years. In scoring the 1908 revision, the concept of mental age was introduced to quantify overall performance.
🔑 Definition — Mental age: A concept introduced in the 1908 revision of the Binet-Simon Intelligence Scale as a way of quantifying an examinee’s overall performance on the test, representing the age level at which a child performs intellectually.
📌 Example: A child who performs at the level of an average 8-year-old would have a mental age of 8, regardless of their chronological age.
Other Pioneers in Psychological Testing
Among other pioneers were Charles Spearman in test theory, Edward L. Thorndike in achievement testing, Lewis Terman in intelligence testing, Robert Woodworth and Hermann Rorschach in personality testing, and Edward Strong in interest measurement. The work of Arthur Otis on group-administered tests of intelligence led directly to the construction of the Army Examinations Alpha and Beta by a committee of psychologists during World War I.
🔑 Definition — Army Examinations Alpha and Beta: Group-administered intelligence tests developed during World War I, based on Arthur Otis's work, used for screening military recruits.
⭐ Key Takeaways
The most critical points from this lecture are that psychological testing originated in ancient China with civil service examinations over 4000 years ago, which later influenced testing systems in Europe and America. Darwin’s theory of evolution and concept of individual differences provided the scientific impetus for studying human abilities, while Galton pioneered sensorimotor tests, hereditarian views, and the statistical technique of correlation. Cattell introduced the term "mental test" but failed to predict intelligence with simple tasks, and it was Binet and Simon who created the first effective intelligence test in 1905, introducing the concept of mental age to identify students needing remedial education.
🧠 Quick Revision Questions
- What was the earliest known formal mental measurement system, and when did it exist?
- How did Charles Darwin’s work influence the study of individual differences in psychology?
- What was Galton’s concept of "hereditary genius," and what testing methods did he develop?
- Why did James McKeen Cattell’s mental tests fail to effectively measure intelligence?
- What was the purpose of the first Binet-Simon Intelligence Scale (1905), and what concept was introduced in its 1908 revision?
📘 Lecture 3 — Historical Background of Psychological Testing
📖 Overview: This lecture traces the historical development of psychological testing from ancient Chinese civil service exams to modern standardized tests. It explains how early pioneers like Esquirol, Seguin, Galton, Cattell, and Binet laid the foundation for intelligence testing, and how World War I accelerated the development of group testing. Understanding this history is crucial for appreciating why tests are designed and used the way they are today.
🗂️ Topics Covered
The lecture begins with early formal assessment systems in China and Greece, then covers 19th-century contributions by French physicians Esquirol and Seguin regarding mental retardation. It discusses Francis Galton's sensorimotor tests and anthropometric laboratory, James McKeen Cattell's introduction of the term "mental tests," and the influence of experimental psychology. The core focus is on Alfred Binet's development of the first intelligence scale (1905-1911), its revision as the Stanford-Binet, and the concept of IQ. Finally, it covers the emergence of group testing during World War I with Army Alpha and Beta, concluding with a timeline of significant milestones in testing history.
📝 Lecture Summary
Historical Background of Psychological Testing
The lecture establishes that while the Chinese had the first formal assessment system for civil servants (2200 B.C.), other societies like the Greeks used tests in education. European universities developed formal examination systems, awarding degrees based on exams. Modern psychological tests began in the 19th century, with key figures being Francis Galton, James McKeen Cattell, and Alfred Binet. Binet developed the first proper intelligence test to identify children who could not benefit from normal schooling, leading to the establishment of the Ministerial Commission for the Study of Retarded Children. Before Binet, professionals were increasingly recognizing the need for diagnostic tools to differentiate between the mentally retarded (born with intellectual deficit) and the "insane" (those with extreme emotional problems).
Esquirol
Esquirol was a French physician who published a two-volume work in 1838 that was a milestone in understanding and treating mental retardation. He described mental retardation as existing along a continuum from "normality" to "low grade idiocy" and found it in various degrees. He proposed that language use was the most reliable criterion for assessing intellectual level after trying other procedures. This idea proved influential, as most modern intelligence tests include verbal ability as a major component.
🔑 Definition — Esquirol (Jean-Étienne Dominique Esquirol): A French physician whose 1838 work established a comprehensive account of mental retardation, proposing that language use is the best criterion for assessing intellectual level.
Seguin
Seguin, another French physician, emigrated to the U.S. in 1848. He established the first school for the mentally retarded in 1837 and developed a "physiological method of training," believing mental retardation was not incurable. He created muscle and sense training techniques, including the Seguin Form Board, which contained variously shaped slots and corresponding blocks that the subject had to insert correctly. This type of performance-based item was later adopted for performance-based intelligence tests.
🔑 Definition — Seguin Form Board: A performance test developed by Seguin containing slots of various shapes and corresponding blocks; the subject inserts the blocks into the correct slots, forming a basis for later performance-based intelligence test items.
💡 Why this matters: Esquirol and Seguin's work established that mental retardation could be systematically assessed and trained, shifting views from incurable to treatable and moving toward objective measurement.
Sir Francis Galton
Sir Francis Galton was the most prominent figure in the 19th-century testing movement. He was primarily interested in the inheritance of genius and developed sensorimotor tests to measure intellectual ability. He believed that the more perceptive the senses are of differences, the larger the field for judgment and intelligence to act. He collected data from over 9,000 people (ages 5-80) at his anthropometric lab at the International Exposition of 1884. His tests included:
- Galton whistle for determining the highest audible pitch
- Galton bar for visual discrimination of length
- Graduated series of weights for measuring kinesthetic discrimination He also introduced rating scales, questionnaire methods, and the free association technique.
🔑 Definition — Anthropometric Lab (Galton's): A laboratory established by Galton at the 1884 International Exposition to collect measurements of physical and sensory traits from thousands of people, aiming to study individual differences and the inheritance of genius.
📐 Formula: Sensory ability → Measure of intellectual ability (Galton's premise). 📌 Example: Galton used a whistle to find the highest pitch a person could hear, assuming that better sensory discrimination indicated higher intelligence.
James McKeen Cattell
James McKeen Cattell was the first person to use the term "mental tests" in psychology, in an 1890 article. He wrote his doctoral dissertation on reaction time under Wilhelm Wundt in Leipzig, Germany. After lecturing in Cambridge, he interacted with Galton, strengthening his interest in individual differences. Cattell's mental tests measured speed of movement, muscular strength, sensitivity to pain, keenness of hearing and vision, memory, weight discrimination, and reaction time. However, these tests were not good predictors of intellectual functioning as their results did not correspond with teachers' ratings or academic grades, limiting their popularity.
🔑 Definition — Mental tests (term origin): The term "mental tests" was first used by James McKeen Cattell in 1890 to describe a series of tests measuring sensory discrimination and reaction time to assess intellectual ability.
The Experimental Psychologists
Experimental psychologists of the 19th century, like those in Wundt's Leipzig lab, contributed indirectly to test development. While their main focus was on uniform patterns of behavior, not individual differences, their experiments revealed individual differences in behavior under the same conditions. They also generated the realization that objectivity required carefully designed measuring instruments and controlled conditions, leading to the formation of standardized procedures adopted in test development and administration.
20th Century: Binet's Scale and the Stanford-Binet
Alfred Binet's scale (first version 1905) was the first formal test of intelligence, consisting of 30 problems arranged in ascending difficulty. It was given to 50 normal children (ages 3-11) and some mentally retarded children and adults. Items covered sensory and perceptual tests as well as verbal content, with special emphasis on judgment, comprehension, and reasoning. In the 1908 revision, unsatisfactory tests were removed and new ones added. The scale was given to 300 normal children (ages 3-13) and was grouped into age levels. Tests passed by 80-90% of normal 3-year-olds were placed at the three-year level. This allowed the scale to determine a child's mental level (also called mental age). A further revision was made in 1911. The most significant translation and adaptation was by L. M. Terman and associates at Stanford University, called the Stanford-Binet. The first Stanford revision was in 1916, with further revisions in 1937 and 1960. The concept of Intelligence Quotient (IQ) was first used in the Stanford-Binet.
🔑 Definition — Intelligence Quotient (IQ): A ratio between a person's Mental Age (MA) and their Chronological Age (CA), multiplied by 100. IQ = (MA/CA) x 100.
📐 Formula: IQ = (Mental Age / Chronological Age) × 100 → This expresses how a person's intellectual performance compares to others of the same age. 📌 Example: A 7-year-old child with a mental age of 10 would have an IQ of (10/7) × 100 ≈ 143. This indicates performance well above average. However comparing a 22-year-old with MA 25 (IQ ≈ 114) vs. a 7-year-old with MA 10 (IQ ≈ 143) is problematic, showing the limits of this ratio.
💡 Why this matters: IQ allowed for a standardized comparison of intellectual ability across ages, but the ratio formula created problems comparing people of different ages, leading to later statistical refinements.
Group Testing
Group testing emerged prominently in early 20th century as a way to administer tests to many people simultaneously, rather than individually. This became critical when the U.S. entered World War I in 1917. The American Psychological Association appointed a committee to help, and the army needed to classify approximately 1.5 million recruits by intellectual level for selection, retention, duty allocation, and training. Arthur S. Otis had prepared an unpublished group intelligence test using multiple-choice questions. Using available tests including Otis's, army psychologists developed Army Alpha (for literate recruits) and Army Beta (for illiterate or non-English speaking recruits) — the first group intelligence tests. After this, test development exploded across all purposes.
🔑 Definition — Group Testing: A method of test administration where a test is given to many individuals simultaneously, allowing for quick screening of large populations, as exemplified by the Army Alpha and Beta tests developed during World War I.
Significant Milestones in the History of Psychological Testing
The lecture provides a detailed timeline of key events in testing history, including:
- 2200 B.C.: Chinese civil-service testing program
- 1879: First psychological laboratory (Wilhelm Wundt, Leipzig)
- 1884: Galton's Anthropometric Laboratory
- 1905: First Binet-Simon intelligence scale
- 1916: Stanford-Binet Intelligence Scale published
- 1917: Army Alpha and Beta constructed
- 1921: Hermann Rorschach's Inkblot Test published
- 1939: Wechsler-Bellevue Intelligence Scale published
- 1947: Educational Testing Service founded
- 1975: Growth of behavioral assessment techniques
- 1989: MMPI-2 published
⭐ Key Takeaways
The key concepts a student must remember are that psychological testing originated from ancient civil service exams but modern tests began in the 19th century with Galton's sensorimotor approach, Cattell's term "mental tests," and the crucial clinical contributions of Esquirol and Seguin distinguishing mental retardation from insanity. The most critical development was Binet's first formal intelligence test (1905), later revised as the Stanford-Binet by Terman, which introduced the concept of IQ as a ratio of mental age to chronological age. The need for mass screening during World War I led to the first group intelligence tests (Army Alpha and Beta), revolutionizing testing by allowing administration to large groups. Finally, the history reveals a shift from sensory/physical measures to verbal and reasoning-based tests, and from individual administration to group formats.
🧠 Quick Revision Questions
- Who was the first person to use the term "mental tests" in psychology, and in what year?
- What was the main criterion Esquirol proposed for assessing intellectual level in people with mental retardation?
- What is the formula for Intelligence Quotient (IQ) as introduced in the Stanford-Binet, and what are its two components?
- What were the names of the first group intelligence tests developed during World War I, and what problem did they solve?
- Who developed the Seguin Form Board, and what type of test items did it influence?
📘 Lecture 04 — Types of Tests and Their Significance
📖 Overview: This lecture establishes the fundamental assumptions underlying all psychological testing and assessment, demonstrating why and how tests are valid tools for measuring human behavior. It then provides a comprehensive taxonomy of test types based on administration method, content, and purpose, and concludes by examining the three major real-world contexts where these tests are applied: education, occupation, and clinical settings.
🗂️ Topics Covered
The lecture begins with twelve basic assumptions of psychological testing and assessment, which explain the theoretical foundations for measuring traits and states. It then categorizes tests into seven types: individual vs. group, ability/achievement/aptitude, intelligence vs. personality, speed vs. ability, structured vs. projective, verbal vs. nonverbal/performance, and commercial vs. available-to-all tests. Finally, it covers the three major contexts of test use: educational, occupational, and clinical and counseling psychology.
📝 Lecture Summary
Assumptions of Psychological Testing and Assessment
Cohen and Swerdlik (1999) enumerate twelve assumptions that form the foundation of psychological testing. First, psychological traits and states exist — a trait is defined by Guilford (1959, p. 6) as "any distinguishable relatively enduring way in which one individual varies from another." Second, these traits and states can be quantified and measured. Quantification refers to gathering and presenting findings in numbers, making the process objective and scores comparable. Measurement is defined as "the act of assigning numbers or symbols to characteristics of subjects according to rules" (Cohen & Swerdlik, 1999, p. 19). A scale is "a set of numbers whose properties model empirical properties of the objects or traits to which numbers are assigned."
Third, various approaches to measuring the same thing can be useful — no single tool exists for measuring any one aspect of behavior. Fourth, assessment can provide answers to momentous life questions, leading to continuous test design and refinement. Fifth, assessment can pinpoint phenomena requiring further attention or study, meaning tests serve diagnostic purposes for therapy, forensic reasons, placement, or educational counseling. Sixth, various sources of data enrich and are part of the assessment process. Seventh, various sources of error are part of the assessment process — many variables may intervene even under controlled conditions. Eighth, tests and measurement techniques have strengths and weaknesses. Ninth, test-related behaviors predict non-test-related behaviors — for example, in the HTP (House, Tree, Person) test, drawing ability is not being tested; rather, the contents present cues to the subject's personality. Tenth, present-day behavior sampling predicts future behavior — test results are assumed to hold true beyond the day of administration. Eleventh, testing can be conducted in a fair and unbiased manner. Twelfth, testing and assessment benefit society.
💡 Why this matters: These assumptions justify the entire enterprise of psychological testing — without them, we could not claim that a test score meaningfully represents a person's intelligence, personality, or aptitude.
Types of Tests
1- Individual Versus Group Tests: Individual tests require one examiner and one subject at a time, such as WAIS, WISC, Stanford-Binet Scales, and Kaufman Scales. Group tests involve one examiner working with many subjects simultaneously, such as Raven's Progressive Matrices (RPM) and Otis Self-Administering Tests of Mental Ability. RPM is available in three forms differing in difficulty level for both individual and group administration. Group tests are quick, easy to administer and score, usually based on multiple choice items, and can even use tape-recorded or computer administration. However, they are unsuitable where subject-examiner rapport is important.
2- Ability, Achievement, or Aptitude Tests: Ability tests measure intellectual ability or cognitive behavior and yield an IQ score. These tests cover a sample of what the person knows at the time of testing and indicate the level of development attained in those abilities. Achievement tests are designed to measure the effect of educational programs or trainings — they are given at the end of a program to determine if specific objectives have been achieved. The SAT is an example. Aptitude tests measure the cumulative influence of multiplicity of experiences in daily living (Anastasi & Urbina, 1997), predicting future performance based on learning under uncontrolled, general conditions. For example, a test measuring how many mathematics problems a child can solve now is an achievement test; a test measuring how well they could solve problems if provided training is an aptitude test. The Differential Aptitude Tests (DAT) are the most widely used aptitude test batteries.
3- Intelligence Versus Personality Tests: These two types are distinct. Intelligence tests yield IQ information, while personality tests assess traits like dispositions. Neither can substitute for the other, and both are available in large numbers.
4- Speed Tests versus Ability Tests: Ability tests measure the level or amount of ability (e.g., IQ), and time for completion is not of utmost importance — some have no time limit, like Raven's Progressive Matrices. Speed tests emphasize speed of performance, such as the Clerical Speed and Accuracy Test in DAT. A related variety is power tests, which have a time limit long enough to allow everyone to complete all items.
5- Structured Personality Tests versus Projective Tests: Structured personality tests (also called objective tests) provide fixed response options — alternate response (true/false), three, four, or more options (MCQs). Scoring is simple with an answer key. These are usually "self-report" type tests, such as the EPPS (Edwards Personal Preference Schedule). Projective tests present a vague or ambiguous stimulus that subjects describe or explain. Examples include Rorschach's Inkblots, TAT (Thematic Apperception Test) where subjects narrate a story about a picture, HTP (House, Tree, Person) where drawings are analyzed, RISB (Rotter's Incomplete Sentence Blank) where subjects complete sentences, and WAT (Word Association Test) where subjects give prompt answers to stimulus words.
🔑 Definition — Structured personality test: A test with fixed response options where examinees choose the option that best describes them.
📌 Example (Structured): "I love animals [ ] Yes [ ] No" or "Animals are one's best friends [ ] Yes [ ] No"
6- Verbal versus Nonverbal/Performance Tests: Most tests are verbal, relying on language use — this is true of IQ tests and all structured personality tests. Nonverbal/performance tests do not involve the use or measurement of verbal ability, such as Raven's Progressive Matrices (RPM).
7- Commercial-Copyrighted Tests Versus "Available To All" or Online Tests: Commercial copyrighted tests like WAIS and WISC cannot be reproduced or photocopied — disclosing items would make them meaningless if examinees are familiar with contents. They must be purchased from the author or authorized agency. "Available to all" tests are found online or in textbooks, primarily for research rather than diagnostic or screening purposes. For example, the Multidimensional Health Locus of Control Scale (Wallston & Wallston) and Self-Efficacy Scale (Schwarzer & Jerusalem) can be downloaded and used freely, yielding valuable empirical evidence.
Major Contexts of Current Test Use
i) Educational Testing: Testing is used in schools at all levels, with the type depending on who uses it (counselor, psychologist, or teacher). Achievement, intelligence, special and multiple aptitude, and personality tests are used. Forms include: a) General Achievement Batteries that generate profiles of scores in major academic areas, such as the Stanford Achievement Test Series with Otis-Lennon School Ability Test, IOWA Test Series, California Achievement Test. b) Tests of Minimum Competency in Basic Skills that measure mastery of basic skills, such as the TABE (Tests of Adult Basic Education) battery covering five difficulty levels across reading, language, and applied mathematics. Results come as competency-based information and norm-referenced scores. c) Teacher-made classroom tests (objective or subjective). d) College level tests like the SAT (Scholastic Assessment Tests) Program including SAT I (Reasoning Test) and SAT II (Subject Tests). e) Graduate School Admission tests like the GRE (Graduate Record Examinations), used for admission to graduate and professional schools, honors, and scholarships, with a general test and subject tests.
ii) Occupational Testing: Tests are used for screening, induction, performance assessment, job analysis, and job performance prediction. Examples include: Academic Intelligence Tests (e.g., Wonderlic Personnel Test), Aptitude Tests (e.g., GATB — General Aptitude Test Battery; ASVAB — Armed Services Vocational Aptitude Battery), and Personality Tests (e.g., Five Factor Model Personality Inventories, MMPI for psychopathology, CPI, and HPI — Hogan Personality Inventory).
iii) Tests Used in Clinical and Counseling Psychology: Tests are used for diagnosis, treatment induction, general assessment, and gauging recovery rate. All intelligence and personality tests may be used — for example, HTP can depict psychopathology. Neuropsychological assessment tests include Bender Visual Motor Gestalt Test (Bender Gestalt Test/BGT) and Benton Visual Retention Test (BVRT). Specific learning disability tests include the Kaufman Test of Educational Achievement (K-TEA).
⭐ Key Takeaways
The twelve assumptions of psychological testing provide the philosophical and scientific justification that traits exist, can be measured, and predict future behavior — without these, testing would lack validity. Tests are classified along multiple dimensions: individual vs. group administration, what they measure (ability, achievement, aptitude), speed vs. power, structured vs. projective format, and verbal vs. nonverbal content. Projective tests use ambiguous stimuli (inkblots, pictures, sentence completions) to reveal personality, whereas structured tests use fixed response options like true/false. The three major real-world contexts — educational, occupational, and clinical — each require different test types for different purposes, from school placement and job screening to diagnosing psychopathology.
🧠 Quick Revision Questions
- According to Guilford (1959), what is a "trait" as distinct from a "state" in psychological testing?
- What is the fundamental difference between a speed test and an ability (power) test?
- Why are projective tests (like Rorschach's Inkblots or TAT) considered more difficult to score and interpret than structured personality tests?
- Name two commercially copyrighted IQ tests and explain why their items cannot be publicly disclosed.
- In the context of educational testing, what is the difference between a general achievement battery and a test of minimum competency in basic skills?
📘 Lecture 05 — The Testing Process: Test Administration and Test Taking
📖 Overview: This lecture examines the entire testing process, focusing on factors that can introduce error variance and reduce test validity. It covers the critical roles of examiner preparation, testing conditions, rapport building, and examinee variables such as test anxiety, coaching, and test sophistication, providing a comprehensive framework for understanding influences on test performance.
🗂️ Topics Covered
The lecture covers the test administration process including advance preparation of examiners and testing conditions. It explores introducing the test through rapport and test-taker orientation, alongside examiner and situational variables. Characteristics of a good examiner are outlined, followed by an in-depth look at examinee variables, focusing on test anxiety including its nature, measurement, and treatment. The lecture concludes by examining the effects of training on test performance, covering coaching, test sophistication, and instructions in broad cognitive skills.
📝 Lecture Summary
Test Administration Process
The main purpose of testing is to generalize results from the test sample to non-test situations. Any influences specific to the test situation constitute error variance and reduce test validity. It is crucial to recognize test-related influences, examiner-related variables, and subject/examinee-related variables as potential confounding variables that need to be controlled.
Advance Preparation Of Examiners
Advance preparation is required for the uniformity of the testing procedure. This preparation takes many forms: verbal instructions must be memorized to prevent misreading and hesitation. The preparation of test materials is another important step to avoid mishandling. Materials should be in easy reach, the testing procedure should be clear beforehand, and training for test administration is essential.
Testing Conditions
The testing environment must be appropriate, including a noise-free room, proper seating, and adequate lighting. It is important to recognize conditions that might affect test scores, such as noise, privacy, and traffic in the testing room. Even the type of answer sheet may affect scores. Subtle conditions affect performance; whether the examiner is a stranger or someone familiar can make a significant difference. The general manner and behavior of the examiner, such as smiling and nodding, can also have a decided effect on test results.
Introducing the Test: Rapport and Test-Taker Orientation
In test administration, rapport refers to the examiner’s efforts to arouse the test taker’s interest, elicit their cooperation, and encourage them to respond appropriately. Techniques for establishing rapport vary with the nature of the test and the age and characteristics of the person tested. Special motivational problems may be encountered with emotionally disturbed persons, prisoners, or juvenile delinquents, who may show suspicion, insecurity, or fear. The experienced examiner must make special efforts to establish rapport under these conditions.
Examiner and Situational Variables
Situational variables have been found to affect test-taking behavior and performance, especially with unstructured and ambiguous stimuli, as well as difficult and novel tasks. Children are more susceptible to examiner and situational influences than adults. Test results may be influenced by the examiner’s behavior, such as a "warm" versus a "cold" interpersonal relationship. The test taker's activities immediately preceding the test may also affect performance, as can feedback regarding test scores; "success" feedback leads to significantly higher performance than "failure" feedback.
🔑 Definition — Rapport: The examiner’s efforts to arouse the test taker’s interest, elicit their cooperation, and encourage them to respond in a manner appropriate to the objectives of the test.
Characteristics of a Good Examiner
A good examiner possesses: understanding of the nature of the test, proper training and experience, knowledge of instructions, clear speech, empathy, sharp senses (hearing and vision), quickness in understanding and responding, and professional honesty and integrity.
Test Anxiety
Test anxiety is a reaction of the test taker in a testing situation, stimulated by the negative effect on test performance. Practices to reduce it include building rapport and reducing the strangeness of the testing situation. The examiner's smooth and well-organized handling of the process also helps reduce test anxiety.
Individual Differences in Test Anxiety
Research shows individual differences in test anxiety responses. Questionnaires were developed to assess these differences, with findings indicating negative correlations between test anxiety and intelligence/achievement test scores. In one study, low-anxious children significantly improved on repeated learning tasks compared to high-anxious children with equal intelligence scores.
Anxious and Relaxed States in Test Performance
Studies comparing "anxious" and "relaxed" states show that ego-involving instructions have a positive effect on low-anxious students but a negative effect on high-anxious students. Chronically high anxiety exerts a detrimental effect on school learning and intellectual development. Research suggests students high on test anxiety obtain lower GPAs and have poorer study habits.
Research on Nature, Measurement, and Treatment Of Test Anxiety
The nature of anxiety is believed to include two components: emotionality and worry. Emotionality includes feelings and physiological reactions (increased heart rate, tension), while worry is the cognitive component involving negative, self-oriented thoughts about failure. These cognitions disrupt test performance by drawing attention away from task-oriented behavior. 💡 Why this matters: Understanding that test anxiety has a cognitive "worry" component, not just a physiological one, is key to developing effective treatments, such as behavior therapy combined with cognitive therapies. 🔑 Definition — Emotionality: The component of anxiety that includes feelings and physiological reactions like increased heart rate and tension. 🔑 Definition — Worry: The cognitive component of anxiety that includes negative self-oriented thoughts about failure and its consequences, which disrupts test performance.
Coaching
Individual with deficient educational backgrounds are more likely to benefit from coaching than those with superior opportunities. How much someone benefits depends on: the ability of the test taker, earlier educational experiences, the nature of the test, and the amount and type of coaching. The more the test material and coaching materials have in common, the greater the improvement.
Test Sophistication
Research shows that test sophistication or test-taking practice has a positive effect on test performance. When alternate forms of the same test were used, the second score tended to be higher. Short orientation and practice sessions can be effective in equalizing test sophistication, as illustrated by booklets like Taking the SAT I: Reasoning Test and GRE familiarization materials. Test familiarization is not limited to print media but includes transparencies, slides, films, videos, and computer software.
Instructions in Broad Cognitive Skills
Some researchers advocate providing education rather than coaching to improve test performance. These programs, often working with educable mentally retarded individuals, are designed to develop effective problem-solving behavior such as careful analysis of problems and consideration of all alternatives. These programs are still in an exploratory stage.
⭐ Key Takeaways
The testing process is vulnerable to error variance from examiner, situational, and examinee variables, all of which reduce test validity. Rapport building and proper orientation are critical to manage test anxiety and maximize the generalizability of results. Test anxiety has a cognitive "worry" component that disrupts performance and can be treated with cognitive-behavioral therapy. Training effects, including coaching and test sophistication, can improve scores, but coaching benefits depend on the test-taker's background and the nature of the test. Finally, a good examiner must be well-trained, empathetic, and aware of how their own behavior and expectations can influence test outcomes.
🧠 Quick Revision Questions
- What are the three main categories of variables that can introduce error variance into the testing process?
- How does the "worry" component of anxiety differ from the "emotionality" component, and how does it affect test performance?
- What is rapport, and why is it particularly challenging to establish with certain populations like juvenile delinquents?
- According to the lecture, what factors determine how much an individual will benefit from coaching?
- What is the main difference between "coaching" and "instructions in broad cognitive skills" as approaches to improving test performance?
📘 Lecture 06 — Test Norms: Interpreting Test Results
📖 Overview: This lecture explains how raw test scores are given meaning through the use of norms, which serve as standards for comparison. It covers the foundational statistical concepts necessary for understanding test score interpretation, including frequency distributions, measures of central tendency, and variability. Understanding norms is critical for evaluating an individual's performance relative to a standardized group.
🗂️ Topics Covered
The lecture begins by defining test norms and explaining their role in interpreting raw scores, introducing the concepts of normative samples and relative standing. It then transitions into essential statistical concepts used in psychological testing, covering frequency distributions, histograms, frequency polygons, the normal bell-shaped curve, and measures of central tendency including mean, median, and mode. Finally, it discusses measures of variability, specifically variance and standard deviation.
📝 Lecture Summary
Test Norms: Interpreting Test Results
Norms are standards used for comparative and normative functions. Test norms are scores on a measure used as standards against which the scores of any test taker are compared. Raw scores on psychological tests have no meaning unless interpreted with additional supporting information. For example, a score of 50 on an intelligence test does not indicate whether a person has an average, below, or above I.Q. level without knowing the average score or norm. Scores become meaningful in the presence of norms developed for that test.
🔑 Definition — Norms: "The test performance data of a particular group of test takers that are designed for use as a reference for evaluating or interpreting individual test scores". Norms provide standards for comparing the results of test takers on different measurements.
Norms are established by administering a test to a normative sample representative of the population of interest. For example, a stress scale for university students is administered to a normative sample of university students. If the average score is 25, this serves as a norm. Another student scoring 30 would be considered above average.
Standardization sample is a group of people whose performance on a test is taken as a standard for comparison. A person's performance is interpreted with reference to the distribution of scores in this sample to discover their relative standing—the position where a person lies in the distribution of scores in relation to other persons.
Raw scores are converted into derived scores for two reasons:
- To learn about an individual's standing in a distribution of scores relative to previously established norms.
- To make a direct comparison of individual scores with standard comparable measures.
These derived scores are expressed in two forms:
- Developmental level
- Relative position within a specific group
💡 Why this matters: Without norms, a raw score is just a number. Norms transform it into meaningful information about a person's performance compared to others.
Statistical Concepts Used In Psychological Testing
The aim of statistical method is to organize and summarize quantitative data. The first step is to tabulate raw scores to give them meaningful shape.
1- Frequency Distribution: A frequency distribution arranges large data in an organized way. Data is grouped into class intervals, and scores are tallied in the appropriate level. The total number of cases is represented by N.
📐 Frequency Distribution Table: For example, raw scores of 100 individuals on an achievement test ranging from 50 to 104 are grouped into class intervals of five points.
| Class Interval | Frequency |
|---|---|
| 50-54 | 2 |
| 55-59 | 5 |
| 60-64 | 7 |
| 65-69 | 10 |
| 70-74 | 14 |
| 75-79 | 13 |
| 80-84 | 20 |
| 85-89 | 11 |
| 90-94 | 4 |
| 95-99 | 6 |
| 100-104 | 8 |
| N=100 |
Frequency distributions can be presented in graphs:
- Histogram: A graph showing distribution of measured scores in the form of class intervals. The horizontal axis (x-axis) presents class interval scores, and the vertical axis (y-axis) presents frequencies. The height of the column presents the number of frequency, and the width covers the length of intervals.
- Frequency Polygon: Scores are presented by taking a mid-point in each class interval. These points are connected by a straight line.
2- Normal Bell-Shaped Curve: Statistical data can be seen in the form of a normal distribution curve. This curve facilitates basic statistical analysis. The curved area covers the average of some value or characteristic (e.g., I.Q., height, weight), and the two sides indicate extreme cases. As the number of people increases, distribution scores are more likely to resemble the original normal curve.
3- Measures of Central Tendency: Measures of central tendency describe raw scores in meaningful ways. They include:
- Mean
- Median
- Mode
- Variability and Standard Deviation
Mean: The mean or average is calculated by adding all scores and dividing by the number of cases (N).
📐 Formula: M = ∑X / N (Where ∑X = sum of all scores, N = number of cases)
📌 Example: Children obtained marks 7, 8, 4, 6, 5, 9, 3 on an English test. Sum of scores (∑X) = 42 Number of cases (N) = 7 Mean (M) = 42 / 7 = 6
Median: The median is the middle-most value or score in a group of data. It is the point that divides the distribution into half above and half below scores.
📐 Formula: Median = (n+1)/2 (where n = number of scores)
📌 Example: For scores 7, 8, 4, 6, 5, 9, 3, first arrange them in order: 3, 4, 5, 6, 7, 8, 9. Median position = (7+1)/2 = 4 The 4th value is 6.
Mode: The mode is the highest frequency value in a set of scores. In a normal distribution curve, it is represented by the highest point.
📌 Example: In scores 7, 8, 7, 7, 9, 4, 7, the mode is 7 because it occurs most often.
Variability and Standard Deviation: Variability is the dispersion of scores around the mean.
Variance: Variance is the "average squared deviation" around the mean.
📐 Formula: Variance = ∑(X - X̄)² / n (Where X = individual score, X̄ = mean, n = number of scores)
To calculate, first subtract the mean from each individual score to find the deviation, then square each deviation value, and finally calculate the mean of these squared deviation values. The sum of deviations around the mean will always equal zero because positive and negative values cancel each other out.
Standard Deviation: Standard deviation (S.D. or σ) is a more adequate measure of variability. It is calculated by taking the square root of the variance.
📐 Formula: Standard Deviation (S.D.) = √Variance
Larger individual differences reveal a larger S.D., while small individual differences indicate low S.D. values. In a normal distribution curve, interpretations of S.D. are very clear.
⭐ Key Takeaways
Norms are essential for interpreting raw test scores, transforming meaningless numbers into meaningful comparisons against a standardization sample. The lecture establishes that understanding basic statistical concepts—including frequency distributions, histograms, frequency polygons, and the normal curve—is foundational for developing and using norms. The three key measures of central tendency (mean, median, and mode) each describe a distribution's center in a different way, with the mean being the most commonly used but sensitive to extreme scores. Variability, measured by variance and standard deviation, describes how spread out scores are around the mean, with a larger standard deviation indicating greater individual differences. For the exam, you must be able to define norms, explain their purpose, and calculate or interpret the mean, median, mode, variance, and standard deviation.
🧠 Quick Revision Questions
- What are test norms, and why are raw scores meaningless without them?
- What is a standardization sample, and what is the purpose of "relative standing"?
- How are a histogram and a frequency polygon different in their graphical representation of data?
- Calculate the mean, median, and mode for the following set of test scores: 10, 12, 12, 15, 18, 20, 20.
- Explain the relationship between variance and standard deviation. If a set of scores has a variance of 36, what is the standard deviation?
📘 Lecture 07 — Types of Norms
📖 Overview: This lecture explains how raw test scores are converted into meaningful derived scores to interpret an individual’s performance. It covers developmental norms (like mental age and grade equivalents) and within-group norms (like percentiles and standard scores), emphasizing their uses, calculations, and limitations in psychological testing.
🗂️ Topics Covered
This lecture covers types of norms including developmental norms (mental age, basal age, grade equivalents, ordinal scales), within-group norms (percentiles, standard scores, normal standard scores, deviation IQ), relativity of norms (interest and longitudinal comparisons), and scales of measurement. Key concepts include how derived scores allow comparison across tests and populations, and the psychometric limitations of each norm type.
📝 Lecture Summary
Types of Norms
Raw test scores need to be converted into relative measures or derived scores to give them meaningful interpretation. These scores reveal an individual’s relative standing and allow comparison across different tests. Derived scores are expressed either in terms of developmental level attained or the relative position of a person within a specific group.
Developmental Norms
Developmental norms are defined as the typical patterns, characteristics, and age-specific tasks or skills of development at any age or stage. They are established based on development and maturation, assuming people can perform at specific levels at different life stages. When most people can perform certain tasks at a given age, this becomes the norm for that age level and is considered the mental age (MA) for that level. For example, if a 13-year-old girl performs tasks that most 13-year-olds can do, her MA is 13. If an adult only performs tasks a 6-year-old can do, his MA is 6, indicating mental deficiency. If a 9-year-old performs tasks meant for a 16-year-old, his MA is higher than his biological age.
However, comparing performance using developmental norms is not always straightforward because people may take tests measuring different abilities, and even subtests of the same test may measure different skills. It is not necessary that everyone attains the same MA across all tests or subtests, making comparison difficult. Although used commonly for descriptive purposes in clinical and research settings, test scores based on developmental norms are not psychometrically sound.
Mental Age
The term mental age became widely used after the development of the Binet-Simon scales, though Binet used the more neutral term “mental level.” Items were grouped in year levels (e.g., items passed by most 8-year-olds were placed in the 8-year level). With frequent use, the problem of scatter was observed — many subjects did not show uniform performance across all subtests, failing items below their age level while passing items above their age level. To address this, the concept of basal age was introduced.
Basal age refers to the highest year at which a person passes all items. For tests passed at higher year levels, the subject is given partial credits in months, which are added to the basal age to yield the child’s mental age. Mental age norms are also used with tests not designed according to year levels. In these tests, the mean raw scores of children of specific age groups in the standardization sample serve as norms. A child’s mental age is determined by comparing her raw score with the age norm. For example, if a 12-year-old girl’s raw score equals the 12-year norm, her mental age is 12.
🔑 Definition — Mental Age (MA): The age level at which a person performs intellectually, based on the average performance of individuals at that chronological age.
A major shortcoming of using mental age is that it does not mean the same thing at different life stages. MA of 4 for a 5-year-old is not equivalent to MA of 24 for a 25-year-old. As age progresses, the unit of MA tends to shrink. A child with MA of 3 at age 4 would be 3 years retarded at age 12, but one year of mental growth from age 3 to 4 is equivalent to three years of growth from age 9 to 12. Therefore, deviation from the norm at different ages is not comparable — deviation at a very young age is much more significant than at older age.
Grade Equivalents
Grade equivalents represent scores on educational achievement tests attained by children in a certain grade. These norms are obtained by calculating the mean raw scores of children in the standardization sample for each grade. If 6th-grade children obtained a mean score of 35 on an arithmetic test, then this raw score has a grade equivalent of 6. A student scoring 35 on the same test is said to have a grade equivalent of 6.
Since most academic years span ten months, successive months can be expressed in decimal points. For example, a grade equivalent of 7.0 refers to average performance of a 7th grader at the beginning of the session, while 7.5 represents average at the middle of the session.
📐 Formula — Grade Equivalent: Mean raw score of children in a specific grade = Grade Equivalent for that raw score.
Grade norms have several limitations:
- Grade units are unequal and these inequalities occur irregularly in different subject matter areas.
- They are only applicable for common subjects taught throughout the grade levels covered by the tests, not for subjects taught for only one to two years.
- Even when the same courses are covered, identical importance and learning cannot be ensured across grades.
- A child may progress more rapidly in one subject than another during the same grade.
- Grade norms tend to be incorrectly regarded as the performance level of all students, ignoring that individual differences in any grade can be so large that scores vary over several grades.
Ordinal Scales
Another approach to developmental norms comes from research in child psychology. Psychologists observed behaviors typical of successive ages in infants and young children, including functions like sensory discrimination, linguistic communication, and concept formation. These observations proved valuable in understanding human development.
An example is the work of Gesell and his associates, who focused on the sequential patterning of early behavior development. The Gesell Developmental Schedules were developed to determine a child’s approximate developmental level in months across four major areas: motor, adaptive, language, and personal-social behavior. Eight key ages from 4 to 36 weeks serve as standards. Gesell claimed children’s development involves: a) orderly progression of behavior changes and b) uniformities of developmental sequence. For example, chronological sequence can be observed in visual fixation and hand/finger movements when reacting to a small object. The scales used are ordinal scales that yield information about the stage where a child stands, with successful performance at one level implying success at all lower levels.
Jean Piaget (1960s) gave his theory of cognitive development, describing stages of cognitive development in a sequence, with age levels being arbitrary. He studied specific concepts including:
- Object permanence: The child is aware of object existence when they are out of sight.
- Conservation: Recognition that an attribute remains constant over changes in perceptual appearance (e.g., liquid quantity remains constant when poured in different shaped containers).
- Perspective: Knowledge that objects appear differently when at a distance.
To assess these, Piagetian tasks are used, designed to reveal the dominant aspect of each developmental stage.
In short, ordinal scales gauge the uniform progression of development through successive stages by measuring attainment of specific functions.
Within-Group Norms
Within-group norms evaluate a person’s performance by comparing their raw score with the most nearly comparable standardization group, such as children of the same age or grade. These norms are so popular that all test scores now provide within-group norms in some form. Within-group scores employ many statistical procedures due to their clearly defined quantitative meaning.
Percentiles
A percentile indicates an individual’s relative position in the standardization sample. The percentage of people in the standardization sample is expressed in terms of percentile scores. For example, if 50% of people obtained a score of 25 on an analytical reasoning test, then this score corresponds to the 50th percentile.
Key points about percentiles:
- They can describe ranks in a group of 100 people — the top person is given rank 1, the bottom person a poorer rank.
- The 50th percentile refers to the median; scores above 50 represent above-average scores, and scores below 50 indicate below-average scores.
- The 25th and 75th percentiles are known as the first and third quartile points (Q1 and Q3), cutting off the lowest and highest quarters of the distribution.
- The difference between percentage and percentile is that percentage is a raw score while percentiles are derived scores.
Standard Scores
Standard scores express an individual’s distance from the mean in terms of the standard deviation of the distribution. Standard scores can be calculated by linear and non-linear transformation of raw scores. Linearly derived scores are known as z-scores.
🔑 Definition — z-score: A standard score obtained by subtracting the mean of the normative sample from the raw score and dividing by the standard deviation of that sample.
📐 Formula — z-score: ( z = \frac{X - M}{SD} )
Where:
- ( X ) = raw score
- ( M ) = mean of normative sample
- ( SD ) = standard deviation of normative sample
📌 Example: If X = 100, M = 80, and SD = 10: ( z = \frac{100 - 80}{10} = \frac{20}{10} = 2.0 )
A raw score equal to the mean results in a z-score of zero. Negative derived scores indicate performance below average; positive scores indicate above-average performance.
Normal Standard Scores
Normal standard scores are standard scores expressed in terms of a distribution transformed to fit a normal curve. They are obtained by finding the percentage of a person in the standardization sample, locating this percentage in the normal curve frequency table, and obtaining the standard score. Normal standard scores can be put in any convenient form.
If normalized standard scores are multiplied by 10 and added to or subtracted from 50, they become T scores (first proposed by McCall, 1922). In a T score scale, a score of 50 corresponds to the mean, 60 corresponds to 1 SD above the mean, and so on.
📐 Formula — T score: ( T = 50 + 10(z) )
Normalized standard scores should be applied when the sample is large and representative and when deviation from normal results is confirmed to be due to drawbacks in the test rather than characteristics of the sample.
Another variation is the Stanine scale (from “standard nine”), developed by the United States Air Force during World War II. Scores run from one to nine, with a mean of 5 and standard deviation of approximately 2.
🔑 Definition — Stanine: A single-digit standardized score system with values from 1 to 9, mean of 5, and SD of approximately 2.
The Deviation IQ
The term IQ (Intelligence Quotient) was introduced in early intelligence tests. The Ratio IQ is obtained by dividing mental age (MA) by chronological age (CA) and multiplying by 100.
📐 Formula — Ratio IQ: ( IQ = \frac{MA}{CA} \times 100 )
If a child’s MA equals CA, IQ = 100 (average performance). IQ below 100 indicates below-average scores; above 100 indicates acceleration.
However, the Ratio IQ has major technical problems — it is not comparable across different age levels unless the SD of the IQ distribution remains constant with age. For example, if a child reads at age 3 (CA) when an average child reads at age 6 (MA), their IQ = 3/6 × 100 = 200. For this reason, ratio IQ was replaced by Deviation IQ.
Deviation IQ is a standard score with a mean of 100 and an SD that approximates the SD of the Stanford-Binet IQ distribution. It compares people of the same age and assumes that IQ is normally distributed.
🔑 Definition — Deviation IQ: A standard score with mean = 100 and SD appropriate to the test, used to compare an individual’s performance to others of the same chronological age.
Relativity of Norms
Interest Comparisons: An IQ score should always be described by the name of the test on which it was obtained. The relative standing of IQs can change when different tests are used. An individual’s relative standing in different functions may be misrepresented by the lack of comparability of test norms. For example, if a verbal test was standardized on a random high school sample but a spatial test was standardized on a selected group, an examiner might erroneously conclude the individual is more verbally able than spatially able.
For longitudinal comparisons (scores on a specific test over time), differences may be due to three reasons:
- Content differences: Intelligence tests with the same label may have different content (e.g., one is verbal, another numerical).
- Scale units incomparability: IQ on one test may have SD of 12 while another has SD of 18.
- Standardization sample composition variation: The same individual will appear to perform better when compared with a less able group than with a more able group.
Scales of Measurement
To describe test scores quantitatively, tests must be designed to yield numeric results — either originally numeric or convertible to that form. Psychological measurement involves rules according to which objects are assigned numbers, and “quality” is expressed in numeric form.
For example, in a personality test, an item asks “do you like to be in the company of young age mates most of the time?” with answer options: “always”, “often”, “could not say”, “rarely”, and “never”. Since the response is qualitative, numbers are assigned (e.g., “never” = 1, “always” = 5). All responses can then be quantified and subjected to statistical treatment. In ability tests where every question has a right answer, the total number of correct responses yields the test score.
💡 Why this matters: Understanding scales of measurement is fundamental because the type of scale (nominal, ordinal, interval, ratio) determines what statistical analyses are appropriate and what interpretations can be made from test scores.
⭐ Key Takeaways
The most critical concepts from this lecture are that raw test scores are meaningless until converted into derived scores using norms, and there are two main categories: developmental norms (mental age, grade equivalents, ordinal scales) and within-group norms (percentiles, standard scores, T scores, Stanines, deviation IQ). Mental age has psychometric limitations because its units shrink with age, making ratio IQ problematic and leading to the adoption of deviation IQ (mean = 100) as a more psychometrically sound alternative. Percentiles provide easily interpretable relative position but are not equal-interval scales, while standard scores (z-scores, T scores, Stanines) allow precise comparison by expressing distance from the mean in standard deviation units. Finally, the relativity of norms means that test scores must always be interpreted in context — considering the specific test used, its standardization sample, and the scale properties — to avoid misinterpretation in both cross-sectional and longitudinal comparisons.
🧠 Quick Revision Questions
- What is the difference between developmental norms and within-group norms, and give one example of each?
- Why is ratio IQ considered technically problematic, and how does deviation IQ address this issue?
- How is a z-score calculated, and what does a z-score of 0, +2, and -1.5 indicate about a person’s performance?
- What are the major limitations of grade equivalents as a type of developmental norm?
- Explain the difference between percentage and percentile, and identify what the 25th, 50th, and 75th percentiles represent in a distribution.
📘 Lecture 8 — Test Norms and Related Concepts
📖 Overview: This lecture explores the concept of standardization in psychological testing through the establishment of test norms. It covers various types of norms—national, national anchor, specific, and fixed reference group norms—and introduces Item Response Theory as an advanced scaling method. Understanding these concepts is critical for correctly interpreting test scores and making valid judgments about test takers.
🗂️ Topics Covered
The lecture covers the process of standardization and its importance, the characteristics and selection of a normative sample including stratified random sampling, national norms based on representative national samples, national anchor norms using the equipercentile method for score equivalence, specific norms for narrowly defined populations, fixed reference group scoring systems using the SAT as a historical example, and finally Item Response Theory as a sample-free measurement approach.
📝 Lecture Summary
Standardization: Test Norms and Related Concepts
Norms are established for the sake of standardization of any test. Standardization is the process whereby a test is administered to a representative sample of population whom the test is meant for, for the sake of establishing norms. A standardized test is one that has normative data, as well as clearly specified administration and scoring procedures.
The Normative Sample
When using test scores, we must ensure that the norms being referred to are representative of the particular population from which the standardization or normative sample was selected. The mean scores of that sample are assumed to represent the parent population. Therefore, if the sample comprised women alone, their mean performance score should be used as a norm for women’s raw scores alone. It is advisable that the standardization sample include maximum characteristics of the population and be of a large enough size to ensure stable values. As the size of sample increases, the chance of making error reduces.
To obtain a true representative sample, careful sampling procedures must be followed. Proportionate stratified random sampling is the best approach, where members from each subgroup/stratum are included in the sample in the same proportion in which they are found in population. A more practical option is to take a purposive sample that we believe contains all characteristics we are interested in catering for. For example, if the standardization sample included children aged 12 to 16 years with six years of schooling, then the norms are meant for a population of children within the same age range and similar educational background. No test provides norms for all sorts of population altogether.
💡 Why this matters: Using tests standardized in western countries with local populations requires great caution, as available norms were not established for such populations.
National Norms
If a test is standardized using a nationally representative sample of the population, it is called a national sample. The sample containing all characteristics of interest is chosen from different geographical regions, communities, socioeconomic status, and institutions. For example, to establish norms for a test measuring achievement of university students in Pakistan, the normative sample must represent university students from all regions of the country.
National Anchor Norms
Many tests measure the same ability or human trait, but a person may obtain different scores on different tests. National anchor norms provide a solution to this problem by creating an equivalency table for scores on various tests of the same ability. Equivalency is calculated using the equipercentile method. Test scores are considered equal only when they have equal percentiles in the group being studied. For example, if a score of 34 on test ABC carries the 85th percentile, and a score of 29 on test XYZ also carries the 85th percentile, then 34 on ABC is equivalent to 29 on XYZ.
🔑 Definition — Equipercentile Method: Test scores are compared and their equivalence is determined with reference to their corresponding percentile scores.
📌 Example: A score of 34 on test ABC carries the 85th percentile, and a score of 29 on test XYZ also carries the same percentile. Therefore, 34 on ABC is equivalent to 29 on XYZ.
💡 Why this matters: National anchor norms should not be used as a single fully dependable source of judgment; every member of the sample must have taken all the tests whose equivalence is being determined.
Specific Norms
An alternate solution to the problem of non-equivalence is to use specific norms, which require that tests be standardized on more narrowly defined populations. Rather than using broadly defined samples, tests can be standardized on narrowly defined samples selected through purposive sampling. In case of using specific samples, the chances of controlling nonequivalence are reduced. When norms for such tests are reported, a clear report of the limits of the normative sample must also be made.
Highly specific norms are useful for most testing purposes. Even when representative norms are available, availability of separately reported subgroup norms is very helpful. For example, while one inventory can measure occupational variables in doctors as one community, separate subscales may be developed for doctors working under different levels of stress and in different working conditions. At times, institutions develop their own local norms. For example, a university may develop norms on its students entering the first year to predict achievement in following years.
Fixed Reference Group
Sometimes non-normative scales are used. One such type is the fixed reference group scoring system. In this system, the distribution of scores obtained from one group of people is used as the basis for calculation of test scores for future administrations. The group from which scores were obtained is called the fixed reference group. This system does not provide normative evaluation of performance but ensures comparability and continuity of scores.
The College Board Scholastic Aptitude Test (SAT) is an example. In 1941, approximately 11,000 candidates took the test. The distribution of scores of this sample was taken as a standard, and all SAT scores were expressed in terms of mean and standard deviation of these candidates. A score of 500 corresponded to the mean of this group; 600 meant one SD above mean, and 400 was one SD below. In 1995, a new fixed reference group began to be used, comprising those who took the SAT in 1990. After April 1, 1995, scores are reported on the "recentered" scale derived from the 1990 reference group.
📐 Formula: SAT scoring system → A score of 500 corresponds to the mean of the fixed reference group; 600 means one SD above the mean; 400 means one SD below the mean.
Item Response Theory
Item Response Theory (IRT) can be understood in terms of latent trait models. Beginning in the 1970s, psychologists became interested in mathematically sophisticated procedures for scaling the difficulty of test items. The availability of high-speed computers made such procedures possible. The basic measure used is the probability that a test taker with a talent trait succeeds on an item of specified difficulty. The latent traits are mathematically derived statistical constructs derived from empirically observed relations among test responses.
The term latent trait model was later replaced by Item Response Theory (IRT) because 'latent trait' created a false impression of a specific trait. The purpose of IRT models is to establish a "sample-free" scale of measurement that is uniform, applicable to individuals and groups having widely varying ability levels, and to test contents varying widely in difficulty levels. Rather than using the mean and standard deviation of a specific reference group, IRT models set origin and unit size in terms of data representing a wide range of ability and item difficulty, obtained from many samples rather than a single sample.
🔑 Definition — Latent Trait: Mathematically derived statistical constructs derived from empirically observed relations among test responses; there is no implication regarding the existence of the trait as such.
⭐ Key Takeaways
The establishment of appropriate norms through a representative standardization sample is fundamental to valid test interpretation. A normative sample must reflect the population characteristics and be of sufficient size, with proportionate stratified random sampling being the ideal method for selection. National anchor norms provide score equivalence across different tests using the equipercentile method, while specific and local norms allow for more precise interpretation within narrowly defined groups. The fixed reference group system, exemplified by the SAT, ensures score comparability over time without requiring normative evaluation. Finally, Item Response Theory offers a sophisticated, sample-free approach to measurement that overcomes limitations of traditional norm-referenced methods by using data from multiple samples across a wide range of ability levels.
🧠 Quick Revision Questions
- What is the purpose of standardization in psychological testing, and what are the key features of a standardized test?
- Describe proportionate stratified random sampling and explain why it is considered the best approach for selecting a representative normative sample.
- How does the equipercentile method work in establishing national anchor norms for score equivalence?
- What is the difference between specific norms and local norms, and when would each be most appropriately used?
- Explain how the fixed reference group scoring system was applied to the SAT, including the significance of the 1941 and 1990 reference groups.
📘 Lecture 09 — Nature and Uses: Domain Referenced Test Interpretation
📖 Overview: This lecture introduces domain-referenced tests as an alternative to norm-referenced tests for interpreting psychological test results. Instead of comparing individuals to a normative group, domain-referenced tests evaluate what a person knows or can do within a specific content domain, making them essential for assessing mastery in education, professional training, and certification contexts.
🗂️ Topics Covered
This lecture covers the definition and purpose of domain-referenced (criterion-referenced) tests compared to norm-referenced tests, including who determines the domain or criterion. It explores content meaning in test interpretation, the process of developing instructional objectives and test items, mastery testing as an all-or-none approach, the relation between domain-referenced and norm-referenced testing, and the use of cutoff scores and expectancy tables for decision-making in minimum qualification contexts.
📝 Lecture Summary
Nature and Uses: Domain Referenced Test Interpretation
So far we have discussed norms for interpreting test results, but norms are not the only way to assess and measure abilities. Tests that use normative data are norm-referenced tests, defined as “a test that evaluates each individual relative to a normative group” (Kaplan & Saccuzzo, 2001). In such tests, a representative sample is tested, raw scores are analyzed, and norms are established. Every individual’s performance is compared to the scores of others — average, below average, or above average relative to the normative sample.
There is another type of test that uses a criterion or standards to describe performance. A person must perform within a domain to be considered proficient in a skill, behavior, or ability. These are called domain-referenced tests. The terms criterion-referenced and domain-referenced are used interchangeably, with the latter being more common. A criterion-referenced test is “a test that describes the specific types of skills, tasks, knowledge of an individual relative to a well-defined mastery criterion. The content of criterion-referenced test is limited to certain well-defined objectives” (Kaplan & Saccuzzo, 2001).
Glaser (1963) first used the term “criterion-referenced testing.” Alternative terms like “domain-referenced” and “content-referenced” were also proposed. The interpretive frame of reference in a domain-referenced test uses a specific content domain rather than a population of persons. Test results are reported in terms of what the test taker knows, how proficient they are, and to what extent they have command over a certain content domain. For example, according to Anastasi (2007), performance may be reported in terms of specific arithmetic operations mastered, estimated vocabulary size, or difficulty level of reading matter comprehended.
🔑 Definition — Norm-Referenced Test: A test that evaluates each individual relative to a normative group. 🔑 Definition — Criterion-Referenced Test: A test that describes specific types of skills, tasks, or knowledge of an individual relative to a well-defined mastery criterion.
Who Decides and Determines the Domain or Criterion?
The domain is primarily derived from the values or standards of an individual or organization. In many real-life and professional scenarios, evaluating a person’s ability relative to a norm is meaningless; what matters is command over the domain. For example, in training surgeons, how much above or below average their skills are is not important — what matters is whether they have acquired the surgical skill to a specified extent. Similarly, in evaluating a pilot, the key question is whether the pilot can fly a plane and how well, not whether they are equal, above, or below others.
Norm-referenced tests assess and describe how well test takers performed in relation to other people. Domain-referenced tests tell us what test takers can do — they focus on potential, while norm-referenced tests focus on comparative performance.
In some contexts, domain-referenced tests are called mastery tests. These are used to assess mastery or achievement of certain skills or contents, with the focus on content rather than a specific population. They became very popular in educational contexts in the 1970s. Performance may be reported in terms of mastery over skills (e.g., specific arithmetic operations, vocabulary size, reading difficulty level), or chances of achieving a designated performance level on an external criterion (educational or occupational).
Domain-referenced tests are commonly used in education, especially in innovations like computer-assisted, computer-managed, and individualized, self-placed instructional systems. According to Anastasi and Urbina (2007), testing is closely integrated in these systems, with instruction introduced before, during, and after each unit to check prerequisite skills, diagnose learning difficulties, and prescribe subsequent instruction. Broad surveys like the National Assessment of Educational Progress in the U.S. also use domain-referenced tests.
Another application area is assessing mastery of a small number of clearly defined job skills, such as in military occupational specialties, or evaluating attainment of minimum requirements (e.g., qualifying for a driver’s or pilot’s license). Familiarity with domain-referenced testing concepts can also help improve traditional teacher-made tests.
💡 Why this matters: Understanding who defines the domain is critical — in high-stakes contexts like surgery or aviation, domain-referenced assessment ensures safety and competence, not just relative ranking.
Content Meaning
The most significant feature of domain-referenced tests is that test performance is interpreted in terms of content meaning. The goal is not relative standing but learning what the person knows and can do. In developing such tests, the content must be carefully chosen, treated, and presented.
A clearly defined domain of knowledge or skills to be assessed is the primary requirement. The content must be important and generally accepted as such. To develop items for assessing mastery, the content is subdivided into smaller units, each defined in performance terms — what behavior will indicate that a content element has been mastered. In educational settings, these units correspond to behaviorally defined instructional objectives. For example: “divides numbers carrying zeros by numbers carrying one zero by canceling one zero at the end,” “can convert grams into ounces,” or “can convert centigrade into Fahrenheit.”
Instructional objectives specify learning outcomes and affect how a course is taught and how assessments are made. After objectives are finalized, the difficult task of item development for each objective follows. Each item must be a good representative of the domain. Careful formulation of objectives, clear statement of concepts and methodologies, and the test developer’s expertise and judgment all matter.
Domain-referenced tests are most useful for testing basic skills like reading and arithmetic at elementary levels. According to Anastasi and Urbina (2007), an ordinal hierarchy is usually adopted for arranging instructional objectives for basic skills — elementary skills must be acquired before higher-level skills. At elementary levels, content and learning sequence are mostly flexible. It is not advisable to formulate highly specific objectives for advanced levels of knowledge in less structured subjects.
🔑 Definition — Instructional Objectives: Behaviorally defined learning outcomes that specify what a student should be able to do, affecting teaching and assessment methods.
Mastery Testing
Mastery testing is an important characteristic of domain-referenced tests. It provides an all-or-none score, meaning the test score indicates the presence or absence of mastery, or the attainment of a pre-established level of mastery. A generally expected level is complete mastery, often set at 80% or 85%.
Another way of reporting mastery is to use a three-way distinction: mastery, non-mastery, and an intermediate, doubtful, or “review” interval.
In mastery tests, individual differences are not a matter of concern. Many educators believe that if suitable instructional methods are used and enough time is given, almost everyone can exhibit complete mastery and achieve instructional objectives. The only area where individual differences may appear is the time taken to learn the content. Therefore, the effect of individual differences can be reduced to a minimum after appropriate training.
Mastery testing is used in many individualized instruction programs. Published domain-referenced tests of basic skills for elementary schools often employ this approach.
Two key issues in constructing mastery tests are: (1) how many items should be used, and (2) what proportion of items must be correct before a reliable assessment is made. Initially, these issues were handled by test developers’ judgment, but now statistical techniques are available.
🔑 Definition — Mastery Testing: A testing procedure that yields an all-or-none score indicating the presence or absence of mastery at a pre-established level (typically 80-85%).
Relation to Norm-Referenced Testing
While mastery testing has advantages, its usefulness is greatest for basic skills at elementary levels. In areas where complete mastery is not the focus, mastery testing is not the best choice. As we go to higher levels with more advanced and less structured subjects (e.g., understanding, critical thinking, appreciation, originality), there is no way to assess complete attainment or mastery — there is no end, limit, or direction to learning. In such cases, norm-referenced tests are useful because there are no cutoff points to show complete absence of these faculties.
Some published tests allow both norm-referenced and domain-referenced applications. For example, Stanford diagnostic tests in reading and mathematics provide both appropriate norms and a system for qualitative analysis of a child’s attainment of detailed instructional objectives.
Some authors note that even when using domain-referenced tests, the concept of norms is still operative — there is an underlying realization that a continuum of abilities exists.
Minimum Qualification and Cutoff Scores
Mastery tests may adopt an all-or-none approach or make a three-way distinction, but there are situations where clear cutoff scores must be specified. These include granting a driving or flying license, selecting workers for a war zone (requiring sharp learning and vision), selecting workers for a nuclear plant, or choosing students for medical school.
Different tests use their own cutoff scores. The following points should be kept in mind when using cutoff scores for decision making (Anastasi & Urbina, 2007): a) Do not use scores from a single test as cutoff — use a band of scores from more than one test. b) Use multiple sources of information for decision making, including relevant performance on other tests from past or present. c) If a panel of judges sets cutoff points, they should be experts in test construction and task performance areas. d) Cutoff scores should be established and verified with the support of empirical information whenever possible. e) Test scores should be obtained from groups that are clearly different from each other on the relevant criterion behavior.
An empirical method for setting cutoff scores is using expectancy tables.
🔑 Definition — Expectancy Table: A table containing the probability of different criterion outcomes for persons who obtain each test score, based on statistical information from past administrations.
📐 Formula/Concept — Expectancy Table Example:
| Score on XYZ | Number of students | Percentage in each grade |
|---|---|---|
| 50-60 | 12 | A: - , B: 1, C: 8, D: 4 |
| 60-70 | 15 | A: 3, B: 6, C: 6, D: - |
| 70-80 | 23 | A: 12, B: 9, C: 2, D: - |
This table shows the relationship between test scores and course grades, helping predict criterion outcomes based on test performance.
⭐ Key Takeaways
Domain-referenced tests focus on content mastery rather than comparative performance, making them essential for assessing specific skills in education, professional training, and certification contexts. The key feature is content meaning — interpreting what a person knows or can do within a well-defined domain. Mastery testing uses an all-or-none approach (typically 80-85% mastery) and is most useful for basic skills at elementary levels. However, for higher-level abilities like critical thinking and creativity, norm-referenced tests remain more appropriate. When cutoff scores are needed for minimum qualification decisions, multiple sources of information and expectancy tables should be used for empirical validation.
🧠 Quick Revision Questions
- What is the fundamental difference between norm-referenced and domain-referenced tests in terms of what they assess?
- According to the lecture, who determines the domain or criterion in domain-referenced testing, and why is this important in professional contexts like surgery or aviation?
- What are instructional objectives, and why must they be defined in performance terms for domain-referenced tests?
- What is mastery testing, and what does an “all-or-none score” mean in this context (including the typical mastery level)?
- What are the recommended guidelines for setting cutoff scores, and how can expectancy tables be used to establish them empirically?
📘 Lecture 10 — Test Development: Test Construction
📖 Overview: This lecture details the systematic process of test development, from initial conceptualization through final revision. It explains the critical steps of constructing, trying out, analyzing, and revising test items to create a psychometrically sound instrument.
🗂️ Topics Covered
The lecture covers the five-step test development process: test conceptualization, test construction (including scaling methods and item writing), test tryout, item analysis, and test revision. It details different types of scales (age, grade, stanine), scaling methods (Likert, paired comparisons, sorting, equal-appearing intervals), item formats (selected response and constructed response), scoring models (cumulative, class, ipsative), and the characteristics of good test items.
📝 Lecture Summary
Test Development Process: Test Conceptualization
Test development begins with a test developer's idea to measure a particular construct. The stimulus for creating a test can come from various sources, such as literature on an already developed test that creates a need for further psychometric work, or the emergence of a social phenomenon that requires measurement.
Once the idea is formed, several questions must be answered: What is the test designed to measure? What is its purpose? Is there a need for it? What should the test content and format be? A critical question is "How will meaning be attributed to scores on the test?" which points to the issue of norm-referenced versus criterion-referenced tests.
- In norm-referenced tests, a good item is one that high scorers get right and low scorers get wrong.
- In criterion-referenced tests, development involves pilot work with two groups: one that has mastered the knowledge and one that has not. Items that best discriminate between these groups are considered "good."
🔑 Definition — Norm-referenced test: A test where an individual's score is interpreted by comparing it to the scores of a normative group. 🔑 Definition — Criterion-referenced test: A test where an individual's score is interpreted by comparing it to a predetermined standard or criterion of mastery. 💡 Why this matters: Understanding this distinction determines the entire approach to item selection and test interpretation.
Pilot Work refers to preliminary research surrounding the creation of a test prototype. Test items may be pilot studied to evaluate their suitability for the final instrument. This process involves creating, revising, and deleting many items. Once pilot work is done, test construction begins, though future pilot research may be needed for updates.
Test Construction
Scaling is defined as the process of setting rules for assigning numbers in measurement. In other words, scaling assigns values to different amounts of the attribute being measured.
Types of scales can be categorized by level of measurement (nominal, ordinal, interval, or ratio) or in other ways:
- Age scale: Performance is of critical interest as a function of age.
- Grade scale: Performance is of critical interest as a function of grade.
- Stanine scale: All raw scores are transformed into scores ranging from 1 to 9.
Scaling Methods include several approaches:
- Likert scale: Used to scale attitudes. Each item presents five alternative responses, usually on an agree/disagree continuum. Likert (1932) concluded that assigning weights of 1 through 5 generally works best.
- Method of paired comparisons: Test takers are presented with pairs of stimuli and must select the more appealing one. The test taker receives a higher score for selecting the option considered more justifiable by a group of judges.
- Sorting tasks: Printed cards, drawings, or other stimuli are presented for evaluation. Comparative scaling entails judgments of a stimulus in comparison with every other stimulus. Categorical scaling places stimuli into two or more alternative categories that differ quantitatively.
- Method of equal-appearing intervals: Described by Thurstone (1929), this method is used to obtain data that are presumed to be interval in nature.
All methods except equal-appearing intervals yield ordinal data.
Writing Items involves three important considerations:
- What range of content should the items cover?
- Which item formats should be employed?
- How many items should be written?
For a standardized test with multiple-choice format, the first draft should contain approximately twice the number of items that the final version will contain. This provides a basis for content validity, as the final version must still adequately sample the domain. Items can be written from personal experience, with help from experts, from the sample to be studied, or through literature searches.
Item Formats are divided into two types:
- Selected response format: Presents the examinee with a choice of answers and requires selection of one alternative. Types include multiple choice, matching, and true/false items.
- Constructed response format: Requires the examinee to provide or create the correct answer. Types include:
- Completion item: Requires providing a word or phrase that completes a sentence.
- Short answer item: Written clearly enough for a brief response.
- Essay: Asks examinees to describe a single topic in detail.
Scoring Items uses several models:
- Cumulative model: The most common model. The higher the score, the higher the ability or trait being measured. Test takers earn cumulative credit for responses made in a particular way.
- Class model: Test takers' responses earn credit toward placement in a particular class or category with others whose score patterns are presumably similar.
- Ipsative scoring: The objective is to compare a test taker's score on one scale within a test with another scale within that same test.
Test Tryout
After creating a pool of items, the test developer tries out the test on the sample for which it is constructed. A general rule is no fewer than five subjects, and preferably as many as ten subjects, for every one item on the test. The more subjects in the tryout, the better. The tryout should be executed under conditions as close as possible to the conditions under which the standardized test will be administered, including test instructions, time limits, and atmosphere.
What is a Good Item?
- A good test item should be valid and reliable.
- A good item helps discriminate test takers; high scorers on the test as a whole get it right.
- An item that high scorers do not get right is probably not a good item.
- A good item is one that low scorers on the test as a whole get wrong.
Item Analysis
Item analysis refers to the statistical procedures employed to select the best items from a pool of tryout items. These procedures include an index of an item's difficulty, an item-validity index, an item-reliability index, and an index of item discrimination.
Qualitative Item Analysis is a non-quantitative method that employs verbal rather than mathematical techniques. Through questionnaires or discussions with test takers, valuable information can be obtained on how the test could be improved.
Test Revision
Based on the information gathered during item analysis, some items are eliminated and others rewritten. The test developer characterizes each item by its strengths and weaknesses and may need to balance these across items. For example, if many good items are somewhat easy, the developer may purposefully include some more difficult items. After balancing concerns, the developer produces a test of improved quality. The revised test is then administered under standardized conditions, and if the item analysis is satisfactory, the test may be considered in its finished form.
⭐ Key Takeaways
The test development process is a systematic, five-step cycle: conceptualization, construction, tryout, item analysis, and revision. A critical early decision is whether the test will be norm-referenced (comparing individuals to a group) or criterion-referenced (comparing to a mastery standard), as this determines item selection criteria. During construction, test developers must choose appropriate scaling methods (like Likert or paired comparisons) and item formats (selected-response vs. constructed-response) based on the test's purpose and the number of examinees. A good item must be reliable, valid, and able to discriminate high scorers from low scorers. Finally, item analysis procedures, both quantitative and qualitative, guide the revision process to produce a psychometrically sound final instrument.
🧠 Quick Revision Questions
- What are the five steps of the test development process, and in what order do they occur?
- How does the definition of a "good item" differ between norm-referenced and criterion-referenced tests?
- What is the recommended minimum number of subjects per item for a test tryout?
- Describe the difference between selected-response and constructed-response item formats, providing two examples of each.
- What are the three scoring models discussed in the lecture, and what is the primary purpose of ipsative scoring?
📘 Lecture 11 — Item Writing
📖 Overview: This lecture focuses on the critical process of writing test items during test construction. It covers three primary considerations—content range, item format, and number of items—and provides a detailed examination of various item formats, including selected-response and constructed-response types. Understanding these principles is essential for developing valid and reliable psychological tests.
🗂️ Topics Covered
The lecture begins with three key considerations for item writing: nature and range of content, type of items, and number of items. It then explores item formats as classified by Kaplan and Saccuzzo (2001), including dichotomous, polytomous, Likert, category, checklists, and Q-sorts. The distinction between selected-response and constructed-response formats is discussed, with detailed explanations of each type, including matching items and correction for guessing.
📝 Lecture Summary
Considerations in Item Writing
The process of test construction requires careful planning, with three important considerations: what range of content the items should cover, which item formats to employ, and how many items to write. The test developer decides the nature and extent of content based on the objectives laid down for measurement, ensuring each item measures some aspect of the content area. The type of item format depends on the test type and construct being measured; for example, a projective item may be suitable for a personality test but not for an achievement test. The number of items affects test length and administration (individual vs. group). Whether measuring a single or multiple traits/abilities/domains also impacts item count—fewer items are needed for a single domain, while multiple domains require a larger number. When developing a standardized test with multiple-choice response format, the first draft should contain approximately twice the number of items that the final version will contain. The test developer may write items from personal experience, with help from experts, and through literature searches.
Item Formats
Kaplan and Saccuzzo (2001) describe several item formats: the dichotomous format, polytomous format, Likert format, category format, and checklists and Q-sorts. Test formats are also classified as recall type vs. recognition type, constructed response type vs. identification type, and objective type vs. essay type. Another common division is between selected-response format (where examinees choose from alternatives) and constructed-response format (where examinees create the answer).
Selected Response Format
This format presents the examinee with a choice of answers and requires selection of one alternative (e.g., on achievement tests, selecting the correct option). Types include multiple choice, matching, and true/false items.
a. The Dichotomous Format: This is the alternate response format providing two response options. One option is right and the other wrong, typically scored as one point per correct answer (e.g., ten right options = score of ten). Common variations include "yes"/"no" options (as in personality inventories) and true/false format. The test taker judges whether statements are true or false based on course texts.
🔑 Definition — Dichotomous format: An alternate response format where the test taker is provided with two response options to choose from, typically right/wrong, yes/no, or true/false.
Advantages: Easy to develop (no need for multiple plausible options), quick to develop, and easy to score. Disadvantages: A test taker can accurately answer at least 50% of answers by chance (probability = ½); if equal true and false items exist, marking all true yields 50% correct. Such items may encourage rote memorization rather than conceptual learning. To control disadvantages, use a large number of items covering a wider range of content.
b. The Polytomous (Polychotomous) Format: This format provides multiple response options rather than two. Each item includes a stem, a correct alternative option, and several incorrect alternatives called distracters or foils.
📌 Example:
- Stem: "A psychological test, an interview, and a case study are:"
- Correct Alt: a. psychological assessment tools
- Distractors: b. standardized behavioral samples, c. reliable assessment instruments, d. theory-linked instruments
The stem should clearly state the question or problem. Distractors should each appear to be the correct answer. Using a four-option format reduces chance correct to 25% (1/4); three options yield 33% chance. Testing experts recommend three or four options per item.
📐 Formula: Correction for guessing — Corrected score = R - W / (n - 1) Where R = number of right responses, W = number of wrong responses, n = number of choices per item. Plain English meaning: This formula adjusts the raw score by subtracting a penalty based on wrong answers divided by the number of options minus one.
📌 Example: A test taker scores 40 out of 100 on a four-option test. Corrected score = 40 - 60/(4-1) = 40 - 60/3 = 40 - 20 = 20. This shows guessing can reduce a score substantially when correction is used. The greater the number of wrong answers, the smaller the corrected score. Test developers should include all response options in the same proportions (not favoring one option like 'C' or 'D' more often).
Matching Item: A matching item presents two columns of responses where the test taker determines which item from the first column matches an item from the second column. If both columns have equal numbers, examinees can deduce the right answer by process of elimination. Providing more options than needed minimizes this possibility.
c. The Likert Format: The Likert scale is a popular tool for measuring personality and attitudes. It provides test takers options to endorse their degree of agreement with a statement. Likert used it as part of his method of attitude scale construction. Typically, five response options range from "strongly disagree" to "strongly agree," with "neutral" in between.
📌 Example: "I like to make friends who are older to me." Options: Strongly disagree, disagree, neutral, agree, strongly agree. To avoid the tendency to mark "neutral," six options may be used: Strongly disagree, moderately disagree, mildly disagree, mildly agree, moderately agree, strongly agree. Responses are summed to determine a person's score; negatively worded items are reverse scored and added to the total.
d. The Category Format: Similar to the Likert scale but with more options (commonly a 10-point scale, though the number may vary). Used to rate something (e.g., team performance). Raters may be inaccurate due to contrast effects—rating a player lower when compared to a superb player, or higher when compared to a poor performer. Some authors recommend showing raters videos of performances rated "10" and "1" to control this factor (Kaplan & Ernst, 1983).
e. Checklists and Q-sorts: Adjective checklists contain a list of adjectives; the test taker checks those true about themselves (or others). Used mostly in personality assessment. In Q-sort, the test taker receives statements and sorts them into nine piles based on how true they are of themselves; this can also be used for rating others.
Constructed Response Format
This format requires the examinee to provide or create the correct answer rather than selecting it. Three types: completion item, short answer, and essay.
- A completion item requires the examinee to provide a word or phrase that completes sentences.
- A good short answer item is written clearly so the test taker can respond briefly; there is no strict rule for length.
- An essay requires the examinee to describe a single topic in detail. Skills measured by essays differ from other formats—essays require recall, organization, planning, and writing ability, while other formats only require recognition.
💡 Why this matters: Understanding the differences between selected-response and constructed-response formats is crucial for selecting the appropriate item type based on the construct being measured and the level of cognitive processing required.
⭐ Key Takeaways
The three main considerations before writing items are content range, item format, and number of items—these must align with test objectives and the construct being measured. Dichotomous formats (true/false, yes/no) are easy to develop and score but suffer from high chance correct rates (50%), while polytomous formats (multiple choice) reduce chance but require well-constructed distractors. Correction for guessing using the formula R - W/(n-1) penalizes wrong answers and discourages random guessing. Likert and category formats measure degree of agreement or ratings but may be affected by response biases like "neutral" tendencies or contrast effects. Constructed-response formats (completion, short answer, essay) measure higher-order skills like recall and organization, unlike selected-response formats which only require recognition.
🧠 Quick Revision Questions
- What are the three important considerations a test developer must address before writing test items?
- What is the chance of answering a dichotomous item correctly by pure guessing, and how can this disadvantage be controlled?
- Using the correction for guessing formula, calculate the corrected score for a test taker who gets 35 right and 45 wrong on a 4-option test.
- What is the key difference between the Likert format and the category format in terms of response options?
- What cognitive skills does an essay item measure that are not measured by multiple-choice items?
📘 Lecture 12 — Item Writing: Guidelines For Item Writing
📖 Overview: This lecture covers the fundamental principles and practical guidelines for writing effective test items. It explains the essential steps in test planning and development, including how to avoid common pitfalls like double-barreled items and response sets. The lecture also addresses the specific considerations required for developing different types of tests, such as educational achievement, personality, intelligence, and screening tests, along with an introduction to various scoring models.
🗂️ Topics Covered
This lecture begins with DeVellis's guidelines for item writing, including clarity, item pool development, reading difficulty, item length, avoiding double-barreled items, and mixing positive and negative items. It then covers the initial test plan and design steps, followed by detailed discussions on developing educational achievement tests using taxonomies of objectives, personality tests based on theoretical constructs, intelligence tests considering multiple variables, and screening tests using job analysis. The lecture concludes with test scoring methods, including cumulative, class, and ipsative models, and important cautions about test construction.
📝 Lecture Summary
Item Writing: Guidelines For Item Writing
Every test developer must keep specific aims and objectives in mind while designing a test. DeVellis (1991) provided several guidelines for writing items. First, whatever is to be measured should be clearly defined and items should be as specific as possible. The test developer should prepare a large number of items first, developing 3-4 items for each one to be included in the final version, all representing the content area to be covered.
The reading difficulty level must be appropriate for the test takers, avoiding complex vocabulary that is not part of a layman's everyday vocabulary. Items should be of reasonable length, as exceptionally long items are rarely good. Test developers must avoid "double-barreled" items that include two or more ideas at the same time, such as "I often help the poor because I believe in serving humanity and I am a follower of my leader who himself is known for social service." Such items should be converted into two or more separate items.
🔑 Definition — Acquiescence response set: A tendency of test takers to agree with most items regardless of their content.
💡 Why this matters: This response set can invalidate test results by introducing systematic bias.
To overcome the acquiescence response set, test developers should mix negatively and positively worded items and use them alternately. If all items are negative (e.g., 'I hate liars', 'Most people betray', 'I feel depressed most of the time'), the test taker may develop a response set and mark every statement the same way. Additional considerations include avoiding double negatives (e.g., 'I do not disagree to the fact that people should not be stopped from playing cricket'), as they make it difficult for test takers to understand what is being asked. The cultural background of test takers should also be kept in mind, and cultural, racial, and gender bias must be completely avoided.
Initial Test Plan and Design:
Test development involves several sequential steps: determining and formulating the observations of the test; deciding the domain and content to be covered; deciding about the format of the test and the test items; developing/writing the items; developing much more items than required; trying out the initial version; analyzing the results of the try out; reshaping and refining the first version; another try out if needed; and norm development if required (standardized tests only). Beyond these common steps, there are considerations specific to certain types of tests.
Development of Educational Achievement Tests:
In educational tests, the most important element is the formulation and use of educational objectives stated in behavioral terms. Educational objectives define behaviors that will indicate whether the content has been learned. Several formal systems, known as taxonomies of educational objectives, help test developers.
The most popular system is Bloom's Taxonomy of Educational Objectives: The Cognitive Domain (Bloom & Krathwohl, 1956), with major categories including: Knowledge, Comprehension, Application, Analysis, Synthesis, and Evaluation. Educational Testing Service's (1965) taxonomy includes: Remembering, Understanding, and Thinking. Gerlach and Sullivan's (1967) system includes: Identifying, Naming, Describing, Constructing, Ordering, and Demonstrating. Ebel's (1979) system covers: Understanding of terminology, Understanding of fact and principle, Ability to explain or illustrate, Ability to calculate, Ability to predict, Ability to recommend appropriate action, and Ability to make an evaluative judgment.
When planning a test, developers often create a table of specification — a two-way table with behavioral objectives on the vertical axis (row headings) and content/topics on the horizontal axis (column headings). This table shows the type of objectives corresponding to specific content areas and the total number of items in each box.
📐 Formula: Table of Specification structure → Behavioral objectives (rows) × Content topics (columns) showing number of items for each objective-topic combination
Development of Personality Tests:
In personality tests, the theoretical approach of the test developer matters most. Items are based on the constructs that are to be measured.
Development of Intelligence Tests:
In planning intelligence tests, several variables must be considered: what aspects of intelligence are to be measured, what criteria will be used (e.g., age, grade, or other), the test format considered suitable, and the target population.
Development of Screening Tests:
Aptitude tests are generally used for screening purposes. The objective is to identify candidates most suitable for the job who fulfill the position's requirements. First, a job analysis or task analysis is conducted, identifying and listing the various components of the job. Test items are based on these tasks and job-related situations or 'critical incidents'.
Test Scoring:
Different tests are scored differently. For objective tests, scoring is done with a scoring key, either manually or using computers with attached scanners. Scoring essay type items is more difficult and can be done using 'holistic' or 'global' scoring (scoring the whole response) or analytic scoring (scoring different components separately), with the latter being the better approach. Involving another person for rescoring can increase objectivity. Standardized tests have their scoring procedures specified in the test manual.
Scoring Items:
There are several scoring models. The most common is the cumulative model, where the higher the score, the higher the ability or trait being measured. For each response made in a particular way, the test taker earns cumulative credit with regard to a particular construct.
The class model is a second model where test takers' responses earn credit toward placement in a particular class or category with others whose score pattern is presumably similar.
A third model is ipsative scoring, where the typical objective is comparing a test taker's score on one scale within a test with another scale within that same test.
🔑 Definition — Ipsative scoring: A scoring model where a test taker's score on one scale is compared with another scale within the same test.
Some Cautions:
Test developers should avoid jargon and use vocabulary that most people can understand. The test should not be too long. It should not be so easy that everyone can do it, nor so difficult that no one can do it. Cultural biases and stereotypical ideas must be avoided. If cultural differences are suspected that may affect test results, nonverbal and culture-free tests should be used. However, even culture-free tests may not be fully culture fair, as people with previous experience and exposure to such items may outperform those from cultures where tasks involving drawings and images are unfamiliar.
⭐ Key Takeaways
The most critical points from this lecture are the specific guidelines for item writing, including clarity of definitions, appropriate reading difficulty, avoiding double-barreled items and double negatives, and mixing positively and negatively worded items to prevent acquiescence response sets. Students must understand the complete test development sequence from determining observations through norm development, and be able to apply Bloom's taxonomy and the table of specification for educational achievement tests. For different test types, remember that personality tests depend on theoretical constructs, intelligence tests require consideration of multiple variables, and screening tests rely on job/task analysis. Finally, know the three scoring models (cumulative, class, and ipsative) and the crucial cautions about test length, difficulty, and cultural bias.
🧠 Quick Revision Questions
- What is a "double-barreled" item and why should it be avoided in test development?
- How does mixing positively and negatively worded items help overcome the acquiescence response set?
- List the six major categories of Bloom's Taxonomy of Educational Objectives for the Cognitive Domain.
- What is the purpose of a table of specification in educational test development?
- What are the three scoring models for test items, and how does each model interpret test taker responses?
📘 Lecture 13 — Reliability
📖 Overview: This lecture introduces the concept of reliability in psychological testing, explaining why it is critical for trustworthy measurement. It covers the definition of reliability, the classical test score theory, sources of error in test scores, and the role of correlation in calculating reliability. Understanding reliability is fundamental because it determines whether a test produces consistent and dependable results for serious decisions.
🗂️ Topics Covered
This lecture begins with the definition of reliability and its importance in psychological testing, distinguishing it from physical sciences. It then explains the classical test score theory (X = T + E) and details various sources of error, including test-related factors, administration process variables, and examinee-related variables. The lecture proceeds to explain the concept of correlation as it relates to reliability, including the coefficient of correlation, its magnitude and direction, and the interpretation of scatter plots for positive, negative, and zero correlations.
📝 Lecture Summary
Reliability
Reliability is one of the three basic qualities of a good psychological test, along with validity and standardization. By definition, reliability means “the consistency of the scores obtained by the same persons when they are reexamined with the same test on different occasions, or with different sets of equivalent items, or under other variable examining conditions” (Anastasi & Urbina, 2007). According to Kaplan & Saccuzzo (2001), reliability is “the extent to which a score or measure is free from measurement error. Theoretically, reliability is the ratio of true score variance to observed score variance.”
The classical test score theory implies that everyone can obtain a true score on any test if the measure is free of error. However, there is no measure that can be considered totally error-free. The scores we obtain are observed scores, which equal the true score plus error. The term error here does not mean a mistake; it refers to the amount and extent of variance that may be expected in results.
🔑 Definition — True Score (T): The score a person would obtain if the measure were completely free of error. 🔑 Definition — Observed Score (X): The actual score obtained on a test, which is the sum of the true score and error. 📐 Formula: X = T + E → The observed score equals the true score plus error. 💡 Why this matters: This formula is the foundation of reliability theory. It means every test score contains some degree of error, and reliability is about estimating how much error is present.
Sources of Error in Test Scores
Psychologists and educationists try to estimate the amount and degree of error that may be expected in measurement. This is why test developers and administrators emphasize uniform testing procedures and controlling possible sources of error.
a. Test Related Factors:
- Difficulty level (too difficult or too easy)
- Length of the test (too long causing fatigue or boredom)
- Domain or content (not suitable for all test takers)
- Items may not be representing the domain or content
- Time limit (may cause stress or a handicap)
b. Test Administration Process:
- Poor and not uniform testing conditions (physical setting and environment)
- Poor, faulty, improperly worded, improperly delivered, not uniform instructions
- Test administrator’s personality (different administrators in different situations)
- Rapport (poor, too much, or too little)
c. Examinee Related Variables:
- Prior learning and experience
- Individual differences (personality styles, stress tolerance, IQ, emotionality, motivation, knowledge)
- Difference from the normative sample
- Within-person differences (changes over time due to life experiences, health, motivation, emotional state)
Test developers try to control or keep these variables constant, but complete consistency is impossible. However, test developers report the coefficient of reliability and characteristics of the normative sample, guiding users about appropriate populations for the measure.
Correlation and Reliability
Reliability is about consistency of scores. Calculation of reliability involves the concept of correlation. The coefficient of correlation is the value yielded by the calculation of correlation, denoted by the letter ‘r’. This coefficient indicates the relationship between two variables and expresses the correspondence between two scores.
The coefficient of correlation tells us two things about a relationship: magnitude (size of correlation) and direction (whether positive or negative). The size of a correlation ranges between -1 and +1, with zero in between. A value of +1 means a perfect positive correlation, -1 means a perfect negative correlation, and zero indicates no correlation. The closer the value is to one, the stronger the correlation.
If scores on two sets increase and decrease together, it is a positive correlation. If scores on set-I increase while set-II decrease, it is a negative correlation. The concept can be understood through a scatter plot. If the lowest scorer on set-I is also lowest on set-II, and highest is highest on both, it is a perfect positive correlation. If the lowest on set-I is highest on set-II, it is a perfect negative correlation.
🔑 Definition — Coefficient of Correlation (r): A numerical value ranging from -1 to +1 that expresses the magnitude and direction of relationship between two variables. 📐 Formula: Pearson Product Moment Correlation (r) → The most commonly used procedure for computing the coefficient of correlation. 📌 Example: Scores from 10 students on IQ and marks in math showed a perfect positive correlation (r = +1.0), as every student’s IQ matched their math marks exactly. Scores on IQ and classes missed showed a perfect negative correlation (r = -1.0), as higher IQ was associated with fewer missed classes. Marks in English and math showed zero correlation (r = 0.0), with no discernible pattern between the two.
The size of the correlation coefficient generally acceptable for reliability is around .80 or .90. A coefficient less than .60 is usually not acceptable. For data from a large number of people, a coefficient less than .80 may also be acceptable.
⭐ Key Takeaways
Reliability is the consistency of test scores and is expressed as the ratio of true score variance to observed score variance, following the formula X = T + E. The three main categories of error sources are test-related factors, administration process variables, and examinee-related variables, which must be controlled to improve reliability. Correlation, specifically the Pearson Product Moment method, is the statistical foundation for calculating reliability, with acceptable coefficients generally above .80. The magnitude and direction of correlation are both important: magnitude ranges from -1 to +1, and direction indicates whether the relationship is positive or negative. A coefficient below .60 is typically unacceptable for reliability purposes.
🧠 Quick Revision Questions
- What is the fundamental formula of classical test score theory, and what does each component represent?
- List three test-related factors that can introduce error into test scores.
- What is the range of possible values for a coefficient of correlation?
- In a scatter plot, what pattern indicates a perfect negative correlation?
- What is the minimum coefficient of correlation generally considered acceptable for reliability?
📘 Lecture 14 — Types of Reliability
📖 Overview: This lecture explores the various methods used to measure the reliability of psychological tests. Understanding these different types is crucial because the choice of method depends on what aspect of consistency is being evaluated—whether stability over time, equivalence between test forms, or internal consistency among items. This knowledge is fundamental for evaluating the quality and trustworthiness of any psychological measure.
🗂️ Topics Covered
This lecture covers five primary types of reliability: test-retest reliability, which assesses stability over time; alternate-form reliability, which examines equivalence between different versions of a test; split-half reliability, a single-administration method measuring internal consistency; Kuder-Richardson reliability, which analyzes inter-item consistency for dichotomously scored items; and coefficient alpha, a general formula for tests with polytomous items. The lecture concludes with a discussion of inter-scorer reliability for subjectively scored tests.
📝 Lecture Summary
Reliability: Definition and Forms
According to Anastasi & Urbina (2007), reliability refers to “the consistency of the scores obtained by the same persons when they are reexamined with the same test on different occasions or with different sets of equivalent items, or under other variable examining conditions”. This definition highlights three key forms of consistency: consistency when the same test is taken again on different occasions, consistency when different but equivalent sets of items are used, and consistency under other variable conditions.
a) Test-Retest Reliability
Test-retest reliability deals with two performances of the same test by the same persons on two different occasions. It measures the consistency and stability of scores over time. The resulting test-retest coefficient is also known as the ‘coefficient of stability’. Scores from the first administration are correlated with scores from the second administration. A major advantage is maximum control over test taker and test item variables, as both the subjects and the items remain the same.
This type of reliability accounts for errors of measurement from several sources:
- Testing conditions may not remain constant (e.g., noise, temperature).
- Changes in the test taker (bodily, emotional, motivational) can occur.
- Fatigue effect or practice effect can introduce error variance.
- The nature of the test may be affected by repeated measurement (e.g., memory of items).
- The length of the time interval between administrations affects the coefficient's magnitude; a shorter interval yields a larger coefficient.
Test-retest reliability is denoted as rₜₜ.
💡 Why this matters: The length of the time interval is a critical factor. A test developer must report the interval between administrations because a reliability coefficient is meaningless without this context.
b) Alternate-Form Reliability
Alternate-form reliability overcomes problems of using the same test twice. The developer creates two alternate or parallel forms that should be completely equivalent in all respects, including same specifications, instructions, time limit, content, number of items, item format, and difficulty level.
When alternate forms are administered in immediate succession (no time gap), the coefficient is the ‘coefficient of equivalence’ or ‘parallel form coefficient’. To compute stability over time (coefficient of stability), the two forms are administered with a time gap. A ‘coefficient of stability and equivalence’ is obtained by counterbalancing the administration of forms:
- Group I takes Form A first, Group II takes Form B first (Occasion I).
- After a time interval, Group II takes Form A, and Group I takes Form B (Occasion II).
This procedure accounts for error from both different sets of items and the time interval. While a good approach, its shortcomings include the difficulty of creating truly equivalent forms and the potential for practice effects across forms.
c) Split-Half Reliability
Split-half reliability avoids problems of constructing parallel forms and the influence of time gaps. In this single-administration approach, the test is divided into two equal halves, and scores on one half are correlated with scores on the other. The obtained value is the ‘coefficient of internal consistency’.
The test can be divided in several ways:
- Dividing by first 50% and last 50% items (problematic due to potential differences in difficulty).
- A better method is dividing into odd-numbered items and even-numbered items, ensuring similar difficulty and content in both halves.
Because the correlation is based on only half the test, it must be adjusted. The Spearman-Brown formula estimates the reliability of the full test.
🔑 Definition — Spearman-Brown formula: A formula used to estimate the internal consistency reliability of a test from the reliability of its two halves, correcting for the fact that the test has been shortened or lengthened.
📐 Formula: rₙₙ = (n * rₜₜ) / (1 + (n-1) * rₜₜ) Where:
- rₙₙ = estimated coefficient for the lengthened/shortened test
- rₜₜ = obtained coefficient from the half-tests
- n = number of times the test has been lengthened or shortened (e.g., for split-half, n=2)
For split-half reliability, the formula becomes: rₜₜ = (2 * rₕₕ) / (1 + rₕₕ) Where rₕₕ is the reliability of the half tests.
d) Kuder-Richardson Reliability
The Kuder-Richardson (K-R) method measures reliability through inter-item consistency using a single test administration. Unlike split-half, it examines performance on each item and is the mean of all possible split-half coefficients. It is only used for tests where items are scored dichotomously (e.g., 0 or 1).
📐 Formula: r = [n / (n-1)] * [1 - (∑pq / SD²ₜ)] Where:
- r = Coefficient of reliability of whole test
- n = number of items
- SD²ₜ = Variance of total scores
- p = proportion of persons who pass each item
- q = proportion of persons who do not pass each item
- ∑pq = sum of the products of p and q for all items
e) Coefficient Alpha
Coefficient alpha (also known as Cronbach’s alpha) is similar to the K-R method but is a general formula that works for tests with more than two scoring weights (e.g., Likert-scale items where 1=never, 2=occasionally, 3=often).
📐 Formula: α = [n / (n-1)] * [1 - (∑s²ᵢ / SD²ₜ)] Where:
- α = Coefficient of reliability of whole test
- SD²ₜ = Variance of total scores
- s²ᵢ = Variance for each item’s scores
- ∑s²ᵢ = sum of the variances of all item scores
Inter-Scorer Reliability / Inter-Rater Reliability
Inter-scorer reliability (or inter-rater reliability) is used when a test requires evaluation by more than one examiner, such as for essay-type or open-ended items. Common procedures include:
- Inter-rater reliability: Two examiners score the test, and the two sets of scores are correlated.
- Intra-class coefficient / coefficient of concordance: Two or more examiners score the performance, and a coefficient is computed from all scores.
⭐ Key Takeaways
The most critical concepts from this lecture are the distinction between methods that require two administrations (test-retest and alternate-form) versus single-administration methods (split-half, Kuder-Richardson, and coefficient alpha). You must remember the specific error sources each method controls: test-retest assesses stability over time but is affected by practice and time interval; alternate-form assesses equivalence across test versions but is difficult to construct; split-half provides a measure of internal consistency using the Spearman-Brown formula. The key difference between Kuder-Richardson and coefficient alpha is that K-R is only for dichotomous items, while alpha is for polytomous items. Finally, inter-scorer reliability is essential for subjectively scored tests.
🧠 Quick Revision Questions
- What is the difference between a "coefficient of stability" and a "coefficient of equivalence"?
- Why is the Spearman-Brown formula necessary in split-half reliability?
- What is the primary limitation of the Kuder-Richardson formula that coefficient alpha overcomes?
- List three potential sources of error variance in test-retest reliability.
- When would an intra-class coefficient be used instead of a simple inter-rater correlation?
📘 Lecture 15 — Reliability in Specific Conditions and Allied Issues
📖 Overview: This lecture examines how reliability estimation methods must be adapted for special testing conditions, particularly oral tests and speed tests. It also covers how sample characteristics affect reliability coefficients, introduces the Standard Error of Measurement (SEM) as an alternative expression of reliability, and discusses acceptable reliability values and methods for enhancing reliability.
🗂️ Topics Covered
This lecture covers reliability issues in oral tests including inter-scorer reliability and methods to improve it; the distinction between speed tests and power tests and why split-half reliability is inappropriate for speed tests; the relationship between sample characteristics (variability and ability level) and reliability coefficients; the concept and calculation of Standard Error of Measurement (SEM); acceptable reliability values for research purposes; and strategies for enhancing test reliability by increasing item numbers and improving item quality.
📝 Lecture Summary
Some Factors Influencing Reliability of a Test:
Reliability of Oral Tests:
In oral tests, inter-scorer reliability becomes a significant concern because examiners must assess responses subjectively. Unlike objective written tests where answers are clearly right or wrong regardless of who marks them, oral test scoring is problematic because different raters may evaluate the same response differently. The reliability of oral tests is usually lower than equivalent written tests. To improve reliability, several measures have been recommended: test takers should be instructed to think before responding (Meredith, 1978), electronic recordings of responses should be made for later reevaluation, oral tests should be carefully designed, model questions should be constructed before administration, and more than one rater should be used. Research shows that when these factors are considered, obtained reliability coefficients range from .60 to .70 (Carter, 1962; Levine & McGuire, 1970; Hitchman, 1966).
Reliability of Speed Tests:
Speed tests focus on how quickly a person completes a test — items are uniformly easy, but the time limit is so short that no one can finish all items. As Anastasi & Urbina (2007, p.116) state: "A pure speed test is one in which individual differences depend entirely on speed of performance." In contrast, power tests have time limits long enough for everyone to attempt all items, but the difficulty is steeply graded so no one can get a perfect score. Most tests contain both features to varying degrees.
💡 Why this matters: The nature of the test (speed vs. power) directly affects which reliability methods are appropriate. Using the wrong method can produce misleading, inflated reliability coefficients.
The split-half procedure using odd-even format is NOT suitable for speed tests. Because speed test items have uniform difficulty, a person who attempts X items correctly will likely attempt equal numbers of odd and even items. This can produce an inflated coefficient (possibly r = +1.00) that does not reflect true reliability. Instead, the following approaches should be used: (a) Test-retest reliability, (b) Equivalent-forms reliability, (c) Variations of split-half method where the two halves are administered separately as if they were two tests, with half the total time for each segment.
📐 Formula Application Example: For a test with 50 items and 60 minutes total time → Odd items (25 items, 30 minutes) and Even items (25 items, 30 minutes) administered separately.
Relationship between the Sample Tested and Reliability:
Two sample characteristics affect reliability coefficients:
Variability: Reliability involves correlation, and correlation requires variance in scores. If scores in a group are identical or very similar, the correlation may be zero. Therefore, the sample should have a wide range of scores and diverse test takers. Test manuals must provide detailed information about the sample used for computing reliability.
Ability level: While lack of variability causes problems, too much variation between subgroups can also introduce error. If the sample contains subgroups very different in ability (especially when difficulty level and prior learning matter), this can distort the score distribution from which correlation is calculated. Subgroups within the larger sample should be chosen with great care.
Standard Error of Measurement (SEM):
SEM is defined as "an index of the amount of error in a test or measure. The standard error of measurement is a standard deviation of a set of observations for the same test" (Kaplan & Saccuzzo, 2001). It is a way of expressing reliability.
📐 Formula: SEM = SD√(1 - rᵵᵵ) where SD = standard deviation of test scores and rᵵᵵ = reliability coefficient
📌 Example: If a test has SD = 12 and reliability coefficient = .90, then SEM = 12√(1 - .90) = 12√0.10 = 12 × 0.316 = 3.79. If a person's true score (mean of hypothetical distribution) is 90, then 68% of people may be expected to score within the range of 90 - 3.79 to 90 + 3.79 (approximately 86.21 to 93.79), because in a normal distribution about 68% of scores fall between -1 and +1 standard deviations from the mean.
Acceptable Values of Reliability:
The closer the reliability coefficient is to 1.00, the better. However, for most research purposes, coefficients within the range of .70 to .80 are considered good enough (Kaplan & Saccuzzo, 2001).
Enhancing Reliability:
Increasing the number of items in a test may improve reliability. Additionally, test items should be carefully phrased and it should be ensured that all items measure the same content that the test is intended to measure.
⭐ Key Takeaways
The most critical points from this lecture are: (1) Oral tests require special attention to inter-scorer reliability, and using multiple raters, recordings, and careful design can achieve coefficients in the .60-.70 range. (2) Split-half reliability methods should never be used for speed tests because uniform item difficulty artificially inflates coefficients; instead use test-retest, equivalent-forms, or separately administered halves. (3) Sample characteristics (variability and ability level) significantly affect reliability coefficients — samples must have adequate score range without extreme subgroup differences. (4) The Standard Error of Measurement (SEM = SD√(1-rᵵᵵ)) expresses reliability as a standard deviation of error scores, allowing you to estimate the range within which a person's true score likely falls (e.g., ±1 SEM contains 68% of observations). (5) For research purposes, reliability coefficients of .70-.80 are acceptable, and reliability can be enhanced by increasing the number of items and ensuring all items measure the same content.
🧠 Quick Revision Questions
- Why is inter-scorer reliability a greater concern for oral tests than for written objective tests?
- What distinguishes a pure speed test from a pure power test in terms of item difficulty and time limits?
- Why does the split-half odd-even method produce an inflated reliability coefficient for speed tests, and what three alternative methods should be used instead?
- How do sample variability and ability level influence the reliability coefficient?
- If a test has a standard deviation of 15 and a reliability coefficient of .84, what is the SEM, and within what range would 68% of examinees' true scores fall?
📘 Lecture 16 — Validity
📖 Overview: This lecture defines validity as an essential characteristic of psychological tests, explaining that a valid test measures what it claims to measure. It explores the three main types of validity evidence—content, criterion, and construct—and details procedures for establishing each type, emphasizing specific reporting rather than general labels.
🗂️ Topics Covered
The lecture begins by drawing analogies from everyday life to explain validity and reliability, then defines validity formally. It distinguishes between achievement and predictive test purposes, introduces the three broad categorizations of validity—content, construct, and criterion—and details content-description procedures, including specific steps and applications, before contrasting content validity with face validity and outlining criterion prediction procedures.
📝 Lecture Summary
Validity
Validity is an essential ingredient of a test. The use of a test that is not valid will not only be a waste of time and energy, but it may also result in erroneous judgments and decisions. Validity, just like reliability, is a characteristic and quality that we seek in most life situations where we have to acquire something, make it a part of our life, and base decision making on that. For example, if we need to buy an air conditioner, we would like to make sure that it cools the room, providing the service that it is meant to provide (validity); we would also like to ensure that every time we switch it on it will turn on and start working (reliability). If the machine fails in cooling, and/or does not turn on when switched on, then we do not need it.
In human relationships, we want our friends to be our ‘friends’ in the true sense—warm, sympathetic, understanding, caring, etc., i.e., what a friend ought to be (validity). On the other hand, we like that our friend is there to support us, help us, and be with us every time we need him/her (reliability). Similarly, psychological tests need these features as an essential integral part.
🔑 Definition — Validity: Traditionally defined as “the extent to which a test measures what it was designed to measure” (Aiken, 1994, p.95). Another definition: “The extent to which a test measures the quality it purports to measure. Types of validity evidence include content validity, criterion validity, and construct validity evidence” (Kaplan & Saccuzzo, 2001, p.640).
From these definitions, one can realize that validity of a test is about the nature of a test with reference to the content of the test as well as the content or domain in which it is rooted. It is about what the test measures and how well it measures. Although the title or name of a test gives us a clue of what the test measures, we may not have an accurate idea of what it actually measures until and unless we have appropriate information. This information may be reported in the test manual or may be calculated by us.
Whether or not a test measures a trait that it claims to measure can be determined only through an examination of the objective sources of information and empirical operations utilized in establishing its validity. A test’s validity must be established with reference to the particular use for which the test is being considered. The validity of a test should not be taken casually, and cannot be reported in general terms either. Validity is not, and should not be measured and reported as “highly” valid or “low” validity. It is reported in specific terms.
Procedures for determining test validity are concerned with the relationships between the test content and the domain that it represents; between the test content and the objectives that lead to test construction; or test performance and performance on other independent measures that are known to measure the same phenomenon/trait/ability/domain, etc.
Psychological tests are developed primarily for two reasons:
- Achievement testing: In order to see if a specific content area has been learned and/or mastery is acquired.
- Predictive function: Predicting future performance on the basis of the present test’s performance.
In both cases, the ultimate objective cannot be achieved if the test is not valid. An achievement test is evaluated by comparing its content with the content domain it is designed to assess. These tests may be used as end-of-course examinations in school and by licensing tests for driving a car or qualifying with a skill for a specified occupation. These focus on content domain. The performance tests used to assess and predict future behavior are designed on certain criterion. To measure the validity of such tests, the correlation coefficient between test scores and a direct and independent measure of that criterion is used.
Another area with reference to validity is about testing the construct in a particular test. Constructs are broad categories, derived from the common features shared by directly observable behavioral variables.
💡 Why this matters: Understanding the three types of validity—content, criterion, and construct—is fundamental because each applies to different test purposes, and a test can be valid for one use but not another. Reporting validity in specific, not general, terms prevents misinterpretation.
Content Validity Evidence
“The evidence that the content of a test represents the conceptual domain it is designed to cover” (Kaplan & Saccuzzo, 2001, p.635).
🔑 Definition — Content Validity Evidence: Evidence that the content of a test represents the conceptual domain it is designed to cover.
Construct Validity Evidence
“A process used to establish the meaning of a test through a series of studies. To evaluate evidence for construct validity, a researcher simultaneously defines some construct and develops the instrumentation to measure it. In the studies, observed correlations between the test and other measures provide evidence for the meaning of the test” (Kaplan & Saccuzzo, 2001, p.635).
🔑 Definition — Construct Validity Evidence: A process used to establish the meaning of a test through a series of studies, where a researcher defines a construct and develops instrumentation to measure it, and observed correlations provide evidence for the test's meaning.
Criterion Validity Evidence
“The evidence that a test score corresponds to an accurate measure of interest. The measure of interest is called the criterion” (Kaplan & Saccuzzo, 2001, p.635).
🔑 Definition — Criterion Validity Evidence: Evidence that a test score corresponds to an accurate measure of interest, where the measure of interest is called the criterion.
Content-Description Procedures
Nature:
- A systematic examination of the test content is made in order to determine whether it covers a representative sample of the behavior domain to be measured.
- Such validation procedures measure how well the individual has mastered a specific skill or course of study.
“Content validity is built into a test from the outset through the choice of appropriate items. For educational tests, the preparation is preceded by a thorough and systematic examination of relevant course syllabi and textbooks, as well as consultation with subject matter experts. On the basis of the information thus gathered, test specifications are drawn up for the item writers” (Anastasi & Urbina, 2007, p. 129)
Certain considerations are needed for content validation measures:
- The test should cover all major aspects and items in correct proportion (overloading and under-representation of elements should be avoided).
- The domain under consideration must be fully described in advance rather than after test preparation.
- Content should be broadly covering the major objectives, application and interpretation, and also factual knowledge.
Specific Procedures: The choice of appropriate items is important in content validity. A number of careful decisions are taken at the time of test development:
- For educational tests, items are prepared from relevant course syllabi and textbooks, and in some cases, consultation with subject experts is used.
- Test specifications are given to item writers. They include instructional objectives, relative importance of topics/processes, clearly indicate the number of items for each topic, and may also provide sample material.
- The process of content validation should include description of all procedures in the manual that ensure the test is appropriate and representative. For example, if subject specialists have participated, the names, qualifications, and number of individuals should be stated.
- The test total score and achievement score of an individual can be checked for grade progress.
Application:
- Content validity is basic to the validity of educational and occupational achievement tests.
- This validity measure is applicable to occupational tests designed for employee selection and classification.
- For aptitude and personality tests, content validation is inappropriate.
Content Validity versus Face Validity
Content validity should not be confused with face validity. Many students of psychology take them to be one and the same, whereas actually they are different. Face validity is about the test appearing to be valid. It is not validity in the technical sense; nor is it technically measured.
🔑 Definition — Face Validity: “The extent to which items on a test appear to be meaningful and relevant, actually not evidence for validity because face validity is not a basis for inference” (Aiken, 1994, p.636).
Face validity may be used where prediction or inference are not involved, e.g., a questionnaire or scale for a survey.
Criterion Prediction Procedures
Criterion validity evidence is “The evidence that a test score corresponds to an accurate measure of interest. The measure of interest is called the criterion” (Kaplan & Saccuzzo, 2001, p.635).
🔑 Definition — Criterion Prediction Procedures: Procedures used to establish criterion validity by correlating test scores with a direct and independent measure of the criterion of interest.
⭐ Key Takeaways
A test is valid only if it measures what it claims to measure, and validity must be established with reference to a specific purpose—not reported in general or vague terms. There are three main types of validity: content validity (how well test content represents a domain), criterion validity (how well test scores predict or correspond to an external measure), and construct validity (how well the test measures an abstract concept through a series of studies). Content validity is built into educational and occupational achievement tests through careful item selection and requiring domain description in advance, while it is inappropriate for aptitude and personality tests. Face validity, which is about a test appearing valid, is technically not a form of validity and should not be confused with content validity.
🧠 Quick Revision Questions
- What is the formal definition of validity according to Aiken (1994)?
- What are the two main reasons for which psychological tests are developed?
- How does content validity differ from face validity?
- What are the specific procedures used to ensure content validity in educational tests?
- What is a construct, and how is construct validity evidence established?
📘 Lecture 17 — Criterion Validity
📖 Overview: This lecture explores criterion-related validity, one of the fundamental approaches to establishing test validity. It explains how test scores are compared against external criteria to determine if a test measures what it claims to measure. This lecture is critical because criterion validity is essential for tests used in selection, diagnosis, and prediction of future performance.
🗂️ Topics Covered
The lecture begins by reaffirming validity as the most essential feature of a test, then introduces the three main approaches to validity assessment (content, criterion-related, and construct). It focuses on criterion-related validity, explaining its conceptual foundation through everyday examples. The two main procedures—predictive validity evidence and concurrent validity evidence—are described in detail with their distinctions and applications. The lecture concludes with common criterion measures and the essential characteristics a criterion must possess, including reliability, validity, relevance, and freedom from contamination.
📝 Lecture Summary
Criterion-related Validity
When planning and designing a test, we have a certain standard in mind that we want to meet. We relate scores on our test with those achieved on the criterion or standard. One approach to assessment of validity is through comparing the test results with those on a criterion.
🔑 Definition — Criterion-related validity: A procedure where scores on a test being used are correlated with scores on a criterion. As defined by Kaplan & Saccuzzo (2001, p.635), "The evidence that a test score corresponds to an accurate measure of interest. The measure of interest is called the criterion."
📌 Example: A teacher develops a test of mathematical ability. To determine its validity, she examines the relationship between scores on her test with another authentic test of the same ability or a test of similar skills. By doing so, she is assessing the criterion-related validity of her test.
The criteria used for this purpose may be in the form of scores on psychological tests, mental and behavioral measurement, classifications, grades, teacher's or supervisor's ratings, etc. According to Aiken (1994, p.96), "Traditionally, however, criterion-related validity has been restricted to validation procedures in which the test scores of a group of examinees are compared with ratings, classifications, or other behavioral or mental measurements."
Criterion Prediction Procedures
Two procedures may be adopted for this purpose:
- Predictive validity evidence
- Concurrent validity evidence
Predictive Validity Evidence
Predictive validity evidence: "The evidence that a test forecasts score on the criterion at some future time" (Kaplan & Saccuzzo, 2001, p.638). Predictive validity pertains to prediction over a time interval. Predictive validation yields information that is most relevant to tests that are used for predicting future performance of an individual.
📌 Example: Tests are used for selecting students for an academic program, for personnel selection and job placement, or for choosing staff members for a specific training or higher education. Tests are also used to predict if a candidate is likely to develop an undesirable behavior, characteristic, mental or emotional problem in future.
Concurrent Validity Evidence
In many situations, estimating validity on predictive validation basis is not too practical. It involves a certain time interval and requires a suitable problem selection sample. In such situations, it is preferable to use the concurrent validation approach. In this approach, the test is administered to such a sample whose data on the criterion are already available.
🔑 Definition — Concurrent validity evidence: "Evidence for criterion validity in which the test and the criterion are administered at the same point in time" (Kaplan & Saccuzzo, 2001, p.635). When the validity of a test is measured with reference to a criterion, and both are administered at nearly the same time, then the resulting validity will be concurrent validity.
📌 Example: One of the best examples of concurrent validity is diagnostic tests such as the Minnesota Multiphasic Personality Inventory (MMPI). The test is administered to people in various categories to determine if there are significant differences between the average scores of different categories of individuals.
💡 Why this matters: Anastasi & Urbina (2007) clearly distinguish between predictive and concurrent validation: "The logical distinction between predictive and concurrent validation is based not on time but on the objectives of testing. Concurrent validation is relevant to tests employed for diagnosis of existing status rather than prediction of future outcomes. The difference can be illustrated by asking 'Does Smith qualify as a satisfactory pilot?' or 'Does Smith have the prerequisites to become a satisfactory pilot?' The first question calls for concurrent validation; the second, for predictive validation."
Criterion Measures
There is no limit to the type of criterion that may be used for the purpose of test validation. Some of the criteria mentioned in the lecture include:
- Academic achievement: used for intellectual tests and similar tests
- Matriculation marks: at the time of admission to higher classes, used for selection and screening
- Years of education: for those who have completed or discontinued education
- Performance in specialized training: used for selection, admission, achievement
- Instructor's ratings
- Grades in a semester
Characteristics of a Criterion
According to Cohen (1999), a criterion should be:
- Reliable: The scores on the criterion as well as those on the test should be reliable. The reliability coefficients of the two sets of scores affect the coefficient of validity. This can be understood from the following pattern of relationship:
📐 Formula: R<sub>xy</sub> (coefficient of validity) is affected by r<sub>xx</sub> (test reliability) and r<sub>yy</sub> (criterion reliability). Weak/little/or no reliability will automatically affect validity.
- Valid: It should measure what it is supposed to measure.
- Relevant: It should be relevant to the purpose for which it is being used.
- Uncontaminated: We should try to use uncontaminated criterion. Criterion contamination occurs when the criterion itself has been, partially or fully, based on predictor measures.
📌 Example: You use a test as a criterion for the diagnosis of patients examined in a psychiatric clinic. A probe into the diagnostic procedure indicates that the diagnosis itself involved the use of the same test. In this case, the criterion will be contaminated and will not be a suitable criterion.
⭐ Key Takeaways
Validity is the most essential feature of a test—a test can be reliable without being valid, but it cannot be valid without being reliable. Criterion-related validity involves correlating test scores with scores on an external criterion, and it has two forms: predictive validity (forecasting future performance over a time interval) and concurrent validity (administering test and criterion at the same time for diagnosis of current status). The key distinction between the two is based on the objective of testing, not on time. For a criterion to be useful, it must be reliable, valid, relevant, and uncontaminated—meaning it should not be based partially or fully on the predictor measure itself.
🧠 Quick Revision Questions
- What is the difference between predictive validity and concurrent validity, and how does the objective of testing determine which one to use?
- Why can a test be reliable without being valid, but a test cannot be valid without being reliable?
- What is criterion contamination, and provide an example where it would invalidate the criterion?
- According to Cohen (1999), what are the four essential characteristics a criterion must possess?
- In the example of the foot ruler and kitchen tiles, what served as the criterion, and what validity concept does this everyday example illustrate?
📘 Lecture 18 — Construct Validity
📖 Overview: This lecture introduces construct validity as a third major approach to establishing test validity, distinct from content and criterion-related validity. It explains that construct validity is established through a process of accumulating evidence from multiple sources, using the example of a science concept acquisition test (SCAT). The lecture emphasizes convergent and discriminant validity as key components, culminating in the multitrait-multimethod approach for rigorous validation.
🗂️ Topics Covered
The lecture begins with a concrete example of developing the SCAT test and forming assumptions to test its validity. It then defines construct validity evidence through multiple scholarly descriptions and outlines its underlying assumptions, specifically convergent and discriminant validity. The steps in the construct validity process are listed, followed by an explanation of the multitrait-multimethod approach and how construct validity is computed using correlations and factor analysis.
📝 Lecture Summary
Construct Validity
The lecture opens with the example of the SCAT (Science Concept Acquisition Test), developed to measure how well children in their final school year applied science concepts learned over the previous 4-5 years. To validate this test, four assumptions were formulated:
- SCAT scores will have a high-positive correlation with scores on a test of similar nature and content.
- SCAT scores will have a high-positive correlation with scores on a test of reasoning ability.
- SCAT scores will have a moderate-positive correlation with marks on school science achievement test.
- SCAT scores will have either no or a negative correlation with scores on a test of rote memorization of meaningless material like non-sense syllables.
Four tests were used to test these assumptions: the Dallas Times Herald Test (scale test), Basic Reasoning Skills Test, General Science Test, and a list of non-sense syllables. A sample of school children took SCAT and the other tests, and correlations were computed. This procedure, where other tests are used systematically to provide evidence, represents a third approach to calculating validity called construct validity, distinguishing it from content and criterion-related validity.
Construct Validity Evidence:
Construct validity is measured by examining whether a test measures the particular construct it is supposed to measure. If it does, the test is considered valid; if it does not, it is not accepted as valid.
🔑 Definition — Construct validity evidence: "A process used to establish the meaning of a test through a series of studies. To evaluate evidence for construct validity, a researcher simultaneously defines some construct and develops the instrumentation to measure it. In the studies, observed correlations between the test and other measures provide evidence for the meaning of the test" (Kaplan & Saccuzzo, 2001, p.635).
🔑 Definition — Construct: "An informed scientific idea developed or constructed to describe or explain behavior." Constructs are unobservable, presupposed (underlying) traits that a test developer may invoke to describe test behavior or criterion performance (Cohen, 1999).
The lecture provides an all-encompassing description from Anastasi & Urbina (2007): "The construct validity of a test is the extent to which the test may be said to measure a theoretical construct or trait. Examples of such constructs are scholastic aptitude, mechanical comprehension, and verbal fluency, speed of walking, neuroticism, and anxiety. Each construct is developed to explain and organize observed response consistencies. It derives from established interrelationships among behavioral measures. Construct validation requires the gradual accumulation of information from a variety of sources."
The significant elements of this description are:
- A theoretical construct or trait
- Each construct is developed to explain and organize observed response consistencies
- It derives from established interrelationships among behavioral measures
- Requires the gradual accumulation of information from a variety of sources
Assumptions Underlying Construct Validity:
Two fundamental assumptions underlie the construct validity approach:
a. Test scores will be highly correlated with tests measuring the same or similar construct. This is convergent validity.
b. Test scores will have weak/low (or at times may be negative) correlation with tests meant to measure constructs different from the one measured by the main test. This is discriminant validity.
These assumptions lead to the calculation of two types of construct validity.
🔑 Definition — Convergent Validity Evidence: This aspect of validity shows the extent to which a test measures the same construct/trait/attribute as do other measures that have been designed and developed to measure the same construct/trait/attribute.
🔑 Definition — Discriminant Validity Evidence: Discriminant validity provides information that the test in question does not measure what other tests measure, or it measures something different from what other available tests measure.
Steps in Construct Validity Process:
- The test in question has to be there
- Clearly defining the construct to be examined
- Assumptions about the construct
- Assumptions about the similar constructs and the different constructs
- Identification of measures that will be used to test the constructs in question
- Test administration, may be in various steps
- Calculation of correlations between the test scores
- Analysis of correlations in order to reach final conclusions regarding construct validity
Multitrait-Multimethod Approach to Construct Validity:
One of the most commonly used models of construct validity is the convergent and discrimination model, the multitrait-multimethod approach proposed by Campbell and Fiske (1959). They proposed that while estimating validity of a test, we should look not only at the measures our test is related with, but also the ones it is not related with.
According to this model, the following relationship analysis can provide information regarding convergent and discriminant validity: a. Correlation between the same construct, using the same method b. Correlation between different constructs using the same method c. Correlation between different methods used for same construct d. Correlation between different constructs using different methods
The validation example of SCAT employed this approach.
Computation of Construct Validity:
Primarily, correlations are calculated for all sets of scores. Where a larger number of tests are involved, factor analysis is applied to the obtained statistics. When the multitrait-multimethod approach is employed, a multitrait-multimethod matrix is developed. This is based on a systematic experimental design (Campbell & Fiske, 1959) and is employed when two or more traits are tested by two or more methods.
⭐ Key Takeaways
Construct validity is a process of gradually accumulating evidence that a test measures the theoretical construct it claims to measure, using multiple sources of information. Its core assumptions are convergent validity (high correlations with tests measuring the same construct) and discriminant validity (low or negative correlations with tests measuring different constructs). The multitrait-multimethod approach is a rigorous model that systematically examines correlations between same and different constructs using same and different methods to establish both convergent and discriminant validity. Construct validity computation primarily uses correlations, with factor analysis applied when many tests are involved.
🧠 Quick Revision Questions
- What are the four assumptions used to validate the SCAT test in the lecture example?
- Define construct validity according to Anastasi & Urbina (2007).
- What is the difference between convergent validity and discriminant validity?
- List the eight steps in the construct validity process.
- What is the multitrait-multimethod approach and what four relationship analyses does it examine?
📘 Lecture 19 — How Much Valid Is “Valid”? Decision Theory
📖 Overview: This lecture examines the practical meaning of test validity—specifically, how large a validity coefficient must be to be useful. It introduces decision theory and the Standard Error of Estimate, and explains how base rates, hit rates, and Taylor-Russell tables help evaluate whether using a test improves selection decisions beyond chance.
🗂️ Topics Covered
The lecture begins by discussing the size of validity and its dependence on the test’s purpose, followed by the Standard Error of Estimate as a measure of prediction error. It then introduces decision theory, explaining base rates and hit rates, and the use of cut-off scores for dichotomous decisions. Finally, it covers Taylor-Russell tables as a method for evaluating the net gain in selection accuracy from using a test.
📝 Lecture Summary
How Much Valid Is “Valid”? — The Size of Validity
The magnitude of a validity coefficient is important but less critical than the significance level. According to Kaplan & Saccuzzo (2001, p.138), “there are no hard-and-fast rules about how large a validity coefficient must be to be meaningful. In practice, one rarely sees a validity coefficient larger than .60, and validity coefficients in the range of .30 to .40 are commonly considered high.” The validity coefficient must be statistically significant at the .01 or .05 level to rule out chance.
The acceptable size of validity depends on the test’s purpose. For admissions to academic programs, lower validity may be acceptable because the consequences of misprediction (e.g., admitting an unsuitable student) are less severe. However, for selecting fighter pilots, anesthetists, cardiac surgeons, or for forensic testimony (e.g., lie detectors), even the smallest doubt about validity is unacceptable.
💡 Why this matters: The same validity coefficient can be deemed acceptable or unacceptable depending on the stakes of the decision being made.
Standard Error of Estimate (SEₑₛₜ)
The Standard Error of Estimate (SEₑₛₜ) quantifies the margin of error in predicting an individual’s criterion score due to imperfect test validity. It is defined as “The error of estimate shows the margin of error to be expected in the individual’s predicted criterion score, as a result of the imperfect validity of the test” (Anastasi & Urbina, 2007, p.157).
🔑 Definition — Standard Error of Estimate (SEₑₛₜ): The estimated standard deviation of prediction errors when using a test to predict a criterion score.
📐 Formula: ( SE_{\text{est}} = SD_y \sqrt{1 - r_{xy}^2} ) → This means: the error in prediction equals the standard deviation of the criterion score multiplied by the square root of the proportion of variance not explained by the test.
📌 Example: If a test has a validity coefficient of ( r_{xy} = 0.60 ) and the criterion’s standard deviation is ( SD_y = 10 ), then ( SE_{\text{est}} = 10 \times \sqrt{1 - 0.36} = 10 \times \sqrt{0.64} = 10 \times 0.8 = 8 ). This means the predicted criterion score has a margin of error of ±8 points (at one SE level).
Decision Theory and Decision Analysis
Decision theory asks: Is it worthwhile to use a test for selection? This requires comparing outcomes with the test versus without the test. Two key concepts emerge:
🔑 Definition — Base Rate: “In decision analysis, the proportion of people expected to succeed on a criterion if they are chosen at random” (Kaplan & Saccuzzo, 2001, p.635). It is the success rate without using the test.
🔑 Definition — Hit Rate: “In test decision analysis, the proportion of cases in which a test accurately predicts success or failure” (Kaplan & Saccuzzo, 2001, p.636). It is the success rate when using the test.
Dichotomous Decisions and Cut-Off Scores
When a test is used for dichotomous decisions (select/reject), a cut-off score is set. Scores above it lead to selection, below to rejection. Four possible outcomes arise:
| Decision on test | Future success | Future failure |
|---|---|---|
| Selected | 1. Valid acceptance (right decision) | 3. False acceptance (wrong decision) |
| Rejected | 2. False rejection (wrong decision) | 4. Valid rejection (right decision) |
The test’s value depends on the proportion of right decisions (cells 1 + 4) versus wrong decisions (cells 2 + 3). The hit rate is the percentage of accurate predictions (cells 1 + 4). The base rate is the success rate among those who would be successful even without the test.
Taylor-Russell Tables
Taylor and Russell (1939) developed tables to evaluate the net gain in selection accuracy from using a test. They show the probability that a person selected based on the test will actually be successful, accounting for the base rate.
🔑 Definition — Selection Ratio: The proportion or percentage of applicants who must be accepted.
To use Taylor-Russell tables, three pieces of information are needed:
- The validity coefficient of the test
- The selection ratio (proportion of applicants to accept)
- The base rate (proportion of successful applicants without using the test)
Additionally, a clear definition of “success” must be established.
📌 Example (from table): With a base rate of .60 (60% of applicants would succeed without the test), if the test’s validity coefficient is .30 and the selection ratio is .10 (only 10% of applicants can be accepted), the expected proportion of successes using the test rises to .82 (82% of those selected will succeed). This represents a net gain from 60% to 82%.
💡 Why this matters: Taylor-Russell tables allow test users to determine whether a test provides enough improvement over chance to justify its cost and effort.
⭐ Key Takeaways
Validity coefficients as low as .30 to .40 are commonly considered high in practice, but the acceptable size depends entirely on the stakes of the decision—higher stakes require higher validity. The Standard Error of Estimate quantifies prediction error and is calculated using the validity coefficient and the criterion’s standard deviation. Decision theory compares base rates (success without the test) to hit rates (success with the test) to determine if the test improves selection accuracy. Taylor-Russell tables provide a practical method to evaluate the net gain from using a test, requiring the validity coefficient, selection ratio, and base rate. The cut-off score creates four possible decision outcomes—valid acceptance, false acceptance, false rejection, and valid rejection—and the test’s value lies in maximizing valid decisions.
🧠 Quick Revision Questions
- What range of validity coefficients is “commonly considered high” in psychological testing?
- What is the formula for the Standard Error of Estimate, and what does each symbol represent?
- Explain the difference between a base rate and a hit rate in decision theory.
- List the four possible outcomes when a test is used for dichotomous decisions with a cut-off score.
- What three pieces of information are needed to use Taylor-Russell tables?
📘 Lecture 20 — Threats to Validity and Related Issues
📖 Overview: This lecture examines the various factors that can threaten the validity of a psychological test. It details the variables affecting test validity, such as group nature and sample heterogeneity, and provides a framework for evaluating validity coefficients. The lecture concludes with a comprehensive review of the different types of validity evidence.
🗂️ Topics Covered
This lecture covers factors affecting test validity including the nature of the group, heterogeneity of the sample, preselection, and the form of the relationship between test and criterion. It then addresses issues for evaluating validity coefficients, such as changes in relationships, criterion meaning, sample size, range restriction, and differential prediction. The lecture ends with a formal review of content, criterion (predictive and concurrent), and construct validity evidence.
📝 Lecture Summary
Factors Affecting Test Validity:
There are several variables that can be a source of error in the validity of a test. One must be aware of these variables and their possible effects. The effect of these variables should be controlled in planning, designing, developing, and administering the test.
a) Nature Of The Group: Tests are not administered to identical or very similar groups every time. The group on which the test was administered may not be the same as other groups with whom the test is used. When using a test on the basis of its validity, we need information regarding the details of the sample on which it was validated. A test with a high validity coefficient obtained from one group may not be an efficient predictor with another sample. For example, a test of mathematical ability validated on a sample of people with mixed abilities might not yield similar results when administered to a group of expert accountants or a group of fine artists who haven't studied math since junior school. Test manuals should specify the details of the sample from whom validity was obtained and the population to which the test can be generalized.
b) Heterogeneity Of The Sample: The validity coefficient is a value of correlation. Statistically, the correlation between two sets of scores will be higher if the scores come from a wider range. Therefore, a correlation will be higher in heterogeneous groups as compared to homogeneous groups.
c) Preselection: Preselection can cause a problem when measuring criterion-related validity for a test used in job selection. If the test is validated on a sample of newly selected employees, the fact that they were selected (and are therefore a more capable, homogeneous group) can limit the range of scores and lower the validity coefficient.
d) The Form Of Relationship Between Test And Criterion: One basic assumption in computing the Pearson correlation coefficient is that the relationship between the test and the criterion is linear and uniform across the whole range of scores. If the relationship is not linear and scores are clustered at certain points, as seen in a scatter plot, it will affect the validity coefficient.
Evaluation of Validity Coefficient:
The joint committee of the American Psychological Association, American Educational Research Association, and the National Council on Measurement in Education (1990) described issues to consider when interpreting validity coefficients.
a) Look for changes in the cause of relationships. The cause of the relationship between the test and the criterion should remain the same in the future, but situations may change, altering this relationship.
b) What does the criterion mean? The criterion should be carefully chosen. A test with unknown or dubious validity should not be chosen as a criterion, and the criterion should relate specifically to the use of the test.
c) Review the subject population in the validity study. One should know the type of population on which the test was validated. The test might be used with groups not represented in the population from which the validity coefficient was obtained.
d) Be sure the sample size was adequate. Tests should not be validated on very small samples; the sample should be reasonably sized.
e) Never confuse the criterion with the predictor. At times, test users may confuse the criterion with the predictor, using the criterion first. For example, universities give admission to students who have not passed a graduate ability test but require them to clear it before results are declared. The test (predictor) is then following the criterion (performance), which has already taken place.
f) Check for restricted range on both predictor and criterion. The range of scores should not be restricted. A wider range of scores yields a better correlation/validity.
g) Review evidence for validity generalization. The test user must know if the test can be used with a different population from the one on which it was validated. If not, the user must reconsider the choice of the test.
h) Consider differential prediction. At times, criterion-related validity is not the right choice. If a good criterion is not available, using a weak or wrong criterion is not advisable. In such situations, construct-related evidence for validity is a better option.
Review of Validity and Its Types:
🔑 Definition — Content Validity Evidence: “The evidence that the content of a test represents the conceptual domain it is designed to cover” (Kaplan & Saccuzzo, 2001, p.635).
🔑 Definition — Criterion-Related Validity: A procedure where scores on a test are correlated with scores on a criterion. It is defined as “The evidence that a test score corresponds to an accurate measure of interest. The measure of interest is called the criterion” (Kaplan & Saccuzzo, 2001, p.635).
🔑 Definition — Construct Validity Evidence: A process used to establish the meaning of a test through a series of studies. The researcher simultaneously defines a construct and develops instrumentation to measure it. Observed correlations between the test and other measures provide evidence for the meaning of the test (Kaplan & Saccuzzo, 2001, p.635).
🔑 Definition — Predictive Validity Evidence: “The evidence that a test forecasts score on the criterion at some future time” (Kaplan & Saccuzzo, 2001, p.638). Predictive validity pertains to prediction over a time interval.
🔑 Definition — Concurrent Validity Evidence: “Evidence for criterion validity in which the test and the criterion are administered at the same point in time” (Kaplan & Saccuzzo, 2001, p.635).
Criterion-related validity can thus take the form of predictive or concurrent validity.
⭐ Key Takeaways
The validity of a test is not a fixed property but is influenced by multiple factors including the nature and heterogeneity of the sample, preselection effects, and the linearity of the relationship between test and criterion. When evaluating a validity coefficient, a student must check for changes in relationships, the appropriateness of the criterion, sample size, range restriction, and the possibility of confusing the predictor with the criterion. Finally, it is crucial to remember the three main types of validity: content, criterion (which includes predictive and concurrent), and construct, as each provides different evidence for what a test measures.
🧠 Quick Revision Questions
- How does the heterogeneity of a sample influence a test’s validity coefficient?
- What is the problem of "preselection," and how can it lower a validity coefficient?
- List three key issues one should review when evaluating a validity coefficient for a test.
- What is the fundamental difference between predictive and concurrent validity?
- According to the lecture, what type of validity evidence should be sought when a good criterion is not available?
📘 Lecture 21 — Item Analysis
📖 Overview: This lecture covers the critical post-administration process of evaluating individual test items through quantitative methods. It explains how item analysis helps test developers identify and refine items that effectively measure the intended construct, ensuring the final test has optimal difficulty and discrimination power.
🗂️ Topics Covered
The lecture covers the definition and purpose of item analysis, followed by detailed explanations of item difficulty index including its calculation, interpretation, and the effects of guessing. It then examines item discrimination index, explaining how to calculate and interpret it using upper and lower scoring groups, along with the shortcut procedure for group selection.
📝 Lecture Summary
Item Analysis
After test items are written and administered in a pilot study, item analysis is conducted to evaluate each item quantitatively. This process answers questions about which items are too easy or too difficult, whether items differentiate between high and low scorers, and if items are arranged by difficulty level. Items that are "too easy" (everyone gets correct) or "too difficult" (no one gets correct) do not differentiate between those who know and those who do not, making them poor items.
🔑 Definition — Item Analysis: "A set of methods used to evaluate test items. The most common techniques involve assessment of item difficulty and item discriminability." (Kaplan & Saccuzzo, 2001, p. 637)
Item Difficulty
Item difficulty assesses how difficult test items are by calculating the percentage or proportion of test takers who respond correctly. The purpose is to examine whether the difficulty level set during construction was appropriate, identify items that are "too easy" (everyone answers correctly), and identify items that are "too difficult" (no one answers correctly). Such items need to be removed, replaced, or revised.
🔑 Definition — Item Difficulty: "A form of item analysis used to assess how difficult items are. The most common index of difficulty is the percentage of test takers who respond with the correct choice" (Kaplan & Saccuzzo, 2001, p. 637).
📐 Formula: Item difficulty index (p) = Number of test takers who answered correctly / Total number of test takers. The value ranges from 0 to 1.00. A large p value (close to 1.00) indicates an easy item; a small p value (close to 0) indicates a difficult item. The common acceptable range is between .3 and .8.
📌 Example: If an item has a difficulty index of .60, this means 60% of test takers answered it correctly (60/100 = .60). An index of .50 means 50% answered correctly (50/100 = .50). An index of 0 means nobody answered correctly (bad item), and 1.00 means everybody answered correctly (bad item).
Average Index of Difficulty of A Test
The average difficulty index for the whole test is calculated by adding all individual item difficulty indices and dividing by the total number of items. For maximum discrimination among test takers' abilities, the optimal average difficulty is approximately .5, with individual items ranging from .3 to .8 (Cohen & Swerdlik, 1999).
📐 Formula: Average difficulty = (Sum of all item difficulty indices) / (Total number of items)
Taking Care of Guessing
For multiple-choice tests where guessing is possible, the optimal average difficulty is set higher than for free-response tests. The optimal average difficulty is calculated by taking the mid-point between chance success proportion and 1.00. Chance success proportion is the likelihood of guessing correctly (e.g., .25 for 4 options, .20 for 5 options, .33 for 3 options).
📐 Formula: Optimal average difficulty = (Chance success proportion + 1.00) / 2
📌 Example: For a test with 4 response options per item, chance success proportion = .25. Optimal item difficulty = (.25 + 1.00) / 2 = 1.25 / 2 = .625. For a test with 5 options, the average proportion correct should be .69 (Lord, 1959).
💡 Why this matters: Adjusting for guessing ensures that the test accurately measures knowledge rather than luck, especially important in high-stakes selection tests.
Item Difficulty Index In Different Types Of Tests
For achievement tests or intellectual ability tests, item difficulty is based on the proportion of people who gave the correct answer. However, for tests with no right or wrong answers (e.g., personality tests), an item-endorsement index may be used instead. This analysis identifies items that nobody replied to or that received the same response from everyone.
Item Discrimination
Item discrimination measures whether a test item differentiates between those who know (high scorers) and those who do not know (low scorers). A good test item should be answered correctly by more high scorers than low scorers. If an item is correctly answered by low scorers but not high scorers, something is wrong with the item.
Item Discrimination Index
The item-discrimination index (d) is a measure of the difference between the proportion of high scorers answering an item correctly and the proportion of low scorers answering the item correctly. The value of d ranges from -1 to +1. A value of +1 means all high scorers answered correctly and no low scorers did (ideal). A value of -1 means all high scorers failed and all low scorers passed (very poor item). A value of 0 means no differentiation between groups.
🔑 Definition — Item Discrimination Index: "A measure of the difference between the proportion of high scorers answering an item correctly and the proportion of low scorers answering the item correctly; the higher the value of d, the greater the number of high scorers answering the item correctly" (Cohen & Swerdlik, 1999).
📐 Formula: d = (Number of high scorers answering correctly – Number of low scorers answering correctly) / n, where n = number of high scorers (or low scorers). Alternatively, d = (U – L) / n.
📌 Example: For item 4 in the table, U = 9, L = 2, n = 10. d = (9 – 2) / 10 = 7/10 = .7. This indicates good discrimination. For item 2, U = 10, L = 0, d = (10 – 0) / 10 = 1.0 (ideal). For item 1, U = 10, L = 10, d = 0 (no discrimination). For item 3, U = 0, L = 10, d = -1 (very poor).
Selecting Upper and Lower Groups
A shortcut procedure (Aiken, 1994) divides examinees into three groups: an upper group consisting of the top 27% of examinees based on total test scores, a lower group of the bottom 27%, and a middle group of the remaining 46%. When the number of examinees is small, upper and lower 50% groups may be used. The item discrimination index is then calculated using only the upper and lower groups.
⭐ Key Takeaways
Item analysis is a crucial quantitative procedure conducted after pilot testing to evaluate individual test items. Item difficulty index (p) ranges from 0 to 1.00, with optimal items falling between .3 and .8; items that are too easy (p=1.00) or too difficult (p=0) should be removed. The average test difficulty for maximum discrimination should be around .5, adjusted upward for multiple-choice tests to account for guessing. Item discrimination index (d) ranges from -1 to +1, with positive values indicating good discrimination between high and low scorers, and values near 0 indicating poor items. The standard procedure uses the top 27% and bottom 27% of scorers to calculate discrimination.
🧠 Quick Revision Questions
- What is the purpose of item analysis in test development?
- How is item difficulty index (p) calculated, and what do values of 0, .50, and 1.00 indicate?
- For a multiple-choice test with 4 options per item, what is the optimal average difficulty after accounting for guessing?
- How is item discrimination index (d) calculated, and what does a value of -1 signify?
- What percentage of examinees are typically used for the upper and lower groups when calculating item discrimination?
📘 Lecture 22 — Item Analysis of Item Distracters
📖 Overview: This lecture extends basic item analysis to evaluate not just whole items, but the individual response options (distracters) within multiple-choice questions. Understanding distracter analysis is crucial for building tests that accurately measure knowledge by ensuring that only high-knowledge examinees can identify the correct answer.
🗂️ Topics Covered
The lecture explains why analyzing item distracters is important for test quality, then introduces the concept of "response valence" as a method for evaluating each alternative option. It presents data patterns that indicate good distracters, poor distracters, and weak items, along with recommendations for improving or removing problematic options.
📝 Lecture Summary
Item Analysis of Item Distracters
Item analysis typically uses the item discrimination index to determine how well an item differentiates between high scorers (those who know the material) and low scorers (those who do not). However, it is equally important to analyze the distracters — the incorrect response options in a multiple-choice item. A good item has distracters that appear plausible to examinees who do not know the correct answer, reducing the chance of guessing correctly. If a distracter is too weak (obviously wrong), test-takers will avoid it; if one option is clearly the best, even those who do not know the material may select it. Therefore, each alternative must be evaluated.
🔑 Definition — Item distracter: An incorrect response option in a multiple-choice test that is intended to appear plausible to examinees who do not know the correct answer.
🔑 Definition — Response valence: The number or percentage of examinees in the high-scoring and low-scoring groups who select each response option (including the keyed correct answer).
The lecture explains that while the 'd' value (item discrimination index) shows whether an item discriminates overall, distracter analysis examines how each non-keyed option is chosen. To determine the valence of each option, percentages of responses to the keyed answer and each distracter are calculated. High scorers should predominantly select the keyed answer, while low scorers should be more evenly distributed across the distracters. If a distracter is selected by many high scorers but few low scorers, it is problematic and needs revision or replacement. If all distracters attract roughly equal numbers of test-takers, they are effective.
📌 Example — Evaluating Distracter Quality: The lecture provides a data table comparing responses from high scorers (n=25) and low scorers (n=25). For a good item where option (a) is correct:
Option (a) ✓ | Option (b) | Option (c) | Option (d) High scorers: 18 | 3 | 3 | 3 Low scorers: 6 | 6 | 7 | 6
Here, the pattern is ideal: most high scorers select the correct answer, and low scorers distribute fairly evenly across all four options.
📌 Example — A Poor Distracter: In this item, distracter (c) attracts more high scorers (10) than low scorers (7), indicating a problem:
Option (a) ✓ | Option (b) | Option (c) | Option (d) High scorers: 13 | 2 | 10 | 2 Low scorers: 6 | 6 | 7 | 6
Distracter (c) is "too attractive" to knowledgeable students, likely because it appears partially correct. This distracter needs improvement or replacement.
📌 Example — A Bad Item: This item shows all distracters being selected by more high scorers than low scorers, and low scorers are not evenly distributed:
Option (a) ✓ | Option (b) | Option (c) | Option (d) High scorers: 4 | 5 | 8 | 8 Low scorers: 9 | 3 | 6 | 7
Here, only 4 of 25 high scorers chose the correct answer, while 9 low scorers did. The item is failing entirely and should be removed or completely rewritten.
The lecture concludes that weak items and poor distracters may be improved, but bad items are better removed from the test.
💡 Why this matters: Distracter analysis prevents the common error of assuming an item is good just because its overall discrimination index is acceptable. Even a discriminating item can have one or more distracters that artificially inflate scores for low-knowledge examinees, undermining test validity.
⭐ Key Takeaways
Distracter analysis is essential for evaluating multiple-choice items beyond the item discrimination index. A good distracter is selected by more low scorers than high scorers, and all distracters should attract roughly equal numbers of low scorers. If a distracter is chosen by many high scorers, it is likely confusing or partially correct and must be revised. If all distracters fail (more high scorers select them than the correct answer), the entire item is bad and should be removed. Effective distracter analysis improves test reliability by ensuring that only knowledgeable examinees consistently select the keyed answer.
🧠 Quick Revision Questions
- What is "response valence" in the context of distracter analysis?
- Why is it important that a distracter appear "plausible" to test-takers who do not know the correct answer?
- In a good item with 25 high scorers and 25 low scorers, what pattern of distracter selection indicates a problem with distracter (b)?
- What does it mean if all distracters are selected by more high scorers than low scorers?
- What should be done with a weak distracter versus a bad item according to the lecture?