PSY631 — Final Term Summary (Lectures 23–45)
📘 Lecture 23 — Item Analysis
📖 Overview: This lecture expands on item analysis by introducing advanced concepts beyond basic difficulty and discrimination indices. It covers Item Response Theory and item-characteristic curves for graphical item evaluation, explains cross-validation for assessing test validity on new samples, and discusses qualitative analysis methods for gathering test-taker feedback. The lecture concludes with a review of core item analysis concepts.
🗂️ Topics Covered
This lecture covers Item Response Theory (IRT) and its alternative names, the construction and interpretation of Item-Characteristic Curves including positive slope, negative slope, and inverted-U patterns, the process and factors affecting Cross Validation, the areas explored in Qualitative Analysis of tests, and a Review of Item Analysis covering item difficulty and item discrimination indices.
📝 Lecture Summary
Item Response Theory
Item response theory (IRT) is an approach that considers the probability of answering each individual item in a test correctly or incorrectly. The information regarding each item is plotted graphically. This approach is also known as Item-Characteristic Curve Theory and Latent Trait Theory. The graph containing information about the items is called the item-characteristic curve. Decisions regarding the items of a test can be based upon this information.
🔑 Definition — Item Response Theory (IRT): An approach that takes into consideration the probability of answering, right or wrong, each individual item in a test, with information plotted graphically.
Item- Characteristic Curves
Item difficulty and item discrimination can be presented graphically using item-characteristic curves. These are graphs where the horizontal axis (X axis) represents the ability being tested, and the vertical axis (Y axis) represents the probability of correct responses or the proportion of examiners responding correctly to the item. According to Kaplan and Saccuzzo (2001), it is “a graph prepared as part of the process of item analysis. One graph is prepared for each item and shows the total test score on the X axis and the proportion of test takers passing the item on Y axis” (p. 637). The shape or slope of the graph or curve indicates whether the item is a good one or not, and how far it discriminates high scorers from low scorers.
A good item has a positive slope. The proportion of high scorers responding correctly is higher than the proportion of low scorers. More of the low scorers are not responding correctly. A steep slope indicates that the test discriminates between the two groups. Scores of a highly discriminating test will yield a very steep slope.
A negative slope means that the item is not a good one. The item does discriminate between the two groups, high scorers and low scorers, but in the opposite way. This graph indicates that more of the low scorers did the item correctly than the high scorers. The probability of answering the item correctly is higher for the low scorers rather than the high scorers. This, therefore, is a bad or poor item that needs to be removed or replaced.
Another type of poor item is one where the majority of neither the top scorers nor the low scorers do the item correctly. It is the middle, moderate, scorers who attain the maximum proportion of correct responses. This type of item is also a bad item.
💡 Why this matters: Item-characteristic curves provide a visual and immediate way to assess whether an item is functioning as intended, distinguishing between knowledgeable and less knowledgeable test-takers.
Cross Validation
The validity of a test may be determined from a sample that was used for item selection. However, to have a better estimate of the validity of the test, the entire test needs to be validated on different samples as well. This process is called cross validation. According to Cohen and Swerdlik (1999, p.246), “The term cross validation refers to a revalidation of a test on a sample of test takers other than the ones on whom test performance was originally found to be a valid predictor of some criterion.” The regression equation is used to predict performance in a sample of test takers who are different from the ones on whom the test was validated.
If validity is computed from the original sample, the validity index may be higher than expected from a new sample due to possible chance variations. It is expected that validity will shrink in the process of cross validation. “The amount of decrease in the strength of the relationship from the original sample to the sample with which the equation is used is known as shrinkage” (Kaplan & Saccuzzo, 2001).
Factors that may affect the amount of shrinkage include:
- The size of the original item pool
- Proportion of test items retained
- Sample size
A high validity coefficient can be expected if the original item pool was large while the proportion of retained items is small. The size of cross validation sample also affects shrinkage. Greater validity shrinkage may be expected if smaller samples are used.
🔑 Definition — Cross Validation: The revalidation of a test on a sample of test takers other than the ones on whom test performance was originally found to be a valid predictor of some criterion.
🔑 Definition — Shrinkage: The amount of decrease in the strength of the relationship from the original sample to the sample with which the equation is used.
Qualitative Analysis
Qualitative analysis of a test may also be conducted along with quantitative analysis. After test administration is over, the test takers may be asked questions about various aspects of the test. These questions can be asked and answered orally or in writing. Different formats can be adopted, such as interviews and discussions. The respondents’ responses can be of great help in improving the individual test items, test format, and the entire test itself. Cohen and Swerdlik (1999) have pinpointed areas that may be explored, including:
- Cultural sensitivity
- Face validity
- Test administrator
- Test environment
- Test fairness
- Test language
- Test length
- Test taker’s guessing
- Test taker’s integrity
- Test taker’s mental/physical state upon entry
- Test taker’s mental/physical state during the test
- Test taker’s overall impressions
- Test taker’s preferences
- Test taker’s preparation
Review of Item Analysis
Methods used for assessing and evaluating characteristics of test items and the test itself. Primarily two characteristics are measured: item difficulty and discriminability.
Item-Difficulty Index: The item difficulty index is either in the form of percentages or proportions of the total number of test takers who attempted an item correctly. Item difficulty is calculated separately for every item. It is denoted by a lowercase italicized ‘p’. A number attached as subscript to this ‘p’ indicates the item number whose difficulty level is described. For example, p₁ indicates item difficulty of item number 1, p₂ is the difficulty level of item number 2, and so on. On occasions, the term ‘facility index’ may be used rather than difficulty index, referring to the percentage of responses to correct choices. Both terms refer to the same procedure.
Item Discrimination: A test is supposed to discriminate between those who know and those who do not know; those who score high and those who score low; those who have acquired a skill and those who have not. A test will not be a good test if the people who are supposed to know the correct answer fail and those who are not supposed to know succeed. A good test differentiates between the high and low scorers. If some items are correctly answered by high scorers and some by low scorers then something is wrong with the test.
Item Discrimination Index: Every test has its discrimination power. To see if the test discriminates between high and low achievers, a certain percentage of the high and low achievers is taken. The discrepancy between their attempts to correct responses is calculated in terms of percentages. The item discrimination index is denoted by a lowercase italicized letter.
⭐ Key Takeaways
You must understand that Item Response Theory (IRT) focuses on the probability of answering each item correctly and uses item-characteristic curves for graphical analysis. The slope of an item-characteristic curve is critical: a positive slope indicates a good discriminating item, a negative slope indicates a poor item where low scorers outperform high scorers, and a curve peaking in the middle is also poor. Cross-validation is essential to estimate a test's true validity on a new sample, and the expected decrease in validity is called shrinkage, which is influenced by the original item pool size, proportion of retained items, and sample size. Finally, qualitative analysis complements quantitative data by gathering test-taker feedback on areas like cultural sensitivity, test fairness, and their mental state, helping improve the test overall.
🧠 Quick Revision Questions
- What are the three alternative names for Item Response Theory mentioned in this lecture?
- In an item-characteristic curve, what does a negative slope indicate about an item?
- Define cross-validation and explain why the validity of a test is expected to shrink during this process.
- Name three factors that affect the amount of shrinkage in cross-validation.
- List five specific areas that can be explored during the qualitative analysis of a test.
📘 Lecture 24 — Assessment of Intellectual and Cognitive Abilities
📖 Overview: This lecture introduces the complex concept of intelligence and its measurement. It explores foundational questions about what intelligence is, why we measure it, and whether tests are reliable and valid. The lecture then provides a comprehensive historical and theoretical overview, from early pioneers like Galton and Binet to modern multi-factor theories by Spearman, Thurstone, Gardner, Sternberg, and Goleman, establishing the essential context for understanding intelligence testing.
🗂️ Topics Covered
The lecture begins by posing fundamental questions about the nature, origin, and measurement of intelligence. It then defines intelligence and reviews its key biological and environmental influences. The majority of the lecture is dedicated to a chronological and theoretical survey of major intelligence theories: Spearman's g-factor, Thorndike's three divisions (social, abstract, concrete), Thurstone's seven primary mental abilities, Cattell's crystallized and fluid intelligence, Guilford's Structure of Intellect model, Gardner's multiple intelligences, Sternberg's triarchic theory, the concept of emotional intelligence, and Piaget's stage theory of cognitive development.
📝 Lecture Summary
What is intelligence?
The lecture begins by stating that intelligence cannot be answered with a single word or sentence. It defines intelligence as "the capacity to understand the world, think rationally, and use resources effectively when faced with challenges" (Feldman, 2002). Intelligence refers to the ability to adapt, reason, solve problems, and think abstractly; it also includes learning from experience. The lecture emphasizes that intellectual ability is based on a constant interaction between environmental factors and inherited potentials, with modern psychology viewing both heredity and environment as influential.
Theories of intelligence
One of the earliest contributions was made by Sir Francis Galton, the "founder of differential psychology." In his work "Hereditary Genius" (1869), he proposed that gifted individuals tend to come from families with other gifted individuals. He attempted to measure intelligence quantitatively to determine the role of heredity and also investigated the relationship between intelligence and head size, though this found no empirical support. Cattell, an American psychologist, focused more on mental processes and first used the term “mental test” for devices measuring intelligence, developing tasks for reaction time, word association, and weight discrimination. Alfred Binet developed the first proper formal intelligence test in 1905, which will be discussed in detail in later sessions.
Spearman’s g-factor theory
Charles Spearman proposed one of the earliest theories, suggesting intelligence is a single, general factor. Using factor analysis, he observed that people scoring high on one mental test also tend to score high on others. He proposed two factors: the “g” factor (general intelligence) and the “s” factor (specific intelligence). The g-factor accounts for the general ability common in all people, while the s-factor accounts for specific abilities that differ among individuals.
🔑 Definition — g-factor: The general intelligence factor, a single, overarching mental ability believed to underlie performance on all cognitive tasks. 🔑 Definition — s-factor: Specific intelligence, referring to unique abilities that vary across different individuals and mental tests.
Thorndike’s Social Intelligence
Edward Thorndike criticized Spearman’s g-factor approach, arguing that intelligence consists of multiple factors expressed in human actions. He proposed three main divisions of intelligence:
- Social intelligence: Enables one to understand and manage relationships.
- Abstract intelligence: Enables one to understand and manage ideas such as algebra, mathematics, or abstract concepts.
- Concrete intelligence: Enables us to manage concrete and mechanical concepts and ideas, e.g., accounting, economics, and architecture.
Thurstone’s approach: Primary Mental Abilities
Louis L. Thurstone (1938) argued that intelligence is not a general factor but is composed of small, independent factors or elements called “primary mental abilities.” With his wife, he prepared a set of 56 tests, administered to 240 college students. Factor analysis yielded seven primary mental abilities:
- Verbal comprehension: Ability to understand and define words.
- Word fluency: Speed of thinking of verbal material, such as rhyming or naming words in a category.
- Spatial visualization: Ability to recognize and manipulate objects in three dimensions, e.g., drafting and blueprint reading.
- Perceptual speed: Quick ability to perceive visual details and differentiate similarities and differences between designs.
- Reasoning/inductive reasoning: A logical ability to derive general ideas from specific information.
- Numbers/arithmetic ability: Capability to work easily with numbers, such as performing simple arithmetic tasks quickly.
- Memory: Capacity to remember and retain material like words and letters, and the ability to recall and associate different words.
Crystallized and Fluid Intelligence: R.B Cattell
R.B. Cattell proposed that intelligence consists of two types:
- Crystallized intelligence: The capability of using information learned through experience. It is affected by education and culture and includes accumulated knowledge, skills, and techniques applied in problem-solving. This type increases with age.
- Fluid intelligence: Largely influenced by biological factors, it is the capability for information-processing and solving problems that depend on neurological development, such as reasoning and memory. This type declines with age.
Guilford’s theory of the Structure of Intellect (SOI)
J.P. Guilford created a model where intelligence is the result of the interaction of operations, contents, and products. He proposed 150 such abilities. The components are:
- Operations: Potential of different ways of thinking, including evaluation, convergent thinking, divergent thinking, memory retention, memory recording, and cognition.
- Contents: Potential of what we think about, including visual, symbolic, semantic, and behavioral content.
- Products: The results obtained by applying certain operations to certain contents, including units, classes, relations, systems, transformations, and implications.
Multiple Intelligences: Howard Gardner’s Approach
Howard Gardner (1985) maintained that intelligence does not consist of a single factor but consists of eight independent intelligences, possessed by all individuals in varying degrees:
- Linguistic
- Logical-mathematical
- Spatial intelligence
- Musical intelligence
- Bodily-kinesthetic
- Interpersonal intelligence
- Intrapersonal intelligence
- Naturalistic intelligence
Sternberg’s Triarchic Theory
Robert Sternberg’s triarchic theory (1980s) posits that intelligence consists of three main components:
- Analytic intelligence
- Creative intelligence
- Practical intelligence
Sternberg emphasized that practical intelligence is related to overall success in living, unlike traditional tests which measure academic success. Practical intelligence is learned through observation of others' behavior, while academic success comes from reading and listening.
💡 Why this matters: Sternberg’s theory highlights a critical gap: traditional IQ tests may not predict career or life success, which requires practical, street-smart intelligence.
Emotional Intelligence
Daniel Goleman (1995) proposed that emotional intelligence is the ability to go along with others. It involves the accurate assessment, evaluation, expression, and regulation of emotions. Emotional intelligence involves the realization and regulation of personal emotions as well as empathy and understanding of others' emotions. Key aspects include social skills, self-awareness, and empathy.
Piaget’s View of Intelligence
Jean Piaget defined intellectual development in terms of qualitative changes in thinking apparent in children of particular ages. His theory emphasized universal patterns of intellectual development and how children acquire knowledge. He proposed four universal and invariant stages of cognitive development:
- Sensorimotor
- Preoperational
- Concrete operational
- Formal operational
⭐ Key Takeaways
The most critical concepts from this lecture are the shift from viewing intelligence as a single, general ability (Spearman's g-factor) to complex, multi-faceted models (Thurstone, Guilford, Gardner, Sternberg). Students must understand the distinction between crystallized and fluid intelligence (Cattell), as this explains age-related changes in ability. Sternberg's introduction of practical intelligence and Goleman's emotional intelligence are crucial for understanding that traditional IQ tests do not capture all dimensions of human capability. Finally, the foundational work of Galton, Binet, and Cattell established the historical trajectory of intelligence testing, leading to these modern theoretical frameworks.
🧠 Quick Revision Questions
- What is the fundamental difference between Spearman's g-factor theory and Thurstone's theory of primary mental abilities?
- Describe the two types of intelligence proposed by R.B. Cattell and explain how they change across a person's lifespan.
- According to Sternberg's triarchic theory, what are the three components of intelligence, and why is "practical intelligence" considered important for career success?
- What are the eight independent intelligences proposed by Howard Gardner, and what is a key characteristic of how they are distributed among individuals?
- How does Daniel Goleman's concept of "emotional intelligence" differ from the traditional understanding of intellectual ability measured by standard IQ tests?
📘 Lecture 25 — Measurement of Intelligence
📖 Overview: This lecture provides a comprehensive historical and practical overview of intelligence testing, beginning with Alfred Binet’s foundational work in France. It traces the evolution of the Binet Scale through multiple revisions to the Stanford-Binet and then details the development and structure of the equally influential Wechsler Scales. Understanding this evolution is critical for appreciating how modern intelligence tests are constructed, standardized, and interpreted.
🗂️ Topics Covered
This lecture covers the historical evolution of intelligence testing starting with the Binet and Simon scale from 1905 through its major revisions in 1908, 1916 (Stanford-Binet), 1937, 1960, 1972, and 1986 (4th Edition). It explains the core concepts of Mental Age and Intelligence Quotient (IQ), including the shift from ratio IQ to deviation IQ. The lecture then transitions to the Wechsler Scales, detailing the three main tests (WAIS-III, WISC-III, WPPSI-R), their verbal and performance subtests, and psychometric properties. Finally, it provides a standard classification of IQ scores.
📝 Lecture Summary
The Binet Scale
The first formal measure of intelligence was developed in 1905 in France by Alfred Binet and Theodore Simon. Its primary purpose was to help the education ministry identify “dull” students in the Paris school system so they could receive remedial training. The core belief was that intelligence could be measured by a child’s performance on age-specific tasks. If a child could perform tasks typical for their age, they were considered of average intelligence. If they could only perform tasks for younger children, they were below average; if they could perform tasks for older children, they were above average.
The Development of Binet and Simon Scale
The test developers first created a number of tasks. These were then presented to groups of students who had been labeled as ‘dull’ or ‘bright’ by their teachers. The tasks that were successfully completed by the ‘bright’ students were retained, as these were considered indicative of intelligence. The key insight was that tasks were age-related; a child who could perform tasks meant for a higher age group was considered above average.
Binet’s scale gained rapid popularity, leading to translations and adaptations worldwide, particularly in the United States. A training school in New Jersey used it as early as 1908, and a modified version published in 1912 extended the age range down to three months. The most significant developments, however, took place at Stanford University.
The 1905 Scale
The original scale consisted of 30 items arranged in increasing order of difficulty. 🔑 Definition — Norms: Performance standards derived from a sample. The norms for this first scale were obtained from a very small sample of only 50 children who were reported to be ‘normal’ based on their average school performance. Although the normative and validity data were insufficient, this scale was a major milestone in psychological measurement.
The 1908 Scale
The 1908 revision was an age scale that retained the principle of age differentiation. Its most important innovation was the introduction of the concept of mental age. 🔑 Definition — Mental Age (MA): The average age of children who achieve a particular score on a test; it represents the typical intelligence level found for people at a given chronological age. A test taker’s performance was compared with the average performance of others of the same chronological age. The standardization sample for this revision included 203 individuals.
The 1916 Revision: The Stanford-Binet Intelligence Scale
In 1916, American psychologist Lewis Terman produced the first Stanford revision, creating the Stanford-Binet Intelligence Scale. This was standardized on an American sample for ages 3 to 14, plus “average and superior adults.” The sample, however, was criticized for containing only white, native Californian children.
This scale had three significant characteristics:
- It was the first American test to use the concept of Intelligence Quotient (IQ).
- It introduced the concept of an ‘alternate item’, an item that could replace a regular item if needed.
- It provided detailed, organized instructions for administration and scoring, a hallmark of standardized testing.
The 1937 Revision
This revision, by Terman and his colleague Merrill, contained new tasks for preschool and adult levels. A major improvement was the creation of two equivalent forms, ‘L’ and ‘M’ (the initials of the authors’ first names). The standardization sample was much larger, with 3,184 individuals from 11 U.S. states. However, it was still criticized for a lack of representativeness, as the subjects were all white and mostly from urban areas.
The Concept of Mental Age
Children taking the Binet-Simon test were assigned a score corresponding to their age group. This score was their “mental age” . 🔑 Definition — Mental Age (reiterated): The average age of children who secure the same score on a test. Mental age can differ from chronological age, reflecting whether a child is performing above, at, or below the level of their peers.
The Concept of Intelligence Quotient or IQ
To address problems with using mental age alone, the Intelligence Quotient was developed, which gives consideration to both mental and chronological age.
📐 Formula: IQ score = MA / CA x 100 → Plain-English meaning: This formula calculates the ratio of a person’s mental age (MA) to their chronological age (CA). Multiplying by 100 removes decimal points. 📌 Example: A 10-year-old child (CA = 10) who has a mental age of 12 (MA = 12) would have an IQ of (12/10) x 100 = 120. If their MA was exactly 10, their IQ would be 100. 💡 Why this matters: This formula provided a single, standardized number to represent intelligence, allowing for easy comparison. An IQ of 100 is the average, scores below 100 indicate below-average performance for one’s age, and scores above 100 indicate above-average performance.
The 1960 Revision
A major change in the 1960 edition was the creation of a single form, ‘L-M’ , containing the best items from the two previous forms. More importantly, it shifted from using ratio IQ to deviation IQ. 🔑 Definition — Deviation IQ: A standard score where a person’s performance is compared to the performance of others of the same age in the standardization sample. The mean is set to 100 and the standard deviation to a constant value (16 for this test). Using deviation IQ allowed for a precise estimation of a test taker’s relative position (e.g., percentile rank) within their age group, solving a key limitation of ratio IQ.
The 1972 Revision
This revision introduced an improved normative sample of 2,100 subjects, which finally included nonwhites. Despite this improvement, the sample was still criticized for not having enough non-white participants.
The 1986 Version: The Stanford Binet (4th Edition)
This version successfully overcame many previous criticisms. It used a large standardization sample of 5,000 subjects from 47 states, stratified by geographic region, community size, ethnic group, age, and gender. It covered four content areas:
- Verbal reasoning (Subtests: Vocabulary, Comprehension, Absurdities, Verbal relations)
- Abstract/visual reasoning (Subtests: Pattern analysis, Copying, Matrices, Paper folding and cutting)
- Quantitative reasoning (Subtests: Quantitative, Number series, Equation building)
- Short-term memory (Subtests: Bead memory, Memory of sentences, Memory of digits, Memory of objects)
Administration of Stanford-Binet Test
The test uses an individual-oral administration format. The examiner establishes a basal age (the lowest level where two consecutive items are passed) and a ceiling (the point where at least three out of four items are missed). 🔑 Definition — Basal age: The lowest level in the test where the examinee passes two consecutive items of approximately equal difficulty. 🔑 Definition — Ceiling: The point in the test at which the examinee fails at least three out of four items. Scores from all 15 subtests are converted into standard age scores (mean = 50, SD = 8), which are then grouped into the four area-content scores (mean = 100, SD = 16). A composite score is also calculated.
Some Sample Items from Early Versions of Simon-Binet Scale
The lecture provides examples of tasks for different ages, showing the increasing complexity:
- Three years: Shows nose, eyes, mouth; repeats two digits.
- Six years: Distinguishes morning from afternoon; copies a shape; counts 13 pennies.
- Eight years: Counts from 20 to 0; indicates omissions in pictures; repeats five digits.
- Fifteen years: Repeats seven digits; gives three rhymes; solves a problem from several facts.
The Wechsler Scales
The Wechsler Scales, developed by psychologist David Wechsler, are perhaps the most commonly used intelligence tests today. The three main scales are:
- WAIS-III (Wechsler Adult Intelligence Scale, 3rd ed.): for ages 16 to 89.
- WISC-III (Wechsler Intelligence Scale for Children, 3rd ed.): for ages 6 to 16.
- WPPSI-R (Wechsler Preschool and Primary Scale of Intelligence-Revised): for ages 3 to 7 years, 3 months.
The first Wechsler scale, the W-B I (Wechsler-Bellevue), was published in 1939. It was a point scale, giving credit for every correct response, unlike the Binet’s age scale. The WAIS was developed in 1955, revised as the WAIS-R in 1981, and the latest version, WAIS-III, in 1997.
The Wechsler scales are divided into two categories: Verbal and Performance scales.
- Verbal Scale subtests include: Vocabulary, Similarities, Arithmetic, Digit span, Information, Comprehension, and Letter-number sequencing.
- Performance Scale subtests include: Picture completion, Digit symbol-coding, Block design, Matrix reasoning, Picture arrangement, Symbol search, and Object assembly.
Psychometric Properties of WAIS-III
The standardization sample for the WAIS-III consisted of 2,450 adult subjects in 13 age groups (from 16-17 to 85-89). Stratification was based on gender, race, education, and geographic region, taken from the 1995 U.S. census.
The Meaning of IQ Test Scores
The lecture provides a standard classification of IQ scores:
- < 70: Retarded
- 85: Borderline
- 100: Average
- Above 115: Superior
- Above 140: Gifted
⭐ Key Takeaways
- The Binet Scale introduced the foundational concepts of Mental Age and Age Differentiation, but its ratio IQ was later superseded by deviation IQ in the Stanford-Binet, which allows for direct comparison of an individual’s performance to their age peers.
- The evolution of the Binet/Stanford-Binet Scale is a textbook example of test development, showing a progressive refinement in standardization, sample size, and representativeness (from 50 white children to 5,000 stratified subjects).
- The Wechsler Scales revolutionized the field by moving from an age scale to a point scale and introducing distinct Verbal and Performance subscales, providing a more detailed profile of cognitive strengths and weaknesses.
- Both major test families rely on individual administration by a trained examiner, establishing a basal age and a ceiling to ensure the test covers the appropriate difficulty range for each test taker.
- The meaning of IQ scores is only interpretable relative to a standardized distribution (mean = 100, SD = 16 for many tests), with labels like “Average,” “Superior,” and “Retarded” defined by specific score ranges from the normative sample.
🧠 Quick Revision Questions
- What was the primary purpose of Alfred Binet and Theodore Simon’s original 1905 intelligence scale?
- How does a ratio IQ differ from a deviation IQ in how it is calculated and what it tells us about a person?
- What is a basal age and a ceiling in the administration of the Stanford-Binet test, and why are they important?
- Name the three main Wechsler scales currently in use and the specific age range for which each is intended.
- According to the standard classification provided, what is the IQ score range that is considered “Superior”?
📘 Lecture 26 — The Kaufman Scales: Intelligence Tests
📖 Overview: This lecture examines several major intelligence test batteries developed by Alan and Nadeem Kaufman, including the K-ABC, K-BIT, and KAIT, alongside the Differential Ability Scales (DAS). It also addresses critical issues of cultural bias in intelligence testing and introduces alternative formulations of intelligence such as moral, social, and emotional intelligence that expand beyond traditional IQ measurement.
🗂️ Topics Covered
The lecture covers the Kaufman Assessment Battery for Children (K-ABC) with its five global scales and 16 subtests, the Kaufman Brief Intelligence Test (K-BIT) as a screening instrument, and the Kaufman Adolescent and Adult Intelligence Scale (KAIT) featuring crystallized and fluid scales. It then presents the Differential Ability Scales (DAS) with its core subtests, diagnostic subtests, and achievement tests across different age groups. The discussion shifts to cultural biases in intelligence testing with culture-fair tests like Raven's Progressive Matrices, followed by significant questions for IQ test use and alternative formulations including moral, social, and emotional intelligence.
📝 Lecture Summary
The Kaufman Scales: Intelligence Tests
The husband and wife duo, Alan and Nadeem Kaufman, have contributed significantly to intelligence testing by developing the K-ABC (1983), K-BIT (1990), and KAIT (1993). Each test serves distinct populations and purposes within intellectual assessment.
K-ABC: Kaufman Assessment Battery for Children
The Kaufman Assessment Battery for Children includes five global scales. Sequential processing involves subtests like hand movement, number recall, and word order. Simultaneous processing includes magic window, face recognition, Gestalt closure, triangles, matrix analogies, spatial memory, and photo series. Mental processing Composite combines sequential and simultaneous processing scales. The Achievement scale includes expressive vocabulary, faces and places, arithmetic, riddles, reading/decoding, and reading/understanding. The Nonverbal scale includes face recognition, hand movements, triangles, matrix analogies, spatial memory, and photo series.
💡 Why this matters: This battery focuses on the information processing approach, examining how individuals process information rather than just what they know.
The scales consist of 16 subtests other than the nonverbal scale. A national sample of 2000 American children aged 2.5 to 12.5 years was used for standardization. The sample considered age, gender, geographic region, parental education, community size, educational placement, and race. Gifted and talented, mentally retarded, and learning disabled individuals were included proportionally to their representation in the general public.
K-BIT: Kaufman Brief Intelligence Test
The Kaufman Brief Intelligence Test, meant for ages 4 to 90 years, is a quick screening instrument. It is an individual test for assessing intellectual functioning through individual administration. The K-BIT yields three scores: verbal, non-verbal, and composite. The verbal subtest includes 45 Expressive Vocabulary items and 37 Definitions. The non-verbal subtest contains 48 matrices. This test used nearly 20% of the KAIT standardization sample.
🔑 Definition — K-BIT: A brief, individually administered intelligence test providing verbal, non-verbal, and composite scores for ages 4 to 90 years
KAIT: Kaufman Adolescent and Adult Intelligence Scale
The KAIT was developed for measuring intelligence of subjects aged 11 to 85 plus. It consists of a Crystallized scale measuring concepts learned from schooling and acculturation, and a Fluid scale measuring the ability to solve new problems. Three subtests are used in each scale in the Core Battery. An Expanded Battery is available for subjects suspected of having neurological damage, where any of four specified subtests are added. A brief mental status test assessing attention and orientation is included for cognitively impaired test takers who cannot complete the whole battery.
Differential Ability Scales
The Differential Ability Scales (DAS) were developed by C. D. Elliot (1990) in Great Britain as a revised form of the British Ability Scales. There are three major components containing 20 subtests in all.
Core Subtests: Block Building, Verbal Comprehension, Picture Similarities, Naming Vocabulary, Early Number Concepts, Copying, Pattern Construction, Recall of Designs, Word Definitions, Matrices, Similarities, and Sequential and Quantitative Reasoning.
Diagnostic Subtests: Matching Letter-Like Forms, Recall of Digits, Recall of Objects, Recognition of Pictures, and Speed of Information Processing.
Achievement Test: Basic Number Skills, Spelling, and Word Reading.
Four core subtests are used with preschoolers aged 2 years 6 months to 3 years 5 months. Six core subtests are for preschoolers aged 3 years 6 months to 5 years 11 months. Six subtests are for school level, ages 6 years to 17 years 11 months.
The standardization sample included 3475 persons aged 2 years 6 months to 17 years 11 months, representing non-institutionalized English-proficient individuals in the U.S. Age, gender, race/ethnic origin, parental education, and geographic region were considered for stratification.
Cultural Biases And Intelligence Tests
Tests used to assess intelligence have been frequently criticized for being biased against particular groups. Culture-fair IQ tests are developed to overcome this problem, such as Raven's Progressive Matrices, which do not discriminate against any minority or cultural group.
Some Significant Questions Pertaining To the Use of IQ Tests
Key questions to consider include: Is the test a validity test? Is it reliable? Was it standardized? Is it being used with people similar to those in the standardization sample? Are the differences between the test taker and standardization sample drastic and serious? Have consequences been anticipated against expected benefits? Can cultural background affect results? Can ethnic origin affect test results? Can administration be problematic due to personal, environmental, or physical reasons?
Alternative Formulations
Alternative formulations of intelligence include Moral intelligence, Social intelligence, and Emotional intelligence.
Moral Intelligence
Given by Coles (1997) and Hass (1998), moral intelligence is the ability to differentiate between right and wrong. More comprehensively, it is the capacity to make right decisions that are beneficial not only for oneself but for others as well.
🔑 Definition — Moral Intelligence: The ability to differentiate right from wrong and make decisions beneficial to oneself and others
Social Intelligence
Given by Hough (2001) and Riggio, Murphy, & Pirozzolo (2002), social intelligence is manifested as SQ. It is the ability to understand and deal with people, exhibited by salesmen, politicians, teachers, clinicians, and religious leaders. It also involves understanding and dealing with oneself by identifying one's thoughts, feelings, attitudes, and behaviors.
🔑 Definition — Social Intelligence (SQ): The ability to understand and deal with others and oneself through identifying thoughts, feelings, attitudes, and behaviors
Emotional Intelligence (EI)
Emotional intelligence is a type of social intelligence involving the ability to cope with one's own and others' emotions, differentiate between them, and use information for guiding thoughts and actions. It is indicated by a person's EQ and includes these aspects:
- Self-awareness
- Managing emotions
- Empathy
- Handling relationships
🔑 Definition — Emotional Intelligence (EI): The ability to cope with one's own and others' emotions, differentiate between them, and use emotional information to guide thoughts and actions
⭐ Key Takeaways
Students must remember that the Kaufman scales (K-ABC, K-BIT, KAIT) target different age ranges with distinct purposes—K-ABC focuses on information processing in children, K-BIT provides quick screening across a wide age range, and KAIT measures crystallized versus fluid intelligence in adolescents and adults. The Differential Ability Scales (DAS) contain 20 subtests organized into core, diagnostic, and achievement components with age-specific administration patterns. Cultural bias remains a critical concern in intelligence testing, addressed through culture-fair tests like Raven's Progressive Matrices, and test users must evaluate validity, reliability, standardization, and cultural appropriateness before administration. Finally, intelligence extends beyond traditional IQ to include moral, social, and emotional intelligence, each with distinct definitions and applications in understanding human capability.
🧠 Quick Revision Questions
- What are the five global scales of the K-ABC, and what does the Mental Processing Composite combine?
- How does the KAIT distinguish between crystallized and fluid intelligence, and what is the purpose of the Expanded Battery?
- How many core subtests are used for preschoolers aged 2 years 6 months to 3 years 5 months on the DAS compared to school-aged children?
- What are three critical questions that should be asked when determining if an IQ test is appropriate for a particular test taker?
- What specific aspects comprise emotional intelligence (EI), and how does it differ from social intelligence (SQ)?
📘 Lecture 27 — Piagetian Approach: Measurement of Cognitive Development
📖 Overview: This lecture introduces the Piagetian approach to measuring cognitive development, which differs from traditional psychological testing by using structured observation rather than standardized tests. It covers Jean Piaget’s theory of cognitive development, his four invariant stages, and the specific tasks used to assess children’s cognitive abilities. Understanding this approach is crucial for appreciating how cognitive development can be measured through qualitative, age-appropriate methods.
🗂️ Topics Covered
The lecture begins by defining cognition and cognitive development, then introduces Jean Piaget, his background, and his clinical method of investigation. It presents Piaget’s four stages of cognitive development (sensorimotor, preoperational, concrete operational, formal operational) with their age ranges and characteristics. The lecture then details specific Piagetian tasks used to measure concept acquisition, including the A-B search task for object permanence and conservation tasks for mass, number, weight, and volume. Finally, it discusses significant influences on cognition, including socio-cultural factors and motivation.
📝 Lecture Summary
Cognition, Knowledge, and Cognitive Development
Cognition refers to the process of knowing as well as what is known. It includes knowledge that is innate or inborn, present in the form of brain structures and functions. We remember the physical environment in which we were raised and develop perceptual constructs accordingly (seeing, hearing, sounds, etc.). Cognition also refers to the mental processes people use to gather or acquire knowledge, and the knowledge that has been gathered is subsequently used in mental processes. Therefore, cognition and knowledge have a circular relationship.
Cognitive development involves the development of language, mental imagery, thinking, reasoning, problem solving, and memory development. It is the process whereby children’s understanding of the world develops as a function of age and experience.
🔑 Definition — Cognitive development: The process whereby the development of children’s understanding of the world as a function of age and experience takes place.
Jean Piaget’s Theory of Cognitive Development
Piaget (1896-1980) was a Swiss psychologist who became interested in epistemology (knowledge and knowing) as a result of his study of philosophy and logic. This interest laid the foundation of his theory of cognitive development. Unlike most psychologists who were impressed by Darwin’s theory of evolution, Piaget was influenced by Henri Bergson’s Creative Evolution, which believed in divine agency instead of chance as the force behind evolution.
After securing a position in Alfred Binet’s laboratory in Paris, Piaget observed children’s performance, their right and wrong answers. This generated an interest in children’s mental processes. The real shift took place when he started observing his own children from birth, keeping records of their behavior and tracing the origins of children’s thoughts to their behavior as babies. Later he became interested in the thought of adolescents as well.
These experiences resulted in two significant consequences: Piaget’s theory of cognitive development and the Piagetian method of study.
Piagetian Method of Investigation
Piaget’s method is known as the clinical approach, which is a form of structured observation. Piaget used to present problems or tasks to children of different ages, asked them to explain their answers, and their explanations were further probed through carefully phrased questions.
Piaget’s Stages of Cognitive Development
Cognitive development takes place in four stages in a set sequence. The sequence of stages is invariant, meaning they always occur in the same order. The age range of each stage is described but not invariant; age specification is arbitrary, and different children may perform the same task at different age levels. The organization of behavior is qualitatively different in different stages. Children throughout the world pass through these four stages in a fixed order.
The four stages are:
- Sensorimotor stage (Infancy: Birth-2 years)
- Preoperational stage (Preschool: 2-7 years)
- Concrete operational stage (Childhood: 7-11 years)
- Formal operational stage (Adolescence and adulthood: 11 years onward)
Sensorimotor Stage (Birth-2 years): The child’s thought is egocentric and confined to action schemes. Development is very rapid, but thought processes are limited to the immediate world of the child. Development of object permanence and development of motor skills takes place. The child has little or no capacity for symbolic representation.
Preoperational Stage (2-7 years): Development of representational thought takes place. The child’s thinking is intuitive, not logical. A significant aspect of development at this stage is the development of language and symbolic thinking. Thinking remains egocentric.
Concrete Operational Stage (7-11 years): The child’s thinking becomes systematic and logical, but only with regard to concrete objects. Development of conservation and mastery of the concept of reversibility takes place.
Formal Operational Stage (11 years onward): Abstract and logical thought develops at this stage. The person can deal with the abstract and the absent.
Some Piagetian Tasks
These tasks measure the acquisition of various concepts. The acquisition of concepts is progressive. Children of different age levels or stages of cognitive development show different levels of acquisition.
The ‘A-B’ Search Task: This task is meant for sensorimotor stage children and concerns the acquisition of the concept of object permanence. Two hiding places (A and B) are used, placed in front of the child (e.g., two place mats or napkins on a table). An object is hidden under either of the two, and the child has to look for it. Children’s responses vary according to their stage of development.
Responses by age:
- 4-8 month olds: These children will not search for the object even when it is hidden under A in front of them.
- 8-12 month olds: They will search for the object and find it under A. When the object is shifted from under A to under B (even when hidden in front of the child), the child will still look for it under “A”. This shows egocentric thinking, also known as the ‘A-not B’ error.
- 12-18 months: The child can accurately search for the object.
- 18-24 months: The child can not only search for the object but can also experiment with it and other similar objects.
🔑 Definition — Object permanence: The understanding that objects continue to exist even when they cannot be seen, heard, or touched.
Conservation Tasks: Conservation is a concept according to which some properties of an object or mass remain unchanged or invariant while others have been changed. For example, the weight of an object will remain the same when its shape has been changed; the number of objects remains unchanged while their arrangement is changed. Children learn conservation of mass and number earlier (around 5-6 years of age) than conservation of weight (around 8-9 years of age).
💡 Why this matters: Conservation tasks reveal whether a child has moved from preoperational to concrete operational thinking. They measure the child’s ability to understand that certain properties remain constant despite superficial changes in appearance.
Conservation of Mass: Using play dough, plasticine, or clay: take two same-sized balls of dough and ask the child if they have the same amount. After the child agrees, flatten one ball like a pancake and ask if they still have the same amount. Children at different cognitive levels will respond differently.
Conservation of Number: Take ten coins and arrange them in two parallel rows with equal spacing. Ask the child if both rows have the same number. After the child agrees, spread the buttons in one row distantly so that row appears longer. Ask if both rows contain the same number. Children who have not acquired conservation of number will say the longer row has more.
Conservation of Weight: Use two same-weight play dough or clay balls. Ask the child if they have the same weight. After agreement, change the shape of one ball into an oblong. Ask if they are still the same weight. Children who have not acquired conservation of weight will say the ball and oblong have different weights.
Conservation of Volume: Put equal amounts of water in two same-sized glasses. After the child agrees they contain the same amount, pour the water into two differently-sized containers (one with a lower water level, one with a higher level). Ask if both containers have the same amount of water. Children at different developmental levels will give different answers.
Perspective: The child is asked to imagine standing at the beginning of a long road with trees on both sides. The child is asked to draw and describe how the road and trees would look from that position.
Significant Influences on Cognition
Socio-Cultural Factor: This approach, debated in the early 1900s, has regained interest among cognitive scientists. It states that cognitive ability does not start only with the anatomy/biology of the individual or only with the environment. The culture and society into which the individual is born provide the most important resources for human cognitive development. They provide the context for the individual’s experience of the world. Social groups help in a person’s cognitive development by placing value on learning certain skills, thereby providing motivation. One perspective suggests there is innate potential within the individual, while another suggests there is potential within the socio-cultural context for development. Knowledge develops largely based on the evolution of intellect within society and culture.
Motivation, Cognition and Learning: Cognitive ability alone cannot account for achievement; motivation is also important in acquiring cognitive skills and abilities. People learn information that corresponds to their view of the world and learn skills that are meaningful to them. For example, children born in a poor family may not value formal education and may pass on similar beliefs to their offspring. Motivation determines whether one is capable of learning. The motivational condition largely depends on how the culture responds to achievements and failures. Culturally developed attitudes about the probability of learning successfully after initial failure can greatly affect future learning.
⭐ Key Takeaways
Students must remember that Piaget’s theory proposes four invariant stages of cognitive development (sensorimotor, preoperational, concrete operational, formal operational), each with qualitatively different thinking patterns. The Piagetian method uses clinical, structured observation with specific tasks like the A-B search task (measuring object permanence and revealing the A-not B error) and conservation tasks (mass, number, weight, volume) to determine a child’s cognitive stage. The concept of conservation—that properties remain invariant despite superficial changes—is a critical milestone achieved during the concrete operational stage. Finally, cognitive development is influenced not only by biological maturation but also by socio-cultural factors and motivation, which determine what skills are valued and learned.
🧠 Quick Revision Questions
- What are the four stages of Piaget’s theory of cognitive development, and at what approximate ages does each stage occur?
- What does the ‘A-not B’ error reveal about a child’s cognitive development during the sensorimotor stage?
- In the conservation of number task, why might a preoperational child say that a longer row has more coins than a shorter row with the same number?
- What is Piaget’s clinical approach, and how does it differ from traditional psychological testing?
- According to the lecture, how do socio-cultural factors and motivation influence cognitive development?
📘 Lecture 28 — Individual Tests of Ability for Specific Purposes
📖 Overview: This lecture explores specialized psychological tests designed for individuals with specific needs, such as learning disabilities, brain damage, and memory problems. It covers the rationale, administration, and interpretation of tests that measure learning disabilities, cognitive-achievement discrepancies, visual-motor skills, and creative thinking, highlighting their practical applications and psychometric limitations.
🗂️ Topics Covered
The lecture begins by defining learning disabilities and the discrepancy model used to identify them. It then examines the Illinois Test of Psycholinguistic Abilities (ITPA) and the Woodcock-Johnson Psycho-Educational Battery – Revised as tools for detecting learning disabilities. Following this, visiographic tests are introduced, including the Benton Visual Retention Test, Bender Visual Motor Gestalt Test, and Memory-for-Designs Test, which assess brain damage and perceptual-motor coordination. Finally, the Torrance Tests of Creative Thinking and the Wide Range Achievement Test-3 are discussed, focusing on creativity assessment and achievement measurement respectively.
📝 Lecture Summary
Learning Disabilities
Learning disabilities are a major concern for psychologists and educationists. In mainstream schools, children may face problems in routine education due to these disabilities. A child’s average achievement scores may be lower than the expected score for their age level because of such a disability. If a child of average cognitive development or IQ cannot perform what other average children of the same intelligence do, the child may be suspected to have a learning disability. The difference between IQ and achievement amounting to 1.5 to 2 standard deviations is considered indicative of learning disability.
💡 Why this matters: This discrepancy model is the foundational principle for identifying learning disabilities in educational and clinical settings.
Illinois Test of Psycholinguistic Abilities (ITPA)
The Illinois Test of Psycholinguistic Abilities (ITPA) is based on modern concepts of information processing. The test is based on the theory that the inability to respond correctly to stimuli does not result from defective output alone; the input has a role to play as well. Input refers to the information-processing system. Our response to an external stimulus takes place in three stages:
- Stage 1: Incoming information is received through senses.
- Stage 2: Analysis or processing of information is done.
- Stage 3: The response takes place.
The Illinois test provides an independent measure of all these three stages. Three subtests measure an individual’s ability to receive visual, auditory, or tactile inputs. Three further subtests provide independent measures of processing in these sensory modalities. There are other measures of motor and verbal output. The test is designed for children aged 2-10 years. ITPA is widely used among educators, psychologists, learning disability specialists, and researchers. However, its psychometric properties are widely criticized, providing no validity and reliability data. The test norms have been obtained from a middle-class population and it contains culturally loaded content, so it may not be appropriate for lower-class or minority groups.
The ITPA subtests include: Auditory Reception, Visual Reception, Auditory Association, Visual Association, Verbal Expression, Manual Expression, Grammatic Closure, Visual Closure, Auditory Sequential Memory, Visual Sequential Memory, Auditory Closure, and Sound Blending.
Woodcock-Johnson Psycho-Educational Battery – Revised
The Woodcock-Johnson Psycho-Educational Battery - Revised is a commonly used measure of children's achievement, measuring various aspects of scholastic ability. The test measures cognitive abilities, aptitudes, achievement, and interests. Learning problems can be identified by comparing subjects’ cognitive ability score with their achievement. If a difference of 1.5 to 2 SD is found between the cognitive ability and achievement of a child, it is considered a major discrepancy that is taken to be indicative of learning disability.
The tests of cognitive ability include: Picture vocabulary, Spatial relations, Memory for sentences, Concept formation, Analogies, and a variety of mathematical problems. The achievement tests include: Letter and word recognition, Reading comprehension, Proofing, Calculation, Science, and Social science and humanities. The tests of interest level cover math, language, and physical and social interests.
The scores can be described in terms of percentiles, which can further be converted into standard scores with a mean of 100 and SD equal to 15. For example, if a child is on the 50th percentile in cognitive score, with an achievement score that is 2 SD below the cognitive score, it can be taken as indicating learning disability. With normative data of more than 4700 individuals including members of different race, gender, urban, and rural status, these tests have good psychometric properties.
Visiographic Tests
Visiographic tests require a subject to copy various designs. These tests are useful for many kinds of brain damage.
Benton Visual Retention Test (BVRT)
The Benton Visual Retention Test (BVRT) is used for measuring brain damage and psychological deficit. It assumes that brain damage easily impairs visual memory ability. The test is designed for individuals aged 8 and older. Subjects have to reproduce geometric designs presented briefly before them and then removed. The subject loses points for mistakes and omission. As the number of errors increases, the subject approaches the organic (brain-damage) range.
🔑 Definition — Organic range: A score range on the BVRT indicating a high probability of brain damage, as inferred from an elevated number of errors in reproducing designs.
Bender Visual Motor Gestalt Test (BVMGT)
The Bender Visual Motor Gestalt Test (BVMGT) is one of the most popular individual tests. The test has nine geometric figures that the subject has to copy. The test is scored on the basis of errors. Norms are available for children aged 5-8 years. One or two errors by the age of 9 years are considered normal. However, if individuals over 9 make more errors, this may be an indication of some deficit. Individuals with more errors can be said to have a mental age less than 9 years (low intelligence), brain damage, or emotional problems. Though the test has a number of scoring systems, the reliability of the test is questioned.
📌 Example: A 12-year-old child is administered the BVMGT and makes 5 errors. According to the norms, individuals over 9 should make 1-2 errors. This high error count may indicate low intelligence, brain damage, or emotional problems.
Memory-for-Designs Test (MFD)
The Memory-for-Designs Test (MFD) is a short-time administered (10 minutes only) simple drawing test. The test measures perceptual motor coordination of individuals from 8 ½ to 60 years of age. The subjects are shown simple designs for a short duration and are then asked to reproduce them. The drawings are given scores from 0 to 3, depending on how close or similar they were to the original designs. The scoring of the 15 drawings can indicate brain injury and brain disease with the help of provided reference tables according to age and intelligence. The test has good psychometric properties with additional needs for validity.
Torrance Tests of Creative Thinking (TTTCT)
Creativity can be defined as “the ability to be original, to combine known facts in new ways, or to find new relationships between known facts”. The Torrance Tests of Creative Thinking (TTTCT) measure different aspects of creativity including fluency, originality, and flexibility.
- Fluency: This is the ability to generate a variety of solutions to a problem. People would score high on fluency if their solutions are distinct. The more distinct the solutions, the higher the score. The individual’s fluency is measured by their ability to provide as many solutions as one can find.
- Originality: This has to do with the uniqueness, novelty, and unusual nature of solutions. One can be said to be a creative person if one can come up with novel ideas or solutions. Unusual and unique solutions, which are different from the usual, conventional, and expected ones, add to the originality score of an individual.
- Flexibility: This is the ability to shift from one standpoint or strategy to another for finding solutions to problems. People can be said to have a flexible approach if they do not mind shifting positions in problem solving. If one strategy does not seem to be working, such people would try other strategies. Flexibility is measured by gauging a person’s ability to switch to different approaches of problem solution.
The test is useful for applied practitioners but needs more research for enhancing its psychometric properties.
🔑 Definition — Creativity: The ability to be original, to combine known facts in new ways, or to find new relationships between known facts.
Wide Range Achievement Test-3 (WRAT-3)
Intelligence tests measure what an individual may achieve, whereas what an individual has actually achieved is measured through achievement tests. IQ tests are about the potential while achievement tests are about the use of that potential. The scores on achievement may be indicative of people’s intelligence test scores, but not necessarily always. Factors such as interest, motivation, training, prior knowledge, or previous exposure may affect the way one has used one’s potential (intelligence). Therefore, both IQ and achievement tests are used for assessing one’s ability, depending on the situation and purpose.
Usually, achievement tests are used in groups, but some individual achievement tests are also available. The Wide Range Achievement Test-3 (WRAT-3) is the most widely used achievement test that measures the grade-level functioning in reading, spelling, and arithmetic. The test can be used for children aged 5 and older. The WRAT-3 is widely criticized for its grade-level reading ability.
⭐ Key Takeaways
The identification of learning disabilities relies on the discrepancy model, which compares IQ and achievement scores, with a difference of 1.5 to 2 standard deviations being a critical indicator. Several tests are designed for this purpose, such as the ITPA for information processing and the Woodcock-Johnson for a broader battery of cognitive and achievement abilities. Visiographic tests like the BVMGT, BVRT, and MFD are crucial for assessing brain damage and perceptual-motor deficits by having subjects copy or reproduce designs. Creativity is a distinct construct measured by the Torrance Tests of Creative Thinking, focusing on fluency, originality, and flexibility. Finally, it is essential to distinguish between potential (IQ tests) and actual achievement (achievement tests like the WRAT-3), as various factors can influence how potential is realized.
🧠 Quick Revision Questions
- What is the standard deviation difference between IQ and achievement scores that typically indicates a learning disability?
- Name the three stages of information processing measured by the Illinois Test of Psycholinguistic Abilities (ITPA).
- Which test uses nine geometric figures that a subject must copy to assess potential brain damage or emotional problems?
- What are the three main aspects of creativity measured by the Torrance Tests of Creative Thinking (TTTCT)?
- What three specific academic skills does the Wide Range Achievement Test-3 (WRAT-3) measure?
📘 Lecture 29 — Group Testing
📖 Overview: This lecture introduces group testing as an alternative to individual testing, detailing when and why it is used. It covers the core characteristics, advantages, and disadvantages of group-administered tests, and provides an overview of several prominent group intelligence test batteries. Understanding these distinctions is crucial for selecting the appropriate testing method based on the purpose of assessment.
🗂️ Topics Covered
This lecture contrasts group and individual testing, explaining their respective advantages and disadvantages. It then defines the key characteristics of group tests, such as their format, administration procedures, and scoring methods. Finally, it reviews several specific group intelligence tests, including the Kuhlmann-Anderson Test, Henmon-Nelson Test, and Cognitive Abilities Test, highlighting their unique features and applications.
📝 Lecture Summary
Group versus Individual Tests
Individual tests are administered one-on-one, allowing for rapport, flexible administration, and adaptation to the examinee's needs. They are ideal for diagnostic purposes. However, group tests are preferred when time and efficiency are critical, such as for screening large numbers of people or for quick data collection in research. The choice between individual and group testing ultimately depends on the specific purpose of the assessment.
🔑 Purpose of the Test: The deciding factor for choosing between individual and group test administration. The mode of administration is decided based on what the test will be used for.
Characteristics of Group Tests
Group tests are designed for simultaneous administration to large numbers of test takers. Instructions are given only once to the entire group and are not repeated. These tests are typically timed, are usually paper-pencil tests, and often use a multiple choice format. Answer sheets are marked using a key (like a stencil) or by a computer. Computerized administration is also becoming common.
🔑 Group Tests: Tests designed for simultaneous administration to many test takers, with instructions given only once to the entire group. They are usually timed, use a multiple-choice format, and are often paper-pencil tests.
Advantages of Group Tests
- Saves time of administration.
- Can examine a large number of people simultaneously.
- Excellent for quick data collection in research projects.
- Useful for making quick decisions like screening or school admissions.
- Easy administration for the examiner with minimal pressure.
- Less impact of the examiner's personality on examinee performance.
- The examiner's role is minimal, and tape-recorded instructions can be used.
Disadvantages of Group Tests
- Very impersonal, lacking the personal touch and rapport of individual testing.
- Not flexible; instructions cannot be repeated or rephrased.
- All subjects attempt all items, unlike individual testing which can use basal and ceiling rules.
- Examinee motivation cannot be judged, maintained, or enhanced.
- The examiner misses verbal and non-verbal cues about examinee anxiety or confusion.
- Most group tests require reading skills and paper-pencil manipulation skills. Examinees lacking these abilities may go undetected, leading to inaccurate scores.
- 💡 Why this matters: This last disadvantage is critical because test scores may reflect a person's reading ability rather than their true intellectual potential, leading to biased or invalid conclusions about the test taker.
Group Tests and Batteries
This section describes several commonly used group-administered intelligence tests.
Kuhlmann-Anderson Test- 8th Edition (KAT)
The KAT is a group intelligence test for kindergarten to 12th grade children. It is primarily a non-verbal test at all grade levels. It has eight levels and is popular as a test of mental ability with strong psychometric properties. Its norms are based on a sample of over 10,000 people, with high reliability and validity coefficients.
Henmon-Nelson Test (H-NT)
The H-NT is a test of mental ability with both grade-wise and age-wise norms. This 90-item test takes about 30 minutes and provides a single score relating to Spearman's g factor. It has high reliability (in the 90s) and validity coefficients (in the 50s to 90s).
Cognitive Abilities Test (COGAT)
The COGAT is a carefully designed test that yields three scores: verbal, nonverbal, and quantitative. It was developed to be less culturally biased, especially for poor readers and those with English as a second language. Items that predicted differentially for white and minority students were removed. It is considered a better tool for assessing culturally diverse, minority, and economically disadvantaged children. It has good psychometric properties with reliability in the 90s.
- Positive Points: Better for culturally diverse and disadvantaged children; can measure verbal underachievement; a sensitive discriminator for giftedness; good predictor of future performance.
🔑 Spearman's g factor: A general intelligence factor that underlies all mental abilities. The Henmon-Nelson Test provides a single score based on this concept.
⭐ Key Takeaways
The choice between individual and group testing hinges entirely on the test's purpose, with group tests being ideal for efficiency and large-scale screening, while individual tests are superior for diagnostic depth and rapport. Group tests are characterized by their impersonal, timed, and often multiple-choice format, and they require examinees to have basic reading and response-recording skills. While offering significant administrative advantages, group tests also present disadvantages, such as an inability to gauge individual motivation or detect specific learning barriers. Memory of this lecture should include the specific characteristics and applications of the three reviewed group intelligence tests: the Kuhlmann-Anderson Test (KAT), the Henmon-Nelson Test (H-NT), and the Cognitive Abilities Test (COGAT), noting that COGAT is especially valuable for culturally diverse populations.
🧠 Quick Revision Questions
- What is the single most important factor that determines whether an individual or group test should be used?
- Name two key characteristics of how instructions are delivered in a group test.
- List three specific disadvantages of group testing compared to individual testing.
- Which of the three group tests discussed (KAT, H-NT, COGAT) is considered the best tool for assessing culturally diverse children, and why?
- What is the primary cognitive ability measured by the Henmon-Nelson Test (H-NT)?
📘 Lecture 30 — Specific Purposes Tests
📖 Overview: This lecture explores a comprehensive range of psychological tests designed for specific purposes, including college entrance exams, graduate school admissions, nonverbal ability testing, and career selection instruments. Understanding these tests is crucial for appreciating how psychometric principles are applied in real-world educational, military, and occupational settings.
🗂️ Topics Covered
The lecture covers the Scholastic Assessment Test (SAT-I) and its components, Graduate Record Examination (GRE) and Miller Analogies Test for graduate admissions, nonverbal group ability tests including Raven Progressive Matrices, Goodenough-Harris Drawing Test, and IPAT Culture Fair Intelligence Test, industrial tests like the Wonderlic Personnel Test, occupational aptitude batteries such as GATB and ASVAB, and career interest inventories including the Strong Vocational Interest Blank and Kuder Occupational Interest Survey.
📝 Lecture Summary
The Scholastic Assessment Test (SAT-I)
The Scholastic Assessment Test (SAT-I), formerly known as the Scholastic Aptitude Test, was first used in 1926 and remains the most commonly used college entrance test in the U.S. The SAT-I contains two main parts: Verbal Reasoning and Mathematical Reasoning, each comprising further subtests. The Verbal Reasoning section has 78 questions to be completed in 75 minutes, distributed as 19 Sentence Completion questions, 40 Critical Reading questions, and 19 Analogies questions. The Mathematical Reasoning section contains 60 questions, including 35 Regular Mathematics multiple choice questions, 10 Student-Produced Responses, and 15 Quantitative Comparisons. Test norms were obtained from a large representative sample. SAT-II is also available and includes Subject Tests such as a direct writing test, tests in Asian languages, and an English-as-a-Second Language Proficiency Test.
Graduate and Professional School Entrance Tests
Graduate school entrance tests are widely used for admission to graduate school and professional degree programs in medicine, art, and law. The Graduate Record Examination Aptitude Test (GRE) is the most commonly used graduate-school entrance test, measuring general scholastic ability along with grade point average and letters of recommendation. The GRE has three sections: verbal (GRE-V) measuring reasoning, antonyms, analogies, and paragraph comprehension; quantitative (GRE-Q) measuring reasoning, algebra, and geometry; and analytic (GRE-A). The GRE also measures general achievement in at least 20 majors including psychology, history, and chemistry. Though psychometric properties are not very impressive, it is used as a relatively strong instrument. Studies on the relationship between Grade Point Average (GPA) and GRE show correlations from .22 to .33, and the GRE combined with GPA has proven to be a good predictor of graduate success.
The Miller Analogies Test (MAT) is the second major widely used scholastic aptitude test. It is a 50-minute verbal test measuring a student's ability to find logical relationships for 100 different analogy problems. The MAT offers special norms for various fields. Research indicates the MAT has an age bias, as its scores over-predicted GPAs for the 25-34 years age group and under-predicted for the 35-44 years age group. While psychometric properties are adequate, the MAT does not predict research ability, creativity, or other important factors in graduate school.
Nonverbal Group Ability Tests
Nonverbal tests evaluate individuals without the use of language. Individuals are typically asked to perform tasks such as drawing, solving mazes, or identifying problem figures from sets of presented figures. The Raven Progressive Matrices (RPM) is a non-verbal multiple-choice measure of general intelligence. In each test item, the subject identifies the missing element that completes a pattern. The test can be administered to groups or individuals from age 5 to older adults, containing 60 matrices with missing parts presented in graded difficulty. The subject selects the appropriate pattern from eight options. Research shows RPM measures general intelligence along with the capacity to think clearly and make sense of complex data. Despite criticism over psychometric properties, RPM has widespread use for children, language-handicapped individuals, and the culturally deprived. The updated manual provides comparison of performance of children from major cities worldwide, having minimized effects of language and culture.
The Goodenough-Harris Drawing Test (G-HDT) is the quickest, easiest, and least expensive nonverbal test for measuring intelligence. The subject draws a whole human figure, and scoring is based on each item included in the drawing, with a maximum of 70 possible points. G-HDT scoring follows the age differentiation principle, where older children tend to get more points due to greater accuracy. The test has good psychometric properties, and scores can be related to Wechsler IQ scores. The test is most appropriately used in combination with other intelligence tests.
The IPAT Culture Fair Intelligence Test is a paper-pencil test designed to remove cultural influences in intelligence and learning. It has three levels: for ages 4-8 and mentally disabled adults, ages 8-12 and randomly selected adults, and high-school age and above-average adults.
Group Tests for Specific Purposes
Along with large group tests for intelligence and academic aptitudes, several tests are available for specific populations. For example, the Black Intelligence Test of Cultural Homogeneity (BITCH) is used as a culture-fair intelligence test for African Americans.
The Wonderlic Personnel Test (WPT) helps in decisions concerning employment, placement, and promotion. Based on the population Otis Self-Administering Tests of Mental Ability, the WPT is a quick 12-minute test of mental ability in adults. While lacking in validity documentation, it is widely used for employee-related decisions in industry.
Tests for Assessing Occupational Aptitude
The General Aptitude Test Battery (GATB) is a widely used ability test measuring aptitude for various occupations. Developed by the U.S. Employment Service for employment decisions in government agencies, the GATB provides scores for motor coordination, perception, and clerical perception, along with verbal, numerical, and spatial aptitudes. The test is criticized for its normative data. Other tests for mechanical ability and clerical competence include the Differential Aptitude Test, the Bennett Mechanical Comprehension Test, and the Revised Minnesota Paper Form Board Test.
The Armed Services Vocational Aptitude Battery (ASVAB) was designed for postsecondary school students and 11th-12th grade students, originally for use by the Defense Department. Test scores can be used in both educational and military settings. The ten subtests include: general science, arithmetic reasoning, word knowledge, paragraph comprehension, numeral operations, coding speed, auto and shop information, mathematics knowledge, mechanical comprehension, and electronics information. These ten subtests are grouped into composites: academic composites (academic ability, verbal, and math); four occupation composites (mechanical and crafts, business and clerical, electronic and electrical, and health and social); and overall general ability. The ASVAB has very good psychometric properties, and recently the military has started using this test adaptively through a new computerized format rather than the traditional paper-based test.
Tests for Choosing Careers
Numerous psychological tests help individuals choose the right career based on their interests and aptitude. The first step of career selection is evaluation of interest. The Carnegie Interest Inventory, introduced in 1921, was the first interest inventory providing measures for 15 different interests. Today, more than 80 interest inventories are available.
The Strong Vocational Interest Blank (SVIB), developed in 1927, is the most widely used interest test. Using the criterion-group approach, the subject's interest was matched with criterion groups of people who were happy in their selected careers. The 399 items relate to 54 occupations for men and 32 occupations for women, presented separately. Items were weighted according to how frequently an interest occurred in a particular occupation group. The test provides strong psychometric characteristics, with the most interesting finding being that patterns of interest remain relatively stable over time. Studies show interest patterns are usually established by age 17. Despite widespread use, the SVIB is criticized for gender bias due to different scales for men and women and lack of theoretical information.
The Kuder Occupational Interest Survey (KOIS) is the second most popular interest test. The test-taker selects the most preferred and least preferred activity among 100 triads of alternative activities. KOIS assesses similarity between the test-taker's interests and those employed in various occupations. Separate norms are available for men and women, along with separate scales for college majors, helping students choose majors. A series of new scales has been added for nontraditional occupations. Despite lack of research data, KOIS is useful for guidance decisions for high-school and college students. 💡 Why this matters: Understanding the distinctions between interest inventories like SVIB and KOIS helps professionals select appropriate tools for career counseling and vocational guidance.
⭐ Key Takeaways
For the exam, students must remember that the SAT-I contains verbal and mathematical reasoning sections with specific item distributions and time limits, while the GRE has three sections measuring verbal, quantitative, and analytic abilities with modest correlations to graduate GPA. Raven Progressive Matrices and Goodenough-Harris Drawing Test represent key nonverbal intelligence measures with specific applications for culturally diverse populations and children respectively. The ASVAB stands out as a comprehensive vocational aptitude battery with excellent psychometric properties now administered in computerized adaptive format. Finally, the SVIB and KOIS represent the two dominant career interest inventories, with SVIB using criterion-group methodology and showing stable interest patterns established by age 17, though criticized for gender bias.
🧠 Quick Revision Questions
- What are the three subtest categories within the SAT-I Verbal Reasoning section, and how many questions does each contain?
- What is the reported correlation range between GRE scores and Grade Point Average (GPA)?
- How does the Goodenough-Harris Drawing Test apply the age differentiation principle in its scoring system?
- Name the ten subtests of the ASVAB and describe how they are grouped into composites.
- What is the criterion-group approach used in developing the Strong Vocational Interest Blank, and why is this test criticized for gender bias?
📘 Lecture 31 — Scales for Infants: Tests for Special Populations
📖 Overview: This lecture covers specialized intelligence and developmental tests designed for infants and special populations, including the mentally retarded. It examines their purposes, scoring methods, strengths, and limitations, and then discusses critical issues surrounding intelligence testing such as nature versus nurture, gender, and personality.
🗂️ Topics Covered
This lecture presents four major infant assessment scales: the Brazelton Neonatal Assessment Scale, Gesell Developmental Schedules, Bayley Scales of Infant Development, and the Cattell Infant Intelligence Scale. It then explains the Vineland Adaptive Behavior Scales for assessing mentally retarded individuals and concludes with a detailed examination of issues in intelligence testing including nature vs. nurture, measurement process, personality, gender, and family environment.
📝 Lecture Summary
Brazelton Neonatal Assessment Scale (BNAS)
The Brazelton Neonatal Assessment Scale (BNAS) is used for infants aged 3 days to 4 weeks. Its aim is to measure a newborn’s competence. The scale has a total of 47 scores (20 elicited responses and 27 behavioral items) and covers social, behavioral, and neurological functioning. Examples of factors assessed include reflexes, motor maturity, ability to habituate to sensory stimuli, startle reactions, cuddliness, responses to stress, and hand-mouth coordination. The BNAS is considered a good assessment and research tool; however, its test-retest reliability is not satisfactory. Also, if prediction of future intelligence is required, this scale cannot be of help.
Gesell Developmental Schedules (GDS)
The Gesell Developmental Schedules target infants and children aged 2.5 to 6 years. The aim is to measure infants’ and children’s intelligence, specifically to measure the subject’s developmental status. The final score is a Developmental Quotient (DQ) that can be used to calculate IQ; the formula is the same as the one for IQ, where the value of MA is replaced by DQ. The areas assessed include gross motor, fine motor, adaptive, language, and personal-social skills. Other names for this tool include Gesell Maturity Scale, the Gesell Norms of Development, and The Yale Tests of Child Development. Originally developed in 1925, this tool has been used popularly; however, it entails certain shortcomings. For example, it cannot be used as a predictor of future intelligence, and it is criticized for poor psychometric properties, an inadequate sample, and weak standardization.
🔑 Definition — Developmental Quotient (DQ): A score derived from the Gesell Developmental Schedules that indicates a child's developmental status, calculated using the same formula as IQ but replacing mental age with developmental age.
📐 Formula: DQ = (Developmental Age / Chronological Age) × 100 → This expresses a child's developmental maturity relative to their actual age.
📌 Example: If a 4-year-old child achieves a developmental age of 3 years, then DQ = (36 months / 48 months) × 100 = 75, indicating developmental delay.
Bayley Scales of Infant Development: 2nd Edition (BSID-II)
The Bayley Scales target infants aged 2 to 30 months. The aim is to measure infants’ cognitive and motor functions. It yields two main scores: mental and motor. The areas assessed include gross motor, fine motor, adaptive, language, and personal-social skills. The Bayley scale is valued and appreciated for adequate and good standardization. Although this scale also cannot predict future intelligence, it is considered a useful tool that can predict well for children who are mentally retarded.
Cattell Infant Intelligence Scale (CIIS)
The Cattell Infant Intelligence Scale is used for infants aged 2 to 30 months. Its aim is to measure infants’ intelligence, covering areas such as gross motor, fine motor, adaptive, language, and personal-social skills. The CIIS is an age scale that uses the concept of mental age and IQ and is designed after the Binet Scale. It is referred to as a downward extension of Binet’s Scale.
Assessment of the Mentally Retarded: Vineland Adaptive Behavior Scales (VABS)
The Vineland Adaptive Behavior Scales (VABS) are the latest version of the Vineland Social Maturity Scale developed by Edger Doll in the 1930s. Doll developed it after observing differences among mentally retarded patients, creating a standardized record form for assessing developmental level. The level was determined by considering looking after their practical needs and taking responsibility in daily living. The VABS is available in three versions and covers the following domains and subdomains:
- Communication: Receptive, Expressive, Written
- Daily Living Skills: Personal, Domestic, Community
- Socialization: Interpersonal relationships, Play and leisure time, Coping skills
- Motor Skills: Gross, Fine
- Adaptive Behavior Composite
- Maladaptive behavior
Issues of Intelligence Testing
No matter how we understand intelligence, several critical issues must be kept in mind: Is intelligence innate or acquired? Is intelligence a stable phenomenon? Can gender, race, culture, or region of subjects affect their scores? Standardized intelligence testing has been called one of psychology's greatest successes but has also been met with issues including race, gender, class and culture; minimizing the importance of creativity, character, and practical know-how; and propagating the idea that people are born with an unchangeable intellectual potential.
Nature versus Nurture: Research has shown the importance of heredity and environment. Twin studies show that identical twins reared apart have similar intelligence test scores. Children born to poverty-stricken parents but adopted into better-educated, middle-class families tend to have higher intelligence test scores. Natural mothers with higher IQs tend to have children with high intelligence regardless of upbringing. Proponents of the nurture view emphasize prenatal and postnatal environment, socioeconomic status, educational opportunities, and parental modeling. The interactionist position propagates that intelligence is the result of interaction between heredity and environment.
The Measurement Process: Issues include the instrument used, standardization sample, test administrator, and accuracy of test scoring. The test-taking attitude, prior coaching of examinee, or test administrator’s training are factors that may affect scores.
Personality: Research shows that several personality and intelligence tests overlap. Wechsler (1958) believed that tests of intelligence measure traits of temperament and personality such as drive, energy level, impulsiveness, persistence, and goal awareness. Higher intelligence scores are associated with high need of achievement, competitive striving, self-confidence, and emotional stability. Lower intelligence is expected among individuals with passivity, dependence, and maladjustment.
Gender: Extensive research has found that gender differences in intelligence are the result of psychosocial and physiological factors.
Family Environment: The family environment includes aspects of both nature and nurture. Twin studies show the importance of heredity and family environment. Issues include parental use of language, parental stress for achievement, access to resources, exposure to the world, and parental influences over discipline and policies. Studies on maternal age and social class have shown that aged mothers tend to have children with higher IQ.
💡 Why this matters: Understanding these issues is crucial for interpreting intelligence test scores accurately and ethically, avoiding misinterpretations based on bias.
⭐ Key Takeaways
The four main infant scales (BNAS, GDS, BSID-II, CIIS) each have specific age ranges, scoring methods, and limitations — most notably, they generally cannot predict future intelligence, with the exception of the Bayley scale predicting outcomes for mentally retarded children. The Vineland Adaptive Behavior Scales provide a comprehensive, multi-domain assessment of adaptive functioning in mentally retarded individuals, measuring practical needs and daily living responsibility. Intelligence testing is affected by numerous issues including the nature vs. nurture debate, measurement process factors, personality traits, gender (psychosocial/physiological factors), and family environment (including maternal age and social class), all of which highlight that intelligence is not a simple, purely innate trait.
🧠 Quick Revision Questions
- What is the age range and primary purpose of the Brazelton Neonatal Assessment Scale, and what is its main psychometric limitation?
- How is the Developmental Quotient (DQ) of the Gesell Developmental Schedules calculated, and what are two major criticisms of this tool?
- Which infant scale is referred to as a "downward extension of the Binet Scale," and what is its age range?
- What are the five domains covered by the Vineland Adaptive Behavior Scales, and who originally developed this tool?
- According to the lecture, what three personality characteristics are associated with higher intelligence scores, and what two types of factors explain gender differences in intelligence?
📘 Lecture 32 — Personality Testing
📖 Overview: This lecture introduces personality testing within psychological assessment, covering the definition of personality and three primary assessment methods: interviews, behavioral observation, and psychological tests. It explains the two main types of psychological tests (objective and projective) and explores how theoretical approaches (psychodynamic, trait, and social-cognitive) shape test content and measurement strategies.
🗂️ Topics Covered
The lecture begins by defining personality from multiple perspectives and outlines three methods of assessment: interview, behavioral observation, and psychological tests. It then distinguishes between objective tests/personality inventories (like the MMPI) and projective tests (like the Rorschach and TAT). The second half examines how theoretical orientation determines test content, focusing on the psychodynamic approach, trait approaches (including Cattell's 16 Personality Factors and Eysenck's dimensions), and the social cognitive approach to personality.
📝 Lecture Summary
Personality Testing
Before discussing assessment approaches, we need to understand what personality is. Personality has been defined in many ways: the sum total of characteristics differentiating people from each other, the stability in a person's behavior across different situations, characteristic ways people behave, and characteristics that are relatively enduring and make us behave in a consistent and predictable way.
Methods of Assessment of Personality:
Three main methods are used: Interview, Observation and behavioral assessment, and Psychological tests.
1. Interview: An interview is a direct, face-to-face encounter and interaction between the psychologist and the person being assessed. Both verbal and non-verbal information is available to the psychologist. Interviews are usually used to supplement information gathered through other sources. The skill of the interviewer is very important since the worth and utility of the interview depends on how well he can draw relevant information from the interviewee.
2. Behavioral Assessment: This refers to direct observation of behavior for investigating, understanding, and describing personality characteristics. Skill and expertise of the observer are the most significant ingredients of the observation process.
3. Psychological Tests: Psychological tests are standard measures devised in order to objectively assess personality and behavior. Like any other type of psychological test, personality tests also have to be valid and reliable. Availability of norms is an additional characteristic.
Psychological tests are generally of two types:
- Objective tests / personality inventories / self-report measures
- Projective tests
Objective Tests / Personality Inventories / Self-Report Measures: These are measures wherein the subjects are asked questions about a sample of their behavior. For example, MMPI (Minnesota Multiphasic Personality Inventory) is the most frequently used personality test. It was initially developed to identify people having specific sorts of psychological difficulties, but it can predict a variety of other behaviors too. It can identify problems and tendencies like Depression, Hysteria, Paranoia, and Schizophrenia.
Projective Tests / Techniques: Tests in which the subject is first shown an ambiguous stimulus and then has to describe it or tell a story about it are known as projective tests. The most famous and frequently used projective tests are the Rorschach test and TAT (Thematic Apperception Test).
How Is The Content Of A Personality Test Decided?
What the test will contain and how it will measure personality will be affected by the theoretical orientation of the test developer. Similarly, the choice of a test to assess personality also depends on how one defines personality.
1. Psychodynamic Approach: This approach focuses upon the unconscious determinants of personality. Psychologists belonging to this approach believe that unconscious forces determine our personality. The unconscious is the part of personality which we are not aware of and contains instinctual drives: infantile wishes, desires, demands, and needs. Therefore, the test based on this orientation will try to unfold and explore the unconscious.
🔑 Definition — Unconscious: the part of personality which we are not aware of, containing instinctual drives, infantile wishes, desires, demands, and needs.
2. Trait Approaches: These are the approaches that propose that there are certain traits that form the basis of an individual's personality. These approaches seek to identify the basic traits necessary to describe and understand personality. Traits are enduring dimensions of personality characteristics that differentiate a person from others. Trait theories do not imply the absence or presence of different traits in different people (i.e., either/or situation). These assume that some people are relatively high on some traits whereas some are low on the same traits.
🔑 Definition — Traits: enduring dimensions of personality characteristics that differentiate a person from others.
Trait theories based upon factor analysis: Factor analysis is a statistical method whereby relationships between a large number of variables are summarized into fewer patterns. These patterns are more general in nature. For example, a researcher prepares a list of traits that people may like in an ideal man. The extensive list is then administered to a large number of people, who are asked to choose traits that may describe an ideal man. Through factor analysis, the responses are statistically combined and the traits associated with one another in the same set (or person) are computed. Thus the most fundamental patterns are identified. These patterns are called factors.
🔑 Definition — Factor analysis: a statistical method whereby relationships between a large number of variables are summarized into fewer patterns, which are more general in nature and called factors.
Raymond Cattell's Sixteen Personality Factors: After using factor analysis, Cattell proposed that two types of characteristics form our personality: surface traits and source traits.
Eysenck's Dimensions of Personality: According to Eysenck, personality can be understood and described in terms of just two major dimensions: Introversion-extroversion and neuroticism-stability. On the first dimension, people can be rated ranging from introverts to extroverts; the rest of the traits fall in between. The second dimension is independent of the first one, and ranges from being neurotic to being stable. Introverts are quiet, passive, and careful people. Extroverts are outgoing, sociable, and active people. Neurotics are moody, touchy, and anxious people. Stable people are calm, care-free, and even-tempered people. Eysenck evaluated a number of people along these dimensions. Using the information thus obtained, he could accurately predict people's behavior in a variety of situations.
3. Social Cognitive Approach to Personality: This approach emphasizes the role of people's cognitions in determining their personalities. Cognitions include: people's thoughts, feelings, expectations, and values. These approaches consider the "inner" variables to be important in determining one's personality. These approaches emphasize the reciprocity between individuals and their environment. There exists a web of reciprocity, consisting of the interaction of environment and people's behavior. Our environment affects our behavior, and our behavior in turn influences our environment and causes modifications in the environment. The modified environment, in turn, affects our behavior.
💡 Why this matters: This approach highlights that personality is not fixed but dynamically shaped through ongoing interactions between internal thoughts and external surroundings.
⭐ Key Takeaways
Personality is defined as stable, enduring characteristics that make behavior consistent across situations, and it can be assessed through interviews, behavioral observation, or psychological tests. Psychological tests fall into two categories: objective tests like the MMPI, which ask direct questions about behavior, and projective tests like the Rorschach and TAT, which use ambiguous stimuli to uncover unconscious content. The content and method of a personality test are directly determined by the test developer's theoretical orientation — psychodynamic approaches focus on the unconscious, trait approaches identify basic dimensions like Cattell's 16 factors or Eysenck's introversion-extroversion and neuroticism-stability, and the social cognitive approach emphasizes the reciprocal interaction between cognitions (thoughts, expectations) and environment. Understanding these approaches is critical for selecting the appropriate test and interpreting results in clinical and research settings.
🧠 Quick Revision Questions
- What are the three major methods used for the assessment of personality?
- How do objective tests (personality inventories) differ from projective tests in terms of procedure and stimulus type?
- According to the psychodynamic approach, what does a personality test aim to explore, and why?
- What is factor analysis, and how did Cattell and Eysenck use it to identify the basic dimensions of personality?
- In the social cognitive approach, what does "reciprocity" between individuals and environment mean, and how does it influence personality?
📘 Lecture 33 — Objective / Structured Tests of Personality
📖 Overview: This lecture covers objective or structured personality measures, which are characterized by clear stimuli and specific response requirements. It explains the advantages of structured tests and provides detailed descriptions of major personality inventories including the Woodworth Personal Data Sheet, Mooney Problem Checklist, MMPI, CPI, and 16PF, along with their shortcomings.
🗂️ Topics Covered
This lecture introduces the concept of structured personality measures and their advantages. It then reviews several popular personality inventories: Woodworth Personal Data Sheet (developed for WWI combat recruit screening), Mooney Problem Checklist (identifying problems), Minnesota Multiphasic Personality Inventory (MMPI, MMPI-2, and MMPI-A versions with validity and clinical scales), California Psychological Inventory-Revised (CPI for normal individuals with four classes of scales), and Cattell’s Sixteen Personality Factor Questionnaire (16PF covering 16 primary source traits). It concludes with the shortcomings of structured tests, including format rigidity and response bias.
📝 Lecture Summary
Objective / Structured Tests of Personality
The objective measures of personality are also known as the structured measures. Structured measures of personality are characterized by structure and lack of ambiguity. A clear and definite stimulus is provided, and the requirements of the subject are evident and specific.
Advantages of Structured Tests:
- They are easily administered.
- The subject can endorse own responses using paper and pencil.
- They are easy to score.
- They can be group administered.
- Their scoring is uniform for all and the scorers’ likes, dislikes, or theoretical orientations do not interfere with the results.
- They are time-economical.
Woodworth Personal Data Sheet
Developed in: During World War I; final form published after the war (Woodworth, 1920). Developed for: Identifying recruits who were likely to break down in combat. The recruits who reported many symptoms were called for interview and were most likely to be rejected. Forms: Single. Format: A paper-pencil psychiatric interview. Items were chosen from psychiatrists’ questions asked in screening interviews and lists of known symptoms of emotional disorders. Items: 116 questions with a ‘yes’ ‘no’ format. Score: A single score was obtained as it was designed as a global measure of functioning. Individual/group: Used for mass screening.
Mooney Problem Checklist
Developed in: 1950. Developed for: Identifying problems experienced/faced by the subject. Items: Checklist of problems chosen from statements of around 4000 high-school students and clinical case history data. Score: Checked items indicate the problems experienced by the respondent. Weak points: One must rely on the reported problems; there is no way to check if reports are true. All that counts is face validity of responses.
Minnesota Multiphasic Personality Inventory (MMPI)
Developed by S. R. Hatahaway and J. C. Mckinley, first published in the 1940s (Hatahaway & Mckinley, 1940, 1942, 1943, 1951). It was intended for people 14 years of age and above, first published by University of Minnesota press in 1943.
Developed in: 1940s. Developed for: Identifying people with specific psychological difficulties or detecting major psychiatric/psychological disorders. It can predict a variety of other behaviors and identify problems like Depression, Hysteria, Paranoia, and Schizophrenia. The main goal is to differentiate abnormal persons from normal. Forms: MMPI, MMPI-2, and MMPI-A. Items: A self-report measure with statements in a true/false format. There are 566 items in MMPI and 567 in MMPI-2. Difference between MMPI and MMPI-2: MMPI had 566 items, 16 of which were repeated. In MMPI-2, the 16 repeated items, 77 items from 399 to 550, and 13 items from clinical scales were dropped. 460 items were retained from the original test. 127 new items were added: two critical items for severe pathology, 81 items for new content scales, and 24 unscored experimental items. MMPI-A is the version for adolescents, containing 478 items (88 less than earlier versions), used in school, educational counseling, and psychiatric settings. Similarities between MMPI and MMPI-2: Use and interpretation are the same for both. MMPI scales: Contains three validity scales and ten clinical scales. Score: Scores on MMPI scales are used to plot a profile showing tendencies/pathologies. Validity scales:
- Lie scale, 15 items, detects naïve attempts to ‘fake good’.
- K scale, 30 items, identifies defensiveness.
- F scale, 64 items, detects attempts to ‘fake bad’. Clinical scales:
- Hypochondriasis, 33 items, for physical complaints.
- Depression, 60 items, for detecting depression.
- Hysteria, 60 items, for detecting immaturity.
- Psychopathic deviate, 50 items, for authority conflict.
- Masculinity-femininity, 60 items, for masculine or feminine interests.
- Paranoia, 40 items, for identifying suspicion and hostility.
- Psychasthenia, 48 items, for detecting anxiety.
- Schizophrenia, 78 items, for detecting alienation and withdrawal.
- Hypomania, 46 items, indicates high energy and elated mood level.
- Social introversion, 70 items, yields score for shyness and introversion. Prerequisite: Subject should have IQ within the normal range. MMPI requires reading ability of grade 6; MMPI-2 requires reading ability of grade 8. Strong points: Includes validity scales that help identify faking. Standardization was done on large samples. Weak points: Lengthy test requiring long administration time.
California Psychological Inventory-Revised Edition (CPI)
CPI (Gough, 1987) follows the pattern of MMPI and many items are the same as MMPI. Developed for: Personality assessment in normally adjusted individuals. Scales: CPI has 20 scales grouped into four classes. Class I scales: Poise, self-assurance, and interpersonal effectiveness. High score: resourceful, active, competitive, outgoing, spontaneous, self-confident, at ease in interpersonal situations. Class II: Socialization, maturity, and responsibility. High score: honest, dependable, conscientious, calm, practical, cooperative, alertness to social issues. Class III: Achievement potential and intellectual efficiency. High score: efficient, organized, capable, forceful, knowledgeable, mature, and sincere. Class IV: Interest modes. High score: socially well adapted, responsive to others’ needs. Items: 462 items. Advantages: Can be used with normal subjects.
Cattell’s Sixteen Personality Factor Questionnaire (16 PF)
At least five editions available. Covers 16 primary source traits: A. Cool-warm B. Concrete thinking – Abstract thinking C. Affected by feelings – Emotionally stable D. Submissive – Dominant E. Sober – Enthusiastic F. Expedient – Conscientious G. Shy – Bold H. Tough minded – Tender minded I. Trusting – Suspicious J. Practical – Imaginative K. Forthright – Shrewd L. Self-assured – Apprehensive Q1: Conservative – Experimenting Q2: Group oriented – Self sufficient Q3: Undisciplined self-conflict – Following self-image Q4: Relaxed – Tense
Shortcomings of Structured Tests
- These tests are of less help if in-depth information about the subject’s personality is required.
- Their format is fixed and cannot be molded according to the respondent’s needs. For example, if the subject has problems understanding wording, no alterations can be made, especially in group administration.
- Accurate endorsement depends on the skill and understanding of the subject and the examiner’s skill.
- Response bias: People may have a tendency to mark all items in the same pattern (e.g., all true or all false).
⭐ Key Takeaways
Structured personality tests provide clear stimuli with specific response requirements, offering advantages in ease of administration, scoring uniformity, and time efficiency. Major inventories include the Woodworth Personal Data Sheet (WWI screen), MMPI with its validity and clinical scales (including MMPI-2 and MMPI-A versions), CPI for normal individuals, and 16PF with 16 source traits. The MMPI is notable for its validity scales (Lie, K, F) that detect faking, though it requires grade 6-8 reading ability. Shortcomings include format rigidity, inability to probe deeply, reliance on subject understanding, and potential response bias where subjects mark items in uniform patterns.
🧠 Quick Revision Questions
- What are the defining characteristics and advantages of structured personality measures?
- What was the Woodworth Personal Data Sheet developed for, and what was its format and scoring method?
- How does the MMPI differ from MMPI-2 in terms of items, and what are the three validity scales and their purposes?
- What do the four classes of scales in the California Psychological Inventory (CPI) measure?
- What are the main shortcomings of structured personality tests, particularly regarding response bias?
📘 Lecture 34 — Projective Personality Tests
📖 Overview: This lecture introduces projective personality tests, focusing on the most famous examples like the Rorschach Inkblot Test and the Thematic Apperception Test (TAT). It explores how ambiguous stimuli can reveal unconscious thoughts, feelings, and motivations, and examines a wide range of related projective techniques including picture-story tests, word association tests, sentence completion tests, and figure drawings.
🗂️ Topics Covered
The lecture covers the definition of projective tests and the Rorschach Inkblot Test with its procedure and administration. It then details the Thematic Apperception Test (TAT) including its development, purpose, scoring, items, forms, strengths, and weaknesses. Other picture-story tests such as the Children's Apperception Test (CAT), The Picture Story Test, The Education Apperception Test, The Michigan Picture Test, and the Make A Picture Story Method are discussed. Tests using pictures as projective stimuli like The Hand Test and The Rosenzweig Picture-Frustration Study are covered, followed by tests using words including Word Association Tests and The Kent-Rosanoff Free Association Test. Sentence completion tests like the Rotter Incomplete Sentence Blank (RISB) and production figure drawings such as the Draw A Person test (DAP) and the House-Tree-Person test (HTP) are also examined, concluding with the advantages and disadvantages of projective tests.
📝 Lecture Summary
Projective Personality Tests
The projective tests are the tests in which the subject is first shown an ambiguous stimulus and then he has to describe it or tell a story about it. Two most famous and frequently used projective tests are the Rorschach test and the TAT or Thematic Apperception Test.
🔑 Definition — Projective test: A test where a subject is shown an ambiguous stimulus and must describe it or tell a story about it.
Rorschach Ink Blot Test
The test consists of inkblot presses. These have no definite shape. The shapes are symmetrical, and are presented to the subject on separate cards. Some cards are black and white and some colored.
Procedure of Rorschach administration: The subject is shown the stimulus card and then asked as to what the figures represent to them. The responses are recorded. Using a complex set of clinical judgments, the subjects are classified into different personality types. The skill and the clinical judgment of the psychologist or the examiner are very important.
Thematic Apperception Test (TAT)
The TAT was developed by Christiana D. Morgan and Henry Murray in 1935 during their working at Harvard Psychological Clinic. Originally the TAT was developed for patients in psychoanalysis to obtain raw data. The main objective of TAT is to evaluate a person's patterns of thought, attitudes, observational capacity, and emotional responses to ambiguous test materials.
The interpretive system of TAT identifies the story with individual/person described in story, needs and demands of environment created by storyteller. Among variety of scoring interoperations the scoring system is based on Murray’s personality theory. The test contains 30 black-and-white picture cards with variety of situations including human figures and situations. Some cards are suggested to use with adult males or females and some are used with children. In clinical practice TAT is used depending on the client’s need and situation; the practitioner may use 20 cards as prescribed number or depending on client’s story-telling capacity, 1-2 or 30 cards may be used.
Strong Points: The test has great intuitive appeal. It helps to identify emotions and motivations of the storyteller projected by unambiguous stimuli. The test can be used with ample liberty of administration and scoring for practitioners.
Weak Points: The psychometric properties of the TAT are debated like other projective techniques. It has general lack of standardization in administration, scoring, and interpretations procedures.
💡 Why this matters: The TAT's strength in revealing unconscious motivations is offset by significant concerns about its reliability and validity due to lack of standardization.
Other Picture-Story Tests
Children’s Apperception Test (CAT) was developed by Leopold Bellak in 1949. The CAT was developed for children ages 3-10 years. The main objective of CAT is to measure the personality traits, attitudes, and psychodynamic processes evident in children. Scoring of the CAT is not based on objective scales; it must be performed by a trained test administrator or scorer. The scorer's interpretation should take into account: the story's primary theme; the story's hero or heroine; the needs or drives of the hero or heroine; the environment in which the story takes place; the child's perception of the figures in the picture; the main conflicts in the story; the anxieties and defenses expressed in the story; the function of the child's superego; and the integration of the child's ego. The test contains black-and-white picture cards and uses animal figures instead of humans. In addition to the original CAT, an alternative version called CAT-H is also published.
The Picture Story Test was developed by Symonds in 1949. The test was developed to use with adolescents. The main objective is to elicit stories related to specific situations like coming home late, leaving home, and planning for the future. The test contains 20 picture cards.
The Education Apperception Test was developed by Thompson and Sones in 1973 and the School Apperception Method was designed by Solomon and Starr in 1968. These two instruments were designed to tap children’s attitude toward school and learning.
The Michigan Picture Test was developed in 1953. It is used to elicit various responses, ranging from conflicts with authority figures to feelings of personal inadequacy. The test was developed for use with children between the ages of 8 and 14 years and contains 16 pictures.
Make A Picture Story Method was developed in 1952. The test helps to indicate the thinking, feelings of test-taker’s projections. The test contains 67 cut-up figures of people and animals that may be presented on any of 22 pictorial backgrounds. The number of figures with blank faces and different background settings like living room, street, nursery, stage, bridge etc. are available to use.
Tests Using Pictures As Projective Stimuli
The Hand Test was developed in 1983 by Wagner. The test contains nine cards with pictures of hands on them and a tenth blank card. The test taker is asked what the hands on each card might be doing. When presented with the blank card, the test taker is instructed to imagine a pair of hands on the card and then describe what they might be doing. The responses are interpreted according to 24 categories such as affection, dependence, and aggression.
The Rosenzweig Picture-Frustration Study was originally developed in 1947 by Rosenzweig. The test is based on the assumption that the test taker will identify with the person being frustrated. The test is available in forms for children, adolescents, and adults. The task of the test is to fill in the response of a cartoon figure being frustrated. The responses are scored in terms of the type of reaction elicited and the direction of the aggression expressed.
Words as Projective Stimuli
The first attempt of using words as a projective measure was made by Galton in 1879. Afterwards Cattell and Bryant in 1889, Kraepelin in 1896, and Jung in 1910 used words as tests. The task of these tests is to interpret the responses of words.
Word Association Tests were developed by Rapaport, Gill, and Schafer in 1946. The Word Association Test tries to evaluate the responses with respect to variables like popularity, reaction time, content, and test-retest responses. The length of response time is recorded each time. The examinee is asked to clarify the relationship existing between the original word and response word. The test consists of 60 words presented in two parts: first, the examinee is asked to respond quickly with the first word that came to mind; in the second part, he is again presented with words and asked to reproduce the original response.
The Kent-Rosanoff Free Association Test was developed in 1910. The test attempts at standardizing the response of individuals to specific words. The purpose of the test is to identify the individuality of response that may be influenced by psychopathology and many other variables. The test consists of 100 stimulus words.
Sentence Completion Tests
Sentence completion tests are some other tests that use verbal material as projective stimuli. A number of standardized tests are available for use.
Rotter Incomplete Sentence Blank (RISB) is a standardized test developed in 1950. The test was developed for use with populations from grade 9 through adulthood. The manual of the RISB suggests that responses be interpreted according to several categories: family attitudes, social and sexual attitudes, general attitudes, and character traits. Each response is evaluated on a 7-point scale ranging from "need for therapy" to "extremely good adjustment". The test consists of 40 incomplete sentences.
Strong Points: The test may be used for obtaining diverse information relating to an individual’s interests, educational aspirations, future goals, fears, conflicts, needs, and so forth.
Weak Points: The sentence completion test is the most vulnerable of all the projective methods to faking on the part of the examinee intent on making a good or bad impression.
Production Figure Drawings
The use of drawings in clinical and research settings has extended beyond the area of personality assessment. It attempts to use artistic productions as a source of information about intelligence, neurological intactness, visual-motor coordination, cognitive development, and learning disabilities. The figure drawings are an appealing source of diagnostic data.
Draw A Person test (DAP) is developed on the working of Karen Machover (1949). The drawings of a person made by the examinee are evaluated for various characteristics including placement, size of the figure, pencil pressure used, symmetry, line quality, shading, the presence of erasures, facial expressions, posture, clothing, and overall appearance. The test needs a simple pencil and 8 ½ by 11 inch paper, and the person is asked to draw a person.
The House-Tree-Person test (HTP) was developed and popularized by Buck in 1948. The drawings of a house, tree, and person are used as a reflective source of psychological functioning. Pencil and white paper are required, and the person is instructed to draw a picture of a house, a tree, and a person.
Advantages of Projective Tests
- In-depth investigation
- Flexible nature
- Subject’s liberty to respond in whatever way
- Psychologist has access to nonverbal cues
- Used for psychodynamic examination
Disadvantages of Projective Tests
- The psychologist has to be highly skillful
- These tests may be very time consuming
⭐ Key Takeaways
The core of this lecture is that projective tests use ambiguous stimuli—inkblots, pictures, words, or drawings—to bypass conscious defenses and reveal unconscious personality dynamics. The Rorschach and TAT are the most prominent, with the TAT being particularly useful for identifying motivations and emotional responses through story-telling. A key limitation across all projective methods is the lack of standardization in administration and scoring, which compromises their psychometric properties (reliability and validity) despite their intuitive appeal for in-depth, qualitative assessment. For the exam, remember the specific names of developers (Morgan & Murray for TAT, Bellak for CAT, Machover for DAP), the populations each test targets (e.g., CAT for children 3-10, RISB for grade 9+), and the crucial weakness of faking as most applicable to sentence completion tests.
🧠 Quick Revision Questions
- What is the fundamental principle behind all projective personality tests?
- Who developed the Thematic Apperception Test (TAT), and what is its primary purpose?
- How does the Children's Apperception Test (CAT) differ from the standard TAT in terms of its stimulus materials and target population?
- Name one specific test that uses pictures of hands as a projective stimulus and state what it is designed to reveal.
- What is a major weakness of the Rotter Incomplete Sentence Blank (RISB) compared to other projective tests?
📘 Lecture 35 — Personality: Measurement of Interests and Attitudes
📖 Overview: This lecture explores the development and application of interest and attitude measurement in psychological testing. It traces the evolution from early vocational interest inventories like the Strong Vocational Interest Blank to more theoretically grounded and gender-fair instruments, while also introducing projective techniques such as sentence completion tests and figure drawings. Understanding these tools is essential for career counseling and personality assessment.
🗂️ Topics Covered
The lecture covers the historical development of interest measurement beginning with the Strong Vocational Interest Blank (SVIB) and its criterion-group approach, followed by the Strong-Campbell Interest Inventory (SCII) which incorporated Holland's theory of vocational choice. It then examines the Kuder Occupational Interest Survey (KOIS), Jackson Vocational Interest Survey (JVIS), Minnesota Vocational Interest Inventory (MVII), Career Assessment Inventory (CAI), and the Self-Directed Search (SDS). The lecture concludes with discussions on eliminating gender bias in interest measurement and introduces projective techniques including the Rotter Incomplete Sentence Blank (RISB), Draw-A-Person test (DAP), and House-Tree-Person test (HTP).
📝 Lecture Summary
Strong Vocational Interest Blank (SVIB)
I. E. K. Strong, Jr., and colleagues studied the activities of persons belonging to different professions. They observed that different professionals had different likes and dislikes for different activities, and their interests followed different patterns. Hobbies of people working in the same profession were also similar, as they tended to indulge in similar pastime activities. Using the criterion-group approach, Strong developed a test to see if the interests of a test taker matched with the interests and values of people who were happy in their chosen careers.
The test is called Strong Vocational Interest Blank (SVIB). The 399 items of SVIB are related to 54 occupations for men and 32 occupations for women, presented separately. Items in the SVIB were weighted according to the frequency of occurrence of an interest in a particular occupational group compared to how frequently it occurred in the general population. The test provides strong psychometric characteristics, with normative samples of around 300 people in each criterion group. Raw scores were converted to standard scores with a mean of 50 and SD of 10.
Research showed that interest patterns usually become established by age 17 years and remain relatively stable over time. In a study of Stanford University students taking this test first in the 1930s and then later, interests remained relatively unchanged even after 22 years. However, the SVIB was criticized for gender bias (having used different scales for men and women) and lack of theoretical information — there was no theoretical basis to explain why people in different professions had similar interests.
🔑 Definition — Criterion-group approach: A test development method where items are selected based on their ability to differentiate between known groups (e.g., people in different professions).
📐 Standard score formula: Raw scores converted to a distribution with mean = 50, SD = 10 → allows comparison across different occupational scales.
Holland's Theory of Vocational Choice
Holland (1975) presented his theory of vocational choice, according to which people's personality is expressed in their interests. Furthermore, considering people's interests, we can classify them into one or more of the following six categories of personality factors:
- a. Realistic
- b. Investigative
- c. Artistic
- d. Social
- e. Enterprising
- f. Conventional
These factors were gender-bias free and could be used with both men and women. There was a similarity between Holland's factors and the patterns of interests yielded by research with SVIB. This appealed to Campbell, who incorporated this theory in his version of SVIB.
💡 Why this matters: Holland's theory provided the theoretical foundation that SVIB lacked, explaining why occupational interests cluster together and allowing for more meaningful interpretation of interest test results.
The Strong-Campbell Interest Inventory (SCII)
The SVIB had certain features for which it was criticized. A new version was developed by D. P. Campbell and named The Strong-Campbell Interest Inventory (SCII).
- Developed in: 1974
- Population: Both men and women — it is a gender-free version
- Special feature: Unlike the SVIB, the SCII does not have separate forms for men and women. The gender bias was removed and items from male and female forms were merged into one form
- Items: There are 325 items with response options including 'like', 'dislike', or 'indifferent'
- Parts: Seven parts including Occupations (131 items), School subjects (36 items), Activities (51 items), Amusements (39 items), Types of people (24 items), Preference between two activities (30 items), Your characteristics (14 items)
- Scoring: Several scores are obtained: General themes based on Holland's six personality types, scoring for administrative indexes, a person's basic interests, and occupational scales
- Strong feature: The test follows Holland's theory of vocational choice and therefore has a theoretical basis, and it is not gender biased
The Kuder Occupational Interest Survey (KOIS)
The Kuder Occupational Interest Survey (KOIS) requires the test taker to select the most preferred and least preferred activity among 100 triads of alternative activities. A triad is a set of three alternatives. The similarity between the test taker's interests and those employed in various occupations is assessed.
This instrument has separate norms for men and women, and separate scales for college majors are also available. The KOIS helps in two major ways: it suggests which occupational group may be best suited to a person's interests, and it can assist students in choosing their majors. A series of new scales has been added to KOIS for nontraditional occupations. Despite lack of research data, the KOIS is useful for guidance decisions for high-school and college students.
The Jackson Vocational Interest Survey (JVIS)
The Jackson Vocational Interest Survey (JVIS) was developed by D. N. Jackson, with a revised version from 1995 in common use. The test is used for career education and counseling of high-school and college students, and can be used for career planning of adults and those seeking mid-life career changes.
- Items: Contains 289 statements about job-related activities, completed in around 45 minutes
- Response format: Forced choice format where the respondent selects one interest they prefer over the other
- Forms: Available in both hand-scored and machine-scored forms
- Strong points: The test has strong psychometric properties and carefully avoided gender bias
The Minnesota Vocational Interest Inventory (MVII)
The Minnesota Vocational Interest Inventory (MVII) is designed for men who are not oriented toward college, emphasizing skilled and semiskilled trades. It has been used extensively by the military and by guidance programs for individuals not going to college. The MVII has nine basic interest areas and 21 specific occupational scales. Basic interest areas include mechanical interests, electronics, and food service. Occupational scales cover occupations like plumber, carpenter, and truck driver.
The Career Assessment Inventory (CAI)
The Career Assessment Inventory (CAI) was developed by Charles B. Johansson in 1976. The test is developed for American citizens with less than four years of postsecondary education, comprising around 80% of U.S. citizens. The test requires a reading ability of grade-6 reading level.
- Items: Test takers are evaluated on Holland's six occupational theme scales, basic interests in 22 areas are also assessed, and 89 occupation scales are used
- Strong points: The test has good validity and reliability, and the developer tried to make the CAI culturally fair and eliminate gender bias
The Self-Directed Search (SDS)
J. L. Holland developed the Self-Directed Search (SDS), which attempts to simulate the counseling process by allowing respondents to list occupational aspirations, indicate occupational preference in six areas, and rate abilities and skills in these areas.
- Items: The test has 228 items. There are six scales with 11 items each describing activities. Competencies are assessed by 66 items. Another six scales with 14 items each evaluate occupations
- Scoring: The SDS is self-administered, self-scored, and self-interpreted. Test takers score the inventory and calculate six summary scores; these scores and further codes reflect areas of highest interest
Eliminating Gender Bias in Interest Measurement
Advocates of women's rights pointed out discrimination against women in early interest inventories. The Associate for Evaluation in Measurement appointed the Commission on Sex Bias in Measurement, which concluded that interest inventories contributed to guiding young men and women into gender-typed careers. The SVIB had separate forms for women, but careers for women tended to be lower in status and commanded lower salaries. Because career choices for many women are complex, interest inventories alone may be inadequate, and more comprehensive approaches are needed.
Rotter Incomplete Sentence Blank (RISB)
Sentence completion tests use verbal material as projective stimuli. The Rotter Incomplete Sentence Blank (RISB) is a standardized test developed in 1950 for populations from grade 9 through adulthood.
- Items: The test consists of 40 incomplete sentences
- Scoring: Responses are interpreted according to several categories: family attitudes, social and sexual attitudes, general attitudes, and character traits. Each response is evaluated on a 7-point scale ranging from "need for therapy" to "extremely good adjustment"
- Strong points: The test may be used for obtaining diverse information relating to an individual's interests, educational aspirations, future goals, fears, conflicts, needs, and so forth
- Weak points: The sentence completion test is most vulnerable of all projective methods to faking on the part of the examinee intent on making a good or bad impression
Production Figure Drawings
The use of drawings in clinical and research settings has extended beyond personality assessment. It attempts to use artistic productions as a source of information about intelligence, neurological intactness, visual-motor coordination, cognitive development, and learning disabilities. Figure drawings are an appealing source of diagnostic data.
Draw-A-Person Test (DAP)
The Draw-A-Person Test (DAP) was developed based on the work of Karen Machover (1949). The test requires a simple pencil and 8½ by 11 inch paper, and the person is asked to draw a person. Drawings are evaluated for various characteristics including placement, size of the figure, pencil pressure used, symmetry, line quality, shading, presence of erasures, facial expressions, posture, clothing, and overall appearance.
The House-Tree-Person Test (HTP)
The House-Tree-Person Test (HTP) was developed and popularized by Buck in 1948. The drawings of a house, tree, and person are used as a reflective source of psychological functioning. The test requires pencil and white paper, and the person is instructed to draw a picture of a house, a tree, and a person.
Advantages and Disadvantages of Projective Tests
Advantages: In-depth investigation, flexible nature, subjects' liberty to respond in whatever way, and used for psychodynamic examination.
Disadvantages: The psychologist has to be highly skillful, and these tests may be very time-consuming.
⭐ Key Takeaways
The measurement of interests evolved from the atheoretical, gender-biased SVIB to theoretically grounded instruments like the SCII that incorporated Holland's six personality types (Realistic, Investigative, Artistic, Social, Enterprising, Conventional), with modern tests striving for gender fairness and cultural appropriateness. Key instruments serve different populations — the KOIS uses triads for high-school/college students, the JVIS uses forced-choice for career planning, the MVII targets non-college men, and the CAI serves those with less postsecondary education. The Self-Directed Search is unique as a self-administered, self-scored inventory. Beyond interest measurement, projective techniques like the Rotter Incomplete Sentence Blank (40 incomplete sentences, scored on adjustment level), Draw-A-Person test, and House-Tree-Person test use verbal and artistic productions to assess personality, though sentence completion tests are particularly vulnerable to faking.
🧠 Quick Revision Questions
- What are Holland's six personality types used in vocational interest measurement, and how did they address the criticism against SVIB?
- What is the criterion-group approach used by Strong in developing the SVIB, and how were items weighted in this test?
- How does the Strong-Campbell Interest Inventory (SCII) differ from the SVIB in terms of gender bias and theoretical foundation?
- What is the response format of the Kuder Occupational Interest Survey (KOIS), and what are its two major uses?
- What is the Rotter Incomplete Sentence Blank (RISB), how many items does it contain, and what is its primary weakness as a projective technique?
📘 Lecture 36 — Measurement of Attitudes, Opinions, Locus of Control, Multidimensional Health Self-efficacy
📖 Overview: This lecture explores how psychologists measure attitudes, opinions, and related personality constructs. It covers classic attitude scaling techniques, the concept and measurement of locus of control, and specialized instruments for health-related beliefs and self-efficacy. Understanding these measurement approaches is essential for research in personality, social psychology, and health psychology.
🗂️ Topics Covered
The lecture begins with the definition and measurement of attitudes and opinions, detailing the Thurstone and Likert scaling methods along with other popular formats. It then explains the concept of locus of control, Rotter's original internal-external dimension, and its evolution into multidimensional health locus of control. Finally, it covers the measurement of self-efficacy, distinguishing between general perceived self-efficacy and domain-specific measures like the GSE scale.
📝 Lecture Summary
Measurement of Attitudes
Besides measuring different personality characteristics, personality dynamics, and interests, psychologists assess and investigate many other aspects of personality too. Attitudes and opinions are two such aspects. An attitude can be defined as “a tendency to react favorably or unfavorably toward a designated class of stimuli, such as a national or ethnic group, a custom, or an institution” (Anastasi & Urbina, 2007, p. 418-419). Another related aspect is ‘opinions’. Attitudes and opinions are closely related since one’s opinions are determined and directed by one’s attitudes. The measures commonly used for the measurement of attitudes are called attitude scales. There are different types of scales available. However, in most attitude or opinion surveys, one needs to develop attitude scales or other instruments according to the theme of the research. One can develop scales following the formats available. Some of the scales whose format is commonly followed are as follows:
🔑 Definition — Attitude: a tendency to react favorably or unfavorably toward a designated class of stimuli, such as a national or ethnic group, a custom, or an institution. 🔑 Definition — Attitude scales: measures commonly used for the measurement of attitudes.
Thurstone (1931) Scale/ Equal Appearing Intervals
In this scale the scale development is most important. A large number of attitude-related statements are developed. The statements can be positive and negative toward the object of attitude. A panel of judges rates each statement from one to eleven. One means highly negative on the subject and eleven indicates highly positive. The ratings of all judges are processed and mean rating for each statement is gauged. When respondents are given scores according to the mean values obtained from the judges, they score the scale value of each item agreed with.
🔑 Definition — Thurstone Scale: an attitude scale where a panel of judges rates statements from 1 (highly negative) to 11 (highly positive), and respondents receive scores based on the mean rating of items they agree with. 📌 Example: A researcher develops 50 statements about recycling. A panel of 20 judges rates each statement. The statement "Recycling is essential for the planet" receives a mean rating of 9.8. A respondent who agrees with this statement earns a score of 9.8 on that item.
Likert (1932) Scale
In this type of scale a number of statements are developed regarding the object of attitude. The statements are both favorable and unfavorable. They are rated according to the following scale:
| 5 | 4 | 3 | 2 | 1 |
|---|---|---|---|---|
| Strongly agree | Agree | Undecided | Disagree | Strongly disagree |
The value of a chosen response is the score of the person on that item. The total or overall score is calculated from the score on individual items.
🔑 Definition — Likert Scale: an attitude scale where respondents rate agreement with statements using a 5-point scale (Strongly agree to Strongly disagree), and the total score is the sum of item scores. 📌 Example: A statement "I enjoy exercising" is presented with the Likert format. If a respondent chooses "Agree" (score 4), that is their item score. If they choose "Strongly disagree" (score 1) for a negative statement "Exercise is boring", that score is reversed before summing.
Some other popular scales include the Guttman (1950) scale, the Semantic Differential Scale (Osgood et al., 1957), and the Social Distance Scale (Bogardus, 1925).
Locus of Control
Locus of control refers to a person’s perceptions or beliefs about the location of responsibility for his or her life circumstances, happenings, events, conditions. The perception of who is in-charge of one’s life, who decides one’s fate, and who is responsible for whatever the person is experiencing, is determined by the person’s locus of control (LOC). People’s perceptions of success and failure, of health and illness, ability or inability, all reflect their locus of control. In other words, the concept of LOC refers to perceived control; the perception of how much a person feels in control of life (Lefcourt, 1982). The earliest formal investigations of the concept were reported by Julian Rotter (1966, 1975). Rotter showed that people have different views about things that happen to them. People have their own beliefs or generalized expectations about where the control of life, or events, resides.
Rotter’s original formulation of LOC had a dualistic approach. People’s beliefs regarding who is responsible for events, and who influences life, were classified along a bipolar dimension. Rotter showed that people have different views about the source of control, and the things happening to them; these beliefs, and therefore people holding them, were seen as falling into two categories namely, internal and external, that could be measured with the I-E scale. Later on Levenson and others added to this construct as well as its measure. However, perhaps the most quoted contribution in this regard is in the assessment of Multidimensional Health Locus of Control.
🔑 Definition — Locus of Control (LOC): a person’s perceptions or beliefs about the location of responsibility for his or her life circumstances, happenings, events, and conditions. 🔑 Definition — Internal LOC: the belief that one's own actions determine outcomes. 🔑 Definition — External LOC: the belief that outcomes are determined by external forces like fate, luck, or powerful others.
Measurement of Multidimensional Health Locus of Control (MHLC)
Health locus of control can be measured by using specifically designed instruments. Its quantification gives an edge to this construct over many other constructs pertaining to health beliefs. The Multidimensional Health Locus of Control (MHLC) scales were developed by B. S. Wallston and K. A. Wallston, and their colleagues (Wallston, Wallston, Kaplan, & Maides, 1976; Wallston, Wallston, & De Vellis, 1978). These are the most popularly used measures for quantifying HLC.
Krischt and his colleagues (Dabbs & Krischt, 1971; Krischt, 1972) were the ones who produced the first published version of an LOC measure specifically meant for use in the domain of health and illness. Due to some inherent flaws in this measure however, it could not gain popularity and a need was felt for more precise measures. B. S. Wallston and K. A. Wallston have so far made highly significant contributions to the measurement of HLC.
🔑 Definition — Health Locus of Control (HLC): a person's beliefs about who or what controls their health outcomes.
Multidimensional Health Locus of Control (MHLC) Scales
The MHLC follows the 6-point Likert response pattern, and includes three scales. In conceiving these scales, Levenson’s (1973, 1981) multidimensional approach was followed, in which the dimension of externality was split into two components. Hence the scales for powerful others health locus of control (PHLC), and chance health locus of control (CHLC). Instead of treating EHLC as a single dimension that covers the influence of powerful others and that of chance as one and the same dimension, the new versions treated PHLC and CHLC as two separate components. PHLC, that measures the powerful others health locus of control, pertains to a person’s beliefs about the influence of other people on her health. The people who are believed to have the power to determine one’s health may include the family, friends, doctors, hospital staff and others. The extent to which a person believes in the influence of chance, luck, or fate is measured by the CHLC Scale. CHLC measures the beliefs about health or illness being related to chance, fate or luck, instead of being related to one’s own responsibility.
The three MHLC scales include:
- Internal Health Locus of Control (IHLC)
- Powerful Others Health Locus of Control (PHLC)
- Chance Health Locus of Control (CHLC)
A six-point rating scale is used with the MHLC items, where one can make a choice from response options ranging from ‘strongly disagree’ to ‘strongly agree’. The scale contains three subscales and eighteen statements in all. Each subscale carries six items/ statements. The items pertaining to the three subscales have been mixed up. The scoring procedure is very simple. The values of the marked responses to each item in a subscale are added up. The sum indicates the person’s score on that subscale. The subscale on which the person obtains the highest score indicates the type of health locus of control that the person has. One may choose any one from forms A or B. The IHLC Scale assesses the internal health locus of control that is the extent to which a person believes that her health or illness is determined by internal factors. K. A. Wallston and B. S. Wallston (1982) have asserted that the dimensions measured by the scales are more or less statistically independent. Therefore a low IHLC score does not necessarily indicate that the person believes in the influence of external factors, powerful others, or chance. A low IHLC score may be understood to mean that the person’s belief in the influence of internal factors is low (or may be non-existent in some cases).
💡 Why this matters: The multidimensional approach allows researchers to distinguish between different types of external beliefs (chance vs. powerful others), providing more nuanced insights into health behaviors and beliefs.
🔑 Definition — IHLC: the extent to which a person believes that her health or illness is determined by internal factors. 🔑 Definition — PHLC: measures a person’s beliefs about the influence of other people (family, friends, doctors, hospital staff) on her health. 🔑 Definition — CHLC: measures the extent to which a person believes in the influence of chance, luck, or fate on health or illness. 📐 Formula: MHLC Score = Sum of responses to items in each subscale (IHLC, PHLC, CHLC), each subscale has 6 items rated on a 6-point scale. 📌 Example: A patient scores 28 on IHLC, 15 on PHLC, and 20 on CHLC. Since IHLC is the highest, this indicates the patient has an internal health locus of control, believing their own actions primarily determine their health.
The Measurement of Self Efficacy
Self-efficacy, as the very name suggests, is the perception of one’s own ability to produce some desired outcomes. Researchers have used the construct of self-efficacy for assessing the impact of people’s perceptions of personal control and capability on their behavior in a variety of situations. A divergent range of self-efficacy measures is available to researchers interested in investigating the relationship between thought and action. Although a considerable majority of studies in this regard have investigated the influence of perceived self-efficacy on people’s health-related behaviors, measures like collective teacher self-efficacy, and teacher self-efficacy scales have also been devised.
For a health psychology researcher, primarily two types of measures of self-efficacy are available. These instruments can be used to assess health related self-efficacy in two ways, including:
- General perceived self-efficacy
- Perceived self-efficacy pertaining to specific health behaviors.
The measures included in the above mentioned categories may be used in their original form as well as with alterations made according to the nature of problem under investigation.
🔑 Definition — Self-efficacy: the perception of one’s own ability to produce some desired outcomes.
General Perceived Self-efficacy (GSE) Scale
The most widely used measure of self-efficacy, General Self-Efficacy (GSE) Scale, was developed by Matthias Jerusalem and Ralph Schwarzer in 1981. The original version, in German language, comprised 20 items, but the later version consisted of only 10 items (Schwarzer & Scholz, 2000; Schwarzer & Jerusalem, 1995). It is this 10 item version that is used by researchers studying self-efficacy in recent researches.
The GSE scale is available in at least 26 different languages. The scale was originally devised for predicting both coping and adaptation; how people coped with daily hassles, and how they adapted after having undergone stressful life events. This scale gauges a generalized sense of self-efficacy, indicating the overall global confidence that a person has about personal ability to cope with a wide range of situations that may be new, novel, taxing or demanding. GSE focuses upon a sense of competence that is broad and stable, rather than being only domain-specific (Schwarzer & Scholz, 2000).
The response format of GSE scale is uniform for all 10 items, consisting of a 4-point scale. The response options range from ‘not at all’ (definitely not) marked as 1, to ‘exactly true’ marked as 4. The final composite score is obtained by adding up the responses to all 10 items. The final score may range from 10 to 40. If the person marks ‘not at all’ in response to the entire range of items, he will be understood to be standing at the lowest possible level of self-efficacy. This can be taken to indicate a lack of self-efficacy. On the contrary a score of 40 will mean the person has the highest possible level of feeling self-efficacious. On average, the scale can be completed in 4 minutes. Some people may take longer, or lesser than the average time.
🔑 Definition — General Self-Efficacy (GSE) Scale: a 10-item scale measuring a generalized sense of self-efficacy, or global confidence in one's ability to cope with a wide range of demanding situations. 📐 Formula: GSE total score = sum of responses to all 10 items (each item scored 1 to 4); possible range = 10 to 40. 📌 Example: A person responds to all 10 GSE items. For item "I can always manage to solve difficult problems if I try hard enough," they select "exactly true" (score 4). Their total score across all items is 35, indicating a relatively high sense of general self-efficacy.
⭐ Key Takeaways
The lecture covers two major areas: attitude measurement and locus of control/self-efficacy assessment. For attitudes, you must know the difference between Thurstone (judges rate items, respondents receive scale values) and Likert (respondents rate agreement on a 5-point scale, total scores are summed) scales. The MHLC scale is the most important tool for health locus of control, with three statistically independent subscales: IHLC (internal), PHLC (powerful others), and CHLC (chance). The GSE scale is a 10-item, 4-point measure of general self-efficacy, with scores ranging from 10-40. For exams, focus on the definitions, structures, and scoring procedures of each instrument, especially the MHLC and GSE scales.
🧠 Quick Revision Questions
- What is the key difference between the Thurstone and Likert attitude scaling methods in terms of how items are scored?
- What are the three subscales of the Multidimensional Health Locus of Control (MHLC) scales and what does each measure?
- According to Rotter, what are the two main categories of locus of control, and how are they measured with the I-E scale?
- How is the General Self-Efficacy (GSE) scale scored, and what is the range of possible total scores?
- Why did the MHLC scales split the external dimension into two separate components (PHLC and CHLC) rather than treating it as a single dimension?
📘 Lecture 37 — Alternate Approaches to Personality Assessment: Behavioral and Cognitive- Behavioral Testing
📖 Overview: This lecture explores alternative approaches to personality assessment beyond traditional psychometric tests. It introduces behavioral assessment, which provides direct, observable evidence of a person's behavior, and cognitive-behavioral assessment, which adds a focus on the individual's cognitions. These approaches are crucial for supplementing test results in decision-making, diagnosis, and treatment planning.
🗂️ Topics Covered
The lecture begins by explaining the need for behavioral assessment as a supplement to psychological testing. It then details the core method of behavioral observation, including self-observation and self-monitoring, and covers various recording methods such as narrative records and behavioral rating scales. Following this, situational performance measures are introduced. The lecture then transitions to cognitive-behavioral assessment, explaining its rationale, the four steps involved, and providing examples of specific methods including the Fear Survey Schedule (FSS), the Irrational Beliefs Test (IBT), and Kanfer and Saslow's Functional Approach.
📝 Lecture Summary
Behavioral Assessment:
Traditional psychological tests measure underlying traits and characteristics, but for more confident judgments, behavioral assessment offers direct, tangible evidence. It samples a person's behavior in specific situations, which is safer to use alongside test results for diagnosis or screening. This approach can be used independently but is best employed as part of a battery of procedures.
🔑 Definition — Behavioral assessment: A method of personality assessment that yields information regarding a sample of behavior, assumed to represent a person's behavior in various situations, providing first-hand, direct, and tangible evidence.
Behavioral Observation:
The most basic method for behavioral assessment is observation, commonly used by developmental, educational, child, and school psychologists. They observe and record the behavior of interest. Technical devices like audio or video recordings can increase accuracy by capturing details the observer might miss. Trained staff may also be employed for record-keeping.
An offshoot of this is self-observation, where a person reports their own behavior. A similar approach is self-monitoring, in which the subject keeps a record of their own behavior as it happens. For example, a smoker might record each cigarette smoked during the day, or a bulimic might record binge eating episodes.
Recording Observed Behavior:
Several approaches can be used to record the behavior in question: a. Narrative records: The observer takes detailed notes, but this can lead to missing information. A better approach is to take short notes and fill in gaps afterward. b. Audio/video recordings: These capture the behavior accurately. c. Behavioral rating scales: Observers use codes instead of narratives to record behavior, saving time and ensuring inter-rater uniformity. These scales can record presence/absence, frequency, intensity, and other aspects. Examples include the Play Performance Scale for Children, the Walker Problem Behavior Identification Checklist, the Behavior Rating Profile, and the Social Skills Rating System.
Situational Performance Measure:
This is a form of observation where a person's behavior is observed under specific, either real or simulated, circumstances. For example, a lecturer candidate might deliver a model lecture to a real class, or an astronaut candidate might perform in a simulated state of weightlessness.
One form of this is a situational stress test, where behavior is observed under stress, anxiety, or frustration. This is used for jobs involving psychological pressure, such as in the armed forces. The U.S. Office of Strategic Services (OSS, 1948) used such tests during World War II for selecting military intelligence candidates.
Cognitive- Behavioral Assessment:
This approach is similar to the behavioral approach but includes a cognitive component. It is based on the cognitive-behavioral model, which is relatively modern.
Why Cognitive Behavioral Testing Is Needed?
Cognitive-behavioral testing focuses on an individual's own cognition, behavior, and related physiological responses. The target is the problem behavior itself, not unconscious determinants or underlying causes. According to Kaplan & Saccuzzo (2001), these tests target 'disordered behavior' rather than the underlying cause. The analysis of this behavior is the goal of cognitive-behavioral assessment.
There are four steps involved in cognitive-behavioral assessment (Kaplan & Saccuzzo, 2001):
- Identification of critical behavior.
- Determining if the critical behaviors are in excess or deficits.
- Evaluation of the frequency, duration, or intensity of the behavior.
- Based on step 3, the frequency, duration, or intensity of the critical behavior is decreased (if in excess) or increased (if in deficit).
The Fear Survey Schedule (FSS):
The Fear Survey Schedule (FSS) is a self-report procedure used for clinical purposes. Primarily, it takes ratings on fear using a rating scale. Initially introduced by Akutagawa (1956) with 50 items, it now has versions with 50 to 122 items, employing 5-point or 7-point scales. The items involve fear-provoking situations and avoidance behaviors, derived from clinical observations (Wolpe & Lang, 1964) and laboratory studies (Geer, 1965).
Irrational Beliefs Test (IBT):
People often hold irrational beliefs that lack a logical or realistic basis. The Irrational Beliefs Test (IBT), developed by R.A. Jones (1968), is one such test. It contains 100 items using a 5-point scale format. The subject indicates their level of agreement or disagreement with each item. Half of the items pertain to the presence, and the other half to the absence, of particular irrational beliefs.
🔑 Definition — Irrational beliefs: Beliefs that do not have a logical or realistic basis but are held by a person and cannot be separated from their cognitive system.
Kanfer and Saslow’s Functional Approach:
Kanfer and Saslow (1969) played a lead role in initiating the cognitive-behavioral approach with their functional approach, a behavior-analytic approach. It focuses on excesses and deficits in people's behavior, rather than traditional psychopathological labels. According to Kaplan & Saccuzzo (2001), "a behavioral excess is any behavior or class of behaviors described as problematic by an individual because of its inappropriateness or because of excesses in its frequency, intensity, or duration." Conversely, "behavioral deficits are classes of behavior described as problematic because they fail to occur with sufficient frequency, with adequate intensity, in appropriate form, or under socially expected conditions."
This approach proposes that the same laws operate in the development of normal and disordered behaviors, with the difference being only in extremes. The psychologist first identifies the excesses and deficits and then helps the client decrease or increase them accordingly.
💡 Why this matters: This approach shifts the focus from diagnostic labels to observable, modifiable behaviors, which is crucial for designing effective behavioral interventions.
⭐ Key Takeaways
A student must understand that behavioral assessment provides direct, observable evidence of behavior to supplement traditional tests, using methods like observation and self-monitoring. The recording of behavior can be done via narrative, audio/video, or with specific behavioral rating scales. The lecture distinguishes cognitive-behavioral assessment, which targets the disordered behavior itself through a four-step process of identification, analysis, and modification. Key tools include the Fear Survey Schedule (FSS) for fear-provoking situations, the Irrational Beliefs Test (IBT) for irrational cognitions, and Kanfer and Saslow's Functional Approach, which analyzes behavior in terms of excesses and deficits rather than psychopathological labels.
🧠 Quick Revision Questions
- What is the fundamental difference between what traditional personality tests measure and what behavioral assessment measures?
- What are the two types of self-report methods in behavioral observation, and how do they differ?
- According to the lecture, what are the four steps involved in cognitive-behavioral assessment as outlined by Kaplan & Saccuzzo (2001)?
- What is the primary purpose of the Fear Survey Schedule (FSS), and from where were its items derived?
- In Kanfer and Saslow's Functional Approach, what are the two key categories into which all problem behaviors are analyzed, and what is the goal for each of these categories in treatment?
📘 Lecture 38 — Testing and Assessment in Health Psychology
📖 Overview: This lecture explores key psychological assessment tools used in health psychology, focusing on measures of anxiety, coping, social support, locus of control, and self-efficacy. Understanding these instruments is crucial for evaluating patient responses to health conditions, stressors, and treatment adherence in clinical and research settings.
🗂️ Topics Covered
This lecture covers the State-Trait Anxiety Inventory (STAI) for measuring anxiety types, the Ways of Coping Scale and Coping Inventory for assessing stress management strategies, and the Social Support Questionnaire (SSQ) for evaluating support networks. It also examines several scales for specific health-related locus of control (drinking, weight, perceived behavioral control, desired control, and health engagement control strategies). Finally, it provides an extensive overview of perceived self-efficacy measures for specific health behaviors including exercise, nutrition, habit cessation, and health-protective behaviors.
📝 Lecture Summary
Testing and Assessment in Health Psychology
Health psychology is a popular field with practical relevance to most people's lives. As research in health-related issues grows, the number of available assessment tools also increases. This section discusses tests and scales that psychologists use when assessing clients or conducting health psychological research.
The State-Trait Anxiety Inventory (STAI)
The STAI is based on Charles D. Spielberger's State-Trait Anxiety theory. The theory assumes anxiety is of two types: state anxiety, which is an emotional reaction varying from situation to situation, and trait anxiety, a personality characteristic stable across situations. The inventory yields two scores: A-State and A-Trait. The STAI uses a 4-point scale format with 20 items per scale (40 items total). It has been used with patients suffering from various health conditions and undergoing surgical or stressful procedures for anxiety assessment.
🔑 State Anxiety: An emotional reaction that varies from situation to situation 🔑 Trait Anxiety: A personality characteristic stable across situations
The Ways of Coping Scale
The Ways of Coping Scale (Lazarus, 1995; Folkman & Lazarus, 1980) assesses how people cope with stress. It is one of the most popular tools in health psychology. It is a checklist where subjects indicate items/thoughts and behaviors that apply to them. It contains 68 items and seven subscales: problem solving, growth, wishful thinking, advice seeking, minimizing threat, seeking support, and self-blame. Research suggests the subscales divide into two broad categories: Problem-focused strategies (cognitive and behavioral strategies for coping, attempting to solve the problem) and Emotion-focused strategies (ways of dealing with emotional response to stress, not resolving the problem).
Coping Inventory
The Coping Inventory (Horowitz & Wilner, 1980) contains items derived from clinical interview data. Its 33 items fall into three categories: activities and attitudes people adopt for avoiding stress, strategies for working through stressful events, and socialization responses.
The Social Support Questionnaire (SSQ)
The SSQ (Sarason et al., 1983) measures social support and related aspects. It contains 27 items, each having two parts. For every item, respondents endorse two things yielding two scores: the Number (N) score, listing people the respondent can count on for support (average calculated from all 26 items), and the Satisfaction (S) score, indicating overall satisfaction with supports (ranging from 1 = very dissatisfied to 6 = very satisfied, average calculated from all items).
Scales for Specific Health Conditions Related Locus of Control
a. Drinking Locus of Control Scale: A 25-item scale measuring drinking locus of control (Donovan & O'Leary, 1978), using a forced choice format involving pairing of internal and external control alternatives.
b. Weight Locus of Control (WLOC) Scale: Assesses internal and external determinants of one's weight (Saltzer, 1982), using a 6-point Likert scale format with 4 items.
c. Perceived Behavioral Control Measure: Developed by Armitage and Connor (1999) to assess perceived behavioral control, including items like "Whether or not I eat a low fat diet is entirely up to me."
d. Desired Control Scale: A 70-item rating scale (Reid & Zeigler, 1981) using a 5-point response scale from 'strongly agree' to 'strongly disagree'. It has two subscales with 35 items each: Desire of outcomes and Beliefs and attitudes.
e. Health Engagement Control Strategies (HECS): A 5-point rating scale comprising 9 items (Worsch, Schulz, & Heckhausen, 2002). Rating options range from "almost never true" to "almost always true," with items like "I invest as much time and energy as possible to improve my health."
Assessment of Perceived Self-Efficacy Pertaining to Specific Health Behaviors
While the GSE (General Self-Efficacy) scale predicts general personal competence across situations, numerous studies use self-efficacy for assessing its impact on specific health behaviors. These include healthy lifestyles, avoiding/quitting unhealthy behaviors, and coping with specific health conditions. Unlike GSE, specific health behavior measures focus on the health condition alone rather than global ability to handle stressful situations. Researchers can replace original items with health-specific ones or devise similar measures. Many studies use brief scales of only 4-5 items, sometimes even single-item measures. Items must be theory-based with appropriate wording, clearly mentioning both the health action and the perceived barrier. The semantic structure recommended is: "I am certain that I can do XX, even if YY (barrier)" (Luszczynska & Schwarzer, 2005).
Specific Health-Condition Related Self-Efficacy Measures
a. Measures for Assessing Exercise-Related Self-Efficacy: The exercise self-efficacy scale (Schwarzer & Renner, 2000) focuses on overcoming barriers to exercising, with a 4-point format from 'definitely not' to 'exactly true'. The Exercise Regularly Scale (Lorig et al., 1996) assesses confidence in regular physical activities with ten response options from 'not at all confident' to 'totally confident'.
b. The Nutrition Self-Efficacy Scales: The Nutrition Self-efficacy Scale (Anderson, Winett, & Wojcik, 2000) offers ten response options from 'very sure I cannot' to 'very sure I can', assessing confidence in using nutritious foods. Schwarzer and Renner's (2000) Nutrition Self-efficacy Scale uses 4 response options from 'definitely not' to 'exactly true', assessing ability to overcome barriers to healthy eating.
c. Habit Cessation and Abstinence Self-Efficacy: The Situational Confidence Questionnaire (Annis, 1987) examines alcohol abstinence self-efficacy using a 6-point scale with response options in percentages (0% to 100%), where 0% = 'not at all confident' and 100% = 'very confident' in resisting heavy drinking. Schwarzer and Renner's (2000) scale assesses control over drinking behavior with 4 response options from 'definitely not' (1) to 'exactly true' (4). The Smoking Cessation Self-efficacy Scale (Dijkstra & De Vries, 2000) uses a 7-point scale from -3 ('not at all sure I am able to') to +3 ('very sure I am able to').
d. Health-Protective Behaviors and Adherence to Medical Advice Self-Efficacy: The Preaction BSE Self-Efficacy Scale (Luszczynska & Schwarzer, 2003) examines a woman's ability to perform regular breast self-examination (BSE) despite odds, with 5 response options from 'definitely not' to 'exactly true'. The Maintenance BSE Self-Efficacy Scale uses the same options to gauge self-efficacy in maintaining regular BSE habit. The Adherence Self-efficacy Scale (Mohr et al., 2001) assesses self-efficacy related to self-injection, with response options from 1 ('I will not have any problems') to 6 ('I will not be able to tolerate it at all').
💡 Why this matters: These specific self-efficacy measures allow researchers and clinicians to target particular health behaviors (exercise, nutrition, smoking cessation, medical adherence) rather than using general measures, enabling more precise assessment and intervention planning.
⭐ Key Takeaways
Students must remember the distinction between state and trait anxiety in the STAI and that it yields separate A-State and A-Trait scores. The Ways of Coping Scale distinguishes problem-focused from emotion-focused coping strategies, a critical framework in health psychology. The SSQ provides both Number (N) and Satisfaction (S) scores for social support assessment. For specific health behaviors, self-efficacy measures must be theory-based with wording including both the health action and the perceived barrier using an "if-then" format. Finally, multiple specific measures exist for exercise, nutrition, smoking cessation, alcohol abstinence, and medical adherence self-efficacy, each with different scale formats and response options.
🧠 Quick Revision Questions
- What are the two types of anxiety measured by the STAI, and how do they differ?
- List the seven subscales of the Ways of Coping Scale and explain the difference between problem-focused and emotion-focused coping strategies.
- What two scores does the Social Support Questionnaire (SSQ) produce, and how is each calculated?
- What is the recommended semantic structure for wording items in specific health behavior self-efficacy measures?
- Name three specific health conditions or behaviors for which researchers have developed self-efficacy scales, and identify the response format for one of them.
📘 Lecture 39 — Measuring Personal Characteristics for Job Placement
📖 Overview: This lecture explores psychological methods for assessing personal characteristics to determine job suitability. It presents two major theoretical approaches — Osipow’s Trait Factor Approach and Roe’s Career-choice Theory — and describes the assessment tools derived from each. Understanding these approaches helps in matching individuals to careers based on their traits, interests, and upbringing.
🗂️ Topics Covered
The lecture begins by asking what factors people consider when choosing a job, such as personal interest, aptitude, skills, workplace setting, and colleagues. It then introduces two key theories: Osipow’s Vocational Dimensions Approach (Trait Factor Approach), which uses a battery of tests like the Kuder Occupational Interest Survey and Strong-Campbell Interest Inventory; and Roe’s Career-choice Theory, which emphasizes person vs. nonperson orientation and the influence of childhood family experiences on career choice, including the California Occupational Preference Survey and the concept of two independent continua.
📝 Lecture Summary
Osipow’s Vocational Dimensions Approach: The Trait Factor Approach
One psychologist best known for the use of the trait factor approach for job decision making is Samuel Osipow. This approach has a global perspective, considering a number of traits or aspects of a person’s personality. A battery of tests is used for assessment. The battery includes a variety of tests such as the Kuder Occupational Interest Survey (Kuder, 1979), Strong-Campbell Interest Inventory (Campbell, 1974), Seashore Measure of Musical Talents (Lezak, 1983), and Purdue Pegboard (Fleishman & Quaintance, 1984). This approach provides quite comprehensive information regarding the traits and interests of the person. However, it is criticized for not taking much into account the work environment. Nevertheless, this approach is found to be very useful in helping people make occupational decisions.
💡 Why this matters: The Trait Factor Approach offers a structured, multi-test method for career counseling, but its limitation is ignoring workplace context, so it should be supplemented with environmental assessments.
Roe’s Career-choice Theory: The California Occupational Preference Survey
The core feature of Roe’s theory is its emphasis on ‘person’ or ‘nonperson’ orientation found in people. According to Roe, this orientation plays a significant role in people’s career choice. In simpler terms, whether one likes to be with other people or not affects one’s career choice. The person/people-oriented people would be looking for jobs where they are in contact with other people, e.g., Arts, entertainment, or other services. The individuals who are not people-oriented would be seeking jobs that involve little interpersonal contact, e.g., lab work, science and technology, field exploration, etc. Roe drew some very interesting conclusions from extensive examination of the personalities of scientists working in different areas of study. Roe proposes that career choices people make in life are a result of their childhood experiences of relationship with their families. According to Roe, whether people, as children, were reared in a warm family environment or a cold and aloof one determines if they are interested in other people or not. Children brought up in a warm and accepting family environment grow into people-oriented adults. On the other hand, children who experienced a cold and aloof environment turn into adults who are interested in things rather than people (Roe & Klos, 1969; Roe & Siegelman, 1964). Roe and Klos (1969) proposed the idea that occupational roles can be divided into two classes according to two independent continua. The first continuum goes from “orientation to purposeful communication” to “orientation to resource utilization”. The second continuum goes from “orientation to interpersonal relations” to “orientation to natural phenomena”. People make career choices according to where they stand on these two continua.
🔑 Definition — Person/Nonperson Orientation: A core concept in Roe’s theory where a person’s preference for being with other people (person-oriented) versus being with things or alone (nonperson-oriented) significantly influences career choice.
📐 Formula: Career choice = f(Childhood family environment → Person/Nonperson orientation → Position on two continua: [Purposeful Communication ↔ Resource Utilization] + [Interpersonal Relations ↔ Natural Phenomena]) → Occupational role
📌 Example: A child raised in a warm, accepting family environment is likely to become a people-oriented adult. This person would be positioned closer to “orientation to interpersonal relations” on the second continuum and might choose a career in Arts or Entertainment (e.g., a theater director). In contrast, a child raised in a cold, aloof environment is likely to become a nonperson-oriented adult, positioned closer to “orientation to natural phenomena” on the second continuum, and might choose a career in science and technology (e.g., a research physicist working in a lab with minimal interpersonal contact).
⭐ Key Takeaways
This lecture introduces two major theoretical approaches for matching personal characteristics to jobs. Osipow’s Trait Factor Approach emphasizes using a comprehensive battery of tests (e.g., Kuder, Strong-Campbell, Seashore, Purdue Pegboard) to assess multiple traits, though it is criticized for neglecting the work environment. Roe’s Career-choice Theory focuses on the person vs. nonperson orientation, arguing that career choices are shaped by childhood family experiences — warm families produce people-oriented adults, while cold families produce thing-oriented adults. Roe further divides occupational roles using two continua: purposeful communication vs. resource utilization, and interpersonal relations vs. natural phenomena. For the exam, remember the names of the psychologists (Osipow and Roe), the specific tests in Osipow’s battery, and the key terms of Roe’s theory including the two continua.
🧠 Quick Revision Questions
- What is the main criticism of Osipow’s Trait Factor Approach?
- List four tests included in Osipow’s battery for vocational assessment.
- According to Roe, how does childhood family environment influence career choice?
- What are the two independent continua proposed by Roe and Klos (1969)?
- According to Roe, what type of career would a nonperson-oriented individual likely choose?
📘 Lecture 40 — Achievement and Educational Tests
📖 Overview: This lecture explores the use of psychological tests in educational settings for various purposes including achievement assessment, diagnosis, selection, and screening. It focuses primarily on achievement tests, discussing their design, types, and key considerations like what, how, and when to assess, while also comparing teacher-made tests with standardized tests and clarifying the distinction between achievement and aptitude tests.
🗂️ Topics Covered
The lecture covers the various purposes of testing in educational settings, then focuses on achievement tests—their definition and the "What, How, and When" decisions. It discusses teacher-made tests, including objective/forced-choice items and subjective/descriptive tests, and their respective advantages and disadvantages. Other varieties of standardized achievement tests (SAT-I, GRE, MAT) are reviewed, followed by a comparison of achievement versus aptitude tests and interpretive issues like grading, percent scores, and percentile ranks.
📝 Lecture Summary
Achievement Tests
Achievement tests are meant to assess if students or trainees have learned whatever they were supposed to learn at the end of a course or program of instruction. These tests measure the students' achievement alongside the effectiveness of a program.
What, How, and When of Achievement Tests
The assessment of achievement involves three basic decisions:
What Is To Be Assessed? This decision pertains to: a. The course content to be covered by the assessment tool. b. The instructional objectives that specify the expected and desired outcomes of the teaching-learning process.
How to Assess? This decision pertains to: a. The type of the assessment tool. b. The administration procedure. c. The number, format, and difficulty level of test items.
When to Assess? This decision involves answers to these questions: a. At what time during the academic session will the assessment take place? b. Once in a term, or more than once? This decision will affect the content area to be covered in assessment.
💡 Why this matters: Of all the above issues and decisions, the most significant is to cover in the test the content area that the students have been taught, keeping in mind the objectives specified for every component of the content.
Teacher Made Achievement Tests
Teacher made achievement tests are the most common type of achievement tests. Teachers have a choice to design and develop their tests the way they like them to be. A teacher made test can be either objective or subjective. On occasions it may be a combination of both.
Objective or Forced Choice Type of Items
These items are difficult to develop but easy to score. They allow the teacher to cover a wide range of content area. There is always a chance of selecting the right answer simply by guessing, but this can be controlled. If the items are MCQs with 4-5 options per item, and the options are carefully developed, then guessing can be controlled to a large extent. Another advantage is uniformity of scoring across examiners—no matter who does the scoring, the students will receive the same score. The essential requirement for availing these benefits is care in writing test items. The stem of every item should be clearly stated, should not be ambiguous, and should convey what the examiner wants to convey. Even more important is the selection of appropriate response options. A good MCQ item is the one in which every option appears to be the right answer, so only those who know the course content can select the right choice.
Subjective or Descriptive Tests
Subjective or descriptive tests also have their advantages. The nature of the items allows the examiner to test in-depth knowledge of the students. However, marking and evaluation of such examiner papers may be problematic. The inter-examiner uniformity of scoring is doubtful in such tests. Examiners' personal or ideological biases may interfere with the objectivity in evaluation required of a just teacher. Ultimately, it is for the examiners to decide what format they prefer to use and what would suit best to the course content.
Other Varieties of Achievement Tests: Standardized Achievement Tests
Other than teacher-made achievement tests, a variety of standardized achievement tests are used at national and international level.
The Scholastic Assessment Test (SAT-I): Previously known as Scholastic Aptitude Test or SAT, first used in 1926, this is the most commonly used college entrance test in the U.S. SAT-I has two parts: Verbal Reasoning and Mathematical Reasoning tests. SAT-II is also available.
Graduate Record Examination Aptitude Test (GRE): GRE is one of the most well-known tests across the globe and is the most commonly used graduate-school entrance test. GRE measures general scholastic ability and contains three sections: Verbal (GRE-V), Quantitative (GRE-Q), and Analytic (GRE-A).
Miller Analogies Test (MAT): MAT is the second major, widely used, scholastic aptitude test. It is a verbal test that measures a student's ability to find logical relationships for 100 different analogy problems.
🔑 Definition — Achievement tests: Tests meant to assess if students or trainees have learned whatever they were supposed to learn at the end of a course or program of instruction.
🔑 Definition — Teacher made achievement tests: The most common type of achievement tests, designed and developed by teachers, which can be objective, subjective, or a combination of both.
🔑 Definition — Objective/Forced Choice items: Items that are difficult to develop but easy to score, allow covering a wide range of content, and provide uniformity of scoring across examiners.
🔑 Definition — Subjective/Descriptive tests: Tests where the nature of items allows testing in-depth knowledge, but marking and inter-examiner uniformity of scoring may be problematic.
Achievement versus Aptitude Tests
A common question arises: what is the difference between achievement tests and aptitude tests? In most situations these terms are used interchangeably. If one analyzes logically, it is not possible to cover in one test all of the content that students from different institutions, regions, and countries have studied. Kaplan and Saccuzzo (2001, p. 343) have given a very good comparative description of the features of the two types of tests.
Grading, Percent Score, And Related Interpretive Issues
School/college tests usually use the grading system. Scores are also given in terms of percent. Grades make it easier to understand the relative position of students. In large scale tests, like GRE or GAT, the results are communicated in terms of percentile ranks. These describe a candidate's position in relation with those scoring above as well as those scoring below him or her.
⭐ Key Takeaways
Achievement tests are designed to assess what students have learned in a program of study, and their construction requires careful decisions about content, instructional objectives, assessment format, and timing. Teacher-made tests can be objective (forced-choice) or subjective (descriptive), each with distinct trade-offs: objective items are harder to construct but easier to score and ensure inter-examiner reliability, while subjective items allow depth but risk evaluator bias. Standardized tests like SAT-I, GRE, and MAT assess general scholastic ability and are widely used for college and graduate admissions. The distinction between achievement and aptitude tests is often blurred in practice, and results are typically interpreted using grades, percent scores, or percentile ranks. The most critical principle is that tests must align with the content taught and specified instructional objectives.
🧠 Quick Revision Questions
- What are the three basic decisions involved in the assessment of achievement?
- What are the advantages and disadvantages of objective/forced-choice test items?
- What are the advantages and disadvantages of subjective/descriptive test items?
- Name three standardized achievement tests discussed in the lecture and briefly describe what each measures.
- How are results typically communicated in large scale tests like GRE or GAT, and what do these measures describe?
📘 Lecture 41 — Multicultural Testing
📖 Overview: This lecture explores the challenges and solutions in psychological testing when test-takers come from diverse cultural backgrounds. It examines how cultural factors can introduce bias into standardized tests and presents specific tests designed to be culture-free or culture-fair, ensuring equitable assessment across different populations.
🗂️ Topics Covered
The lecture introduces the concept of multicultural testing and why it is necessary when working with diverse populations. It details four key factors that can cause cultural bias in testing: language, reading ability, speed, and familiarity with test formats. Finally, it describes four specific multicultural tests designed to minimize or eliminate cultural bias: the Leiter International Performance Scale-Revised, Raven Progressive Matrices, Goodenough-Harris Drawing Test, and the IPAT Culture Fair Intelligence Test.
📝 Lecture Summary
Multicultural Testing
Psychologists often work with subjects or clients from diverse cultural backgrounds, where cultural background may interfere with test performance. There is a need for tests that can be used with all people and that are neither biased against nor in favor of any specific cultural origin. Multicultural testing refers to tests and testing procedures that are not affected by the cultural background of the test taker. Such situations arise when immigrants from different cultural backgrounds settle in developed countries and need to be tested on the same variables using the same tests. Examples include measurement of IQ or personality, screening or selection for jobs, and diagnosis of maladjustment or mental illness. Even people belonging to subcultures within a larger society may experience cultural disadvantage. The content, administration, scoring procedures, or scores of multicultural tests are not affected by the cultural origin of the test taker. Multicultural testing is also known as transcultural testing or cross-cultural testing.
🔑 Definition — Multicultural Testing: Tests and testing procedures that are not affected by the cultural background of the test taker.
Factors That May Cause Cultural Bias
Cultural differences become significant when people have to take psychological tests developed in cultures other than their own. Anastasi and Urbina (2007) describe four parameters along which cultural differences may be found: a. Language: People are at a disadvantage if they can only use the language of their own culture, not the one used in the culture where they must adjust. They will be handicapped if psychological tests are administered in an unfamiliar language. b. Reading Ability: People are still at a disadvantage if they cannot read. Most tests require a certain level of reading ability. People may be familiar with the language the test is designed in, but they will remain handicapped if they cannot read the test items. c. Speed: The speed required for completing a test may also cause problems. In some cultures, life is fast and people are familiar with a sense of urgency to meet deadlines. In other cultures, the tempo of life is slower (e.g., rural and agricultural societies), and people are used to patiently waiting rather than striving for immediate outputs. Persons from such cultures may find it difficult to cope with speed-based tests. d. Familiarity With The Format, Style, And Contents Of Tests: People may not be familiar with certain forms of test items and formats. They may find it hard to attempt certain types of items, take longer than allocated time, and make mistakes because they cannot understand what they are supposed to do. For example, people may find it difficult to attempt MCQ-type questions if they have not seen such items previously. Problems involving figures for assessing spatial reasoning may be a totally new experience for test takers who have never seen geometric drawings.
Multicultural tests generally do not involve reading or writing, verbal ability, or test-taking speed. Familiarity with format, style, and contents is controlled by using performance and drawing-based items, and avoiding the use of designs and patterns that most people cannot relate to.
🔑 Definition — Cultural Disadvantage: The disadvantage experienced by certain segments of the population because the nature of many standardized tests favors specific cultural backgrounds. 💡 Why this matters: Understanding these factors is crucial for identifying when a test score reflects genuine ability versus cultural unfamiliarity.
Some Multicultural Tests
The Leiter International Performance Scale- Revised (LIPS-Revised)
The LIPS-Revised (Roid & Miller, 1997) is an individually administered test measuring intellectual ability. It was first developed in 1940 and has undergone many revisions. It can be used with all age groups from 2 to 20 years. Its 1997 version was standardized on a sample of 2000, both atypical and normal subjects from the U.S. Special features:
- Does not involve verbal instructions.
- Individually administered, following a difficulty level sequence (easiest item first).
- No time limit.
- Easels are used to present graphic stimulus materials; picture cards are placed in a response tray.
The scale covers four domains: Reasoning, Visualization, Attention, and Memory. Tasks for Reasoning and Visualization include matching, form completion, design analogies, sequential ordering, paper folding, figure rotation, and classification. Tasks for Attention and Memory include sustained and divided attention measures, and immediate and delayed memory tasks.
Raven Progressive Matrices
One of the most popularly used nonverbal and culture-free tests of general intelligence is the Raven Progressive Matrices (RPM). It uses a multiple-choice format. In each test item, the subject is asked to identify the missing element that completes a pattern. The test can be administered to groups or individuals from 5 years old to older adults. There are 60 matrices with a missing part presented in graded difficulty. The subject selects the appropriate pattern from a group of eight options.
Goodenough-Harris Drawing Test
The Goodenough-Harris Drawing Test (G-HDT) is the quickest, easiest, and least expensive nonverbal test for measuring intelligence. The subject is asked to draw a whole human figure. The test is scored for each item included in the drawing. The subject gets credit for inclusion of elements such as individual body parts, proportion, perspective, and clothing details. The scoring follows the age differentiation principle; older children tend to get more points because of greater accuracy. It is not a test of drawing or artistic skill, but of the development of conceptual thinking and accuracy of observation. In the revised scale, the subject is asked to draw a picture of a man, a woman, and of one's own self. The self-scale is used as a projective test of personality. Scores on the G-HDT can be related to Wechsler IQ scores. The test can be more appropriately used in combination with other tests of intelligence.
🔑 Definition — Age Differentiation Principle: The principle that older children tend to get more points on the Goodenough-Harris Drawing Test because of greater accuracy in their drawings. 📐 Key Relationship: G-HDT scores → can be related to Wechsler IQ scores.
IPAT Culture Fair Intelligence Test
R. B. Cattell directed the development of this test. The IPAT Culture Fair Intelligence Test is a paper-pencil test for three levels:
- Age levels 4-8 years and mentally disabled adults.
- Age levels 8-12 and randomly selected adults.
- High-school age and above-average adults.
⭐ Key Takeaways
The core of this lecture is that cultural background can introduce significant bias into psychological testing, rendering scores invalid for individuals from diverse backgrounds. Four key factors causing bias are language, reading ability, speed, and familiarity with test formats, and multicultural tests aim to minimize these by avoiding reading and writing, verbal ability, and time limits. The Leiter International Performance Scale-Revised is a nonverbal, untimed test for intellectual ability that uses pantomimed instructions. The Raven Progressive Matrices is a classic, culture-free test of general intelligence using pattern completion. The Goodenough-Harris Drawing Test is a quick, inexpensive measure of intelligence based on drawing a human figure, scored for detail and accuracy. Finally, the IPAT Culture Fair Intelligence Test by Cattell is designed to measure intelligence with minimal cultural influence across different age levels.
🧠 Quick Revision Questions
- What is the definition of multicultural testing, and what are three specific testing scenarios where it is needed?
- List and briefly explain the four factors that can cause cultural bias in psychological testing as described by Anastasi and Urbina.
- What are the four domains measured by the Leiter International Performance Scale-Revised, and what are three of its special features that make it culture-free?
- Describe the task in the Raven Progressive Matrices. How many matrices are there, and from how many options does the subject choose the correct pattern?
- What is the Goodenough-Harris Drawing Test measuring, and what is the age differentiation principle that governs its scoring?
📘 Lecture 42 — Adaptive Testing and Other Issues: Computer Based Administration
📖 Overview: This lecture explores innovative test administration procedures that adapt to individual test-taker characteristics, moving beyond one-size-fits-all testing. It covers two adaptive testing models (two-stage and pyramidal) and the role of computers in modern psychometrics. The lecture also introduces interviews as important alternatives to psychological tests, detailing their types, features, and required skills.
🗂️ Topics Covered
The lecture begins by acknowledging that test takers have different attitudes, motivations, abilities, and response characteristics. It then explains adaptive testing as a procedure that adjusts test item coverage according to individual response patterns, detailing the Two-Stage Adaptive Testing model and the Pyramidal Testing Model. The role of Computer Based Administration in scoring, item analysis, and large-scale testing is discussed. Finally, the lecture presents interviews as alternatives to psychological tests, covering their similarities with tests, four specific types (evaluation, structured clinical, case history, mental status, and employment), and essential interviewing skills for psychologists.
📝 Lecture Summary
Adaptive Testing
Psychologists have been working on tailor-making tests according to the individual response characteristics of test takers. The core idea is that people should not be at a disadvantage because of their specific response characteristics. Normally, test takers attempt all items with increasing difficulty, regardless of whether they answered easier items correctly. As a consequence, many people score worse than they could have if the test had been adapted to their individual response characteristics.
Adaptive testing refers to the testing procedure whereby test item coverage is adjusted according to the response characteristics of individual subjects.
🔑 Definition — Adaptive Testing: A testing procedure where the selection and presentation of test items is adjusted based on the individual response characteristics of each test taker.
📐 Two-Stage Adaptive Testing Model (Anastasi & Urbina, 2007): This model uses three measurement levels. A hypothetical test comprises 70 items in total. Ten items are placed in a routing test, while the remaining 60 items are divided into three measurement tests of 20 items each. These three measurement tests are of varying difficulty levels: easy, intermediate, and difficult.
📌 Example: All subjects first attempt the 10-item routing test, but not all of the other 60 items. Depending on their performance on the routing test, they take only one of the three measurement tests (easy, intermediate, or difficult). Therefore, everyone is given 30 items in total, but the last 20 items vary from person to person. If someone answers the difficult items correctly on the routing test, they get the 'difficult' measurement test. If someone only answers the easy items correctly, they get the 'easy' measurement test.
🔑 Definition — Pyramidal Testing Model: An alternative adaptive testing model where everyone begins with an item of intermediate difficulty level.
📌 Example: If a test taker answers the first intermediate item correctly, they are given the next item of a higher difficulty level. If they fail the first item, they are routed downward to an item of lower difficulty level. This procedure is repeated until the test taker manages to answer the desired number of items.
💡 Why this matters: In both models, test takers are treated according to their response pattern. While these adaptive testing procedures can be done manually with paper and pencil, they are quite tedious. Computerized adaptive testing is a convenient option that provides facility to psychologists, requiring only suitable software and the psychologist's skill.
Computer Based Administration
Like all other fields of life, computers play a very significant role in psychometrics. Availability of computers has facilitated psychologists in a number of ways, making things possible and easier whether it is computerized administration, scoring, item analysis, analysis of data obtained from large standardization samples, or adaptive testing. Computers are especially useful in group testing involving large numbers of participants or tests with tedious procedures.
Computers can be used for administration and scoring in cases such as multilevel batteries, different forms of educational testing, aptitude testing (on its own or for career guidance), and achievement testing.
📌 Example: Today, all major achievement and ability tests are administered and/or scored with the help of computers, e.g., GRE (Graduate Record Examination), SAT (Scholastic Assessment Test), MAT (Miller Analogies Test), GAT (Graduate Assessment Test), IELTS (International English Language Testing System).
Alternatives to Psychological Tests - Interviews as Assessment Tools
Psychological tests are just one form of instrument for assessment. Interviews are another important instrument, providing an opportunity for direct, face-to-face interaction with the person being examined. Interviews can be used as an alternative to psychological tests.
According to Kaplan & Saccuzzo (2001), tests and interviews share these common features:
- Method for gathering data
- Used to make predictions
- Evaluated in terms of reliability
- Evaluated in terms of validity
- Can be group or individual
- Can be structured or unstructured
Types of Interviews That Are Used For Assessment
🔑 Definition — Evaluation Interview: This interview helps the psychologist assess and understand why the student/client/individual has come to them.
🔑 Definition — Structured Clinical Interview: Structured interviews follow a fixed and set pattern of questions and procedures. This pattern may be decided by the clinic/hospital/institution or recommended by another agency, e.g., the use of DSM (Diagnostic and Statistical Manual of Mental Disorders) according to a sequence of steps.
🔑 Definition — Case History Interview: This interview is more detailed compared to other types as it aims at in-depth information. It usually takes a developmental approach and is relatively flexible though focused.
🔑 Definition — Mental Status Examination: This interview is more fixed and focused, and is used more commonly in psychiatric settings. It is typically used when some psychiatric, neurological, or emotional problem is suspected.
🔑 Definition — Employment Interview: These interviews are used by employers for the selection of right people for available jobs. Such interviews may be both structured and/or unstructured, depending upon the nature of the organization, the employer, and the position.
Interviewing Skills Required In Psychologists
- Practice and training
- Command over language and vocabulary
- Overcoming personal complexes
- Empathy
- Flexibility and acceptance of the other person's opinion
- Control over own emotional reactions
- Cultural sensitivity
- Note taking skills and technological assistance
⭐ Key Takeaways
Adaptive testing allows tests to be tailored to individual response characteristics using models like the two-stage (routing test followed by a difficulty-matched measurement test) or pyramidal (starting with an intermediate item and moving up or down based on performance). Computers have revolutionized psychometric testing by enabling efficient administration, scoring, and adaptive testing for major exams like the GRE and IELTS. Beyond formal tests, interviews serve as a valid alternative assessment tool, sharing key features with tests (reliability, validity, structure). Psychologists must master five specific interview types (evaluation, structured clinical, case history, mental status, and employment) and develop crucial skills including empathy, cultural sensitivity, and command of language.
🧠 Quick Revision Questions
- What is the fundamental purpose of adaptive testing, and what problem does it solve for test takers?
- Describe the two-stage adaptive testing model. How many items are in the routing test and each measurement test, and how is the measurement test assigned?
- Explain how the pyramidal testing model works, starting from the first item a test taker receives.
- List at least three major standardized tests that are now administered and/or scored with the help of computers.
- Name four of the eight interviewing skills required in psychologists, and describe the key differences between a case history interview and a mental status examination.
📘 Lecture 43 — Social and Ethical Considerations in Testing
📖 Overview: This lecture examines the social and ethical issues surrounding psychological testing, emphasizing the importance of adhering to professional codes of conduct. It explains multiple agencies and documents that provide ethical standards, and details specific ethical issues that test developers, users, and administrators must consider to protect test takers' rights and well-being.
🗂️ Topics Covered
The lecture begins by introducing the major ethical codes and agencies regulating psychological testing, including the APA Ethics Code, SIOP guidelines, and the Board on Testing and Assessment. It then provides a detailed breakdown of eight specific ethical issues: the training and eligibility of test users, human rights in testing, invasion of privacy, confidentiality and openness, careful labeling, divided loyalties, issues for test developers, and testing across diverse populations.
📝 Lecture Summary
Ethical Issues in Psychological Testing
A number of agencies have proposed ethical standards covering all aspects of testing, from test development to real-world application. These standards are essential because psychological testing involves significant social and psychological consequences. Most psychologists follow the APA standards (American Psychological Association), though some guidelines may also appear in test manuals, particularly regarding who may administer the test.
🔑 Definition — APA Ethics Code: A comprehensive document by the American Psychological Association covering confidentiality, development, and use of psychological assessment techniques, including legal and forensic contexts. 🔑 Definition — Principles for the Validation and Use of Personnel Selection Procedures: Guidelines developed in 1987 by the Society for Industrial and Organizational Psychology (SIOP) for assessment procedures used in hiring. 🔑 Definition — The RUST Statement: "Responsibilities of Users of Standardized Tests," adopted by the American Counseling Association (ACA) in 1989. 🔑 Definition — Board on Testing and Assessment (BoTA): A board established in 1993 in the U.S., supported by the Departments of Defense, Education, and Labor, primarily working on the use of psychological tests as tools of public policy.
1. The Training and Eligibility of the Test User
The person administering a test must be properly trained and experienced. The personality and qualifications of the psychologist are critical, though standards vary by institution or region. Academic and professional associations specify minimum qualifications. The administrator should have completed sufficient supervised training hours before working independently. Training is especially important for intelligence tests (especially individually administered ones) and personality tests, and becomes even more significant for interpreting projective tests. For achievement tests (especially objective ones), a compromise on qualifications may be acceptable. The APA Ethics Code (1992, p. 1599) states that psychologists should "provide only those services and use only those techniques for which they are qualified by education, training, or experience."
💡 Why this matters: Proper training prevents misinterpretation of test results, which could harm test takers through incorrect diagnoses or inappropriate recommendations.
2. Human Rights and Test Use
All individuals have the right to decide if they want to be tested or not. People should not be tested if they refuse. The only exception is legal or forensic scenarios where testing is directed by a court of law.
3. Invasion of Privacy
Confidentiality is essential in counseling and clinical encounters. Unless the test taker allows it, the psychologist must not disclose:
- That the person was tested
- The person's scores
- The interpretation of test results
- Any diagnosis
Legal situations are an exception. Also, for achievement tests, test results are generally understood to be made public.
🔑 Definition — Invasion of Privacy: The unauthorized disclosure of information about a person being tested, including the fact of testing, scores, interpretations, or diagnoses, without the test taker's consent.
4. Confidentiality, Honesty, and Openness
The psychologist must explain the nature and purpose of the test to the subject, and inform them about the possible use of test results. Test results cannot be used for purposes other than those mentioned to the subject. However, explaining the true nature of some tests beforehand can be problematic. For projective tests like Rorschach, TAT (Thematic Apperception Test), or WAT (Word Association Test), the required information may be provided immediately after the test is over, as foreknowledge could affect performance.
🔑 Definition — Informed Consent: The process of explaining the nature, purpose, and use of test results to the test taker before testing begins, ensuring they understand and agree.
5. Care with Labeling
Psychologists should avoid labeling unless essential. People may be diagnosed or show certain tendencies, but labeling should not occur unless it was the main objective of testing. Certain labels have social stigma attached, such as "schizophrenic" or PWA (Patient with AIDS). Such labels can damage the test taker socially, psychologically, and financially (e.g., being refused employment). Therefore, labeling should be avoided if possible.
💡 Why this matters: Labels can follow a person for life, affecting their job prospects, social relationships, and self-image, even if the original diagnosis or tendency is no longer relevant.
6. The Issue of Divided Loyalties
Psychologists may be hired and paid by an organization, but their profession demands care and protection of the test taker as well. When organizational and test taker interests clash, the psychologist must make rational decisions. According to Kaplan & Saccuzzo (2001), the following steps should be taken:
- Inform clients beforehand about the purpose of test results
- Inform clients about the limits of confidentiality
- Explain results to the client or their representative
- Provide the organization only the required information
- In adverse decisions, the person's right to know results should be preferred over test security
🔑 Definition — Divided Loyalties: A situation where a psychologist faces a conflict between the interests of the organization that hired them and the interests of the test taker, requiring careful ethical decisions.
7. Issues Pertaining To the Test Developers
Test developers must be fair and objective. They should use test content that is not gender biased or culturally biased. If the test is to be used with diverse populations, it should be culture free.
8. Issues Pertaining To the Test User In Diverse Populations
Test users or administrators must be fair to test takers. Tests developed and standardized in other cultures should not be used blindly with subjects from very different cultures. Tests must either be culture free, or translated, adapted, and standardized for the indigenous culture.
⭐ Key Takeaways
The core of ethical testing rests on several critical principles: test users must have proper training and qualifications, especially for intelligence and personality tests; test takers retain the right to refuse testing and must give informed consent; confidentiality must be maintained unless legally required or the test is an achievement test; labeling with stigmatizing terms must be avoided unless absolutely necessary; and tests developed in one culture cannot be blindly applied to another. The most important takeaway is that psychologists must balance their duty to the organization paying for testing with their ethical obligation to protect the test taker, and in cases of conflict, the test taker's right to know results takes precedence over test security.
🧠 Quick Revision Questions
- What are the four main documents/agencies that provide ethical standards for psychological testing?
- In which types of tests is the training and qualification of the test user most critical?
- Under what circumstances can a test be administered without the test taker's consent?
- What information must a psychologist NOT disclose without the test taker's permission?
- What is "divided loyalties" in psychological testing, and what should a psychologist do when it occurs?
📘 Lecture 44 — Assessment and Psychological Testing in Clinical & Counseling Settings
📖 Overview: This lecture examines the role of psychological testing within the broader context of clinical and counseling assessment. It covers specific tests and batteries used for neuropsychological evaluation, outlines the components and tools of behavioral assessment, and discusses the advantages and limitations of various assessment techniques.
🗂️ Topics Covered
The lecture first distinguishes psychological testing from psychological assessment, then lists common purposes of testing in clinical settings. It details neuropsychological testing, including specific tests like the Bender-Gestalt and batteries like the Halstead-Reitan. The lecture then explains behavioral assessment, covering self-report tools like the BDI, direct observation, and physiological measures, and concludes with a brief evaluation of assessment techniques.
📝 Lecture Summary
Assessment and Psychological Testing in Clinical & Counseling Settings
Psychological testing is a part of psychological assessment, which is more comprehensive. Assessment involves behavioral observation, interviews, and examination of case history in addition to testing. In clinical and counseling settings, tests can be used independently or as part of a complete assessment package. Tests are used for diagnosis, treatment induction, general assessment, and gauging recovery rates. All intelligence and personality tests may be used; for example, the House-Tree-Person (HTP) test can depict psychopathology. The Kaufman Test of Educational Achievement (K-TEA) is used for diagnosing specific learning disabilities.
Common purposes for testing include:
- General assessment of ability/IQ
- General assessment of personality
- Diagnosis of intellectual deficits
- Diagnosis of mental disorders
- Assessment of aptitude
- Neuropsychological assessment
- Assessment of learning disabilities
Neuropsychological Testing
This is a complicated area of psychological assessment, particularly regarding the diagnosis of brain damage. It may involve an extensive battery of tests assessing cognitive ability, verbal ability, spatial relations, and more. A single test is often not accurate enough, so test batteries are preferred.
Some tests used include:
- Bender-Gestalt test and Benton Visual Retention Test: These are commonly used but are more reliable as part of a battery.
- Halstead-Reitan Neuropsychological Test Battery (HRB) and the Luria-Nebraska Neuropsychological Battery: These are preferred because they provide information in a variety of areas.
🔑 Definition — Neuropsychological Testing: An area of psychological assessment that uses tests to evaluate cognitive, verbal, and spatial abilities, often to diagnose brain damage. 💡 Why this matters: A single test may not be accurate for diagnosing brain damage, so standardized batteries that measure multiple skills are preferred.
According to Anastasi & Urbina (2007), these batteries are useful because they:
- Provide measures of all significant neuropsychological skills.
- Can detect brain damage with a high degree of success.
- Help identify and localize impaired brain areas.
- Allow differentiation between particular syndromes associated with cerebral pathology.
Behavioral Assessment
Behavioral assessment procedures include self-report by the client, direct observation of behavior, and physiological measures.
Self-report by the client can take various forms, such as inventories, checklists, and clinical interviews. A commonly used tool is the Beck Depression Inventory (BDI). 🔑 Definition — Beck Depression Inventory (BDI): A self-report instrument where clients make self-ratings on 21 items to help assess the severity of depression. 📌 Example: A client completes the BDI by rating 21 items related to depressive symptoms, such as sadness, loss of pleasure, and changes in appetite, on a scale from 0 to 3. The sum of these ratings provides an overall score indicating the severity of depression (e.g., minimal, mild, moderate, or severe).
Another instrument is the Alcohol Use Inventory. Some instruments involve multiple informants. The Social Skills Rating System (SSRS) uses separate forms for parents, teachers, and students to evaluate positive and problematic behaviors in educational and family settings. The Behavior Assessment System for Children (BASC) is a comprehensive instrument that includes:
- Behavior rating scales for teachers and parents.
- A self-report questionnaire for children.
- A form for coding and recording classroom behavior.
- An additional structured interview for taking developmental history from parents.
Direct observation of behavior can be recorded by the psychologist, parents, or teachers in naturalistic settings. Observations may be recorded using narratives, checklists, rating scales, or record forms.
Physiological measures are used depending on the nature of the problem, especially in cases of anxiety or sleep disorders. These measures may include assessments of cardiovascular activity, cerebral functioning, electrodermal activity, and electro-ocular activity.
Evaluation of Various Assessment Techniques
All techniques have their advantages and disadvantages. Psychologists may choose any method that best serves their purpose and has minimum limitations.
⭐ Key Takeaways
Psychological testing is a component of the broader process of psychological assessment, which also includes behavioral observation, interviews, and case history review. In clinical settings, tests are used for diagnosis, treatment planning, and monitoring progress. For neuropsychological assessment, comprehensive test batteries like the Halstead-Reitan are preferred over single tests for their ability to detect and localize brain damage across multiple functional areas. Behavioral assessment relies on multiple methods, including self-report inventories (e.g., BDI), direct observation in natural settings, and physiological measures, each with its own strengths and limitations.
🧠 Quick Revision Questions
- What is the key difference between psychological testing and psychological assessment?
- Name two test batteries used for neuropsychological assessment.
- According to Anastasi & Urbina, what is one advantage of using a neuropsychological test battery over a single test?
- What are the three main procedures included in behavioral assessment?
- Provide one specific example of a self-report instrument and describe its purpose.
📘 Lecture 45 — Overview of the Course
📖 Overview: This is the final lecture of the series, providing a comprehensive review of all topics covered in the course on Psychological Testing and Measurement. It summarizes the fundamental concepts of psychological testing, including test types, reliability, validity, norms, and item analysis, emphasizing the importance of ethical standards in test use. This lecture serves as a crucial revision for students to consolidate their understanding of the entire course material.
🗂️ Topics Covered
This lecture reviews the entire course, beginning with the definition and purpose of psychological tests as tools for quantifying behavior. It covers the categorization of tests by purpose, administration, and structure, and their major contexts of use in education, occupation, and clinical settings. The lecture then summarizes the essential characteristics of good tests—validity, reliability, and norms—and details the specific methods for assessing reliability (test-retest, alternate-form, split-half, coefficient alpha) and validity (content, criterion, and construct). It concludes with a recap of norms, item analysis (difficulty and discrimination), and the future of psychological testing, emphasizing the need for ethical standards.
📝 Lecture Summary
What is a Psychological Test?
Psychology is the scientific study of behavior and mental processes. Psychologists use carefully designed tools for data collection, and psychological tests are one of those tools. In all professional settings where psychologists work, some form of testing and assessment is used.
🔑 Definition — Test: “A measurement device or technique used to quantify behavior or aid in the understanding and prediction of behavior” (Kaplan & Saccuzzo, 2001).
Tests are a part of the assessment package and procedure.
Types of Tests
Tests can be categorized on the basis of:
- The purpose or the type of behavior/characteristics to be measured: personality, aptitude, intelligence, achievement etc.
- The administration procedure: individual versus group tests
- Speed versus ability tests
- Aptitude tests, achievement tests, or intelligence tests
- Ability versus personality tests
- Structured/objective tests versus projective tests
- Original versus translated and adapted tests
- Translated tests
Major Contexts of Current Test Use
Psychological tests are designed for use in numerous life situations. The three major contexts of test use are:
- Educational testing
- Occupational testing
- Clinical and Counseling Psychology (Anastasi & Urbina, 1997).
Test Construction
The development of a good test takes place after going through a number of stages, keeping in view the established principles of test construction.
Essential Characteristics of Psychological Tests
A good psychological test should have these qualities:
- Validity: A test should measure what it is intended to measure.
- Reliability: A test should give consistent results. It should give the same or similar results every time it is administered to the same subjects in the same conditions.
- Norm development and standardization
Reliability
According to the definition given by Anastasi & Urbina (2007), reliability refers to “the consistency of the scores obtained by the same persons when they are reexamined with the same test on different occasions, or with different sets of equivalent items, or under other variable examining conditions.”
Reliability of a measure can be measured in a number of ways, such as the stability of scales over time or the consistency between items.
a) Test-Retest Reliability:
- Test-retest reliability deals with two performances of the same test by the same persons on two different occasions.
- If reliability refers to the consistency and stability of scores over time, then it will be measured using this method.
- The test-retest coefficient is also known as the ‘coefficient of stability’.
- The test takers’ scores on the first administration of the test are correlated with their scores obtained on the second administration of the same test.
b) Alternate-Form Reliability: In this approach, the test developer develops two alternate or parallel forms of the same test. Ideally, the two forms should be independently developed and completely parallel or equivalent forms of the same measure. They should match each other in all respects, including:
- Same specifications
- Same instructions
- Same time limit
- Same content
- Same number of items
- Same item format
- Same difficulty level
c) Split-Half Reliability: Alternate form reliability is popular but has some obvious limitations. Construction of two completely parallel forms is not easy and requires a lot of time and effort; even when two alternate forms are available, the practice effect and prior experience may affect performance on the second occasion. The time gap between two administrations is another intervening variable. To overcome these problems, split-half reliability is computed.
e) Coefficient Alpha:
- Another approach, quite similar to the Kuder-Richardson technique, is to calculate ‘coefficient alpha.’ The Kuder-Richardson formula can be used only for tests in which items are scored as either zero or one.
- It is not applicable to tests where answers to items are assigned two or more scoring weights, e.g., personality inventories where a number of response options with attached weights are available (never=0, occasionally=1, often=2, and so on).
- Coefficient alpha is a general formula that caters for such tests.
Validity
“Traditionally, the validity of a test has been defined as the extent to which a test measures what it was designed to measure” (Aiken, 1994, p.95). “The extent to which a test measures the quality it purports to measure. Types of validity evidence include content validity, criterion validity, and construct validity evidence” (Kaplan & Saccuzzo, 2001, p.640).
Validity is measured in several ways, divided into three categories:
Content Validity Evidence: “The evidence that the content of a test represents the conceptual domain it is designed to cover” (Kaplan & Saccuzzo, 2001, p.635).
Construct Validity Evidence: “A process used to establish the meaning of a test through a series of studies. To evaluate evidence for construct validity, a researcher simultaneously defines some construct and develops the instrumentation to measure it. In the studies, observed correlations between the test and other measures provide evidence for the meaning of the test” (Kaplan & Saccuzzo, 2001, p.635).
Criterion Validity Evidence: “The evidence that a test score corresponds to an accurate measure of interest. The measure of interest is called the criterion” (Kaplan & Saccuzzo, 2001, p.635).
Specific Procedures for Content Validity: A number of careful decisions are taken at the time of test development when test specifications are made. Test specifications include:
- Instructional objectives, relative importance of topics/processes.
- It should clearly indicate the number of items for each topic.
- It may also provide sample material. The process of content validation should include a description of all procedures in the manual that ensure the test is appropriate and representative.
Criterion-Related Validity:
- Whenever we plan and design a test we have a certain standard in mind that we want to meet.
- We relate the score on our test with those achieved on the criterion or standard.
- One approach to assessment of validity is through comparing the test results with those on a criterion, a procedure where scores on a test being used are correlated with scores on a criterion.
Criterion Prediction Procedures: Two procedures may be adopted for this purpose:
- Predictive validity evidence
- Concurrent validity evidence
🔑 Definition — Predictive Validity Evidence: “The evidence that a test forecasts score on the criterion at some future time” (Kaplan & Saccuzzo, 2001, p.638).
🔑 Definition — Concurrent Validity Evidence: “Evidence for criterion validity in which the test and the criterion are administered at the same point in time” (Kaplan & Saccuzzo, 2001, p.635).
Anastasi & Urbina (2007) have provided an all-encompassing description of construct validity: “The construct validity of a test is the extent to which the test may be said to measure a theoretical construct or trait.”
Assumptions Underlying Construct Validity:
- Test scores will be highly correlated with tests measuring the same or a similar construct. This is convergent validity.
- Test scores will have weak/low (or at times may be negative) correlation with tests meant to measure constructs which are different from the one measured by the main test. This is discriminant validity.
Norms
The scores on any test become meaningful in the presence of norms that have been developed for that test.
🔑 Definition — Norms: “The test performance data of a particular group of test takers that are designed for use as a reference for evaluating or interpreting individual test scores.”
Norms provide standards to which the results of the test takers on different measurements can be compared. Norms are the test performance of the standardization sample, a group of people whose performance on a specific test is taken as a standard or norm for comparison. All other individuals’ performance on this specific test is compared with the scores of this standardization sample.
Item Analysis
Item analysis is “a set of methods used to evaluate test items. The most common techniques involve assessment of item difficulty and item discriminability” (Kaplan & Saccuzzo, 2001, p.637). It may involve other things as well, such as item response valence or content analysis.
🔑 Definition — Item Difficulty: “A form of item analysis used to assess how difficult items are. The most common index of difficulty is the percentage of test takers who respond with the correct choice.”
Item Discrimination: A test is supposed to discriminate between those who know and those who do not know; those who score high and those who score low. “The item-discrimination index is a measure of the difference between the proportion of high scorers answering an item correctly and the proportion of low scorers answering the item correctly; the higher the value of d, the greater the number of high scorers answering the item correctly” (Cohen & Swerdlik, 1999).
Types of Tests
A variety of psychological tests were discussed in this course:
- Intelligence tests: individual and group tests, verbal and performance tests, speed and ability tests
- Personality tests: Personality inventories and projective tests
- Achievement tests: Ordinary school tests and standardized, internationally used tests
- Tests used for occupational settings
- Tests for special populations
- Tests used in clinical and counseling settings
- Aptitude tests
- Tests of interests
- Variations of psychological tests: the Piagetian tasks.
Future of Psychological Testing
- The application of psychological tests is becoming broader every day, as the areas of application of psychology are also increasing.
- BUT one should always keep the ethical standards in mind; these standards may have been set by the APA, the government, or any other professional and/or regulatory agency.
⭐ Key Takeaways
For the exam, a student must remember that a good psychological test must possess validity (measuring what it intends to), reliability (yielding consistent results), and established norms for meaningful score interpretation. The three main types of validity are content, criterion (predictive and concurrent), and construct (which involves convergent and discriminant evidence). Reliability can be assessed through test-retest, alternate-form, split-half, and coefficient alpha methods, each addressing different sources of measurement error. Finally, item analysis using indices of difficulty and discrimination is crucial for evaluating individual test items, and ethical standards must guide all test development and use.
🧠 Quick Revision Questions
- What are the three essential characteristics that a good psychological test must possess?
- Distinguish between test-retest reliability and alternate-form reliability.
- Define construct validity and explain the two assumptions (convergent and discriminant validity) that underlie it.
- What is the purpose of item analysis, and what are the two most common indices used in this analysis?
- List the three major contexts of test use as discussed in the course.