ENG520 — Midterm Summary (Lectures 1–22)
📘 Lecture 1 — Concept of Assessment-I
📖 Overview: This introductory lecture establishes the foundational concepts of educational measurement, assessment, and evaluation, clarifying their distinct yet interrelated purposes. It also introduces the classification systems for classroom assessment and the different types of assessments based on nature and format, providing a framework for understanding how student learning is measured and improved.
🗂️ Topics Covered
The lecture covers the definitions and differences between Measurement, Assessment, and Evaluation with a classroom example; the concept of Classroom Assessment and its four classifications by nature, format, use in instruction, and method of interpreting results; and the Types of Assessment by nature (Maximum Performance vs. Typical Performance) and by format (Fixed Choice vs. Complex Performance).
📝 Lecture Summary
Topic- 001: Measurement, Assessment and Evaluation
Measurement is the process by which the attributes or dimensions of some object (both physical and abstract) are quantified. The tools used for this purpose may include test, observation, checklist, homework, portfolios, project, etc. Measurement is easily understood when applied to physical objects like height or distance because these are physically present and can be directly measured with a scale. However, in education, variables are not physical and cannot be directly measured — e.g., attitude, behavior, and achievement — these are abstract, so their measurement is relatively more difficult than those with physical existence. The tool used for measuring abstract variables cannot measure exactly like a scale or thermometer. In this course, whenever the word "measurement" is used, it means the tool used for measuring student abilities that then converts them into numerical form.
Assessment means an appraisal of something to improve the quality of teaching and learning process for deciding what more can be done to improve the teaching, learning, and outcomes.
Evaluation is the process of making a value judgment against intended learning outcomes and behavior, to decide the quality and extent of learning. Evaluation is always related to your purpose — you align your purpose of teaching with what students achieved at the end, with their quality and quantity of learning.
🔑 Definition — Measurement: The process of quantifying attributes or dimensions of an object (physical or abstract) using tools, converting them into numerical form.
🔑 Definition — Assessment: An appraisal of something to improve the quality of the teaching and learning process.
🔑 Definition — Evaluation: The process of making a value judgment against intended learning outcomes and behavior to decide the quality and extent of learning.
📌 Example: In a classroom, when a teacher teaches a chapter or unit, first the objectives are made (either by the teacher or from a curriculum document). These Student Learning Outcomes (SLOs) are checked by two aspects: assessment and evaluation. When a teacher reflects daily on their teaching — whether yesterday's teaching was good, whether students understood, and then decides changes for improvement — this is ASSESSMENT. At the end of assessment, the opinion is not about the individual; it is about the process for the betterment of that process. Evaluation is when the learning process is complete and the teacher wants to see to what extent students achieved the targets or objectives — tools are made and measures come that tell how much students learned and the quality of their learning.
Difference in Measurement, Assessment and Evaluation
These terms are not different words for the same concept but represent a process, serving as prerequisites to each other and having unique purposes. Every mechanism or process starts with measurement. In education, measurements generate tools (test, observation, quiz, checklist, homework, portfolios). They can give information about students’ learning. The information obtained from the tool is MEASUREMENT. If that measurement is used for making the teaching-learning process better, then it is ASSESSMENT. EVALUATION is a systematic determination of a subject's merit, worth, and significance, using criteria governed by a set of standards. Assessment and evaluation do not exist in a hierarchy; they are parallel and different in purpose. Measurement is the source that provides the base and evidence to quantify the teaching-learning process. The quantified number has no meaning until we do an assessment or evaluation. Assessment's purpose is to make the teaching-learning process better so that student learning improves, and measurement's purpose is to align the learning with a purpose.
💡 Why this matters: Understanding the distinction between these three terms is critical because they are often used interchangeably in casual conversation, but they serve fundamentally different roles in educational practice. Measurement provides the raw data, assessment uses that data to improve the process, and evaluation makes final judgments about the quality of outcomes.
Topic – 002: Classroom Assessment
The process of gathering, recording, interpreting, using, and communicating information about a child's progress and achievement during the development of knowledge, concepts, skills, and attitudes is called classroom assessment.
When teaching in a classroom, four things are being done:
- Developing students' knowledge
- Improving their concepts
- Teaching the skills
- Making their attitudes
While doing so, information is collected about the development of these things, and then it is recorded, interpreted, used, and communicated about the learning progress of students. This whole procedure is called classroom assessment.
Classification of Assessment
Assessment can be classified in four ways:
- Nature of Assessment — Maximum Performance Assessment, Typical Performance Assessment
- Format of Assessment — Fixed Choice Assessment, Complex Performance Assessment
- Use in Classroom Instruction — Placement Assessment, Formative Assessment, Diagnostic Assessment, Summative Assessment
- Method of Interpreting Results — Norm Referenced Assessment, Criterion Referenced Assessment
🔑 Definition — Classroom Assessment: The process of gathering, recording, interpreting, using, and communicating information about a child's progress and achievement during the development of knowledge, concepts, skills, and attitudes.
Topic – 003: Types of Assessment
1. Types of Assessment by Nature
i. Maximum Performance Assessment
Maximum performance assessment determines what individuals can do when performing at their best, e.g., assessing students in an environment where they exhibit their best performance. The procedure is concerned with how well individuals perform when they are motivated to obtain as high a score as possible. This type includes aptitude tests and achievement tests. In an achievement test, students learn by themselves or are taught, and at the end, we want to see how much they learned against our target — so a test is made to determine their best abilities. It is designed to indicate the degree of success in any past learning activity. An aptitude test is used when we want to predict success in future learning activity, e.g., to see the interest of students in a particular field like medicine, sport, or teaching. Different abilities are needed for different professions, so a test is made depending on these abilities to assess in which abilities students perform well.
🔑 Definition — Maximum Performance Assessment: Assessment that determines what individuals can do when performing at their best, under motivated conditions, to obtain the highest possible score. Includes achievement tests (past learning) and aptitude tests (future learning prediction).
ii. Typical Performance Assessment
Typical performance assessment determines what an individual will do under natural conditions. This type includes attitude, interest, personality inventories, observational techniques, and peer appraisal. The emphasis here is on what students will do rather than what they can do.
🔑 Definition — Typical Performance Assessment: Assessment that determines what an individual will do under natural conditions, focusing on typical behavior rather than maximum capability.
2. Types of Assessment by Format
i. Fixed Choice Assessment
Fixed choice assessment is used to measure the skills of people efficiently (meaning measure more skills in less time) using fixed choice items: multiple choice questions, matching exercises, fill in the blanks, and true/false. It is called "fixed choice" because the person attempting the paper does not need to write the answer — they just need to choose the answer. From these, abilities of lower-level learning can be assessed. Fixed choice assessment is used for efficient measurement of knowledge and skills.
🔑 Definition — Fixed Choice Assessment: Assessment using items where the test-taker selects from given options (e.g., multiple choice, true/false, matching), used for efficient measurement of knowledge and lower-level skills.
ii. Complex Performance Assessment
Complex performance assessment is used for measurement of performance in contexts and problems valued in their own right. This includes hands-on laboratory experiments, projects, essays, and oral presentations. For example, if wanting to measure a student's ability to write an essay, this cannot be judged by fixed response items.
🔑 Definition — Complex Performance Assessment: Assessment used to measure performance in real-world contexts using tasks like experiments, projects, essays, and oral presentations, which cannot be measured by fixed-choice items.
💡 Why this matters: The classification of assessments by nature and format directly impacts test design. Maximum performance assessments (like final exams) require different construction than typical performance assessments (like behavioral checklists). Similarly, fixed-choice assessments are efficient for testing knowledge but cannot measure complex skills like writing or experimentation.
⭐ Key Takeaways
The most critical distinction from this lecture is that measurement is the quantification of student abilities into numerical form, assessment uses that measurement to improve the teaching-learning process, and evaluation makes value judgments about the quality of learning against objectives. Classroom assessment is defined by a five-step cycle of gathering, recording, interpreting, using, and communicating information about student development. Assessments are classified by nature (maximum vs. typical performance), format (fixed choice vs. complex performance), use in instruction (placement, formative, diagnostic, summative), and interpretation method (norm-referenced vs. criterion-referenced). Maximum performance assessment measures best performance under motivation (achievement and aptitude tests), while typical performance measures natural behavior (attitude and personality inventories). Fixed choice items efficiently measure lower-level skills, but complex performance assessments are required for higher-order skills like writing and experimentation.
🧠 Quick Revision Questions
- What is the key difference between assessment and evaluation in an educational context?
- List the four ways assessment can be classified according to this lecture.
- What is the difference between maximum performance assessment and typical performance assessment? Give one example of each.
- Name the two types of assessment by format and explain what kinds of skills each is best suited to measure.
- In the classroom example, which activity represents assessment and which represents evaluation?
📘 Lecture 02 — Concept of Assessment-II
📖 Overview: This lecture explores the different classifications and uses of assessment in classroom instruction. It details how assessment serves placement, diagnostic, formative, and summative purposes, and explains the critical distinction between norm-referenced and criterion-referenced methods for interpreting results. Understanding these concepts is essential for teachers to effectively design, implement, and utilize assessments to improve student learning.
🗂️ Topics Covered
The lecture covers three main topics: first, the use of assessment in classroom instruction for placement and diagnostic purposes; second, the distinction between formative and summative assessments during and after instruction; and third, the two primary methods of interpreting assessment results: norm-referenced and criterion-referenced assessment. Each topic explains the purpose, application, and key characteristics of the respective assessment type, supported by examples.
📝 Lecture Summary
Topic- 001: Use of Assessment in Classroom Instruction
This section classifies assessment based on its use, specifically for placement and diagnosis.
Placement Assessment is used at the beginning of instruction to determine a student's prerequisite skills, degree of mastery of course goals, and mode of learning. It assesses a student's entry-level performance to see if they have sufficient knowledge for a particular course. This helps the teacher know where a student should be placed according to their present knowledge or skills and plan the lesson accordingly. It also determines a student's interest and aptitude regarding a subject, helping in selecting the correct path for the future.
🔑 Definition — Placement Assessment: An assessment that determines a student's prerequisite skills, degree of mastery of course goals, and preferred mode of learning at the beginning of an instructional session.
📌 Example: Common examples include readiness tests (to determine if a student has the foundational concept for a course), aptitude tests (used for admission into a specific program), pretests (made according to course objectives to determine a student's present knowledge), and self-report inventories (which determine student level through interviews or discussion).
Diagnostic Assessment is used to determine the causes of persistent learning difficulties. A diagnosis does not start on the first day; it is for constant or continuous problems. For example, if a student continues to experience failure in reading or mathematics despite the use of prescribed alternative methods, then a diagnosis is indicated. The goal of diagnostic assessment is to find the root cause of a student's failure, which can be intellectual, physical, emotional, or environmental.
🔑 Definition — Diagnostic Assessment: An assessment used to identify the underlying causes (intellectual, physical, emotional, or environmental) of persistent and continuing learning difficulties in a student.
📌 Example: The lecture uses a medical analogy: if a headache persists after initial self-treatment and then a doctor's prescription, the doctor will order tests (e.g., blood test) to find the root cause. Similarly, in education, when a student continues to fail a subject like mathematics despite using alternative teaching methods, a diagnostic assessment is indicated to find the root of the failure.
💡 Why this matters: Placement helps teachers start instruction at the right level, while diagnosis provides targeted intervention for ongoing learning problems that standard instruction cannot solve.
Topic- 002: Use of Assessment in Classroom Instruction – Formative and Summative Assessments
This section explains the two main types of assessment based on when they are conducted during instruction.
Formative Assessment is conducted during the teaching-learning process. Its purpose is to determine learning progress, provide feedback to reinforce learning, and correct learning errors. It is an ongoing process used to modify teaching strategies based on student needs. It provides feedback to teachers about the weakness and strength of the learning process, allowing them to modify their teaching practices to improve the teacher-learning process. The main difference from summative assessment is that in formative assessment, the goal is improvement in the process of learning rather than to certify students. Common tools include teacher-made tests and custom-made tests from textbooks.
🔑 Definition — Formative Assessment: An ongoing process of evaluation conducted during instruction to monitor student learning progress, provide feedback, and modify teaching strategies to improve the learning process.
📌 Example: A teacher gives a short quiz halfway through a unit. The results show many students are struggling with a specific concept. The teacher then re-teaches that concept using a different method. The quiz and the subsequent adjustment of teaching are part of formative assessment.
Summative Assessment comes at the end of an instructional session, such as at the end of a course or unit. It is designed to measure the extent of achievement of intended learning outcomes. The primary utility is to assign grades and certify the level of mastery and expertise in a certain subject. It usually compares student learning either with other students' learning (norm-referenced) or a standard for a grade level (criterion-referenced). It can include teacher-made achievement tests or alternative techniques like portfolios. In a semester system, both midterms and final terms are considered summative assessments.
🔑 Definition — Summative Assessment: An assessment conducted at the end of an instructional period (unit, course, semester) to measure and certify the extent of student achievement of intended learning outcomes.
📌 Example: A final exam at the end of a semester that determines the student's final grade for the subject is a classic example of summative assessment. A final portfolio submitted at the end of a course to summarize overall performance is another example.
Topic- 003: Types of Assessment: Methods of Interpreting Results
This section distinguishes between two approaches to interpreting assessment results: norm-referenced and criterion-referenced.
Norm-referenced Assessment (NRT) measures a student's performance according to their relative position within a known group. It reports how a student stands among other students, not their actual achievement level. It is utilized to discriminate between different groups of students and is never used for certification or issuing grades (though it is used for selection like in NTS or CSS exams). The position of the student is generally represented by a percentile score. To achieve a large spread of scores for easy discrimination, a norm-referenced test includes items of average difficulty with high discriminating power, meaning items should not be very easy or very difficult.
🔑 Definition — Norm-referenced Assessment (NRT): A method of interpreting test results that measures a student's performance by comparing it to the performance of a specific group, reporting their relative standing (e.g., percentile rank) rather than their mastery of a subject.
📐 Formula: Percentile Score → A score indicating the percentage of students in the norm group who scored the same or lower. For example, a 97th percentile means 96% of students scored lower.
📌 Example: The National Testing Service (NTS) exam and Central Superior Services (CSS) exams are examples of norm-referenced tests. If a student achieves a 97% percentile on the NTS, it means 96% of the other test-takers scored lower than that student.
Criterion-referenced Assessment (CRT) describes student performance according to a specific domain of clearly defined learning tasks. It does not compare a student's performance with other students; instead, it compares all students' performance against a fixed criterion (e.g., learning outcomes). It is most commonly used in schools to report the achievement of learning outcomes. Student grades represent their mastery over content, and a cut point is determined to distinguish between failed and successful students.
🔑 Definition — Criterion-referenced Assessment (CRT): A method of interpreting test results that describes a student's performance by comparing it to a pre-defined standard, domain, or set of learning objectives, without comparing the student to others.
📐 Formula: Cut Score/Cut Point → A pre-determined score that separates those who have achieved mastery (e.g., passed) from those who have not (e.g., failed), based on the defined criteria.
📌 Example: A teacher sets a learning objective: "Students will be able to add single-digit whole numbers." The teacher then creates a test with 10 addition problems. A student who answers 8 correctly has demonstrated 80% mastery of the objective, regardless of how the rest of the class performed. The cut score might be 7/10 (70%) to pass.
⭐ Key Takeaways
A student must remember that assessment can be classified by its use (placement and diagnosis) and by its timing (formative and summative). Placement assessment is a pre-assessment to determine entry-level knowledge and skills, while diagnostic assessment identifies the root causes of persistent learning difficulties. Formative assessment is an ongoing, in-process evaluation used to improve teaching and learning through feedback, whereas summative assessment is a final evaluation to certify achievement and assign grades. Finally, the interpretation of assessment results can be norm-referenced, comparing a student's performance to a group, or criterion-referenced, comparing it to a fixed standard or learning objective.
🧠 Quick Revision Questions
- What is the primary purpose of a placement assessment, and when is it typically used in a course?
- Explain the key difference between a formative and a summative assessment, providing an example of each.
- Describe the main goal of a diagnostic assessment and what type of student problem requires it.
- How does a norm-referenced test differ from a criterion-referenced test in how it interprets a student's score?
- If a test is designed with items of "average difficulty" to create a large spread of scores, is it more likely to be a norm-referenced or criterion-referenced test? Why?
Here is the summary of Lecture 3, formatted exactly as requested.
📘 Lecture 03 — Assessment, Testing and National Curriculum
📖 Overview: This lecture explains the hierarchical structure of Pakistan's national curriculum, detailing how student learning is classified from broad competencies down to specific Student Learning Outcomes (SLOs). It clarifies how these four levels connect to form a coherent curriculum and outlines the specific modes and purposes of assessment recommended within this framework.
🗂️ Topics Covered
This lecture begins by defining the four levels of student learning classification in the national curriculum: Competency, Standards, Benchmarks, and Student Learning Outcomes (SLOs). It then demonstrates how these levels are connected, using an example from the English curriculum to show the progression from a competency to specific SLOs. Finally, it covers the modes of assessment recommended in the curriculum, including formative and summative approaches, and lists suitable assessment tools like MCQs, constructed response items, and performance tasks.
📝 Lecture Summary
Topic- 007: Role of National Curriculum in Assessment
In the national curriculums of Pakistan, student learning is classified into four distinct levels. These levels create a hierarchy that moves from the broadest learning area to the most specific, measurable outcome for a single grade. Understanding this hierarchy is crucial for aligning teaching and assessment with national standards.
🔑 Definition — Competency: A key learning area, for example, algebra, arithmetic, geometry, etc. in mathematics and vocabulary, grammar, composition, etc. in English.
🔑 Definition — Standards: These define the competency by specifying broadly, the knowledge, skills and attitudes that students will acquire, should know and be able to do in a particular key learning area during twelve years of schooling.
🔑 Definition — Benchmarks: The benchmarks further elaborate the standards, indicating what the students will accomplish at the end of each of the five developmental levels in order to meet the standard.
🔑 Definition — Student Learning Outcomes (SLOs): These are built on the descriptions of the benchmarks and describe what students will accomplish at the end of each grade. It is the lowest level of hierarchy.
Topic- 008: Connecting all Four Levels in Curriculum
The four levels are not isolated; they are interconnected in a pyramid-like structure. SLOs are at the bottom (the most specific level), and all SLOs combine to form a benchmark. Multiple benchmarks then convert into a standard, and finally, standards combine to define a competency.
📌 Example: This connection is illustrated with an example from the English subject curriculum.
- Competency 1: Reading and thinking skills
- Standard 1: All students will discover and understand a variety of text types through tasks which require multiple reading and thinking strategies for comprehension, fluency and enjoyment.
- Benchmark 1: Use reading readiness strategies
- Student Learning Outcome:
- Articulate, identify and differentiate between the sounds of individual letters, digraphs and trigraphs in initial and final positions in a word.
- Identify paragraph as a graphical unit of expansion, know that word in a sentence join to make sense in relation to each other.
Topic- 009: Modes of Assessment in Curriculum
The curriculum document provides specific guidelines for assessment. It recommends two primary forms of assessment to measure student progress and achievement.
🔑 Definition — Periodic/formative assessment: This is done through homework, quizzes, class tests and group discussions.
🔑 Definition — End of term/ summative assessment: This is done through a final examination.
Purpose of Assessment and Curriculum-English 2006 The assessment system for the present curriculum should include:
- A clear statement of the specific purpose(s) for which the assessment is being carried out.
- A wide variety of assessment tools and techniques to measure student’s ability to use language effectively.
- Criteria to be used for determining performance levels for the SLOs for each grade level.
- Procedures for interpretation and use of assessment results to evaluate the learning outcomes.
Form of Suitable Assessment Tools- English 2006 The curriculum recommends the following forms of assessment tools: ➢ MCQs (Multiple Choice Questions) ➢ Constructed response
- Restricted response
- Extended response ➢ Performance tasks
⭐ Key Takeaways
The core of this lecture is the hierarchical structure of Pakistan's national curriculum, which moves from broad Competencies down to specific Student Learning Outcomes (SLOs) for each grade. You must remember that SLOs are the smallest, most granular level, and they aggregate to form Benchmarks, which then form Standards, which ultimately define a Competency. The curriculum explicitly recommends two modes of assessment: periodic formative assessment and end-of-term summative assessment, and emphasizes using a variety of tools like MCQs, constructed response items, and performance tasks to measure student learning effectively.
🧠 Quick Revision Questions
- Name the four levels of student learning classification in Pakistan's national curriculum, starting from the broadest to the most specific.
- What is the key difference between a Standard and a Benchmark in the curriculum hierarchy?
- In the English curriculum example provided, what is the specific Competency that the Standard "use reading readiness strategies" falls under?
- What are the two recommended modes of assessment mentioned in the curriculum document?
- List two of the three forms of suitable assessment tools specified in the English 2006 curriculum.
📘 Lecture 04 — TAXONOMIES OF EDUCATIONAL OBJECTIVES AND ASSESSMENT –I
📖 Overview: This lecture introduces foundational frameworks for classifying educational objectives and assessing student learning outcomes. It covers three major taxonomies—Bloom’s, SOLO, and DOK—explaining their hierarchical levels, how they categorize learning complexity, and why they are essential for designing effective assessments.
🗂️ Topics Covered
The lecture begins by identifying three pillars of assessment and introduces three popular taxonomies: Bloom’s taxonomy (cognitive, affective, psychomotor domains), SOLO taxonomy (five levels of competency), and DOK (four levels of depth). It then explains each taxonomy in detail, including levels, indicative verbs, and examples for SOLO and DOK.
📝 Lecture Summary
Taxonomies of Educational Objectives and Assessment
Every assessment, regardless of its purpose, rests on three important pillars: a model for how students present knowledge and develop competence, tasks or situations that allow observation of student performance, and inferences from performance evidence about the quality of learning. In developing a test to assess student learning, a taxonomy provides a framework of categories with different hierarchical levels of outcomes.
Popular Taxonomies:
- Bloom’s taxonomy of educational objectives
- Structure of Observed Learning Outcomes (SOLO)
- Depth of Knowledge (DOK)
Bloom’s Taxonomy and SOLO Taxonomy
The SOLO taxonomy (Structure of Observed Learning Outcomes) was initially developed by Biggs and Collis in 1982, and well described by Biggs and Tang in 2007. It carries five different levels of competency for learners.
Levels of SOLO:
- Pre-structural
- Uni-structural
- Multi-structural
- Relational
- Extended Abstract
Depth of Knowledge (DOK) was presented by Webb in 1997, giving four levels of learning activities.
Levels of DOK:
- Recall
- Skill/Concept
- Strategic Thinking
- Extended Thinking
Bloom’s Taxonomy of Learning Objectives was presented by Benjamin Bloom in 1956, consisting of a framework with the most common objectives of classroom instructions across three domains:
- Cognitive
- Affective
- Psychomotor
Cognitive Domain levels: Knowledge, Comprehension, Application, Analysis, Synthesis, Evaluation.
Affective Domain levels: Receiving, Responding, Valuing, Organization, Characterization.
Psychomotor Domain levels: Perception, Set, Guided Response, Mechanism, Complex Covert Response, Adaptation, Origination.
SOLO Taxonomy
The five levels of SOLO are:
1. Pre-structural Students are simply able to acquire bits of unconnected information and respond to a question in a meaningless way. 🔑 Definition — Pre-structural: Students respond with irrelevant or no meaningful connections. 📌 Example: Question: "What is your name?" Answer: "What is your name?"
2. Uni-structural Student shows concrete understanding of the topic but is only able to respond with one relevant element from the stimuli or an item provided. Indicative verbs: identify, memorize, do simple procedure.
3. Multi-structural The student can understand several components, but the understanding of each remains discreet. A number of connections are made, but the significance of the whole is not determined. Ideas and concepts around an issue are disorganized and aren't related together. Indicative verbs: enumerate, classify, describe, list, combine, do algorithms.
4. Relational At this level, the learner is able to understand the significance of the parts in relation to the whole. Ideas and concepts are linked, and they provide a coherent understanding of the whole. Indicative verbs: compare/contrast, explain causes, integrate, analyze, relate, and apply.
5. Extended Abstract At this level, the learner is able to think hypothetically and can synthesize material logically. Students make connections not only within the given subject area, but understanding is transferable and generalizable to different areas. Indicative verbs: theorize, generalize, hypothesize, reflect, generate.
Depth of Knowledge (DOK)
The four levels of DOK measure the degree to which the knowledge brought about from students on assessments is as complex as what students are expected to know and do as stated in the curriculum.
Level 1: Recall Recall of a fact, information, or procedure. The subject matter at this level usually involves working with facts, terms, and/or properties of objects. Key words: list, enlist, name, define, etc.
Level 2: Skill/Concept It includes the engagement of some mental processing beyond recalling or reproducing a response. Use information or conceptual knowledge, two or more steps, not just recalling. Key words: graph, separate, relate, contrast, narrate, compare, etc.
Level 3: Strategic Thinking Items falling in this category demand a short-term use of higher order thinking processes, such as analysis and evaluation, to solve real-world problems with predictable outcomes. Key words: argue, critique, formulate.
Level 4: Extended Thinking Learning outcomes at this level demand extended use of higher order thinking processes such as synthesis, reflection, assessment, and adjustment of plans over time. Key words: create, synthesize, design, and reflection.
💡 Why this matters: DOK helps ensure assessments match the cognitive complexity expected in the curriculum, preventing tests from being too easy or too hard.
⭐ Key Takeaways
The three major taxonomies—Bloom’s, SOLO, and DOK—provide complementary frameworks for classifying educational objectives and assessment complexity. SOLO taxonomy’s five levels (pre-structural to extended abstract) describe how student understanding progresses from disconnected facts to transferable, abstract thinking. DOK’s four levels (recall to extended thinking) measure the cognitive depth required by assessment tasks. Using these taxonomies helps educators design assessments that properly align with learning outcomes and accurately measure student competency.
🧠 Quick Revision Questions
- What are the three pillars on which every assessment rests?
- List the five levels of SOLO taxonomy in order from lowest to highest.
- What is the key difference between a multi-structural and relational level in SOLO?
- Name the four levels of DOK and give one keyword example for each.
- Which taxonomy has three domains (cognitive, affective, psychomotor), and who presented it?
📘 Lecture 5 — TAXONOMIES OF EDUCATIONAL OBJECTIVES AND ASSESSMENT –II
📖 Overview: This lecture explores the classification systems used to define and assess educational objectives, focusing on Bloom’s Taxonomy of the cognitive domain. It covers both the original and revised versions of Bloom’s Taxonomy, introduces the concept of instructional objectives, and explains how to write specific learning outcomes. Understanding these taxonomies is crucial for teachers to design effective lessons, create appropriate assessments, and ensure students achieve desired learning levels.
🗂️ Topics Covered
The lecture begins with an introduction to the three main domains of learning: cognitive, affective, and psychomotor. It then delves into the original Bloom’s Taxonomy for the cognitive domain, detailing the six levels from Knowledge to Evaluation. The revised version of Bloom’s Taxonomy is presented, with the levels reordered to Remembering, Understanding, Applying, Analyzing, Evaluating, and Creating, along with key verbs for each. A comparison is made between Bloom, SOLO, and DOK taxonomies. Finally, the lecture covers instructional objectives as learning outcomes, sources for objectives, criteria for selecting final objectives, general objectives, and the steps for stating specific learning outcomes.
📝 Lecture Summary
Topic- 014: Bloom‘s Taxonomy-I
There are three main domains of learning that all teachers should know and use to construct lessons: the Cognitive Domain (thinking), Affective Domain (feeling), and Psychomotor Domain (doing). In 2000-01, revisions to the cognitive taxonomy were led by Bloom's former student Lorin Anderson and his original partner David Krathwohl. One major change between the old and new versions is that the two highest forms of cognition were reversed.
The Old Cognitive Domain levels are:
-
Knowledge: Defined as the remembering of previously learned material. This involves the recall of a wide range of facts, procedures, principles, and generals. 🔑 Definition — Knowledge: The remembering of previously learned material, involving the recall of facts, procedures, and processes. 📌 Example: Sample question: “Define the 6 levels of Bloom's taxonomy of the cognitive domain.”
-
Comprehension: Defined as the ability to grasp the meaning of the material. An individual can use the content without necessarily relating it to other content. 🔑 Definition — Comprehension: The ability to grasp the meaning of the material. 📌 Example: Sample question: “Explain the purpose of Bloom's taxonomy of the cognitive domain.”
-
Application: Refers to the ability to use previously learned material in new and concrete situations. Abstractions may be in the shape of universal ideas or rules of methods. 🔑 Definition — Application: The ability to use previously learned material in new and concrete situations. 📌 Example: Sample question: “Write an instructional objective for each level of Bloom's taxonomy.”
Topic- 015: Bloom‘s Taxonomy-II
-
Analysis: The breakdown of a concept into its constituent parts so that the relative hierarchy of the concept becomes easy to understand, or the relation between parts is elaborated. 🔑 Definition — Analysis: The breakdown of a concept into its constituent parts to understand its hierarchy and the relations between parts. 📌 Example: Sample question: “Compare and contrast the cognitive and affective domains.”
-
Synthesis: There is a collection of the constituents or parts of a concept to make a whole. An individual works with pieces and groups them to formulate a pattern not clearly there before. 🔑 Definition — Synthesis: The collection of parts of a concept to form a new whole, creating a pattern or structure. 📌 Example: Sample question: “Design a classification scheme for writing educational objectives that combines the cognitive, affective, and psychomotor domains.”
-
Evaluation: Concerned with the ability to judge the value of material for a given purpose. Judgments are made on definite criteria. 🔑 Definition — Evaluation: The ability to judge the value of material for a given purpose based on definite criteria. 📌 Example: Sample question: “How far are the different BISEs and universities developing papers using Bloom's taxonomy? Support your answer with arguments.”
Topic- 016: Revised version of Bloom‘s Taxonomy
The revised cognitive domain levels are:
-
Remembering: Exhibit memory of previously learned material by recalling facts, terms, and basic concepts. 🔑 Key Verbs: Choose, define, find, how, label, list, match, name, omit, recall, relate, select, show, spell, tell, what, when, where, which, who, why.
-
Understanding: Constructing meaning from different types of functions, whether written or graphic messages, or activities. 🔑 Key Verbs: Classify, compare, contrast, demonstrate, explain, extend, illustrate, infer, interpret, outline, relate, rephrase, show, summarize, and translate.
-
Applying: Solve problems in new situations by applying acquired knowledge, facts, techniques, and rules in a different way. 🔑 Key Verbs: Apply, build, choose, construct, develop, experiment with, identify, interview, make use of, model, organize, plan, select, solve and utilize.
-
Analyzing: Breaking materials or concepts into parts, determining how the parts relate to one another or to an overall structure or purpose. 🔑 Key Verbs: Analyze, assume, categorize, classify, compare, conclusion, contrast, discover, dissect, distinguish, divide, examine, function, inference and inspect.
-
Evaluating: Making judgments based on criteria and standards through checking and critiquing. 🔑 Key Verbs: Agree, appraise, assess, award, choose, compare, conclude, criteria, criticize, decide, deduct, defend, determine, disprove and estimate.
-
Creating: Putting elements together to form a coherent or functional whole; reorganizing elements into a new pattern or structure through generating, planning, or producing. 🔑 Key Verbs: Adapt, build, change, choose, combine, compile, compose, construct, create, delete, design, develop, discuss, elaborate, estimate and formulate.
These categories range from simple to complex and from concrete to abstract levels of student learning. The taxonomy represents a cumulative hierarchy, so mastery of each simpler category is a prerequisite for mastery of the next, more complex one. 💡 Why this matters: This hierarchy guides teachers in scaffolding instruction and assessments, ensuring students build foundational knowledge before tackling higher-order thinking skills.
A Comparison of Bloom, SOLO and DOK taxonomies is presented, showing their alignment and distinctions for assessing different levels of cognitive complexity.
Topic- 017: Instructional Objectives
Instructional Objectives as Learning Outcome: Instructional goals and objectives are stated in terms of actions to be taken. When viewing instructional objectives in terms of learning outcomes, we are concerned with products rather than the process of learning.
Sources for Lists of Objectives include:
- Professional Associations standards
- State Content Standards
- Methods Books
- Year Books
- Encyclopedia of Educational Research
- Curriculum Frameworks
- Test Manuals
Criteria for Selecting the Final List of Objectives:
- Prepare a tentative list of instructionally relevant learning outcomes.
- Review the list for: Completeness, Appropriateness, Soundness, and Feasibility.
General Objectives:
- Stating the general objectives requires selecting the proper level of generality.
- Objectives should be specific enough to provide direction for instruction but not so specific that instruction is reduced to training.
- Stating general objectives in general terms provides for the integration of specific facts and skills into complex responses.
- General statements give teachers freedom in selecting the method and materials of instruction.
- A list of general objectives shows the desired level of generality, for example: “Knows basic terminology,” “Understand concepts,” “Relates concepts to everyday observations,” “Applies principles to new situations,” “Interpret graphs,” and “Demonstrate scientific attitude.”
Specific Learning Outcomes: Each general objective must be defined by a sample of specific learning outcomes to clarify how students can demonstrate they have achieved the general objective. Until general objectives are further defined, they will not provide adequate direction for assessment.
Steps for Stating Specific Outcomes:
- List below each general objective a representative sample of specific learning outcomes that describe the terminal performance students are expected to demonstrate.
- Begin each specific learning outcome with an action verb that specifies observable performance.
- Make sure that each specific learning outcome is relevant to the general objective it describes.
- Include enough SLOs to adequately describe the performances of students who have attained the objectives.
- Keep the SLOs sufficiently free of course content so that the list can be used with various units of study.
- Consult reference materials for the specific components of those complex outcomes that are difficult to define.
- Add a third level of specificity to the list of outcomes, if needed.
⭐ Key Takeaways
The most critical point is that the cognitive domain of Bloom’s Taxonomy exists in two versions: the original (Knowledge, Comprehension, Application, Analysis, Synthesis, Evaluation) and the revised one (Remembering, Understanding, Applying, Analyzing, Evaluating, Creating), with the key difference being the reversal of the top two levels. It is essential to memorize the specific definitions and sample verbs for each level of the revised taxonomy, as they guide the creation of learning objectives and assessment items. Furthermore, instructional objectives must be stated as observable learning outcomes, moving from broad general objectives to specific, action-verb-driven specific learning outcomes. Teachers must master the skill of writing measurable specific learning outcomes that align with general objectives to ensure both instruction and assessment are effective and focused. Finally, the cumulative hierarchy of Bloom’s Taxonomy means simpler levels must be mastered before students can succeed at more complex ones.
🧠 Quick Revision Questions
- What are the three main domains of learning, and on which domain does Bloom’s Taxonomy focus?
- What is the single most important structural difference between the original and the revised version of Bloom’s Taxonomy for the cognitive domain?
- List the six levels of the revised Bloom’s Taxonomy in order from simplest to most complex, and provide one key verb for each level.
- According to the lecture, what is the purpose of writing specific learning outcomes (SLOs) under each general objective?
- What does it mean that Bloom’s Taxonomy represents a “cumulative hierarchy”?
📘 Lecture 6 — Purpose of Testing-I
📖 Overview: This lecture covers the fundamental purposes of educational testing, beginning with the various types of educational decisions that test data inform. It then classifies different types of written tests used in classroom assessment and concludes by contrasting two major assessment frameworks: Norm-Referenced Testing (NRT) and Criterion-Referenced Testing (CRT).
🗂️ Topics Covered
The lecture first explores nine types of educational decisions made at different levels (administrative, school, and classroom), including instructional, grading, diagnostic, selection, placement, program/curriculum, and administrative decisions. It then categorizes written tests into eight types based on format (verbal/non-verbal), scoring (objective/subjective), construction (teacher-made/standardized), and time constraints (power/speed). Finally, it introduces Norm-Referenced Assessment (NRT), which compares a student’s performance to a peer group, and Criterion-Referenced Assessment (CRT), which measures mastery against a fixed standard or criterion.
📝 Lecture Summary
Topic-018: Educational Decisions Making
Educational decisions are classified into nine types taken at different levels: board/administrative, school management, or classroom by teachers. These decisions rely on test data.
🔑 Instructional Decisions: The nuts and bolts decisions made daily by classroom teachers, such as deciding to spend more time on specific units, regroup students in class for better management, or modify instructional plans.
🔑 Grading Decisions: Made by the classroom teacher but much less frequently than instructional decisions. For most students, grading decisions are the most influential decisions made about them.
🔑 Diagnostic Decisions: Decisions made about a student’s strengths and weaknesses and the reasons behind them. Teachers make these based on information from informal, teacher-made tests, but standardized tests can also provide diagnostic information.
🔑 Selection Decisions: Involve test data used in part for accepting or rejecting applicants for admission into a group, program, or institution.
🔑 Placement Decisions: Made after an individual has been accepted into a program. They determine where in a program someone is best suited to begin.
🔑 Program or Curriculum Decision: A policy-level decision about whether a lesson, unit, or subject will continue or be abandoned for the next academic session, based on national objectives of education.
🔑 Administrative Decisions: Policy decisions made at school, district, state, or national level. Based on measurement data, this includes financial decisions of schools.
Topic-019: Types of Test
In classroom assessment, different forms are utilized, each with its own benefits and disadvantages. The most common type is written assessment.
🔑 Types of Written Tests: Verbal, Non-verbal, Objective, Subjective, Teacher Made, Standardized, Power, and Speed.
🔑 Verbal tests: Emphasize reading, writing, or speaking. Most tests in education are verbal tests.
🔑 Non-verbal tests: Do not require reading, writing, or speaking ability. Tests composed of numerals or drawings are examples.
🔑 Objective tests: Refers to scoring where two or more scorers can easily agree on whether an answer is correct or incorrect. True/false, multiple choice, and matching tests are examples.
🔑 Subjective tests: When it is difficult for two scorers to agree on whether an item is correct or incorrect. Essay tests are the example.
🔑 Teacher Made tests: Constructed solely by a teacher to be used in the classroom. They are custom-designed according to the needs and issues of a specific class.
🔑 Standardized tests: Constructed by measurement experts over a period of years. They measure broad national objectives and have a uniform set of instructions adhered to during each administration. Mostly, they have tables of norms to which a student’s performance may be compared to determine where they stand in relation to a national sample of students at the same age or grade level.
🔑 Power tests: Have liberal time limits that allow each student to attempt each item. Items tend to be difficult.
🔑 Speed tests: Have time limits so strict that no one is expected to complete all items. Items tend to be easy.
Topic-020: Norm Referenced Assessment (NRT)
The general purpose of assessment is to gather information to make better and more informed decisions. The utility of that information differentiates among the types of assessments.
🔑 Norm-Referenced Test (NRT): A type of test that tells us where a student stands compared to other students. It helps to determine a student’s place or rank among a group of similar students.
Dimensions of NRT:
- It provides an estimate of ability in a variety of skills in a much shorter time.
- NRT tends to be general. It measures a variety of skills at the same time but fails to measure them thoroughly.
- It is hard to make decisions regarding the mastery of a student’s skill in a subject.
- NRT is much more difficult for students to solve. On average, only 50% of students are able to get an item right in a test.
Topic-021: Criterion Referenced Assessment (CRT)
🔑 Criterion-Referenced Test (CRT): A type of test that tells us about a student’s level of proficiency in or mastery of some skill or set of skills. This is achieved by comparing a student’s performance to a standard mastery called a criterion.
Dimensions of CRT:
- CRT tends to be specific. It measures a particular set of skills at one time and focuses on the level of achievement of that skill. CRT gives a clear picture regarding the mastery of a student’s skill in a subject.
- It measures skill more thoroughly, so naturally, it takes more time compared to NRT in measuring the mastery of said skill.
- Items included in CRT are relatively easier. Around 80% of the students are expected to respond to an item correctly in the test.
- CRT compares students’ performance to the standards indicative of mastery.
- The breadth of content sampled is narrow and covers very few objectives.
⭐ Key Takeaways
A student must understand that test data serves different educational decision-making purposes at various levels (classroom, school, policy). They must be able to classify tests into their eight written types and distinguish between them. The most critical distinction is between Norm-Referenced Tests (NRT), which rank students against each other (e.g., “50% get it right”), are general, and assess broad skills, and Criterion-Referenced Tests (CRT), which measure mastery of a specific skill against a standard (e.g., “80% get it right”), are specific, and provide a clear picture of individual proficiency. Remember also that teacher-made tests are custom for a class, while standardized tests have national norms.
🧠 Quick Revision Questions
- What is the main difference between a diagnostic decision and a placement decision?
- List three types of written tests that are based on who constructs them (rather than format or scoring).
- Explain the key difference between a power test and a speed test, including the typical difficulty of items in each.
- A test reports that a student scored better than 85% of other students in their grade. Is this likely from a Norm-Referenced Test (NRT) or a Criterion-Referenced Test (CRT)? Why?
- If a teacher wants to know if a student has fully mastered the skill of dividing fractions, which type of assessment (NRT or CRT) would be more appropriate, and why?
📘 Lecture 7 — Purpose of Testing –II
📖 Overview: This lecture explores the distinctions between criterion-referenced and norm-referenced testing, examining their unique characteristics, item selection, and scoring approaches. It also introduces formative assessment as a tool for ongoing feedback during the instructional process, which is crucial for both student and instructor growth.
🗂️ Topics Covered
The lecture covers three main topics: the characteristics of criterion-referenced assessment (CRT), a detailed comparison between criterion-referenced tests (CRT) and norm-referenced tests (NRT) across six bases of comparison, and the definition and types of formative assessment used during the learning process.
📝 Lecture Summary
Topic- 022: Characteristics of Criterion Referenced Assessment
Criterion Referenced Assessment (CRT) evaluates an examinee’s performance against a predefined external standard of competence, rather than comparing it to other examinees. Sampled content in CRT is much more comprehensive, usually three or more items are used to cover a single objective. The meaning of the score does not depend upon comparison with other scores; it flows directly from the connection between the items and the criterion. Items are chosen to reflect the criterion behavior, with emphasis placed upon the domain of relevant responses. Scoring uses the number succeeding or failing or range of acceptable performance.
🔑 Definition — Criterion Referenced Assessment (CRT): A test where an examinee's performance is compared to an external standard of competence, focusing on mastery of specific objectives.
💡 Why this matters: CRT tells you exactly what a student knows and can do relative to a set of standards, not how they compare to peers.
📌 Example: "90% proficiency achieved" means the student answered 90% of items correctly on the criterion, or "80% class reached 90% proficiency" means that 80% of students in the class achieved the 90% mastery level.
Topic- 023: Difference between NRT and CRT
This topic compares Norm Referenced Tests (NRT) and Criterion Referenced Tests (CRT) across six bases of comparison:
-
Comparison Targets: In CRT, the examinee’s performance is compared to an external standard of competence. While in NRT, examinee’s performance is typically compared to that of other examinees.
-
Selection of Items: Items included in CRT are of specific nature and designed for the student skilled in particular subjects. In NRT, items are of general knowledge nature; the student should be able to answer it, but superficial knowledge is sufficient to respond to the item correctly.
-
Meaning of Success: In CRT, an examinee is classified as a master or non-master. There is no limit to the number of pass or fail. In NRT, an examinee’s opportunity for success is relative to the performance of the other individuals who take the test.
-
Average Item Difficulty: In CRT, the average item difficulty is fairly high because examinees are expected to show mastery. In NRT, the average item difficulty is lower because tests are designed to spread out examinees and provide a reliable ranking.
-
Score Distributions: In CRT, a plot of the resulting score distribution will show most of the scores clustering near the high end of the score scale. In NRT, a broader spread of scores is expected, with a few examinees earning very low or high scores and many earning medium scores.
-
Reported Scores: In CRT, classification of the examinee is measured as master/non-master or pass/fail. In NRT, percentile ranks or scale scores are frequently used.
🔑 Definition — Norm Referenced Test (NRT): A test where an examinee's performance is compared to that of other examinees, designed to rank students and spread out scores.
🔑 Definition — Master/Non-master: In CRT, an examinee is classified as a master (has met the criterion standard) or non-master (has not met the criterion standard), with no limit on how many can pass or fail.
Topic- 024: Formative Assessment
Formative assessment provides feedback and information during the instructional process, while learning is taking place and occurring. It measures student progress, but it can also assess your own progress as an instructor.
💡 Why this matters: Unlike summative assessment (which happens at the end), formative assessment happens during learning, allowing for real-time adjustments to teaching and learning strategies.
Types of Formative Assessment include:
- Observations during in-class activities, including student’s non-verbal feedback during lecture
- Homework exercises as review for exams and class discussions
- Reflection journals that are reviewed periodically during the semester
- Question and answer sessions, both formal (planned) and informal (spontaneous)
- Conferences between the instructor and student at various points in the semester
- In-class activities where students informally present their results
- Student feedback collected by periodically answering specific questions about the instruction and their self-evaluation of performance and progress
⭐ Key Takeaways
CRT assesses mastery against a fixed standard using specific, comprehensive items with high difficulty, reporting pass/fail results. NRT compares performance to other test-takers using general items with lower difficulty to create a wider score distribution, reporting percentile ranks. Formative assessment is a continuous feedback loop during instruction, using various methods like observations, homework, journals, Q&A sessions, and conferences to improve both student learning and teaching effectiveness.
🧠 Quick Revision Questions
- On which basis of comparison do CRT and NRT differ regarding the selection of test items?
- In a CRT, what does it mean if a class reports "80% of students achieved 90% proficiency"?
- Why is average item difficulty fairly high in CRT but lower in NRT?
- What is the primary purpose of formative assessment during instruction?
- List three types of formative assessment activities mentioned in the lecture.
📘 Lecture 8 — Purpose of Testing-III
📖 Overview: This lecture continues exploring the purpose of testing by focusing on formative and summative assessment. It explains the functions of each, how they differ in timing, focus, and purpose, and why understanding these distinctions is crucial for effective teaching and learning.
🗂️ Topics Covered
This lecture covers the functions of formative assessment, the definition and types of summative assessment, and the functions of summative assessment. It explains how formative assessment focuses on the learning process with ongoing feedback, while summative assessment evaluates the final product for grading and certification.
📝 Lecture Summary
Functions of Formative Assessment
Formative assessment focuses on a predefined segment of instruction, meaning it only tests what was just taught. A limited sample of learning tasks are addressed, not the entire course content. The difficulty of items varies with each segment of instruction, so early segments may have easier items. Formative assessment is conducted periodically during the instructional process, not just at the end. Its results are used to improve and direct learning through ongoing feedback, helping students correct mistakes immediately.
🔑 Definition — Formative Assessment: Assessment conducted during instruction to monitor student learning and provide ongoing feedback for improvement.
Summative Assessment
Summative assessment takes place after the learning has been completed and provides information and feedback that sums up the teaching and learning process. Typically, no more formal learning is taking place at this stage, other than incidental learning through completing projects and assignments. Summative assessment is more product-oriented and assesses the final product, whereas formative assessment focuses on the process toward completing the product. Once the project is completed, no further revisions can be made. If students are allowed to make revisions, the assessment becomes formative.
🔑 Definition — Summative Assessment: Assessment that evaluates student learning at the end of an instructional unit by comparing it against some standard or benchmark.
📌 Example: A final exam is summative because it tests the entire course after instruction ends. If a teacher allows students to revise their answers after seeing the correct ones, it becomes formative.
Types of Summative Assessment include:
- Examinations (major, high-stakes exams)
- Final examinations (a truly summative assessment)
- Term papers (drafts submitted during the semester would be a formative assessment)
- Projects (project phases submitted at various completion points could be formatively assessed)
- Portfolios (could also be assessed during its development as a formative assessment)
- Performances
💡 Why this matters: The key distinction is that the same assignment (e.g., a term paper) can be either formative or summative depending on when and how it is assessed. Drafts are formative; final submissions are summative.
Functions of Summative Assessment
The focus of measurement in summative assessment is on course or unit objectives. A broad sample of all objectives is used, unlike formative assessment which samples only a segment. This type of assessment uses a wide range of difficulty while selecting items for the test. Summative assessment is done at the end of the unit or the course. The most important functions of summative assessment are to assign grades, provide certification of accomplishment, and allow evaluation of teaching.
📌 Example: A final exam for a semester-long course will include easy, medium, and hard questions covering all 10 units, not just the last unit. The results determine the final grade and whether the student can proceed to the next level.
⭐ Key Takeaways
Formative assessment is conducted during instruction to improve learning through feedback, while summative assessment occurs after instruction to evaluate final outcomes. Formative assessment focuses on a limited segment of instruction with varying item difficulty, while summative assessment covers all objectives with a broad range of difficulty. The same activity (like a term paper or portfolio) can be either formative or summative depending on whether revisions are allowed. The main functions of summative assessment are grading, certification, and teaching evaluation, whereas formative assessment aims to guide ongoing learning.
🧠 Quick Revision Questions
- What is the key difference in when formative and summative assessments are conducted?
- List three functions of formative assessment.
- Can a term paper be both a formative and summative assessment? Explain.
- What is the focus of measurement in summative assessment?
- Why does item difficulty vary in formative assessment but use a wide range in summative assessment?
📘 Lecture 9 — Table of Specification
📖 Overview: This lecture introduces the Table of Specification, a critical tool for developing test blueprints that ensure content validity and balanced assessment. It explains how teachers systematically allocate questions across content areas and Bloom's Taxonomy cognitive levels, providing practical examples for weighting objectives and constructing a balanced test.
🗂️ Topics Covered
The lecture covers the concept and definition of Table of Specification, its two main functions in test construction, the six elements and appropriateness checklist for developing a comprehensive table of specification, the process of balancing learning objectives through instruction time calculations, and a practical step-by-step example demonstrating how to construct a Table of Specification with proper weightage distribution.
📝 Lecture Summary
Topic- 028: Table of Specification
The Table of Specification is a tool used by teachers to develop a blueprint for the test. It is the first formal step in test construction and serves as the technical name for the test blueprint.
The blueprint is meant to ensure content validity, which is the most important factor in constructing an achievement test. A unit test or comprehensive exam is based on several lessons/chapters and should reflect a balance between content areas and learning levels (objectives).
🔑 Definition — Table of Specification: A two-way chart or grid relating instructional objectives to instructional content.
💡 Why this matters: Without a Table of Specification, teachers may unintentionally over-emphasize certain topics while neglecting others, compromising the test's validity.
The Table of Specification performs two important functions:
- It ensures the balance and proper emphasis across all content areas covered by the teacher.
- It ensures the inclusion of items at each level of the cognitive domain of Bloom's Taxonomy.
Topic- 029: Concept of Table of Specification
The Table of Specification helps a teacher in allotting questions to different content areas and Bloom's learning categories in a systematic manner. It ensures that every content area receives appropriate representation and that questions target various cognitive levels appropriately.
Topic- 030: Elements and Appropriateness in Table of Specification
Carey (1988) listed six major elements that should be attended to in developing a Table of Specifications for a comprehensive end-of-unit exam:
i. Balance among the goals selected for the exam (weighing objectives) ii. Balance among the levels of learning (higher order and lower order) iii. The test format (objective and subjective) iv. The total number of items v. The number of test items for each goal and level of learning vi. The enabling skills to be selected from each goal framework
A Table of Specifications incorporating these six elements will result in a "comprehensive posttest that represents each unit and is balanced by goals and levels of learning."
Checklist for Appropriateness of Table of Specification:
- Are the specifications in harmony with the purpose of the test?
- Do the specifications indicate the nature and limits of the achievement domain?
- Do the specifications indicate the types of learning outcomes to be measured?
- Do the specifications indicate the sample of learning outcomes to be measured?
- Is the number of test items indicated for the total test and for each subdivision?
- Are the types of items to be used appropriate for the outcomes to be measured?
- Is the difficulty of the items appropriate for the types of interpretation to be made?
- Is the distribution of items adequate for the types of interpretation to be made?
- If sample items are included, do they illustrate the desired attributes?
- Do the specifications, as a whole, indicate a representative sample of instructionally relevant tasks that fits the use to be made of the results?
Topic- 031: Balance Among Learning Objectives and Their Weighting in Table of Specification
In developing a test blueprint, it is necessary to select learning objectives. Some objectives are more important because more instruction time is spent on them, while others are less important. Therefore, we need to weigh the learning objectives for calculating their relative weightage in the test.
Step 1: Instruction Time To calculate instruction time for columns of the Table of Specifications, the teacher must use the following formula:
📐 Formula: Percentage of instruction time = Time spent on objective (min) / Total time for instruction being examined (min)
Example: If 250 minutes were spent on a particular objective out of 1000 total minutes of instruction: Percentage of instruction time = 250/1000 Percentage of instruction time = 25%
Step 2: Examination Value The instructor should determine the number of test items/score to be allocated to that objective. Let us assume total marks of the test are 100. Then 25 marks should be allocated to questions related to that objective.
Step 3: Validation Percent of instruction time = Percent of examination value (within ±2 percent, if not, redo test) 25 ± 2 = 25 ± 2
If total marks of the test are 50, then 25% of 50 = 12.5 marks.
📐 Formula: Point total of questions for objective / Total points on examination = % of examination value
Topic- 032: Balance Among Learning Objectives and Their Weight in Table of Specification: Example
This section provides a practical example of developing a Table of Specification.
Initial Table with Time Weightage:
| Topics/Level | Knowledge | Comprehension | Application | Marks |
|---|---|---|---|---|
| Pakistan Movement (100/500×100 = 20%) | ||||
| Geography of Pakistan (150/500×100 = 30%) | ||||
| Climate Change (150/500×100 = 20%) | ||||
| Industries (50/500×100 = 10%) | ||||
| Economy (50/500×100 = 10%) | ||||
| Total (Time: 500/Marks: 50) |
Step 1: Distribution of Marks for Each Topic (50-mark test):
| Topics/Level | Knowledge | Comprehension | Application | Marks |
|---|---|---|---|---|
| Pakistan Movement | 10 (20%) | |||
| Geography of Pakistan | 15 (30%) | |||
| Climate Change | 15 (30%) | |||
| Industries | 5 (10%) | |||
| Economy | 5 (10%) | |||
| Total (Time: 500/Marks: 50) | 50 (100%) |
Step 2: Distribution by Cognitive Level According to Bloom's Taxonomy:
| Topics/Level | Knowledge | Comprehension | Application | Marks |
|---|---|---|---|---|
| Pakistan Movement | 5 (50%) | 2 (20%) | 3 (30%) | 10 (20%) |
| Geography of Pakistan | 2 (10%) | 6 (40%) | 7 (50%) | 15 (30%) |
| Climate Change | 7 (50%) | 8 (50%) | 15 (30%) | |
| Industries | 1 (10%) | 1 (20%) | 3 (70%) | 5 (10%) |
| Economy | 1 (20%) | 1 (20%) | 3 (60%) | 5 (10%) |
| Total | 9 (18%) | 17 (34%) | 24 (48%) | 50 (100%) |
📌 Example Explanation: In the Pakistan Movement topic (worth 10 marks), the teacher allocated 5 marks (50%) to Knowledge level questions, 2 marks (20%) to Comprehension level, and 3 marks (30%) to Application level. The overall test has 18% Knowledge questions, 34% Comprehension questions, and 48% Application questions.
💡 Why this matters: This systematic distribution ensures the test measures both lower-order and higher-order thinking skills appropriately across all content areas.
⭐ Key Takeaways
A Table of Specification is a blueprint that ensures content validity by systematically distributing test items across content areas and Bloom's Taxonomy cognitive levels. The six elements from Carey (1988) provide a comprehensive framework including balance among goals, learning levels, test format, total items, items per goal, and enabling skills. The weighting of objectives is determined by calculating the percentage of instruction time spent on each topic using the formula (time on objective / total instruction time), and this percentage should be reflected in the examination marks with a ±2% tolerance. The practical example demonstrates that for a 50-mark test with 500 total minutes of instruction, each topic receives marks proportional to its instruction time, and these marks are further distributed among Knowledge, Comprehension, and Application levels according to the teacher's judgment of each topic's cognitive demands. Finally, the appropriateness checklist helps teachers evaluate whether their Table of Specification aligns with the test purpose, achievement domain, learning outcomes, item types, difficulty levels, and overall representativeness of instruction.
🧠 Quick Revision Questions
-
What are the two main functions of a Table of Specification in test construction?
-
How do you calculate the percentage of instruction time for a specific objective, and what is the acceptable tolerance when converting this to examination marks?
-
List the six major elements that Carey (1988) identified for developing a comprehensive Table of Specifications.
-
In the practical example, if the Geography of Pakistan topic received 30% of instruction time and the test is worth 50 marks, why did it receive 15 marks, and how were these marks distributed across the three cognitive levels?
-
According to the appropriateness checklist, what questions should a teacher ask to ensure the Table of Specification's item types are suitable for the outcomes being measured?
📘 Lecture 10 — SELECTION OF TEST
📖 Overview: This lecture explains how to select appropriate published tests for educational purposes. It covers the types of published tests available, the standards for selecting them, and the importance of fairness in test selection. This is critical because using the wrong test can lead to invalid assessments and unfair outcomes.
🗂️ Topics Covered
The lecture covers selecting pre-designed published tests, including achievement tests, aptitude tests, reading tests, readiness tests, and placement tests. It then details standards for selecting appropriate tests across two parts, focusing on defining purpose, investigating sources, reading materials, understanding test development, reading evaluations, examining specimen sets, and ensuring skill availability. Finally, it addresses fairness in test selection, including content sensitivity, performance review, identifying bias, and accommodating handicapped test takers.
📝 Lecture Summary
Topic- 033: Selecting Pre-designed
Published tests are designed and conducted in such a manner that each and every characteristic is pre-planned and known. There are many published tests available for school use. The two most valuable to the instructional program are: i. Achievement tests ii. Aptitude tests
There are hundreds of tests available for each type. Selecting the most appropriate one is an important task. In some cases, published tests are used by teachers. But more frequently, these are used by provincial or national testing programs.
In classrooms, the most used published tests are: i. Achievement tests ii. Reading tests
Published tests commonly used by provincial or national testing programs are:
- Aptitude tests
- Readiness tests
- Placement tests
Topic- 034: Standards for Selecting Appropriate Test –I
Test users should select tests that meet the purpose for which they are to be used and that are appropriate for the intended population.
- First, define the purpose for testing and the population to be tested and select the test accordingly.
- Investigate the potentially useful sources of information, in addition to the test scores, to validate the information provided by tests.
- Read the materials provided by test developers and avoid using tests for which unclear or incomplete information is provided.
- Become familiar with how and when the test was developed and tried out.
🔑 Definition — Standards for Selecting Appropriate Test (Part I): The process of selecting tests by defining purpose and population, validating information from other sources, reading developer materials, and knowing test development history.
Topic- 035: Standards for Selecting Appropriate Test –II
Test users should select tests that meet the purpose for which they are to be used and that are appropriate for the intended population.
- Read independent evaluations of a test and of possible alternative measures.
- Examine specimen sets, disclosed tests or sample questions, directions, answer sheets, manuals, and score reports before selecting the tests.
- Select and use only those tests for which the skills needed to administer the test and interpret scores correctly are available.
🔑 Definition — Standards for Selecting Appropriate Test (Part II): The process of selecting tests by reading independent evaluations, examining specimen sets, and ensuring available skills for administration and score interpretation.
Topic- 036: Fairness in Selecting Appropriate Test
- Evaluate the procedures used by test developers to avoid potentially insensitive content or language.
- Review the performance of test takers of different races, gender, and ethnic groups when samples of sufficient size are available.
- Evaluate the extent to which the performance differences may have been caused by inappropriate characteristics of the test.
- Use appropriately modified forms of tests or administration procedures for test takers with handicapping conditions.
🔑 Definition — Fairness in Selecting Appropriate Test: The practice of evaluating test content for sensitivity, reviewing performance across demographic groups, identifying bias, and accommodating test takers with disabilities.
💡 Why this matters: Fairness ensures that test scores reflect true abilities rather than irrelevant factors like language bias or disability limitations, leading to more valid and equitable assessments.
⭐ Key Takeaways
Selecting the right test requires defining the testing purpose and population first. Teachers must investigate validation sources, read developer materials, and understand the test’s development history. They should read independent evaluations and examine full specimen sets before deciding. Fairness is essential: test users must review content for insensitivity, analyze performance across groups, identify bias, and accommodate handicapped test takers with modified forms or procedures. Remember: the most commonly used published tests in classrooms are achievement and reading tests, while provincial/national programs use aptitude, readiness, and placement tests.
🧠 Quick Revision Questions
- What are the two most valuable types of published tests for the instructional program?
- List the three standards given in Topic-034 for selecting an appropriate test.
- According to Topic-035, what should test users examine before selecting a test?
- What four steps must be taken to ensure fairness in test selection?
- Name the three types of published tests commonly used by provincial or national testing programs.
📘 Lecture 11 — Characteristics of a Good Test-I
📖 Overview: This lecture introduces the three essential characteristics of a good test: validity, reliability, and usability. It explains the nature of validity in detail and discusses content validity as one of the three main evidences of validity. Understanding these concepts is critical for evaluating and designing effective assessments.
🗂️ Topics Covered
The lecture covers the three essential characteristics of a good test—validity, reliability, and usability—followed by an in-depth examination of the nature of validity, including its definition as a matter of degree, specificity, and unitary concept. It then discusses content validity as evidence of validity, including its meaning, procedure, and methods for establishing it through classroom instruction, achievement domains, and instructional priorities.
📝 Lecture Summary
Topic- 037: Characteristics of Good Test: Validity, Reliability and Usability
The most essential characteristics of a good test are validity, reliability, and usability. Validity is an evaluation of the adequacy and appropriateness of the interpretation and uses of results. It determines if a test is measuring what it intended to measure. Reliability refers to the consistency of assessment results. A key relationship exists between reliability and validity: reliability of measurement is needed to obtain valid results, but we can have reliability without validity. Reliability is a necessity but not a sufficient condition for validity.
🔑 Definition — Validity: An evaluation of the adequacy and appropriateness of the interpretation and uses of results; determines if a test is measuring what it intended to measure.
🔑 Definition — Reliability: The consistency of assessment results.
💡 Why this matters: A test can be reliable (producing consistent scores) but not valid (not measuring the right thing), but a valid test must always be reliable.
Usability refers to the practical requirements of an assessment procedure beyond validity and reliability. These include feasibility, administration environment, and availability of results for decision makers.
🔑 Definition — Usability: Practical requirements of an assessment procedure including feasibility, administration environment, and availability of results.
Topic- 038: Nature of Validity
The nature of validity is described by five key points. First, validity is referred to as "validity of test" but it is in fact validity of the interpretation and use to be made of the results. Second, validity is a matter of degree; it does not exist on an all or none basis. It is best considered in terms of categories that specify degree, such as high, moderate, or low validity. Third, validity is specific to some particular use or interpretation; no assessment is valid for all purposes. For example, an arithmetic test may have a high degree of validity for computational skill and a low degree for arithmetical reasoning. Fourth, validity is a unitary concept; it does not have different types but is viewed as a unitary concept based on different kinds of evidences. Fifth, validity involves an overall evaluative judgment; it requires an evaluation in terms of the consequences of interpretations and uses of assessment results.
🔑 Definition — Validity as a matter of degree: Validity is not all-or-none; it exists on a continuum (high, moderate, or low).
🔑 Definition — Validity as specific to use: No single assessment is valid for all purposes; validity depends on the intended use or interpretation.
📌 Example: An arithmetic test may have high degree of validity for computational skill and low degree for arithmetical reasoning.
Topic- 039: Evidences of Validity: Content Validity
There are three evidences of validity: content, construct, and criterion. Content validity refers to how well the sample of assessment tasks represents the domain of the tasks to be measured. The procedure for establishing content validity involves comparing the assessment tasks to the specifications describing the task domain under consideration. The method for establishing content validity involves three steps:
- Classroom instruction determines which intended learning outcomes (objectives) are to be achieved by students.
- Achievement domain specifies and delimits a set of instructionally relevant learning tasks to be measured by an assessment.
- Instructional and assessment priorities specify the relative importance of learning objectives to be assessed.
🔑 Definition — Content validity: How well the sample of assessment tasks represents the domain of the tasks to be measured.
💡 Why this matters: Content validity ensures that a test actually covers the material it was designed to measure, not just a limited or irrelevant subset.
📌 Example: If a teacher teaches four units (A, B, C, D) and gives equal instructional time to each, but the test has 80% of questions from unit A and only 5% from each other unit, the test would have low content validity because it does not represent the full domain of instruction.
⭐ Key Takeaways
The three essential characteristics of a good test are validity, reliability, and usability, with validity being the most critical for ensuring a test measures what it intends to measure. Reliability is necessary but not sufficient for validity—a test can be reliable without being valid, but cannot be valid without being reliable. Validity is a matter of degree, specific to particular uses, and a unitary concept requiring overall evaluative judgment. Content validity is established by comparing assessment tasks to the full domain of learning outcomes, and its method involves aligning classroom instruction, achievement domains, and instructional priorities.
🧠 Quick Revision Questions
- What are the three essential characteristics of a good test?
- What is the relationship between reliability and validity?
- Why is validity considered "a matter of degree" rather than all-or-none?
- Explain why an arithmetic test could have high validity for computational skill but low validity for arithmetical reasoning.
- What are the three steps in the method for establishing content validity?
📘 Lecture 12 — Characteristics of a Good Test-II
📖 Overview: This lecture continues the exploration of test quality by focusing on three essential types of validity: construct validity, criterion validity, and consequence validity. These concepts are crucial for ensuring that assessments accurately measure what they claim to measure and produce fair, meaningful results for both students and educators.
🗂️ Topics Covered
This lecture covers three main types of validity evidence. First, it explains construct validity, including its meaning, procedure for development using a test framework, and methods of confirmation through expert judgment and factor analysis. Second, it explores criterion validity, distinguishing between concurrent and predictive validity, along with the procedure and statistical method for verification. Finally, it discusses consequence validity, focusing on how assessment results impact teaching and learning, including considerations and factors within the test itself that can undermine validity.
📝 Lecture Summary
Topic- 040: Evidences of Validity: Construct Validity
Construct validity addresses how well a test measures up to its claims. For example, a test designed to measure depression must only measure that particular construct, not closely related ideas such as anxiety or stress.
The procedure for establishing construct validity involves developing a test framework, which includes:
- Defining the construct
- Identifying sub-constructs
- Listing indicators of each sub-construct
- Writing test items for each indicator
A detailed example is provided for the construct of Essay Writing:
| Sub-construct | Meaning/Scope | Indicators |
|---|---|---|
| Introduction Paragraph | It introduces the main idea, captures the interest of reader, and tells why topic is important. | 1. Single sentence called the thesis statement is written. 2. Background information about your topic provided. 3. Definitions of important terms written. |
| Supporting Paragraphs | Supporting paragraphs make up the main body of your essay. | 1. List the points about the main idea of the essay. 2. Write a separate paragraph for each supporting point. 3. Develop each supporting point with facts, details, and examples. |
| Summary Paragraph | Concluding paragraph comes after you have finished developing your ideas. | 1. Restate the strongest points. 2. Restate the main idea. 3. Give your personal opinion or suggest a plan for action. |
There are two methods to confirm construct validity of a test:
- Expert judgment: Experts in the field assess the construct validity of the table, and the table is revised under their guidance.
- Factor analysis: Questions are grouped by keeping in view the responses of respondents on them.
Topic- 041: Evidences of Validity: Criterion Validity
Criterion validity demonstrates the degree of accuracy of a test by comparing it with another test, measure, or procedure which has been demonstrated to be valid.
There are two types of criterion validity:
- Concurrent validity: This approach allows one to show the test is valid by comparing it with an already valid test.
- Predictive validity: It involves testing a group of subjects for a certain construct, and then comparing them with results obtained at some point in the future.
The procedure compares assessment results with another measure of performance obtained at a later date (for prediction) or with another measure of performance obtained concurrently (for estimating present status).
The method for determining criterion validity involves statistically correlating the two sets of scores. The resulting correlation coefficient provides a numerical summary of the relationship.
📐 Formula: Correlation coefficient → A numerical summary indicating the degree of relationship between two sets of scores. 📌 Example: If a new test for job aptitude is administered, and the scores are correlated with actual job performance ratings obtained six months later, the resulting correlation coefficient indicates the predictive validity of the test.
Topic- 042: Evidences of Validity: Consequence Validity
Consequence validity concerns how well the use of assessment results accomplishes intended purposes and avoids unintended effects.
The procedure involves evaluating the effects of the use of assessment results on teachers and students. Both the intended positive effects (e.g., increased learning) and possible unintended negative effects (e.g., dropout of school) need to be evaluated.
Key considerations include:
- Does the assessment artificially constrain the focus of students' study?
- Does the assessment encourage or discourage exploration and creative modes of expression?
Factors in the test or assessment itself that can undermine validity include:
- Unclear directions
- Reading vocabulary and sentence structure too difficult
- Ambiguity
- Inadequate time limits (construct irrelevant variance)
- Overemphasis of easy-to-access aspects of domain at the expense of important, but hard-to-access aspects
- Test items inappropriate for the outcomes being measured
- Poorly constructed test items
- Test too short
- Improper arrangement of items
- Identifiable pattern of answers
💡 Why this matters: Consequence validity reminds us that even if a test is technically accurate, its use can have harmful side effects that must be monitored and addressed.
⭐ Key Takeaways
The most critical points to remember from this lecture are that construct validity requires a systematic framework defining the construct, its sub-constructs, and indicators before writing test items, and it is confirmed through expert judgment or factor analysis. Criterion validity is established by comparing a test against an already valid measure, with concurrent validity for present status and predictive validity for future performance, using a correlation coefficient to describe the relationship. Consequence validity evaluates both the intended positive and unintended negative effects of assessment use on students and teachers, and many factors within the test itself, such as unclear directions, ambiguity, and inadequate time limits, can undermine overall validity. Finally, a good test must satisfy all three types of validity evidence to be considered truly valid for its intended purpose.
🧠 Quick Revision Questions
- What is the difference between construct validity and criterion validity?
- What are the two methods used to confirm construct validity, and how do they differ?
- Explain the difference between concurrent validity and predictive validity, and provide an example of each.
- What is the purpose of consequence validity, and what two types of effects must be evaluated?
- List at least four factors within the test itself that can undermine validity according to the lecture.
📘 Lecture 13 — Characteristics of a Good Test-III
📖 Overview: This lecture focuses on the concept of reliability in educational assessment, explaining its nature as consistency of measurement. It details various methods for estimating reliability, including test-retest, equivalent forms, and internal consistency approaches, highlighting how each method addresses a specific type of consistency (stability, equivalence, or internal consistency).
🗂️ Topics Covered
The lecture covers three main topics: the nature of reliability, methods of estimating reliability including stability, equivalence, and internal consistency, and the test-retest method in detail. It explains how reliability is a statistical concept ranging from +1 to -1, and emphasizes that while reliability is necessary for validity, it is not sufficient. Seven specific methods for estimating reliability are listed, with the test-retest method discussed in depth, including the importance of the time interval between test administrations.
📝 Lecture Summary
Topic- 043: Nature of Reliability
Reliability refers to the consistency of measurement. It is important to note that reliability refers to the results obtained with an assessment instrument, not to the instrument itself. An estimate of reliability always refers to a particular type of consistency, which can be stability (over time), equivalence (across different forms), or internal consistency (within the assessment itself). Reliability is a necessary but not sufficient condition for validity, meaning a test can be reliable without being valid, but a valid test must be reliable. Reliability is primarily a statistical concept, with coefficients ranging from +1 (perfect positive reliability) to -1 (perfect negative reliability).
💡 Why this matters: Understanding that reliability is a property of test scores, not the test itself, helps teachers interpret assessment results correctly and avoid false conclusions about student performance.
🔑 Definition — Reliability: The consistency of measurement obtained with an assessment instrument.
Topic- 044: Method of Estimating Reliability
There are three main types of consistency that reliability estimates address:
- Stability: Consistency over a period of time.
- Equivalence: Consistency over different forms of assessment.
- Internal consistency: Consistency within the assessment itself.
In determining reliability, it is desirable to obtain two sets of measures under identical conditions and then compare the results. The reliability coefficient resulting from each method must be interpreted according to the type of consistency being investigated.
The following methods are used to estimate reliability:
- Test-Retest (measures stability)
- Equivalent Forms (measures equivalence)
- Test-Retest with Equivalent Forms (measures both stability and equivalence)
- Split Half (measures internal consistency)
- Kuder-Richardson (measures internal consistency)
- Cronbach Alpha (measures internal consistency)
- Inter-rater reliability (measures consistency of rating)
🔑 Definition — Reliability coefficient: A statistical index ranging from +1 to -1 that indicates the degree of consistency in measurement.
Topic- 045: Test-retest Method
The test-retest method is a measure of stability. It involves giving the same test twice to the same group with any time interval between tests. The time interval can range from several minutes to several years.
Example of Test-Retest Procedure:
- September 25: Administer Form A (Item a: yes, Item b: no, Item c: yes)
- October 15: Administer Form A again (Item a: yes, Item b: no, Item c: yes)
- The responses are then correlated to determine stability.
The time interval is a key point in this method:
- A short interval will provide an inflated coefficient of reliability (because students remember their previous answers).
- A very long interval will influence results by instability and actual changes in students over time (e.g., learning or forgetting).
📐 Formula: No explicit formula is provided in the lecture, but the reliability coefficient is computed by correlating the two sets of scores. 📌 Example: A teacher gives a math test on Monday and gives the same test to the same students on Friday. If most students score similarly both times, the test-retest reliability is high. If the interval is only 10 minutes, the correlation will be artificially high (inflated). If the interval is 6 months, the correlation will be low because students have learned and changed.
💡 Why this matters: The choice of time interval directly affects the reliability estimate, so teachers must select an appropriate interval to get a meaningful measure of stability.
⭐ Key Takeaways
The most critical points from this lecture are: reliability refers to consistency of measurement, not the test itself, and it is necessary but not sufficient for validity. There are three main types of consistency: stability, equivalence, and internal consistency. Seven methods exist for estimating reliability, each targeting a specific type of consistency. The test-retest method specifically measures stability by administering the same test twice, and the time interval between administrations is crucial — too short inflates the coefficient, while too long introduces irrelevant changes. Finally, reliability coefficients range from +1 (perfect) to -1 (perfect negative), and all estimates must be interpreted in the context of the type of consistency being investigated.
🧠 Quick Revision Questions
- What does reliability refer to in educational measurement — the instrument itself or the results obtained?
- Name the three types of consistency that reliability estimates can address.
- Why is time interval a critical factor in the test-retest method?
- What is the difference between a short interval and a very long interval in test-retest reliability?
- List four methods that estimate internal consistency.
📘 Lecture 14 — Characteristics of a Good Test-IV
📖 Overview: This lecture continues exploring methods for estimating test reliability, focusing on equivalence, internal consistency, and rater consistency. It covers Equivalent Forms method, Split Half method, Kuder-Richardson methods, and Inter-Rater method, each providing a different approach to establishing that a test yields consistent results.
🗂️ Topics Covered
The lecture covers four main methods of estimating reliability: Equivalent Forms method (including Test-Retest with Equivalent Forms), Split Half method (using the Spearman-Brown formula), Kuder-Richardson methods and Coefficient Alpha (for dichotomous and polytomous scoring), and the Inter-Rater method for judgmental scoring. Each method is classified as a measure of either equivalence, stability and equivalence, internal consistency, or consistency of ratings.
📝 Lecture Summary
Topic- 046: Method of Estimating Reliability: Equivalent Form Method
The Equivalent Forms method is a measure of equivalence. It involves giving two different forms of the same test (Form A and Form B) to the same group of students in close succession, typically on the same day. The correlation between the scores on the two forms provides an estimate of reliability based on equivalence.
🔑 Definition — Equivalent Forms method: A reliability estimation technique where two different but equivalent test forms are administered to the same group in close succession to measure equivalence. 📐 Formula: Correlation between Form A scores and Form B scores → Measures how equivalent the two forms are. 📌 Example: On September 25, a class takes Form A (items a, b, c) and immediately after takes Form B (items d, e, f). The correlation between scores on Form A and Form B estimates equivalence reliability.
The Test-Retest with Equivalent Forms method is a measure of both stability and equivalence. It involves giving two forms of the test to the same group with an increased time interval between administrations, typically days or weeks apart. This approach allows the reliability coefficient to reflect both the consistency of the test forms and the stability of the trait over time.
📌 Example: Students take Form A on September 25 and then take Form B on October 15. A student scores 82 on Form A and 74 on Form B. The correlation between these scores estimates both stability and equivalence.
💡 Why this matters: Equivalent Forms method isolates consistency between different test versions, while Test-Retest with Equivalent Forms adds a time dimension to measure if the trait remains stable over the interval.
Topic- 047: Method of Estimating Reliability: Split Half Method
The Split Half method is a measure of internal consistency. It requires administering the test only once. After scoring, the test is divided into two equivalent halves—typically odd-numbered items and even-numbered items. The correlation between the scores on these two halves is then corrected to estimate the reliability of the full test using the Spearman-Brown formula.
🔑 Definition — Split Half method: A reliability estimation technique where a single test administration is split into two halves (odd/even) and the correlation between halves is corrected to estimate the reliability of the whole test. 📐 Formula: Spearman-Brown formula: ( r_{full} = \frac{2r_{half}}{1 + r_{half}} ) → Corrects the half-test correlation to estimate the full-test reliability. 📌 Example: On September 25, a test is administered once. Items are numbered 1 through 6. The sum of scores on odd items (1, 3, 5) = 40, and the sum on even items (2, 4, 6) = 42. Total score = 82. The correlation between odd and even halves is computed, then corrected using the Spearman-Brown formula to estimate the full test reliability.
Split half reliabilities tend to be higher than equivalent form reliabilities because the split half method is based on a single administration of the assessment, eliminating time-related and form-variability errors.
Topic- 048: Method of Estimating Reliability: Kuder-Richardson Method
The Kuder-Richardson methods and Coefficient Alpha are measures of internal consistency. Like the Split Half method, they require only one test administration and score the total test. However, unlike Split Half, these formulas do not require splitting the assessment into halves for scoring purposes.
The KR20 formula is applicable only when student responses are scored dichotomously (0 or 1), meaning items are scored as correct or incorrect. It is most useful with traditional test items that have right/wrong answers.
The generalization of KR20 for assessments that have more than dichotomous, right-wrong scores (e.g., Likert scales or partial credit) is called Coefficient Alpha (also known as Cronbach's alpha).
🔑 Definition — KR20 (Kuder-Richardson Formula 20): A measure of internal consistency reliability for dichotomously scored items, based on the total test score and item variance. 🔑 Definition — Coefficient Alpha: A generalization of KR20 for assessments with polytomous (more than two) scoring categories, measuring internal consistency.
The Inter-Rater method is a measure of consistency of ratings. It is used when test responses require judgmental scoring (e.g., essays, performances). This method gives a set of student responses to two or more raters and has them independently score the responses. The correlation between the ratings from different raters provides an estimate of inter-rater reliability.
🔑 Definition — Inter-Rater method: A reliability estimation technique where two or more raters independently score the same set of student responses to measure consistency of judgments across raters.
⭐ Key Takeaways
The four methods of estimating reliability serve different purposes: Equivalent Forms measures equivalence of two test versions; Test-Retest with Equivalent Forms measures both stability over time and equivalence; Split Half and Kuder-Richardson methods (including Coefficient Alpha) measure internal consistency from a single administration; and Inter-Rater method measures consistency across different scorers. The Spearman-Brown formula is essential for correcting split half correlations to estimate full test reliability. KR20 is specifically for dichotomous (right/wrong) scoring, while Coefficient Alpha extends this to polytomous scoring. Split half reliabilities are typically higher than equivalent form reliabilities because they avoid errors from different test forms and time intervals. Understanding which reliability method to apply depends on the test design, scoring type, and whether stability, equivalence, internal consistency, or rater consistency needs to be measured.
🧠 Quick Revision Questions
- What is the difference between the Equivalent Forms method and Test-Retest with Equivalent Forms?
- Why are split half reliabilities typically higher than equivalent form reliabilities?
- What is the Spearman-Brown formula used for, and what does it assume about the two halves?
- For what type of scoring is KR20 applicable, and what is its generalization called?
- When would you use the Inter-Rater method instead of the other reliability estimation methods?
Here is the summary of Lecture 15, formatted exactly as requested.
📘 Lecture 15 — TYPES OF ASSESSMENT TOOLS-I
📖 Overview: This lecture introduces assessment tools that go beyond standard paper-pencil tests to evaluate learning outcomes in behavioral and social domains. It focuses on Anecdotal Records, explaining their purpose, effective use, advantages, and limitations for observing and recording student behavior in natural settings.
🗂️ Topics Covered
The lecture covers the types of assessment tools for non-cognitive outcomes, including observation, peer appraisal, self-appraisal, and portfolios. It then delves into Anecdotal Records as a method for factual observation, followed by guidelines for their effective use, and concludes with a discussion of their advantages and limitations.
📝 Lecture Summary
Topic- 049: Anecdotal Records
Many learning outcomes in the cognitive domain are measured by paper-pencil tests, but other outcomes require informal observation of natural interactions. Learning outcomes can be assessed by: observing students (Anecdotal record), asking peers (Peer appraisal), questioning directly (Self-appraisal), and measuring progress by recorded work (portfolio). Since impressions from observation can be biased, an anecdotal record provides a method for accurate documentation. These records are factual descriptions of meaningful incidents and events that a teacher observes.
🔑 Definition — Anecdotal Records: Factual descriptions of meaningful incidents and events that the teacher observes.
Topic- 050: Effective Use of Anecdotal Records
To use anecdotal records effectively, one must keep these points in mind:
- Determine in advance what to observe but remain alert to unusual behavior.
- Analyze observational records for possible sources of bias.
- Observe and record enough of the situation to make the behavior meaningful.
- Record the incident as soon after the observation as possible.
- Limit each anecdote to a brief description of a single incident.
- Keep the factual description of the incident and your interpretation of it separate.
- Record both positive and negative behavioral incidents.
- Collect a number of anecdotes on a student before drawing inferences about typical behavior.
- Obtain practice in writing anecdotal records.
Topic- 051: Advantages and Limitations of Anecdotal Records
The advantages of anecdotal records are:
- It depicts actual behaviors in natural situations.
- It facilitates gathering evidence on events that are exceptional but significant.
- It is beneficial for students with less communication skills.
The limitations of anecdotal records are:
- It takes a long time to maintain.
- It is subjective in nature.
- Anxiety may lead to wrong observation.
⭐ Key Takeaways
The most critical points from this lecture are that anecdotal records are factual, not interpretive, descriptions of observed student behavior in natural settings. To be effective, one must separate the factual record from their own interpretation, record both positive and negative behaviors, and collect multiple anecdotes before drawing conclusions. While invaluable for assessing students with poor communication skills and capturing significant events, anecdotal records are time-consuming, subjective, and can be influenced by the observer's anxiety.
🧠 Quick Revision Questions
- What is the primary purpose of using an anecdotal record instead of just relying on general impressions?
- List three key guidelines for making the use of anecdotal records effective.
- Why is it important to separate the factual description from the interpretation of the incident in an anecdotal record?
- Name one significant advantage of using anecdotal records for students who struggle with communication.
- What are the three main limitations of anecdotal records mentioned in the lecture?
📘 Lecture 16 — Types of Assessment Tools-II
📖 Overview: This lecture explores two major types of assessment tools: peer appraisal (including the guess-who and sociometric techniques) and student portfolios. It explains how these tools can be used for both instructional and assessment purposes, highlighting their strengths, weaknesses, and practical implementation guidelines.
🗂️ Topics Covered
This lecture covers three main topics: first, peer appraisal techniques including the guess-who technique and sociometric technique; second, the portfolio as a systematic collection of student work, including its key steps, strengths, and weaknesses; third, the purposes of portfolios, distinguishing between instructional and assessment purposes, and explaining showcase versus documentation portfolios and finished versus working portfolios.
📝 Lecture Summary
Topic- 052: Peer Appraisal
In this procedure, students rate their peers on the same rating device used by their teacher. It depends on greatly simplified procedures. There are two widely used techniques in this area: guess who technique and sociometric technique.
🔑 Definition — Peer Appraisal: A procedure in which students rate their peers using the same rating device as the teacher.
Guess Who Technique
In this technique, a teacher uses a positive or negative behavior of a student as an example. Other students from the same group try to guess the statement with that characteristic correctly. Generally, behaviors used are positive in nature to avoid any adverse effect on the student pointed out in the example. The guess-who technique is based on the nomination method of obtaining peer ratings and is scored by simply counting the number of mentions each student receives on each description.
📌 Example: A teacher might say, "This student always helps others when they are struggling with a task." Students then guess which classmate fits this description. The guess is scored by counting how many times each student is mentioned positively.
Sociometric Technique
Sociometric technique is a method for assessing the social acceptance of individual students and the social structure of the group. It is based on students' choice of companion for any group situation. This form was used to measure student's acceptance as seating companions, work companions, and play companions.
There are few important principles of sociometric choosing:
- The choices should be real choices that are a natural part of classroom activities.
- The basis for the choice and restriction on the choosing should be made clear.
- All students should be equally free to participate in the activity or situation.
- The choice of each student must be kept confidential.
- The choices should actually be used to organize or rearrange the group.
Topic- 053: Portfolio
Systematic collection of students' work into portfolios can serve a variety of instructional and assessment purposes. The value of portfolios depends heavily on the clarity of purpose, the guidelines for the inclusion of materials, and the criteria to be used in evaluating the portfolio.
🔑 Definition — Portfolio: A purposeful collection of pieces of student's work selected to serve a particular purpose, such as documentation of student growth.
Key steps in defining and using portfolios:
- Specify purpose
- Provide guidelines for selecting portfolios
- Define student's role in selection and self-evaluation
- Specify evaluation criteria
- Use portfolios in instruction and communication
Strengths of portfolios:
- They can be readily integrated with instruction.
- Provide opportunity for students to show what they can do.
- Encourage students to become reflective learners.
- Help in setting goals and self-evaluation.
- Help teacher and student to collaborate and reflect on student's progress.
- Effective way to communicate with parents.
- Provide mechanism for student-centered and student-directed conferences with parents.
- Provide concrete examples of student's development and current skills.
Weaknesses of portfolios:
- Can be time consuming to assemble.
- Hard to use in summative assessment.
- Difficult to compare results.
- Very low reliability.
Topic- 054: Purpose of Portfolio
Fundamentally, there are two global purposes for creating portfolios of student work: for student's assessment and instruction. It can be used to showcase student's accomplishment and document the progress.
Instructional purposes: When the primary purpose is instruction, the portfolio might be used as a means of:
- Helping students develop and refine self-evaluation skills.
- Providing teacher with more reflecting information regarding students' progress.
- Setting criteria of excellence between teacher and student.
- Student-directed conferences with parents.
- Access to student thought process and awareness of standards.
- Teaching students to communicate with different audiences.
Assessment purposes: When emphasis is on assessment, it is important to distinguish between formative and summative roles of assessment.
- It can be used for formative purposes to measure progress.
- Basis for certifying accomplishment.
- For system accountability mechanism.
Current accomplishment and progress: When the focus is on accomplishments, portfolios usually are limited to finished work and may cover only a relatively small period of time. When focus is on demonstrating growth and development, the time frame is longer. It will include multiple versions of the same work over time to measure progress.
🔑 Definition — Showcase portfolio: Contains student-selected entries demonstrating students' ability to choose their best work which demonstrates their ability to do a task. It provides evidence about breadth as well as depth of learning.
🔑 Definition — Documentation portfolio: A more inclusive portfolio not just limited to special strengths of the student, intended to show growth over time.
🔑 Definition — Finished portfolio: Implies that work is complete for a specific audience (e.g., a job application portfolio).
🔑 Definition — Working portfolio: Contains multiple versions of the same work over time to show progress, with a longer time frame.
Guidelines for portfolio entries: Guidelines should specify:
- The uses that will be made of the portfolio.
- Who will have access to it.
- What type of work is appropriate to include.
- What criteria will be used in evaluating the work.
- Should define timeline for the portfolios.
- Minimum and maximum numbers of entries.
💡 Why this matters: Understanding the distinction between different types of portfolios (showcase vs. documentation, finished vs. working) is crucial for teachers to select the appropriate assessment tool based on whether the goal is to measure current accomplishment or demonstrate growth over time. This directly impacts grading, reporting, and student motivation.
⭐ Key Takeaways
Peer appraisal techniques (guess-who and sociometric) allow students to evaluate each other using simplified procedures, with the guess-who technique based on nomination counting and sociometric technique based on assessing social acceptance through companion choices. Portfolios are purposeful collections of student work that can serve both instructional and assessment purposes, with strengths including integration with instruction and fostering reflective learning, but weaknesses including time consumption and low reliability. Portfolios serve two main purposes: instructional (developing self-evaluation, reflecting on progress) and assessment (formative progress measurement, certification, accountability). Teachers must clearly specify purpose, guidelines, and evaluation criteria for portfolios, and they must decide whether the focus is on current accomplishments or progress over time when designing portfolio assignments.
🧠 Quick Revision Questions
- What are the two widely used techniques of peer appraisal, and how do they differ in what they measure?
- List the five important principles of sociometric choosing.
- What are the key steps in defining and using portfolios?
- What are the four weaknesses of portfolios mentioned in the lecture?
- What is the difference between a showcase portfolio and a documentation portfolio?
📘 Lecture 17 — Creating Fixed-Choice Test Items: MCQs-I
📖 Overview: This lecture introduces the principles of selecting and constructing fixed-choice test items, focusing specifically on Multiple Choice Questions (MCQs). It explains the criteria for choosing item formats, the structure and characteristics of MCQs, and their various uses in measuring different levels of learning outcomes, from basic knowledge to more complex understanding.
🗂️ Topics Covered
This lecture covers the selection of item format based on the test blueprint and skills to be measured, distinguishing between selection and supply type objective items. It details the characteristics of MCQs including the stem and alternatives (answer and distractors), the formats of direct questions versus incomplete statements, and the best answer vs. correct answer type. The lecture concludes by exploring the uses of MCQs for measuring knowledge outcomes, specifically terminology, principles, and methods.
📝 Lecture Summary
Topic- 055: Selection of Item in a Test
The selection of item format is determined by the table of specifications and the test blueprint, which indicate the kinds of skills and the balance of test content. A teacher decides the most appropriate item type (objective vs. subjective) based on the nature of the content, cognitive processes, and student mental level. The choice must be based on the behavior to be tested, not personal preference. For instance, MCQs are suitable for testing knowledge of English mechanics in large groups but not for directly measuring writing skill, which requires an essay format.
There are two main types of objective type items:
- Selection Type: Includes true-false, matching exercises, and multiple choice items.
- Supply Type: Includes completion (fill in the blanks) and short answer items.
The selection of any one or combination of these items is based on their suitability for measuring the desired learning outcome.
Topic- 056: Characteristics of MCQs
Multiple choice items are recognized as the most widely applicable and useful type of objective test item, capable of measuring knowledge, comprehension, and application level learning outcomes. An MCQ consists of two parts:
- Stem: The problem stated as a direct question or an incomplete statement.
- Alternatives (Options/Choices): The list of suggested solutions. The correct alternative is called the answer, while all incorrect or less appropriate alternatives are called distractors or foils.
The format of the stem can be a direct question or an incomplete statement. The direct question format is easier to write, more natural for younger students, and more likely to present a clearly formulated problem.
🔑 Definition — Stem: The part of a multiple-choice item that presents the problem, which can be a direct question or an incomplete statement. 🔑 Definition — Distractors: The incorrect or less appropriate alternatives in a multiple-choice item, designed to distract students who have not achieved the learning outcome.
📌 Example of Direct Question: Which of the following cities is the capital of Pakistan? a) Islamabad b) Karachi c) Lahore d) Quetta
📌 Example of Incomplete Statement: The capital of Pakistan is: a) Islamabad b) Karachi c) Lahore d) Quetta
A common procedure is to start each stem as a direct question and shift to the incomplete statement form only when clarity can be retained and greater conciseness achieved.
For measuring more complex learning outcomes (understanding, application, interpretation) the best answer type MCQ is used. This type is useful when answers have varying degrees of acceptability, and the student must select the best one.
📌 Example of Best Answer Type MCQ: Which of the following factors contributed most to the selection of Islamabad as capital of Pakistan? a) Central location b) Good climate c) Large population d) Good highways
💡 Why this matters: The best answer type MCQs are more difficult than the correct answer type because they require students to evaluate and compare options, measuring a more complex level of learning.
Topic- 057: Uses of MCQs –I
The MCQ is the most versatile type of test item, capable of measuring learning outcomes from simple to complex and adaptable to most types of content. However, it cannot measure all outcomes, such as the ability to organize and present ideas. For such skills, constructed and restricted response questions are used.
Uses of MCQs in measuring knowledge outcomes: MCQs can measure a variety of knowledge outcomes, which are prominent in all school subjects.
- Knowledge of Terminology: Students can show their knowledge of a term by selecting a word with the same meaning or by identifying the best definition of a given term.
📌 Example (Synonym): Which of the following words has same meaning as the word Egress? a) Depress b) Enter c) Exit d) Regress
📌 Example (Definition): Which of the following statements best defines the word egress? a) An expression of disapproval b) An act of leaving an enclosed place. c) Proceeding to higher level
Topic- 058: Uses of MCQs –II
MCQs can also be used to measure more complex types of knowledge.
- Knowledge of Principles: This is an important learning outcome in most school subjects. MCQs can be constructed to measure knowledge of principles as easily as they measure facts.
📌 Example: The principle of capillary action helps explain how fluids: a) Enter solution of lower concentration b) Rise in fine tube c) Escape through small openings
- Knowledge of Methods: This includes diverse areas such as knowledge of lab procedures, communication methods, and computational skills. It can be used to measure procedural knowledge before practice or as an important learning outcome in its own right.
📌 Example: To make legislation the prime minister of Pakistan must have the consent of: a) Parliament b) Ministry of law c) Military command d) Supreme court
These examples illustrate the range of MCQs for measuring knowledge outcomes, but MCQs can also be used to measure much more complex types of knowledge.
⭐ Key Takeaways
The choice of item format (such as multiple-choice, essay, or true-false) must be driven by the specific learning outcomes being measured and the test blueprint, not personal preference. An MCQ consists of a stem (a direct question or incomplete statement) and several alternatives, including one correct answer and several distractors. For complex outcomes requiring understanding, the "best answer" type MCQ is more appropriate than the simple "correct answer" type. MCQs are highly versatile and can effectively measure knowledge outcomes like terminology, principles, and methods, but they are not suitable for assessing skills like organizing and presenting ideas.
🧠 Quick Revision Questions
- What are the two main types of objective test items, and which category do MCQs fall into?
- Name the two parts of a multiple-choice item and define each.
- What is the key difference between a "correct answer" MCQ and a "best answer" MCQ?
- Provide one example of an MCQ that measures "knowledge of terminology" and one that measures "knowledge of principles."
- Why is the direct question format often preferred over the incomplete statement format for the stem?
📘 Lecture 18 — Creating Fixed-Choice Test Items: MCQs-II
📖 Overview: This lecture explores the uses, advantages, and limitations of Multiple-Choice Questions (MCQs) in educational assessment. It explains how MCQs can effectively measure knowledge of specific facts and provides a balanced analysis of their benefits and drawbacks compared to other test item formats.
🗂️ Topics Covered
The lecture covers three main areas: the use of MCQs for testing specific factual knowledge with examples from history and science; the advantages of MCQs including reduced ambiguity and guessing compared to short-answer and true/false items; and the limitations of MCQs, such as time-consuming construction and the risk of testing only recall.
📝 Lecture Summary
Topic- 059: Uses of MCQs-III
Another learning outcome basic to all school subjects is the knowledge of specific facts. It provides the basis for developing understanding, thinking skills and other complex learning. MCQs designed to measure specific facts can take many forms, that questions of who, what, when, and where variety are most common.
🔑 Definition — Knowledge of specific facts: The foundational learning outcome that involves recalling concrete information such as names, dates, and events, providing the basis for higher-order thinking.
📌 Example 1 (Who): Who was 1st astronaut to land on moon? a) Buzz Aldrin b) Neil Armstrong c) Yuri Gagarin
📌 Example 2 (What): What was the name of space shuttle that landed on the moon? a) Apollo b) Atlas c) Midas d) Polaris
📌 Example 3 (When): When did first man landed on the moon? a) 1962 b) 1966 c) 1967 d) 1969
💡 Why this matters: Factual knowledge MCQs test the recall of specific details, which is essential for building a foundation before students can engage in analysis or application.
Topic- 060: Advantages and Limitations of MCQs –I
The MCQ is one of the most widely applicable test items for measuring knowledge, achievement. It can effectively measure various types of knowledge and complex learning outcomes. It is free from some of the common shortcomings which are characteristics of the other test items like ambiguity and vagueness usually associated with short questions.
🔑 Definition — Ambiguity: A flaw in test items where multiple interpretations or correct answers are possible, making the item unclear and invalid for measurement.
📌 Example of vague question (poorly constructed): Quid -e -Azam was born in ______. Problem: There can be multiple correct answers for this question (e.g., 1876, in Karachi).
📌 Example of good MCQ (resolves ambiguity): Quid -e -Azam was born in: a) Karachi b) Lahore c) Peshawar d) Dehli
💡 Why this matters: By providing fixed answer choices, MCQs eliminate the problem of vague short-answer items where students might give different but equally correct responses.
Topic- 061: Advantages and Limitations of MCQs –II
MCQs reduce the risk of guessing the correct answer. You have to know the correct answer. There is a high chance of wrong answer if you solely depend on guesses.
🔑 Definition — Risk of guessing: The probability that a student can obtain a correct answer by chance rather than knowledge, which is higher in true/false items than in MCQs with multiple options.
📌 Example of True/False (high guessing risk): Quid -e -Azam was born in 1867. True/False Problem: The student will receive score even if he didn‘t know the correct year of birth if he/she tick false (50% chance of being correct).
📌 Example of MCQ (reduced guessing risk): In which year Quid- e- Azam was born? a) 1867 b) 1876 c) 1878 d) 1887 Effect: MCQs items reduced the probability of guessing as compared to other form of item (only 25% chance of random correct answer).
💡 Why this matters: With four options instead of two, MCQs significantly lower the influence of lucky guessing, making scores more reflective of actual knowledge.
Topic- 062: Advantages and Limitations of MCQs –III
Advantages of MCQs:
- Ensure objectivity, reliability and validity; preparations of questions with colleagues provide constructive criticism.
- Increase significantly the range and variety of facts that can be sampled in given time.
- Provide precise and unambiguous measurement of the higher intellectual processes.
- Provide detailed feedback for both students and teachers.
- MCQs are easy and rapid to score.
Limitations of MCQs:
- Take long time to construct in order to avoid arbitrary and ambiguous questions.
- Require careful preparation to avoid multitude of questions testing only recall.
🔑 Definition — Objectivity: The quality of a test where results are not influenced by the scorer's subjective judgment; MCQs have clear right/wrong answers. 🔑 Definition — Constructive criticism: Feedback from colleagues during item development that helps improve question quality and eliminate flaws.
💡 Why this matters: While MCQs offer strong measurement properties and efficiency, their construction demands significant expertise and time to prevent them from becoming superficial recall tests.
⭐ Key Takeaways
The MCQ is a highly efficient and objective tool for measuring factual knowledge and higher-order thinking, but it is not without challenges. Its key strength lies in eliminating the ambiguity and vagueness common in short-answer items while significantly reducing the guessing advantage seen in true/false items. However, constructing high-quality MCQs requires meticulous preparation and time to avoid over-reliance on simple recall. The format's ability to sample a wide range of content quickly makes it ideal for comprehensive assessments, though collaboration with colleagues is essential for maintaining validity and reliability.
🧠 Quick Revision Questions
- Why are MCQs considered better than true/false items for reducing the impact of guessing?
- What is the main problem with the vague short-answer item "Quid-e-Azam was born in ______"?
- List three specific types of factual knowledge that MCQs can test (using the who/what/when framework).
- State two major advantages and two major limitations of MCQs mentioned in this lecture.
- How does providing multiple answer choices in an MCQ improve the measurement of student knowledge compared to a fill-in-the-blank format?
📘 Lecture 19 — Creating Fixed-Choice Test Items: MCQs-III
📖 Overview: This lecture provides essential guidelines for constructing high-quality multiple-choice questions (MCQs). It focuses on improving item stems, avoiding negative phrasing, ensuring grammatical consistency, and removing irrelevant clues to enhance test validity and reliability.
🗂️ Topics Covered
The lecture covers four key suggestions for constructing effective MCQs: making the stem meaningful and self-contained, including as much content as possible in the stem while avoiding irrelevant material, avoiding negative statements unless necessary, and ensuring all alternatives are grammatically consistent with the stem.
📝 Lecture Summary
Topic- 063: Suggestions for Constructing MCQs -I
The general applicability and superior qualities of multiple choice test items are realized most fully when care is taken in their construction. This involves formulating clearly stated problems, identifying plausible alternatives, and removing irrelevant clues to the answer. The stem of the item should be meaningful by itself and should present a definite problem. Often, stems placed in MCQ form are incomplete statements that make little sense until all alternatives have been read; this is not a proper MCQ but rather a true-false question placed in MCQ form.
🔑 Definition — Stem: The introductory part of a multiple-choice item that presents the problem or question to be answered.
📌 Example (Poor item): South America a) Is flat, arid country b) Imports coffee from the United States c) Has larger population than Europe d) Was settled by colonists from Spain
📌 Example (Better item): Most of South America was settled by colonists from a) England b) France c) Holland d) Spain
💡 Why this matters: A self-contained stem reduces ambiguity and allows students to understand the question before reading the options, improving measurement accuracy.
Topic- 064: Suggestions for Constructing MCQs –II
The item stem should include as much of the item as possible and should be free of irrelevant material. A clear stem increases the probability of the item being understood correctly and reduces the reading time required.
📌 Example (Poor item): Most of the Indian subcontinent was settled by colonists from Britain. How would you account for the large number of colonists settling there? a) They are adventurers b) They were in search of wealth c) They wanted lower taxes d) They were seeking religious freedom
📌 Example (Better item): Why did Britishers settle in India? a) For adventures b) For wealth c) For lower taxes d) For religious freedom
💡 Why this matters: Including all necessary information in the stem and removing irrelevant material reduces cognitive load and prevents confusion, making the item more efficient and valid.
Topic- 065: Suggestions for Constructing MCQs –III
Try to avoid negative statements, unless the significant learning outcome requires it. Negative statements increase the possibility of students overlooking words like "no" or "least" and similar words used in negative items.
📌 Example (Poor item): Which of the following cities is not located in north Islamabad? a) Abbottabad b) Gilgit c) Lahore d) Mingora
📌 Example (Better item — rephrased to test the same knowledge positively): Which of the following cities is located in south Islamabad? a) Abbottabad b) Gilgit c) Lahore d) Mingora
💡 Why this matters: Negative phrasing can confuse students and introduce measurement errors unrelated to the actual knowledge being tested. Use negatives only when the learning outcome specifically requires it.
Topic- 066: Suggestions for Constructing MCQs –IV
All alternatives should be grammatically consistent with the stem of the item. The main function of this rule is to prevent irrelevant clues from entering the item. In the following examples, note how the better version results from a change in the alternatives to obtain grammatical consistency.
📌 Example (Poor item — grammatical inconsistency): An electric transformer can be used a) For strong electricity b) To increase the voltage of alternating current c) It converts electrical energy into mechanical energy d) Alternating current is changed to direct current
📌 Example (Better item — grammatically consistent): An electric transformer can be used to a) Produce strong electricity b) Increase the voltage of alternating current c) Convert electrical energy into mechanical energy d) Change alternating current to direct current
💡 Why this matters: Grammatical inconsistency can give away the correct answer (or eliminate incorrect options) based on grammar alone, rather than on content knowledge, reducing the item's validity.
⭐ Key Takeaways
Students must remember that high-quality MCQs require a self-contained stem presenting a clear problem, with as much content as possible included in the stem and irrelevant material removed. Negative statements should be avoided unless the learning outcome demands them, as they introduce confusion and potential errors. All alternatives must be grammatically consistent with the stem to prevent irrelevant grammatical clues from revealing the answer. The core principle is that every aspect of item construction should focus on measuring student knowledge of the content, not their ability to decode poorly written items. Applying these four suggestions significantly improves the reliability and validity of MCQ-based assessments.
🧠 Quick Revision Questions
-
Why should the stem of an MCQ be meaningful by itself, and what is the problem with incomplete stems that require reading all alternatives first?
-
What is the benefit of including as much of the item as possible in the stem and removing irrelevant material?
-
Under what circumstances is it acceptable to use negative statements in MCQs, and why should they generally be avoided?
-
How does grammatical inconsistency between the stem and alternatives create irrelevant clues in an MCQ?
-
Rewrite the following poorly constructed MCQ to follow the guidelines in this lecture: "Electricity a) Is dangerous b) Can be used to power lights c) It flows through wires d) Voltage is measured in volts"
📘 Lecture 20 — Creating Fixed-Choice Test Items: MCQs-IV
📖 Overview: This lecture continues the series on constructing effective multiple-choice questions (MCQs), providing five additional practical suggestions for improving item quality. These guidelines focus on ensuring only one correct answer, making distractors plausible, avoiding verbal clues, and equalizing answer lengths, all of which are critical for creating fair and valid assessments.
🗂️ Topics Covered
The lecture covers four main topics: ensuring only one correct or clearly best answer per item (Topic 067), making all distractors plausible to confuse the uninformed (Topic 068), avoiding verbal association between the stem and the correct answer (Topic 069), and ensuring the relative length of alternatives does not provide a clue to the answer (Topic 070). Each topic is illustrated with a poor example and a better example.
📝 Lecture Summary
Topic- 067: Suggestions for Constructing MCQs –V — An item should contain only one correct or clearly best answer.
An item should contain only one correct or clearly best answer. Including more than one correct answer and asking students to select all correct alternatives has two shortcomings: (a) such items are usually no more than a collection of true-false items presented in MCQ form, and (b) the number of alternatives selected as correct answers varies from one student to another, making scoring inconsistent.
🔑 Definition — Single Correct Answer: A test item should have exactly one correct or clearly best answer among the alternatives. 📌 Example (Poor): "Pakistan borders on: a) India b) Tajikistan c) Saudi Arabia d) China" — this has two correct answers (India and China). 📌 Example (Better): The same options are presented as separate True/False statements: "Pakistan borders on: a) India T/F b) Tajikistan T/F c) Saudi Arabia T/F d) China T/F"
Topic- 068: Suggestions for Constructing MCQs –VI — All distractors should be plausible.
All distractors should be plausible. The purpose of a distractor is to confuse the uninformed student. To the student who has not achieved the learning outcome being tested, the distractor should be as attractive as the correct answer. If properly constructed, each distractor will be selected by some students. If a distractor is not selected by anyone, it is not contributing to the functioning of the item and should be eliminated or revised.
🔑 Definition — Plausible Distractor: An incorrect alternative that appears attractive to a student who has not mastered the learning outcome. 📌 Example (Poor): "Who wrote national anthem of Pakistan? a) Allama Iqbal b) Christopher Columbus c) Hafeez Jullandhuri d) Ibrar ul Haq" — "Christopher Columbus" is not plausible. 📌 Example (Better): "Who wrote national anthem of Pakistan? a) Allama Iqbal b) Habib Jalib c) Hafeez Jullandhuri d) Munir Niazi" — all are Pakistani poets, making each distractor plausible. 💡 Why this matters: A non-functioning distractor (selected by no one) wastes a potential slot for a more effective distractor and reduces item discrimination.
Topic- 069: Suggestions for Constructing MCQs –VII — Verbal association between the stem and the correct answer should be avoided.
Verbal association between the stem and the correct answer should be avoided. Frequently, a word in the correct answer provides an alternative clue because it looks or sounds like a word in the stem. However, words similar to those in the stem might be included in the distractors to increase their plausibility. Students who depend on rote memory and verbal association will then be led away from, rather than to, the correct answer.
🔑 Definition — Verbal Association: A clue where the correct answer contains a word that looks or sounds like a word in the stem, allowing students to guess correctly without knowing the content. 📌 Example (Poor): "Which of the following agencies should you contact to find about a flood? a) National flood relief b) Local radio station c) Pakistan post d) Pakistan weather bureau" — "flood" in stem matches "flood" in option (a). 📌 Example (Better): "Which of the following agencies should you contact to find about a flood? a) Disaster management office b) Radio station c) Post office d) Weather bureau" — no word matches the stem. 💡 Why this matters: Avoiding verbal association forces students to rely on actual knowledge rather than surface-level word matching.
Topic- 070: Suggestions for Constructing MCQs –VIII — The relative length of the alternatives should not provide a clue to the answer.
The relative length of the alternatives should not provide a clue to the answer. The best we can hope for is to make alternatives approximately equal in length. Because the correct answer usually needs to be qualified, it tends to be longer than the distractors unless a special effort is made.
🔑 Definition — Equal Length Alternatives: The correct answer should not be systematically longer or shorter than the distractors, as test-takers may guess that the longest option is correct. 📌 Example (Poor): "What is the major purpose of United Nations? a) To maintain peace among people of the world b) To establish international law c) To provide military control d) To form new governments" — option (a) is noticeably longer. 📌 Example (Better): "What is the major purpose of United Nations? a) To maintain peace among people of the world b) To develop new system of international law c) To provide military control of new nations d) To establish democratic forms of governments" — all options are similar in length.
⭐ Key Takeaways
For effective MCQs, always ensure a single correct or clearly best answer—avoid collections of true-false items. Make every distractor plausible and attractive to uninformed students; if a distractor is never selected, revise or delete it. Prevent verbal association clues by not repeating stem words in the correct answer; instead, put similar words in distractors. Finally, equalize the length of all alternatives so the correct answer is not systematically longer or shorter. These five rules (Topics 067-070) collectively improve item validity, reduce guessing cues, and ensure that test scores reflect true learning outcomes.
🧠 Quick Revision Questions
- Why should an MCQ contain only one correct answer, and what are two shortcomings of items with multiple correct answers?
- What is a "plausible distractor," and what should you do if a distractor is never selected by any student?
- Give an example of verbal association between a stem and correct answer, and explain how to fix it.
- Why might the correct answer tend to be longer than distractors, and how can this be avoided?
- For Topic 067, the "poor" item "Pakistan borders on: a) India b) Tajikistan c) Saudi Arabia d) China" had two correct answers. What is the recommended solution for turning this into a better item?
📘 Lecture 21 — Creating Fixed-Choice Test Items: True/ False and Its Uses-I
📖 Overview: This lecture introduces True/False (alternate-response) test items as a type of fixed-choice assessment. It covers their definition, various uses for measuring simple learning outcomes, advantages and limitations, and specific guidelines for constructing high-quality items that avoid ambiguity and trivial content.
🗂️ Topics Covered
The lecture covers True/False items as declarative statements requiring students to select between two alternatives (true/false, yes/no, fact/opinion). It discusses multiple uses including measuring factual knowledge, distinguishing fact from opinion, recognizing cause-effect relationships, and simple logic. The advantages (efficiency for lower-level outcomes, broad coverage) and limitations (guessing, trivial content, limited to lower levels) are presented, followed by essential construction guidelines focusing on avoiding broad generalizations and trivial statements.
📝 Lecture Summary
Topic- 071: True/ False Items
Alternate-response test items consist of a declarative statement that the student is asked to mark true or false, right or wrong, correct or incorrect, or the like.
These items are used for measuring relatively simple learning outcomes by using a single declarative statement with various response methods.
📌 Example (True/False format): Directions: Read each of the following statements. If the statement is true, encircle T and if false, encircle F.
- A river is bigger than a stream. T F
- Founded is the past tense of found. T F
- Dozen is equivalent to 20. T F
📌 Example (Yes/No format): Directions: Read each statement. If the answer is yes, encircle Y; if no, encircle N.
- Is 51% of 38 more than 19? Y N
- Is 50% of 4/10 equal to 2/5? Y N
- If 60% of a number is 9, is the number smaller than 9? Y N
One of the most useful functions of True/False items is in measuring the students’ ability to distinguish fact from opinion.
📌 Example (Fact/Opinion format): Directions: Read each statement. If it is a fact, encircle F; if it is an opinion, encircle O.
- Current constitution of Pakistan is written in 1973. F O
- Pakistan progressed most under dictator rule. F O
- 18th amendment decentralized the ministry of education. F O
Topic- 072: Uses of True/ False Items
Another aspect of understanding measured by True/False items is the ability to recognize cause and effect relationship. This item type usually contains true propositions in both parts, and the student judges whether the relationship between them is true or false.
📌 Example (Cause-effect format): Directions: Both parts of each statement are true. Decide whether the second part explains why the first part is true. If it does, encircle Yes; if not, encircle No.
- Leaves are essentials because they shade the tree trunk. Yes No
- Some plants do not need sunlight because they get food from other plants. Yes No
True/False items can also measure simple aspects of logic, requiring students to evaluate both the original statement and its converse.
📌 Example (Logic format with converse): Directions: For each statement, if true encircle T; if false encircle F. Also, if the converse is true circle CT; if the converse is false circle CF. Give 2 answers for each statement.
- All trees are plants. T F CT CF
- All parasites are animals. T F CT CF
- All eight-legged animals are spiders. T F CT CF
Topic- 073: Advantages and Limitations of True-False
Alternative form questions require students to select any one of two given categories (true-false, yes-no, correct-incorrect, fact-opinion). These items are most suitable for measuring lower level learning outcomes.
Advantages of Alternative Form (True-False):
- Well suited for testing lower-level outcomes, such as:
- Identifying correctness of factual statements (e.g., "Earth is a planet").
- Definitions of terms (e.g., "Photosynthesis is the process by which leaves make food for plants").
- Statement of principles (e.g., "Earth is revolving around the sun").
- Distinguishing facts from opinion (e.g., "Islam is the official religion of Pakistan").
- Recognizing cause-and-effect relationships.
- From a teacher's perspective, useful when:
- A lot of content must be covered in a short time.
- Time available for scoring is very short.
Limitations of Alternative Form (True-False):
- Not all learning outcomes can be measured; generally limited to lower level learning outcomes.
- Ease of guessing correct answers — with only two choices, a student can expect to guess correctly on half of the items for which answers are not known.
- Tendency to take quotations from text with minor wording changes.
- Tendency to include trivial material from the text.
- Items are prone to high guessing and can only be used for measuring lower-level outcomes.
Topic- 074: Suggestions for Constructing True-False Items – I
The most important task in formulating statements is ensuring they are free from ambiguity and irrelevant clues. There is a list of things to avoid when phrasing statements.
1. Avoid broad general statements if they are to be judged true or false.
- Explanation: Most broad generalizations are false unless qualified, and the use of qualifiers provides clues to the answer.
- 📌 Example (Poor item): The president of Pakistan is usually elected to his/her office.
- 📌 Example (Better item): According to the constitution of Pakistan, the president of Pakistan is elected by parliament.
💡 Why this matters: Broad statements create confusion about what criteria to use and often include clue words like "usually" that signal a false answer, compromising item validity.
Topic- 075: Suggestions for Constructing True-False Items – II
2. Avoid trivial statements.
- Explanation: To obtain statements that are clearly true or false, teachers often turn to specific statements of fact that fit the criterion but have little significance from a learning point of view. Such items test recall of unimportant details rather than meaningful learning.
- 📌 Example (Poor item): Mamnoon Hussain is the 12th president of Pakistan.
- 📌 Example (Poor item): India declared war on Pakistan on September 3rd, 1965.
💡 Why this matters: Trivial items reward rote memorization of isolated facts and do not assess understanding of important concepts, principles, or relationships.
⭐ Key Takeaways
True/False items are alternate-response questions best suited for measuring lower-level learning outcomes such as factual knowledge, definitions, principles, distinguishing fact from opinion, and recognizing cause-effect relationships. Their main advantages are efficiency in covering broad content and quick scoring, but they suffer from high guessing probability (50% chance correct by chance) and are limited to lower cognitive levels. Teachers must avoid broad generalizations that introduce qualifier clues and must avoid trivial statements that test unimportant details. Proper construction requires clear, unambiguous statements that focus on significant learning outcomes, and the item format can be varied (true/false, yes/no, fact/opinion, cause-effect, or logic with converse).
🧠 Quick Revision Questions
- What are the five specific learning outcomes that True/False items are well-suited to measure?
- Why is guessing a major limitation of True/False items, and what is the probability of guessing correctly?
- What is the problem with using broad general statements in True/False items, and what alternative is recommended?
- Why should teachers avoid trivial statements when constructing True/False items?
- List three different response formats (category pairs) that can be used with alternate-response test items along with examples of each.
📘 Lecture 22 — Creating Fixed-Choice Test Items: True/False and Its Uses-II
📖 Overview: This lecture provides detailed suggestions for constructing effective true-false test items. It covers seven key principles, from avoiding negative phrasing and complex sentences to ensuring statements are factual, balanced in length, and focused on single ideas. These guidelines help educators create clear, fair, and valid assessments that accurately measure student learning.
🗂️ Topics Covered
The lecture covers seven specific topics (076–080) offering practical suggestions for constructing true-false items. Topics include avoiding negative and double negative statements, avoiding long and complex sentences, avoiding two ideas in one statement unless measuring cause-effect relationships, avoiding opinion statements not attributed to a source, and avoiding unequal length between true and false statements. Each topic includes explanations, poor examples, and better examples for clarity.
📝 Lecture Summary
Topic- 076: Suggestions for Constructing True-False Items –III — Avoid the use of negative especially double negative statements.
Students often overlook negative words like "no" or "not," and double negatives create ambiguity. If a negative statement is absolutely necessary, the negative word should be underlined or written in italic to ensure students do not overlook it.
📌 Example (Poor item): "None of the steps in the experiment was unnecessary." (This is confusing because of the double negative.)
📌 Example (Better item): "All the steps in the experiment were necessary." (Clear and straightforward.)
Topic- 077: Suggestions for Constructing True-False Items –IV — Avoid long and complex sentences.
A test item should measure whether a student has achieved the intended knowledge or understanding, not their reading comprehension. Long, complex sentences introduce an extraneous factor of reading comprehension. If simplifying the sentence is impossible, consider changing to a different item format.
📌 Example (Poor item): "Despite the theoretical and experimental difficulties of determining the exact pH value of a solution, it is possible to determine whether a solution is acid by the red color formed on the litmus paper when it is inserted into the solution." (Too long and complex.)
📌 Example (Better item): "Litmus paper turns red in an acid solution." (Simple and direct.)
Topic- 078: Suggestions for Constructing True-False Items –V — Avoid including two ideas in one statement unless cause and effect relationship are being measured.
When two ideas are in one statement, students can get confused and answer based on only one idea, while the other may have a different truth value. It is better to use two simple statements. An exception is when measuring cause-and-effect relationships.
📌 Example (Poor item): "Pakistan could not qualify for world cup hockey because of poor resources, low talent and government support." (Should be false because "government support" is not a reason.)
📌 Example (Better item — simple statements):
- "Pakistan could not qualify for world cup hockey because of poor resources, low talent." (False)
- "Pakistan could not qualify for world cup hockey because of poor resources." (True)
Topic- 079: Suggestions for Constructing True-False Items –VI — Avoid using opinion that is not attributed to some source.
A statement of opinion cannot be marked true or false. It is unfair to expect students to guess how the teacher will score such items or to treat opinion statements as statements of fact.
📌 Example (Poor item): "All anti-state activists should be hanged." (This is an opinion, not a fact.)
📌 Example (Better item): "National action plan allows the justice system to give death penalty to anti-state activists if proven guilty." (This can be verified as true or false based on law.)
Topic- 080: Suggestions for Constructing True-False Items –VII — Avoid using true statements and false statements that are unequal in length.
There is a natural tendency for true statements to be longer because they must be precisely phrased to be absolutely true. This can be overcome by making false statements longer by using qualifying phrases similar to those found in true statements.
🔑 Definition — Key Principle: To avoid giving clues to test-takers, ensure true and false statements are approximately equal in length. Lengthen false statements by adding appropriate qualifying phrases.
⭐ Key Takeaways
- Avoid negative and especially double-negative statements; if negatives are necessary, use visual emphasis like underlining or italics.
- Keep sentences short and simple to avoid measuring reading comprehension instead of the intended knowledge.
- Avoid including two separate ideas in one statement unless testing a cause-effect relationship; use separate statements instead.
- Never use opinion statements unless attributed to a verifiable source; true-false items must measure factual knowledge only.
- Ensure true and false statements are approximately equal in length to prevent students from using length as a clue to the correct answer.
🧠 Quick Revision Questions
- Why should double negatives be avoided in true-false items, and what should be done if one is necessary?
- What extraneous factor is introduced by using long and complex sentences in true-false items?
- When is it acceptable to include two ideas in one true-false statement?
- Why is it unfair to include opinion statements without attributing them to a source?
- How can test-writers overcome the natural tendency for true statements to be longer than false statements?