Educational evaluation - test, measurement, assessment, evaluation - principles of evaluation - test characteristics: objectivity, reliability, validity
In the field of education, understanding how to gauge student learning is paramount. This involves a set of interconnected concepts: test, measurement, assessment, and evaluation. While often used interchangeably, they represent distinct stages and purposes in the process of understanding student progress.
1. Test
A test is a specific tool or instrument used to collect data about a student's knowledge, skills, or abilities. It's a sample of behavior or performance under controlled conditions. Tests can take many forms, such as multiple-choice questions, essays, practical demonstrations, or even interviews. The primary purpose of a test is to elicit a response from the learner that can be observed and quantified.
Types of Tests:
- Objective Tests: These tests have specific correct answers, and scoring is not influenced by the examiner's judgment. Examples include multiple-choice, true/false, and matching items.
- Subjective Tests: These tests require the student to construct their answers, and scoring involves some degree of interpretation by the examiner. Essays and short-answer questions are common examples.
- Diagnostic Tests: Used to identify specific strengths and weaknesses of students.
- Aptitude Tests: Measure a person's potential to learn a new skill.
- Achievement Tests: Measure what a student has learned in a specific subject or area.
2. Measurement
Measurement is the process of assigning numerical values or scores to the performance or characteristics of an individual based on a test. It's about quantifying the results obtained from a test. For instance, if a student scores 80 out of 100 on a math test, 80 is the measurement. Measurement focuses on the "how much" or "how many" aspect. It provides a score but doesn't necessarily interpret the meaning of that score.
Scales of Measurement:
- Nominal Scale: Used for labeling or categorizing data (e.g., gender, student ID). No order or magnitude is implied.
- Ordinal Scale: Used for ranking data, indicating relative position (e.g., 1st, 2nd, 3rd place). The difference between ranks is not necessarily equal.
- Interval Scale: Data can be ordered, and the differences between values are equal and meaningful (e.g., temperature in Celsius or Fahrenheit). However, there is no true zero point.
- Ratio Scale: Has all the properties of an interval scale plus a true zero point, allowing for meaningful ratios (e.g., height, weight, test scores where zero means absence of the trait).
3. Assessment
Assessment is a broader term that encompasses the entire process of gathering and interpreting information about student learning. It involves using various tools and techniques (including tests and measurements) to understand what students know, understand, and can do. Assessment is about collecting evidence of learning. It's not just about assigning a score but about understanding the learning process and outcomes.
Purposes of Assessment:
- Formative Assessment: Conducted during the learning process to provide ongoing feedback to both students and teachers. It helps identify areas where students are struggling and informs instructional adjustments. Examples include quizzes, class discussions, and observation.
- Summative Assessment: Conducted at the end of a learning period (e.g., end of a unit, semester, or year) to evaluate overall achievement. Examples include final exams, standardized tests, and major projects.
- Diagnostic Assessment: Identifies students' prior knowledge, skills, and potential learning difficulties before instruction begins.
4. Evaluation
Evaluation is the final step, involving the interpretation and judgment of the data collected through assessment and measurement. It's about making decisions based on the evidence. Evaluation answers the question: "How well did the student learn?" or "How effective was the teaching?" It involves comparing the results against certain standards, criteria, or objectives.
Key Aspects of Evaluation:
- Judgment: Assigning value or worth to student performance.
- Decision-Making: Using the judgment to make decisions about student progress, curriculum effectiveness, or teaching strategies.
- Interpretation: Understanding the meaning of the measured data in the context of learning goals.
For example, if a student scores 85% on a history test (measurement), and the passing score is 70% (standard), the evaluation would be that the student has passed the test. If the teacher observes that many students scored below 70%, the evaluation might be that the teaching method was not effective for this particular topic.
- Test: The tool.
- Measurement: The number.
- Assessment: Gathering evidence.
- Evaluation: Judging the worth.
Principles of Educational Evaluation
Effective educational evaluation is guided by several fundamental principles. Adhering to these principles ensures that the process is fair, meaningful, and contributes positively to the learning environment.
- Purposeful: Evaluation should always have a clear objective. What do we want to find out? What decisions will be made based on the results?
- Comprehensive: It should cover all important aspects of learning, not just rote memorization. This includes cognitive, affective, and psychomotor domains.
- Continuous: Evaluation is not a one-time event but an ongoing process that happens throughout the learning journey.
- Scientific: Evaluation tools and procedures should be objective, reliable, and valid.
- Diagnostic: It should help identify the strengths and weaknesses of learners.
- Guidance-Oriented: The results should be used to guide students and teachers in improving the learning process.
- Democratic: Students should be involved in understanding the evaluation criteria and process.
- Practical: The methods used should be feasible within the given time, resources, and context.
- Ethical: Confidentiality, fairness, and respect for individuals must be maintained.
Test Characteristics
For any test to be useful in educational settings, it must possess certain desirable characteristics. These characteristics ensure that the test accurately measures what it intends to measure and that the results are trustworthy. The three primary characteristics are objectivity, reliability, and validity.
1. Objectivity
Objectivity refers to the extent to which a test is free from the examiner's personal bias or subjective judgment in scoring. An objective test can be scored consistently by different scorers, or even by the same scorer at different times, yielding the same result. This is particularly important for tests with predetermined correct answers, like multiple-choice or true/false questions.
Characteristics of an Objective Test:
- Unambiguous Questions: The questions should be clear and have only one correct answer.
- Uniform Scoring: Scoring is based on a key or rubric, eliminating personal opinion.
- Ease of Administration: Often designed for mass administration.
Example: In a multiple-choice question asking "What is the capital of France?", the options are Paris, London, Berlin, Rome. There is only one correct answer, Paris. Any scorer, regardless of their personal feelings, will mark Paris as correct. This makes the scoring objective.
Conversely, an essay question asking students to "Discuss the impact of the Industrial Revolution" can be subjective. Different evaluators might give different scores based on their interpretation of the essay's quality, depth, or writing style.
2. Reliability
Reliability refers to the consistency and stability of a test's results. A reliable test will produce similar scores for the same individual if administered again under similar conditions, assuming no significant learning or forgetting has occurred. It's about the precision of the measurement. A test can be objective but not reliable, or reliable but not objective. However, a test cannot be truly valid if it is not reliable.
How Reliability is Assessed:
- Test-Retest Reliability: Administering the same test to the same group of students on two different occasions and correlating the scores. A high correlation indicates good stability over time.
- Internal Consistency Reliability: Measures how consistent the items within a single test are with each other.
- Split-Half Method: Dividing the test into two halves (e.g., odd vs. even items) and correlating the scores on the two halves.
- Cronbach's Alpha: A statistical measure that estimates the average correlation among all possible split-halves.
- Parallel-Forms Reliability: Creating two equivalent forms of a test and administering both to the same group. Correlating the scores from the two forms.
- Inter-Rater Reliability: Used for subjective tests, it measures the degree of agreement between two or more raters who score the same test.
Example: Imagine a weighing scale. If you step on it multiple times in a row and it shows the exact same weight each time, the scale is reliable. If it shows different weights each time, it's unreliable, even if the displayed weights seem plausible. Similarly, a reliable educational test consistently measures a student's ability.
3. Validity
Validity refers to the extent to which a test measures what it is intended to measure. It's about the accuracy and appropriateness of the inferences made from test scores. A test can be reliable (consistent) but not valid (not measuring the right thing). For example, a test might consistently measure a student's reading speed (reliable) but fail to measure their comprehension (validity issue if comprehension is the goal).
Types of Validity:
- Content Validity: The test items adequately represent the entire domain or content area they are supposed to cover. This is crucial for achievement tests. For example, a final exam for a biology course should cover all the major topics taught during the semester, not just one chapter.
- Criterion-Related Validity: Measures how well a test score predicts future performance or correlates with another established measure (the criterion).
- Predictive Validity: The extent to which the test score predicts future performance. Example: Entrance exam scores predicting success in a college program.
- Concurrent Validity: The extent to which the test score correlates with a criterion measure obtained at the same time. Example: A new, shorter math test correlating highly with a well-established, longer math test.
- Construct Validity: The extent to which the test measures the underlying theoretical construct or trait it is designed to measure (e.g., intelligence, anxiety, creativity). This is the most complex type of validity and often involves gathering various types of evidence.
- Face Validity: The extent to which a test *appears* to measure what it is supposed to measure, based on surface-level examination. While not a technical measure of validity, it's important for user acceptance and motivation.
Example: A test designed to measure a student's ability to solve algebraic equations is valid if it actually assesses algebraic problem-solving skills. If it primarily measures reading comprehension or arithmetic fluency instead, it lacks construct and content validity for algebra.
- A test can be reliable but not valid. (e.g., a faulty scale consistently showing 5kg too much is reliable but not valid).
- A test cannot be valid if it is not reliable. (If scores fluctuate randomly, they cannot accurately measure anything).
- The highest level of test quality is achieved when a test is both reliable and valid.