Item analysis - Index of difficulty and discrimination, Reliability - Question Bank

1. A test designer wants to create a test that accurately reflects a specific curriculum. Which type of validity is most crucial?
A) Concurrent validity
B) Predictive validity
C) Construct validity
D) Content validity
2. Which of the following is a major drawback of the split-half reliability method if not corrected by the Spearman-Brown formula?
A) It underestimates the test's true reliability.
B) It overestimates the test's true reliability.
C) It is only suitable for speed tests.
D) It requires multiple test administrations.
3. An item with an Index of Difficulty (ID) of 0.75 and an Index of Discrimination (ID) of 0.30 suggests:
A) The item is too easy and a weak discriminator.
B) The item is easy and discriminates moderately.
C) The item is difficult and discriminates poorly.
D) The item is moderately difficult and discriminates fairly.
4. The purpose of ensuring high reliability in a test is to:
A) Maximize the validity of the test.
B) Minimize measurement error.
C) Ensure the test is easy for most students.
D) Guarantee that the test covers all relevant content.
5. Which coefficient indicates the highest level of reliability?
A) 0.25
B) 0.50
C) 0.75
D) 0.95
6. When an item analysis shows that an item has an Index of Difficulty (ID) of 0.15 and an Index of Discrimination (ID) of 0.40, what might be the interpretation?
A) The item is too easy and a poor discriminator.
B) The item is difficult but discriminates moderately well.
C) The item is too difficult and does not discriminate.
D) The item is excellent.
7. Content validity is assessed by:
A) Correlating test scores with an external criterion.
B) Examining the internal consistency of the test.
C) Determining the extent to which test items represent the domain of content they are supposed to cover.
D) Administering the test to different groups.
8. Which of the following is a primary assumption for using the test-retest method of reliability?
A) The trait being measured is unstable.
B) The test items are constantly changing.
C) The trait being measured is relatively stable over the time interval.
D) The scoring is subjective.
9. If the top 25% of test-takers answered an item correctly at a rate of 80%, and the bottom 25% answered it correctly at a rate of 20%, what is the Index of Discrimination (ID) using this common method?
A) 0.20
B) 0.50
C) 0.60
D) 0.80
10. An item that is answered correctly by 90% of the students and incorrectly by 10% of the students has an Index of Difficulty (ID) of:
A) 0.10
B) 0.50
C) 0.90
D) 1.00
11. When considering the Index of Difficulty (ID), an item with an ID of 0.30 is considered:
A) Easy
B) Moderately difficult
C) Difficult
D) Ambiguous
12. A high reliability coefficient suggests that:
A) The test measures what it is supposed to measure.
B) The test scores are free from random error.
C) The test is easy to administer.
D) The test content is relevant.
13. The standard error of measurement (SEM) is related to reliability and indicates:
A) The average difficulty of the test items.
B) The consistency of item discrimination.
C) The amount of error expected in an individual's score.
D) The proportion of correct answers.
14. Which of the following is a method for estimating reliability by assessing the consistency of responses to items within a single test administration?
A) Test-retest method
B) Alternate-forms method
C) Split-half method
D) CRITERION-REFERENCED method
15. What does a low Index of Discrimination (ID) generally suggest about a test item?
A) The item is too easy.
B) The item is too difficult.
C) The item is not effectively differentiating between students with high and low overall knowledge.
D) The item is measuring a specific, advanced concept.
16. If an item has an Index of Discrimination (ID) of 0.50 and an Index of Difficulty (ID) of 0.50, what is the assessment of this item?
A) Excellent item, highly discriminating and moderately difficult.
B) Poor item, too easy and not discriminating.
C) Good item, discriminating well and at an ideal difficulty level.
D) Fair item, but could be improved.
17. An item analysis is most useful during which stage of test development?
A) Initial conceptualization
B) Pilot testing and revision
C) Final scoring
D) Reporting results
18. Parallel-forms reliability (also known as equivalent-forms reliability) involves:
A) Administering the same test twice.
B) Administering two different but equivalent forms of a test.
C) Having multiple raters score the same test.
D) Splitting a test into two halves.
19. Which statement best describes the relationship between reliability and validity?
A) A test must be valid to be reliable.
B) A test can be reliable without being valid.
C) A test must be reliable to be valid.
D) Reliability and validity are unrelated concepts.
20. What is the implication of a very low Index of Difficulty (ID) for a test item?
A) The item is too easy and may not differentiate well among students.
B) The item is too difficult and may not differentiate well among students.
C) The item is perfectly discriminating.
D) The item is unreliable.
21. Which of the following is a measure of validity, not reliability?
A) Cronbach's Alpha
B) Test-retest coefficient
C) Content validity ratio
D) KR-20
22. The Kuder-Richardson Formula 20 (KR-20) is used to calculate internal consistency reliability for tests with:
A) Essay items
B) True-false or multiple-choice items (dichotomous scoring)
C) Rating scales
D) Performance tasks
23. An item analysis reports that an item has an ID of 0.95 and an Index of Discrimination of 0.10. What action should likely be taken?
A) Keep the item as it is.
B) Revise the item for clarity and difficulty.
C) Discard the item.
D) Use it only for high-achieving students.
24. If the Index of Difficulty (ID) for an item is 0.50, it means:
A) The item is very difficult.
B) The item is very easy.
C) Approximately half of the students answered the item correctly.
D) The item is ambiguous.
25. Which type of reliability is most appropriate for a multiple-choice test scored by a computer?
A) Test-retest reliability
B) Inter-rater reliability
C) Parallel-forms reliability
D) Internal consistency reliability
26. A test that yields consistent scores for the same individual under similar conditions is considered:
A) Valid
B) Reliable
C) Objective
D) Standardized
27. The Spearman-Brown Prophecy Formula is used to estimate:
A) The reliability of a test if its length is changed.
B) The difficulty of test items.
C) The discrimination power of items.
D) The content validity of a test.
28. Which factor can decrease the reliability of a test?
A) Clear instructions
B) Homogeneous test content
C) Ambiguous item wording
D) Objective scoring
29. A reliability coefficient of 0.90 indicates:
A) Very low reliability.
B) Moderate reliability.
C) High reliability.
D) Perfect reliability.
30. Inter-rater reliability is important when:
A) The test is administered over a long period.
B) The scoring of the test involves subjective judgment.
C) The test is designed for a single administration.
D) The test measures a single skill.
31. Split-half reliability is a method of internal consistency where:
A) The test is administered twice to the same group.
B) Two different forms of the test are administered.
C) The test is divided into two halves, and scores on the halves are correlated.
D) Scores from different raters are compared.
32. Internal consistency reliability assesses the extent to which items within a test measure the same construct. Which statistic is commonly used for this?
A) Cronbach's Alpha
B) Pearson Correlation Coefficient
C) Spearman-Brown Prophecy Formula
D) Kuder-Richardson Formula 20 (KR-20)
33. If a student gets a similar score when taking the same test on two different occasions, the test is considered to have high:
A) Content validity
B) Test-retest reliability
C) Criterion validity
D) Face validity
34. Test-retest reliability measures the consistency of scores over:
A) Different parts of the same test.
B) Different scorers.
C) Time.
D) Different versions of the test.
35. Which of the following is NOT a type of reliability estimation?
A) Test-retest reliability
B) Internal consistency reliability
C) Content validity
D) Inter-rater reliability
36. Reliability, in the context of educational measurement, refers to:
A) The extent to which a test measures what it is intended to measure.
B) The consistency or stability of test scores.
C) The fairness and equity of the test items.
D) The relevance of the test content to the curriculum.
37. When an item analysis reveals an item with a very low Index of Difficulty (close to 0) and a negative Index of Discrimination, it should typically be:
A) Retained as it is.
B) Revised or discarded.
C) Used for advanced learners only.
D) Published without modification.
38. An item with an Index of Discrimination of 0.00 means:
A) It is a very good item.
B) It discriminates perfectly between high and low scorers.
C) It does not differentiate between high and low scorers.
D) It is too difficult for everyone.
39. Which method is commonly used to calculate the Index of Discrimination?
A) Comparing the proportion of correct answers between the top and bottom score groups.
B) Calculating the average score of all test-takers.
C) Measuring the time taken to answer the item.
D) Assessing the item's grammatical correctness.
40. An Index of Discrimination (ID) of -0.10 suggests that the item:
A) Is a good discriminator.
B) Is a poor discriminator.
C) Functions negatively, with lower scorers more likely to answer correctly.
D) Is too easy.
41. An Index of Discrimination (ID) of 0.45 is generally considered:
A) Poor
B) Fair
C) Good
D) Excellent
42. A positive Index of Discrimination (ID) indicates that:
A) Students who answered the item correctly tend to score lower on the total test.
B) Students who answered the item correctly tend to score higher on the total test.
C) The item is equally difficult for high and low scorers.
D) The item is not related to the test's content.
43. The Index of Discrimination (ID) measures:
A) How many students answered the item correctly.
B) How well an item differentiates between high-scoring and low-scoring students.
C) The average score on the test.
D) The time required to complete the test.
44. If an item has an Index of Difficulty (ID) of 1.00, it means:
A) No student answered it correctly.
B) All students answered it correctly.
C) The item is highly discriminating.
D) The item is unreliable.
45. Which of the following is generally considered the ideal range for the Index of Difficulty (ID) for most achievement tests?
A) 0.00 to 0.20
B) 0.20 to 0.80
C) 0.80 to 1.00
D) 0.00 to 1.00 for all items
46. An Index of Difficulty (ID) of 0.20 suggests that the item is:
A) Easy
B) Moderately difficult
C) Difficult
D) Perfectly reliable
47. A high Index of Difficulty (ID) value, close to 1.00, indicates that the item is:
A) Very difficult for most students.
B) Very easy for most students.
C) Ambiguous and poorly worded.
D) Unrelated to the test's objective.
48. The Index of Difficulty (ID) for a test item represents:
A) The average score of students who answered the item correctly.
B) The proportion of students who answered the item correctly.
C) The correlation between the item score and the total test score.
D) The time taken by students to answer the item.
49. What is the primary purpose of item analysis in test construction?
A) To determine the chronological age of test-takers.
B) To evaluate the effectiveness and quality of individual test items.
C) To standardize the administration procedures of a test.
D) To establish the norms for a particular population.