Why Item Quality Matters
NBME 2024 · Ch.1–2A poorly written question does not test knowledge — it tests test-taking skill, reading speed, or the ability to interpret an item writer's ambiguous intent. The goal is a level playing field where performance reflects mastery alone.
A test-taker's probability of answering correctly should be determined entirely by their content expertise, not by structural flaws that add irrelevant difficulty or by cueing strategies that benefit the "testwise" examinee.
Item Format Families
RECOMMENDEDOne-Best-Answer (A-type)
Select the single best option from 4–5 choices. Options need not be wholly wrong — they are ranked on a single continuum from "least correct" to "most correct." This format best assesses clinical judgment, synthesis, and application.
Advantages
- Assesses judgment and reasoning, not just recall
- Distractors don't need to be completely false
- Same stem can support multiple lead-ins (diagnosis + management)
- Currently the only format used on USMLE/NBME exams
AVOID IF POSSIBLETrue-False Family
Test-takers must select all "true" options. Requires absolute true/false judgments — any ambiguity collapses the item. Reviewers alter answer keys far more frequently than for one-best-answer items.
Pitfalls
- Forces recall of isolated facts rather than application
- True/false distinctions are often ambiguous among experts
- Test-takers must guess the item writer's intent
- Frequently rewritten or discarded after review
The Five Core Rules
Ch.5Item Writing Workflow
Rule 1Focus on an Important Concept
Every item must trace back to a specific testing point on the exam blueprint. Ask: "What do I want test-takers to be able to do at the next stage of their training?" Avoid trivia. Prioritise common presentations and potentially catastrophic conditions.
Rule 2Assess Application, Not Isolated Recall
Recall items ask: "What is the treatment for disease X?" — answerable from a single textbook paragraph. Application items require identifying relevant findings, integrating them, and reaching a decision. Use a clinical or experimental vignette to force this higher-order thinking.
Which of the following medications is used to decrease preload in systolic heart failure?
A 58-year-old man with a 3-week history of exertional dyspnoea and bilateral ankle oedema… [full vignette] …Which of the following is the most appropriate initial management?
Rule 3Closed, Focused Lead-in ("Cover-the-Options" Rule)
The lead-in must be a single, clear, closed question. After reading the vignette and lead-in, a knowledgeable test-taker should be able to cover the options and answer correctly. This is the single most reliable indicator of a well-written item.
Rule 4Homogeneous and Plausible Options
All options — including distractors — must address the lead-in in the same manner, belong to the same category, and be ranked on a single dimension from least to most correct. Distractors must be plausible enough to attract test-takers who don't know the answer.
Rule 5Review for Technical Flaws
After drafting, step back and read critically. Remove any structural flaw that adds irrelevant difficulty or gives away the answer. Ask a colleague to review for content accuracy and structural integrity before finalising.
Constructing the Clinical Vignette
Ch.6The vignette is the engine of the item. A rich, authentic clinical scenario forces the test-taker to synthesise information — distinguishing those who know the material from those who are merely familiar with the facts.
TemplateStandard Vignette Structure
Present elements in this order. Not all are required — include only those relevant to the testing point.
- Age, gender (self-identified where relevant)
- Site of care (office, ED, inpatient ward, clinic)
- Chief concern / presenting symptom
- Duration of symptoms
- Relevant history — past medical, family, psychosocial, medications, allergies
- Physical examination findings
- Laboratory / diagnostic study results
- Initial treatment and subsequent findings (for management items)
Vignette Length and Window Dressing
APPROPRIATENecessary Complexity
Longer vignettes that require test-takers to sift relevant from irrelevant information are valid and intentional — they mirror real clinical reasoning. Including some window dressing (incidental findings) is acceptable and authentic.
Use: "The family history is noncontributory" to efficiently dispose of irrelevant details rather than omitting them entirely.
AVOIDRed Herrings and Pure Verbosity
Do not include information designed to deliberately mislead (red herrings) or information that serves no purpose other than increasing reading time. Every sentence should either support the correct answer or make a distractor more plausible.
Avoid basing items too closely on real patients — real cases have atypical features that confuse introductory learners.
Effect of Vignette Length on Performance
Research Finding — Vignette Depth Separates Competence Levels
NBME data on the same testing point across three formats shows clearly that longer vignettes expose the gap between strong and weak performers:
| Format | High Group (Top 20%) | Low Group (Bottom 20%) | Conclusion |
|---|---|---|---|
| No vignette | 99% | 90% | Minimal discrimination — recall only |
| Short vignette | 98% | 82% | Moderate discrimination |
| Long vignette | 98% | 66% | Strong discrimination — best for summative exams |
Implication: For summative, high-stakes assessments, longer clinical vignettes are preferable. For formative in-class quizzes, shorter formats are acceptable.
Patient Characteristics — Inclusive and Unbiased Vignettes
Ch.6 · NBME 2024Patient characteristics (PC) — including sex/gender identity, race/ethnicity, sexual orientation, disability, and socioeconomic status — should be included thoughtfully. They must never reinforce stereotypes or cue the correct answer. Always report gender and race/ethnicity as self-identified.
Include PCs When
- They are clinically relevant to the diagnosis or distractor quality
- Omitting them would make the item unreasonably difficult
- They add to exam-level representativeness (even if not clinically essential)
- Geographic/demographic data is necessary for reasoning (e.g., infectious disease endemic areas)
Omit or Change PCs When
- The PC cues too strongly to the diagnosis (e.g., Swedish patient + sarcoidosis)
- Including the PC perpetuates racial or ethnic stereotypes
- The condition occurs across multiple populations with similar frequency
- The PC is not needed for clinical reasoning and risks bias
Writing the Lead-in
The Lead-in Must Be a Single, Closed Question
A closed lead-in ends in a question mark and specifies exactly what is being asked. An open lead-in (e.g., "The diagnosis in this patient is:") forces test-takers to search options before they know what they're looking for.
The diagnosis in this patient is:
The most appropriate management would be:
Which of the following is the most likely diagnosis?
Which of the following is the most appropriate initial management?
High-Yield Lead-in Phrases by Competency
| Competency | Sample Lead-in Phrases |
|---|---|
| Diagnosis | Which of the following is the most likely diagnosis? / Which of the following is the most likely working diagnosis? |
| Mechanism / Foundational Science | Which of the following is the most likely cause/mechanism? / Which of the following is the most likely explanation for these findings? |
| Diagnostic Studies | Which of the following is the most appropriate diagnostic study at this time? / Which of the following laboratory studies is most likely to confirm the diagnosis? |
| Management | Which of the following is the most appropriate initial management? / Which of the following is the most appropriate pharmacotherapy? |
| Prognosis / Complications | Which of the following is the most likely complication? / Based on these findings, this patient is most likely to develop which of the following? |
| Prevention / Screening | Which of the following is the most appropriate screening test? / Which of the following is most likely to have prevented this patient's condition? |
| Communication / Ethics | Which of the following is the most appropriate provider response? / Which of the following is the most appropriate next step? |
Principles of Option Construction
Ch.5–6Options are where most item flaws originate. The correct answer requires equal craftsmanship as the distractors — in many ways the distractors are harder to write well.
- Homogeneous — all in the same category, rankable on a single dimension
- Plausible — distractors must genuinely attract those who don't know the answer
- Parallel — same grammatical structure and approximately equal length
- Clear — no vague frequency terms, no absolute terms, no ambiguous language
- Closed — all options should follow grammatically from the lead-in
Anatomy of a Well-Built Option Set
CORRECTHomogeneous Options on a Single Dimension
All five options are medications — they answer the lead-in in exactly the same way. A knowledgeable test-taker can cover the options and identify "an NSAID" before seeing the list.
FLAWEDHeterogeneous Options — Multiple Dimensions
The options span four different dimensions. A test-taker must decide whether "hereditary" is more or less true than "responds to allopurinol" — an impossible comparison.
Writing Effective Distractors
Start with the Correct Answer
Generate the keyed answer first, then build distractors that are plausible alternatives within the same category. For a diagnosis item (e.g., correct = community-acquired pneumonia), reasonable distractors include pulmonary embolism, lung cancer, and pneumothorax — not unrelated conditions.
Distractors Need Not Be Wholly Wrong
In one-best-answer format, distractors are "less correct" than the keyed answer — they sit lower on the same dimension. This makes items fairer and harder to write, but correctly separates those who understand gradations from those who only recognise extremes.
Language Precision in Options
| Avoid | Why | Instead |
|---|---|---|
| "Is associated with" | Imprecise — almost anything is associated with almost anything | Specific causal or mechanistic language |
| "Usually," "often," "frequently" | Interpreted differently by different people; creates ambiguity | Quantify: "in >50% of cases" or restructure the option |
| "Always," "never" | Absolute terms are almost always wrong; testwise examinees eliminate them | Remove absolute terms; restructure the lead-in |
| "May," "could be," "might" | Cueing language — almost anything "may" be true | Commit to specific, definitive statements |
| "None of the above" | Converts item to true-false; allows guessing outside the option set | Replace with a specific fifth option or "No intervention is indicated" |
| Open-ended options (e.g., "Treat the underlying cause") | Vague; multiple options could qualify | Name the specific treatment or intervention |
Sequential Item Sets (F-type)
Ch.6Rules for Writing F-type Sets
F-type sets present 2–3 items around an unfolding clinical scenario. Test-takers cannot return to earlier items once they advance.
Three Rules
- Include enough clinical richness upfront to support the case as it unfolds — don't introduce the key information only in later scenarios.
- Advance the case in time and provide updated patient information to restore context between items.
- Imply a response to the previous item but do not turn the next item into a pure recall question — the next clinical decision must still require reasoning.
Two Categories of Technical Flaws
Ch.3Technical flaws undermine the validity of every item they appear in. They shift what the item measures — from content knowledge to reading ability, test-taking skill, or luck. There are exactly two categories.
Category 1Flaws Adding Irrelevant Difficulty
These confuse all test-takers, adding construct-irrelevant variance to scores. They make items harder for reasons unrelated to the testing point — often affecting weak and strong students equally.
Category 2Flaws Cueing the Testwise Examinee
These help savvy test-takers guess correctly without knowing the content. They inflate scores for students who have learned structural test-taking strategies rather than the clinical content.
Flaw Reference — Click to Expand
Category 1 — Irrelevant Difficulty
Problem: Overly long options increase reading load, shifting what is measured from content knowledge to reading speed. This flaw applies to options only — a long vignette is acceptable and often desirable.
Fix: Move common text into the stem. Use parallel construction. Shorten options to the minimum necessary phrase. Options should ideally be 1–5 words.
A. Reasonable and timely notice, an impartial panel empowered to make a decision, a chance to hear evidence and to confront witnesses, and the ability to present evidence in defence
A. Timely notice and an impartial hearing panel
Problem: Words like "often," "usually," "frequently," and "sometimes" are interpreted differently by different readers — even content experts. This creates multiple defensible answers and makes the keyed answer arbitrary.
Fix: Replace vague qualifiers with specific figures (e.g., ">60% of cases") or restructure the option so no qualifier is needed.
A. Often is related to endocrine disorders · D. Usually responds dramatically to dietary regimens · E. Usually responds to pharmacotherapy
Problem: Mixing numeric formats (ranges vs. specific values) or listing ranges that overlap or contain other options creates logical impossibilities and allows elimination of options without content knowledge.
Fix: Use a single consistent format (all ranges or all specific values). List numbers in ascending order. Ensure no range contains another option as a specific value.
A. Less than 20% · B. 20 to 30% · C. Greater than 50% · D. 75% · E. 90%
Options D and E fall within C — a testwise student eliminates D and E immediately.
Problem: Stems that ask test-takers to rank multiple items using Roman numerals, perform multiple sequential logical operations, or interpret confusingly encoded information add cognitive load unrelated to clinical knowledge.
Fix: Reduce to a single decision point. Put the ranked/enumerated options in the option set itself. Focus on the single most important clinical judgement.
Problem: "Each of the following is true EXCEPT:" reverses the cognitive task. Test-takers who are accustomed to positive phrasing frequently miss the negative qualifier, even when it is bolded or capitalised.
Fix: Rephrase positively. If testing knowledge of what is not indicated, use an active positive lead-in: "Which of the following is contraindicated in this patient?"
Each of the following statements about cholesterol is true EXCEPT:
Which of the following statements about cholesterol is incorrect?
Problem: Options that differ in grammatical structure, level of specificity, or format force test-takers to perform additional cognitive work to compare them — work that is irrelevant to the testing point.
Fix: Edit all options to the same grammatical form, similar length, and consistent level of detail. All options should complete the lead-in in the same way.
Category 2 — Cueing the Testwise Examinee
Problem: An option that does not follow grammatically from the lead-in is immediately eliminated by a testwise student. This occurs when item writers focus on crafting the correct answer and write distractors hastily.
Fix: Read every option aloud as a continuation of the lead-in. Use closed lead-ins — they prevent grammatical mismatches by forcing all options to complete the same sentence.
Lead-in: "Her diagnosis is most likely to be an:"
A. asthma attack ✓ · B. costochondritis ✗ · C. pleurisy ✗
Only option A follows grammatically ("an asthma attack").
Problem: When a word from the vignette appears in only one option, that option is cued as the correct answer — even etymologically (e.g., "bone pain" in the stem + "osteitis" in one option).
Fix: Scan every option for words that echo the vignette. Either remove the repeated word from the option or include it in all options.
Stem: "…experiencing the world as unreal."
A. Depersonalisation · B. Derailment · C. Derealisation ← cued · D. Focal memory deficit
Problem: The correct answer is noticeably longer than distractors, contains two components, or includes parenthetical teaching notes. Item writers often over-elaborate the correct answer because they want to explain it.
Fix: Review all options for length consistency. Strip instructional caveats and parenthetical information from the correct option. The rationale belongs in an answer key, not the question.
A. A complication of a variety of illnesses and tends to prolong many (>3) of them ← correct
B. A frequent problem in obsessive-compulsive disorder
C. Never seen in organic brain damage
D. Synonymous with malingering
Problem: A subset of options covers all logically possible outcomes. A testwise student identifies this subset, eliminates the non-subset options, and narrows to a 1-in-3 guess without any content knowledge.
Fix: Replace at least one option in the exhaustive subset with a different type of option. When revising, confirm no new exhaustive subset is created.
Administration of furosemide results in:
A. a decrease in urine potassium · B. an increase in urine potassium ✓ · C. improved glucose control · D. no change in urine potassium · E. requires decreasing dose with renal failure
Urine K can only increase, decrease, or stay the same → one of A, B, or D must be correct. C and E are distractors without merit.
Problem: Options containing "always" or "never" are almost always incorrect in medicine. Testwise students immediately eliminate them, reducing the effective option set.
Fix: Remove absolute terms. Move the verb into the closed lead-in and use short, parallel options without modifiers.
C. is never seen in patients with neurofibrillary tangles at autopsy
D. is never severe
Both eliminated immediately by any testwise examinee.
Problem: The correct answer is the option with the most elements in common with other options. Because item writers often write distractors as permutations of the correct answer, the keyed answer ends up being the "most average" option — identifiable by counting term frequencies.
Fix: Review all options and count how many times each unique term appears. Balance term distribution so no option is distinguishable purely by frequency of shared language.
A. anionic, inside · B. cationic, inside ✓ · C. cationic, outside · D. uncharged, inside · E. uncharged, outside
3/5 options say "inside" → correct answer is likely in that group. 2/5 say "cationic" → converges on B.
Quick Reference — Flaws and Fixes
| Flaw | Category | Solution |
|---|---|---|
| Long/complex options | Irrelevant difficulty | Move common text to stem; shorten to parallel phrases |
| Vague terms (often, usually) | Irrelevant difficulty | Quantify or restructure; avoid frequency qualifiers |
| Inconsistent numeric data | Irrelevant difficulty | Consistent format; ascending order; no overlapping ranges |
| Complicated stem (Roman numerals) | Irrelevant difficulty | Reduce to single decision; put options in option set |
| Negatively phrased lead-in (EXCEPT) | Irrelevant difficulty | Rephrase positively (e.g., "contraindicated") |
| Nonparallel options | Irrelevant difficulty | Edit all options to same grammatical form and length |
| "None of the above" | Irrelevant difficulty | Replace with specific fifth option |
| Grammatical cues | Cueing | Use closed lead-ins; read each option after the lead-in |
| Clang clues (word repetition) | Cueing | Remove repeated word or use it in all options |
| Correct answer stands out | Cueing | Equalise option length; remove teaching rationales |
| Collectively exhaustive subset | Cueing | Replace one option in the exhaustive subset |
| Absolute terms (always/never) | Cueing | Eliminate; move verb to closed lead-in |
| Convergence | Cueing | Balance frequency of repeated terms across all options |
Using Item Analysis Data
Ch.4After administration, every item can be evaluated statistically. Item analysis provides objective evidence about whether an item worked as intended — and flags problems that expert review may have missed.
P-value (Item Difficulty)
The proportion of all test-takers who answered correctly. A higher p-value = easier item.
Interpretation Guide
- p > 0.95: Too easy — little discriminating power; consider revision
- p = 0.50–0.80: Ideal range for most assessments
- p < 0.30: Very difficult — review for miskeying or structural flaws
- Very low p + negative discrimination: Almost certainly miskeyed
Always compare the observed p-value to your expectation. A mismatch signals something unexpected is happening in the item.
Discrimination Index (Item-Total Correlation)
Correlation between performance on the item and performance on the overall test. Ranges from −1.0 to +1.0.
Interpretation Guide
- Positive (+0.2 or higher): Item discriminates well — high performers answered correctly more often
- Near zero: Item measures something unrelated to overall test construct
- Negative: Low performers answered correctly more than high performers — likely miskeyed, or flaw benefiting guessing
Option Analysis — Reading the Data
What to Look for in Each Option
- Was any option never selected? → Not plausible enough; rewrite as a more attractive distractor
- Was a distractor selected more often than the key? → Likely miskeyed, or the item has two defensible answers
- Did High performers choose a distractor more than Low performers? → That distractor may be more correct than the key
- Are Low performers spread evenly across distractors? → Healthy pattern — they are guessing genuinely
Worked Examples — Five Common Scenarios
Pattern 1Miskeyed Item
| Group | A | B* (keyed) | C | D | E |
|---|---|---|---|---|---|
| High (top 25%) | 1 | 1 | 91 | 4 | 1 |
| Low (bottom 25%) | 20 | 6 | 51 | 14 | 6 |
| Total | 9 | 2 | 76 | 8 | 3 |
p-value: 2 · Discrimination: −0.21
Diagnosis: Classic miskeying pattern. Only 2% answered the keyed answer (B); 91% of high performers chose C. The correct answer is almost certainly C. Rekeying to C gives p = 76, discrimination = +0.40. No further revision needed before using.
Pattern 2Well-Functioning Item
| Group | A | B | C* (keyed) | D | E |
|---|---|---|---|---|---|
| High (top 25%) | 0 | 1 | 90 | 3 | 3 |
| Low (bottom 25%) | 0 | 1 | 60 | 25 | 8 |
| Total | 0 | 1 | 74 | 12 | 7 |
p-value: 74 · Discrimination: +0.33
Diagnosis: Excellent item. Appropriate difficulty, strong discrimination. Options A and B attract almost no one — consider rewriting them to be more plausible distractors, though be aware this may shift difficulty unpredictably.
Pattern 3Difficult Item — Possible Second Answer
| Group | A | B | C* (keyed) | D | E |
|---|---|---|---|---|---|
| High (top 25%) | 44 | 1 | 50 | 2 | 1 |
| Low (bottom 25%) | 20 | 15 | 21 | 22 | 20 |
| Total | 32 | 7 | 34 | 14 | 11 |
p-value: 34 · Discrimination: +0.30
Diagnosis: 44% of High performers chose option A rather than the key (C). This is a red flag — A may be defensible as a correct answer. A content expert should review before scoring. If the key is confirmed correct, A needs revision to be clearly wrong.
Pattern 4Difficult but Structurally Sound
| Group | A | B | C* (keyed) | D | E |
|---|---|---|---|---|---|
| High (top 25%) | 18 | 10 | 51 | 17 | 2 |
| Low (bottom 25%) | 24 | 24 | 21 | 25 | 4 |
| Total | 22 | 17 | 34 | 22 | 3 |
p-value: 34 · Discrimination: +0.30
Diagnosis: Same difficulty as Pattern 3 but healthier structure. Low performers are spread fairly evenly across distractors — they are genuinely guessing. High performers correctly choose C at 51% vs. only 21% of Low performers. Good discrimination; content review still warranted for distractors A, B, and D.
Pattern 5Two Correct Answers / Structural Flaw
| Group | A | B | C | D* (keyed) | E |
|---|---|---|---|---|---|
| High (top 25%) | 10 | 43 | 5 | 40 | 2 |
| Low (bottom 25%) | 23 | 36 | 12 | 26 | 3 |
| Total | 17 | 43 | 7 | 31 | 2 |
p-value: 31 · Discrimination: −0.09
Diagnosis: Both High and Low performers prefer option B over the keyed answer D. Classic two-correct-answers pattern. Do not score this item until a content expert resolves the conflict. The item should be either rekeyed to B, revised to make D clearly superior, or withdrawn.
- Low p + negative discrimination: Check keying first. Likely miskeyed → rekey and re-check.
- High p + positive discrimination: Item is functioning well → keep, consider rewriting weak distractors.
- High performers favour a distractor: Possible second correct answer → expert review before scoring.
- Near-zero discrimination: Item measuring a different construct, or structural flaw → review and revise.
- Option never selected: Implausible distractor → rewrite to be more attractive.
Pre-Submission Item Quality Checklist
Appendix A · NBME 2024Use this checklist before finalising any item. For each question, check every element systematically. Having a colleague independently apply the same checklist identifies blind spots the item writer cannot see.
Vignette Quality
- The vignette contains all information needed to answer the question — nothing essential is deferred to the options
- The structure follows the standard order: demographics → symptoms → history → examination → investigations → treatment
- Every sentence either supports the correct answer or makes a distractor more plausible
- No "red herrings" designed purely to mislead — window dressing is acceptable; deliberate traps are not
- The vignette is clinically authentic — the presentation would be encountered in practice
- Patient characteristics (age, gender, ethnicity) are included appropriately and do not cue the diagnosis unless clinically essential
- Gender and race/ethnicity are reported as self-identified
- The vignette represents a level of complexity appropriate to the test-taker population
Lead-in Quality
- The lead-in is a single, closed question ending with a question mark
- "Cover-the-options" test passes: a knowledgeable test-taker can answer without seeing the option set
- The lead-in is positively phrased — no "EXCEPT," "NOT," or other negative qualifiers
- The lead-in does not contain open-ended phrasing (e.g., "The diagnosis in this patient is:")
- The lead-in contains an active verb specifying the clinical task (diagnose, manage, investigate, etc.)
Option Set Quality
- All options are homogeneous — they all belong to the same category and address the lead-in in the same way
- All options are plausible — each one could attract a test-taker who doesn't know the answer
- All options are parallel in grammatical structure and approximately equal in length
- The correct answer is not longer than distractors and contains no teaching parentheticals
- Numeric options are in a single consistent format (all ranges or all specific values) listed in ascending order
- No option uses vague frequency terms (often, usually, frequently, sometimes)
- No option uses absolute terms (always, never)
- "None of the above" is not used
Technical Flaw Screening
- No grammatical cues: Every option follows grammatically from the lead-in when read aloud
- No clang clues: No word from the vignette appears exclusively in the correct answer option
- No collectively exhaustive subset: No group of options covers all logically possible outcomes
- No convergence: No individual option is identifiable as correct simply by counting shared terms across options
- No complicated stem: The item requires exactly one decision, not a ranking or sequential logical chain
Final Steps Before Submission
- You have identified and confirmed the single defensible correct answer with a specific reference or rationale
- You have written a brief answer rationale explaining why the keyed answer is correct and why each distractor is incorrect
- A colleague with content expertise has reviewed the item for accuracy and appropriateness
- A colleague without content expertise has read the item for clarity and structural integrity
- The item addresses a testing point on the exam blueprint — it is not testing trivia or an exotic "zebra" diagnosis
- The item is appropriate for the competency level of your test-taker population
- If media is included: it is essential to answering the question, it is high quality, and it has no identifying patient information or distracting visual cues
Recall vs. Application Decision Guide
| Context | Recommended Format | Vignette Length | Acceptable Item Types |
|---|---|---|---|
| Formative mid-lecture quiz | Recall acceptable | None to short | A-type (single item) |
| End-of-module formative | Application preferred | Short to medium | A-type; F-type (2 items) |
| End-of-year summative | Application required | Medium to long | A-type; F-type (2–3 items) |
| High-stakes licensing/promotion | Application required | Full clinical vignette | A-type; F-type only |
| Specialty/postgraduate certification | Application; may include atypical features | Full vignette; may include chart format | A-type; F-type; G-type |
A high-quality MCQ presents an authentic clinical problem, asks a single focused question that a knowledgeable test-taker can answer without seeing the options, offers five homogeneous and plausible choices of equal length, and contains no structural feature that benefits a test-taker who does not know the content.