Nick Item-Writing Academy
About 50 minutes · Academy module: Item-Analysis Literacy: p-Values, Discrimination, and Distractor Analysis.
Every item report starts with N (how many examinees answered) and the p-value (the proportion who answered correctly). The p-value is difficulty, nothing more: 0.91 is very easy, 0.45 is moderately hard. Small N means wide uncertainty — with 20 responses, a p-value of 0.70 could easily be 0.55 or 0.85 in a larger group. Never make keep/revise decisions on tiny samples; flag and re-administer.
Discrimination asks whether the right examinees got the item right: typically the difference in correct-response rates between the top and bottom scoring groups. Ebel's long-standing guidelines: 0.40 and above — very good; 0.30–0.39 — reasonably good; 0.20–0.29 — marginal, needs improvement; below 0.20 — poor, review or discard. A negative discrimination — high performers missing it more than low performers — is the loudest alarm in item analysis: it smells like a miskey or a subtle flaw.
But data never diagnose; they only flag. A negative discrimination with a very high p-value says the strongest examinees are missing an easy item — the classic miskey signature. A low discrimination on a very easy or very hard item may just be range restriction. The rule: statistics flag, faculty decide.
Option-level data shows the percentage of examinees choosing each option. Read it like this: an option nobody chose is a passenger — replace it with a near-miss. An option that pulled disproportionately from the top group is a second-key suspect — read the item for ambiguity. A distractor that pulled evenly across all groups is doing its job: attracting the partially-prepared without fooling the well-prepared.
Our read: distractor analysis is where item writers improve fastest, because it shows exactly which of your wrong answers were too wrong and which were accidentally right. One revision cycle on pulled data teaches more than a semester of theory.
No statistic, however alarming, justifies deleting or rekeying an item automatically. Negative discrimination means faculty review: read the item, check the key against the content, look for the ambiguity or the miskey, then decide — keep, revise and re-administer, or retire. Auto-delete throws away the diagnostic information (what did the item's failure teach you about your teaching?) and risks discarding a good item that had a bad administration.
(intended key: C; illustrative data)
Stem: A faculty member sees pooled stats for one of her NCLEX-style items (112 responses; illustrative data): p-value 0.93, discrimination index −0.16, option pull A 3%, B 2%, C 93% (key), D 2%. What is the correct next step?
A. Delete the item — a negative discrimination means it is broken B. Rekey it to whatever the top students chose C. Flag it for faculty review; decide by reading the item, not by the numbers alone D. Do nothing — the high p-value shows students learned the material
treating statistics as a verdict instead of a flag.
A negative discrimination with a very high p-value says the strongest examinees are missing an easy item — that smells like a miskey or a subtle flaw. But data never diagnose; they only flag. Deleting (A) destroys the diagnostic information, rekeying by vote (B) is not how keys are determined, and doing nothing (D) ignores the alarm. The rule: negative discrimination = faculty review, never auto-delete or auto-rekey.
— not applicable; this drill tests the decision rule itself.
C — flag for faculty review. Read the item with the statistics in hand: check the key against the content, look for an ambiguity that would trip strong examinees specifically (the classic cause: a subtle flaw that only careful readers notice), then decide to keep, revise, or retire. Both independent cold solvers selected C.
Recording it adds the module to your Academy completion record on this device — module, track, level, and date, ready to download from your account page for faculty-development documentation.