Nick Item-Writing Academy
About 40 minutes · Academy module: Peer-Review Workshop.
Every item gets the same pass, in the same order: (1) Tested point — can you state it in one sentence? If not, the item is unfocused. (2) Key — solve it cold, without knowing the author's answer. Do you converge on the key? (3) Distractors — is each one plausible, same-family, and single-best? (4) Flaws — run the Flaw Clinic taxonomy: cueing, negatives, AOTA/NOTA, absolutes, two-correct. (5) Mapping — one test-plan category; if you and the author map it differently, flag it. (6) Bias/sensitivity — the five-minute pass. Reviewers who free-read find the flaws they already look for; reviewers who checklist find the flaws that are actually there.
Good review feedback is specific ("option C is a grammatical outlier — 'An assessment of' breaks the parallel noun phrases"), content-grounded ("the near-miss distractor should be hyperkalemia, the confusable opposite, not hypernatremia"), and prioritized (the key dispute first, the wording polish last). Bad feedback is vague ("this seems confusing"), personal ("you always write tricky items"), or a rewrite delivered as a verdict. Our read: the reviewer's job is to diagnose, not to rewrite. Hand the author the diagnosis; let them own the repair — they will defend a repair they authored and resent one imposed.
The author's job is to listen like a scientist: the review is data about how the item reads to a competent stranger, which is exactly what the examinee is. Do not defend the draft — if a reviewer misread your item, an examinee will too, and "they should have read more carefully" is not a validity argument. Record the verdict per checklist item (accept / revise / dispute with reason), revise, and re-solve cold. Disputes about the key go to a third solver, never to a vote.
A review panel needs three things: the checklist (same one every time), cold solves (reviewers solve before seeing the key — never after), and a quorum rule (one key dispute = revise; two reviewers flagging the same flaw = fix before administration). Timebox ruthlessly: fifteen minutes per item keeps panels honest. Rubber-stamp reviews — "looks good to me" in ninety seconds — are worse than no review, because they certify what they never examined.
(a review transcript)
Reviewer: "Okay, item 12, the potassium one. Looks fine — the answer's obviously hypokalemia. Anybody disagree? No? Moving on. Item 13…"
rubber-stamp review — no cold solve, no checklist, ninety seconds.
The reviewer announced the key instead of solving for it, asked for disagreement instead of independent judgment, and checked nothing — not the distractors (are they same-family?), not the flaws (any cueing?), not the mapping. "Obviously hypokalemia" is exactly what a cueing flaw feels like from the inside. A real review of this item takes fifteen minutes: cold solve, checklist pass, mapping check, bias pass. The transcript below shows the repair.
— the same review, done properly:
Reviewer: "Item 12. Solving cold: the stem gives furosemide, muscle weakness, potassium 2.9 — I converge on hypokalemia, option A. Tested point: one sentence — 'loop diuretics waste potassium; recognize hypokalemia signs.' Distractors: B is hyperkalemia, good near-miss; C is hyponatremia, plausible; D is hypermagnesemia — different family, passenger. Suggest replacing D with hypocalcemia. No cueing found; options parallel. Maps to Physiological Adaptation. Bias pass: clean. Verdict: accept with D revised."
The proper review solves cold before naming the key, walks the checklist aloud, grounds the one criticism in content (D is the wrong family — replace it with a specific better option), and ends with an explicit verdict. Fifteen minutes, one actionable revision, no rubber stamp.
Recording it adds the module to your Academy completion record on this device — module, track, level, and date, ready to download from your account page for faculty-development documentation.