Item Psychometrics: How to Read the Facility and Discrimination Indices
At the end of every term the same conversation repeats in department meetings. A student challenges an item. The instructor defends it. The department head looks for some basis on which to settle the matter. In most cases the discussion resolves into impressions: the item was clear, the students did not prepare, the wording was ambiguous.
The irony is that a resolution is available and requires no measurement specialist. Two numbers computed from the exam's own data — the facility index and the discrimination index — are enough to move the argument from impression to evidence. A third, distractor analysis, usually tells you why an item failed rather than merely that it did.
This article explains how each indicator is calculated, how to read it, where interpretation commonly goes wrong, and includes a worked example from a realistically sized cohort, a decision matrix a department can adopt as written policy, and five cautions that will stop you from retiring a perfectly sound item.
Key Takeaways
- Facility measures how many answered correctly. Discrimination measures who answered correctly. The second matters far more, and is used far less.
- Difficulty is not a virtue. An item everyone gets wrong measures nothing, exactly like an item everyone gets right.
- A negative discrimination index is the single most alarming figure in the report. It usually means a keying error or ambiguous wording — not weak students.
- An option nobody selects does not make an item harder. It reduces the effective number of options and raises the guessing floor.
- These are screening tools, not verdicts. Read them against sample size and test purpose before acting.
Why the Author's Intuition Is Not Enough
When an instructor tags an item "hard" during authoring, they are estimating difficulty from the position of someone who knows the answer. That is an educated guess, not a measurement. Data from any active item bank shows that a substantial share of items labelled hard turn out to be answered easily, and the reverse.
The difference is that psychometric indicators are computed from actual learner behavior rather than authorial impression. Their value also compounds: the more often an item is used, the more precise its estimate becomes.
The Facility Index
How it is calculated
The proportion of candidates who answered correctly:
p = correct responses ÷ number of candidates
The value runs from 0 to 1. An item with p = 0.85 was answered correctly by 85% of candidates — meaning it is easy. Note the linguistic trap: the higher the number, the lower the difficulty. This reverses intuition and causes recurring confusion in committee discussions.
How to read it
| Range | Reading | Usual action |
|---|---|---|
| Above 0.90 | Very easy | Useful as a warm-up or in formative testing, not for discrimination |
| 0.40 – 0.80 | The productive band | Keep |
| 0.20 – 0.40 | Hard | Review; may stay if it discriminates well |
| Below 0.20 | Very hard | Check the key and the wording first |
For multiple-choice items, watch the guessing floor. On a four-option item, a candidate who knows nothing scores 25% by chance. A facility index near 0.25 therefore does not necessarily mean the item is extremely hard — it may mean candidates are guessing, which is an entirely different diagnosis.
The common error
Treating difficulty as evidence of quality or of "raising standards." An item that 95% of candidates get wrong does not separate the prepared candidate from the unprepared one; everyone failed it. Psychometrically it carries no information, however rigorous it looks.
The Discrimination Index
This is the indicator that answers the real question: does this item separate those who learned from those who did not?
Method one: upper and lower groups
Rank candidates by total score, then take the top and bottom slices (27% at each end is the common convention — a statistical compromise between group size and contrast):
D = proportion correct in upper group − proportion correct in lower group
The value runs from −1 to +1.
Method two: point-biserial correlation
A correlation between performance on a single item and the total score. It is more precise than the upper/lower method because it uses every candidate rather than only the extremes, and it is what modern platforms compute by default. Practical reading: 0.30 and above is good; below 0.15 warrants review.
Technical note: on short tests, use the corrected correlation — with the item's own score removed from the total — otherwise the item correlates partly with itself and reports inflated discrimination.
How to read it
| Range | Reading | Action |
|---|---|---|
| 0.40 and above | Excellent | Keep as is |
| 0.30 – 0.39 | Good | Keep; improve distractors if possible |
| 0.20 – 0.29 | Marginal | Needs revision |
| Below 0.20 | Weak | Revise substantially or retire |
| Negative | Alarm | Suspend immediately and check the key |
A negative index means high performers got the item wrong more often than low performers. This does not happen by chance. The three explanations, in order of likelihood: an error in the answer key; wording that carries two readings, with the stronger candidate picking up the subtler one; or an item that rewards shallow recall and penalizes deeper understanding.
Why the Two Numbers Are Linked
There is a mathematical constraint many people miss: discrimination is bounded by facility. If every candidate answers correctly (p = 1.0), the upper group's proportion equals the lower group's, and discrimination is necessarily zero. The same holds if everyone fails.
Discrimination therefore reaches its maximum when facility sits in the middle range. This is why the two are never read separately. An item with p = 0.95 and D = 0.05 is not a poorly discriminating item; it is a very easy item that was never given room to discriminate.
Distractor Analysis: Locating the Fault
The two indices tell you an item is unwell. The distribution of option selections tells you why.
- Dead distractor — an option chosen by fewer than 5% of candidates. It is doing no work, and the item is effectively a three-option question, which raises the guessing floor from 25% to 33%.
- A distractor more attractive than the key — if a wrong option outdraws the key, you are looking at one of two things: a keying error, or a misconception widespread enough in the cohort to deserve explicit teaching. Both are valuable findings.
- A distractor favoured by high performers — when a wrong option draws the upper group more than the lower group, the problem is precision of wording, not student ability.
A Worked Example: Three Items From a 200-Student Cohort
Upper and lower groups = 54 candidates each (27%).
Item 14 — a healthy item
120 correct out of 200, so facility is 0.60. Upper group: 48 of 54 correct (0.89). Lower group: 16 of 54 (0.30). So D = 0.89 − 0.30 = 0.59.
Squarely in the productive band with excellent discrimination. Keep it and leave it alone.
Item 7 — a red flag
84 correct out of 200, facility 0.42 — a figure that looks ideal at first glance. But the upper group scored 19 of 54 (0.35) and the lower group 26 of 54 (0.48). So D = −0.13.
A perfectly respectable facility index concealing a fundamental fault. This item should be suspended and its key reviewed before results are finalized. It is precisely why facility alone is insufficient.
Item 22 — failing distractors Option distribution: key (A) 42%, option (B) 51%, option (C) 5%, option (D) 2%. Option B outdraws the key itself — check the key first. If the key proves correct, you have identified a misconception shared by half the cohort, which is a more consequential finding than the grade. Options C and D are dead and should be replaced.
A Decision Matrix for the Department
Adopt this as written policy rather than individual judgment:
- Negative discrimination → suspend the item immediately, review the key, and rescore the cohort if an error is confirmed.
- Discrimination 0.40+ with facility between 0.40 and 0.80 → approved. Move it into the stable core of the item bank.
- Discrimination between 0.20 and 0.39 → revise the distractors. The fix is usually in the options, not the stem.
- Facility above 0.90 → move to formative use, or keep it deliberately as an opening item to settle candidates.
- Facility below 0.20 → check in order: the key, the wording, then whether the content was actually taught. The third possibility is not a fault in the item; it is information about the course.
- Any edit → saved as a new version. Never overwrite a version already used in a delivered assessment.
Five Cautions Before You Judge an Item
1. Sample size. Indices computed on 20 candidates are unstable. Treat them as a preliminary signal only, and do not base a retirement decision on fewer than 50 responses — preferably 100 or more.
2. Test purpose. On mastery tests, where everyone is expected to clear a threshold, a high facility index is the intended outcome rather than a defect. Discrimination will be low because variance is restricted, and that is expected.
3. Criterion circularity. Discrimination compares an item against the total test score. If the test as a whole is poorly constructed, the criterion itself is unreliable. Read item-level indicators only after confirming the test's overall reliability.
4. Item position. An item at the end of a time-pressured exam may show low facility because candidates never reached it, not because it was hard. Check the proportion of non-responses before drawing conclusions.
5. An indicator is not a verdict. These are screening tools that direct a human reviewer's attention to items worth examining. Whether an item stays or goes remains an academic judgment the department makes.
What This Requires From the Platform
Reading these indicators by hand after every exam is unrealistic in an institution running dozens of courses. To become a standing practice, the system needs to:
- Compute indices automatically from live response data rather than relying on a manual difficulty tag.
- Accumulate values across repeated administrations of the same item instead of computing from a single sitting.
- Surface unwell items in the department dashboard automatically, rather than waiting for someone to go looking.
- Tie every edit to a versioning cycle that preserves the integrity of historical reports.
- Rescore an entire cohort in a single reliable transaction when a key is corrected.
How EvaliX Addresses This
EvaliX computes facility and discrimination for every item automatically from actual responses, and accumulates those values across every administration rather than settling for a single-exam snapshot. The department dashboard surfaces unwell items — negative discrimination, dead distractors, and options outdrawing the key — without anyone having to search for them.
When an answer key is corrected, the platform runs an atomic rescore across the whole cohort and stores the change as a new item version, so previously issued reports remain exactly as issued.
Request a demo to see the item analysis report running on assessment data from your own institution.
FAQs
How many candidates are needed to compute these indices?
There is no hard threshold, but values become acceptably stable at around 100 responses or more. Between 50 and 100, read them as a signal worth investigating. Below 50, treat them as preliminary and do not base a retirement decision on them.
Do these indicators apply to essay items?
Discrimination does, in its correlational form, because it works with graded rather than binary scores. Facility is replaced by the mean score on the item as a proportion of its maximum. Distractor analysis does not apply by nature; the equivalent is examining how scores distribute across the rubric bands.
An item discriminates excellently but measures a marginal outcome — should I keep it?
No. Psychometric indicators measure an item's statistical efficiency, not its content validity. An item with excellent discrimination that measures a peripheral outcome raises the test's reliability while weakening its validity. Alignment to learning outcomes is a prior condition, not a substitute for statistical analysis.
