See ../NOTICE.md for provenance and licensing: the questions and the
derived files have different legal statuses.
| File | Contents |
|---|---|
questions.json |
the full bank: 4920 questions, 41 sessions, 2006–2026 |
schema.json |
JSON Schema for a single entry |
taxonomy.json |
28 topic tags with their textbook chapter mapping |
statistics.csv |
psychometrics and tags, keyed by session and number, without question text |
Entries are sorted by (session, number).
correct is null for the 20 questions annulled by CEM, which always also carry
annulled: true. Exclude them from scoring; do not count them as wrong.
statistics is null for 1217 questions from sessions 2018–2022, for which CEM
published no item analysis. Where present it contains:
difficulty: CEM's facility index, 0–1; higher means more candidates answered correctlydiscrimination: item discrimination indexanswer_distribution: percentage choosing each option, A–ERPBI: point-biserial correlation for each option, A–E
All four are reproduced as CEM published them, directly from the source.
A caution on RPBI. Point-biserial correlation measures whether an item separates
stronger candidates from weaker ones. A negative value for the keyed option means better
candidates were less likely to choose it.
640 questions in this corpus have a negative point-biserial for the key, or a key that
was not the most-chosen option. These keys were correct on real exam and were published by CEM. Six
of them were checked individually against the CEM database to confirm this. Do not use
RPBI or difficulty to infer that a key is wrong, and do not "correct" keys on
statistical grounds.
| Sessions | Questions | With item statistics | Source |
|---|---|---|---|
| 2006–2017 | 2880 | 2880 | web, CEM |
| 2018–2020 | 720 | 0 | PDF, NIL release |
| 2021–2022 | 480 | 0 | scans + OCR, NIL release |
| 2023–2026 | 840 | 823 | web, CEM |
| total | 4920 | 3703 |
Every session contains exactly 120 questions numbered 1–120, with no gaps - although some have been annulled.
python src/qa/final_validation.py # structure, schema conformance, tag validity
python src/qa/full_audit.py # cross-session duplicates, key distribution, formatting
python src/qa/stress_test.py # round-trip reconstruction against source filesThe last requires the original source material, which is not redistributed here.
Set PES_SOURCES to point at it. Published output from all three is in ../reports/.
UTF-8, normalised to Unicode NFC. Private-use-area characters, soft hyphens,
decomposed diacritics and typographic spaces have been removed; see
../docs/quality-assurance.md. Normalise search input to NFC before matching against
this data, or text pasted from other sources may fail to match.
{ "specialization": "Endokrynologia", "question": "Podwyższone stężenie hormonu tyreotropowego (TSH) we krwi może świadczyć o:", "answers": { "A": "…", "B": "…", "C": "…", "D": "…", "E": "…" }, "correct": "E", // "A"–"E", or null if annulled "statistics": { … }, // object, or null where CEM published no item analysis "number": 1, // 1–120 "session": "20061", // YYYY1 = spring, YYYY2 = autumn "tags": ["tarczyca-diagnostyka-i-guzki"], // always present, 1–3 entries "annulled": true, // optional; only alongside correct: null "needs_review": "…" // optional; internal provenance note, not for display }