Skip to content

Latest commit

 

History

History
85 lines (65 loc) · 3.68 KB

File metadata and controls

85 lines (65 loc) · 3.68 KB

The data files

See ../NOTICE.md for provenance and licensing: the questions and the derived files have different legal statuses.

File Contents
questions.json the full bank: 4920 questions, 41 sessions, 2006–2026
schema.json JSON Schema for a single entry
taxonomy.json 28 topic tags with their textbook chapter mapping
statistics.csv psychometrics and tags, keyed by session and number, without question text

Entry format

{
  "specialization": "Endokrynologia",
  "question": "Podwyższone stężenie hormonu tyreotropowego (TSH) we krwi może świadczyć o:",
  "answers": { "A": "…", "B": "…", "C": "…", "D": "…", "E": "…" },
  "correct": "E",                    // "A"–"E", or null if annulled
  "statistics": { … },               // object, or null where CEM published no item analysis
  "number": 1,                       // 1–120
  "session": "20061",                // YYYY1 = spring, YYYY2 = autumn
  "tags": ["tarczyca-diagnostyka-i-guzki"],   // always present, 1–3 entries
  "annulled": true,                  // optional; only alongside correct: null
  "needs_review": "…"                // optional; internal provenance note, not for display
}

Entries are sorted by (session, number).

Two fields that are nullable

correct is null for the 20 questions annulled by CEM, which always also carry annulled: true. Exclude them from scoring; do not count them as wrong.

statistics is null for 1217 questions from sessions 2018–2022, for which CEM published no item analysis. Where present it contains:

  • difficulty: CEM's facility index, 0–1; higher means more candidates answered correctly
  • discrimination: item discrimination index
  • answer_distribution: percentage choosing each option, A–E
  • RPBI: point-biserial correlation for each option, A–E

All four are reproduced as CEM published them, directly from the source.

A caution on RPBI. Point-biserial correlation measures whether an item separates stronger candidates from weaker ones. A negative value for the keyed option means better candidates were less likely to choose it.

640 questions in this corpus have a negative point-biserial for the key, or a key that was not the most-chosen option. These keys were correct on real exam and were published by CEM. Six of them were checked individually against the CEM database to confirm this. Do not use RPBI or difficulty to infer that a key is wrong, and do not "correct" keys on statistical grounds.

Coverage

Sessions Questions With item statistics Source
2006–2017 2880 2880 web, CEM
2018–2020 720 0 PDF, NIL release
2021–2022 480 0 scans + OCR, NIL release
2023–2026 840 823 web, CEM
total 4920 3703

Every session contains exactly 120 questions numbered 1–120, with no gaps - although some have been annulled.

Validation

python src/qa/final_validation.py    # structure, schema conformance, tag validity
python src/qa/full_audit.py          # cross-session duplicates, key distribution, formatting
python src/qa/stress_test.py         # round-trip reconstruction against source files

The last requires the original source material, which is not redistributed here. Set PES_SOURCES to point at it. Published output from all three is in ../reports/.

Text encoding

UTF-8, normalised to Unicode NFC. Private-use-area characters, soft hyphens, decomposed diacritics and typographic spaces have been removed; see ../docs/quality-assurance.md. Normalise search input to NFC before matching against this data, or text pasted from other sources may fail to match.