OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.
-
Updated
Sep 11, 2026 - TypeScript
OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.
Measure whether your model judge agrees with human raters. Chance-corrected agreement statistics, bootstrap confidence intervals, and calibration gates for Swift Testing and CI. Zero dependencies.
Cross-model LLM-as-judge eval harness: validate AI judges with Fleiss' kappa / Krippendorff's alpha, not accuracy. Ships a real 7-model panel (Claude, GPT, Gemini, Grok, Qwen, DeepSeek, GLM) you can replay in ~30s, no API key. MIT.
Read and analyse the qualitative coding exported by the zotQDA and qdaZ Zotero plugins, or from REFI-QDA projects
Statistical validity checks for human-graded AI evaluations
Tool-agnostic inter-coder reliability (Krippendorff alpha, Cohen/Fleiss kappa) and disagreement adjudication for qualitative coding
Classificazione dei commenti YouTube di Breaking Italy e La Repubblica tramite il modello User Needs (Shishkin & SmartOcto, 2021). Tesi magistrale in Comunicazione, ICT e Media — UniTO.
A reliability and DIF report card for LLM-judge and human-rater scoring instruments.
Measure how much your LLM judges actually agree. Inter-judge agreement metrics for LLM-as-a-judge evaluations.
Browser-based inter-rater reliability calculator: Krippendorff's alpha, Fleiss's kappa, Cohen's kappa, percent agreement, verified against R and Python
Browser-based inter-rater reliability calculator for systematic literature reviews. Computes Krippendorff's Alpha using a pooled coincidence matrix across all (Paper, RQ, Field) units. No installation required — single HTML file, fully client-side. Built for the GenAI Evidence Hub at Learning Data Insights.
Validate LLM-as-judge reliability with inter-rater agreement, not accuracy. Reproduces a TypeScript reference implementation on a real 7-model panel.
Measures whether your graders and your rubric are trustworthy: inter-annotator agreement, anchor diagnostics, grader calibration against gold, and an adjudication queue.
Reliability, rogue-rater, drift and leakage diagnostics for labelled data — on the raw GoEmotions ratings, 27 of 28 emotions fall below the 0.667 agreement floor.
Chance-corrected inter-rater agreement statistics in pure numpy: Cohen and Fleiss kappa, Krippendorff alpha, ICC, and bootstrap confidence intervals, validated against published reference values.
Analyse the qualitative coding exported by the zotQDA and qdaZ Zotero plugins, or from REFI-QDA projects
Evaluation toolkit for multi-annotator human annotation research with agreement metrics, disagreement analysis, reports, and plots.
To associate your repository with the krippendorff-alpha topic, visit your repo's landing page and select "manage topics."