quica is a tool to run inter coder agreement pipelines in an easy and effective ways. Multiple measures are run and results are collected in a single table than can be easily exported in Latex
-
Updated
Nov 9, 2020 - Python
quica is a tool to run inter coder agreement pipelines in an easy and effective ways. Multiple measures are run and results are collected in a single table than can be easily exported in Latex
Qualitative Coding Assistant for Google Sheets
Kupper-Hafner inter-rater agreement calculation library
The official Crowd Deliberation data set.
A python script to compute kappa-coefficient, which is a statistical measure of inter-rater agreement.
Assessment of radiological abnormalities and prediction of lung function with AI-determined lung opacity
Evaluation and agreement scripts for the DISCOSUMO project. Each evaluation script takes both manual annotations as automatic summarization output. The formatting of these files is highly project-specific. However, the evaluation functions for precision, recall, ROUGE, Jaccard, Cohen's kappa and Fleiss' kappa may be applicable to other domains too.
A study in "Agent Trajectory Evaluation": judging an AI agent's full Execution Path, Not just its final Answer. A small ReAct Agent generates Trajectories; ~40 are hand-labelled as Ground truth; 'Rule-based and LLM-judge scorers' are measured against 'those labels'. The Deliverable is the agreement analysis, and where Automated Evaluation breaks.
Python tool for calculating inter-rater reliability metrics and generating comprehensive reports for multi-rater datasets. Optionally have an LLM create an interpretation report.
Cross-Family LLM-Judge Agreement for Institutional RAG: 5 families, 9 judges. Validated on TREC RAG 2024 (kappa=0.4941) + BEIR scifact.
[MICCAI ISIC 2024] Code for "Segmentation Style Discovery: Application to Skin Lesion Images"
A calculator for two different inter-rater agreement statistics, generalized to any numbers of categories
Three LLM judges measured against human expert labels on FaithBench. Chance correction removes about 34 points of the agreement the field reports, all three land below a rule that flags every summary unread, and across a 40.6-point band of release thresholds every judge ships what the human labels would block.
Measure how much your LLM judges actually agree. Inter-judge agreement metrics for LLM-as-a-judge evaluations.
Replication package for the Archetypal Analysis conducted in the paper: Evaluating the Agreement among Technical Debt Measurement Tools: Building an Empirical Benchmark of Technical Debt Liabilities accepted at Springer's EMSE Journal.
Build, adjudicate, freeze and audit golden sets for LLM evaluation: stratified sampling with reported shortfalls, inter-annotator agreement, adjudication that refuses to guess, tamper-evident freezing, drift detection. Dependency-free.
Turns raw production logs into a labeled, deduplicated, statistically validated eval dataset. Human-vs-judge agreement measured double-blind (Cohen's κ + bootstrap CI95): a first run blocked at κ=0.26, drove a guideline fix, then cleared at κ=0.80 on the target domain. Honest dedup, HDBSCAN coverage, sha256 provenance, 677 offline tests.
Pairwise rating CLI for AI responses with per-axis scoring (helpfulness/harmlessness/accuracy/instruction-following), JSONL in/out, inter-rater Cohen's kappa
To associate your repository with the inter-rater-agreement topic, visit your repo's landing page and select "manage topics."