Skip to content
#

cohens-kappa

Here are 41 public repositories matching this topic...

🤟 Enhance sign language interpretation using transfer learning and multimodal features for accurate gesture recognition and robust evaluation methods.

  • Updated Sep 12, 2026
  • Python

Does a CLAUDE.md actually change how Claude behaves? An ablation harness: run adversarial traps with the rules and without them, grade blind, and test whether the difference is real.

  • Updated Jul 23, 2026
  • Python

Evaluation and agreement scripts for the DISCOSUMO project. Each evaluation script takes both manual annotations as automatic summarization output. The formatting of these files is highly project-specific. However, the evaluation functions for precision, recall, ROUGE, Jaccard, Cohen's kappa and Fleiss' kappa may be applicable to other domains too.

  • Updated Feb 10, 2017
  • Python

Three LLM judges measured against human expert labels on FaithBench. Chance correction removes about 34 points of the agreement the field reports, all three land below a rule that flags every summary unread, and across a 40.6-point band of release thresholds every judge ships what the human labels would block.

  • Updated Sep 7, 2026
  • Python

Add this topic to your repo

To associate your repository with the cohens-kappa topic, visit your repo's landing page and select "manage topics."

Learn more