Automated Essay Scoring on The Hewlett Foundation dataset on Kaggle
-
Updated
Apr 26, 2018 - Jupyter Notebook
Automated Essay Scoring on The Hewlett Foundation dataset on Kaggle
Contributed to a vision-driven accessibility tool translating sign language into text
🤟 Enhance sign language interpretation using transfer learning and multimodal features for accurate gesture recognition and robust evaluation methods.
Measure whether your model judge agrees with human raters. Chance-corrected agreement statistics, bootstrap confidence intervals, and calibration gates for Swift Testing and CI. Zero dependencies.
Python | scikit-learn, SVM, SMOTE | Automates systematic literature review screening with 95%+ recall, reducing manual workload by 64%.
💬 Advanced NLP with Spacy Course
Labeling queue library for managing human labeling workflows
Calibrate an LLM-as-judge against human labels: Cohen's kappa, accuracy CI, position and verbosity bias probes, and an honest reliable-for-X-not-Y verdict. Stdlib only, $0 offline path.
In this final project, you will explore the interdisciplinary domain of computational social science. You will study how large language models can support qualitative coding, a social science research method that involves assigning categorical labels to open-ended text data.
Does a CLAUDE.md actually change how Claude behaves? An ablation harness: run adversarial traps with the rules and without them, grade blind, and test whether the difference is real.
Evaluation and agreement scripts for the DISCOSUMO project. Each evaluation script takes both manual annotations as automatic summarization output. The formatting of these files is highly project-specific. However, the evaluation functions for precision, recall, ROUGE, Jaccard, Cohen's kappa and Fleiss' kappa may be applicable to other domains too.
Cross-Family LLM-Judge Agreement for Institutional RAG: 5 families, 9 judges. Validated on TREC RAG 2024 (kappa=0.4941) + BEIR scifact.
Tiny zero-dependency evaluation kit for LLM-as-a-Judge: agreement %, Cohen's kappa, drift, and a judge-prompt bias linter.
Statistical validation of labeling consistency across three independent raters for a handwritten digit classification dataset.
Chance-corrected inter-rater agreement statistics in pure numpy: Cohen and Fleiss kappa, Krippendorff alpha, ICC, and bootstrap confidence intervals, validated against published reference values.
Three LLM judges measured against human expert labels on FaithBench. Chance correction removes about 34 points of the agreement the field reports, all three land below a rule that flags every summary unread, and across a 40.6-point band of release thresholds every judge ships what the human labels would block.
AI-powered ECGT compliance screening tool for SME environmental claims
Measure how much your LLM judges actually agree. Inter-judge agreement metrics for LLM-as-a-judge evaluations.
AI agent evaluation harness with Scout - CSV data analyst with dual-judge validation, CI regression testing, and adversarial red-teaming
To associate your repository with the cohens-kappa topic, visit your repo's landing page and select "manage topics."