← Back to all projectsAgentic AI & LLM Systems

Golden Dataset & Evaluation Framework

Human-in-the-loop golden-dataset annotation workflow for conversational AI, with Cohen's/Fleiss' kappa inter-annotator agreement and Expected Calibration Error implemented from scratch, plus a CI-enforced evaluation gate that fails a build when a model or prompt version regresses against the golden set.

EvaluationAnnotationCalibrationCI/CD Gate
View source on GitHub