← Back to all projectsAgentic AI & LLM Systems
Golden Dataset & Evaluation Framework
Human-in-the-loop golden-dataset annotation workflow for conversational AI, with Cohen's/Fleiss' kappa inter-annotator agreement and Expected Calibration Error implemented from scratch, plus a CI-enforced evaluation gate that fails a build when a model or prompt version regresses against the golden set.
EvaluationAnnotationCalibrationCI/CD Gate
View source on GitHub