A local-first MCP host with pluggable skills, multi-provider routing, and multi-channel I/O. Kotlin. Plus a Python eval harness with rule + LLM-judge scorers over a human-authored golden set.
-
Updated
Jul 17, 2026 - Python
A local-first MCP host with pluggable skills, multi-provider routing, and multi-channel I/O. Kotlin. Plus a Python eval harness with rule + LLM-judge scorers over a human-authored golden set.
Evaluation dataset quality auditor for LLM and RAG applications. Checks golden sets for conflicting labels, duplicate prompts, weak reference answers, ambiguous questions, over-easy examples, and category coverage gaps.
Minimal Python script to evaluate your RAG pipeline against a golden set.
Add a description, image, and links to the golden-set topic page so that developers can more easily learn about it.
To associate your repository with the golden-set topic, visit your repo's landing page and select "manage topics."