Claim listing
yangstar89/ground-truth-evals
An LLM eval harness graded by computed ground truth, not an LLM judge. Worked example: poker, served to models as MCP tools — gpt-4o-mini goes from 26/75 to 75/75 with them.
Claim your listing to add a tagline, logo, and category. Verified maintainers get a Verified Publisher badge and priority placement on the AgentRank index.
Leave your email to claim this listing. GitHub verification coming soon.