ground-truth-evals MCP Server
yangstar89/ground-truth-evals
An LLM eval harness graded by computed ground truth, not an LLM judge. Worked example: poker, served to models as MCP tools — gpt-4o-mini goes from 26/75 to 75/75 with them.
claude mcp add agentrank -- npx -y agentrank-mcp-server Overview
yangstar89/ground-truth-evals is a JavaScript tool licensed under MIT. An LLM eval harness graded by computed ground truth, not an LLM judge. Worked example: poker, served to models as MCP tools — gpt-4o-mini goes from 26/75 to 75/75 with them.
Ranked #8906 out of 24840 indexed tools.
Actively maintained with commits in the last week.
Ecosystem
Score Breakdown
1 stars → early stage
Last commit 2d ago → actively maintained
No issues filed → no history to score
0 contributors → solo project
No dependents → no downstream usage
Weights: Freshness 20% · Issue Health 20% · Dependents 22% · Stars 10% · Contributors 8% · How we score →
How to Improve
Matched Queries
Get the weekly AgentRank digest
Top movers, new tools, ecosystem insights — straight to your inbox.