Claim listing
mbsdeepak/gauntlet
A rigorous agentic tool-use eval set for LLMs: deterministic simulated tool environments, state/trajectory/LLM-judge grading, pass@k with Wilson confidence intervals, and cost/latency tracking.
Claim your listing to add a tagline, logo, and category. Verified maintainers get a Verified Publisher badge and priority placement on the AgentRank index.
Leave your email to claim this listing. GitHub verification coming soon.