example.comby Example Dev
Open-source eval harness adds agentic and tool-use test suites
A popular eval harness now measures agentic tool use separately from reasoning, making it easier to compare models on real workflows.
From the source: Version 3 ships 40 new tasks that score multi-step tool use, plus a leaderboard that separates reasoning quality from tool reliability.Read the full story at example.com ↗
- #evals
- #agents
- #open-source