BenchGen
Provider: DtCloud
BenchGen is the learning infrastructure for AI agents — an open platform where developers discover benchmarks and RL environments, connect and evaluate their complete agent systems, and continuously improve them against verifiable rewards. Most evaluation tools only trace what already happened. BenchGen closes the loop: benchmark → evaluate → fine-tune → re-evaluate, in one place. And unlike model-only leaderboards, BenchGen evaluates the whole agent — Model + Harness (memory, skills, tools, orchestration) — because most real-world agent failures live in the harness, not the model weights. Every benchmark on BenchGen carries a verifiable reward: a machine-checkable pass condition, not a subjective score. This is Test-Driven Development applied to agents — define the benchmark and the verifiable reward first, then build or fine-tune the agent until it consistently beats it. BenchGen ingests full agent decision trajectories (State → Action → Tool Response → Outcome → Reward), scores t
Freemium
Use Agent