AI Agents Marketplace

AI Testing Assistant AI Agents in AI Development

3 AI Testing Assistant AI agent(s) in AI Development.

BenchGen

Provider: DtCloud

BenchGen is the learning infrastructure for AI agents — an open platform where developers discover benchmarks and RL environments, connect and evaluate their complete agent systems, and continuously improve them against verifiable rewards. Most evaluation tools only trace what already happened. BenchGen closes the loop: benchmark → evaluate → fine-tune → re-evaluate, in one place. And unlike model-only leaderboards, BenchGen evaluates the whole agent — Model + Harness (memory, skills, tools, orchestration) — because most real-world agent failures live in the harness, not the model weights. Every benchmark on BenchGen carries a verifiable reward: a machine-checkable pass condition, not a subjective score. This is Test-Driven Development applied to agents — define the benchmark and the verifiable reward first, then build or fine-tune the agent until it consistently beats it. BenchGen ingests full agent decision trajectories (State → Action → Tool Response → Outcome → Reward), scores t

Freemium

Use Agent

Maxim AI

Provider: Maxim AI

Maxim is an agent simulation, evaluation, and observability platform that empowers modern AI teams to deploy agents with quality, reliability, and speed. Maxims end-to-end evaluation and data management stack covers every stage of the AI lifecycle, from prompt engineering to pre & post release testing and observability, data-set creation & management, and fine-tuning. Use Maxim to simulate and test your multi-turn workflows on a wide variety of scenarios and across different user personas before taking your application to production. Features: Agent Simulation Agent Evaluation Prompt Playground Logging/Tracing Workflows Custom Evaluators- AI, Programmatic and Statistical Dataset Curation Human-in-the-loop

$29

Use Agent

RagaAI Inc.

Provider: RagaAI

AI platform for robust, multimodal testing of AI systems.

Freemium

Use Agent