Explore 6+ Testing and Evaluation AI agents. Compare features, pricing, and ratings. Security-validated by One9Founders.
A comprehensive testing and evaluation framework for voice agents across language models, prompts, and agent personas. 
an open source RAG evaluation framework that does not require golden answers, and can be used to evaluate performance of RAG tools connected to an AI Agent (Agentic RAG)
EvoAgentX is building a Self-Evolving Ecosystem of AI Agents, it will give you automated framework for evaluating and evolving agentic workflows. 
Arize-Phoenix is an open source library for agent testing, evaluation and observability. 
Open-source, real-time cost observability platform for AI agents. Track tokens, costs, messages, and model usage with a local-first dashboard. Supports 28+ LLM models, OTLP ingestion, self-hosted. 
Self-improving agentic QA harness for web and mobile tests. Write tests in natural language, use memory to adapt to UI changes, and catch regressions before releases ship. 