AI Evaluation & Testing
Braintrust
Braintrust is a ai agent observability and evaluation platform for inspect production agent behavior and evaluate agent quality before release.
Starter has $0 platform fee with included free usage; usage applies beyond included limits.approx.
AI Evaluation & Testing
OpenAI Evals
OpenAI Evals is an open-source framework and benchmark registry for evaluating LLMs and LLM systems.
Open-source framework; API/model usage costs depend on the evaluation run.approx.
AI Evaluation & Testing
DeepEval
DeepEval is an open-source LLM evaluation framework that supports evaluating LLM applications, unit-testing LLM outputs and running end-to-end evals.
Open-source framework plus Confident AI cloud plans; exact public prices not captured.approx.
AI Evaluation & Testing
Ragas
Ragas is an AI evaluation library for systematic eval loops, metrics, datasets, RAG tests, agent checks and prompt experiments.
Open-source library; pricing not captured in this pass.approx.
Stack Tribune may earn a commission from some outbound links. Category rankings are not sold.