Artificial Analysis has launched Optima, a platform for building AI benchmarks around a team’s own data, workflows, and task definitions. The goal is to make model comparisons more useful than broad public leaderboards for production decisions.

Users can upload evaluation datasets, connect agent traces from tools such as Arize, Braintrust, or Langfuse, or describe a use case with sample inputs and outputs. Optima can then suggest test inputs, evaluation criteria, and example tasks for review.

The platform supports rubric-based scoring and pairwise comparisons, where users judge sample response pairs and the system derives a ranking across the test set. It also reports cost per task and time per task, not just answer quality.

That matters for agentic applications. A model with cheaper tokens can cost more overall if it needs retries, fails tasks, or requires human cleanup. Optima does not guarantee that a custom benchmark captures business value, but it gives teams a more relevant starting point for model selection than generic scores alone.