Skip to content

Benchmark Dashboard

CCX aggregates model capability, cost, and multi-source comparison data from several public benchmark sources and publishes this interactive chart.

  • The upper section shows the capability-cost frontier, with switches for mean / median cost and source scope.
  • The lower section compares raw scores for the same model across multiple benchmark sources.
  • Because benchmark task sets, scoring rules, and scales differ, cross-source scores should not be treated as a strictly normalized ranking.
Data updated at:unknown

Methodology notes

  • pass@1 and raw scores come from the corresponding upstream benchmark results or public APIs.
  • The cost frontier currently uses the single-task cost data that can be mapped from DeepSWE and CodexRadar.
  • The multi-source comparison chart juxtaposes raw scores from DeepSWE, BenchLM.ai, CodexRadar, Artificial Analysis, and related public sources.
  • For the generation timestamp, see the generatedAt field in benchmark-viz-data.json; that file is generated alongside the chart HTML in the docs static asset directory.

Data sources and acknowledgments

Artificial Analysis Logo Benchmark data on this page includes Artificial Analysis free API data. Attribution to artificialanalysis.ai is required when using that data. Intelligence Index scores are currently interpreted against v4.1.1; Coding Index and Agentic Index are derived subsets of the same evaluation set and are not separately versioned.
Data CCX benchmark visualizations and registry updates also use publicly available data from BenchLM.ai, DeepSWE, and CodexRadar. Pricing and context metadata are additionally refreshed from LiteLLM model_prices_and_context_window.json.