Benchmark Dashboard
CCX aggregates model capability, cost, and multi-source comparison data from several public benchmark sources and publishes this interactive chart.
- The upper section shows the capability-cost frontier, with switches for mean / median cost and source scope.
- The lower section compares raw scores for the same model across multiple benchmark sources.
- Because benchmark task sets, scoring rules, and scales differ, cross-source scores should not be treated as a strictly normalized ranking.
Data updated at:unknown
Methodology notes
pass@1and raw scores come from the corresponding upstream benchmark results or public APIs.- The cost frontier currently uses the single-task cost data that can be mapped from DeepSWE and CodexRadar.
- The multi-source comparison chart juxtaposes raw scores from DeepSWE, BenchLM.ai, CodexRadar, Artificial Analysis, and related public sources.
- For the generation timestamp, see the
generatedAtfield inbenchmark-viz-data.json; that file is generated alongside the chart HTML in the docs static asset directory.
Data sources and acknowledgments
| Benchmark data on this page includes Artificial Analysis free API data. Attribution to artificialanalysis.ai is required when using that data. Intelligence Index scores are currently interpreted against v4.1.1; Coding Index and Agentic Index are derived subsets of the same evaluation set and are not separately versioned. | |
| Data | CCX benchmark visualizations and registry updates also use publicly available data from BenchLM.ai, DeepSWE, and CodexRadar. Pricing and context metadata are additionally refreshed from LiteLLM model_prices_and_context_window.json. |