Skip to main content
Verified model research

AI model benchmarks for practical decisions

Compare 68 models released since 2024 across 5 task families. Model records include linked official sources, while benchmark results retain their compatible evaluation snapshot.

68 models13 providers5 task familiesDataset 2026-08-14Verified 2026-08-14

Model explorer

Choose one task family, then filter its validated inventory. Charts and comparisons use compatible benchmark snapshots and pricing units; missing values remain visible as Not reported.

Choose a task

Compare models only within one compatible task family.

Decision summary

Facts from one compatible snapshot, measured 2026-06-25.

LiveBench · 2026-06-25

Highest score on LiveBench Overall

GPT-5.5

79.91 points on LiveBench Overall, snapshot 2026-06-25, measured 2026-06-25. Eligible models (4): GPT-5.4, GPT-5.5, Claude Opus 4.8, Claude Sonnet 5. Price unit: USD per 1M input tokens.

Methodology

Cost and quality frontier

3 nondominated models

LiveBench Overall, snapshot 2026-06-25, measured 2026-06-25, compared with USD per 1M input tokens. Eligible models (4): GPT-5.4, GPT-5.5, Claude Opus 4.8, Claude Sonnet 5.

Methodology
More filters
Providers

110 of 51 models

0 of 4 selected

AI model inventory with pricing, selected benchmark, provenance, and comparison controls
Grok 4.6CurrentProviderxAIRelease2026-08-12AccessAPICapabilitiesreasoningOperational limitsContext 500,000 · max output Not reportedPrice / 1M input tokens$2 / 1M input tokens · standardPrice / 1M output tokens$6 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • Grok 4.6 is positioned for coding, agentic tasks, and knowledge work; its API documentation also lists function calling and structured outputs
Cautions
  • Published token rates increase for all tokens in a request once its prompt reaches xAI's 200,000-token long-context threshold
Provenance
Gemini 3.5 Flash-LiteCurrentProviderGoogleRelease2026-07-21AccessAPICapabilitiesreasoningOperational limitsContext 1,048,576 · max output 65,536Price / 1M input tokens$0.3 / 1M input tokens · standardPrice / 1M output tokens$2.5 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • Gemini 3.5 Flash-Lite is Google's low-latency, cost-efficient 3.5 model for high-throughput subagent and document-processing tasks
Cautions
  • The represented prices are standard API rates; batch, flex, and priority consumption use different rates
Provenance
Gemini 3.6 FlashCurrentProviderGoogleRelease2026-07-21AccessAPICapabilitiesreasoningOperational limitsContext 1,048,576 · max output 65,536Price / 1M input tokens$1.5 / 1M input tokens · standardPrice / 1M output tokens$7.5 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • Gemini 3.6 Flash supports multimodal inputs and a 1,048,576-token input limit for agentic and coding workflows
Cautions
  • This pinned LiveBench snapshot predates Gemini 3.6 Flash, so no compatible overall result is reported here
Provenance
Grok 4.5CurrentProviderxAIRelease2026-07-16AccessAPICapabilitiesreasoningOperational limitsContext 500,000 · max output Not reportedPrice / 1M input tokens$2 / 1M input tokens · standardPrice / 1M output tokens$6 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • Grok 4.5 is positioned for coding, agentic tasks, and knowledge work with a 500,000-token context window
Cautions
  • The published base rates increase for requests at or above the provider long-context threshold
Provenance
GPT-5.6 LunaCurrentProviderOpenAIRelease2026-07-09AccessAPICapabilitiesreasoningOperational limitsContext 1,050,000 · max output 128,000Price / 1M input tokens$0.2 / 1M input tokens · standardPrice / 1M output tokens$1.2 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • GPT-5.6 Luna is the cost-sensitive GPT-5.6 tier for efficient high-volume workloads
Cautions
  • This pinned LiveBench snapshot predates GPT-5.6, so no compatible overall result is reported here
Provenance
GPT-5.6 SolCurrentProviderOpenAIRelease2026-07-09AccessAPICapabilitiesreasoningOperational limitsContext 1,050,000 · max output 128,000Price / 1M input tokens$5 / 1M input tokens · standardPrice / 1M output tokens$30 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • GPT-5.6 Sol is the flagship GPT-5.6 tier for complex reasoning, coding, and professional work
Cautions
  • This pinned LiveBench snapshot predates GPT-5.6, so no compatible overall result is reported here
Provenance
GPT-5.6 TerraCurrentProviderOpenAIRelease2026-07-09AccessAPICapabilitiesreasoningOperational limitsContext 1,050,000 · max output 128,000Price / 1M input tokens$2 / 1M input tokens · standardPrice / 1M output tokens$12 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • GPT-5.6 Terra is the balanced GPT-5.6 tier for capability and token cost
Cautions
  • This pinned LiveBench snapshot predates GPT-5.6, so no compatible overall result is reported here
Provenance
Claude Sonnet 5CurrentProviderAnthropicRelease2026-06-30AccessAPICapabilitiesreasoningOperational limitsContext Not reported · max output Not reportedPrice / 1M input tokens$2 / 1M input tokens · standardPrice / 1M output tokens$10 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • Claude Sonnet 5 is positioned for agentic execution, tool use, coding, and knowledge work
Cautions
  • The launch price is introductory through 2026-08-31; refresh pricing before evaluating workloads after that date
Provenance
Claude Fable 5CurrentProviderAnthropicRelease2026-06-09AccessAPICapabilitiesreasoningOperational limitsContext 1,000,000 · max output 128,000Price / 1M input tokens$10 / 1M input tokens · standardPrice / 1M output tokens$50 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • Claude Fable 5 is Anthropic's highest-capability widely released model for demanding reasoning and long-horizon agentic work
Cautions
  • Fable requires 30-day data retention, and classifier-triggered requests fall back to Claude Opus 4.8
Provenance
Claude Opus 4.8CurrentProviderAnthropicRelease2026-05-28AccessAPICapabilitiesreasoningOperational limitsContext 1,000,000 · max output Not reportedPrice / 1M input tokens$5 / 1M input tokens · standardPrice / 1M output tokens$25 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • Claude Opus 4.8 is positioned for coding, agentic tasks, and professional work with a published 1,000,000-token context window
Cautions
  • Maximum output length is not represented because the approved source set does not establish one unambiguously
Provenance

Decision charts

Compare declared coverage and compatible operational values for the current task.

Capability coverage

Boolean provider-declared coverage for 10 current-task models. Cells are not scores and do not compare capability quality.

Legend: ✓ Supported · × Unsupported

Capability coverage matrixRows are models and columns are task-relevant capabilities. Each status chip explicitly states whether the model supports that capability.

Grok 4.6

ReasoningSupported

Gemini 3.5 Flash-Lite

ReasoningSupported

Gemini 3.6 Flash

ReasoningSupported

Grok 4.5

ReasoningSupported

GPT-5.6 Luna

ReasoningSupported

GPT-5.6 Sol

ReasoningSupported

GPT-5.6 Terra

ReasoningSupported

Claude Sonnet 5

ReasoningSupported

Claude Fable 5

ReasoningSupported

Claude Opus 4.8

ReasoningSupported
View exact capability data
Capability coverage data
ModelReasoning
Grok 4.6Supported
Gemini 3.5 Flash-LiteSupported
Gemini 3.6 FlashSupported
Grok 4.5Supported
GPT-5.6 LunaSupported
GPT-5.6 SolSupported
GPT-5.6 TerraSupported
Claude Sonnet 5Supported
Claude Fable 5Supported
Claude Opus 4.8Supported

Operational comparisons

Each panel has its own unit, direction, and fixed displayed range. Values are never combined across panels.

LiveBench Overall

points · Higher is better · Range 0100

LiveBench Overall5 models compared in points. Higher is better. Exact values are in the table below.Claude Sonnet 574.85Claude Opus 4.878.93Gemini 3.5 Flash74.64GPT-5.579.91GPT-5.477.97

Price per 1M input tokens

USD / 1M input tokens · Lower is better for cost · Range 010

Price per 1M input tokens17 models compared in USD / 1M input tokens. Lower is better for cost. Exact values are in the table below.Grok 4.62Gemini 3.5 Flash-Lite0.3Gemini 3.6 Flash1.5Grok 4.52GPT-5.6 Luna0.2GPT-5.6 Sol5GPT-5.6 Terra2Claude Sonnet 52Claude Fable 510Claude Opus 4.85GPT-5.55Grok 4.201.25GPT-5.42.5GPT-4.12GPT-4.1 mini0.4GPT-4.1 nano0.1Grok 4.31.25

Price per 1M output tokens

USD / 1M output tokens · Lower is better for cost · Range 050

Price per 1M output tokens17 models compared in USD / 1M output tokens. Lower is better for cost. Exact values are in the table below.Grok 4.66Gemini 3.5 Flash-Lite2.5Gemini 3.6 Flash7.5Grok 4.56GPT-5.6 Luna1.2GPT-5.6 Sol30GPT-5.6 Terra12Claude Sonnet 510Claude Fable 550Claude Opus 4.825GPT-5.530Grok 4.202.5GPT-5.415GPT-4.18GPT-4.1 mini1.6GPT-4.1 nano0.4Grok 4.32.5

Context window

tokens · Higher supports larger inputs · Range 01,050,000

Context window23 models compared in tokens. Higher supports larger inputs. Exact values are in the table below.Grok 4.6500,000Gemini 3.5 Flash-Lite1,048,576Gemini 3.6 Flash1,048,576Grok 4.5500,000GPT-5.6 Luna1,050,000GPT-5.6 Sol1,050,000GPT-5.6 Terra1,050,000Claude Fable 51,000,000Claude Opus 4.81,000,000Gemini 3.5 Flash1,048,576GPT-5.51,050,000Grok 4.201,000,000GPT-5.41,050,000Kimi K2 Base128,000Kimi K2 Instruct128,000GPT-4.11,047,576GPT-4.1 mini1,047,576GPT-4.1 nano1,047,576Kimi-VL-A3B-Instruct128,000Gemini 2.5 Pro1,048,576Qwen2.5-VL 72B Instruct128,000DeepSeek-R1128,000Grok 4.31,000,000
View exact operational data
Operational metric data
MetricModelValueDirection
LiveBench OverallClaude Sonnet 574.85 pointsHigher is better
LiveBench OverallClaude Opus 4.878.93 pointsHigher is better
LiveBench OverallGemini 3.5 Flash74.64 pointsHigher is better
LiveBench OverallGPT-5.579.91 pointsHigher is better
LiveBench OverallGPT-5.477.97 pointsHigher is better
Price per 1M input tokensGrok 4.62 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGemini 3.5 Flash-Lite0.3 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGemini 3.6 Flash1.5 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGrok 4.52 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-5.6 Luna0.2 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-5.6 Sol5 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-5.6 Terra2 USD / 1M input tokensLower is better for cost
Price per 1M input tokensClaude Sonnet 52 USD / 1M input tokensLower is better for cost
Price per 1M input tokensClaude Fable 510 USD / 1M input tokensLower is better for cost
Price per 1M input tokensClaude Opus 4.85 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-5.55 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGrok 4.201.25 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-5.42.5 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-4.12 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-4.1 mini0.4 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-4.1 nano0.1 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGrok 4.31.25 USD / 1M input tokensLower is better for cost
Price per 1M output tokensGrok 4.66 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGemini 3.5 Flash-Lite2.5 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGemini 3.6 Flash7.5 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGrok 4.56 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-5.6 Luna1.2 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-5.6 Sol30 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-5.6 Terra12 USD / 1M output tokensLower is better for cost
Price per 1M output tokensClaude Sonnet 510 USD / 1M output tokensLower is better for cost
Price per 1M output tokensClaude Fable 550 USD / 1M output tokensLower is better for cost
Price per 1M output tokensClaude Opus 4.825 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-5.530 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGrok 4.202.5 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-5.415 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-4.18 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-4.1 mini1.6 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-4.1 nano0.4 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGrok 4.32.5 USD / 1M output tokensLower is better for cost
Context windowGrok 4.6500,000 tokensHigher supports larger inputs
Context windowGemini 3.5 Flash-Lite1,048,576 tokensHigher supports larger inputs
Context windowGemini 3.6 Flash1,048,576 tokensHigher supports larger inputs
Context windowGrok 4.5500,000 tokensHigher supports larger inputs
Context windowGPT-5.6 Luna1,050,000 tokensHigher supports larger inputs
Context windowGPT-5.6 Sol1,050,000 tokensHigher supports larger inputs
Context windowGPT-5.6 Terra1,050,000 tokensHigher supports larger inputs
Context windowClaude Fable 51,000,000 tokensHigher supports larger inputs
Context windowClaude Opus 4.81,000,000 tokensHigher supports larger inputs
Context windowGemini 3.5 Flash1,048,576 tokensHigher supports larger inputs
Context windowGPT-5.51,050,000 tokensHigher supports larger inputs
Context windowGrok 4.201,000,000 tokensHigher supports larger inputs
Context windowGPT-5.41,050,000 tokensHigher supports larger inputs
Context windowKimi K2 Base128,000 tokensHigher supports larger inputs
Context windowKimi K2 Instruct128,000 tokensHigher supports larger inputs
Context windowGPT-4.11,047,576 tokensHigher supports larger inputs
Context windowGPT-4.1 mini1,047,576 tokensHigher supports larger inputs
Context windowGPT-4.1 nano1,047,576 tokensHigher supports larger inputs
Context windowKimi-VL-A3B-Instruct128,000 tokensHigher supports larger inputs
Context windowGemini 2.5 Pro1,048,576 tokensHigher supports larger inputs
Context windowQwen2.5-VL 72B Instruct128,000 tokensHigher supports larger inputs
Context windowDeepSeek-R1128,000 tokensHigher supports larger inputs
Context windowGrok 4.31,000,000 tokensHigher supports larger inputs

Latest compatible benchmark · Measured 2026-06-25

Benchmark observations by release

LiveBench Overall, snapshot 2026-06-25, measured 2026-06-25. Scores are shown from newest release to oldest. Eligible models (5): Claude Sonnet 5, Claude Opus 4.8, Gemini 3.5 Flash, GPT-5.5, GPT-5.4.

Legend: bars use the fixed range 0100 points; longer is better. Exact values remain visible.

Source snapshot: 2026-06-25, measured 2026-06-25.

LiveBench Overall observations by releaseHorizontal bars show 5 models from newest release to oldest using the 0 to 100 points scale. Values are printed beside every bar.0100Claude Sonnet 5, 74.85 pointsClaude Sonnet 574.85Claude Opus 4.8, 78.93 pointsClaude Opus 4.878.93Gemini 3.5 Flash, 74.64 pointsGemini 3.5 Flash74.64GPT-5.5, 79.91 pointsGPT-5.579.91GPT-5.4, 77.97 pointsGPT-5.477.97
View exact benchmark observation data
Benchmark observation data
ModelScoreMeasuredSource/version
Claude Sonnet 574.85 points2026-06-25
Claude Opus 4.878.93 points2026-06-25
Gemini 3.5 Flash74.64 points2026-06-25
GPT-5.579.91 points2026-06-25
GPT-5.477.97 points2026-06-25

Sources/versions:

Missing scores are omitted.

Value frontier

Takeaway: Claude Sonnet 5, GPT-5.5, GPT-5.4 are not dominated on both published 1M input tokens price and this direct benchmark score among the 4 eligible models. Higher benchmark scores are better.

Frontier modelOther eligible model

Source snapshot: 2026-06-25, measured 2026-06-25. Price unit: 1M input tokens.

The score axis is zoomed to 7382 points; it does not start at 0.

1M input tokens price versus LiveBench Overall scoreScatter plot of 4 models priced per 1M input tokens. Circles identify nondominated frontier models and diamonds identify other eligible models. Higher benchmark scores are better. The score axis covers 73 to 82 points. Every plotted value is also listed in the table below.8280787573$0.00$2.50$5.00Published price (USD per 1M input tokens)LiveBench Overall (points)Claude Sonnet 5: $2 / 1M input tokens · standard, 74.85 pointsClaude Sonnet 5$2 / 1M input tokens · standard · 74.85 pointsClaude Opus 4.8: $5 / 1M input tokens · standard, 78.93 pointsGPT-5.5: $5 / 1M input tokens · standard, 79.91 pointsGPT-5.5$5 / 1M input tokens · standard · 79.91 pointsGPT-5.4: $2.5 / 1M input tokens · standard, 77.97 pointsGPT-5.4$2.5 / 1M input tokens · standard · 77.97 points
View exact value frontier data
Value frontier data
ModelPrice (1M input tokens)ScoreStatusMeasuredSource/version
Claude Sonnet 5$2 / 1M input tokens · standard74.85 pointsFrontier2026-06-25
Claude Opus 4.8$5 / 1M input tokens · standard78.93 pointsEligible2026-06-25
GPT-5.5$5 / 1M input tokens · standard79.91 pointsFrontier2026-06-25
GPT-5.4$2.5 / 1M input tokens · standard77.97 pointsFrontier2026-06-25

Sources/versions:

Prices are provider list prices verified 2026-08-14.

Side-by-side comparison

Compare 2–4 models using directly sourced capability, cost, and access facts.

Select models from the active task. A shared benchmark score is shown only when an audited snapshot covers every selected model.

How to use this comparison

  1. 01Filter or search for the models that fit your workload.
  2. 02Pick them in the selector above, or Select Compare on a row in the inventory.
  3. 03Read the matrix, source evidence, and trade-offs before choosing a default.

0 models selected for comparison

Pick 2 more models above to build the matrix.

Methodology: models with missing scores or prices are omitted from the applicable visual. Prices are provider list prices at the dataset verification date. Benchmark results should not be compared across dataset versions.

Van Data Team analysis

Recommended by use case

Editorial shortlists for evaluation, not automatic winners. Dataset 2026-08-14, last verified 2026-08-14.

AI agents and tool use

GPT-5.6 Sol, Claude Sonnet 5, Gemini 3.6 Flash

Start a current agent-workflow evaluation with GPT-5.6 Sol, Claude Sonnet 5, and Gemini 3.6 Flash.

Review rationale and evidence

Why

  • All three active API records are positioned by their providers for agentic or tool-using workflows.
  • GPT-5.6 Sol and Gemini 3.6 Flash publish context windows above one million tokens, while Claude Sonnet 5 launches with a lower introductory input price than GPT-5.6 Sol.

Watch for

  • The pinned LiveBench snapshot predates GPT-5.6 Sol and Gemini 3.6 Flash and does not measure tool-use success, latency, uptime, or end-to-end agent reliability; Claude Sonnet 5 pricing changes after 2026-08-31.

Fits when: Test the models with the actual tools, permissions, retry rules, and human approval steps used in production.

Evidence

Apply this shortlist

Coding and repository work

GPT-5.6 Sol, Claude Fable 5

Test GPT-5.6 Sol and Claude Fable 5 as current coding-focused API options before choosing a model for repository work.

Review rationale and evidence

Why

  • Both providers position these active records for complex coding or software-engineering workflows.
  • Their published million-token-class context windows support evaluation on large repositories and long-running work.

Watch for

  • The catalog does not include a compatible repository-level coding benchmark for either model, and the pinned LiveBench snapshot predates both models.

Fits when: Evaluate repository navigation, code edits, test execution, and review quality on your own languages and tooling.

Evidence

Apply this shortlist

Complex reasoning and knowledge work

GPT-5.5, Claude Fable 5

Compare GPT-5.5 and Claude Fable 5 for current high-capability analysis, then test both against your review standard.

Review rationale and evidence

Why

  • Both providers position these active API records for demanding reasoning or professional knowledge work.
  • Both records publish context windows of at least 1,000,000 tokens.

Watch for

  • The pinned LiveBench snapshot does not provide a compatible Claude Fable 5 result and does not establish factuality, domain expertise, citation quality, latency, or production reliability.

Fits when: Use representative documents, expected citations, and a defined human review rubric for the production trial.

Evidence

Apply this shortlist

Multimodal workflows

GPT-5.6 Terra, Gemini 3.6 Flash, Grok 4.5

Compare current API options from OpenAI, Google, and SpaceXAI when the workload combines text with visual inputs.

Review rationale and evidence

Why

  • Gemini 3.6 Flash records text, image, audio, and video inputs; GPT-5.6 Terra and Grok 4.5 record text and image inputs.
  • The three current records span published context windows from 500,000 to 1,050,000 tokens and different input-price points.

Watch for

  • The dataset has no compatible multimodal quality benchmark for this shortlist and does not record modality-specific pricing.

Fits when: Build an evaluation set from the image, audio, or video formats and document layouts the system will receive.

Evidence

Apply this shortlist

High-volume, cost-sensitive processing

Gemini 3.5 Flash-Lite, GPT-4.1 nano

Compare current Gemini 3.5 Flash-Lite with GPT-4.1 nano for cost-sensitive API trials using published list prices.

Review rationale and evidence

Why

  • Gemini 3.5 Flash-Lite records $0.30 input and $2.50 output per million tokens, while GPT-4.1 nano records $0.10 input and $0.40 output.
  • Both active records accept multimodal input and publish context windows above one million tokens.

Watch for

  • List price does not include every production cost, and this dataset has no compatible quality score for either record.

Fits when: Estimate full task cost with representative input size, output size, retries, caching, review, and supporting infrastructure.

Evidence

Apply this shortlist
Methods and provenance

Evidence library

Open only the methodology, production guidance, release history, sources, or answers you need. Complete records remain available in the page HTML.

How to read benchmark data

How to read the benchmark data

Use the catalog to screen candidates, then verify them on your workload. Linked official sources support each record as a whole; they are not a field-by-field audit trail. Benchmark results retain their publisher and compatible snapshot.

Keep task families separate

Scores and operational metrics should not be compared across task families. Text reasoning, image quality, video quality, and speech error metrics answer different questions.

Match one benchmark snapshot

Compare results only when the task family, benchmark version, measurement date, and eligible model set match. Different snapshots may use different prompts, evaluators, or datasets.

Follow the score direction

Higher-is-better metrics reward a larger result; lower-is-better metrics such as error rate reward a smaller result. The direction belongs to the named benchmark, not to every task.

Match the pricing unit

A price per million tokens, image, audio minute, video second, or realtime minute is a different cost basis. Compare only the same unit and tier, then add retries, review, tooling, and infrastructure.

Check lifecycle and source status

Current models are active provider offerings; Legacy models remain for historical context. Official records come from providers, while independent results come from benchmark publishers, each with a visible verification date.

Keep missing data missing

Not reported is not zero. The explorer excludes missing values from rankings and does not infer them from a related model.

Scope begins 2024-01-01. Dataset version 2026-08-14. Last verified 2026-08-14. 91 source records.

Production fit

Why a benchmark winner may not fit production

Benchmark rank and production fit measure different things. A result can support a shortlist without deciding the system design.

  • A public score may cover one narrow task while your system combines retrieval, tools, permissions, and review.
  • Production acceptance also depends on latency, reliability, safety, data handling, regional availability, and total cost.
  • Van Data Team recommendations are transparent editorial shortlists. They use published fields and results, apply no hidden weighting, and must be tested on representative work.
Release timeline

Model release timeline since 2024

Generated from exact explorer release dates. Models without a verified exact date remain in the inventory and are omitted here rather than assigned an inferred date.

  1. 2026 Q3

    • Grok 4.62026-08-12
    • Gemini 3.5 Flash-Lite2026-07-21
    • Gemini 3.6 Flash2026-07-21
    • Grok 4.52026-07-16
    • GPT-5.6 Luna2026-07-09
    • GPT-5.6 Sol2026-07-09
    • GPT-5.6 Terra2026-07-09
  2. 2026 Q2

    • Claude Sonnet 52026-06-30
    • Claude Fable 52026-06-09
    • Claude Opus 4.82026-05-28
    • Gemini 3.5 Flash2026-05-19
    • GPT-5.52026-04-23
  3. 2026 Q1

    • Grok 4.202026-03-10
    • GPT-5.42026-03-05
    • Scribe v22026-01-09
  4. 2025 Q4

    • Scribe v2 Realtime2025-11-11
  5. 2025 Q3

    • Kimi K2 Base2025-07-11
    • Kimi K2 Instruct2025-07-11
  6. 2025 Q2

    • Kimi-Dev-72B2025-06-17
    • GPT-4.12025-04-14
    • GPT-4.1 mini2025-04-14
    • GPT-4.1 nano2025-04-14
    • Kimi-VL-A3B-Instruct2025-04-09
  7. 2025 Q1

    • Gemini 2.5 Pro2025-03-25
    • Claude 3.7 Sonnet2025-02-24
    • Grok 32025-02-19
    • Qwen2.5-VL 72B Instruct2025-01-26
    • DeepSeek-R12025-01-20
  8. 2024 Q4

    • DeepSeek-V32024-12-26
    • Gemini 2.0 Flash2024-12-11
    • Ministral 8B2024-10-16
  9. 2024 Q3

    • Llama 3.2 11B Vision2024-09-25
    • Qwen2.5 32B Instruct2024-09-19
    • Pixtral 12B2024-09-17
    • Grok 22024-08-13
    • Mistral Large 22024-07-24
    • Llama 3.1 405B2024-07-23
  10. 2024 Q2

    • DeepSeek-Coder-V2 Instruct2024-06-17
    • Qwen2 72B Instruct2024-06-07
    • Qwen2 7B Instruct2024-06-07
    • Codestral 22B2024-05-29
    • Gemini 1.5 Flash2024-05-14
    • GPT-4o2024-05-13
    • DeepSeek-V2 Chat2024-05-06
    • Llama 3 70B2024-04-18
    • Llama 3 8B2024-04-18
    • Grok-1.5V2024-04-12
  11. 2024 Q1

    • Grok-1.52024-03-28
    • Claude 3 Haiku2024-03-13
    • Claude 3 Opus2024-03-04
    • Claude 3 Sonnet2024-03-04
    • Gemini 1.5 Pro2024-02-15
Source catalog · 91 records

Source catalog

Official sources come from model providers. Independent sources come from a separate benchmark publisher. Each row shows its publisher, verification date, version, license status, and coverage; missing metadata is not filled or inferred.

Official provider sources

OpenAI
  • GPT-5.4 model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model documentation verifies the dated snapshot, token limits, modalities, and exact standard token prices represented here.

  • Introducing GPT-5.5
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official release announcement verifies model identity and the 2026-04-23 release date.

  • GPT-5.5 model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model documentation verifies token limits, modalities, and exact standard token prices represented here.

  • GPT-5.6: Frontier intelligence that scales with your ambition
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official release announcement verifies the Sol, Terra, and Luna identities and 2026-07-09 general-availability date.

  • OpenAI GPT-5.6 model catalog
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model catalog verifies the Sol, Terra, and Luna token limits, modalities, roles, and exact standard token prices represented here.

  • Hello GPT-4o
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Introducing GPT-4.1 in the API
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • OpenAI API pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    Official pricing page used only for the exact token prices represented in model records.

  • GPT-4o model
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official provider documentation used to verify exact model specifications.

  • GPT-4.1 model
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official provider documentation used to verify exact model specifications.

  • GPT-4.1 mini model
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official provider documentation used to verify exact model specifications.

  • GPT-4.1 nano model
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official provider documentation used to verify exact model specifications.

  • GPT Image 2 model
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • GPT-Realtime-1.5 model
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • gpt-audio-1.5 model
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • GPT-4o Transcribe model
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • TTS-1 model
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

Anthropic
  • Introducing Claude Opus 4.8
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official release documentation verifies model identity, release date, the published context window, and standard input/output pricing.

  • Introducing Claude Sonnet 5
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official release documentation verifies model identity, release date, API availability, and introductory pricing through 2026-08-31.

  • Claude Fable 5 and Claude Mythos 5
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official release announcement verifies Claude Fable 5 identity, 2026-06-09 release date, API access, standard token pricing, and general availability after redeployment.

  • Claude models overview
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official current-model documentation verifies Fable's text-and-image inputs, 1,000,000-token context window, 128,000-token output limit, and current API availability.

  • Introducing the next generation of Claude
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official provider documentation used to verify model identity and release details.

  • Claude 3.7 Sonnet and Claude Code
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Claude pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    Official pricing documentation used only for the exact token prices represented in model records.

  • Claude 3 Haiku
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify exact model specifications.

  • Claude models overview
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify exact model specifications.

Google
  • Gemini API release notes
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Official release notes and deprecation schedule verify the 2026-05-19 Gemini 3.5 Flash release date and current availability.

  • Gemini 3.5 Flash model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Official model documentation verifies model code, modalities, and token limits; exact token prices remain unreported here.

  • Gemini 3.6 Flash model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Official model documentation verifies model code, 2026-07-21 update, modalities, and token limits.

  • Using the latest Gemini models
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity, pricing

    Official latest-model guidance verifies the exact standard input/output prices represented for Gemini 3.6 Flash.

  • Gemini 3.5 Flash-Lite model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Official model documentation verifies the stable model code, multimodal inputs, token limits, capabilities, and July 2026 update.

  • Gemini 3.5 Flash-Lite pricing
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    pricing

    Official pricing table verifies the standard paid-tier input, cached-input, and output rates represented here.

  • Our next-generation model: Gemini 1.5
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Gemini 1.5 Flash
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Introducing Gemini 2.0
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Gemini 2.5: Our most intelligent AI model
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Gemini Developer API pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    Official pricing documentation consulted for price provenance; tiered values are not collapsed into model records.

  • Gemini API models
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify exact model specifications.

  • Gemini 2.5 Pro
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify exact model specifications.

  • Gemini 3.1 Flash Image pricing
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party pricing table used to verify Gemini 3.1 Flash Image standard token and resolution-specific image prices.

  • Gemini native image generation capabilities
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party native image guide used to verify Gemini 3.1 Flash Image generation, editing, image input, and mixed text/image output capabilities.

  • Gemini Developer API pricing and native media model identifiers
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

SpaceXAI
  • Grok 4.3 model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model documentation verifies Grok 4.3 modalities, context window, reasoning modes, aliases, and standard short-context prices.

  • xAI inference models API reference
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official API reference exposes Grok 4.3 modalities, aliases, and token pricing fields; its object creation timestamp is not treated as a public release date.

  • Skills in web, iOS, and Android
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official 2026-05-18 announcement confirms Grok 4.3 was live across grok.com, iOS, and Android by that date; it is not treated as the model's exact release date.

  • SpaceXAI API release notes
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official release notes verify Grok 4.20 and Grok 4.20 Multi-agent API availability on 2026-03-10.

  • Grok 4.20 System Card
    Verified
    2026-08-07
    Dataset version
    2026-04-07
    License
    Not stated
    Covers
    identity

    Official system card verifies Grok 4.20 identity, single-agent and multi-agent deployment modes, supported input modalities, and intended uses.

  • SpaceXAI model pricing
    Verified
    2026-08-14
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official current pricing table verifies active Grok 4.6, Grok 4.5, and Grok 4.20 variants, their documented context windows, and standard and long-context token rates.

  • Grok model retirement on May 15, 2026
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official lifecycle notice verifies Grok 3 retirement and redirect to Grok 4.3, and identifies Grok 4.3 as the recommended general replacement.

  • Introducing Grok 4.5
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official announcement verifies the model identity, release date, and intended coding, agentic, and knowledge-work scope.

  • Grok 4.5 model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model documentation verifies modalities, context window, and standard short-context token prices; the model caution discloses higher long-context rates.

  • Introducing Grok 4.6
    Verified
    2026-08-14
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official announcement verifies the 2026-08-12 release date and positioning for long-running agents, coding, and knowledge work.

  • Grok 4.6 model
    Verified
    2026-08-14
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model documentation verifies the exact model ID, text and image input, text output, 500,000-token context window, reasoning, function calling, structured outputs, and standard token prices.

  • grok-imagine-image model
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • Video generation
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • Imagine API pricing
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

Meta
  • Introducing Meta Llama 3
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Introducing Llama 3.1
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Llama 3.2
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

xAI
  • Announcing Grok-1.5
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Grok-1.5 Vision Preview
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Grok-2 Beta Release
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Grok 3 Beta
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • xAI models and pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model catalog used to verify documented model availability and pricing fields, not historical release dates.

DeepSeek
  • DeepSeek-V2
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • DeepSeek-Coder-V2
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • DeepSeek-V3
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • DeepSeek-R1
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • DeepSeek API pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    Official pricing documentation consulted for price provenance; unavailable exact values remain unreported.

Mistral AI
  • Codestral
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Large Enough
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Announcing Pixtral 12B
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Un Ministral, des Ministraux
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Mistral AI pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    Official pricing page consulted for price provenance; unavailable exact values remain unreported.

Qwen Team
  • Hello Qwen2
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Qwen2.5: A Party of Foundation Models
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Qwen2.5-VL
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

Moonshot AI
  • Kimi-VL
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Kimi-Dev
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Kimi K2
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

Runway
  • Available AI Models
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

ElevenLabs
  • ElevenLabs models
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • Introducing Scribe v2
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • Introducing Scribe v2 Realtime
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

Stability AI
  • Stability AI Developer Platform pricing
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

Black Forest Labs
  • FLUX image generation endpoints
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • FLUX API pricing
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

Independent benchmark sources

LiveBench
  • LiveBench category-weighted averaging implementation
    Verified
    2026-08-07
    Dataset version
    LiveBench-2026-06-25@19e766a5de4de07d672ed5bf9f0a69ceed1d39bf
    License
    Apache-2.0
    Covers
    benchmark

    Pinned implementation averages task scores within each of the seven categories, then averages the category scores; the paired categories_2026_06_25.json file defines category membership.

  • LiveBench 2026-06-25 task results
    Verified
    2026-08-07
    Dataset version
    LiveBench-2026-06-25@19e766a5de4de07d672ed5bf9f0a69ceed1d39bf
    License
    Apache-2.0
    Covers
    benchmark

    Snapshot and measurement date: 2026-06-25. Overall scores are imported only for unambiguous model mappings and are recomputed with the pinned category-weighted averaging implementation.

  • LiveBench datasheet and methodology
    Verified
    2026-08-06
    Dataset version
    LiveBench-2024-11-25
    License
    Apache-2.0
    Covers
    benchmark

    Primary benchmark documentation confirms objective scoring, public distribution, and Apache 2.0 reuse terms.

  • LiveBench 2024-11-25 task results
    Verified
    2026-08-06
    Dataset version
    LiveBench-2024-11-25@347e3d0b6a4916cd66f8bda31ce7ecdd7436ecc3
    License
    Apache-2.0
    Covers
    benchmark

    Snapshot and measurement date: 2024-11-25. Exact published snapshot; only direct Web of Lies V2 percentages with unambiguous model mappings are imported.

FAQ

AI model benchmark FAQ

What does an AI model benchmark score mean?

A benchmark score reports performance on one defined evaluation. Compare scores only when the task family, benchmark version, measurement date, scoring method, and eligible model set match; the result is a screening signal, not a prediction for every workload.

How fresh are the model and source records?

Each source shows a verification date, and the page shows the dataset version and last verified date. A verification date confirms when Van Data Team checked the linked record; it does not promise that a provider page has stayed unchanged since then.

Why are some model values marked Not reported?

Not reported means the catalog does not have a compatible, directly sourced value for that exact model record. Missing values remain empty rather than being estimated, copied from a related model, or treated as zero.

How should I compare AI model pricing?

Compare only the same pricing unit and tier, such as a million tokens, image, audio minute, video second, character, or realtime minute. Prices are provider list prices checked on the source verification date, not live quotes; retries, review, tooling, and infrastructure also affect production cost.

Does open-weight access mean a model is free to run?

No. Open weights means the model parameters are available for deployment under the provider license. This can provide infrastructure control, but compute, serving, monitoring, security, and engineering still carry costs.

How should I choose models for a production evaluation?

Start with the use case, select a small evidence-backed shortlist, and test it on representative work. Set acceptance thresholds for quality, latency, cost, reliability, safety, and human review before comparing results.

How often is this AI model comparison updated?

Van Data Team updates the versioned dataset when a review verifies material model, price, or benchmark changes. The published last verified date is the update record; this is a curated comparison, not a live provider feed.

How can Van Data Team help with model evaluation?

Van Data Team can turn a shortlist into a production evaluation with representative cases, measurable acceptance rules, model routing, workflow controls, observability, and cost tracking tailored to your system.

Production evaluation

Turn a model shortlist into a production decision

We design representative evaluations and the workflow controls needed to operate the selected model with clear quality, cost, and review boundaries.

  • Evaluation cases and acceptance thresholds tied to real work
  • Model routing, tool permissions, observability, and human review design