Highest score on LiveBench Overall
GPT-5.5
79.91 points on LiveBench Overall, snapshot 2026-06-25, measured 2026-06-25. Eligible models (4): GPT-5.4, GPT-5.5, Claude Opus 4.8, Claude Sonnet 5. Price unit: USD per 1M input tokens.
MethodologyCompare 68 models released since 2024 across 5 task families. Model records include linked official sources, while benchmark results retain their compatible evaluation snapshot.
Choose one task family, then filter its validated inventory. Charts and comparisons use compatible benchmark snapshots and pricing units; missing values remain visible as Not reported.
Compare models only within one compatible task family.
Facts from one compatible snapshot, measured 2026-06-25.
Highest score on LiveBench Overall
79.91 points on LiveBench Overall, snapshot 2026-06-25, measured 2026-06-25. Eligible models (4): GPT-5.4, GPT-5.5, Claude Opus 4.8, Claude Sonnet 5. Price unit: USD per 1M input tokens.
MethodologyCost and quality frontier
LiveBench Overall, snapshot 2026-06-25, measured 2026-06-25, compared with USD per 1M input tokens. Eligible models (4): GPT-5.4, GPT-5.5, Claude Opus 4.8, Claude Sonnet 5.
Methodology1–10 of 51 models
0 of 4 selected
| Compare | Model | Provider | Release | Access | Capabilities | Operational limits | Price / 1M input tokens | Price / 1M output tokens | LiveBench Web of Lies V2 | Details and provenance |
|---|---|---|---|---|---|---|---|---|---|---|
| Grok 4.6Current | ProviderxAI | Release2026-08-12 | AccessAPI | Capabilitiesreasoning | Operational limitsContext 500,000 · max output Not reported | Price / 1M input tokens$2 / 1M input tokens · standard | Price / 1M output tokens$6 / 1M output tokens · standard | LiveBench Web of Lies V2Not reported | View detailsStrengths, cautions, and provenance for Grok 4.6Strengths
| |
| Gemini 3.5 Flash-LiteCurrent | ProviderGoogle | Release2026-07-21 | AccessAPI | Capabilitiesreasoning | Operational limitsContext 1,048,576 · max output 65,536 | Price / 1M input tokens$0.3 / 1M input tokens · standard | Price / 1M output tokens$2.5 / 1M output tokens · standard | LiveBench Web of Lies V2Not reported | View detailsStrengths, cautions, and provenance for Gemini 3.5 Flash-LiteStrengths
| |
| Gemini 3.6 FlashCurrent | ProviderGoogle | Release2026-07-21 | AccessAPI | Capabilitiesreasoning | Operational limitsContext 1,048,576 · max output 65,536 | Price / 1M input tokens$1.5 / 1M input tokens · standard | Price / 1M output tokens$7.5 / 1M output tokens · standard | LiveBench Web of Lies V2Not reported | View detailsStrengths, cautions, and provenance for Gemini 3.6 FlashStrengths
| |
| Grok 4.5Current | ProviderxAI | Release2026-07-16 | AccessAPI | Capabilitiesreasoning | Operational limitsContext 500,000 · max output Not reported | Price / 1M input tokens$2 / 1M input tokens · standard | Price / 1M output tokens$6 / 1M output tokens · standard | LiveBench Web of Lies V2Not reported | View detailsStrengths, cautions, and provenance for Grok 4.5Strengths
| |
| GPT-5.6 LunaCurrent | ProviderOpenAI | Release2026-07-09 | AccessAPI | Capabilitiesreasoning | Operational limitsContext 1,050,000 · max output 128,000 | Price / 1M input tokens$0.2 / 1M input tokens · standard | Price / 1M output tokens$1.2 / 1M output tokens · standard | LiveBench Web of Lies V2Not reported | View detailsStrengths, cautions, and provenance for GPT-5.6 LunaStrengths
| |
| GPT-5.6 SolCurrent | ProviderOpenAI | Release2026-07-09 | AccessAPI | Capabilitiesreasoning | Operational limitsContext 1,050,000 · max output 128,000 | Price / 1M input tokens$5 / 1M input tokens · standard | Price / 1M output tokens$30 / 1M output tokens · standard | LiveBench Web of Lies V2Not reported | View detailsStrengths, cautions, and provenance for GPT-5.6 SolStrengths
| |
| GPT-5.6 TerraCurrent | ProviderOpenAI | Release2026-07-09 | AccessAPI | Capabilitiesreasoning | Operational limitsContext 1,050,000 · max output 128,000 | Price / 1M input tokens$2 / 1M input tokens · standard | Price / 1M output tokens$12 / 1M output tokens · standard | LiveBench Web of Lies V2Not reported | View detailsStrengths, cautions, and provenance for GPT-5.6 TerraStrengths
| |
| Claude Sonnet 5Current | ProviderAnthropic | Release2026-06-30 | AccessAPI | Capabilitiesreasoning | Operational limitsContext Not reported · max output Not reported | Price / 1M input tokens$2 / 1M input tokens · standard | Price / 1M output tokens$10 / 1M output tokens · standard | LiveBench Web of Lies V2Not reported | View detailsStrengths, cautions, and provenance for Claude Sonnet 5Strengths
| |
| Claude Fable 5Current | ProviderAnthropic | Release2026-06-09 | AccessAPI | Capabilitiesreasoning | Operational limitsContext 1,000,000 · max output 128,000 | Price / 1M input tokens$10 / 1M input tokens · standard | Price / 1M output tokens$50 / 1M output tokens · standard | LiveBench Web of Lies V2Not reported | View detailsStrengths, cautions, and provenance for Claude Fable 5Strengths
| |
| Claude Opus 4.8Current | ProviderAnthropic | Release2026-05-28 | AccessAPI | Capabilitiesreasoning | Operational limitsContext 1,000,000 · max output Not reported | Price / 1M input tokens$5 / 1M input tokens · standard | Price / 1M output tokens$25 / 1M output tokens · standard | LiveBench Web of Lies V2Not reported | View detailsStrengths, cautions, and provenance for Claude Opus 4.8Strengths
|
Compare declared coverage and compatible operational values for the current task.
Boolean provider-declared coverage for 10 current-task models. Cells are not scores and do not compare capability quality.
Legend: ✓ Supported · × Unsupported
| Model | Reasoning |
|---|---|
| Grok 4.6 | Supported |
| Gemini 3.5 Flash-Lite | Supported |
| Gemini 3.6 Flash | Supported |
| Grok 4.5 | Supported |
| GPT-5.6 Luna | Supported |
| GPT-5.6 Sol | Supported |
| GPT-5.6 Terra | Supported |
| Claude Sonnet 5 | Supported |
| Claude Fable 5 | Supported |
| Claude Opus 4.8 | Supported |
Each panel has its own unit, direction, and fixed displayed range. Values are never combined across panels.
points · Higher is better · Range 0–100
USD / 1M input tokens · Lower is better for cost · Range 0–10
USD / 1M output tokens · Lower is better for cost · Range 0–50
tokens · Higher supports larger inputs · Range 0–1,050,000
| Metric | Model | Value | Direction |
|---|---|---|---|
| LiveBench Overall | Claude Sonnet 5 | 74.85 points | Higher is better |
| LiveBench Overall | Claude Opus 4.8 | 78.93 points | Higher is better |
| LiveBench Overall | Gemini 3.5 Flash | 74.64 points | Higher is better |
| LiveBench Overall | GPT-5.5 | 79.91 points | Higher is better |
| LiveBench Overall | GPT-5.4 | 77.97 points | Higher is better |
| Price per 1M input tokens | Grok 4.6 | 2 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | Gemini 3.5 Flash-Lite | 0.3 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | Gemini 3.6 Flash | 1.5 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | Grok 4.5 | 2 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | GPT-5.6 Luna | 0.2 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | GPT-5.6 Sol | 5 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | GPT-5.6 Terra | 2 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | Claude Sonnet 5 | 2 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | Claude Fable 5 | 10 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | Claude Opus 4.8 | 5 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | GPT-5.5 | 5 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | Grok 4.20 | 1.25 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | GPT-5.4 | 2.5 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | GPT-4.1 | 2 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | GPT-4.1 mini | 0.4 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | GPT-4.1 nano | 0.1 USD / 1M input tokens | Lower is better for cost |
| Price per 1M input tokens | Grok 4.3 | 1.25 USD / 1M input tokens | Lower is better for cost |
| Price per 1M output tokens | Grok 4.6 | 6 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | Gemini 3.5 Flash-Lite | 2.5 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | Gemini 3.6 Flash | 7.5 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | Grok 4.5 | 6 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | GPT-5.6 Luna | 1.2 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | GPT-5.6 Sol | 30 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | GPT-5.6 Terra | 12 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | Claude Sonnet 5 | 10 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | Claude Fable 5 | 50 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | Claude Opus 4.8 | 25 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | GPT-5.5 | 30 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | Grok 4.20 | 2.5 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | GPT-5.4 | 15 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | GPT-4.1 | 8 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | GPT-4.1 mini | 1.6 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | GPT-4.1 nano | 0.4 USD / 1M output tokens | Lower is better for cost |
| Price per 1M output tokens | Grok 4.3 | 2.5 USD / 1M output tokens | Lower is better for cost |
| Context window | Grok 4.6 | 500,000 tokens | Higher supports larger inputs |
| Context window | Gemini 3.5 Flash-Lite | 1,048,576 tokens | Higher supports larger inputs |
| Context window | Gemini 3.6 Flash | 1,048,576 tokens | Higher supports larger inputs |
| Context window | Grok 4.5 | 500,000 tokens | Higher supports larger inputs |
| Context window | GPT-5.6 Luna | 1,050,000 tokens | Higher supports larger inputs |
| Context window | GPT-5.6 Sol | 1,050,000 tokens | Higher supports larger inputs |
| Context window | GPT-5.6 Terra | 1,050,000 tokens | Higher supports larger inputs |
| Context window | Claude Fable 5 | 1,000,000 tokens | Higher supports larger inputs |
| Context window | Claude Opus 4.8 | 1,000,000 tokens | Higher supports larger inputs |
| Context window | Gemini 3.5 Flash | 1,048,576 tokens | Higher supports larger inputs |
| Context window | GPT-5.5 | 1,050,000 tokens | Higher supports larger inputs |
| Context window | Grok 4.20 | 1,000,000 tokens | Higher supports larger inputs |
| Context window | GPT-5.4 | 1,050,000 tokens | Higher supports larger inputs |
| Context window | Kimi K2 Base | 128,000 tokens | Higher supports larger inputs |
| Context window | Kimi K2 Instruct | 128,000 tokens | Higher supports larger inputs |
| Context window | GPT-4.1 | 1,047,576 tokens | Higher supports larger inputs |
| Context window | GPT-4.1 mini | 1,047,576 tokens | Higher supports larger inputs |
| Context window | GPT-4.1 nano | 1,047,576 tokens | Higher supports larger inputs |
| Context window | Kimi-VL-A3B-Instruct | 128,000 tokens | Higher supports larger inputs |
| Context window | Gemini 2.5 Pro | 1,048,576 tokens | Higher supports larger inputs |
| Context window | Qwen2.5-VL 72B Instruct | 128,000 tokens | Higher supports larger inputs |
| Context window | DeepSeek-R1 | 128,000 tokens | Higher supports larger inputs |
| Context window | Grok 4.3 | 1,000,000 tokens | Higher supports larger inputs |
Latest compatible benchmark · Measured 2026-06-25
LiveBench Overall, snapshot 2026-06-25, measured 2026-06-25. Scores are shown from newest release to oldest. Eligible models (5): Claude Sonnet 5, Claude Opus 4.8, Gemini 3.5 Flash, GPT-5.5, GPT-5.4.
Legend: bars use the fixed range 0–100 points; longer is better. Exact values remain visible.
Source snapshot: 2026-06-25, measured 2026-06-25.
Sources/versions:
Missing scores are omitted.
Takeaway: Claude Sonnet 5, GPT-5.5, GPT-5.4 are not dominated on both published 1M input tokens price and this direct benchmark score among the 4 eligible models. Higher benchmark scores are better.
Source snapshot: 2026-06-25, measured 2026-06-25. Price unit: 1M input tokens.
The score axis is zoomed to 73–82 points; it does not start at 0.
Compare 2–4 models using directly sourced capability, cost, and access facts.
Select models from the active task. A shared benchmark score is shown only when an audited snapshot covers every selected model.
0 models selected for comparison
Methodology: models with missing scores or prices are omitted from the applicable visual. Prices are provider list prices at the dataset verification date. Benchmark results should not be compared across dataset versions.
Editorial shortlists for evaluation, not automatic winners. Dataset 2026-08-14, last verified 2026-08-14.
AI agents and tool use
Start a current agent-workflow evaluation with GPT-5.6 Sol, Claude Sonnet 5, and Gemini 3.6 Flash.
Fits when: Test the models with the actual tools, permissions, retry rules, and human approval steps used in production.
Coding and repository work
Test GPT-5.6 Sol and Claude Fable 5 as current coding-focused API options before choosing a model for repository work.
Fits when: Evaluate repository navigation, code edits, test execution, and review quality on your own languages and tooling.
Complex reasoning and knowledge work
Compare GPT-5.5 and Claude Fable 5 for current high-capability analysis, then test both against your review standard.
Fits when: Use representative documents, expected citations, and a defined human review rubric for the production trial.
Multimodal workflows
Compare current API options from OpenAI, Google, and SpaceXAI when the workload combines text with visual inputs.
Fits when: Build an evaluation set from the image, audio, or video formats and document layouts the system will receive.
High-volume, cost-sensitive processing
Compare current Gemini 3.5 Flash-Lite with GPT-4.1 nano for cost-sensitive API trials using published list prices.
Fits when: Estimate full task cost with representative input size, output size, retries, caching, review, and supporting infrastructure.
Open only the methodology, production guidance, release history, sources, or answers you need. Complete records remain available in the page HTML.
Use the catalog to screen candidates, then verify them on your workload. Linked official sources support each record as a whole; they are not a field-by-field audit trail. Benchmark results retain their publisher and compatible snapshot.
Scores and operational metrics should not be compared across task families. Text reasoning, image quality, video quality, and speech error metrics answer different questions.
Compare results only when the task family, benchmark version, measurement date, and eligible model set match. Different snapshots may use different prompts, evaluators, or datasets.
Higher-is-better metrics reward a larger result; lower-is-better metrics such as error rate reward a smaller result. The direction belongs to the named benchmark, not to every task.
A price per million tokens, image, audio minute, video second, or realtime minute is a different cost basis. Compare only the same unit and tier, then add retries, review, tooling, and infrastructure.
Current models are active provider offerings; Legacy models remain for historical context. Official records come from providers, while independent results come from benchmark publishers, each with a visible verification date.
Not reported is not zero. The explorer excludes missing values from rankings and does not infer them from a related model.
Scope begins 2024-01-01. Dataset version 2026-08-14. Last verified 2026-08-14. 91 source records.
Benchmark rank and production fit measure different things. A result can support a shortlist without deciding the system design.
Generated from exact explorer release dates. Models without a verified exact date remain in the inventory and are omitted here rather than assigned an inferred date.
Official sources come from model providers. Independent sources come from a separate benchmark publisher. Each row shows its publisher, verification date, version, license status, and coverage; missing metadata is not filled or inferred.
Official model documentation verifies the dated snapshot, token limits, modalities, and exact standard token prices represented here.
Official release announcement verifies model identity and the 2026-04-23 release date.
Official model documentation verifies token limits, modalities, and exact standard token prices represented here.
Official release announcement verifies the Sol, Terra, and Luna identities and 2026-07-09 general-availability date.
Official model catalog verifies the Sol, Terra, and Luna token limits, modalities, roles, and exact standard token prices represented here.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official pricing page used only for the exact token prices represented in model records.
Official provider documentation used to verify exact model specifications.
Official provider documentation used to verify exact model specifications.
Official provider documentation used to verify exact model specifications.
Official provider documentation used to verify exact model specifications.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
Official release documentation verifies model identity, release date, the published context window, and standard input/output pricing.
Official release documentation verifies model identity, release date, API availability, and introductory pricing through 2026-08-31.
Official release announcement verifies Claude Fable 5 identity, 2026-06-09 release date, API access, standard token pricing, and general availability after redeployment.
Official current-model documentation verifies Fable's text-and-image inputs, 1,000,000-token context window, 128,000-token output limit, and current API availability.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official pricing documentation used only for the exact token prices represented in model records.
Official provider documentation used to verify exact model specifications.
Official provider documentation used to verify exact model specifications.
Official release notes and deprecation schedule verify the 2026-05-19 Gemini 3.5 Flash release date and current availability.
Official model documentation verifies model code, modalities, and token limits; exact token prices remain unreported here.
Official model documentation verifies model code, 2026-07-21 update, modalities, and token limits.
Official latest-model guidance verifies the exact standard input/output prices represented for Gemini 3.6 Flash.
Official model documentation verifies the stable model code, multimodal inputs, token limits, capabilities, and July 2026 update.
Official pricing table verifies the standard paid-tier input, cached-input, and output rates represented here.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official pricing documentation consulted for price provenance; tiered values are not collapsed into model records.
Official provider documentation used to verify exact model specifications.
Official provider documentation used to verify exact model specifications.
First-party pricing table used to verify Gemini 3.1 Flash Image standard token and resolution-specific image prices.
First-party native image guide used to verify Gemini 3.1 Flash Image generation, editing, image input, and mixed text/image output capabilities.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
Official model documentation verifies Grok 4.3 modalities, context window, reasoning modes, aliases, and standard short-context prices.
Official API reference exposes Grok 4.3 modalities, aliases, and token pricing fields; its object creation timestamp is not treated as a public release date.
Official 2026-05-18 announcement confirms Grok 4.3 was live across grok.com, iOS, and Android by that date; it is not treated as the model's exact release date.
Official release notes verify Grok 4.20 and Grok 4.20 Multi-agent API availability on 2026-03-10.
Official system card verifies Grok 4.20 identity, single-agent and multi-agent deployment modes, supported input modalities, and intended uses.
Official current pricing table verifies active Grok 4.6, Grok 4.5, and Grok 4.20 variants, their documented context windows, and standard and long-context token rates.
Official lifecycle notice verifies Grok 3 retirement and redirect to Grok 4.3, and identifies Grok 4.3 as the recommended general replacement.
Official announcement verifies the model identity, release date, and intended coding, agentic, and knowledge-work scope.
Official model documentation verifies modalities, context window, and standard short-context token prices; the model caution discloses higher long-context rates.
Official announcement verifies the 2026-08-12 release date and positioning for long-running agents, coding, and knowledge work.
Official model documentation verifies the exact model ID, text and image input, text output, 500,000-token context window, reasoning, function calling, structured outputs, and standard token prices.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official model catalog used to verify documented model availability and pricing fields, not historical release dates.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official pricing documentation consulted for price provenance; unavailable exact values remain unreported.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official pricing page consulted for price provenance; unavailable exact values remain unreported.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
Official provider documentation used to verify model identity and release details.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.
Pinned implementation averages task scores within each of the seven categories, then averages the category scores; the paired categories_2026_06_25.json file defines category membership.
Snapshot and measurement date: 2026-06-25. Overall scores are imported only for unambiguous model mappings and are recomputed with the pinned category-weighted averaging implementation.
Primary benchmark documentation confirms objective scoring, public distribution, and Apache 2.0 reuse terms.
Snapshot and measurement date: 2024-11-25. Exact published snapshot; only direct Web of Lies V2 percentages with unambiguous model mappings are imported.
A benchmark score reports performance on one defined evaluation. Compare scores only when the task family, benchmark version, measurement date, scoring method, and eligible model set match; the result is a screening signal, not a prediction for every workload.
Each source shows a verification date, and the page shows the dataset version and last verified date. A verification date confirms when Van Data Team checked the linked record; it does not promise that a provider page has stayed unchanged since then.
Not reported means the catalog does not have a compatible, directly sourced value for that exact model record. Missing values remain empty rather than being estimated, copied from a related model, or treated as zero.
Compare only the same pricing unit and tier, such as a million tokens, image, audio minute, video second, character, or realtime minute. Prices are provider list prices checked on the source verification date, not live quotes; retries, review, tooling, and infrastructure also affect production cost.
No. Open weights means the model parameters are available for deployment under the provider license. This can provide infrastructure control, but compute, serving, monitoring, security, and engineering still carry costs.
Start with the use case, select a small evidence-backed shortlist, and test it on representative work. Set acceptance thresholds for quality, latency, cost, reliability, safety, and human review before comparing results.
Van Data Team updates the versioned dataset when a review verifies material model, price, or benchmark changes. The published last verified date is the update record; this is a curated comparison, not a live provider feed.
Van Data Team can turn a shortlist into a production evaluation with representative cases, measurable acceptance rules, model routing, workflow controls, observability, and cost tracking tailored to your system.
Production evaluation
We design representative evaluations and the workflow controls needed to operate the selected model with clear quality, cost, and review boundaries.