Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-6 Sol

Basis

GPT-6 Sol is strong in long context and reasoning; capable in factuality, math, and multimodal tasks; and behind the leaders in agentic tasks, coding, and instruction following. Too few results yet to rate safety or multilingual tasks.

Price per million tokens

$2.00input$10.00output

Price from Artificial Analysis · 3 providers tracked · All prices

Costs more than 83% of 323 priced models · 3:1 input-to-output blend, log scale
Evidence

126results on51benchmarks

  • 37 independently verified
  • 66 aggregator
  • 3 vendor-reported
  • 20 cross-referenced

From 12 sources · latest Sep 29, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

Capability profile

Bars show the model's median result as a share of the leading model's, per capability.

Strong2 capabilities
  1. Long Context−5.7%1 of 3
  2. Reasoning−9.5%5 of 6
Capable3 capabilities
  1. Factuality−19.3%2 of 4
  2. Math−22.3%3 of 5
  3. Multimodal−22.6%1 of 6
Limited3 capabilities
  1. Agentic−26.9%2 of 7
  2. Coding−28.1%4 of 10
  3. Instruction Following−41.3%1 of 3

Too few results to rate: SafetyMultilingual

GPT-6 Sol benchmark results

126 results on 51 benchmarks, grouped by capability. Every score links to its source; a bar is the result as a share of the capability leader's.

Long Context

Strong1 of 3 ranked benchmarks measuredFull long context ranking
  • AA-LCR
    aa_lcr
    83.6794%
    High · Max
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
4 more long context results
  • AA-LCR
    aa_lcr
    81.33—
    xHigh
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • AA-LCR
    aa_lcr
    64.00—
    No reasoning
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • AA-LCR
    aa_lcr
    82.33—
    Medium
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • AA-LCR
    aa_lcr
    79.33—
    Low
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗

Reasoning

Strong5 of 6 ranked benchmarks measuredFull reasoning ranking
27 more reasoning results

Factuality

Capable2 of 4 ranked benchmarks measuredFull factuality ranking
11 more factuality results

Math

Capable3 of 5 ranked benchmarks measuredFull math ranking

Multimodal

Capable1 of 6 ranked benchmarks measuredFull multimodal ranking
  • MMMU-Pro
    aa_mmmu_pro
    83.2995%
    Max
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
5 more multimodal results
  • MMMU-Pro
    aa_mmmu_pro
    68.15—
    No reasoning
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • MMMU-Pro
    aa_mmmu_pro
    82.43—
    xHigh
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • MMMU-Pro
    aa_mmmu_pro
    80.64—
    Medium
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • MMMU-Pro
    aa_mmmu_pro
    81.21—
    High
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • MMMU-Pro
    aa_mmmu_pro
    78.84—
    Low
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗

Agentic

Limited2 of 7 ranked benchmarks measuredFull agentic ranking
11 more agentic results

Coding

Limited4 of 10 ranked benchmarks measuredFull coding ranking
5 more coding results
  • SciCode
    aa_scicode
    55.09—
    xHigh
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • SciCode
    aa_scicode
    47.34—
    No reasoning
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • SciCode
    aa_scicode
    53.82—
    Medium
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • SciCode
    aa_scicode
    54.86—
    High
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • SciCode
    aa_scicode
    50.23—
    Low
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗

Instruction Following

Limited1 of 3 ranked benchmarks measuredFull instruction following ranking

More results

40 more results
  • AA Intelligence
    aa_intelligence_index
    28.09—
    No reasoning
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • AA Intelligence
    aa_intelligence_index
    44.10—
    xHigh
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • AA Intelligence
    aa_intelligence_index
    47.53—
    Max
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • AA Intelligence
    aa_intelligence_index
    42.82—
    High
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • AA Intelligence
    aa_intelligence_index
    39.78—
    Medium
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • AA Intelligence
    aa_intelligence_index
    33.90—
    Low
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • AA-Omniscience
    aa_omniscience
    26.67—
    xHigh
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • AA-Omniscience
    aa_omniscience
    -0.85—
    No reasoning
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • AA-Omniscience
    aa_omniscience
    27.12—
    Max
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • AA-Omniscience
    aa_omniscience
    26.82—
    High
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • AA-Omniscience
    aa_omniscience
    27.02—
    Medium
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • AA-Omniscience
    aa_omniscience
    26.52—
    Low
    artificialanalysis.ai
    AggregatorSep 29, 2026view ↗
  • Agentic Safe Completions - Chat prod - Chat Plugins
    0.92—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • 56.40—
    api.llm-stats.comscores
    VendorSep 29, 2026view ↗
  • 92.67—
    xHigh
    arcprize.orgdata
    VerifiedSep 29, 2026view ↗
  • 95.50—
    Max
    arcprize.orgdata
    VerifiedSep 29, 2026view ↗
  • 68.80—
    api.llm-stats.comscores
    VendorSep 29, 2026view ↗
  • Dynamic Benchmarks - Emotional reliance
    0.97—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • Dynamic Benchmarks - Mental health
    1.00—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • Dynamic Benchmarks - Self-harm
    0.97—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • GDPval-AA 2.1
    GDPval-AA v2.1
    1487.00—
    www-cdn.anthropic.comClaude Sonnet 5.5 System Card
    Cross-refSep 28, 2026view ↗
  • 96.20—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • 30.10—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • HealthBench length-adjusted
    HealthBench length-adjusted
    53.20—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • 60.80—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • Image input evaluations - extremism
    0.97—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • Image input evaluations - harms-erotic
    1.00—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • Image input evaluations - hate
    1.00—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • Image input evaluations - self-harm
    0.98—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • 64.40—
    api.llm-stats.comscores
    VendorSep 29, 2026view ↗
  • Production Benchmarks - Gore
    0.90—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • U18 evaluations - Age-restricted goods, services, and dangerous challenges / activities
    0.86—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • U18 evaluations - Eating Disorders
    0.85—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • U18 evaluations - Emotional Reliance
    0.95—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • U18 evaluations - Gore
    0.90—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • U18 evaluations - Self Harm
    0.99—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • U18 evaluations - Sexual Content
    0.97—
    deploymentsafety.openai.comgpt-6-1-sol
    Cross-refSep 29, 2026view ↗
  • 100.00—
    raw.githubusercontent.comREADME
    VerifiedSep 23, 2026view ↗
  • vectara_avg_summary_length
    Average Summary Length (Words)
    71.40—
    raw.githubusercontent.comREADME
    VerifiedSep 23, 2026view ↗
  • vectara_factual_consistency
    Factual Consistency Rate
    93.50—
    raw.githubusercontent.comREADME
    VerifiedSep 23, 2026view ↗

About this record

Verification: 126 scores · 37 independently verified · 66 aggregator-attributed · 20 vendor cross-reference · 3 vendor-reported. How these tiers are assigned

Also listed as
gpt-6-sol-lowgpt-6-sol-highgpt-6-sol-mediumgpt-6-sol-non-reasoninggpt-6-sol-xhighgpt-6-sol:batchgpt-6 sol (low)gpt-6 sol (high)gpt-6 sol (medium)gpt-6 sol (max)gpt-6 sol (non-reasoning)gpt-6 sol (xhigh)gpt-6-sol-maxgpt sol 6gpt-6 sol (none)
Tracked since
Sep 16, 2026
Newest source mention
Sep 29, 2026