GPT-6 Sol
GPT-6 Sol is strong in long context and reasoning; capable in factuality, math, and multimodal tasks; and behind the leaders in agentic tasks, coding, and instruction following. Too few results yet to rate safety or multilingual tasks.
- Price per million tokens
$2.00input$10.00output
Price from Artificial Analysis · 3 providers tracked · All prices
Costs more than 83% of 323 priced models · 3:1 input-to-output blend, log scale - Evidence
126results on51benchmarks
- 37 independently verified
- 66 aggregator
- 3 vendor-reported
- 20 cross-referenced
From 12 sources · latest Sep 29, 2026 · How verification works
- API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Capability profile
Bars show the model's median result as a share of the leading model's, per capability.
Too few results to rate: SafetyMultilingual
GPT-6 Sol benchmark results
126 results on 51 benchmarks, grouped by capability. Every score links to its source; a bar is the result as a share of the capability leader's.
Long Context
Strong1 of 3 ranked benchmarks measuredFull long context ranking- AA-LCR83.6794%aa_lcr
Reasoning
Strong5 of 6 ranked benchmarks measuredFull reasoning ranking- LiveBench · Reasoning88.6596%livebench_reasoning@2026-06-25
- 30.8696%
- 89.5894%
- 73.1083%
- Humanity's Last Exam47.9174%aa_hle
27 more reasoning results
- 1.67—
- 31.53—
- 57.78—
- 68.89—
- 78.06—
- 1.21—
- 1.76—
- 4.93—
- 5.86—
- 9.45—
- 23.03—
- 0.28—
- 0.13—
- 0.42—
- 0.90—
- 1.80—
- 4.62—
- 28.00—
- 4.00—
- 24.57—
- 25.43—
- 16.29—
- Humanity's Last Exam46.29—aa_hle
- Humanity's Last Exam18.40—aa_hle
- Humanity's Last Exam40.96—aa_hle
- Humanity's Last Exam44.11—aa_hle
- Humanity's Last Exam34.94—aa_hle
Factuality
Capable2 of 4 ranked benchmarks measuredFull factuality ranking- AA-Omniscience · Accuracy54.4881%omniscienceAccuracy
- AA-Omniscience · Non-hallucination39.8845%omniscienceNonHallucination
11 more factuality results
- AA-Omniscience · Accuracy45.18—omniscienceAccuracy
- AA-Omniscience · Accuracy53.85—omniscienceAccuracy
- AA-Omniscience · Accuracy53.72—omniscienceAccuracy
- AA-Omniscience · Accuracy53.45—omniscienceAccuracy
- AA-Omniscience · Accuracy51.25—omniscienceAccuracy
- AA-Omniscience · Non-hallucination16.02—omniscienceNonHallucination
- AA-Omniscience · Non-hallucination41.10—omniscienceNonHallucination
- AA-Omniscience · Non-hallucination43.22—omniscienceNonHallucination
- AA-Omniscience · Non-hallucination41.88—omniscienceNonHallucination
- AA-Omniscience · Non-hallucination49.26—omniscienceNonHallucination
- 6.50—
Math
Capable3 of 5 ranked benchmarks measuredFull math ranking- LiveBench · Mathematics96.3699%livebench_math@2026-06-25
Multimodal
Capable1 of 6 ranked benchmarks measuredFull multimodal ranking- MMMU-Pro83.2995%aa_mmmu_pro
Agentic
Limited2 of 7 ranked benchmarks measuredFull agentic ranking- Terminal-Bench 4.043.9469%
- 49.3556%
11 more agentic results
- 49.44—
- 46.83—
- 36.27—
- 41.01—
- 43.80—
- 33.78—
- Terminal-Bench 4.013.13—
- Terminal-Bench 4.030.30—
- Terminal-Bench 4.018.69—
- Terminal-Bench 4.026.26—
- Terminal-Bench 4.09.09—
Coding
Limited4 of 10 ranked benchmarks measuredFull coding ranking- 1680.9797%
- LiveBench · Coding81.7789%livebench_coding@2026-06-25
- SciCode57.6486%aa_scicode
- LiveBench · Agentic Coding52.8868%livebench_agentic_coding@2026-06-25
Instruction Following
Limited1 of 3 ranked benchmarks measuredFull instruction following ranking- LiveBench · Instruction Following68.5784%livebench_instruction_following@2026-06-25
More results
- livebench_language85.30—livebench_language@2026-06-25
- livebench_data_analysis81.19—livebench_data_analysis@2026-06-25
- 29.33—
- 72.17—
- 83.67—
- 91.00—
40 more results
- AA Intelligence28.09—aa_intelligence_index
- AA Intelligence44.10—aa_intelligence_index
- AA Intelligence47.53—aa_intelligence_index
- AA Intelligence42.82—aa_intelligence_index
- AA Intelligence39.78—aa_intelligence_index
- AA Intelligence33.90—aa_intelligence_index
- AA-Omniscience26.67—aa_omniscience
- AA-Omniscience-0.85—aa_omniscience
- AA-Omniscience27.12—aa_omniscience
- AA-Omniscience26.82—aa_omniscience
- AA-Omniscience27.02—aa_omniscience
- AA-Omniscience26.52—aa_omniscience
- Agentic Safe Completions - Chat prod - Chat Plugins0.92—
- 56.40—
- 92.67—
- 95.50—
- 68.80—
- Dynamic Benchmarks - Emotional reliance0.97—
- Dynamic Benchmarks - Mental health1.00—
- Dynamic Benchmarks - Self-harm0.97—
- GDPval-AA 2.11487.00—GDPval-AA v2.1
- 96.20—
- 30.10—
- HealthBench length-adjusted53.20—HealthBench length-adjusted
- 60.80—
- Image input evaluations - extremism0.97—
- Image input evaluations - harms-erotic1.00—
- Image input evaluations - hate1.00—
- Image input evaluations - self-harm0.98—
- 64.40—
- Production Benchmarks - Gore0.90—
- U18 evaluations - Age-restricted goods, services, and dangerous challenges / activities0.86—
- U18 evaluations - Eating Disorders0.85—
- U18 evaluations - Emotional Reliance0.95—
- U18 evaluations - Gore0.90—
- U18 evaluations - Self Harm0.99—
- U18 evaluations - Sexual Content0.97—
- vectara_answer_rate100.00—Answer Rate
- vectara_avg_summary_length71.40—Average Summary Length (Words)
- vectara_factual_consistency93.50—Factual Consistency Rate
About this record
Verification: 126 scores · 37 independently verified · 66 aggregator-attributed · 20 vendor cross-reference · 3 vendor-reported. How these tiers are assigned
- Also listed as
- gpt-6-sol-lowgpt-6-sol-highgpt-6-sol-mediumgpt-6-sol-non-reasoninggpt-6-sol-xhighgpt-6-sol:batchgpt-6 sol (low)gpt-6 sol (high)gpt-6 sol (medium)gpt-6 sol (max)gpt-6 sol (non-reasoning)gpt-6 sol (xhigh)gpt-6-sol-maxgpt sol 6gpt-6 sol (none)
- Tracked since
- Sep 16, 2026
- Newest source mention
- Sep 29, 2026