Mercury 2
Mercury 2 is behind the leaders in instruction following, reasoning, and coding. Too few results yet to rate agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Agentic, Safety, Long Context, Math, Multimodal, Multilingual or Factuality.
Price
$0.25input$0.75outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
29results on26benchmarks
- 5 independently verified
- 18 aggregator
- 6 vendor-reported
From 4 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
1 paper reference Mercury 2Mercury 2 benchmark results
29 results on 26 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
29.1% behind the leader1 of 3 ranked benchmarks measured
- IFBench69.80Oct 8, 2026aa_ifbench
Show 1 more instruction following resultHide 1 instruction following result
- 71.00Oct 7, 2026
37.8% behind the leader3 of 6 ranked benchmarks measured
- GPQA Diamond76.97Oct 8, 2026gpqa
- Humanity's Last Exam17.15Oct 8, 2026aa_hle
- 0.85Oct 8, 2026
Show 1 more reasoning resultHide 1 reasoning result
- GPQA Diamond74.00Oct 7, 2026GPQA
40.6% behind the leader4 of 10 ranked benchmarks measured
- SciCode37.73Oct 8, 2026aa_scicode
- Terminal-Bench Hard26.52Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.127.34Oct 8, 2026terminalbenchV21
- 1166.43May 22, 2026
Show 1 more coding resultHide 1 coding result
- 38.00Oct 7, 2026
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
- τ-Bench V3 · Banking9.48Oct 8, 2026tauBanking
0 of 3 ranked benchmarks measured
- 43.67Oct 8, 2026
0 of 4 ranked benchmarks measured
- 12.30May 2, 2026
- AA-Omniscience · Accuracy21.18Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination8.78Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- vectara_avg_summary_length149.10May 2, 2026Average Summary Length (Words)
- vectara_answer_rate100.00May 2, 2026Answer Rate
- vectara_factual_consistency87.70May 2, 2026Factual Consistency Rate
- -50.72Oct 8, 2026
- AA Intelligence13.77Oct 8, 2026aa_intelligence_index
- τ²-Bench Telecom (AA run)70.76Oct 8, 2026aa_tau2
Show 5 more resultsHide 5 results
- 4.02Sep 9, 2026
- 91.10Oct 7, 2026
- Artificial Analysis Coding Index31.11Sep 9, 2026aa_coding_index
- 67.00Aug 23, 2026
- 53.00Oct 7, 2026
Mercury 2: common questions
Who makes Mercury 2?
Mercury 2 is made by Inception.
When was Mercury 2 released?
Mercury 2 was released on Feb 20, 2026, according to Artificial Analysis.
What is Mercury 2 good at?
Mercury 2 is behind the leaders in instruction following, reasoning, and coding. Too few results yet to rate agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or factuality.
How much does Mercury 2 cost?
Mercury 2 costs $0.25 per million input tokens and $0.75 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it is cheaper than 66% of the 331 priced models we track.
How many benchmarks has Mercury 2 been tested on?
We track 29 results for Mercury 2 on 26 benchmarks from 4 sources, 5 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Mercury 2 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Mercury 2.
About this record
Where Mercury 2's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- May 2, 2026
Where the results come from
Verification: 29 scores · 5 independently verified · 18 aggregator-attributed · 6 vendor-reported. How these tiers are assigned
From 4 sources on 4 sites. Artificial Analysis supplies 18 of them; the 5 independently verified results come from 2 sites. Bars are coloured by trust tier.
- artificialanalysis.ai18
- api.llm-stats.com6
- raw.githubusercontent.com4
- datasets-server.huggingface.co1