Celeris-1 vs K2 Horizon 3.7B
Wins 0 of 5 areas
—
Wins 5 of 5 areas
Coding · Agents · Reasoning · Facts · Long documents
K2 Horizon 3.7B is the stronger all-rounder.
Scores updated · 13 tests both models report · How we compare
Where each one wins
Tests won in each of the five areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- AgentsCarrying out multi-step tasks on its own02K2 Horizon 3.7B2 of 3 tests · 1 tie
- ReasoningHard problems that need careful thinking02K2 Horizon 3.7B2 of 3 tests · 1 tie
- CodingWriting and fixing software01K2 Horizon 3.7B1 of 2 tests · 1 tie
- FactsGetting facts right instead of making them up01K2 Horizon 3.7B1 of 2 tests · 1 tie
- Long documentsFinding answers in very long texts01K2 Horizon 3.7B1 of 1 test
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where K2 Horizon 3.7B pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+48.7points ahead
- Reasons across sets of long documentsAA-LCR+24.3points ahead
- Command-line tasks in a real terminalTerminal-Bench 2.1+16.9points ahead
Where Celeris-1 pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 13 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingK2 Horizon 3.7B
- Terminal-Bench 2.1K2 Horizon 3.7B by 16.911.228.1+16.9
- SciCodetie21.622tie
AgentsK2 Horizon 3.7B
- GDPValK2 Horizon 3.7B by 19.2019.2+19.2
- τ-Bench V3 · BankingK2 Horizon 3.7B by 15.73.919.6+15.7
- Terminal-Bench 4.0tie00tie
ReasoningK2 Horizon 3.7B
- Humanity's Last ExamK2 Horizon 3.7B by 7.16.813.9+7.1
- GPQA DiamondK2 Horizon 3.7B by 6.163.169.2+6.1
- CritPttie00tie
FactsK2 Horizon 3.7B
- AA-Omniscience · Non-hallucinationK2 Horizon 3.7B by 48.77.255.9+48.7
- AA-Omniscience · Accuracytie1110.8tie
Other results2 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AA-OmniscienceK2 Horizon 3.7B by 43-71.6-28.6+43
- AA IntelligenceK2 Horizon 3.7B by 9.36.315.6+9.3
Questions people ask
Which is better, Celeris-1 or K2 Horizon 3.7B?
K2 Horizon 3.7B wins all five areas where both have results: coding, agents, reasoning, facts and long documents. Celeris-1 wins none.
Which is better for coding?
K2 Horizon 3.7B. It wins 1 of the 2 coding tests both models report; Celeris-1 wins none, and 1 is a tie.
How do you compare the two?
We use the 13 benchmark tests both models have published scores on. The verdict counts the 11 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 2 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.