Skip to content
ORIEL
Cat Board

Cat Board

A self-funded, independently run community benchmark — not official, not run by Oriel. The question set is private and rotates monthly, so this is one narrow view of long-term trends, not an authoritative or complete ranking. Don't take it as gospel.

Source: llm2014/llm_benchmark · Original site

This edition 2026-09
Peak × median — Peak / MedianPeakMedian
1GPT-6 Astra (xhigh)2.76%
2Fable 5.1 (xhigh ~Estimated)3.82%
3Opus 5.5 (xhigh ~Estimated)3.17%
4GPT-5.6 Sol (xhigh)13.36%
5GPT-6 Sol (xhigh)16.33%
6Opus 5 (xhigh)10.63%
7Seed-2.1-pro 0915 (high)16.89%
8GLM-5.3 (max)17.46%
9Opus 5.5 (xhigh w/ refuse)3.79%
10Kimi-K3 (max)10.52%
11Grok 4.7 (xhigh)11.62%
12DeepSeek V4.1 Flash (max)10.47%
13Gemini 3.8 Flash (high)11.37%
14Gemini 3.1 Pro (high)17.82%
15Muse Spark 1.3 (max)18.47%
5095Gap

Top 15 of 56 by peak score

ModelPeak scoreMedian scorePeak–median gapChangeAvg. time (s)TokensCost (CNY)Price (CNY/1M)ReleasedReasoning
GPT-6 Astra (xhigh)
91.2788.752.76%+25.1%24112198¥117.83¥345.0026-09-05On
Fable 5.1 (xhigh ~Estimated)
84.5681.333.82%-48141434¥400.25¥345.0026-09-02On
Opus 5.5 (xhigh ~Estimated)
82.9980.363.17%+26.3%29533096¥127.88¥138.0026-09-23On
GPT-5.6 Sol (xhigh)
78.9068.3613.36%-47821006¥81.17¥138.0026-07-10On
Opus 5.5 (xhigh w/ refuse)
69.4866.853.79%-29533096¥127.88¥138.0026-09-23On
Opus 5 (xhigh)
71.2063.6310.63%-50029868¥144.26¥172.5026-07-25On
Fable 5.1 (xhigh w/ refuse)
66.7063.474.84%-48141434¥400.25¥345.0026-09-02On
GPT-6 Sol (xhigh)
74.0361.9416.33%-22018451¥35.65¥69.0026-09-23On
Kimi-K3 (max)
69.1361.8610.52%+32.6%124138577¥113.42¥105.0026-07-16On
DeepSeek V4.1 Flash (max)
68.7461.5410.47%+18.9%58485949¥19.25¥8.0026-09-10On
Grok 4.7 (xhigh)
68.9960.9711.62%+21.9%142179365¥73.60¥33.1226-09-22On
Gemini 3.8 Flash (high)
68.0860.3411.37%+9.4%39539333¥28.50¥25.8826-09-03On
Seed-2.1-pro 0915 (high)
70.1258.2816.89%+44.8%120043884¥36.86¥30.0026-09-16On
GLM-5.3 (max)
69.6957.5217.46%+16.8%118764363¥50.46¥28.0026-08-19On
Gemini 3.1 Pro (high)
67.5055.4717.82%+5.8%23823891¥55.39¥82.8026-02-19On
GLM-5.3-Flash (max)
64.0655.1613.89%-102544240¥3.47¥2.8026-08-26On
Muse Spark 1.3 (max)
66.7254.4018.47%+38%43966822¥54.87¥29.3326-09-02On
DeepSeek V4 Pro 0813 (max)
64.7353.7017.04%+26.4%164360698¥45.89¥27.0026-08-13On
Qwen3.8-Max (xhigh)
59.8053.5510.45%-141970320¥70.88¥36.0026-08-03On
DeepSeek V4.1 Flash (high)
63.2652.0917.66%-64068561¥15.36¥8.0026-09-10On
Grok 4.6 (high)
64.5449.9922.54%+5.7%83436510¥42.32¥41.4026-08-13On
Step 5 Preview (high)
62.8049.8720.59%+413.8%122452901¥29.62¥20.0026-09-20On
Qwen3.8-Flash (xhigh)
59.0948.1418.53%-89762634¥4.74¥2.7026-08-26On
GPT-5.6 Luna (xhigh)
59.9344.2926.10%-30839417¥9.14¥8.2826-07-10On
Seed-2.1-pro (high)
50.7840.2320.78%+7.7%213969582¥58.45¥30.0026-06-23On
Gemini 3.7 Flash (low)
48.8339.7218.66%-579377¥6.79¥25.8826-08-13On
Qwen3.8-27B (xhigh)
50.5339.4821.87%-218071692¥24.09¥12.0026-08-14On
Muse Spark 1.2 (xhigh)
51.6439.3923.72%-34039800¥32.68¥29.3326-08-05On
Hy3 (high)
46.5237.6918.98%+23.2%111346845¥5.25¥4.0026-07-06On
MiniMax-M3
38.4233.1613.69%+33.3%103067827¥15.95¥8.4026-06-01On
Qwen3.7-Plus (high)
43.4131.5927.23%+24.9%122757472¥12.87¥8.0026-06-02On
Sonnet 5 (xhigh)
40.3431.5521.79%-64551921¥100.31¥69.0026-07-01On
GPT-6 Luna (xhigh)
37.7825.5532.37%-28126202¥2.53¥3.4526-09-23On
Seed-2.0-lite 0428 (high)
29.8022.5224.43%-63923332¥2.35¥3.6026-05-07On
Opus 5
30.5520.5932.60%+14%20113862¥66.95¥172.5026-07-25Standard
Gemini 3.5 Flash Lite (high)
26.7119.7626.02%+27.1%18516583¥8.01¥17.2526-07-21On
Spark X2.5
32.4517.7345.36%+317%64545239¥7.60¥6.0026-09-06On
Ling-3.0-flash
22.3913.4739.84%-72782850--26-07-24On
MiMo-V2.5-Pro
24.8412.5249.60%+18.9%63130724¥5.16¥6.0026-04-23On
Gemma 4 31B
15.3812.5118.66%-70614569¥1.13¥2.7626-04-03On
GPT-5.5 Instant
18.7112.1135.28%-291599¥9.27¥207.0026-05-06Standard
Step-3.7-Flash
23.3712.0548.44%-31745857¥10.40¥8.1026-05-29On
openPangu-2.0-Pro
15.1511.1926.14%-70729159¥11.84¥14.5026-07-31On
Qwen3.7-Plus
14.9410.8127.64%-3058554¥1.92¥8.0026-06-02Standard
Mistral Medium 3.5
13.1110.6119.07%-20626507¥38.41¥51.7526-04-30On
Dots3-Note Preview
15.8210.3234.77%-21823936--26-08-14On
DeepSeek V4.1 Flash
13.949.9628.55%+56.6%293838¥0.86¥8.0026-09-10Standard
openPangu-2.0-Flash
15.589.0541.91%-59327215¥1.22¥1.6026-06-30On
Qwen3.8-27B
12.238.9027.23%-1948565¥2.88¥12.0026-08-14Standard
ERNIE 5.1
13.077.6241.70%+2.4%47923797¥11.99¥18.0026-05-09On
Gemini 3.1 Flash Lite
8.717.2616.65%-11712¥0.21¥10.3526-03-03Standard
Sonnet 5
12.946.5749.23%+21%1247196¥13.90¥69.0026-07-01Standard
LongCat-2.0
13.586.5251.99%-37913041¥2.92¥8.0026-06-30On
Seed-2.1-pro 0915
11.506.2345.83%+41.2%381317¥1.11¥30.0026-09-16Standard
Seed-2.1-pro
13.354.7464.49%-3388751¥7.35¥30.0026-06-23Standard
Mistral Medium 3.5
6.064.4127.23%-172069¥3.00¥51.7526-04-30Standard

The question set rotates monthly; only the latest edition is shown. Scores are not comparable across months, so this site does not chart a trend for them.

This edition 2026-09
ModelSimple Model(H)iOS+Server(I)Animation(J)Data Process(K)Metal(L)Science(M)UnpromptedScaffoldReasoning
Opus 5.5 (max)
PassPass4/A+(341.72)Perfect(211.83)Pending2/A+(1365.55)5Claude CodeOn
GPT-6 Astra (max)
5/A+(41.53)3/A+(64.74)4/A+(75.84)5/A(61.32)8/B+(114.59)10/A(91.22)4Codex CLIOn
Fable 5 (high)
2/A+(90.52)3/A+(103.95)7/A(59.02)SkippedSkippedSkipped1Claude CodeOn
Opus 5 (max)
4/A(123.94)1/A+(122.57)12/B+(573.18)16/D(304.84)18/C(233.48)15/B(990.08)4Claude CodeOn
GPT-6 Sol (max)
PassPass16/B(27.74)12/C(90.31)Skipped14/B(42.82)1Codex CLIOn
GPT-5.6 Sol (max)
4/A(49.32)8/A(20.08)23/B(75.54)8/B(607.30)13/C(98.19)Skipped3Codex AppOn
Kimi-K3 (max)
6/A(29.59)5/A(43.22)22/B(43.23)28/D+(383.49)25/D(193.42)26/C+(267.49)2Claude CodeOn
Grok 4.7 (high)
8/B(53.87)3/A+(15.68)14/C+(32.91)12/C+(139.47)SkippedSkipped0Grok BuildOn
GLM-5.3 (max)
8/B+(11.48)7/A(25.25)20/C+(90.76)30/D(83.29)24/D(343.68)Failed0Claude CodeOn
Qwen3.8-Max-0902 (max)
8/B(28.26)7/A(109.76)21/C(102.71)13/D(211.18)Pending26/D(347.42)0Claude CodeOn
Gemini 3.7 Flash (high)
14/B(9.38)8/B(16.84)14/B(14.28)34/D(28.93)SkippedSkipped4OpenCode CLIOn
MiMo-V2.6-Pro (high)
16/C(10.11)5/B+(1.55)14/B(7.23)FailedSkippedFailed0Claude CodeOn
Grok 4.6 (high)
16/C+(17.67)7/B+(18.69)20/C+(23.19)SkippedSkippedSkipped1Grok BuildOn
Seed-2.1-Pro 0915 (high)
11/B(20.89)5/A(50.68)Failed30/D(71.77)PendingPending0Claude CodeOn
DeepSeek V4.1 Flash (max)
12/B(2.79)7/B(2.92)Failed20/C(5.88)PendingSkipped0Claude CodeOn
GLM-5.3-Flash (max)
9/B(4.21)13/B(6.63)FailedSkippedSkippedSkipped0Claude CodeOn
DeepSeek V4 Pro 0813 (max)
14/B(13.46)16/C(12.73)Failed16/B(18.64)27/D(69.58)Skipped0Claude CodeOn
Muse Spark 1.2 (xhigh)
26/C(15.82)11/B(17.76)FailedSkippedSkippedSkipped0Claude CodeOn
DeepSeek V4 Flash 0731 (max)
17/B(3.95)24/D+(2.72)FailedFailedSkippedSkipped2Claude CodeOn
Sonnet 5 (high)
22/C(154.97)16/C+(72.34)FailedSkippedSkippedSkipped1Claude CodeOn
Hy3 (high)
10/B(2.01)20/C+(13.29)FailedFailedSkippedSkipped0Claude CodeOn
GPT-5.6 Luna (max)
27/C(3.84)21/C(4.90)FailedSkippedSkippedSkipped0CodexOn
MiniMax-M3
30/D(2.42)17/C+(8.28)FailedSkippedSkippedSkipped1Claude CodeOn

The code methodology was revised to v3: no longer a single score, but per-project Pass / Skip / Pending / Failed, plus a "correction rounds / final grade" pair. The question set rotates monthly; only the latest edition is shown. Scores are not comparable across months, so this site does not chart a trend for them.

This edition 2025-11This category hasn't been updated in 10 months
Peak × median — Peak / MedianPeakMedian
1Gemini 3 Pro4.69%
2Gemini 2.5 Pro3.58%
3GPT-5 (high)12.86%
4Seed-1.6-vision (Think)7.82%
5qwen3-vl-235b-a22b (Think)8.10%
6Seed-1.6-thinking-07158.15%
7GLM-4.5V7.22%
8Seed-1.6-vision10.53%
9Step-37.19%
10Qwen3-vl-235b-a22b4.72%
11Sonnet 4 (Think)18.37%
12Sonnet 49.48%
13ERNIE 4.5 Turbo VL 082911.74%
14Qwen-vl-max-2025-08-139.83%
15hunyuan-t1-vision 061912.02%
3080Gap

Top 15 of 20 by peak score

ModelPeak scoreMedian scorePeak–median gapPrice (CNY/1M)Avg. tokensCost (CNY)Avg. time (s)ReleasedReasoning
Gemini 3 Pro
73.9470.474.69%¥86.410039¥16.4813825-11-19On
Gemini 2.5 Pro
64.0961.803.58%¥72.07277¥9.957925-05-06On
GPT-5 (high)
61.7353.7912.86%¥72.06271¥8.5818225-08-07On
Seed-1.6-vision (Think)
60.5855.857.82%¥8.06695¥1.028325-08-15On
qwen3-vl-235b-a22b (Think)
58.2153.508.10%¥20.05031¥1.9110125-09-24Standard
Seed-1.6-thinking-0715
55.3950.878.15%¥8.06104¥0.9311825-07-14On
GLM-4.5V
53.2649.427.22%¥6.02814¥0.325125-08-11On
Seed-1.6-vision
48.8943.7410.53%¥8.02263¥0.342925-08-15Standard
Step-3
48.8545.347.19%¥8.06033¥0.9229425-07-31On
Qwen3-vl-235b-a22b
47.8845.624.72%¥8.01590¥0.243425-09-24Standard
Sonnet 4 (Think)
47.5738.8318.37%¥108.02895¥5.944225-05-23On
Sonnet 4
45.5341.219.48%¥108.0931¥1.911725-05-23Standard
ERNIE 4.5 Turbo VL 0829
45.3139.9911.74%¥9.03433¥0.5924525-08-29On
Qwen-vl-max-2025-08-13
44.3640.009.83%¥4.02520¥0.1911525-08-13Standard
hunyuan-t1-vision 0619
42.6837.5512.02%¥9.06009¥1.0311525-06-19On
SenseNova-V6-Reasoner
40.7137.388.18%¥16.02316¥0.709025-04-10On
Seed-1.6
39.3035.3210.15%¥8.01006¥0.152325-06-11Standard
hunyuan-turbos-vision 0619
38.0534.549.23%¥9.01188¥0.203125-06-19Standard
SenseNova-V6-Pro
34.4126.6822.47%¥9.0689¥0.121825-04-10Standard
MiniMax-Text-01
25.1723.387.10%¥8.0587¥0.092425-01-13Standard

The question set rotates monthly; only the latest edition is shown. Scores are not comparable across months, so this site does not chart a trend for them.