大模型排行榜

更新时间:2026-08-04

模型 总分 专业能力 高难度提示词 编程 数学 创意写作 指令遵循 长文本
claude-fable-5 1 2 1 1 3 1 1 2
claude-opus-4-6-thinking 2 3 2 3 5 2 2 1
claude-opus-4-7-thinking 3 6 4 2 7 4 3 4
claude-opus-4-6 4 4 3 4 6 8 4 3
qwen3.8-max 5 9 6 9 3 9 5
claude-opus-4-7 6 5 7 5 13 6 7 8
claude-opus-5-high 7 1 5 7 2 11 5 7
claude-opus-5-max 8 8 9 17 1 9 8 12
muse-spark-1.1 9 23 10 10 16 32 17 31
muse-spark 10 41 14 15 44 16 31 38
gemini-3-pro 11 26 15 26 21 5 20 15
gemini-3.1-pro-preview 12 16 11 21 12 7 12 10
claude-opus-4-8-thinking 13 7 8 6 11 14 6 6
gpt-5.6-sol-xhigh 14 10 13 12 18 10 11 17
gemini-3.6-flash 15 21 17 20 4 17 16 20
gpt-5.5-high 16 12 18 22 14 24 13 18
gpt-5.4-high 17 13 20 19 10 41 18 24
gpt-5.5 18 18 24 42 9 31 21 28
gpt-5.2-chat-latest-20260210 19 35 22 29 51 49 38 41
gemini-3.5-flash-high 20 54 23 38 8 13 23 22
qwen3.7-max-preview 21 17 25 16 15 33 25 11
claude-opus-4-8 22 14 12 11 31 18 14 9
grok-4.20-beta1 23 63 40 44 54 19 47 53
gemini-3.5-flash-medium 24 40 28 46 26 15 30 34
gpt-5.5-instant 25 53 33 31 43 22 39 37
gemini-3-flash 26 32 29 43 25 21 37 36
claude-opus-4-5-20251101-thinking-32k 27 22 19 8 36 12 10 13
grok-4.20-beta-0309-reasoning 28 60 35 39 39 34 53 49
claude-sonnet-4-6 29 19 16 13 42 25 15 14
grok-4.20-multi-agent-beta-0309 30 51 42 41 52 26 55 58
glm-5.2-max 31 43 36 48 29 39 27 32
claude-opus-4-5-20251101 32 24 21 18 41 20 19 16
grok-4.5 33 27 32 24 17 29 32 26
glm-5.1 34 30 31 27 19 27 26 25
gpt-5.6-terra-xhigh 35 11 34 23 24 66 29 44
ernie-5.1 36 38 38 33 22 58 43 51
grok-4.1-thinking 37 70 50 58 62 56 73 74
mimo-v2.5-pro 38 20 26 28 28 51 24 21
gpt-5.4 39 36 37 32 45 48 35 33
qwen3.5-max-preview 40 29 30 35 37 30 22 30
claude-sonnet-5-high 41 15 27 14 33 55 34 29
kimi-k2.6 42 25 41 30 20 53 42 35
qwen3.6-max-preview 43 31 44 45 30 45 46 39
grok-4.1 44 86 57 66 85 50 72 68
qwen3.7-plus 45 42 48 40 35 44 45 43
gemini-3-flash (thinking-minimal) 46 73 53 67 50 35 56 55
deepseek-v4-pro 47 48 45 52 60 37 41 40
glm-5 48 46 49 61 63 36 52 45
gemini-3.5-flash-lite 49 82 54 47 98 47 61 54
hy3 50 33 55 56 60 51 46
dola-seed-2.0-pro 51 57 46 37 55 87 66 67
claude-sonnet-4-5-20250929-thinking-32k 52 28 39 25 49 28 28 19
claude-sonnet-4-5-20250929 53 44 43 34 86 23 33 27
gpt-5.1-high 54 50 58 70 47 54 48 56
deepseek-v4-pro-high-preview 55 65 59 75 40 43 50 52
gemma-4-31b 56 47 60 59 32 64 44 48
gpt-5.6-luna-xhigh 57 37 56 54 23 67 49 65
kimi-k2.5-thinking 58 45 61 55 34 61 62 57
claude-opus-4-1-20250805-thinking-16k 59 49 47 36 61 38 36 23
ernie-5.0-preview-1203 60 92 67 97 119 65 88 95
gpt-5.3-chat-latest 61 67 62 64 90 78 68 60
gpt-5.4-mini-high 62 52 63 63 68 84 65 71
mimo-v2-pro 63 34 52 51 53 59 54 47
claude-opus-4-1-20250805 64 72 51 49 77 42 40 42
ernie-5.0-0110 65 102 66 71 72 63 80 91
gemini-2.5-pro 66 77 79 103 66 40 60 59
minimax-m3 67 59 65 60 69 79 59 61
gpt-4.5-preview-2025-02-27 68 120 105 116 118 46 63 82
qwen3.6-plus 69 56 64 65 48 75 64 62
chatgpt-4o-latest-20250326 70 119 84 99 128 62 75 90
grok-4.3 71 89 82 74 101 57 102 77
glm-4.7 72 101 69 79 83 81 76 66
qwen3.5-397b-a17b 73 55 68 69 57 77 71 63
inkling 74 61 71 62 27 114 86 89
gpt-5.1 75 84 85 91 95 71 77 79
gemma-4-26b-a4b 76 62 70 83 38 82 58 69
deepseek-v4-flash-high-preview 77 68 80 84 64 74 69 72
gpt-5.2-high 78 58 77 72 46 111 78 99
deepseek-v4-flash 79 78 76 80 89 73 74 70
gpt-5.2 80 76 74 81 78 105 82 86
longcat-flash-chat-2602-exp 81 74 75 53 75 107 99 94
qwen3-max-preview 82 69 78 82 71 99 79 81
gpt-5-high 83 80 96 101 76 128 107 127
mimo-v2.5 84 64 72 68 65 97 70 64
glm-5v-turbo 85 81 87 73 56 90 83 83
gemini-3.1-flash-lite-preview 86 103 100 119 70 70 110 106
kimi-k2.5-instant 87 87 73 50 67 108 67 78
grok-4-1-fast-reasoning 88 105 104 110 103 69 117 113
o3-2025-04-16 89 100 106 114 58 116 118 134
mimo-v2-omni 90 83 81 76 74 95 81 73
kimi-k2-thinking-turbo 91 79 86 77 73 102 91 100
mistral-medium-3.5 92 107 98 88 79 96 84 104
gpt-5-chat 93 96 88 108 112 103 92 92
nvidia-nemotron-3-ultra-550b-a55b-nvfp4 94 71 101 92 59 123 113 98
amazon-nova-experimental-chat-26-02-10 95 39 89 78 92 163 95 107
deepseek-v3.2 96 94 93 98 81 88 85 87
glm-4.6 97 104 103 118 100 86 101 101
deepseek-v3.2-exp-thinking 98 93 94 90 80 98 93 103
claude-opus-4-20250514-thinking-16k 99 95 83 57 99 52 57 50
qwen3-max-2025-09-23 100 134 91 93 88 94 97 102

数据来源:LMSYS Chatbot Arena (arena.ai) © Open-source research project by LMSYS Org.