LinkWord
Home
Directory
Articles
AI models
Tools
Pixel Plaza
Settings
ContactRSSFriend linksSubmit site
Privacy Policy·Disclaimer
陕ICP备2025083618号-2

Hot channels

AI ToolsDeveloper ToolsProductivity ToolsEntertainment & MediaJobs & Careers
DirectoryArticlesTools
Back to home

AI model leaderboard · LinkWord

Compare chat, image, and video models by composite score and category-specific metrics. Rankings are editorially maintained for reference.

RankModelProvider
ScoreEditorial composite score; higher ranks higher
Context (K)Context window size (thousand tokens)
Input $Input token price (USD per 1M tokens)
Output $Output token price (USD per 1M tokens)
MMLUMassive Multitask Language Understanding accuracy (%)
HumanEvalCode generation benchmark pass rate (%)
EloArena Elo from human preference battles; higher is stronger
1
G
GPT-5.5
OpenAI's latest flagship model with 1050K context, leading AI reasoning and coding capabilities
OpenAI98.0105053091.5091.401420
2
C
Claude Fable 5
Anthropic's most capable model, positioned above Opus tier. Public Mythos-class model with 1M context, scoring >10% higher than Claude Opus 4.8 on key benchmarks. Adaptive thinking only.
Anthropic97.010001050———
3
G
Gemini 3.5 Pro
Google DeepMind's flagship model with 1000K context window, comprehensive capabilities
Google DeepMind96.010001.5099189.501400
4
G
GPT-5.6 Sol
OpenAI's flagship GPT-5.6 model. HealthBench Professional 60.5%, High on cybersecurity and biosecurity evals. New API Fast mode: 2.5x faster than Standard at 2x price, no change in intelligence (2026-07-30).
OpenAI96.010502100—0
5
C
Claude Fable 5.1
Anthropic's latest Fable series model, same $10/$50 pricing with incremental performance gains
Anthropic96.0100010500—0
6
G
Grok 4.6
xAI's flagship model released Aug 12, 2026 (now operating as SpaceXAI). Scores 61 on the AA Intelligence Index, level with GPT-5.6 Sol. Strong in agentic work: GDPval-AA v2 Elo 1753, Terminal-Bench v2.1 88.4%, τ³-Banking 50.7%. Priced at $2/$6 per 1M tokens with 500K context; ~$0.84 per task, on par with Kimi K3.
xAI95.0500260—1753
7
C
Claude Opus 4.8
Anthropic's most advanced flagship model with industry-leading reasoning and safety features
Anthropic95.010005250—0
8
C
Claude Opus 5
Anthropic's latest Opus flagship, within 0.5% of Fable 5 frontier intelligence at half the cost, $5/$25 per M tokens, #1 on AA Intelligence Leaderboard
Anthropic94.0100052589.20—1320
9
K
Kimi K3
Moonshot AI's latest flagship model. 2.8T-parameter MoE (16/896 experts active), 1M context window, native vision. AA Intelligence Index #4, trailing only Fable 5 and GPT-5.6 Sol. Open weights releasing July 27.
Moonshot AI94.01049315———
10
G
GPT-5.6 Terra
Balanced GPT-5.6 model for everyday work. API price cut 20% on 2026-07-30 ($2.00/M input, $12.00/M output, cached input $0.20/M).
OpenAI93.510502120—0
11
G
Gemini 3.7 Flash
Google's Flash-tier workhorse model released Aug 13, 2026 for coding and agents, intro-priced at half of 3.6 Flash's original rate ($0.75/$3.75 per 1M tokens). FrontierCode 1.1: 43.6%, DeepSWE v1.1: 65.3%.
Google DeepMind93.010480.753.750—0
12
C
Claude Opus 4.7
Anthropic's premium reasoning model, 1000K context, for complex tasks and enterprise use
Anthropic93.010005259192.501400
13
Q
Qwen3.8-Max-0902
Alibaba Qwen's upgraded snapshot of Qwen3.8-Max. Leads Code Arena at 1691 points, surpassing Claude Opus 5 max. Enhanced coding and multi-tool orchestration. 1M context, priced at $2/$6.
Alibaba (Qwen)92.01000260—0
14
G
GPT-6 Astra
OpenAI's flagship model released September 2026. 1.05M context. 98.6% on ARC-AGI-3. Major gains in Artificial Analysis Coding Agent Index. Priced at $10/$50 per 1M tokens.
OpenAI92.0105010500—0
15
M
Muse Spark 1.3
Meta's latest Muse Spark, 1M context, multimodal, priced at $1.25/$4.25
Meta92.010481.254.250—0
16
H
Hy4 Preview
Tencent Hunyuan Hy4 Preview: 770B total / 49B active MoE with 1M context, Apache 2.0 open weights. Uses Gated DSA sparse attention with IndexCache. Self-evaluation shows slight edge over GLM-5.3 and Kimi K3. OpenRouter pricing $0.83/$2.50 per 1M tokens.
Tencent Hunyuan92.010000.832.500—0
17
M
Muse Spark 1.2
Meta Muse Spark 1.2 multimodal model with 1M context, successor to Muse Spark 1.1
Meta92.010481.254.250—0
18
G
GLM-5.3
Z.ai's flagship GLM model, 1.3M context, open-weight released, 1171+805 HN points
Z.ai92.013101.404.400—0
19
C
Claude Opus 5 Fast
Fast-mode variant of Claude Opus 5 with identical capabilities, ~2.5x faster output at 2x base pricing ($10/$50 per M tokens), 1M context, ideal for latency-sensitive long-horizon agentic work
Anthropic92.01000105089.20—1320
20
G
Gemini 3.6 Flash
Google's latest Gemini Flash series flagship for coding, knowledge work, and multimodal tasks. 17% less output tokens, DeepSWE 49%, OSWorld-Verified 83.0%. API $1.50/$7.50.
Google DeepMind92.010491.507.500—0
21
M
Muse Spark 1.1
Meta's flagship multimodal reasoning model with 1M-token context, built for agentic tasks, coding, and computer use
Meta92.010241.254.250—0
22
C
Claude Sonnet 5
Anthropic Claude mid-range model, balancing performance and cost
Anthropic92.01000210——1312
23
D
DeepSeek V4 Pro 0813
Upgraded DeepSeek V4 Pro, 1M context, 1041 HN points, community says beats GPT-5.5 Pro on precision
DeepSeek91.010481.123.350—0
24
Q
Qwen 3.8 2.4T A95B
Qwen 3.8 flagship MoE model with 2.4T total params and 95B active, 1M context window
Alibaba (Qwen)91.01049260—0
25
G
GLM-5.2
Zhipu AI's latest flagship model with strong bilingual understanding and reasoning performance
Z.ai (Zhipu AI)91.010001.404.40———
26
S
Sakana Fugu Ultra
Fugu Ultra is Sakana AI's flagship multi-agent orchestration model. Rather than a single monolithic model, it dynamically orchestrates a pool of expert models to tackle complex multi-step tasks. Benchmarks competitive with Fable 5 and Mythos Preview. 1M context window, text+image input.
Sakana AI91.01000530———
27
G
GPT-5.6 Luna
Fastest, most affordable GPT-5.6 model. API price cut 80% on 2026-07-30 ($0.20/M input, $1.20/M output), near-frontier performance at roughly 6 cents on the dollar per task. Supports tools and multi-step workflows.
OpenAI91.010500.201.200—0
28
Q
Qwen3.8 Max
Alibaba's newest flagship in the Qwen family: 2.4T-parameter MoE (95B active), focused on coding and agentic cowork long-horizon tasks. Open weights Qwen3.8-2.4T-A95B (BF16/FP8) released on Hugging Face Aug 13, 2026 — the largest open-weight model ever; the open version is text-only with mandatory thinking and 262K native context. PaperBench 93.0, the highest published score. API $2/$6 per M tokens.
Alibaba (Qwen)90.01000260—0
29
C
Claude Opus 4.6
Anthropic's previous-generation flagship model with strong reasoning and long-context capabilities
Anthropic90.010005250—0
30
G
GPT-5.4 Pro
OpenAI's professional-grade model offering advanced reasoning for enterprise workloads
OpenAI90.0105030180———
31
G
Gemini 3.5 Flash
Google lightweight flagship model with built-in Computer Use, function calling, Search/Maps Grounding — ideal for agent scenarios
Google DeepMind90.010491.50992.3086.801370
32
G
GPT-5.5 Instant
GPT-5.5 low-latency version, fast responses ideal for chat scenarios
OpenAI88.09220.75389.5088.201350
33
G
Gemini 3.8 Flash
Google's latest Flash model optimized for long-horizon software engineering and autonomous agents. Outperforms larger frontier models on DeepSWE v1.1, 54.9% on HLE-Verified. 1M context, priced at $0.75/$3.75 (promo through end of 2026).
Google DeepMind88.010000.753.750—0
34
Q
Qwen3.8 Flash
Qwen Flash series, 1M context, ultra-cheap $0.15/$0.47, open-weight
Alibaba (Qwen)88.010000.150.470—0
35
G
GLM-5.3 Flash
Lightweight GLM-5.3, 1.3M context, ultra-cheap $0.15/$0.50, 1132 HN points, community says better than DeepSeek
Z.ai88.013100.150.500—0
36
D
DeepSeek V4 Flash
DeepSeek V4 Flash, 1.3M context, price dropped to $0.05/$0.16 (was $0.22/$0.66)
DeepSeek88.013100.220.660—0
37
V
VibeThinker-3B
3B dense reasoning model. AIME26: 94.3, based on Qwen2.5. Uses Spectrum-to-Signal post-training. No tool calling support, focused on math and code reasoning.
Weibo AI88.032—————
38
D
DeepSeek V4 Pro
DeepSeek's flagship reasoning model. Official release deepseek-v4-pro-0813 shipped Aug 13, 2026, replacing V4-Pro-Preview. 1M context, up to 384K output, $0.435/$0.87 per 1M tokens. Big agent gains: Terminal Bench 2.1 87.9, DeepSWE 62.7, Cybergym 83.3 (preview: 72.1/12.8/52.7). Natively supports the Responses API (Codex-compatible).
DeepSeek87.010480.661.980—0
39
D
DeepSeek V4 Flash Vision
DeepSeek V4 Flash Vision experimental, multimodal input support, extremely low pricing
DeepSeek86.010490.220.660—0
40
Q
Qwen3.8-Flash-Next
Alibaba lightweight chat model, 1M context, pricing $0.15/$0.47 per 1M tokens
Alibaba (Qwen)85.010000.150.470—0
41
G
GLM-5.3-Flash
Z.ai lightweight chat model, 1.3M context, ultra-low pricing $0.07/$0.25 per 1M tokens
Z.ai85.013110.070.250—0
42
K
Kimi K2.7 Code
Moonshot AI's specialized coding model based on Kimi K2.6 architecture
Moonshot AI85.02560.743.50———
43
Q
Qwen3.7 Max
Alibaba Qwen's most powerful model, 1000K context, MoE architecture
Alibaba (Qwen)85.010001.253.7587871300
44
G
Gemini 3.1 Pro
Google DeepMind previous-gen flagship model with 1049K context
Google DeepMind85.0104921287.50851300
45
G
Grok 4.5
xAI's latest coding and agentic model based on 1.5T V9 foundation, 500K context, $2/$6/M tokens, coding performance competitive with GPT-5.5
xAI83.0500260—0
46
Q
Qwen 3.8 27B
Qwen 3.8 27B open-weight model with 1M context, strong reasoning capabilities
Alibaba (Qwen)83.010000.4030—0
47
S
Seed 2.1 Turbo
ByteDance Seed 2.1 Turbo model, 262K context, fast speed and low pricing
ByteDance82.02620.502.500—0
48
G
Grok 4.20 Multi-Agent
xAI multi-agent reasoning model built on Grok 4.20, supports multi-agent collaborative orchestration with 2M context, ideal for complex task decomposition and parallel execution
xAI82.020001.252.5086851275
49
C
Cursor Composer 2.5
Cursor AI IDE built-in Agent coding mode, 256K context
Cursor82.02560085861260
50
K
Kimi K2.6
Moonshot AI's latest Kimi model, 262K context, strong Agent capabilities
Moonshot AI82.02620.683.4285.5084.501280
51
G
GPT-5.4
OpenAI GPT-5.4 flagship model with 1050K context
OpenAI82.010502.501588.2087.501320
52
G
Grok 4.20
xAI's latest model featuring multi-agent coordination and enhanced reasoning
xAI81.020001.252.50———
53
C
Claude Sonnet 4.6
Anthropic Claude Sonnet series, mid-to-high-end reasoning model
Anthropic80.0100031586.50881280
54
M
MiniMax M3
MiniMax's latest flagship model with competitive performance across reasoning tasks
MiniMax80.010000.301.20———
55
W
Windsurf SWE-1.6
Windsurf full-stack AI coding assistant, 200K context
Windsurf (Codeium)80.020000000
56
G
Grok 4.3
xAI Grok latest flagship with 1000K context
xAI80.010001.252.5086851270
57
M
MiMo-V2.5 Pro
Xiaomi MiMo-V2.5 Pro on-device LLM
Xiaomi78.010000.440.8885841260
58
S
Sakana Namazu
Sakana AI's multimodal model, 262K context, image input, $0.95/$4.00
Sakana AI78.02620.9540—0
59
Q
Qwen3.8-27B
Qwen3.8 series open model (released Aug 14, 2026, Apache 2.0): 27B dense multimodal model with native image/video understanding, 262K native context (extendable to 1M via YaRN), Gated DeltaNet hybrid architecture, thinking mode on by default and toggleable per request. SWE-bench Pro 61.7, QwenSWEBench 79.0, DeepSWE 1.1 42.2, LiveCodeBench v6 90.3, OSWorld 84.3 - beats Claude Opus 4.6 Max on multiple coding and agent benchmarks. Runs locally on consumer hardware (~48 tps q4km on a 4090).
Alibaba (Qwen)78.010000.453.200—0
60
L
LongCat-2.0
1.6 trillion parameter MoE model from Meituan (LongCat), ~48B activated per token, trained on domestic AI ASIC superpods, 1M context window, MIT license
Meituan78.0104900———
61
M
Muse Glimmer
Meta Superintelligence Labs' open 30B on-device agentic model (Apache 2.0), distilled from Muse Spark, quantized to under 20GB for single-GPU local runs, with multimodal input and tool calling.
Meta (Superintelligence Labs)78.0128000—0
62
Q
Qwen3.6 Plus
Alibaba Qwen3.6 Plus mid-to-high-end model, 1000K context
Alibaba (Qwen)76.010000.331.9584841250
63
G
GPT-4o
OpenAI GPT-4o multimodal model with 128K context
OpenAI75.01282.501088.7090.201287
64
K
K2 Horizon 375B-A23B
IFM open-source MoE flagship, 375B total/23B active params, Apache 2.0, full lifecycle openness
IFM (Institute of Foundation Models)75.0128000—0
65
M
Muse Glimmer 30B
Meta's mid-size multimodal model, 131K context, $0.30/$1.10
Meta75.01310.301.100—0
66
G
Gemini 3.5 Flash-Lite
Google's fastest 3.5-series model at 350 tokens/s, designed for low-latency, high-throughput agentic search and document processing. Priced $0.30/$2.50. SWE-Bench Pro 54.2%, OSWorld 74.0%. 1M context.
Google DeepMind75.010000.302.500—0
67
Q
Qwen3.7 Plus
Alibaba's flagship Qwen model with improved instruction following and tool use
Alibaba (Qwen)75.010000.321.28———
68
X
Xiaomi-Robotics-1
Xiaomi's embodied foundation model, 100K+ hours real-world pretraining, natural language instructions, 1600+ scenario generalization
Xiaomi75.0128000—0
69
G
GLM-5.1
Zhipu AI GLM series previous-gen flagship, 200K context
智谱AI (Zhipu)75.02000.401.2083821240
70
G
Gemini Robotics 2
Google DeepMind's next-gen embodied intelligence family: Gemini Robotics 2 (VLA whole-body control), Gemini Robotics ER 2 (embodied reasoning, on AI Studio) and On-Device 2 (on-device). Fine dexterity, multi-robot collaboration, adapts to new robot bodies in hours.
Google DeepMind74.0128000—0
71
C
Cursor Composer 2
Cursor Composer 2 AI coding assistant
Cursor72.02560082821220
72
M
MiMo-V2.5
Xiaomi MiMo-V2.5 on-device LLM
Xiaomi72.010490.150.29———
73
M
MiniMax-M2.7
MiniMax M2.7 chat model
MiniMax72.02050.281.2082811220
74
K
Kimi K2.5
Moonshot AI Kimi K2.5 chat model, 262K context
Moonshot AI72.02620.401.9082821220
75
G
Gemini 3 Flash
Google DeepMind lightweight Gemini model with 1000K context
Google DeepMind70.010000.150.608280.501220
76
G
GLM-5
Zhipu AI GLM-5 model, previous-gen flagship
智谱AI (Zhipu)70.02000.300.9081791210
77
Q
Qwen3.5 397B
Alibaba Qwen3.5 397B parameter large model
Alibaba (Qwen)68.02620.451.3580.5080.501200
78
I
Inkling
Open-weight MoE model by Thinking Machines Lab, 975B total / 41B active params. Multimodal (text, image, audio), 1M context. AA Intelligence Index 41 - leading US open weights model. Strong agent performance: Elo 1238 on GDPval-AA v2, beating Kimi K2.6 and DeepSeek V4 Flash.
Thinking Machines Lab68.0104914.050—1238
79
G
GPT-5.4 Mini
OpenAI's compact model in the GPT-5.4 series, optimized for efficiency
OpenAI67.04000.754.50———
80
Q
Qwen3 Coder 480B A35B
Qwen's most powerful open-source coding model. 480B MoE with 35B active params, native 256K context (YaRN scalable to 1M). Strong SWE-Bench performance. Apache 2.0 licensed. Ships with Qwen Code CLI.
Alibaba (Qwen)66.02560.221.80———
81
I
Inkling Small
Open-weights MoE from Thinking Machines Lab, 276B total/12B active, comparable performance to Inkling at a quarter of the size. Native audio+image reasoning, variable thinking effort, up to 1M context. Highest-scoring open-weight model on ARC Prize.
Thinking Machines Lab66.0524000—0
82
G
Gemini 2.5 Pro
Google DeepMind previous-gen Gemini high-end model
Google DeepMind65.010000.351.4080.50781180
83
G
Grok 3
xAI previous-gen Grok model with 1000K context
xAI65.010000.150.6080801180
84
H
Hunyuan Hy3 Preview
Tencent Hunyuan Hy3 Preview model, 256K context
Tencent Hunyuan65.02560.060.1879781180
85
S
Seed 2.0 Code
ByteDance Seed 2.0 Code model optimized for frontend development, multilingual coding, and agentic coding
ByteDance Seed65.02620.5030—0
86
N
Nemotron 3.5 Lightning
NVIDIA's lightweight inference model, 262K context, ultra-cheap $0.08/$0.20, 262 HN points
NVIDIA65.02620.080.200—0
87
P
Poolside Laguna S 2.1
Poolside coding-specialized agent model, 118B-A8B MoE, 1M context, Terminal-Bench 70.2%, DeepSWE 40.4%, focused on autonomous long-horizon engineering work
Poolside65.010000.100.20———
88
D
DeepSeek V3.2
DeepSeek's capable mid-range model offering strong performance at lower cost
DeepSeek63.01640.230.340—0
89
N
Nemotron 3 Ultra
NVIDIA's most capable open-weight model with strong benchmark performance
NVIDIA62.010000.502.20———
90
G
GPT-Live-1
OpenAI real-time voice conversation model, supports interruption and continuation, simultaneous voice chat and reasoning
OpenAI62.032000—0
91
C
Claude 4.5 Haiku
Anthropic lightweight fast model with 200K context
Anthropic60.02000.80478751150
92
K
K2 Horizon 36B-A4B
IFM MoVA sparse attention model, 36B total/4B active, near-dense-32B performance
IFM60.0128000—0
93
M
Mercury 2.5 Preview
Inception's diffusion language model, ultra-fast inference, 260K context, $0.04/$0.15
Inception60.02600.040.150—0
94
G
Gemini 3.1 Flash Lite
Google's lightweight Gemini model optimized for speed and efficiency on edge devices
Google DeepMind60.010490.251.500—0
95
C
Codestral 2508
Mistral AI's specialized coding model, optimized for code generation, completion, and refactoring across multiple programming languages
Mistral AI60.02560.300.900—0
96
S
Step 3.7 Flash
StepFun's latest multimodal MoE model with 196B parameter language backbone and vision encoder for native image/video understanding
StepFun60.02560.201.157875—
97
G
GPT-5.4 Nano
OpenAI's compact model in the GPT-5.4 lineup, designed for rapid inference and cost-efficient deployment
OpenAI58.04000.201.25———
98
L
Laguna XS 2.1
Poolside Laguna XS 2.1 is a 33B-A3B MoE coding agent model optimized for local deployment. Builds on XS.2 with improved SWE-bench Multilingual (63.1%) and stronger terminal-style task performance. Supports 256K context, runs locally in vLLM/SGLang.
Poolside58.02560.060.12———
99
N
Nex AGI Nex-N2-Pro
Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active / 397B total parameters. Built on Qwen3.5 architecture, supports text and image input, 262K context window. Extremely cost-effective at $0.25/M input tokens.
Nex AGI58.02620.251———
100
C
Cursor Composer 1.5
Cursor Composer 1.5 earlier version
Cursor58.02000076741150

Disclaimer

This leaderboard is editorially curated and updated by LinkWord. Composite scores, benchmark figures, pricing, and capability metrics are provided for browsing and rough comparison only—not as purchase advice, performance guarantees, or endorsements. Data may lag official vendor releases; results vary by evaluation method, model version, and use case. Third-party trademarks belong to their respective owners. You assume all risk from relying on this page; if you cite or republish, please attribute the source.