| 1 | GGPT-5.5 OpenAI's latest flagship model with 1050K context, leading AI reasoning and coding capabilities | OpenAI | 98.0 | 1050 | 5 | 30 | 91.50 | 91.40 | 1420 |
| 2 | CClaude Fable 5 Anthropic's most capable model, positioned above Opus tier. Public Mythos-class model with 1M context, scoring >10% higher than Claude Opus 4.8 on key benchmarks. Adaptive thinking only. | Anthropic | 97.0 | 1000 | 10 | 50 | — | — | — |
| 3 | GGemini 3.5 Pro Google DeepMind's flagship model with 1000K context window, comprehensive capabilities | Google DeepMind | 96.0 | 1000 | 1.50 | 9 | 91 | 89.50 | 1400 |
| 4 | GGPT-5.6 Sol OpenAI's flagship GPT-5.6 model. HealthBench Professional 60.5%, High on cybersecurity and biosecurity evals. New API Fast mode: 2.5x faster than Standard at 2x price, no change in intelligence (2026-07-30). | OpenAI | 96.0 | 1050 | 2 | 10 | 0 | — | 0 |
| 5 | CClaude Fable 5.1 Anthropic's latest Fable series model, same $10/$50 pricing with incremental performance gains | Anthropic | 96.0 | 1000 | 10 | 50 | 0 | — | 0 |
| 6 | GGrok 4.6 xAI's flagship model released Aug 12, 2026 (now operating as SpaceXAI). Scores 61 on the AA Intelligence Index, level with GPT-5.6 Sol. Strong in agentic work: GDPval-AA v2 Elo 1753, Terminal-Bench v2.1 88.4%, τ³-Banking 50.7%. Priced at $2/$6 per 1M tokens with 500K context; ~$0.84 per task, on par with Kimi K3. | xAI | 95.0 | 500 | 2 | 6 | 0 | — | 1753 |
| 7 | CClaude Opus 4.8 Anthropic's most advanced flagship model with industry-leading reasoning and safety features | Anthropic | 95.0 | 1000 | 5 | 25 | 0 | — | 0 |
| 8 | CClaude Opus 5 Anthropic's latest Opus flagship, within 0.5% of Fable 5 frontier intelligence at half the cost, $5/$25 per M tokens, #1 on AA Intelligence Leaderboard | Anthropic | 94.0 | 1000 | 5 | 25 | 89.20 | — | 1320 |
| 9 | KKimi K3 Moonshot AI's latest flagship model. 2.8T-parameter MoE (16/896 experts active), 1M context window, native vision. AA Intelligence Index #4, trailing only Fable 5 and GPT-5.6 Sol. Open weights releasing July 27. | Moonshot AI | 94.0 | 1049 | 3 | 15 | — | — | — |
| 10 | GGPT-5.6 Terra Balanced GPT-5.6 model for everyday work. API price cut 20% on 2026-07-30 ($2.00/M input, $12.00/M output, cached input $0.20/M). | OpenAI | 93.5 | 1050 | 2 | 12 | 0 | — | 0 |
| 11 | GGemini 3.7 Flash Google's Flash-tier workhorse model released Aug 13, 2026 for coding and agents, intro-priced at half of 3.6 Flash's original rate ($0.75/$3.75 per 1M tokens). FrontierCode 1.1: 43.6%, DeepSWE v1.1: 65.3%. | Google DeepMind | 93.0 | 1048 | 0.75 | 3.75 | 0 | — | 0 |
| 12 | CClaude Opus 4.7 Anthropic's premium reasoning model, 1000K context, for complex tasks and enterprise use | Anthropic | 93.0 | 1000 | 5 | 25 | 91 | 92.50 | 1400 |
| 13 | QQwen3.8-Max-0902 Alibaba Qwen's upgraded snapshot of Qwen3.8-Max. Leads Code Arena at 1691 points, surpassing Claude Opus 5 max. Enhanced coding and multi-tool orchestration. 1M context, priced at $2/$6. | Alibaba (Qwen) | 92.0 | 1000 | 2 | 6 | 0 | — | 0 |
| 14 | GGPT-6 Astra OpenAI's flagship model released September 2026. 1.05M context. 98.6% on ARC-AGI-3. Major gains in Artificial Analysis Coding Agent Index. Priced at $10/$50 per 1M tokens. | OpenAI | 92.0 | 1050 | 10 | 50 | 0 | — | 0 |
| 15 | MMuse Spark 1.3 Meta's latest Muse Spark, 1M context, multimodal, priced at $1.25/$4.25 | Meta | 92.0 | 1048 | 1.25 | 4.25 | 0 | — | 0 |
| 16 | HHy4 Preview Tencent Hunyuan Hy4 Preview: 770B total / 49B active MoE with 1M context, Apache 2.0 open weights. Uses Gated DSA sparse attention with IndexCache. Self-evaluation shows slight edge over GLM-5.3 and Kimi K3. OpenRouter pricing $0.83/$2.50 per 1M tokens. | Tencent Hunyuan | 92.0 | 1000 | 0.83 | 2.50 | 0 | — | 0 |
| 17 | MMuse Spark 1.2 Meta Muse Spark 1.2 multimodal model with 1M context, successor to Muse Spark 1.1 | Meta | 92.0 | 1048 | 1.25 | 4.25 | 0 | — | 0 |
| 18 | GGLM-5.3 Z.ai's flagship GLM model, 1.3M context, open-weight released, 1171+805 HN points | Z.ai | 92.0 | 1310 | 1.40 | 4.40 | 0 | — | 0 |
| 19 | CClaude Opus 5 Fast Fast-mode variant of Claude Opus 5 with identical capabilities, ~2.5x faster output at 2x base pricing ($10/$50 per M tokens), 1M context, ideal for latency-sensitive long-horizon agentic work | Anthropic | 92.0 | 1000 | 10 | 50 | 89.20 | — | 1320 |
| 20 | GGemini 3.6 Flash Google's latest Gemini Flash series flagship for coding, knowledge work, and multimodal tasks. 17% less output tokens, DeepSWE 49%, OSWorld-Verified 83.0%. API $1.50/$7.50. | Google DeepMind | 92.0 | 1049 | 1.50 | 7.50 | 0 | — | 0 |
| 21 | MMuse Spark 1.1 Meta's flagship multimodal reasoning model with 1M-token context, built for agentic tasks, coding, and computer use | Meta | 92.0 | 1024 | 1.25 | 4.25 | 0 | — | 0 |
| 22 | CClaude Sonnet 5 Anthropic Claude mid-range model, balancing performance and cost | Anthropic | 92.0 | 1000 | 2 | 10 | — | — | 1312 |
| 23 | DDeepSeek V4 Pro 0813 Upgraded DeepSeek V4 Pro, 1M context, 1041 HN points, community says beats GPT-5.5 Pro on precision | DeepSeek | 91.0 | 1048 | 1.12 | 3.35 | 0 | — | 0 |
| 24 | QQwen 3.8 2.4T A95B Qwen 3.8 flagship MoE model with 2.4T total params and 95B active, 1M context window | Alibaba (Qwen) | 91.0 | 1049 | 2 | 6 | 0 | — | 0 |
| 25 | GGLM-5.2 Zhipu AI's latest flagship model with strong bilingual understanding and reasoning performance | Z.ai (Zhipu AI) | 91.0 | 1000 | 1.40 | 4.40 | — | — | — |
| 26 | SSakana Fugu Ultra Fugu Ultra is Sakana AI's flagship multi-agent orchestration model. Rather than a single monolithic model, it dynamically orchestrates a pool of expert models to tackle complex multi-step tasks. Benchmarks competitive with Fable 5 and Mythos Preview. 1M context window, text+image input. | Sakana AI | 91.0 | 1000 | 5 | 30 | — | — | — |
| 27 | GGPT-5.6 Luna Fastest, most affordable GPT-5.6 model. API price cut 80% on 2026-07-30 ($0.20/M input, $1.20/M output), near-frontier performance at roughly 6 cents on the dollar per task. Supports tools and multi-step workflows. | OpenAI | 91.0 | 1050 | 0.20 | 1.20 | 0 | — | 0 |
| 28 | QQwen3.8 Max Alibaba's newest flagship in the Qwen family: 2.4T-parameter MoE (95B active), focused on coding and agentic cowork long-horizon tasks. Open weights Qwen3.8-2.4T-A95B (BF16/FP8) released on Hugging Face Aug 13, 2026 — the largest open-weight model ever; the open version is text-only with mandatory thinking and 262K native context. PaperBench 93.0, the highest published score. API $2/$6 per M tokens. | Alibaba (Qwen) | 90.0 | 1000 | 2 | 6 | 0 | — | 0 |
| 29 | CClaude Opus 4.6 Anthropic's previous-generation flagship model with strong reasoning and long-context capabilities | Anthropic | 90.0 | 1000 | 5 | 25 | 0 | — | 0 |
| 30 | GGPT-5.4 Pro OpenAI's professional-grade model offering advanced reasoning for enterprise workloads | OpenAI | 90.0 | 1050 | 30 | 180 | — | — | — |
| 31 | GGemini 3.5 Flash Google lightweight flagship model with built-in Computer Use, function calling, Search/Maps Grounding — ideal for agent scenarios | Google DeepMind | 90.0 | 1049 | 1.50 | 9 | 92.30 | 86.80 | 1370 |
| 32 | GGPT-5.5 Instant GPT-5.5 low-latency version, fast responses ideal for chat scenarios | OpenAI | 88.0 | 922 | 0.75 | 3 | 89.50 | 88.20 | 1350 |
| 33 | GGemini 3.8 Flash Google's latest Flash model optimized for long-horizon software engineering and autonomous agents. Outperforms larger frontier models on DeepSWE v1.1, 54.9% on HLE-Verified. 1M context, priced at $0.75/$3.75 (promo through end of 2026). | Google DeepMind | 88.0 | 1000 | 0.75 | 3.75 | 0 | — | 0 |
| 34 | QQwen3.8 Flash Qwen Flash series, 1M context, ultra-cheap $0.15/$0.47, open-weight | Alibaba (Qwen) | 88.0 | 1000 | 0.15 | 0.47 | 0 | — | 0 |
| 35 | GGLM-5.3 Flash Lightweight GLM-5.3, 1.3M context, ultra-cheap $0.15/$0.50, 1132 HN points, community says better than DeepSeek | Z.ai | 88.0 | 1310 | 0.15 | 0.50 | 0 | — | 0 |
| 36 | DDeepSeek V4 Flash DeepSeek V4 Flash, 1.3M context, price dropped to $0.05/$0.16 (was $0.22/$0.66) | DeepSeek | 88.0 | 1310 | 0.22 | 0.66 | 0 | — | 0 |
| 37 | VVibeThinker-3B 3B dense reasoning model. AIME26: 94.3, based on Qwen2.5. Uses Spectrum-to-Signal post-training. No tool calling support, focused on math and code reasoning. | Weibo AI | 88.0 | 32 | — | — | — | — | — |
| 38 | DDeepSeek V4 Pro DeepSeek's flagship reasoning model. Official release deepseek-v4-pro-0813 shipped Aug 13, 2026, replacing V4-Pro-Preview. 1M context, up to 384K output, $0.435/$0.87 per 1M tokens. Big agent gains: Terminal Bench 2.1 87.9, DeepSWE 62.7, Cybergym 83.3 (preview: 72.1/12.8/52.7). Natively supports the Responses API (Codex-compatible). | DeepSeek | 87.0 | 1048 | 0.66 | 1.98 | 0 | — | 0 |
| 39 | DDeepSeek V4 Flash Vision DeepSeek V4 Flash Vision experimental, multimodal input support, extremely low pricing | DeepSeek | 86.0 | 1049 | 0.22 | 0.66 | 0 | — | 0 |
| 40 | QQwen3.8-Flash-Next Alibaba lightweight chat model, 1M context, pricing $0.15/$0.47 per 1M tokens | Alibaba (Qwen) | 85.0 | 1000 | 0.15 | 0.47 | 0 | — | 0 |
| 41 | GGLM-5.3-Flash Z.ai lightweight chat model, 1.3M context, ultra-low pricing $0.07/$0.25 per 1M tokens | Z.ai | 85.0 | 1311 | 0.07 | 0.25 | 0 | — | 0 |
| 42 | KKimi K2.7 Code Moonshot AI's specialized coding model based on Kimi K2.6 architecture | Moonshot AI | 85.0 | 256 | 0.74 | 3.50 | — | — | — |
| 43 | QQwen3.7 Max Alibaba Qwen's most powerful model, 1000K context, MoE architecture | Alibaba (Qwen) | 85.0 | 1000 | 1.25 | 3.75 | 87 | 87 | 1300 |
| 44 | GGemini 3.1 Pro Google DeepMind previous-gen flagship model with 1049K context | Google DeepMind | 85.0 | 1049 | 2 | 12 | 87.50 | 85 | 1300 |
| 45 | GGrok 4.5 xAI's latest coding and agentic model based on 1.5T V9 foundation, 500K context, $2/$6/M tokens, coding performance competitive with GPT-5.5 | xAI | 83.0 | 500 | 2 | 6 | 0 | — | 0 |
| 46 | QQwen 3.8 27B Qwen 3.8 27B open-weight model with 1M context, strong reasoning capabilities | Alibaba (Qwen) | 83.0 | 1000 | 0.40 | 3 | 0 | — | 0 |
| 47 | SSeed 2.1 Turbo ByteDance Seed 2.1 Turbo model, 262K context, fast speed and low pricing | ByteDance | 82.0 | 262 | 0.50 | 2.50 | 0 | — | 0 |
| 48 | GGrok 4.20 Multi-Agent xAI multi-agent reasoning model built on Grok 4.20, supports multi-agent collaborative orchestration with 2M context, ideal for complex task decomposition and parallel execution | xAI | 82.0 | 2000 | 1.25 | 2.50 | 86 | 85 | 1275 |
| 49 | CCursor Composer 2.5 Cursor AI IDE built-in Agent coding mode, 256K context | Cursor | 82.0 | 256 | 0 | 0 | 85 | 86 | 1260 |
| 50 | KKimi K2.6 Moonshot AI's latest Kimi model, 262K context, strong Agent capabilities | Moonshot AI | 82.0 | 262 | 0.68 | 3.42 | 85.50 | 84.50 | 1280 |
| 51 | GGPT-5.4 OpenAI GPT-5.4 flagship model with 1050K context | OpenAI | 82.0 | 1050 | 2.50 | 15 | 88.20 | 87.50 | 1320 |
| 52 | GGrok 4.20 xAI's latest model featuring multi-agent coordination and enhanced reasoning | xAI | 81.0 | 2000 | 1.25 | 2.50 | — | — | — |
| 53 | CClaude Sonnet 4.6 Anthropic Claude Sonnet series, mid-to-high-end reasoning model | Anthropic | 80.0 | 1000 | 3 | 15 | 86.50 | 88 | 1280 |
| 54 | MMiniMax M3 MiniMax's latest flagship model with competitive performance across reasoning tasks | MiniMax | 80.0 | 1000 | 0.30 | 1.20 | — | — | — |
| 55 | WWindsurf SWE-1.6 Windsurf full-stack AI coding assistant, 200K context | Windsurf (Codeium) | 80.0 | 200 | 0 | 0 | 0 | 0 | 0 |
| 56 | GGrok 4.3 xAI Grok latest flagship with 1000K context | xAI | 80.0 | 1000 | 1.25 | 2.50 | 86 | 85 | 1270 |
| 57 | MMiMo-V2.5 Pro Xiaomi MiMo-V2.5 Pro on-device LLM | Xiaomi | 78.0 | 1000 | 0.44 | 0.88 | 85 | 84 | 1260 |
| 58 | SSakana Namazu Sakana AI's multimodal model, 262K context, image input, $0.95/$4.00 | Sakana AI | 78.0 | 262 | 0.95 | 4 | 0 | — | 0 |
| 59 | QQwen3.8-27B Qwen3.8 series open model (released Aug 14, 2026, Apache 2.0): 27B dense multimodal model with native image/video understanding, 262K native context (extendable to 1M via YaRN), Gated DeltaNet hybrid architecture, thinking mode on by default and toggleable per request. SWE-bench Pro 61.7, QwenSWEBench 79.0, DeepSWE 1.1 42.2, LiveCodeBench v6 90.3, OSWorld 84.3 - beats Claude Opus 4.6 Max on multiple coding and agent benchmarks. Runs locally on consumer hardware (~48 tps q4km on a 4090). | Alibaba (Qwen) | 78.0 | 1000 | 0.45 | 3.20 | 0 | — | 0 |
| 60 | LLongCat-2.0 1.6 trillion parameter MoE model from Meituan (LongCat), ~48B activated per token, trained on domestic AI ASIC superpods, 1M context window, MIT license | Meituan | 78.0 | 1049 | 0 | 0 | — | — | — |
| 61 | MMuse Glimmer Meta Superintelligence Labs' open 30B on-device agentic model (Apache 2.0), distilled from Muse Spark, quantized to under 20GB for single-GPU local runs, with multimodal input and tool calling. | Meta (Superintelligence Labs) | 78.0 | 128 | 0 | 0 | 0 | — | 0 |
| 62 | QQwen3.6 Plus Alibaba Qwen3.6 Plus mid-to-high-end model, 1000K context | Alibaba (Qwen) | 76.0 | 1000 | 0.33 | 1.95 | 84 | 84 | 1250 |
| 63 | GGPT-4o OpenAI GPT-4o multimodal model with 128K context | OpenAI | 75.0 | 128 | 2.50 | 10 | 88.70 | 90.20 | 1287 |
| 64 | KK2 Horizon 375B-A23B IFM open-source MoE flagship, 375B total/23B active params, Apache 2.0, full lifecycle openness | IFM (Institute of Foundation Models) | 75.0 | 128 | 0 | 0 | 0 | — | 0 |
| 65 | MMuse Glimmer 30B Meta's mid-size multimodal model, 131K context, $0.30/$1.10 | Meta | 75.0 | 131 | 0.30 | 1.10 | 0 | — | 0 |
| 66 | GGemini 3.5 Flash-Lite Google's fastest 3.5-series model at 350 tokens/s, designed for low-latency, high-throughput agentic search and document processing. Priced $0.30/$2.50. SWE-Bench Pro 54.2%, OSWorld 74.0%. 1M context. | Google DeepMind | 75.0 | 1000 | 0.30 | 2.50 | 0 | — | 0 |
| 67 | QQwen3.7 Plus Alibaba's flagship Qwen model with improved instruction following and tool use | Alibaba (Qwen) | 75.0 | 1000 | 0.32 | 1.28 | — | — | — |
| 68 | XXiaomi-Robotics-1 Xiaomi's embodied foundation model, 100K+ hours real-world pretraining, natural language instructions, 1600+ scenario generalization | Xiaomi | 75.0 | 128 | 0 | 0 | 0 | — | 0 |
| 69 | GGLM-5.1 Zhipu AI GLM series previous-gen flagship, 200K context | 智谱AI (Zhipu) | 75.0 | 200 | 0.40 | 1.20 | 83 | 82 | 1240 |
| 70 | GGemini Robotics 2 Google DeepMind's next-gen embodied intelligence family: Gemini Robotics 2 (VLA whole-body control), Gemini Robotics ER 2 (embodied reasoning, on AI Studio) and On-Device 2 (on-device). Fine dexterity, multi-robot collaboration, adapts to new robot bodies in hours. | Google DeepMind | 74.0 | 128 | 0 | 0 | 0 | — | 0 |
| 71 | CCursor Composer 2 Cursor Composer 2 AI coding assistant | Cursor | 72.0 | 256 | 0 | 0 | 82 | 82 | 1220 |
| 72 | MMiMo-V2.5 Xiaomi MiMo-V2.5 on-device LLM | Xiaomi | 72.0 | 1049 | 0.15 | 0.29 | — | — | — |
| 73 | MMiniMax-M2.7 MiniMax M2.7 chat model | MiniMax | 72.0 | 205 | 0.28 | 1.20 | 82 | 81 | 1220 |
| 74 | KKimi K2.5 Moonshot AI Kimi K2.5 chat model, 262K context | Moonshot AI | 72.0 | 262 | 0.40 | 1.90 | 82 | 82 | 1220 |
| 75 | GGemini 3 Flash Google DeepMind lightweight Gemini model with 1000K context | Google DeepMind | 70.0 | 1000 | 0.15 | 0.60 | 82 | 80.50 | 1220 |
| 76 | GGLM-5 Zhipu AI GLM-5 model, previous-gen flagship | 智谱AI (Zhipu) | 70.0 | 200 | 0.30 | 0.90 | 81 | 79 | 1210 |
| 77 | QQwen3.5 397B Alibaba Qwen3.5 397B parameter large model | Alibaba (Qwen) | 68.0 | 262 | 0.45 | 1.35 | 80.50 | 80.50 | 1200 |
| 78 | IInkling Open-weight MoE model by Thinking Machines Lab, 975B total / 41B active params. Multimodal (text, image, audio), 1M context. AA Intelligence Index 41 - leading US open weights model. Strong agent performance: Elo 1238 on GDPval-AA v2, beating Kimi K2.6 and DeepSeek V4 Flash. | Thinking Machines Lab | 68.0 | 1049 | 1 | 4.05 | 0 | — | 1238 |
| 79 | GGPT-5.4 Mini OpenAI's compact model in the GPT-5.4 series, optimized for efficiency | OpenAI | 67.0 | 400 | 0.75 | 4.50 | — | — | — |
| 80 | QQwen3 Coder 480B A35B Qwen's most powerful open-source coding model. 480B MoE with 35B active params, native 256K context (YaRN scalable to 1M). Strong SWE-Bench performance. Apache 2.0 licensed. Ships with Qwen Code CLI. | Alibaba (Qwen) | 66.0 | 256 | 0.22 | 1.80 | — | — | — |
| 81 | IInkling Small Open-weights MoE from Thinking Machines Lab, 276B total/12B active, comparable performance to Inkling at a quarter of the size. Native audio+image reasoning, variable thinking effort, up to 1M context. Highest-scoring open-weight model on ARC Prize. | Thinking Machines Lab | 66.0 | 524 | 0 | 0 | 0 | — | 0 |
| 82 | GGemini 2.5 Pro Google DeepMind previous-gen Gemini high-end model | Google DeepMind | 65.0 | 1000 | 0.35 | 1.40 | 80.50 | 78 | 1180 |
| 83 | GGrok 3 xAI previous-gen Grok model with 1000K context | xAI | 65.0 | 1000 | 0.15 | 0.60 | 80 | 80 | 1180 |
| 84 | HHunyuan Hy3 Preview Tencent Hunyuan Hy3 Preview model, 256K context | Tencent Hunyuan | 65.0 | 256 | 0.06 | 0.18 | 79 | 78 | 1180 |
| 85 | SSeed 2.0 Code ByteDance Seed 2.0 Code model optimized for frontend development, multilingual coding, and agentic coding | ByteDance Seed | 65.0 | 262 | 0.50 | 3 | 0 | — | 0 |
| 86 | NNemotron 3.5 Lightning NVIDIA's lightweight inference model, 262K context, ultra-cheap $0.08/$0.20, 262 HN points | NVIDIA | 65.0 | 262 | 0.08 | 0.20 | 0 | — | 0 |
| 87 | PPoolside Laguna S 2.1 Poolside coding-specialized agent model, 118B-A8B MoE, 1M context, Terminal-Bench 70.2%, DeepSWE 40.4%, focused on autonomous long-horizon engineering work | Poolside | 65.0 | 1000 | 0.10 | 0.20 | — | — | — |
| 88 | DDeepSeek V3.2 DeepSeek's capable mid-range model offering strong performance at lower cost | DeepSeek | 63.0 | 164 | 0.23 | 0.34 | 0 | — | 0 |
| 89 | NNemotron 3 Ultra NVIDIA's most capable open-weight model with strong benchmark performance | NVIDIA | 62.0 | 1000 | 0.50 | 2.20 | — | — | — |
| 90 | GGPT-Live-1 OpenAI real-time voice conversation model, supports interruption and continuation, simultaneous voice chat and reasoning | OpenAI | 62.0 | 32 | 0 | 0 | 0 | — | 0 |
| 91 | CClaude 4.5 Haiku Anthropic lightweight fast model with 200K context | Anthropic | 60.0 | 200 | 0.80 | 4 | 78 | 75 | 1150 |
| 92 | KK2 Horizon 36B-A4B IFM MoVA sparse attention model, 36B total/4B active, near-dense-32B performance | IFM | 60.0 | 128 | 0 | 0 | 0 | — | 0 |
| 93 | MMercury 2.5 Preview Inception's diffusion language model, ultra-fast inference, 260K context, $0.04/$0.15 | Inception | 60.0 | 260 | 0.04 | 0.15 | 0 | — | 0 |
| 94 | GGemini 3.1 Flash Lite Google's lightweight Gemini model optimized for speed and efficiency on edge devices | Google DeepMind | 60.0 | 1049 | 0.25 | 1.50 | 0 | — | 0 |
| 95 | CCodestral 2508 Mistral AI's specialized coding model, optimized for code generation, completion, and refactoring across multiple programming languages | Mistral AI | 60.0 | 256 | 0.30 | 0.90 | 0 | — | 0 |
| 96 | SStep 3.7 Flash StepFun's latest multimodal MoE model with 196B parameter language backbone and vision encoder for native image/video understanding | StepFun | 60.0 | 256 | 0.20 | 1.15 | 78 | 75 | — |
| 97 | GGPT-5.4 Nano OpenAI's compact model in the GPT-5.4 lineup, designed for rapid inference and cost-efficient deployment | OpenAI | 58.0 | 400 | 0.20 | 1.25 | — | — | — |
| 98 | LLaguna XS 2.1 Poolside Laguna XS 2.1 is a 33B-A3B MoE coding agent model optimized for local deployment. Builds on XS.2 with improved SWE-bench Multilingual (63.1%) and stronger terminal-style task performance. Supports 256K context, runs locally in vLLM/SGLang. | Poolside | 58.0 | 256 | 0.06 | 0.12 | — | — | — |
| 99 | NNex AGI Nex-N2-Pro Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active / 397B total parameters. Built on Qwen3.5 architecture, supports text and image input, 262K context window. Extremely cost-effective at $0.25/M input tokens. | Nex AGI | 58.0 | 262 | 0.25 | 1 | — | — | — |
| 100 | CCursor Composer 1.5 Cursor Composer 1.5 earlier version | Cursor | 58.0 | 200 | 0 | 0 | 76 | 74 | 1150 |