Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Β 

History

174 Commits
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

AI Models Matrix Logo

Awesome AI Models Matrix 🧠

Awesome License: CC BY-NC 4.0 Last Updated Star History Forks Issues Pull Requests Contributors

Research-based list of AI models, development tools, and automation resources. Use it to compare releases, pricing, benchmarks, and deployment options from official sources.

πŸš€ Quick Start

Looking for the best AI models? Check out our Top Models by Category table below for quick recommendations. For detailed analysis, explore the comprehensive sections on frontier models, open-source alternatives, and development tools.

πŸ”₯ Highlights

  • 138+ AI models from major providers (OpenAI, Anthropic, Google, Meta, xAI, Microsoft, and others)
  • Latest pricing and performance benchmarks (latest data through 2026-07-19 16:54 UTC)
  • GPT-5.6 reached general availability on 2026-07-09 across ChatGPT, Codex, and the OpenAI API; Sol, Terra, and Luna remain the three-tier family.
  • Grok-4.5 (xAI, 2026-07-08): new frontier multimodal LLM with a 500K context window, text+image input, function calling, structured outputs, and reasoning; API pricing $2.00 / $0.50 cached / $6.00 per 1M tokens. Now the xAI flagship route over Grok 4.3. Source: πŸ”—
  • Kimi K3 (Moonshot AI, 2026-07-16): new 2.8T-parameter open-weight frontier model with 1M context, native vision, $3.00 / $15.00 API pricing, and 90% cache-hit discount; open weights promised by 2026-07-27 under a modified MIT license. Source: πŸ”—
  • Meta Muse Spark 1.1 (Meta, 2026-07-09): new closed-weights frontier multimodal reasoning model for agentic tool and computer use, 1M context, available via Meta Model API in public preview for US developers. Source: πŸ”—
  • Inkling (Thinking Machines Lab, 2026-07-15): new 975B-parameter open-weight model with 1M context, native multimodal (text/image/audio), Apache 2.0 license, $1.87/$4.68 API pricing. Source: πŸ”—
  • Qwen 3.8 (Alibaba, 2026-07-19): new 2.4T-parameter open-weight model in preview on Token Plan/Qoder/QoderWork; claims to be "second only to Fable 5" on coding benchmarks. Source: πŸ”—
  • Codex CLI 0.144.5 (2026-07-16) improved dangerous-command detection and clearer rejection reasons; 0.144.4 (2026-07-14) had no user-facing changes; 0.144.3 (2026-07-13) was version-only.
  • Gemma 4 models were released through the Gemini API and AI Studio on 2026-07-06, including Gemma 4 31B and Gemma 4 26B-A4B.
  • Claude Sonnet 5 launched 2026-06-30 00:00 UTC with 1M context, 128K output, $2.00 / $10.00 introductory pricing through 2026-08-31 00:00 UTC, and 92.4% SWE-bench Verified
  • Claude Fable 5 and Mythos 5 access restored 2026-07-01 00:00 UTC after export controls were lifted; Fable 5 globally available while Mythos 5 remains restricted to Project Glasswing partners
  • Cartesia Sonic-3.5 and Ink-2 are now official production speech models for low-latency voice agents
  • MiniMax Speech 2.8 official page verified: native sound tags, high-fidelity cloning, and studio-grade multilingual TTS
  • GPT-Live (OpenAI, 2026-07-08): full-duplex voice model family powering ChatGPT Voice; GPT-Live-1 default for paid tiers, GPT-Live-1 mini for Free; no API yet (soon); GPT-5.5 backend delegates complex work. GPT-Realtime 2.1 / 2.1-mini shipped 2026-07-06 in the Realtime API with function calling, MCP, and SIP support
  • Cursor 3.11, Kiro 1.0.138, and Codex CLI 0.144.5 refreshed from official release notes
  • Codex CLI 0.142.5 patched WebSocket trace logging, and GitHub Copilot added enterprise agent session streaming plus GITHUB_TOKEN support for Copilot CLI in GitHub Actions
  • Browser Use CLI 3.0, Browser Harness, and Vercel agent-browser refreshed in Browser Automation from official project docs and releases
  • Vercel AI Gateway routing rules, eve Agent Runs in Vercel MCP/CLI, and Vercel Sandbox FUSE support added from official July 2026 changelogs
  • Cloudflare Browser Run /json Quick Action, Jina v5 Omni embeddings, and official MCP Registry references refreshed from official docs
  • MiniMax M2.7 self-evolving LLM with SWE-Pro 56.22%
  • MAI-1, MAI-1 Plan, and MAI-Frontier Tunings launched July 2026 with Microsoft Frontier Tuning framework reaching GA status
  • Cheapest verified direct APIs now tracked from official pricing pages: Groq Llama 3.1 8B Instant [$0.05/$0.08], Gemini 2.5 Flash-Lite [$0.10/$0.40], DeepSeek-V4-Flash [$0.14/$0.28], and OpenAI GPT-5.4 nano [$0.20/$1.25]
  • Google AI Plus price drop: $7.99 β†’ $4.99/month with storage doubled to 400GB (effective June 8, 2026)
  • Self-hosting guides for open-source models
  • Development tools for AI application building

πŸ“Š Top Models by Category

Category #1 #2 #3
Coding Claude Opus 4.8 GPT-5.5 Claude Sonnet 5
Reasoning Gemini 3.1 Pro GPT-5.5 Pro Qwen3.7-Max
Open Source Kimi K3 GLM-5.2 MiniMax M3
Cost Efficiency Llama 3.1 8B Instant Gemini 2.5 Flash-Lite DeepSeek-V4-Flash
Free & Budget Gemini 2.5 Flash-Lite Command A+ Mistral Small 4
Agentic Performance Claude Opus 4.8 GPT-5.5 Gemini 3.1 Pro
Context Window Gemini 3 Flash (10M) Llama 4 Scout (10M) Claude Opus 4.8 (1M)

πŸ“š Documentation


πŸ“‹ At a Glance

πŸ” Search & Navigation

  • Ctrl+F - Use your browser's search to find specific models or topics
  • Jump to sections using the table of contents above
  • Model search - Use your browser's search to find specific models (e.g., "GPT", "Claude", "Gemini")

Contents

Models 🧠

Comprehensive documentation of Large Language Models (LLMs), Small Language Models (SLMs), and specialized AI models available today.

Frontier Models πŸš€

State-of-the-art proprietary AI models with cutting-edge capabilities from leading AI labs.

Model Company Context GPQA Diamond Arena Elo SWE-bench Verified AIME 2025 Pricing Verified
Claude Opus 4.8 Anthropic 1M 93.6% 1506 88.6% β€” $5.00 / $25.00 2026-05-28
Claude Sonnet 5 ⭐ Anthropic 1M β€” β€” 92.4% β€” $2.00 / $10.00 (intro) 2026-06-30 00:00 UTC ⭐
Claude Opus 4.7 Anthropic 1M 94.2% 1505 87.6% ~95% $5.00 / $25.00 2026-04-26
GPT-5.5 OpenAI 1M 93.6% 1506 88.5% 99.9% $5.00 / $30.00 2026-04-26
GPT-5.5 Instant OpenAI 1M 91.0% β€” 88.7% β€” $0.20 / $1.00 2026-05-05
GPT-5.5 Pro OpenAI 1.05M 94.2% 1551 92.3% 100% $30.00 / $180.00 2026-05-08
Gemini 3.1 Pro Google 1M 94.3% 1505 80.6% 100% $2.00 / $12.00 2026-04-26
Claude Fable 5 ⚠️ Anthropic 1M 94.5% 1510 95.0% β€” $10.00 / $50.00 2026-06-09 ⭐
Claude Mythos 5 ⚠️ Anthropic 1M ~94.1% β€” 95.5% β€” $10.00 / $50.00 2026-06-09 ⭐
NVIDIA Nemotron 3 Ultra NVIDIA 1M 87.0% β€” β€” β€” Free (NVIDIA NIM) 2026-06-05
Kimi K3 ⭐ Moonshot AI 1M 93.5% β€” 88.3% (Terminal-Bench 2.1) β€” $3.00 / $15.00 2026-07-16 ⭐
Meta Muse Spark 1.1 ⭐ Meta 1M β€” β€” 61.5% (SWE-Bench Pro) β€” API-only (Meta Model API) 2026-07-09 ⭐
Inkling ⭐ Thinking Machines Lab 1M β€” β€” β€” β€” $1.87 / $4.68 2026-07-15 ⭐
Qwen 3.8 ⭐ Alibaba 1M β€” β€” β€” β€” API-only (preview) 2026-07-19 ⭐
Kimi K2.7 Code ⭐ Moonshot AI 262K 65.8% β€” 78.2% 91.5% $0.95 / $4.00 2026-06-12 ⭐
GLM-5.2 ⭐ Zhipu AI 1M ~91.2% β€” β€” 99.2% (AIME 2026) $1.40 / $4.40 2026-06-13 ⭐
MAI-Thinking-1 ⭐ Microsoft 256K 84.2% β€” β€” 97% TBD (Private Preview) 2026-06-02 ⭐
MAI-Code-1-Flash ⭐ Microsoft β€” β€” β€” β€” β€” Free (Copilot) 2026-06-02 ⭐
MiniMax M3 ⭐ MiniMax 1M β€” β€” 80.5% β€” $0.30 / $1.20 2026-06-01 ⭐
MiniMax M2.7 ⭐ MiniMax 205K β€” β€” 56.22% (SWE-Pro) β€” $0.30 / $1.20 2026-03-18 ⭐
Qwen3.7-Plus ⭐ Alibaba 1M β€” β€” β€” β€” $0.40 / $2.40 2026-06-03 ⭐
Step 3.7 Flash ⭐ StepFun 1M β€” β€” β€” β€” Free (limited) 2026-05-28 ⭐
GPT-5.6 Sol ⭐ OpenAI 1M+ β€” β€” β€” β€” $5.00 / $30.00 2026-07-09 ⭐
GPT-5.6 Terra ⭐ OpenAI 1M+ β€” β€” β€” β€” $2.50 / $15.00 2026-07-09 ⭐
GPT-5.6 Luna ⭐ OpenAI 1M+ β€” β€” β€” β€” $1.00 / $6.00 2026-07-09 ⭐
Grok-4.5 ⭐ xAI 500K β€” β€” β€” β€” $2.00 / $6.00 2026-07-08 ⭐

⭐ Kimi K3 (Moonshot AI, 2026-07-16): New 2.8T-parameter open-weight frontier model, 1M context, native vision, $3.00/$15.00 per 1M tokens with 90% cache-hit discount ($0.30 cached input). Open weights promised by 2026-07-27 under a modified MIT license. Source: πŸ”—

⭐ Meta Muse Spark 1.1 (Meta, 2026-07-09): New closed-weights frontier multimodal reasoning model built for agentic tool and computer use, 1M context, native multimodal perception, available via Meta Model API in public preview for US developers. Source: πŸ”—

⭐ Claude Fable 5 and Mythos 5 redeployed: Anthropic says export controls were lifted on 2026-06-30 00:00 UTC, and access to Fable 5 and Mythos 5 was restored on 2026-07-01 00:00 UTC. Fable 5 is available globally on Claude Platform, Claude.ai, Claude Code, and Claude Cowork, with cloud partner re-enablement in progress. Mythos 5 access is restored for approved US organizations and remains limited to Project Glasswing partners. Fable 5 has a 30-day data retention requirement, uses a different tokenizer (+30% token overhead vs Opus), and is unavailable for Zero Data Retention (ZDR) agreements. Source: πŸ”—

⭐ GPT-5.6 Sol/Terra/Luna (OpenAI, 2026-07-09): General availability release following the limited preview. Three-tier family: Sol ($5.00/$30.00), Terra ($2.50/$15.00), Luna ($1.00/$6.00) per 1M tokens. Source: πŸ”—

⭐ Grok-4.5 (xAI, 2026-07-08): New frontier multimodal LLM β€” the xAI flagship route over Grok 4.3. 500K context window, text+image input and text output, with function calling, structured outputs, and reasoning. API pricing $2.00 input / $0.50 cached input / $6.00 output per 1M tokens. Released 2026-07-08; available in Palantir AIP from 2026-07-14. Source: πŸ”—

⚠️ GPT-5.5 Instant pricing correction: The API pricing for GPT-5.5 Instant is $0.20/$1.00 per 1M tokens (input/output), not $5.00/$30.00 which is the standard GPT-5.5 rate. GPT-5.5 Instant replaced GPT-5.3 Instant as the free-tier ChatGPT default on May 5, 2026.

Model Specifications πŸ“‹

Detailed technical specifications, pricing, and capabilities for frontier models and notable API models. Data as of 2026-07-19 16:54 UTC.

Output Token Limits

Maximum output tokens per single API request.

Model Max Output Context Window Notes
Claude Opus 4.6 128K (300K via beta) 1M Extended output via output-128k-2025-02-19 beta header
Claude Opus 4.7 128K (300K via beta) 1M Extended output via output-128k-2025-02-19 beta header
Claude Sonnet 4.6 64K 1M β€”
Claude Sonnet 5 ⭐ 128K (300K via beta) 1M Released 2026-06-30 00:00 UTC; adaptive thinking on by default; batch API supports 300K with output-300k-2026-03-24 beta header
Claude Sonnet 4.5 64K 200K β€”
GPT-5.4 128K 1.05M β€”
GPT-5.4 mini 128K 400K β€”
GPT-5.4 nano 128K 400K β€”
GPT-5.3-Codex 128K 400K β€”
GPT-5.5 Instant 128K 1M Default ChatGPT model since 2026-05-05
Gemini 3.1 Pro 64K 1M β€”
Gemini 3.5 Flash 64K 1M Released 2026-05-19 at Google I/O 2026; 4x faster than 3.1 Pro
Gemini 3 Pro 64K 2M β€”
Gemini 3 Flash 64K 1M β€”
Gemini 3.1 Flash-Lite 32K 1M β€”
DeepSeek-V4-Flash 384K 1M API model ID deepseek-v4-flash
DeepSeek-V4-Pro 384K 1M API model ID deepseek-v4-pro; 75% discount made permanent 2026-05-31; standard price $0.435 / $0.87
DeepSeek-V3.2 8K / 64K (reasoner) 128K Reasoner mode unlocks 64K output
Qwen3.5-Max 65K 1M β€”
Qwen3.7-Max 65K 1M Extended thinking by default; proprietary, API-only
GLM-5 128K 200K β€”
GLM-5.1 131K 200K β€”
MiniMax M3 512K 1M (guaranteed 512K min) Native multimodal (text/image/video); open weights pending ~2026-06-11
MiniMax M2.7 262K 205K Self-evolution capabilities; Agent Teams native; SWE-Pro 56.22%
MiniMax-M2.5 131K 1M β€”
Kimi K2.6 262K 262K Max output = 262,144 tokens (MIT open-weight)
Kimi K2.7 Code 262K 262K Coding-focused variant of K2.6; 30% fewer thinking tokens; Modified MIT
Kimi K3 ⭐ 128K 1M 2.8T MoE (16 of 896 experts active); open weights promised by 2026-07-27; $3.00 / $15.00
Meta Muse Spark 1.1 ⭐ β€” 1M Closed-weights frontier multimodal reasoning model; agentic tool and computer use; API-only (Meta Model API)
Inkling ⭐ β€” 1M 975B/41B active MoE; Apache 2.0; multimodal (text/image/audio); released 2026-07-15
Qwen 3.8 ⭐ β€” 1M 2.4T parameters; open-weight promised; preview on Token Plan/Qoder/QoderWork; released 2026-07-19
Step-3.5-Flash 66K 256K β€”
Grok 4 β€” 256K Not publicly specified
Grok 4 Fast 30K 128K Now aliased to Grok 4.3; $0.20 / $0.50
Grok 4.20 30K 2M All variants: $1.25 / $2.50
Grok 4.3 30K 1M Always-on reasoning; released 2026-05-01
Grok-4.5 ⭐ β€” 500K Multimodal (text+image); function calling, structured outputs, reasoning; $2.00 / $6.00; released 2026-07-08
Mistral Small 4 8K 256K 119B MoE; multimodal (text+image); $0.15 / $0.60
Mistral Medium 3.5 8K 256K Released 2026-04-29; 128B dense; $1.50/$7.50
Mistral Large 3 8K 256K β€”
Llama 4 Scout 16K 10M β€”
Llama 4 Maverick 16K 1M β€”
Claude Fable 5 128K 1M Mythos-class; $10.00 / $50.00; released 2026-06-09
Claude Mythos 5 128K 1M Mythos-class (restricted); $10.00 / $50.00; released 2026-06-09
Kimi K3 ⭐ β€” 1M 2.8T MoE; open weights promised by 2026-07-27; $3.00 / $15.00; released 2026-07-16
Meta Muse Spark 1.1 ⭐ β€” 1M Closed-weights; agentic tool and computer use; API-only; released 2026-07-09
GPT-4.1 ⭐ 128K 1M Knowledge cutoff June 2024; $2.00 / $8.00; retired from ChatGPT Feb 13, 2026; API still available; released 2026-04
GPT-4.1 mini ⭐ 128K 1M Knowledge cutoff June 2024; $0.40 / $1.60; retired from ChatGPT Feb 13, 2026; API still available; released 2026-04
GPT-4.1 nano ⭐ 128K 1M Knowledge cutoff June 2024; $0.10 / $0.40; retired from ChatGPT Feb 13, 2026; API still available; released 2026-04
GPT-5.6 Sol ⭐ 128K 1M+ GA model; $5.00 / $30.00; released 2026-07-09
GPT-5.6 Terra ⭐ 128K 1M+ GA model; $2.50 / $15.00; released 2026-07-09
GPT-5.6 Luna ⭐ 128K 1M+ GA model; $1.00 / $6.00; released 2026-07-09
Gemini 3.1 Flash Image ⭐ β€” 131K Image + text + audio understanding; generation; $0.50/1M text+image input tokens; output images priced by resolution ($0.045/512px, $0.067/1Kpx, $0.101/2Kpx); GA June 2026

Cached & Batch Pricing

Discounted pricing tiers for high-volume usage. All prices in USD per million tokens.

Model Standard Input Cached Input Batch Discount Notes
Claude Opus 4.7 $5.00 $0.50 (hit) / $6.25 (5m write) 50% off Batch: $2.50 in / $12.50 out
Claude Opus 4.8 $5.00 $0.50 (hit) / $6.25 (5m write) 50% off Same pricing as 4.7; fast mode $10/$50
Claude Opus 4.6 $5.00 $0.50 (hit) / $6.25 (5m write) 50% off Batch: $2.50 in / $12.50 out
Claude Sonnet 4.6 $3.00 $0.30 (hit) / $3.75 (5m write) 50% off Batch: $1.50 in / $7.50 out
Claude Sonnet 5 ⭐ $2.00 (intro) / $3.00 standard $0.20 hit / $2.50 5m write (intro) 50% off Intro pricing through 2026-08-31 00:00 UTC: $2.00 / $10.00; standard from 2026-09-01 00:00 UTC: $3.00 / $15.00
Claude Sonnet 4.5 $3.00 $0.30 (hit) / $3.75 (5m write) 50% off Batch: $1.50 in / $7.50 out
GPT-5.4 $2.50 $0.25 50% off Data residency +10%
GPT-5.4 mini $0.75 $0.075 50% off β€”
GPT-5.4 nano $0.20 $0.02 50% off β€”
GPT-5.3-Codex $1.75 $0.175 50% off β€”
GPT-5.5 Instant $0.20 $0.02 50% off Free-tier ChatGPT default since 2026-05-05; API pricing $0.20/$1.00
Gemini 3.1 Pro $2.00 $0.20–$0.40 + $4.50/hr storage 50% off Tiered by input length
Gemini 3.5 Flash $1.50 $0.15 50% off Released 2026-05-19; 4x faster than 3.1 Pro
Gemini 3 Flash $0.50 $0.05 + $1.00/hr storage 50% off β€”
Gemini 3.1 Flash-Lite $0.25 $0.025 + $0.25/hr storage 50% off Cached input $0.025/1M (Vertex standard tier, ≀200K ctx)
DeepSeek-V4-Flash $0.14 $0.0028 (hit) β€” Cache-hit pricing from 2026-04-26 12:15 UTC
DeepSeek-V4-Pro $0.435 $0.003625 (hit) β€” 75% discount made permanent (2026-05-31); list price now $0.435 / $0.87
DeepSeek-V3.2 $0.28 $0.028 β€” No formal batch API
Qwen3.5-Max $0.40 Available 50% off β€”
Qwen3.7-Max $2.50 $0.25 β€” 90% cached discount; proprietary API-only
GLM-5 / GLM-5.1 $1.00 $0.20 β€” β€”
Grok 4 $3.00 $0.75 β€” β€”
Grok 4 Fast $0.20 $0.05 β€” Aliased to Grok 4.3; output $0.50/1M
Grok 4.20 $1.25 $0.20 β€” All variants: reasoning, non-reasoning, multi-agent
Grok 4.3 $1.25 $0.20 β€” ~38% cheaper input than 4.20; released 2026-05-01
Grok-4.5 ⭐ $2.00 $0.50 (hit) β€” Multimodal text+image; function calling, structured outputs, reasoning; released 2026-07-08
Mistral Small 4 $0.15 $0.015 (estimated) β€” 119B MoE; multimodal; released 2026-03-16
Command A+ ⭐ $0.40 β€” β€” 25B active (218B total MoE); faster than GPT-5.4 nano, Claude Haiku, Grok 4.3; released 2026-05
Mistral Large 3 $0.50 $0.05 (cached) β€” Cached input $0.05/1M; 256K context
Step-3.5-Flash $0.10 β€” β€” β€”
Claude Fable 5 $10.00 $1.00 (hit) / $12.50 (5m write) 50% off Batch: $5.00 in / $25.00 out; fast mode $20/$100
Claude Mythos 5 $10.00 $1.00 (hit) / $12.50 (5m write) 50% off Same pricing as Fable 5; restricted access via Project Glasswing
MiniMax M2.7 $0.30 β€” β€” Self-evolution capabilities; Agent Teams native; SWE-Pro 56.22%; 205K context; released 2026-03-18
MiniMax M3 (≀512K) $0.30 $0.06 (hit) β€” Extended context (>512K): $0.60/$2.40; released 2026-06-01
Kimi K3 ⭐ $3.00 $0.30 (hit, 90% cached) β€” 2.8T MoE; 1M context; open weights promised by 2026-07-27; released 2026-07-16
Meta Muse Spark 1.1 ⭐ API-only β€” β€” Closed-weights; 1M context; agentic tool and computer use; released 2026-07-09
Inkling ⭐ $1.87 β€” β€” 975B/41B active MoE; 1M context; Apache 2.0; multimodal; released 2026-07-15
Qwen 3.8 ⭐ API-only β€” β€” 2.4T parameters; open-weight promised; preview on Token Plan/Qoder/QoderWork; released 2026-07-19
Qwen3.7-Plus $0.40 β€” β€” Released 2026-06-03; 1M context
Step 3.7 Flash Free β€” β€” Free tier available; 1M context; released 2026-05-28
NVIDIA Nemotron 3 Ultra Free β€” β€” Free via NVIDIA NIM; 1M context (NVFP4 Blackwell)
GPT-4.1 ⭐ $2.00 $0.20 β€” 1M context; knowledge cutoff June 2024; retired from ChatGPT Feb 13, 2026; API still available
GPT-4.1 mini ⭐ $0.40 $0.04 β€” 1M context; retired from ChatGPT Feb 13, 2026; API still available
GPT-4.1 nano ⭐ $0.10 $0.01 β€” 1M context; retired from ChatGPT Feb 13, 2026; API still available; OpenAI's cheapest model
GPT-5.6 Sol ⭐ $5.00 $0.50 (90% cached) / $6.25 (cache write) 50% off GA; cache write 1.25x input; cached read 90% discount; released 2026-07-09
GPT-5.6 Terra ⭐ $2.50 $0.25 (90% cached) / $3.125 (cache write) 50% off GA; cache write 1.25x input; cached read 90% discount; released 2026-07-09
GPT-5.6 Luna ⭐ $1.00 $0.10 (90% cached) / $1.25 (cache write) 50% off GA; cache write 1.25x input; cached read 90% discount; released 2026-07-09

Speed & Latency

Output throughput and time-to-first-token from Artificial Analysis and provider benchmarks.

Model Output Speed (tok/s) TTFT Notes
Gemini 3.1 Flash-Lite ~250 ~2.1s Fastest budget Google model
Step-3.5-Flash 85–350 β€” Variable by provider; peak ~350 tok/s
Gemini 3 Flash ~193 ~4.16s β€”
MiniMax-M2.5 Lightning ~100 β€” Faster tier
GPT-5.3-Codex ~86 ~77.86s High TTFT due to extended reasoning
Grok 4 ~56 ~8.96s β€”
MiniMax-M2.5 Standard ~50 β€” β€”

Most frontier models (Claude Opus/Sonnet 4.6, GPT-5.4, Gemini 3.1 Pro, etc.) have not yet been benchmarked on Artificial Analysis as of April 2026. Anthropic notes that fast mode on Claude Opus 4.7 is deprecated and scheduled for removal on 2026-07-24 00:00 UTC; use Claude Opus 4.8 for current Opus fast-mode workflows. Source: πŸ”—

Training Data Cutoffs

Knowledge cutoff dates β€” the point after which a model has no training data.

Model Training Cutoff Notes
Claude Sonnet 4.6 Jan 2026 Most recent cutoff among frontier models
Claude Sonnet 5 ⭐ Jan 2026 Reliable knowledge cutoff Jan 2026; released 2026-06-30 00:00 UTC
Claude Opus 4.6 Aug 2025 Reliable knowledge: May 2025
GPT-5.4 / mini / nano Aug 31, 2025 β€”
GPT-5.3-Codex Aug 31, 2025 β€”
Grok 4 Fast Jul 2025 β€”
Grok 4.20 / 4.3 Jul 2025 Approximate; not publicly disclosed
Grok-4.5 ⭐ ~Nov 2024 Same cutoff family as Grok 4; approximate; not separately disclosed
DeepSeek-V4 (Flash/Pro) May 2025 β€”
Gemini 3.1 Flash-Lite Jan 2025 β€”
Gemini 3.1 Pro / 3 Pro / 3 Flash Jan 2025 β€”
Grok 4 ~Nov–Dec 2024 Approximate
DeepSeek-V3.2 Jul 2024 β€”
Llama 4 Scout / Maverick Aug 2024 β€”
DeepSeek-R1 ~Oct 2023 Based on base model
GPT-4.1 / mini / nano June 2024 Retired from ChatGPT Feb 13, 2026; API still available
GPT-4.5 June 2024 Retired from API July 2025; retiring from ChatGPT 2026-06-27
Kimi K3 ⭐ ~Jul 2025 Approximate; not separately disclosed; released 2026-07-16
Meta Muse Spark 1.1 ⭐ ~Jan 2026 Approximate; not separately disclosed; released 2026-07-09
Inkling ⭐ ~Jan 2026 Approximate; not separately disclosed; released 2026-07-15
Qwen 3.8 ⭐ ~Jan 2026 Approximate; not separately disclosed; released 2026-07-19

Models not listed (Qwen, GLM, MiniMax, Kimi, Step, Mistral): training cutoff not publicly disclosed.

Multilingual Support

Model Languages Details
Qwen3.5-Max 201 Largest language coverage
Llama 4 Scout 200 Pre-training languages
Qwen3-Max-Thinking 119 Qwen3 series
Gemini 3 Flash 100 91.8% MMMLU score across 100 languages
Gemini 3.1 Pro / 3 Pro 100+ β€”
Gemini 3.1 Flash-Lite 100 91.3% MMMLU score
Llama 4 Maverick 12 Output languages
Claude (all) Many English-optimized; broad multilingual; Sonnet 5 supports multilingual capabilities and vision
GPT-5.4 (all) Many Broad multilingual coverage
DeepSeek (all) Many Chinese + English focused
Grok (all) Many β€”
GLM-5 / GLM-5.1 Many 28.5T token training data
Inkling ⭐ Many Multimodal (text/image/audio); released 2026-07-15
Qwen 3.8 ⭐ Many Multilingual support; released 2026-07-19

Structured Output & Function Calling

All frontier models support structured JSON output and function/tool calling except where noted.

Capability Supported Models Not Supported
Structured Output (JSON mode) All models listed in Frontier table Gemini 3 Deep Think (no API)
Function Calling / Tool Use All models listed in Frontier table Gemini 3 Deep Think (no API)

Gemini 3 Deep Think is available only via Gemini's in-app Think mode β€” no API access for structured output or function calling.

Regional Availability

Provider API Availability Cloud Partners Notes
Anthropic Global AWS Bedrock, GCP Vertex AI US-only inference at 1.1x via inference_geo
OpenAI Global Azure OpenAI Data residency endpoints +10% (post-3/5/26)
Google Global Google AI Studio, Vertex AI Some regional restrictions per Google terms
DeepSeek Global Azure (R1 only, select regions) China-based servers
Alibaba (Qwen) Global Alibaba Cloud Model Studio China-based; globally accessible
Zhipu AI (GLM) Global Z.AI API MIT license enables self-hosting anywhere
MiniMax Global MiniMax API β€”
Moonshot AI (Kimi) Global platform.kimi.ai MIT open-weight
xAI (Grok 4 / 4.20 / 4.3 / 4.5 / 4 Fast) US-focused Oracle OCI (East/Midwest/West), Amazon Bedrock (June 2026) Major price cuts: Grok 4.20 now $1.25/$2.50; Grok 4 Fast aliased to 4.3; Grok 4.3 on Bedrock at $1.25/$2.50; Grok-4.5 flagship at $2.00/$6.00 (500K context, text+image)
Mistral Global Azure AI Foundry, AWS, GCP β€”
Meta (Llama) Global (self-host) All major cloud providers Llama 4 Community License
Meta (Muse Spark 1.1) US-only (public preview) Meta Model API Closed-weights frontier multimodal reasoning model; public preview for US developers; released 2026-07-09
Thinking Machines Lab (Inkling) US-focused Inkling API Open-weight model (Apache 2.0); 975B/41B active MoE; 1M context; released 2026-07-15
Alibaba (Qwen 3.8) Global (preview) Token Plan / Qoder / QoderWork Open-weight model; 2.4T parameters; preview access; released 2026-07-19
StepFun Global HuggingFace Apache 2.0 open-source

Free-Source Models πŸ†“

Self-hostable models with permissive licenses or open weights for privacy, cost control, and customization.

Model Company Params Context License
DeepSeek-V4-Flash DeepSeek 1.6T / 49B active (MoE) 1M Open Weight
Qwen3.5-Max Alibaba 397B / 17B active (MoE) 262K Apache 2.0
Qwen3-Max-Thinking Alibaba 1T+ 128K Apache 2.0
Qwen3.6-27B Alibaba 27B dense 262K Apache 2.0
Qwen3.5-122B Alibaba 397B / 17B active (MoE) 262K Apache 2.0
Qwen3.5-27B Alibaba 27B dense 262K Apache 2.0
Mistral Small 4 Mistral AI 119B / 6B active (MoE, 128 experts) 256K Apache 2.0
Mistral Large 3 Mistral AI 675B / 41B active (MoE) 256K Apache 2.0
Command A+ Cohere 25B active (218B total MoE) 128K Apache 2.0
Llama 4 Scout Meta 109B 10M Community
Llama 4 Maverick Meta 400B 128K Community
GPT-OSS-120B OpenAI 117B 128K Apache 2.0
GPT-OSS-20B OpenAI 21B 128K Apache 2.0
Qwen3-Coder Alibaba 480B 262K Apache 2.0
GLM-5.1 Zhipu AI 754B / 40B active (MoE) 200K MIT
GLM-4.7 Zhipu AI 400B+ MoE 205K Open Weight
Gemma 4 31B Google 31B dense 256K Apache 2.0
Gemma 4 26B-A4B Google 26B MoE (4B active) 256K Apache 2.0
Gemma 4 E4B Google 4B dense 256K Apache 2.0
Gemma 4 E2B Google 2B dense 256K Apache 2.0
Gemma 4 12B Google 12B dense 256K Apache 2.0
Qwen3-Coder 7B Alibaba 7B dense 128K Apache 2.0
Qwen 2.5 Coder 32B Alibaba 32B dense 128K Apache 2.0
DeepSeek Coder-V2 DeepSeek 236B / 2.4B active 128K MIT
Step-3.5-flash StepFun 196B / 11B active (MoE) 256K Open Weight
Yi-Coder 01.AI 9B/1.5B 128K Apache 2.0
Lizzy-7B Flower Labs 7B β€” MIT
MiMo-V2.5 Xiaomi 309B / 15B active 1M MIT
MiMo-V2.5-Pro Xiaomi 1.02T / 42B active 1M MIT
NVIDIA Nemotron 3 Ultra NVIDIA 550B / 55B active (MoE) 262K (BF16) / 1M (NVFP4 Blackwell) Open Weight (NVIDIA)
MiniMax M3 ⭐ MiniMax ~428B / ~23B active (MoE) 1M MiniMax Community
MiniMax M2 ⭐ MiniMax 230B / 10B active (MoE) 1M Modified MIT
MiniMax M2.1 ⭐ MiniMax 229B / 10B active (MoE) 1M Apache 2.0
GLM-5.2 ⭐ Zhipu AI 753B / 40B active (MoE) 1M MIT
Kimi K2.7 Code ⭐ Moonshot AI 1T / 32B active (MoE) 262K Modified MIT
MiniMax M2.7 ⭐ MiniMax ~400B+ (MoE) 205K MiniMax Community
Kimi K2.6 Moonshot AI 1T / 32B active (MoE) 262K Modified MIT
Kimi K3 ⭐ Moonshot AI 2.8T / 16 active (MoE) 1M Modified MIT
LongCat-2.0 ⭐ Meituan 1.6T / ~48B active (MoE) 1M MIT
Hy3 ⭐ Tencent 295B / 21B active (MoE) 256K Apache 2.0
Ornith-1.0 ⭐ DeepReinforce 9B / 31B / 35B-MoE / 397B-MoE β€” MIT

Deployment Options

Local Inference Tools:

  • Ollama - Easy local deployment
  • LM Studio - User-friendly GUI
  • llama.cpp - Efficient CPU inference
  • vLLM - High-throughput serving
  • SGLang - Structured generation

Cloud Deployment:

  • Hugging Face Inference - Managed deployment
  • AWS SageMaker - Full control
  • Google Cloud Vertex - Integrated
  • RunPod - GPU rental

Coding Models πŸ’»

Specialized AI models optimized for software development tasks.

SWE-bench Verified Leaderboard

Rank Model Company SWE-bench Verified
πŸ₯‡ #1 Claude Mythos 5 ⚠️ Anthropic 95.5%
πŸ₯ˆ #2 Claude Fable 5 ⚠️ Anthropic 95.0%
πŸ₯‰ #3 Claude Mythos Preview Anthropic 93.9%
#4 Claude Sonnet 5 ⭐ Anthropic 92.4%
#5 GPT-5.5 Pro OpenAI 92.3%
#6 GPT-5.5 OpenAI 88.7%
#7 Claude Opus 4.8 Anthropic 88.6%
#8 Claude Opus 4.7 Anthropic 87.6%
#9 GPT-5.3-Codex OpenAI 85.0%
#10 Ornith-1.0-397B ⭐ DeepReinforce 82.4%
#11 Claude Opus 4.5 Anthropic 80.9%
#12 Claude Opus 4.6 Anthropic 80.8%
#13 DeepSeek-V4-Pro (Max) DeepSeek 80.6%
#13 Gemini 3.1 Pro Google 80.6%
#14 MiniMax M3 ⭐ MiniMax 80.5%
#14 Qwen3.7-Max Alibaba 80.4%
#15 Kimi K2.6 Moonshot AI 80.2%
#16 MiniMax-M2.5 MiniMax 80.2%
#17 GPT-5.2 OpenAI 80.0%
#18 Claude Sonnet 4.6 Anthropic 79.6%
#19 DeepSeek-V4-Flash (Max) DeepSeek 79.0%
#20 Qwen3.6 Plus Alibaba 78.8%
#21 Kimi K2.7 Code ⭐ Moonshot AI 78.2%
#22 Gemini 3 Flash Google 78.0%
#22 MiMo-V2-Pro Xiaomi 78.0%
#24 GLM-5 Zhipu AI 77.8%
#25 Mistral Medium 3.5 Mistral AI 77.6%
#26 Claude Sonnet 4.5 Anthropic 77.2%

Commercial Coding Models

Model Developer Pricing Best For
Claude Fable 5 ⚠️ ⭐ Anthropic $10.00 / $50.00 per 1M Mythos-class coding; SWE-bench 95.0%; restored globally 2026-07-01 00:00 UTC with stronger safeguards
Claude Opus 4.8 Anthropic $5.00 / $25.00 per 1M Agentic coding, complex tasks, best available Opus
Claude Sonnet 5 ⭐ Anthropic $2.00 / $10.00 per 1M intro through 2026-08-31 00:00 UTC Agentic coding at Sonnet latency and price; 92.4% SWE-bench Verified; 63.2% SWE-bench Pro
Kimi K2.7 Code ⭐ Moonshot AI $0.95 / $4.00 per 1M Open-weight coding agent; 30% fewer thinking tokens vs K2.6; now available in GitHub Copilot model picker (July 1, 2026)
Claude Opus 4.6 Anthropic $5.00 / $25.00 per 1M Agentic coding, complex tasks
GPT-5.5 Pro OpenAI $30.00 / $180.00 per 1M Highest benchmark coding
GPT-5.3-Codex OpenAI $1.75 / $14.00 per 1M Agentic coding, 7+ hour autonomy
Claude Haiku 4.5 Anthropic $1.00 / $5.00 per 1M Low-latency coding, sub-agents, computer use
GLM-5-Code Zhipu AI $1.20 / $5.00 per 1M Code generation, refactoring
MiniMax M3 MiniMax $0.30 / $1.20 per 1M Frontier coding + 1M context + native multimodal; open weights pending
MiniMax M2 ⭐ MiniMax $0.30 / $1.20 per 1M Open-weight MoE (230B/10B active); 1M context; 10x cheaper than Sonnet; released 2026-07-02
MiniMax M2.1 ⭐ MiniMax $0.30 / $1.20 per 1M Open-weight MoE (229B/10B active); 1M context; Apache 2.0; multilingual coding; released 2026-07-02
MiniMax M2.7 ⭐ MiniMax $0.30 / $1.20 per 1M Self-evolution; Agent Teams native; SWE-Pro 56.22%; 205K context; released 2026-03-18
MiniMax-M2.5 MiniMax $0.30 / $1.20 per 1M Code generation, refactoring
Claude Sonnet 4.5 Anthropic $3.00 / $15.00 per 1M Code review, refactoring
Mistral Small 4 Mistral AI $0.15 / $0.60 per 1M Unified reasoning + coding + multimodal; open-weight
Codestral Mistral AI $0.30 / $0.90 Real-time completion
North Mini Code Cohere Free (OpenRouter) / API 30B MoE (3B active); Apache 2.0; released 2026-06-09
Qwen3.7-Max Alibaba $2.50 / $7.50 per 1M Agentic coding, long-horizon autonomy, reasoning
Grok 4.3 xAI $1.25 / $2.50 per 1M Flagship reasoning, lowest hallucination
MAI-Code-1-Flash Microsoft $0.75 / $4.50 per 1M GitHub Copilot integration, fast everyday coding

Open-Source Coding Models

Model Developer License Hardware
Kimi K2.7 Code ⭐ Moonshot AI Modified MIT 80-160 GB VRAM
Kimi K3 ⭐ Moonshot AI Modified MIT 160-320 GB VRAM
LongCat-2.0 ⭐ Meituan MIT 320+ GB VRAM
GLM-5.2 ⭐ Zhipu AI MIT 80-160 GB VRAM
MiniMax M2 ⭐ MiniMax Apache 2.0 48-80 GB VRAM
MiniMax M2.1 ⭐ MiniMax Apache 2.0 48-80 GB VRAM
MiniMax M3 ⭐ MiniMax MIT 80-160 GB VRAM
MiniMax M2.7 ⭐ MiniMax MiniMax Community 80-160 GB VRAM
Ornith-1.0 ⭐ DeepReinforce MIT 160-320 GB VRAM
GPT-OSS-120B OpenAI Apache 2.0 80-160 GB VRAM
Qwen3-Coder Alibaba Apache 2.0 160-320 GB VRAM
DeepSeek-Coder-V2 DeepSeek MIT 48-80 GB VRAM
GLM-5.1 Zhipu AI MIT 80-160 GB VRAM
Phi-4 Microsoft MIT 24-48 GB VRAM
Qwen3-Coder 7B Alibaba Apache 2.0 24-48 GB VRAM

Reasoning Models 🧠

Models optimized for step-by-step reasoning, mathematical problem-solving, and complex logical inference.

AIME 2025 Leaderboard

Rank Model AIME 2025 ARC-AGI-2 Notes
πŸ₯‡ #1 MAI-Thinking-1 ⭐ 97% β€” Microsoft's first in-house reasoning model; private preview
#2 Kimi K2.5 (Reasoning) 96.1% β€” Moonshot AI reasoning variant
#2 Kimi K2.5 96.1% β€” Open-weight reasoning
#4 GLM-4.7 95.7% β€” Zhipu AI open-weight
#5 MiMo-V2-Flash 94.1% β€” Xiaomi open-weight
#6 Claude Sonnet 4.5 87% β€” β€”
#7 Exaone 4.0 32B 85.3% β€” LG AI Research
#8 Nemotron 3 Nano Omni 30B 82.1% β€” NVIDIA
#9 LFM2.5-8B-A1B 42.5% β€” LiquidAI
#10 MiniCPM5-1B 40.4% β€” OpenBMB
#12 Step-3.5-Flash 97.3% β€” Best efficiency ratio
#12 GLM-5.2 ⭐ 99.2% (AIME 2026) β€” Zhipu AI open-weight; leads AIME 2026 leaderboard
#13 Kimi K2.6 96.4% β€” Strong multimodal reasoning
#14 GLM-4.7 95.7% β€” Zhipu AI strong performer
#15 Claude Sonnet 4.6 ~95%–95.6% 58.3% Near-Opus performance
#16 GLM-5 92.7% β€” Thinking mode
#17 o4-mini 92.7% β€” Efficient OpenAI reasoning

Reasoning Model Details

Model Type Context Pricing
Gemini 3 Deep Think Reasoning 1M+ Ultra subscription
Qwen3-Max-Thinking Reasoning/Coding 128K $1.20 / $6.00
o3 / o1-Pro Reasoning 128K $2-150 / $8-600
GPT-5.5 Pro Reasoning 1.05M $30.00 / $180.00
Gemini 3 Pro General/Multimodal 1M+ $2.00 / $12.00
DeepSeek-R1 Reasoning 128K $0.50 / $2.15
Claude Sonnet 4.5 Hybrid 200K $3.00 / $15.00
Claude Sonnet 5 ⭐ Hybrid 1M $2.00 / $10.00 intro; $3.00 / $15.00 standard
GPT-Rosalind Life Sciences Reasoning 128K Pay-per-token (Research Preview)
MAI-Thinking-1 Reasoning 256K TBD (Private Preview)

Use Cases

  • Mathematical Problem Solving: Qwen3-Max-Thinking, GPT-5.5 Pro, Gemini 3 Pro
  • Scientific Analysis: Claude Sonnet 5, Claude Opus 4.6, GPT-5.5, Gemini 3 Pro
  • Strategic Planning: o3/o1-Pro, Claude Sonnet 4.5, DeepSeek-R1
  • Code Debugging: Claude Sonnet 5, Claude Sonnet 4.5, GPT-5.3-Codex, DeepSeek-V3.2

Multimodal Models 🎨

Models capable of processing and generating multiple types of content: text, images, audio, and video.

Leading Multimodal Models

Model Developer Context Key Features
GPT-5.4 OpenAI 1M Unified multimodal, audio
Gemini 3 Pro Google 1M+ Native multimodal, video
Claude Sonnet 4.5 Anthropic 200K Document understanding
Claude Sonnet 5 ⭐ Anthropic 1M Text and image input, 128K output, agentic coding and tool use
Llama 4 Maverick Meta 128K Open multimodal
MiniMax M3 MiniMax 1M First open-weight model: frontier coding + 1M context + native image & video understanding
Kimi K3 ⭐ Moonshot AI 1M Open-weight frontier model: 2.8T MoE, native vision, $3.00/$15.00 API pricing; released 2026-07-16
Meta Muse Spark 1.1 ⭐ Meta 1M Closed-weights frontier multimodal reasoning model; agentic tool and computer use; API-only; released 2026-07-09
Gemini Omni Flash Google β€” Native multimodal (text+image+audio+video input, video generation), any-to-any creation
Nemotron 3 Nano Omni NVIDIA 30B (3B active) Vision, audio, language unified, 9x throughput

Vision Capabilities

Model MMMU / MMMU-Pro MathVista DocVQA
Gemini 3.1 Pro 95% (MMMU-Pro) β€” β€”
GPT-5.4 94% (MMMU-Pro) β€” β€”
Gemini 3 Pro 81% (MMMU-Pro) β€” β€”
Gemini 3 Flash 80% (MMMU-Pro) β€” β€”
Claude Sonnet 4.5 77.8% (MMMU) β€” β€”
Llama 4 Maverick 73.4% (MMMU) β€” β€”

Audio & Video

Model Speech-to-Text Text-to-Speech Video Input
Gemini 3 Pro βœ… βœ… βœ…
GPT-5 βœ… βœ… ⚠️
Whisper v3 βœ… ❌ βœ…

Image Generation

Model Developer License Best For
Gemini 3.1 Flash-Lite Image ⭐ Google Proprietary Gemini image model card updated 2026-06-30 00:00 UTC; image + text + audio understanding and generation; output images priced by resolution
Nano Banana 2 Lite ⭐ Google Proprietary Fastest Gemini Image model; 4s generation; GA 2026-06-30; $0.034 per 1K images; rolling out to Gemini app, Search, NotebookLM, Photos, Stitch, Flow, Ads
GPT Image 2 OpenAI Proprietary ChatGPT Images 2.0; improved text rendering, multilingual support, advanced visual reasoning; released 2026-04-21; API pricing varies by quality/resolution
MAI-Image-2-Efficient Microsoft Proprietary Production-ready quality, 41% lower cost
Flux 2 Black Forest Labs Apache 2.0 (Dev); Proprietary (Pro) Released Nov 2025; exceptional photorealism + natural language; Flux 2 Pro/Dev/Schnell variants
Flux.1 Black Forest Labs Apache 2.0 High-fidelity art (original Flux family)
Stable Diffusion 3.5 Stability AI Community License Fine-tuning; MMDiT architecture, 2.5B params
GLM-Image Zhipu AI (Z.ai) API Fast image generation
CogView-4 Zhipu AI (Z.ai) API Creative image generation
Firefly AI Assistant Adobe Public Beta (2026-04-27) Creative agent, 60+ tools, Photoshop/Premiere integration

Hardware Requirements πŸ–₯️

Comprehensive hardware specifications for self-hosting AI models.

Quick Reference by Model Size

Model Params Q4 Size Min VRAM Rec VRAM Min RAM
Phi-4 14B 8 GB 24 GB 48 GB 32 GB
GPT-OSS-20B 21B 12 GB 24 GB 48 GB 32 GB
Llama 4 Scout 109B 66 GB 48 GB 80 GB 96 GB
GPT-OSS-120B 117B 70 GB 80 GB 160 GB 128 GB
DeepSeek-Coder-V2 236B 143 GB 48 GB 80 GB 192 GB
Llama 4 Maverick 400B 242 GB 160 GB 320 GB 320 GB
DeepSeek-V4 family 671B 404 GB 80 GB 320 GB 512 GB
Qwen3-Max-Thinking 1T+ 600+ GB 160 GB 640 GB 768 GB

By Hardware Tier

Consumer/Entry Level (24-48 GB VRAM):

  • Phi-4, GPT-OSS-20B, Yi-Coder, Qwen2.5-Coder
  • Recommended GPUs: RTX 3090 (24GB), RTX 4090 (24GB)

Professional (80-160 GB VRAM):

  • Llama 4 Scout, GPT-OSS-120B, DeepSeek-Coder-V2
  • Recommended GPUs: A100 80GB, 2x A100 40GB

Enterprise (320+ GB VRAM):

  • Llama 4 Maverick, GLM-4.7, DeepSeek-V4 family, Qwen3-Max-Thinking
  • Recommended GPUs: 4x A100 80GB, 8x A100 80GB

Quantization Explained

Level Bits Size vs FP16 Quality Use Case
FP16/BF16 16 100% Best Training
Q8_0 8 ~50% Excellent High-quality inference
Q4_K_M 4 ~25% Good Recommended for deployment
Q3_K_M 3 ~19% Fair Limited resources

Comprehensive Benchmark Reference πŸ“ˆ

Detailed benchmark scores across all major evaluations. Scores are percentages (%) unless noted. Arena Elo scores are integers. β€” = not publicly reported. Data as of 2026-07-19 16:54 UTC.

Full Benchmark Table

Model GPQA Diamond MMLU-Pro Arena Elo (Text) HLE SWE-bench Verified SWE-bench Pro LiveCodeBench AIME 2025 ARC-AGI-2 MMMU-Pro IFEval FrontierMath
Claude Opus 4.6 91.3% β€” 1500 40.0–53.0% 80.8% β€” β€” 99.8% 68.8% β€” β€” β€”
Claude Opus 4.8 93.6% β€” β€” β€” 88.6% 69.2% β€” β€” β€” β€” β€” β€”
Claude Mythos Preview 94.5% β€” β€” β€” 93.9% β€” β€” β€” β€” β€” β€” β€”
GPT-5.5 93.6% β€” 1495 42.1–55.0% 88.5% β€” β€” 99.9% 71.2% β€” β€” 52%
GPT-5.5 Instant 91.0% β€” β€” β€” 88.7% β€” β€” β€” β€” β€” β€” β€”
NVIDIA Nemotron 3 Ultra β€” β€” β€” β€” β€” β€” β€” β€” β€” β€” β€” β€”
Ornith-1.0-397B ⭐ β€” β€” β€” β€” 82.4% β€” β€” β€” β€” β€” β€” β€”
GPT-5.5 Pro 95.1% 96% 1520 48.5–62.0% 92.3% β€” β€” 100% 78.5% β€” 97% 58%
Claude Sonnet 4.6 89.9% β€” ~1438 33.2–49.0% 79.6% β€” β€” ~95% 58.3% β€” β€” β€”
Claude Sonnet 5 ⭐ β€” β€” β€” 43.2% no tools / 57.4% tools 92.4% 63.2% β€” β€” β€” β€” β€” β€”
Claude Sonnet 4.5 83.4% 88.0% β€” β€” 77.2% β€” β€” 87–100% β€” β€” β€” β€”
Claude Fable 5 94.5% β€” β€” β€” 95.0% 80.3% 29.3% (FrontierCode) β€” β€” β€” β€” β€”
Claude Mythos 5 ~94.1% β€” β€” β€” 95.5% β€” β€” β€” β€” β€” β€” β€”
Mistral Small 4 ~71.2% β€” β€” β€” β€” β€” β€” β€” β€” β€” β€” β€”
Mistral Medium 3.5 β€” β€” β€” β€” 77.6% β€” β€” 86.3% β€” β€” β€” β€”
GPT-5.4 92.0% 94% 1484 36.6–41.6% ~80% 57.7% 84–88% 88% 73.3% 94% β€” 50% (Pro)
GPT-5.4 mini 87.5% β€” β€” β€” β€” 54.4% β€” β€” β€” β€” β€” β€”
GPT-5.3-Codex 91.5% β€” β€” β€” β€” 56.8% 85% β€” β€” β€” β€” β€”
GPT-5.2 92.4% β€” 1479 35.2% 80.0% 55.6% β€” 100% 52.9% β€” 95.6% ~40.3%
Gemini 3.1 Pro 94.3% 92% 1494 44.4–51.4% 80.6% 54.2–72% 71% 100% 77.1% 95% 95% β€”
DeepSeek-V4-Pro (Max) 90.1% β€” β€” β€” 80.6% 55.4% β€” β€” β€” β€” β€” β€”
Gemini 3.5 Flash ~90.4% β€” β€” β€” 55.1% β€” β€” 72.1% 83.6% 84.2% 40.2% β€”
Gemini 3 Pro 91.9–93.8% 83% 1486 37.5% 76.2% 43.3% 49% 98–100% 31.1–45.1% 81% 88% 38%
Gemini 3 Flash 90.4% 72% 1474 33.7% 78.0% 44% β€” β€” β€” 80% 85% β€”
Gemini 3 Deep Think ~97% 81% β€” 48.4% ~58% 63% 58% β€” 84.6% β€” β€” β€”
DeepSeek-V3.2 87.1% 85.0% β€” 25.1% 67.8% β€” β€” 89.3% β€” β€” β€” β€”
DeepSeek-V4-Flash (Max) 88.1% β€” β€” β€” 79.0% 52.6% β€” β€” β€” β€” β€” β€”
DeepSeek-R1 71.5% 84.0% β€” 8.5% 49.2% β€” 63.5% 70.0% β€” β€” β€” β€”
Qwen3.5-Max 89.3% β€” β€” β€” 76.4% β€” β€” 91.3% β€” 79% β€” β€”
Qwen3.7-Max 92.4% β€” β€” 41.4% 80.4% β€” β€” β€” β€” β€” β€” β€”
Qwen3.6 Plus 90.4% β€” β€” β€” 78.8% β€” β€” β€” β€” β€” β€” β€”
Qwen3-Max-Thinking 86.1% β€” β€” 26.2% β€” β€” β€” β€” β€” β€” β€” β€”
GLM-5 82.0% β€” ~1451 10.4% 77.8% β€” β€” 92.7% β€” β€” β€” β€”
MiMo-V2-Pro 87.0% β€” β€” β€” 78.0% β€” β€” β€” β€” β€” β€” β€”
GLM-5.1 β€” β€” β€” β€” ~80.4% (est.) β€” β€” β€” β€” β€” β€” β€”
Kimi K2.6 90.5% 87.1% β€” 31.5–50.2% 80.2% β€” 85.0% 96.4% β€” 78.5% β€” β€”
Kimi K3 ⭐ 93.5% β€” β€” β€” 88.3% β€” β€” β€” β€” β€” β€” β€”
Meta Muse Spark 1.1 ⭐ β€” β€” β€” β€” 61.5% β€” β€” β€” β€” β€” β€” β€”
GLM-5.2 ⭐ ~91.2% β€” β€” β€” β€” 62.1% 76.8% (MCP-Atlas) 99.2% β€” β€” β€” β€”
MiniMax M3 ~92.9% β€” β€” β€” β€” β€” β€” β€” β€” β€” β€” β€”
MiniMax-M2.5 85.2% β€” β€” β€” 80.2% 55.4% β€” 86.3% β€” β€” β€” β€”
Step-3.5-Flash 83.1% β€” β€” β€” 74.4% β€” 86.4% 97.3% β€” β€” β€” β€”
Grok 4 ~91.5% 91.5% ~1493 50.7% β€” β€” β€” 100% β€” β€” β€” β€”
Grok 4.20 β€” β€” β€” β€” β€” β€” β€” β€” β€” β€” β€” β€”
Grok 4.3 β€” β€” β€” β€” β€” β€” β€” β€” β€” β€” β€” β€”
Grok-4.5 ⭐ β€” β€” β€” β€” β€” β€” β€” β€” β€” β€” β€” β€”
Llama 4 Maverick 69.8% 80.5% β€” β€” β€” β€” 43.4% β€” β€” β€” β€” β€”
Llama 4 Scout 57.2% 74.3% β€” β€” β€” β€” 32.8% β€” β€” β€” β€” β€”

FrontierMath Scores

FrontierMath is a benchmark of 350 original, exceptionally challenging mathematics problems created by expert mathematicians (Epoch AI). Problems span number theory, analysis, algebraic geometry, and category theory. Tier 4 problems can take research mathematicians multiple days.

Model Tiers 1–3 Tier 4 Source
GPT-5.4 Pro 50% ~36–38% Epoch AI
GPT-5.2 Pro ~40.3% 31% Epoch AI
Gemini 3 Pro 38% 19% Epoch AI
GPT-5.1 Thinking ~25% β€” llm-stats

Notes on Limited Availability Models

Claude Mythos Preview is not publicly available and is only accessible through Project Glasswing, an invitation-only partner program for cybersecurity applications. Benchmarks for this model are not publicly disclosed as of May 2026.

Benchmark Glossary

Benchmark Description Source
GPQA Diamond Graduate-level science questions (PhD difficulty) Google Research
MMLU-Pro Extended multi-task language understanding (harder than MMLU) TIGER-Lab
Arena Elo Crowdsourced human preference ranking lmarena.ai
HLE Humanity's Last Exam β€” expert-level questions Scale AI
SWE-bench Verified Real GitHub issue resolution (human-verified subset) SWE-bench
SWE-bench Pro More challenging subset of SWE-bench SWE-bench
LiveCodeBench Live competitive programming problems (not in training data) LiveCodeBench
AIME 2025 American Invitational Mathematics Examination MAA
ARC-AGI-2 Abstract reasoning challenge (fluid intelligence) ARC Prize
MMMU / MMMU-Pro Multi-discipline multimodal understanding MMMU
IFEval Instruction-following evaluation Google Research
FrontierMath Expert-level research mathematics (Epoch AI) Epoch AI

Development Tools πŸ› οΈ

AI-powered tools for software development, from IDEs and CLI tools to API providers and IDE extensions.

IDEs πŸ’»

Integrated Development Environments with built-in AI capabilities.

Agentic IDEs

IDE Platform Version Release Date Pricing Key Features GitHub
Firebase Studio Web - - Free (3 workspaces, up to 30 with Google Developer Program) Cloud-based, Gemini, MCP πŸ”—
Lingma IDE (ι€šδΉ‰η΅η ) Windows, macOS - - Free (download) Built-in agent, MCP tool use, terminal command execution ❌
Tonkotsu Windows, macOS - - Free (during early access) Team of agents, workflow πŸ”—
OpenCode Windows, macOS, Linux - - Free (OSS) Terminal, desktop, IDE extension, multi-provider πŸ”—
Codex app Windows - 2026-03-04 00:00 UTC Included with Codex plans Multiple agents, isolated worktrees, reviewable diffs, CLI and IDE interop πŸ”—
Visual Studio Windows, macOS 17.14.12+, 18.1.0+ 2026-01-06 00:00 UTC Free / $250/yr Gemini 3 Flash integration, faster performance, zero-migration upgrades, real-time profiler agent ❌
IntelliJ IDEA Windows, macOS, Linux 2025.3.2 2026-01 Free / $149/yr Java 24 support, Kotlin K2 mode, performance and memory improvements ❌
IBM Bob Cross-platform GA (April 28, 2026) 2026-04-28 Free trial + Enterprise plans Multi-model orchestration, full SDLC, 45% productivity gain πŸ”—
PolyAI ADK PolyAI GA (April 22, 2026) 2026-04-22 Enterprise CX AI-native dev, Cursor/Claude Code integration πŸ”—
Superset Windows, macOS, Linux - 2026-05-22 Free (OSS, YC P26) Agentic IDE for running coding agents in parallel, git worktree management, remote workspaces, task tracking πŸ”—
Creor Windows, macOS, Linux - 2026 Free / Paid AI-native IDE with specialized agents (Build, Plan, Explore), agent auto-routing, MCP support πŸ”—
Sinew Windows, macOS, Linux - 2026-05-13 Free (MIT) Desktop AI coding harness, Tauri 2, multi-provider, agent swarm (2-8 agents), MCP, skills πŸ”—
JAT Windows, macOS, Linux - 2026-04-14 Free (MIT) Self-contained agentic IDE, 20+ parallel agents, task management, unified environment πŸ”—
Qoder 1.0 ⭐ Windows, macOS, Linux 1.0 2026-06-16 00:00 UTC Free / API Alibaba Cloud Autonomous Development Desktop; Quest window for agent command; editor + agent parallel workspaces; GA 2026-06-16 πŸ”—

Native AI Editors

Editor Platform Version Release Date Pricing Key Features GitHub
Zed macOS, Windows, Linux 0.226.3 2026-03-03 00:00 UTC Free (OSS) + Copilot $10/mo Fast, collaboration, Gemini and Claude, Zeta AI, agent thread history, edit prediction providers, self-hosted OpenAI-compatible servers πŸ”—
Dyad Windows, macOS, Linux - - Free (OSS) Local generation, BYO keys πŸ”—
Memex macOS, Windows - - Freemium (Free + $10/mo) Agentic, browser↔desktop πŸ”—

VS Code Forks

IDE Platform Version Release Date Pricing Autonomous MCP GitHub
Cursor ⭐ Windows, macOS, Linux 3.11 2026-07-10 00:00 UTC ⭐ Free / Pro $20/mo / Pro+ $60/mo / Ultra $200/mo; side chats, conversation search, and simpler project/repo pickers βœ… βœ… ❌
Windsurf Windows, macOS, Linux Quota plans 2026-03-19 00:00 UTC Free / Pro $20/mo / Max $200/mo / Teams from $40 flex seat or $80 full user per month βœ… βœ… ❌
Trae macOS, Windows - - Free ❌ ❌ πŸ”—
PearAI Windows, macOS, Linux - - Free (OSS) βœ… ❌ πŸ”—
Void Windows, macOS, Linux - - Free (OSS) βœ… βœ… πŸ”—
Kiro ⭐ Windows, macOS, Linux 1.0.138 2026-07-13 00:00 UTC ⭐ Free preview; faster session startup, PowerShell trust on Windows, compaction-loop fix for long sessions, and improved MCP tool recovery; latest 1.0.x builds download directly while auto-updates roll out gradually βœ… βœ… πŸ”—
VS Code Agents Windows, macOS, Linux Insiders 2026-04-21 Free βœ… βœ… πŸ”—

Cursor vs. Windsurf Pricing (July 2026): Cursor Pro is $20/month, Pro+ is $60/month, and Ultra is $200/month. Windsurf Pro is $20/month and Max is $200/month, with Teams offering $40 flex seats or $80 full users per month. Cursor's current differentiator is broader cloud-agent and customization surface, while Windsurf's post-March 2026 quota plans simplify spend predictability. For a full alternatives analysis, see Cursor Alternatives & Better Value ↓.

Cursor Alternatives & Better Value 🎯

Cursor Pro is $20/month, while several official pricing pages and product docs show lower-cost or BYOK alternatives with similar agent workflows, local-model support, or broader IDE coverage.

Pricing Reality Check (July 2026)
Tool Free Tier Pro / Paid Premium / Max Notes
Cursor 2K completions + 50 premium req $20/mo (Pro) / $60/mo (Pro+) $200/mo (Ultra) Heavy users report $40–60/mo effective cost; quotas tightened June 2025
Claude Code Limited $20/mo (Pro) $100/mo (Max 5x) / $200/mo (Max 20x) Significantly higher limits than Cursor Pro at same price
GitHub Copilot Limited AI Credits $10/mo base + $5 included AI Credits (Pro) $39/mo base + $31 included AI Credits (Pro+) / $39/user (Enterprise) Usage-based billing since June 1, 2026; bundled credits are consumed by chat, agent mode, cloud agent, CLI, and apps
Cline (BYOK) Free (bring API key) Free (pay API costs ~$5–15/mo) Unlimited 5M+ installs; zero subscription cost
Windsurf Light quota $20/mo (Pro) / Teams from $40 flex seat or $80 full user $200/mo (Max) Quota-based plans replaced older credit plans in March 2026
Continue.dev Free (OSS) Free (BYOK) Unlimited No subscription; supports Ollama, LM Studio, all cloud APIs
Void Free (OSS) Free (BYOK) Unlimited VS Code fork + MCP; privacy-first, no telemetry
Aider Free (BYOK) Free (BYOK) Unlimited Git-native; used by serious OSS contributors
OpenHands Free (OSS) Free (self-host) Unlimited Full SDLC agent; MCP support
Value-For-Money Ranking
Rank Tool $/Month (Real) SWE-bench Limits Verdict
πŸ₯‡ 1 GitHub Copilot Pro $10 base + $5 credits 55% Usage-based AI Credits (with bundled allowance) Best value for light-moderate users
πŸ₯ˆ 2 Cline (BYOK) ~$5–15 (API) 80.8% Unlimited Best value for power users
πŸ₯‰ 3 Windsurf Pro $20 β€” Quota-based plan Best value paid Pro tier
4 Claude Code Pro $20 80.8% Higher than Cursor Pro Best autonomy at $20/mo
5 Continue.dev $0 β€” Unlimited Best free option
6 Cursor Pro $20–60 (effective) Unpublished Caps hit easily Good UX, poor value at scale
7 Claude Code Max $100 80.8% 5Γ— limits Outclasses Cursor Ultra ($200)
Official Source Basis

This comparison uses official pricing pages, official changelogs, official extension marketplaces, and official GitHub repositories. Unofficial social feedback is not used for eligibility, pricing, or benchmark claims.

Why Specific Alternatives Beat Cursor

πŸ₯‡ Cline (BYOK) β€” Best Overall Value Free VS Code extension with zero subscription cost. You pay only for your API usage (~$5–15/month for typical developer workloads). Achieved 80.8% SWE-bench Verified β€” identical to Claude Code. Supports 100+ providers including local Ollama/LM Studio models, meaning you can run it for $0/month with a local model like Qwen3-Coder or DeepSeek-Coder. Over 5 million installs make it the most-adopted free agent tool. Full MCP support.

πŸ₯ˆ Claude Code β€” Best Paid Alternative Anthropic's purpose-built autonomous coding CLI. At $20/mo (Pro), usage limits are substantially higher than Cursor Pro at the same price. Max 5Γ— ($100/mo) offers more headroom than Cursor Ultra ($200/mo) at half the cost. Claude Code produces code described as ~30% less rework vs. Cursor. Won 67% of blind quality comparisons against Cursor. Integrates natively into VS Code and JetBrains via extension, plus terminal. Multi-session support, computer use (research preview), sub-agents.

πŸ₯‰ GitHub Copilot β€” Best for Teams & IDE-Native Work At $10/month base with $5 in bundled AI Credits (Pro), GitHub Copilot requires no fork, no new editor, and no learning curve. Works inside existing VS Code, JetBrains, Vim, Neovim, and more. Since June 1, 2026, billing switched to usage-based AI Credits; chat, agent mode, cloud agent, CLI, and apps all consume the same credit pool. Pro+ starts at $39/month base with larger included credits. Best for developers who want AI assistance without changing their workflow.

Continue.dev β€” Best Free Fully-Featured Option 100% open-source, works in VS Code and JetBrains. BYOK means no subscription cost ever. Supports Ollama, LM Studio, and all major cloud APIs. Chat, autocomplete, and agent mode. Active development community. No telemetry by default.

Void β€” Best Privacy-First Fork Open-source VS Code fork with full MCP support, BYOK, and no telemetry. Feature-comparable to Cursor with zero subscription cost. Best for teams that need data sovereignty or work in air-gapped environments.

🎯 Bottom Line for Sharif's Use Case: If you use LM Studio and prefer local models, Cline or Continue.dev (both BYOK, both free) give you unlimited usage at zero subscription cost. For cloud-based work, Claude Code Pro ($20) remains the strongest autonomy-first alternative at the same entry price as Cursor Pro ($20). GitHub Copilot is still the easiest low-friction option if you want to stay inside your current IDE and manage against a usage-credit budget.

Web-Based IDEs

IDE Platform Version Release Date Pricing Self-Hostable Best For GitHub
Replit 3 Web - - Free Starter, Core $20/mo, Pro $100/mo ❌ Learning/Prototyping ❌
Bolt.new Web - - Free, Pro $20-25/mo, Teams $30/user/mo ❌ Quick apps ❌
Bolt.diy Self-hosted - - Free (MIT), bring your own API βœ… Self-hosted πŸ”—
Lovable Web - - Free (5 credits/day), Pro $25/mo, Business $50/mo ❌ UI/Full-stack ❌
v0 Web - - Free ($5 credits/mo), Premium $20/mo, Teams $30/user ❌ React components ❌
Gitpod Web - - Free + Paid ❌ Cloud dev environments ❌
Rork Web - - Free & Paid (credits) ❌ Mobile apps (iOS/Android) ❌
Google Stitch Web - 2026-03 Free (Google account, 550 gen/mo) ❌ UI design, Figma/React export ❌
Google Antigravity 2.0 Windows, macOS, Linux (CLI), Web (IDE) 2.0 2026-05-19 00:00 UTC Google AI Pro $19.99/mo / Ultra $99.99/mo (5Γ—) or $200/mo (20Γ—) Agent-first platform: desktop app + CLI (agy) + SDK + Managed Agents API; multi-agent orchestration, parallel tasks, voice control πŸ”—
Jules Web - 2025-05-20 00:00 UTC Free beta, higher limits on Google AI Pro / Ultra Async repo agent, reviewable diffs, GitHub integration ❌

CLI Tools πŸ–₯️

Command-line AI tools for autonomous coding and terminal enhancement.

Autonomous Coding Agents

Tool Platform Pricing Key Features GitHub
Aider Windows, macOS, Linux Free Gold standard, Architect mode, thinking tokens πŸ”—
Claude Code 2.2.1+ macOS, Linux, Windows Free + API Sonnet 5 default support via Claude Code 2.1.197+; Opus 4.8 support; Fable 5 access restored globally 2026-07-01 00:00 UTC; simple mode file editing, multi-session support, /autofix-pr, plan mode, computer use (research preview), CLAUDE.md skills & sub-agents πŸ”—
Codex CLI ⭐ Windows, macOS, Linux Included with ChatGPT Plus/Pro/Business/Edu/Enterprise or API key Rust-native local coding agent; sandbox, approval modes, subagents, command autocomplete, workspace search; v0.144.5 (2026-07-16) improved dangerous-command detection and denial messaging; v0.144.4 (2026-07-14) had no user-facing changes; v0.144.2 restored Guardian auto-review policy πŸ”—
Junie CLI ⭐ Windows, macOS, Linux Free (BYOK) LLM-agnostic, JetBrains IDE integration (ACP), GA 2026-06-17, agentic debugging, PR review, local model support (LiteLLM/LMStudio/Ollama) πŸ”—
Goose Windows, macOS, Linux Free (Apache-2.0) MCP, extensible, desktop app, 25+ providers πŸ”—
Grok Build Windows, macOS, Linux $300/mo (SuperGrok/X Premium Plus) Plan mode, Arena mode, 2M context window, security-focused ❌
GPT-Pilot Windows, macOS, Linux Free Full dev team simulation πŸ”—
OpenHands Windows, macOS, Linux Free Cloud agents, MCP πŸ”—
Mentat Windows, macOS, Linux Free Multi-file coordination πŸ”—
SERA Linux, macOS Free (Apache 2.0) Open-source coding agent, 200K synthetic trajectories πŸ”—
Kimi K2.7 Code ⭐ Windows, macOS, Linux $0.95 / $4.00 per 1M Coding-focused agent; 30% fewer thinking tokens vs K2.6; Modified MIT; 262K context; released 2026-06-12 πŸ”—
Kilo for GitHub ⭐ GitHub (Cloud) Free (Kilo credits) @kilocode-bot coding agent in GitHub issues/PRs; Cloud Agent backend; reads full thread context; released 2026-06-18 πŸ”—
Hermes Agent v0.17.0 ⭐ Windows, macOS, Linux Free (OSS) Major desktop app update; live subagent watch-windows; VS Code theme support; 199K GitHub stars; released 2026-06-19 πŸ”—
Kimi Code CLI Windows, macOS, Linux Free (MIT) TypeScript CLI by Moonshot AI; successor to kimi-cli; powered by Kimi K2.6 (1T MoE); subagent support (coder/explore/plan); 262K context; SWE-bench Pro 58.6%; ACP (Zed, JetBrains); released 2026-06-06 πŸ”—
MiMo Code ⭐ Windows, macOS, Linux Free (MIT) Xiaomi open-source terminal coding agent; 62% SWE-bench Pro; integrated with Claude Code, OpenClaw, Hermes; released 2026-06-11 πŸ”—
Vibe (Mistral) Windows, macOS, Linux Free (Le Chat) / API Agentic coding + work agent, multi-model Mistral harness, VS Code extension, Code Mode web surface, parallel sessions πŸ”—
Google Agents CLI ⭐ Windows, macOS, Linux Free Unified ADLC for Google Cloud Agent Platform; scaffold/run/evaluate/deploy agents; released 2026-04-22 πŸ”—
Peri ⭐ Windows, macOS, Linux Free (Apache-2.0) Rust terminal coding agent; Claude Code compatible; multi-LLM; MCP & ACP support; 95%+ prompt cache; released 2026-03-20 πŸ”—
AI Dev Kit Cross-platform Free 59 skills, 33 agents, TDD, security audit, CI/CD πŸ”—

Assisted CLI Tools

Tool Developer Pricing Best For GitHub
Gemini CLI ⚠️ Google Free (now shut down) Google ecosystem & Gemini models in-terminal; stopped serving requests June 18, 2026 for AI Pro/Ultra/free users; migrate to Antigravity CLI πŸ”—
Antigravity CLI (agy) ⭐ Google Free (~20 req/day) / Pro/Ultra higher limits Agentic coding, multi-agent orchestration, Go-based replacement for Gemini CLI; released 2026-06-18 πŸ”—
Google Colab CLI ⭐ Google Free (Colab account) Run Python on remote Colab GPUs/TPUs from local terminal; AI agent compatible; released 2026-06-06 πŸ”—

⚠️ Anthropic June 15 Billing Split: As of 2026-06-15, Anthropic moved Claude Agent SDK, claude -p, and Claude Code GitHub Actions to separate metered credits billed at full API rates. Pro subscribers get $20/mo credit; Max 20x gets $200/mo. Interactive Claude Code and Claude.ai remain on standard subscription. Fable 5 was included on Pro/Max/Team/Enterprise subscription plans at no extra cost through June 22; after that date, usage credits are required.

| Cursor CLI | Cursor | Free tier | Terminal + IDE bridge | πŸ”— | | Qwen Code | Alibaba | Free | Qwen optimization | πŸ”— | | Qodo CLI | Qodo | Free tier | Testing, review & agent workflows | πŸ”— |

Gemini CLI (deprecation & security notes): As of June 18, 2026, Gemini CLI and the Gemini Code Assist IDE extensions have stopped serving requests for Google AI Pro, Ultra, and free Gemini Code Assist for individuals, per Google's official announcement. Google ecosystem users should migrate agentic workflows to Antigravity CLI (agy). Enterprise Code Assist Standard/Enterprise licensees retain access. Current CLI generations add offline search, bundled ripgrep, optional color-accessible themes, interactive shell tool invocation, and an A2A-style agent registry surface in the tool (see Gemini CLI changelogs). Treat upgrades through v0.39.1+ as mandatory maintenance: public advisories described a critical remote-code path affecting earlier builds (widely reported as CVSS 10.0); confirm your installed build against the release notes for your package channel. Source: Google Developers Blog

CLI Tools by Programming Language

AI coding CLI tools categorized by their primary language support. All tools below accept plain English prompts.

Tool Primary Languages Multi-Language Local LLM Cloud API Pricing GitHub
Aider Python, JS, TS, Go, Rust, Ruby, Java, C/C++ βœ… (100+ langs) βœ… (Ollama, LM Studio) βœ… Free (OSS) πŸ”—
Claude Code All (polyglot) βœ… ❌ Claude API ($3–$15/M) Free tool / API cost ❌
Codex CLI Python, JS, TS, Bash βœ… ❌ OpenAI API Free OSS / API cost πŸ”—
OpenHands Python, JS, TS, Go, Rust, Java βœ… βœ… βœ… Free (OSS) πŸ”—
Goose (Block) Polyglot (25+ providers) βœ… βœ… (Ollama, LM Studio) βœ… Free (OSS) πŸ”—
Continue Polyglot (VS Code, JetBrains) βœ… βœ… (Ollama, LM Studio) βœ… Free (OSS) πŸ”—
Qwen Code Python, JS, TS, Go, Java βœ… βœ… (Qwen models) βœ… Free (OSS) πŸ”—
Devstral CLI Python, JS, TS, Go, Rust βœ… βœ… Mistral API Free OSS model / API cost ❌
OpenCode Polyglot βœ… βœ… βœ… Free (OSS) πŸ”—
Mentat Python, JS, TS, Go βœ… ❌ OpenAI API Free (OSS) πŸ”—
Amp (Sourcegraph) All βœ… ❌ βœ… Free / Enterprise ❌

CLI for Programming Languages & Multiple Use

Purpose-built CLI tools for coding across specific languages or polyglot multi-stack workflows.

Tool Language Focus Platform Pricing Key Features GitHub
Aider Polyglot (Python, JS, TS, Go, Rust, any) All Free (BYOK) Git-native multi-file edits, Architect mode, repo maps, thinking tokens πŸ”—
Claude Code Polyglot All Free + API Computer use, sub-agents, CLAUDE.md skills, Opus 4.7, multi-session πŸ”—
Codex CLI Python, JS, TS All Free (OpenAI account) Sandbox execution, approval modes, OpenAI models πŸ”—
OpenHands Python, JS, TS, Go, Rust All Free (OSS) Full SDLC agent, MCP, local LLM via Ollama πŸ”—
Goose Polyglot All Free (Apache-2.0) 25+ providers, MCP, extensible extensions, desktop app πŸ”—
Continue Polyglot All Free (OSS) VS Code + JetBrains, custom models via Ollama/LM Studio πŸ”—
Qwen Code Python, JS, TS, Go All Free Optimized for Qwen3-Coder 480B, Apache 2.0 ❌
Mentat Polyglot All Free Multi-file coordination, context-aware diffs πŸ”—
AI Dev Kit Polyglot All Free 59 skills, 33 agents, TDD, security audit, CI/CD pipeline πŸ”—
Devstral CLI Python, JS, TS, Go All Free (Mistral free tier) Mistral's open coding model, OpenRouter free access ❌
Junie CLI Polyglot All Free (BYOK) LLM-agnostic, JetBrains IDE integration, MCP πŸ”—
SERA Python, JS, TS Linux, macOS Free (Apache 2.0) Open-source coding agent, 200K synthetic trajectories πŸ”—

Terminal Enhancers

Tool Platform Pricing Key Features
Warp Terminal macOS, Linux, Windows Free AI Agents, workflow sharing, Warp Drive automation
Microsoft Intelligent Terminal Windows Free (MIT) AI agent pane (Agent Client Protocol), automatic error detection, GitHub Copilot CLI default, Windows Terminal fork β€” 2026-06-02 ⭐
Starship Cross-platform Free (OSS) Fast cross-shell prompt; rich plugins & AI-adjacent ecosystem
Fig macOS, Linux Free Autocomplete, AI suggestions
Google Colab CLI ⭐ Cross-platform Free (Colab account) Run Python on remote Colab GPUs/TPUs from local terminal; AI agent compatible; released 2026-06-06

IDE Add-ons 🧩

Extensions and plugins that add AI capabilities to existing IDEs.

Universal (Cross-Platform)

Add-on Platform Pricing Context Best For GitHub
GitHub Copilot ⭐ VS Code, JetBrains, Vim Usage-based AI Credits since June 1, 2026 Large General coding; Kimi K2.7 Code available in model picker (July 1, 2026); Copilot Vision GA (July 1, 2026); enterprise session streaming/API, AI credit pools, and Copilot-only Gemini 2.5 Pro / Gemini 3 Flash deprecation on 2026-07-31 πŸ”—
Supermaven VS Code, JetBrains, Neovim Free / $10/mo 1M Large codebases ❌
Codeium VS Code, JetBrains, Vim Free / $15/mo / $60/mo Medium Free alternative ❌
Continue VS Code, JetBrains Free (OSS) Custom Self-hosted πŸ”—
Cody VS Code, JetBrains, Web Free (discontinued) / Enterprise Starter $19/mo / Enterprise $59/mo Enterprise Code search πŸ”—
Tabnine VS Code, JetBrains, VS, Eclipse Free / $39/mo Local Privacy ❌
Tabby VS Code, JetBrains, Vim, Neovim Free (OSS) Self-hosted Self-hosted code completion πŸ”—

VS Code Specific

Add-on Pricing Autonomous MCP Best For GitHub
Codex Free (with ChatGPT Plus $20/mo or Pro $200/mo) βœ… βœ… OpenAI's official coding agent πŸ”—
Cline Free βœ… βœ… Full agent πŸ”—
GitHub Copilot (Agent Mode) ⭐ Usage-based AI Credits since June 1, 2026 ⚠️ ❌ Guided agent workflows; auto-opens PRs with self-review; Kimi K2.7 Code available; Copilot CLI in GitHub Actions now supports GITHUB_TOKEN with copilot-requests: write πŸ”—
RooCode Free/Pro ⚠️ ❌ Complex tasks πŸ”—
JetBrains AI Assistant $10/mo (Pro) βœ… ❌ JetBrains-quality AI now in VS Code (public preview); multi-file edits, Mellum LLM ❌
ReSharper for VS Code Included with JetBrains subscription ❌ ❌ C# code analysis + refactoring inside VS Code/Cursor; released 2026-03-05 ❌
Keploy OSS/Enterprise ❌ ❌ Testing ❌
Microsoft Foundry Toolkit for VS Code ⭐ Free (Azure account) βœ… βœ… End-to-end Hosted Agent lifecycle; GA 2026-04-16; consolidated AI Toolkit + Foundry extension; LangGraph samples; released 2026-06-02 πŸ”—
VS Code ACP Client ⭐ Free (MIT) βœ… βœ… Agent Client Protocol client; connects 11+ AI agents (Claude Code, Codex, Cursor, etc.) to VS Code; released 2026-06-10 πŸ”—
Prompt Foundry ⭐ Free ❌ βœ… VS Code/Cursor extension; modular prompt blocks with Liquid syntax; MCP server for context updates; released 2026-06-18 πŸ”—
GrackerAI ⭐ Free ❌ ❌ Validates llms.txt, robots.txt, AI visibility signals in VS Code; released 2026-06-16 ❌
A11yResolver ⭐ Free (beta) ❌ ❌ VS Code extension for accessibility remediation; AI agent flags WCAG issues; released 2026-06-18 ❌

Claude Tag (Anthropic, June 2026): Not an IDE extension but worth noting here. Claude Tag brings an always-on AI teammate to Slack. Each channel gets an isolated Claude identity with persistent memory. Available on Team and Enterprise plans. Source: πŸ”—

JetBrains Specific

Add-on Pricing Claude Agent Junie (GA) Best For
JetBrains AI Assistant $10/mo (Pro), $249/yr (Ultimate) βœ… βœ… Deep IDE integration; Junie GA 2026-06-17
JetBrains Claude Agent Included in subscription βœ… β€” Native agent
JetBrains Junie ⭐ Included in subscription β€” βœ… GA 2026-06-17; ACP-based IDE integration; agentic debugging; BYOK; local model support; IntelliJ/WebStorm/PyCharm/GoLand/CLion

JetBrains Marketplace Security (June 2026): 15 malicious AI plugins removed from JetBrains Marketplace; 7 publisher accounts terminated; remote kill-switch triggered. If you installed any affected plugins before June 17, 2026, revoke exposed API keys immediately.

API Providers πŸ”Œ

Services for accessing AI models via API. This pass prioritizes the cheapest verified public text APIs and the free or trial access that is explicitly documented by each provider.

Model Labs (Direct)

Provider Cheapest verified public text model Direct price Free or trial access Notes Source
Groq Llama 3.1 8B Instant $0.05 input / $0.08 output GroqCloud Free plan ($0) Lowest verified public direct text API price in this pass πŸ”—, πŸ”—
Google Gemini Gemini 2.5 Flash-Lite Free tier: free of charge; paid tier $0.10 / $0.40 Gemini Developer API free tier Batch and Flex drop to $0.05 / $0.20 πŸ”—
DeepSeek DeepSeek-V4-Flash $0.14 input / $0.28 output No standing free tier listed on pricing page Cache-hit input is $0.0028 πŸ”—
Mistral AI Mistral Small 4 $0.15 input / $0.60 output Free mode enabled by default, no credit card required Good low-cost coding and multimodal option πŸ”—, πŸ”—
Cohere Command R $0.15 input / $0.60 output Trial API keys are free but limited; Command A+ is free until rate limits are reached Low-cost RAG and tool use option πŸ”—, πŸ”—, πŸ”—
OpenAI GPT-5.4 nano $0.20 input / $1.25 output OpenAI quickstart includes one free test API request Batch and Flex drop to $0.10 / $0.625 πŸ”—, πŸ”—
Anthropic Claude Haiku 4.5 $1.00 input / $5.00 output New users receive a small amount of free credits Lowest-cost Claude API model πŸ”—, πŸ”—
xAI Grok-4.5 $2.00 input / $0.50 cached / $6.00 output No public free tier listed on pricing page New flagship (2026-07-08); 500K context, text+image; Grok 4.3 remains $1.25 / $2.50 πŸ”—
Meta Meta Muse Spark 1.1 API-only (public preview) US-only public preview via Meta Model API Closed-weights frontier multimodal reasoning model; agentic tool and computer use; 1M context; released 2026-07-09 πŸ”—

Unified APIs & Aggregators

Provider Models Key Features
OpenRouter 400+ Unified API, crypto/fiat, model rankings, fallback routing, Speech & Transcription APIs, Model Fusion (launched 2026-06-12), private models, enterprise workspace controls; new models: Claude Opus 4.8, Step 3.7 Flash, MiniMax M3, Qwen3.7-Plus, NVIDIA Nemotron 3 Ultra, Microsoft MAI models, Nano Banana 2, Nano Banana Pro, GLM 5.2, Cohere North Mini Code (free), Nex-N2-Pro (free), GPT-Image-1.5, GPT-Image-2, Poolside Laguna, Qwen3-Coder (June 2026)
Hugging Face Thousands Serverless inference, free tier
LiteLLM 100+ Open-source proxy; single OpenAI-compatible API for any provider
Portkey 200+ Gateway: load balance, fallbacks, caching, observability

Model Router Services πŸ”€

Services that aggregate multiple models through a unified API, often with load balancing, caching, and cost optimization features.

Provider Models Pricing Key Features
Together AI 200+ $0.10-$2.50/1M Broadest open-weight catalog (text/image/video/audio/embeddings), NVIDIA Blackwell-optimized
Fireworks AI 50+ $0.10-$1.50/1M Production-grade open-source inference, 90% cheaper than closed providers, NVIDIA Blackwell
Groq 20+ $0.05-$0.90/1M Fastest LPU inference: 476 tok/s on 120B, <100ms TTFT, deterministic latency
Cerebras 10+ $0.10-$1.00/1M Wafer-scale engine: ~3,000 tok/s, 5Γ— faster than NVIDIA Blackwell, 80–150ms TTFT
NVIDIA NIM 91+ Free (cloud endpoints) / self-host Broadest model variety: LLMs + vision + audio + bio + climate; self-hostable Docker containers
Vercel AI Gateway ⭐ 100+ $5/mo free credits; paid credits at provider list rates, zero markup Available on all plans; unified API for text, image, video, realtime, speech, embeddings, and reranking; BYOK on paid tier; model fallbacks, provider preferences, budgets, observability, and Routing Rules beta
Mercury 2 (Inception Labs) ⭐ 1 $0.25/$0.75/1M Diffusion LLM; 1,009 tok/s on Blackwell; tunable reasoning; 128K context; released 2026-05-12
Anyscale 100+ $0.20-$2.00/1M OpenAI-compatible, HIPAA/SOC 2/EU data residency, enterprise contracts
Replicate 50+ $0.20-$2.00/1M Model-as-a-service, versioning, GPU rental
OctoAI 20+ $0.25-$2.00/1M High-performance inference
Predibase 12+ $0.30-$2.50/1M Fine-tuning + multi-LoRA serving
RunPod Serverless 10+ $0.20-$2.00/1M GPU serverless, flexible pricing
CoreWeave 8+ Custom Enterprise-grade, dedicated GPUs
Infron ⭐ 100+ Up to 35% off direct pricing Enterprise-grade unified API; provisioned throughput plan; broad model selection; released 2026-06-10
Aethir Mesh ⭐ 10+ open-source Pay-per-use Decentralized GPU infrastructure; single API for open-source models (DeepSeek V4, Kimi K2.6, GLM-5.1, etc.); released 2026-06-12

Budget Model Aggregators

Provider Models Pricing Best For
DeepInfra 50+ Free tier + $0.10-1.00/1M Free tier, diverse models
TensorZero 30+ Free tier + $0.15-1.50/1M Cost optimization, caching
Scale AI 25+ Custom pricing Enterprise, managed services
AssemblyAI 15+ Free tier + pay-per-use Speech + text models

GPU Clouds

Provider Type A100 80GB ($/hr) H100 80GB ($/hr) Best For
Vast.ai GPU Marketplace ~$0.09–$0.59 ~$0.50–$1.50 Cheapest market pricing, diverse GPU options
RunPod GPU Rental $0.34–$0.69 (RTX 4090) / $0.60–$1.19 (A100) $1.99–$3.29 Flexibility, per-second billing, cost-effective fine-tuning & inference
Lambda Labs Cloud GPU $1.29–$2.79 $2.49–$3.99 Reliable on-demand, academic discounts, 1-Click Clusters
DigitalOcean GPU Droplets $1.29 $2.99 Simple fine-tuning workflows
Vultr Global Cloud $1.29–$2.80 $1.99–$2.99 Hourly GPU instances, global regions
Hyperbolic Decentralized ~$1.50 ~$3.50 Crypto/Fiat payments
Cerebrium Serverless GPU ~$2.10 ~$3.40 Python-native ML inference & fine-tuning
Modal Labs Serverless GPU ~$2.50–$3.40 ~$3.95 Fine-tuning with LoRA, distributed training; per-second billing
Replicate Model-as-a-Service $5.04 $5.49 Quick deployment, serverless inference
Together AI AI-Native Cloud β€” β€” Fast inference & fine-tuning for open models
Fireworks AI Inference & Fine-tuning β€” β€” Fast inference, RFT for model shaping
CoreWeave Enterprise GPU Cloud ~$2.70 ~$6.16 Enterprise-grade, dedicated GPUs, Kubernetes-native
Databricks Mosaic AI Integrated ML Platform β€” β€” Enterprise fine-tuning, governed serving, RAG
NVIDIA DGX Cloud Managed AI Training Custom Custom Co-engineered clusters, maximum ROI for training

GPU pricing as of July 2026. RunPod/Vast.ai prices vary by community vs. secure cloud and spot vs. on-demand. RTX 4090 pricing from RunPod starts at $0.34/hr. Vast.ai A100 40GB from $0.52/hr. Modal Labs production workloads cost 3.75x base rates (3x non-preemptible multiplier x 1.25x US regional multiplier). CoreWeave A100/H100 pricing normalized from 8-GPU nodes. Lambda Labs A100 40GB at $1.29/hr, A100 80GB at $2.49-$2.79/hr, H100 SXM at $3.99/hr (8x) or $4.29/hr (1x). Vast.ai prices are market-set medians.

Inference Clouds

Provider Specialization Speed
Together AI Llama/Qwen/Mistral Fast
Fireworks AI FireAttention Low-latency, 6 free models
Groq LPU >500 T/s
Cerebras Wafer-Scale >2000 T/s
NVIDIA NIM 91 free endpoints, DGX Cloud 20Γ— faster than NVIDIA GPU

Hosted Open Source Models πŸ†“

Services that provide hosted access to open source models with API endpoints, offering easier deployment than self-hosting.

Provider Models Pricing Best For
Hugging Face Inference 1000+ Free tier + $0.10-1.00/1M Largest selection, easy deployment
Replicate 50+ Free tier + $0.20-2.00/1M Model-as-a-service, versioning
Together AI 30+ $0.20-3.00/1M Optimized for open models
Fireworks AI 25+ Free tier + $0.10-1.50/1M Low-latency, 6 free models
Anyscale 100+ $0.20-2.00/1M Distributed inference, scaling
OctoAI 20+ $0.25-2.00/1M High-performance GPU inference
Banana.dev 15+ $0.15-1.80/1M Fast deployment, auto-scaling
Predibase 12+ $0.30-2.50/1M Fine-tuning + inference
RunPod Serverless 10+ $0.20-2.00/1M GPU serverless, flexible pricing
CoreWeave 8+ Custom pricing Enterprise-grade, dedicated GPUs

Free Tier APIs

Provider Official free access What happens after free or trial Best For Source
Google Gemini Gemini Developer API free tier is free of charge Gemini 2.5 Flash-Lite paid tier starts at $0.10 / $0.40 Best official always-on free multimodal API πŸ”—
Groq GroqCloud Free plan is $0 for build and test use Developer plan is pay-per-token; Llama 3.1 8B Instant starts at $0.05 / $0.08 Fast low-cost iteration πŸ”—, πŸ”—
Cohere Trial API keys are free but limited; Command A+ is free until rate limits are reached Command R starts at $0.15 / $0.60 RAG, tool use, multilingual testing πŸ”—, πŸ”—, πŸ”—
Mistral AI Free mode is enabled by default with no credit card required Mistral Small 4 starts at $0.15 / $0.60 on Scale Cheap coding and multimodal experiments πŸ”—, πŸ”—
Anthropic New users receive a small amount of free credits Haiku 4.5 starts at $1.00 / $5.00 Testing Claude API behavior πŸ”—, πŸ”—
Vercel AI Gateway ⭐ $5/month included on the free tier for eligible models Paid credits use provider list rates with zero markup Multi-provider testing behind one endpoint πŸ”—
OpenAI Developer quickstart includes one free test API request Standard API pricing starts with GPT-5.4 nano at $0.20 / $1.25 SDK evaluation and first-request testing πŸ”—, πŸ”—

Automation πŸ€–

AI-powered tools for automating browser and desktop tasks.

Browser Automation 🌐

Tools and frameworks for AI-powered browser automation.

Standalone AI Browsers

Browser Platform Pricing Open Source Local AI Agent/Computer Use API Access Multi-Agent Parallel Sessions Best For GitHub
Norton Neo Windows, macOS, iOS, Android Free ❌ ❌ ⚠️ ❌ ❌ ❌ AI-native browser, Magic Box, Peek & Summarize, Smart Tab Groups ❌
Deta Surf macOS, Linux Free (alpha) βœ… βœ… ⚠️ ❌ βœ… βœ… Browser + file manager + AI notebook, Surf Memory ❌
Firefox (AI Controls) Windows, macOS, Linux, iOS, Android Free βœ… ❌ ⚠️ ❌ ❌ ❌ AI Controls dashboard (v148+), ChatGPT/Claude/Mistral sidebars ❌
Dendrite Windows, macOS, Linux Free (OSS) βœ… βœ… βœ… ❌ βœ… βœ… Developer AI browser client for agents ❌
Perplexity Comet Windows, macOS, iOS, Android Free / Pro $20/mo ❌ ❌ βœ… ❌ ❌ ❌ Research + background tasks, voice mode, Computer Max agent ❌
ChatGPT Agent Mode Web, iOS, Android Plus $20/mo, Pro $200/mo ❌ ❌ βœ… ❌ ❌ ❌ Full computer use: browse, code, fill forms, book travel ❌
Dia macOS (M1+ / macOS 14+) Free / Pro $20/mo ❌ ❌ ⚠️ ❌ ❌ ❌ Tab intelligence, Skills, browsing history AI context ❌
Google Chrome (Auto Browse 2) Windows, macOS, Linux, ChromeOS Free / Gemini Pro $19.99/mo ❌ ❌ βœ… ❌ ❌ ❌ Gemini 3 built-in, auto browse 2 agentic tasks, Chrome Skills saveable workflows, Universal Commerce Protocol, WebMCP support ❌
Microsoft Edge (Copilot Agent) Windows, macOS, iOS, Android Free / Copilot Pro $20/mo ❌ ❌ βœ… ❌ ❌ ❌ Cross-tab context, voice commands, form automation, bookings ❌
Genspark Windows, macOS, Web, iOS, Android Free / Plus $25/mo / Pro $249/mo ❌ βœ… (169 local models) βœ… ❌ βœ… βœ… Super Agent, AI slides, AI websites, deep research, Call For Me, Genspark Claw desktop app ❌
Brave Leo (AI Browser) Windows, macOS, Linux, iOS, Android Free / Premium $14.99/mo βœ… (Chromium) βœ… (Leo local) ⚠️ ❌ ❌ ❌ Privacy-first, zero-log AI, Skills, Memories, local models ❌
SigmaOS (Airis) macOS Free / Pro (subscription) ❌ ❌ ⚠️ ❌ ❌ ❌ NL commands: "Book Airbnb in Iceland", cross-tab AI, YC-backed ❌
Opera Neon Windows, macOS $19.90/mo ❌ ❌ βœ… βœ… MCP Connector (Mar 2026) ❌ ❌ Agentic browsing, Aria assistant, MCP API for Claude/ChatGPT/n8n integration ❌
Opera One (Aria) Windows, macOS, Linux, iOS, Android Free ❌ ❌ ⚠️ ❌ ❌ ❌ Built-in Aria AI assistant, sidebar AI tools ❌
Firefox (AI Sidebar) Windows, macOS, Linux, iOS, Android Free βœ… ❌ ⚠️ ❌ ❌ ❌ AI Controls dashboard (v148+), ChatGPT/Claude/Mistral sidebars ❌
BrowserOS Linux, macOS Free βœ… βœ… βœ… ❌ βœ… βœ… Privacy-focused, built-in MCP, agentic πŸ”—
Manus AI Web (Cloud) Free 300 credits/day / Plus $20/mo / Pro $200/mo ❌ ❌ βœ… ❌ βœ… βœ… Cloud agent, full computer: code, deploy, files, search ❌
Sigma AI Browser Windows, macOS, Linux Free / Pro $29/mo ❌ βœ… βœ… ❌ ❌ ❌ Built-in local AI agent, offline, no tracking πŸ”—
Fellou Windows, macOS Free 4 tasks/day / Pro $20/mo ❌ ❌ βœ… ❌ ❌ ❌ Complex multi-step automation, agentic tasks πŸ”—
Arc Max macOS, Windows Free ❌ ❌ ⚠️ ❌ ❌ ❌ AI-enhanced browsing, pinch-to-summarize, Ask on Page ❌
Maxthon Windows, macOS, iOS, Android Free / Premium ❌ ❌ ⚠️ ❌ ❌ ❌ MaxAsk AI answers, built-in VPN, ad-blocker, resource sniffer ❌
ChatGPT Atlas macOS Free (with ChatGPT subscription) ❌ ❌ βœ… ❌ ❌ ❌ OpenAI integration, macOS computer use overlay πŸ”—
AnythingLLM Windows, macOS, Linux Free (OSS) βœ… βœ… ⚠️ βœ… (local API) ❌ ❌ All-in-one desktop AI, document chat, local + API πŸ”—
BrowserGPT iOS, Android Free / Premium ❌ ❌ ⚠️ ❌ ❌ ❌ Mobile-first AI browser ❌
Sidekick Browser Windows, macOS, Linux Free / Pro $10/mo ❌ ❌ βœ… ❌ ❌ ❌ AI assistant, natural language tab management, summarize, automate tasks ❌
Browser Operator Windows, macOS Free (OSS) βœ… βœ… (Ollama, 100+ models) βœ… ❌ βœ… βœ… Multi-agent automation, privacy-first local processing, persistent memory πŸ”—
Browserless Agent ⭐ Cloud (MCP) Free tier / paid ❌ ❌ βœ… βœ… (MCP-native) ❌ ❌ Fastest MCP browser agent; stateful sessions; command batching; released 2026-06-12 πŸ”—
Kane CLI ⭐ Windows, macOS, Linux Free ❌ ❌ βœ… ❌ ❌ ❌ Browser automation testing for AI agents; deterministic pass/fail; released 2026-05-24 πŸ”—
Browsewright ⭐ Python Free (MIT) βœ… βœ… βœ… βœ… (MCP) ❌ ❌ Open-source browser agent; LLM-driven form filling; structured JSON output; released 2026-06-16 πŸ”—
Lightpanda Agent ⭐ Windows, macOS, Linux Free (OSS) βœ… βœ… βœ… ❌ ❌ ❌ Native headless browser agent; PandaScript reproducible scripts; LLM at build-time not runtime; released 2026-06-17 πŸ”—
Veilbrowser ⭐ TypeScript Free (MIT) βœ… βœ… βœ… βœ… (MCP) ❌ ❌ Stealth browser for AI agents; real Chrome over raw CDP; passes sannysoft 57/57; released 2026-06-09 πŸ”—
Dev Browser ⭐ Claude Code plugin Free ❌ βœ… βœ… ❌ ❌ ❌ Browser automation plugin for Claude Code; stateful Playwright wrapper; released 2026-06-19 πŸ”—
Intuned ⭐ Python/Cloud Free tier / paid ❌ βœ… βœ… βœ… ❌ ❌ YC-backed browser automation as code; AI agent generates Playwright code; auto-healing; released 2026-06-08 ❌
Webwright (Microsoft) ⭐ Python Free (MIT) βœ… βœ… βœ… ❌ ❌ ❌ SOTA on long-horizon web tasks; terminal-based; Codex/Claude Code plugins; released 2026-04-08 πŸ”—
SimuLang ⭐ Python/TypeScript Free (OSS) βœ… βœ… βœ… ❌ ❌ ❌ Playwright for the entire desktop; browser + native apps + OS workflows; released 2026-05-19 πŸ”—

Browser Extensions

Extension Pricing Free Multi-Agent Best For GitHub
OpenDia Free (OSS) βœ… ❌ Open alternative to Dia/Comet; anti-detection for X/LinkedIn/Facebook; works with Chrome/Firefox/Edge/Brave πŸ”—
Monica.im Freemium (Free + ~$9/mo) ❌ βœ… Chrome extension, no browser switch ❌
Harpa AI Free βœ… ❌ Automation recipes πŸ”—
MultiOn Free/Paid ⚠️ βœ… Complex tasks πŸ”—
NanoBrowser Free βœ… βœ… Local control, Ollama πŸ”—
Neobrowser Free (OSS) βœ… ❌ Local LLMs via Ollama, privacy-first, Chrome/Edge ❌
Open Operator Free βœ… ❌ Browserbase-powered, open NL browser control πŸ”—
Openator Free (OSS) βœ… ❌ Docker-based headless NL browser agent πŸ”—

Developer Libraries

Consumer-focused AI browsers (for example Perplexity Comet) typically bundle chat with browsing but enforce vendor-side quotas on usage tiers. The stacks below target developers: scale via your LLM keys (BYOK where offered), metered cloud APIs, or self-hosted / OSS runtimes so automation limits follow your infrastructure and providersβ€”not the same caps as end-user browser apps.

Library Language Pricing Best For API Access Multi-Agent Parallel Sessions GitHub
Chrome DevTools MCP TypeScript Free (OSS) AI web debugging, 29 DevTools ❌ ❌ ❌ πŸ”—
Cloudflare Browser Run Cloud API Free Workers / paid browser hours CDP + MCP, WebMCP, Live View, Human-in-the-Loop, session recordings; /json Quick Action extracts schema-shaped data with Workers AI or BYOK (2026-07-01) βœ… ❌ βœ… πŸ”—
Browser-use ⭐ Python / Cloud API Free OSS / Cloud free tier: 3 concurrent sessions and 10 agent tasks/mo; paid cloud from $40/mo or PAYG credits Browser Use CLI 3.0 powered by Browser Harness; browser-use skill installs agent skills for Claude Code, Codex, Cursor, Gemini, OpenCode, and related skill directories; managed cloud adds stealth browsers, CAPTCHA solving, scheduling, persistent memory, skills, and integrations βœ… βœ… βœ… πŸ”—
Browser Harness ⭐ Python / CDP Free (OSS) Thin self-healing CDP harness for direct LLM control of real Chrome; packaged skill for coding agents; pinned by Browser Use CLI 3.0 βœ… ❌ βœ… πŸ”—
Vercel agent-browser ⭐ Rust / Node CLI Free (OSS) Browser automation CLI for AI agents; compact text snapshots, ref-based actions, Chrome and Lightpanda engines, packaged agent skill βœ… ❌ βœ… πŸ”—
Stagehand v3.6 ⭐ TypeScript/Python Free (OSS) v3.6 (2026-06-19): WebMCP support; Claude Fable 5 support; Microsoft Entra ID auth; full CDP rewrite, 44% faster, action caching, hybrid deterministic + AI βœ… ❌ βœ… πŸ”—
LaVague Python Free (OSS) NL to code ❌ ❌ βœ… πŸ”—
Skyvern Python Free tier / $29–$149/mo CV-based automation, Ollama support βœ… βœ… βœ… πŸ”—
Notte Python/Cloud Free tier / $29/mo+ Deterministic replay, demoβ†’script βœ… ❌ βœ… πŸ”—
Firecrawl Python / CLI Free tier / $49/mo+ LLM-powered crawling & scraping βœ… ❌ βœ… πŸ”—
Playwright MCP TypeScript Free (OSS) Cross-browser automation, VS Code βœ… ❌ βœ… πŸ”—
Langflow Python Free (OSS) / Cloud $29/mo Visual multi-agent & RAG workflows βœ… βœ… βœ… πŸ”—
LlamaIndex Python Free (OSS) / Cloud $29/mo Document-heavy RAG, retrieval quality βœ… βœ… βœ… πŸ”—
Haystack Python Free (OSS) / Cloud $49/mo Regulated deployments, structured pipelines βœ… βœ… βœ… πŸ”—
AgentQL TypeScript/Python Free (1K req/mo) / $49/mo / $149/mo Natural language web querying/automation βœ… βœ… βœ… πŸ”—
ScrapeGraphAI Python Free OSS / Cloud $29/mo Natural language web scraping βœ… βœ… βœ… πŸ”—
WebVoyager Python Free (OSS) Autonomous web browsing research ❌ ❌ βœ… πŸ”—
Anchor Browser TypeScript, Python, REST Free credits tier / usage-based plans Managed browser fleets + AI steps for computer-use-style web tasks; see docs.anchorbrowser.io βœ… βœ… βœ… πŸ”—
Hyperbrowser TypeScript, Python, REST Free credits + pay-as-you-go (see hyperbrowser.ai/pricing) Hosted agents: Browser-Use, Claude/OpenAI/Gemini computer use, Stagehand, HyperAgent; optional BYOK for LLM βœ… βœ… βœ… πŸ”—
Steel Python, Node.js, REST Hobby credits + paid tiers; OSS self-host via Docker (steel.dev, docs.steel.dev) Sessions API for Playwright/Puppeteer/Selenium; self-hosted steel-browser avoids vendor browser caps; agent cookbooks βœ… ⚠️ βœ… πŸ”—

Cloud Automation

Service Platform Pricing Best For GitHub
ChatGPT agent ChatGPT Plus $20/mo / Pro $200/mo / Team Guided browser tasks, research, forms, spreadsheets, code execution ❌
Gemini Spark ⭐ Google AI Ultra Included with Google AI Ultra ($99.99/mo, US beta from 2026-05-26) 24/7 always-on personal AI agent; runs on Google Cloud VMs while device is off; MCP connections to Canva, OpenTable, Instacart at launch; Adobe, GitHub, Notion, Slack confirmed summer 2026 ❌
Claude Cowork Anthropic (desktop app) Included with Claude paid plans (GA 2026-04-09) Desktop agent: works in local files & apps, multi-step task completion, Dispatch (assign by text/voice), runs while user is away ❌
Project Mariner Google AI Ultra Included with Google AI Ultra ($99.99/mo) Multi-step browser tasks, shopping, and reservations ❌
Skyvern Cloud Cloud API Free 1K credits / Hobby $29/mo / Pro $149/mo Resilient CV-based automation, Ollama support πŸ”—
Browserbase Cloud API Free credits / paid tiers Stealth mode, session recording, CDP access ❌
Amazon Nova Act API (AWS) Pay-per-use Autonomous web agent, multi-step browser workflows ❌
Amazon Bedrock AgentCore ⭐ AWS Pay-per-use Platform for building, connecting, securing, tracing, and operating production agents with any model or framework; includes Runtime, Gateway, Memory, Browser tool, Code Interpreter, Identity, and Observability πŸ”—
Vercel Sandbox ⭐ Vercel Usage-based Vercel platform Agent/runtime sandbox for code execution; FUSE mounts for S3, network, and custom filesystems added 2026-07-03; pairs with AI SDK and eve workflows πŸ”—
OpenAI Workspace Agents ChatGPT Business, Enterprise, Edu, Teachers Credit-based after preview Research-preview workplace agents usable in ChatGPT and Slack channels; role-based admin controls πŸ”—
Claude Science ⭐ Anthropic Applications through 2026-07-15 00:00 UTC AI workbench for scientists; auditable artifacts, package/tool integration, flexible compute access, up to $30K Claude credits for selected projects πŸ”—
Databricks Genie One ⭐ Databricks Free $10/user/mo GA 2026-06-16; data-smart agentic coworker; iOS/Android/web; 50+ app integrations; Genie Ontology context layer ❌
Thoughtworks Agent/works ⭐ Multi-cloud Enterprise GA 2026-06-16; governed runtime for enterprise agents; single control plane; any cloud deployment ❌
Kore.ai Artemis ⭐ Azure (initial) Enterprise GA 2026-05-21; AI-native agent platform; governance + observability; 40+ channels; Microsoft Foundry integration ❌
Huawei Cloud AgentArts ⭐ Huawei Cloud Enterprise GA 2026-06-08; enterprise-grade agent platform; open-source edition (openJiuwen); AgentArts Orchard portal ❌
Zensar ZenseAI.AgentMesh ⭐ Multi-cloud Enterprise GA 2026-06-19; 80+ pre-built agents; governed orchestration; EU AI Act alignment ❌
Sedai AI Agent Optimization ⭐ Multi-cloud Early access GA planned late 2026; smart model routing; governance + observability; supports OpenAI/Bedrock/Vertex/Azure ❌
GMI Cloud AgentBox ⭐ GMI Cloud Usage-based GA 2026-06-08; marketplace for AI agents; 100+ models via single key; MaaS inference ❌
ClickHouse Agents ⭐ ClickHouse Free (public beta) GA 2026-06-09; Claude-powered agentic analytics; no-code agent builder; MCP-compatible ❌
Google Data Agents ⭐ Google Cloud Free (preview) GA 2026-06-16; Data Science Agent, Database Agent, Looker Dashboard Agent, Data Insights Agent, Deep Research Agent ❌
C1 Autonomous Worker ⭐ C1 Enterprise GA 2026-06-15; agent that runs work loops, carries state, executes code; available via Slack ❌

Autonomous Agents β€” Plain English Prompts πŸ€–

Control a computer or cloud sandbox using plain English text β€” no coding required. Just describe the task and the agent handles everything: clicking, typing, navigating, running code, and completing multi-step workflows.

Legend: πŸ–₯️ = runs on your physical computer | ☁️ = cloud/sandbox computer | 🌐 = controls a browser | πŸ” = multi-agent/parallel | πŸ’¬ = simple English prompt | πŸ”“ = free/open-source | πŸ’° = paid


☁️ Cloud Sandbox Computer Use (English Prompts)

These services run in a cloud sandbox (virtual Linux/Windows desktop), control the computer for you, and are driven purely by natural language instructions.

Agent Interface Pricing Multi-Agent Parallel Sessions Local LLM English Prompt GitHub
Suna (Kortix) Web dashboard Free (OSS self-host); Cloud from $20/mo; Pro $29/mo; Team $199/mo βœ… βœ… ❌ βœ… πŸ”—
OpenManus Web dashboard Free OSS βœ… βœ… ❌ βœ… πŸ”—
Manus AI Web dashboard Free (300 credits/day) / Plus $20/mo / Pro $200/mo βœ… βœ… ❌ βœ… ❌
ChatGPT Agent ChatGPT Web/App Plus $20/mo / Pro $200/mo ❌ ❌ ❌ βœ… ❌
Gemini Computer Use API / AI Studio Gemini Pro $19.99/mo / API metered ❌ ❌ ❌ βœ… ❌
Devin Web dashboard Core $20/mo ($2.25/ACU) / Team $500/seat/mo ❌ ❌ ❌ βœ… ❌
OpenHands Web UI / CLI Free OSS / Cloud Individual free βœ… βœ… βœ… (any API) βœ… πŸ”—
E2B Desktop Sandbox API / SDK Hobby free / Pro $150/mo ❌ βœ… (via code) βœ… βœ… πŸ”—
Cua (trycua) CLI / Python SDK Free (OSS) ❌ βœ… βœ… βœ… πŸ”—
Airtop Web dashboard / API Starter $26/mo (3 sessions) / Pro $80/mo (30 sessions) ❌ βœ… ❌ βœ… ❌
Skyvern Cloud Web dashboard / API Free 1K credits / Hobby $29/mo / Pro $149/mo ❌ βœ… βœ… (Ollama) βœ… πŸ”—
Convergence Proxy Web / API Free tier / Pro $20/mo (acquired by Salesforce) ❌ ❌ ❌ βœ… ❌
Amazon Nova Act API (AWS) Pay-per-use (AWS pricing) ❌ βœ… ❌ βœ… ❌
Project Mariner Google AI Ultra Included ($100/mo Ultra 5Γ— plan) ❌ ❌ ❌ βœ… ❌
Perplexity Computer Web dashboard Perplexity Pro $20/mo ❌ ❌ ❌ βœ… ❌
OpenAI Computer Use (API) API / ChatGPT $15/M input, $60/M output βœ… βœ… ❌ βœ… ❌
OpenHands Agent Canvas ⭐ Web / Self-hosted Free (OSS) Workspace for creating automations integrating Slack/GitHub; self-hostable; released 2026-06-16 πŸ”—

πŸ–₯️ Local Machine / Physical Computer Use

These agents run on your own machine, see your screen, and control your keyboard/mouse β€” no cloud required.

Agent Windows macOS Linux Dashboard/UI CLI API/LLM Multi-Agent Parallel Sessions Pricing GitHub
Claude Computer Use βœ… βœ… βœ… ❌ βœ… (API) Claude API ($3–$15/M) ❌ ❌ Claude API ($3–$15/M tokens) Commercial
Agent TARS (ByteDance) βœ… βœ… βœ… βœ… Web UI βœ… npx @agent-tars/cli@latest Any LLM ❌ βœ… Free (OSS) πŸ”—
Aiden ❌ ❌ βœ… βœ… Desktop βœ… Local / BYO API ❌ ❌ Free (OSS) πŸ”—
Simular Agent S2 βœ… βœ… βœ… ❌ βœ… Any LLM API ❌ ❌ Free (OSS) πŸ”—
UI-TARS Desktop (ByteDance) βœ… βœ… βœ… βœ… Desktop app ❌ UI-TARS-2 model ❌ ❌ Free (OSS) πŸ”—
Open Interpreter βœ… βœ… βœ… βœ… Web βœ… interpreter Any (OpenAI, Claude, local) ❌ ❌ Free (OSS) πŸ”—
Open-Interface βœ… βœ… βœ… ❌ βœ… GPT-4V / any vision LLM ❌ ❌ Free (OSS) πŸ”—
Agent S / S2 βœ… βœ… βœ… ❌ βœ… Any LLM API ❌ ❌ Free (OSS) πŸ”—
UFO (Microsoft) βœ… ❌ ❌ βœ… UI βœ… GPT-4V / Azure ❌ ❌ Free (OSS) πŸ”—
Windows-Use βœ… ❌ ❌ ❌ βœ… Any vision LLM ❌ ❌ Free (OSS) πŸ”—
Bytebot ❌ ❌ βœ… βœ… (Docker) βœ… Any LLM ❌ ❌ Free (OSS) πŸ”—
OpenCUA βœ… βœ… βœ… ❌ βœ… Any ❌ ❌ Free (OSS) πŸ”—
Windows Agent Framework (Microsoft) βœ… ❌ ❌ Agent Designer in VS 2026 βœ… wagent Any LLM ❌ ❌ Free (MIT) πŸ”—
Khoj βœ… βœ… βœ… βœ… Web UI βœ… Any (Ollama, LM Studio, OpenAI) ❌ ❌ Free (OSS) / Cloud $10/mo πŸ”—
Eigent (CAMEL-AI) βœ… βœ… βœ… ❌ βœ… Any LLM βœ… βœ… Free (OSS) πŸ”—

🌐 Browser-Only Agents

Control a browser with natural language β€” click, fill forms, scrape, automate. No script writing needed.

Agent Type Pricing Dashboard CLI Multi-Agent Parallel Sessions Local LLM GitHub
Browser-use ⭐ OSS Python lib + Cloud Free OSS / Cloud free tier: 3 concurrent sessions and 10 tasks/mo; paid cloud from $40/mo or PAYG credits βœ… Cloud βœ… CLI 3.0 βœ… βœ… βœ… (Ollama) πŸ”—
Stagehand OSS TypeScript Free (OSS) ❌ βœ… ❌ βœ… βœ… πŸ”—
NanoBrowser Chrome extension Free (OSS) βœ… Extension ❌ βœ… ❌ βœ… (Ollama) πŸ”—
Skyvern Python / Cloud Free tier / $29–$149/mo βœ… Cloud βœ… βœ… βœ… βœ… (Ollama) πŸ”—
Openator Python Free (OSS) ❌ βœ… ❌ βœ… βœ… πŸ”—
Open Operator Web UI Free βœ… ❌ ❌ ❌ ❌ πŸ”—
Airtop Web / API $26–$80/mo βœ… ❌ βœ… βœ… ❌ ❌
MultiOn API / Chrome ext Free / Paid βœ… ❌ βœ… ❌ ❌ πŸ”—

πŸ” Multi-Agent / Parallel Agent Platforms (Plain English Orchestration)

Coordinate multiple AI agents in parallel to complete complex workflows β€” driven by plain English goals.

Platform Type Dashboard CLI Cloud Local LLM Parallel Pricing GitHub
CrewAI Multi-agent OSS + Cloud βœ… AMP Studio βœ… crewai βœ… AMP βœ… βœ… Free OSS / Starter $99/mo / Pro $299/mo / Enterprise custom πŸ”—
Microsoft Agent Framework (AutoGen + Semantic Kernel unified) Multi-agent conversations (v1.0 GA April 2026) βœ… AG2 Studio βœ… Python βœ… Azure βœ… βœ… Free (OSS) / Azure pay-per-token πŸ”—
LangGraph Stateful agent graphs βœ… LangSmith βœ… Python βœ… Cloud βœ… βœ… Free OSS / Professional $99/mo πŸ”—
OpenHands Dev-focused multi-agent βœ… Web UI βœ… βœ… Cloud βœ… βœ… Free (OSS + Cloud free tier) πŸ”—
OWL (Camel-AI) Distributed multi-agent ❌ βœ… Python ❌ βœ… βœ… Free (OSS) πŸ”—
Manus AI Cloud multi-agent βœ… Web ❌ βœ… ❌ βœ… Free 300 credits/day / $20–$200/mo ❌
n8n Workflow + AI agents βœ… Visual canvas βœ… n8n βœ… Cloud βœ… (Ollama node) βœ… Free OSS / Starter $24/mo / Pro $60/mo πŸ”—
Devin Software engineering βœ… Web ❌ βœ… ❌ ❌ Core $20/mo ($2.25/ACU) / Team $500/seat ❌
Smolagents (HuggingFace) Lightweight code agents ❌ βœ… Python ❌ βœ… ⚠️ Free (OSS) πŸ”—
C3 Code (C3 AI) Enterprise AI dev platform βœ… Web ❌ βœ… ❌ βœ… Enterprise pricing Full SDLC agents, governed deployment, natural language to production apps
Dify Visual LLM platform βœ… Web UI βœ… βœ… Cloud βœ… βœ… Free OSS / Cloud plans πŸ”—
OpenAI Agents SDK Agent handoffs + tool orchestration ❌ βœ… Python / TS βœ… OpenAI API ❌ βœ… Free (OSS) / OpenAI API costs πŸ”—
Google ADK Hierarchical agent tree, A2A protocol βœ… Vertex AI βœ… Python βœ… Vertex AI ❌ βœ… Free (OSS) / Vertex AI pay-per-token πŸ”—
ADK Go 2.0 ⭐ Graph-based agent workflow engine β€” βœ… go run βœ… Google Cloud ❌ βœ… Free (OSS) πŸ”—
Genkit Agents ⭐ Full-stack agent framework βœ… Developer UI β€” β€” βœ… βœ… Free (OSS) πŸ”—
AI SDK 7 ⭐ Production agent platform (TypeScript) β€” βœ… @ai-sdk/tui βœ… Vercel βœ… βœ… Free (OSS) πŸ”—
Mastra ⭐ Agent framework with file-based agents, skills, workspace, and subagents β€” βœ… mastra βœ… Vercel βœ… βœ… Free (OSS) πŸ”—
Ruflo Multi-agent orchestration ❌ βœ… Python ❌ βœ… βœ… Free OSS / Cloud plans πŸ”—
Vercel eve ⭐ Open-source agent framework with Agent Runs traces via Vercel MCP/CLI βœ… Web βœ… vercel agent-runs βœ… Vercel βœ… βœ… Free (OSS) / Vercel hosting costs πŸ”—
Omnigent (Databricks) ⭐ Meta-harness for agents βœ… Web/App βœ… βœ… βœ… βœ… Free (Apache 2.0) πŸ”—
GitKraken Kepler ⭐ Agentic Development Environment βœ… Desktop βœ… ❌ βœ… βœ… Free tier / paid plans πŸ”—
LangChain Deep Agents ⭐ Long-horizon agent harness ❌ βœ… Python ❌ βœ… βœ… Free (MIT) πŸ”—
Wheelie (Continua) ⭐ Agentic dev OS + service runtime βœ… Web βœ… βœ… VMs βœ… βœ… Free (waiting list) πŸ”—
Cloudflare Flue ⭐ Agent framework on Workers βœ… Dashboard βœ… βœ… Cloudflare ❌ βœ… Free (OSS) / Workers pricing πŸ”—
C1 Autonomous Worker ⭐ Enterprise agent βœ… (Slack) ❌ βœ… ❌ ❌ Enterprise ❌
ClickHouse Agents ⭐ Agentic analytics βœ… ❌ βœ… ❌ βœ… Free (public beta) ❌
Google Data Agents ⭐ Data agents βœ… ❌ βœ… ❌ βœ… Free (preview) ❌
Salesforce Agentforce 3 ⭐ MCP-native agent platform βœ… Web ❌ βœ… ❌ βœ… Enterprise pricing ❌
Kore.ai Artemis ⭐ AI-native agent platform βœ… ❌ βœ… (Azure) ❌ βœ… Enterprise ❌
Konecta Kolibri ⭐ Agentic AI orchestration βœ… ❌ βœ… ❌ βœ… Enterprise ❌
Zensar ZenseAI.AgentMesh ⭐ Enterprise agentic AI βœ… ❌ βœ… ❌ βœ… Enterprise ❌
Sedai AI Agent Optimization ⭐ Agent optimization βœ… ❌ βœ… ❌ βœ… Early access ❌

Multi-Agent & Parallel Execution Summary

Tools supporting parallel agent orchestration (βœ…) vs single-agent only (❌):

Category Supports Parallel Agents Tools
Cloud Sandbox βœ… Manus AI, OpenHands, E2B Desktop Sandbox, Cua (trycua), Airtop, Skyvern Cloud, Amazon Nova Act, Perplexity Computer, OpenAI Computer Use (API)
Cloud Sandbox ❌ ChatGPT Agent, Gemini Computer Use, Devin, Convergence Proxy, Project Mariner
Local Machine βœ… Agent TARS, E2B Desktop Sandbox, Cua (trycua)
Local Machine ❌ Claude Computer Use, UI-TARS Desktop, Open Interpreter, Open-Interface, Agent S/S2, UFO, Windows-Use, Bytebot, OpenCUA, Khoj
Browser-Only βœ… Browser-use, Skyvern, Airtop, MultiOn, Browser Operator
Browser-Only ❌ Stagehand, NanoBrowser, Openator, Open Operator
Developer Libraries βœ… Browser-use, Browser Harness, Vercel agent-browser, Skyvern, Cloudflare Browser Run, Langflow, LlamaIndex, Haystack, AgentQL, ScrapeGraphAI, WebVoyager, Hyperbrowser, Anchor Browser, Steel
Developer Libraries ❌ Chrome DevTools MCP, Stagehand, LaVague, Notte, Firecrawl, Playwright MCP
Multi-Agent Platforms βœ… CrewAI, Microsoft Agent Framework, LangGraph, OpenHands, OWL, Manus AI, n8n, Smolagents, Dify, OpenAI Agents SDK, Google ADK, ADK Go 2.0, Genkit Agents, AI SDK 7, Mastra, Vercel eve
Multi-Agent Platforms ❌ Devin

AI Infrastructure πŸ—οΈ

Tools, frameworks, and specialized models for building production AI systems β€” from embeddings and video generation to safety, evaluation, and model routing.

Embedding & Reranking Models 🧲

Specialized models for converting text (or images) into dense vector representations and for reranking retrieval results. Essential infrastructure for RAG pipelines and semantic search. Prices as of July 2026.

Embedding Models

Model Developer Dimensions Max Tokens Pricing Best For GitHub
Gemini Embedding 2 Google 3,072 8,192 text / 6 images / 120s video / 180s audio $0.20/1M text tokens; GA 2026-04-22 First natively multimodal embedding β€” text, image, video, audio, PDF in one space β€”
Granite Embedding Multilingual R2 (311M) ⭐ IBM 768 32,768 Free (Apache 2.0) MTEB Multilingual Retrieval 65.2 (#2 open <500M); 200+ languages; 32K context (64x R1) πŸ”—
Granite Embedding Multilingual R2 (97M) ⭐ IBM 384 32,768 Free (Apache 2.0) Best sub-100M multilingual embedder (MTEB 60.3); 200+ languages πŸ”—
pplx-embed-v1 (4B) Perplexity 3,072 8,192 API pay-per-use MTEB Multilingual leader; web-scale retrieval; contextual variant available πŸ”—
LFM2.5-Embedding-350M ⭐ Liquid AI 1,024 8,192 Free (open-source) 350M-param bidirectional embedder; 11 languages; best-in-class multilingual; runs on CPU/edge πŸ”—
LFM2.5-ColBERT-350M ⭐ Liquid AI Per-token vectors 8,192 Free (open-source) First bidirectional ColBERT; word-by-word matching for higher accuracy; 11 languages; runs on CPU/edge πŸ”—
ML-Embed-0.6B (CodeFuse) ⭐ CodeFuse AI 1,024 8,192 Free (Apache 2.0) 3D Matryoshka Learning; multilingual; ICML 2026 paper πŸ”—
Qwen3-Embedding-8B Alibaba / Qwen 4,096 (flex) 32K Free (Apache 2.0) MTEB leader (70.6); 100+ languages, instruction-aware, flexible dimensions; 0.6B/4B/8B sizes πŸ”—
jina-embeddings-v5-omni-small ⭐ Jina AI / Elastic 1,024 (Matryoshka 32-1,024) 32K text; image, video, audio, PDF API pay-per-use Multimodal embedding model for text, image, video, audio, and PDF search; frozen v5 text backbone keeps text outputs aligned with v5-text-small πŸ”—
jina-embeddings-v4 Jina AI 32–2,048 (flex) 8,192 API pay-per-use Built on Qwen2.5-VL-3B (3.8B); 3 LoRA adapters (query/passage/matching) πŸ”—
text-embedding-3-small OpenAI 1,536 8,191 $0.02/1M tokens Cost-effective English embeddings β€”
text-embedding-3-large OpenAI 3,072 8,191 $0.13/1M tokens Highest-quality English retrieval β€”
Embed v4 Cohere 1,536 128K $0.12/1M (text), $0.47/1M (image) Multimodal text + image RAG β€”
voyage-3-large Voyage AI 256–2,048 (flex) 32K ~$0.18/1M tokens Highest-quality retrieval, long context β€”
BGE-M3 BAAI 1,024 8,192 Free (open-source) Multi-functional: dense + sparse + ColBERT πŸ”—
Nomic Embed v2 (MoE) Nomic AI 256–768 (flex) 512 Free (open-source) Multilingual, MoE efficiency (305M active) πŸ”—
text-embedding-005 Google (Vertex AI) 768 2,048 $0.10/1M tokens GCP-native semantic search β€”

Reranking Models

Model Developer Max Tokens Pricing Best For GitHub
Rerank 4.0 Pro Cohere 32K $1.00/1K queries High-accuracy domain-specific reranking β€”
Rerank 4.0 Fast Cohere 32K $0.50/1K queries Low-latency production reranking β€”
rerank-2.5 Voyage AI 32K API pay-per-use Instruction-following, multilingual β€”
Qwen3-Reranker-4B Alibaba / Qwen 32K Free (Apache 2.0) MTEB-R 69.76, MMTEB-R 72.74; 100+ languages; instruction-following πŸ”—
BGE Reranker v2-m3 BAAI 8,192 Free (open-source) Open-source cross-encoder reranking πŸ”—
Jina Reranker v2 Jina AI 8,192 API pay-per-use Multilingual, long-context reranking β€”

Video Generation Models 🎬

Text-to-video and image-to-video generation models for creating short clips from prompts. The field is moving rapidly β€” resolutions, durations, and pricing change frequently. Specs as of July 2026.

Model Developer Resolution Duration Pricing Open Source Best For GitHub
Sora 2 (deprecated) OpenAI Up to 1080p Up to 20s App shut down 2026-04-26; API shutdown 2026-09-24 No Deprecated β€” use Veo 3 or Luma Ray3 β€”
Veo 3.1 Google DeepMind 1080p Up to 8s (extendable) ~$0.20–$0.40/s No Native 48kHz audio + video, realistic physics; only model with synced dialogue β€”
Gemini Omni Flash ⭐ Google DeepMind 1080p Public preview $0.10/s No High-quality video generation + conversational editing; text/image/video input; released 2026-06-30 β€”
Seedance 2.0 ByteDance 1080p Up to 16s API pay-per-use No #1 Artificial Analysis leaderboard (Elo 1213, June 2026); accepts 9 images + 3 clips + 3 audio β€”
Luma Ray3.14 Luma AI Up to 1080p Up to 10s $29–$99/mo No Native 1080p, 4Γ— faster at 720p vs Ray2, 3Γ— lower cost β€”
LTX-2 Lightricks Up to 4K Up to 20s @ 50fps Free (open-source; commercial license for >$10M ARR) Yes (Apache 2.0) First open-source 4K audio+video sync model, 19B DiT params, runs on consumer GPU πŸ”—
Runway Gen-4 / Gen-4.5 Runway Up to 4K Up to 16s $12–$76/mo No Professional creative workflows β€”
Kling 3.0 Kuaishou 1080p Up to 15s Free / $5.99–$66/mo No Multi-Shot Storyboard, best value-for-money; released 2026-02-04 β€”
Pika 2.0 Pika Labs 1080p Up to 5s Free / $8–$58/mo No Social media, creative effects β€”
MiniMax Video-01 MiniMax 720p Up to 6s ~$0.40/video No Strong text-motion responsiveness β€”
HunyuanVideo Tencent 720p–2K Up to 16s Free (self-host; ~60GB VRAM) Yes (Apache 2.0) High per-frame fidelity, long clips πŸ”—
Wan 2.2 (14B) Alibaba 480p–1080p Up to 10s ~$0.10–$0.30/clip (API) Yes (Apache 2.0) Motion quality, VBench #1 benchmark πŸ”—
MiniMax Hailuo 2.3 ⭐ MiniMax 1080p Up to 10s API pay-per-use No Released 2026-06-22; successor to Hailuo 02 β€”
MiniMax Hailuo 02 ⭐ MiniMax 1080p Up to 10s API pay-per-use No 3x params vs predecessor, 4x training data, best complex instruction adherence; #2 Artificial Analysis Video Arena; released 2026-06-20 β€”
Kling 3.0 Turbo ⭐ Kuaishou 480p–720p 1–15s preview Free / paid No Fast-preview mode for rapid iteration; released 2026-06-17 β€”
MiniMax Video-01 ⭐ MiniMax 720p Up to 6s ~$0.40/video No First AI-native video model from MiniMax; text-to-video + image-to-video; released 2026-06-20 β€”
Grok Imagine Video 1.5 ⭐ xAI 720p Up to 15s $0.06/s ($4.20/min at 720p) No #1 Image-to-Video Arena leaderboard (Elo 1,473); native audio+video sync; Fast variant 25s for 6s clip; GA 2026-06-16 β€”
HappyHorse-1.0 ⭐ Alibaba (ATH) 1080p Up to 5s Free (open-source) Yes (open-source) #1 on Artificial Analysis video leaderboard; unified Transformer; simultaneous audio/video; 15B params; API coming 2026-04-30 πŸ”—
Seedance 2.1 ⭐ ByteDance 1080p Up to 16s API pay-per-use No Sharper motion, tighter character consistency; available on Atlas Cloud Day 0 β€”
Microsoft Mirage ⭐ Microsoft Research β€” β€” Free (research) Research Video world model with persistent spatial memory; 10.57x faster generation; built on Wan2.2 πŸ”—
MoVerse ⭐ Academic/MSFT β€” β€” Free (research) Research Real-time video world model; panoramic Gaussian scaffold; 8 FPS on RTX 4090 β€”
Echo-Infinity ⭐ JD/Echo Team β€” 24h+ rollouts Free (research) Research Real-time infinite video generation; learnable evolving memory; 1.3M+ frames πŸ”—
JoyAI-Echo ⭐ JD Up to 1080p Minute-level Free (research) Research Multi-shot audio-video generation; 7.5x speedup via DMD distillation πŸ”—
Mochi 1 Genmo 480p Up to 5.4s @ 30fps Free (open-source) Yes (Apache 2.0) High-quality open text-to-video πŸ”—
CogVideoX Zhipu AI / Tsinghua 720p ~6s Free (open-source) Yes (Apache 2.0) Image-to-video quality, LoRA fine-tuning πŸ”—

Speech & TTS Models πŸ”Š

Text-to-speech (TTS) and speech-to-text (STT / ASR) models for voice generation, transcription, and real-time audio. Prices as of July 2026.

Artificial Analysis Speech Arena β€” Top 5 (July 2026)

Rank Model Elo Pricing Languages
πŸ₯‡ #1 Sonic-3.5 ⭐ ~1,220 ~$39 / 1M chars 40+
πŸ₯ˆ #2 Gemini 3.1 Flash TTS ~1,216 $18.30 / 1M chars 70+
πŸ₯‰ #3 Realtime TTS-2 (Research Preview) ~1,208 $25–$35 / 1M chars 100+
#4 Sonic 4 ~1,210 ~$46.70 / 1M chars 40+
#5 Realtime TTS 1.5 Max ~1,200 $35 / 1M chars 100+

Arena Elo scores shift continuously. Treat rankings as point-in-time readings. Source: Artificial Analysis, June 2026. Cartesia positions Sonic-3.5 and Ink-2 as its latest production speech stack for real-time voice agents. Source: πŸ”—

Text-to-Speech (TTS) β€” Proprietary & API

Model Developer Languages Real-time Open Source Pricing Best For GitHub
Sonic-3.5 ⭐ Cartesia 42 Yes (sub-90ms) No See Cartesia pricing Fastest, most natural Cartesia TTS model; ranked #1 for naturalness; GA in 2026 changelog πŸ”—
Gemini 3.1 Flash TTS Google 70+ Yes No $18.30 / 1M chars Audio tags for granular style/pace control, SynthID watermarking β€”
Realtime TTS-2 (Research Preview) Inworld AI 100+ Yes No $25–$35 / 1M chars Realtime conversation, cross-lingual voice identity, emotional perception β€”
Sonic 4 ⭐ Cartesia 40+ Yes No ~$46.70 / 1M chars Sonic 4 Turbo ~40ms TTFA (May 2026); 40+ languages, 95% world pop; instant 3s voice clone β€”
Realtime TTS 1.5 Max Inworld AI 100+ Yes No $35 / 1M chars Realtime conversational agents, low latency + low cost β€”
Eleven v3 ElevenLabs 70+ No No Subscription / API (up to 55% price cut May 2026) Most expressive TTS, inline audio tags [whispers], [laughs], multi-speaker dialogue β€”
Eleven Flash v2.5 ElevenLabs 32 Yes (~75ms) No Subscription / API Ultra-fast real-time, same voice library as offline β€”
Eleven Multilingual v2 ElevenLabs 30 Yes No Subscription / API Emotionally-aware multilingual synthesis β€”
gpt-4o-mini-tts OpenAI 50+ Yes No $0.60 / $12 per 1M (in/out) Natural-language voice steering, 13 built-in voices β€”
GPT-Realtime-2 OpenAI β€” Yes No β€” Speech-to-speech with GPT-5-class reasoning, tool calls, interruptions β€” released 2026-05-07 β€”
GPT-Realtime 2.1 / 2.1-mini ⭐ OpenAI β€” Yes No β€” Newer Realtime API speech-to-speech models with function calling, MCP servers, and SIP; shipped 2026-07-06 πŸ”—
gpt-realtime-mini OpenAI β€” Yes No β€” Cost-efficient speech-to-speech β€”
GPT-Live-1 / 1 mini ⭐ OpenAI β€” Yes (full-duplex) No Bundled in ChatGPT tiers (API pending) Consumer voice model powering ChatGPT Voice; full-duplex, listens while speaking; GPT-5.5 backend delegates complex work; announced 2026-07-08, API "soon" πŸ”—
gpt-audio-mini OpenAI β€” Yes No β€” Cost-efficient audio generation β€”
OpenAI TTS / TTS HD OpenAI 57 Yes No $15 / $30 per 1M chars Enterprise, seamless GPT integration β€”
Grok TTS xAI 25+ Yes No $15.00 / 1M chars Speech tags, 80+ voices, SOC 2 / HIPAA compliant β€”
Deepgram Aura-2 Deepgram 7 Yes (~90–200ms) No $0.030 / 1K chars ($30/1M) Enterprise voice agents, unified STT+TTS stack, on-prem deployment β€”
Hume Octave 2 Hume AI 11+ Yes (~100–200ms) No Varies (contact sales) Emotional intelligence, voice conversion, phoneme editing β€”
Supertonic 3 ⭐ Supertone 31 Yes No Subscription / API Fast, cost-efficient multilingual narration β€”
Lightning V3.1 / V3.2 Smallest.ai 15 Yes No Pay-as-you-go Conversational TTS, MOS 3.89, auto language detection, mid-sentence switching β€”
StepAudio 2.5 Realtime ⭐ StepFun Chinese, English Yes No β€” End-to-end real-time speech LLM, paralinguistic comprehension, persona RLHF β€”
VibeVoice-Realtime-0.5B ⭐ Microsoft English (9 experimental langs) Yes (~300ms) Yes (MIT) Free (self-host) 0.5B real-time TTS; streaming text input; 10-min robust generation; Qwen2.5 base; <300ms TTFA; released 2025-12-03; HuggingFace integration 2026-03-06 πŸ”—
Qwen3-TTS πŸ‡¨πŸ‡³ Alibaba 10 Yes (streaming) Yes (Apache 2.0) Free (self-host) / API Voice design, voice cloning, instruction control, multilingual πŸ”—
Stability Audio 3.0 Stability AI Music/SFX Yes (small/medium OSS) Yes (small/medium, Apache 2.0) Free (small/medium OSS weights); commercial via API Professional-grade music >6 min, open-weight small/medium variants πŸ”—
Sesame CSM Sesame AI Labs English Yes Yes Free Conversational, emotionally expressive (4.7 MOS) πŸ”—
MAI-Voice-2 ⭐ Microsoft 15 Yes No API (Foundry) Most expressive Microsoft TTS; granular emotion control (whisper/sad/etc.); 15 languages; custom voice from 5-60s clip; released 2026-06-02 β€”
Higgs Audio v3 TTS ⭐ Boson AI 100+ Yes No API / free (self-host) Voice chat optimized; inline emotion/style/prosody tags; zero-shot voice cloning; released 2026-06-04 πŸ”—
Chatterbox Multilingual v3 ⭐ Resemble AI 25 Yes Yes (MIT) Free (self-host) / API 0.5B Llama backbone; PerTh watermarking; improved speaker similarity; released 2026-06-10 πŸ”—
MisoTTS ⭐ Miso Labs β€” Yes Yes (Modified MIT) Free (self-host) / API pending 8B emotive TTS; RVQ scales vocabulary; conditions on text + audio; released 2026-06-03 πŸ”—
MiniMax Speech 2.8 MiniMax 40+ Yes No API pay-per-use Native sound tags, high-fidelity cloning, studio-grade clarity; introduced 2026-01-23 00:00 UTC πŸ”—
ZONOS2 ⭐ Zyphra Multilingual Yes Yes (Apache 2.0) Free (self-host) / API 8B MoE TTS; first open-source MoE TTS; multilingual + code-switched; zero-shot voice cloning; ECAPA-TDNN speaker embeddings; released 2026-06-12 πŸ”—
Mistral Voxtral TTS ⭐ Mistral AI 9 Yes Yes (CC BY-NC 4.0) $0.016/1K chars 4B parameter TTS; 70ms latency; voice cloning from 3s audio; released 2026-06-18 πŸ”—
MiniMax Speech-01-HD ⭐ MiniMax 17 Yes No API 300+ pre-built voices; high-fidelity voice cloning from 10s audio; released 2026-06-19 β€”
dots.tts ⭐ RedNote (Xiaohongshu) β€” Yes Yes (Apache 2.0) Free 2B fully continuous AR TTS; 48kHz AudioVAE; no discrete tokens; released 2026-06 πŸ”—

Speech-to-Text (STT / ASR)

Model Developer Languages Real-time Open Source Pricing Best For GitHub
Ink-2 ⭐ Cartesia 40+ Yes (~100ms) No Credit-based Real-time transcription model paired with Sonic-3.5 for voice-agent workflows; launched 2026-06-30 00:00 UTC πŸ”—
MAI-Transcribe-1.5 ⭐ Microsoft 43 Yes No API (Foundry) SOTA on FLEURS; #3 Artificial Analysis; 5x faster than Gemini 3.1; keyword biasing; 1hr audio in <15s; released 2026-06-02 β€”
Cohere Transcribe ⭐ Cohere 14 No Yes (Apache 2.0) Free (self-host) / Model Vault #1 HuggingFace Open ASR Leaderboard (5.42% WER); 2B Conformer; released 2026-03-26
VibeVoice-ASR-7B ⭐ Microsoft 50+ Yes (~15s for 60min audio) Yes (MIT) Free (self-host) 60-min single-pass ASR; speaker diarization + timestamps; 50+ languages; structured Who/When/What output; released 2026-01-21; HuggingFace transformers integration 2026-03-06 πŸ”—
Cohere Transcribe v2 ⭐ Cohere 14 Yes Yes (Apache 2.0) Free (self-host) Real-time streaming STT; open-source; released 2026-05 πŸ”—
Gladia Solaria-3 ⭐ Gladia 5 (EU) Yes No API #1 on business audio; optimized for noisy/real-world European languages; released 2026-06-10 β€”
Speechmatics Melia ⭐ Speechmatics 55+ Yes (preview) No From $0.129/hr (10hr free) Code-switching across 55+ languages; lowest-priced Speechmatics model; released 2026-06-17 β€”
Voxtral Realtime ⭐ Mistral AI 13 Yes Yes (Apache 2.0) Free (self-host) Natively streaming ASR; 480ms delay matches Whisper quality; 4B params; released 2026-02 πŸ”—
Soniox v5 Async ⭐ Soniox β€” No No API Structured speech-to-text; speaker separation; language IDs; normalized structured entities; released 2026-06-11 β€”
NVIDIA Nemotron 3.5 ASR Streaming 0.6B ⭐ NVIDIA 40 Yes Yes (Apache 2.0) Free (self-host) / API 40 language-locales; language-ID prompt conditioning; punctuation/capitalization; released 2026-06-04 πŸ”—
Gnani Prisma v2.5 ⭐ Gnani AI 9 (Indian) Yes No API #1 in 8/9 Indian languages on real-world benchmarks; 15% lower WER for rural Hindi; 14M hours training data; released 2026-06-19 β€”
Grok STT xAI 25+ Yes No $0.10/hr (batch) / $0.20/hr (streaming) Entity recognition (medical/legal/financial), diarization, multichannel, SOC 2 / HIPAA β€”
Whisper large-v3 OpenAI 100+ No Yes (MIT) $0.006/min (API) Open-source multilingual baseline πŸ”—
GPT-4o Transcribe OpenAI 50+ Yes No $0.006/min High-accuracy managed STT β€”
GPT-4o-mini-transcribe OpenAI 50+ Yes No β€” Cost-efficient managed STT β€”
Deepgram Nova-3 Deepgram 36+ Yes No $0.0043/min Ultra-low latency, production STT β€”
Deepgram Nova-3 Fast Deepgram 36+ Yes No β€” Lowest-latency Deepgram tier β€”
AssemblyAI Universal-2 AssemblyAI Multilingual Yes No $0.0025/min Accurate, feature-rich transcription β€”
AssemblyAI Universal-3 Pro ⭐ AssemblyAI Multilingual Yes No API (pay-per-use) New flagship STT; Universal-3 Pro + Multilingual Streaming (May 2026); Voice Agent product; LLM Gateway β€”
Inworld AI Multilingual Streaming ⭐ Inworld AI Multilingual Yes No API Real-time multilingual STT for voice agents; launched May 7, 2026 alongside Universal-3 Pro β€”
ElevenLabs Scribe v2 ElevenLabs 90+ Yes No β€” State-of-the-art transcription, word-level timestamps, diarization β€”
ElevenLabs Scribe v2 Realtime ElevenLabs 90+ Yes (~150ms) No β€” Live transcription, ultra-low latency β€”

TTS Pricing Comparison (Per 1M Characters)

Provider Model Cost / 1M chars Languages Voice Cloning
VibeVoice-Realtime vibevoice-realtime-0.5b Free English (9 experimental) βœ…
Mistral Voxtral TTS voxtral-tts $16 9 βœ…
Fish Audio S2 Pro s2-pro $15 80+ βœ…
OpenAI tts-1 $15 57 ❌
Grok TTS grok-tts $15 25+ βœ…
Gemini 3.1 Flash TTS gemini-3-1-flash-tts $18.30 70+ ❌
Cartesia sonic-3.5 ~$39 40+ βœ… (instant)
Deepgram Aura-2 aura-2 $30 7 ❌
Inworld realtime-tts-1.5-max $35 100+ βœ…
ElevenLabs flash-v2.5 ~$60 (effective) 32 βœ…
MiniMax speech-2.8-turbo $60 40+ βœ…
OpenAI tts-1-hd $30 57 ❌

Use-Case Recommendations

Use Case Top Picks
Real-time voice agents Cartesia Sonic-3.5 (~100ms TTFA), Ink-2 (~100ms STT), Inworld Realtime TTS-2, Deepgram Aura-2 (~90ms)
Long-form narration / audiobooks ElevenLabs v3, Gemini 3.1 Flash TTS, Fish Audio S2 Pro
Multilingual content Gemini 3.1 Flash TTS (70+), ElevenLabs v3 (70+), Fish Audio S2 Pro (80+), Sonic-3.5 (40+)
Emotional fidelity Hume Octave 2 (reads for meaning), ElevenLabs v3 (audio tags), Fish Audio S2 Pro
On-device / low cost Kokoro-82M, MOSS-TTS-Nano, Qwen3-TTS (0.6B), Voxtral TTS (4B), VibeVoice-Realtime (0.5B)
Video dubbing (duration control) IndexTTS-2 (precise duration control + emotion)
Enterprise / compliance Deepgram Aura-2 (on-prem, HIPAA), Grok TTS (SOC 2 / HIPAA)
Open-source production Fish Audio S2 Pro, MOSS-TTS-v1.5, IndexTTS-2, Chatterbox Multilingual v3, ZONOS2

AI Safety & Guardrails πŸ›‘οΈ

Tools and frameworks for detecting unsafe content, preventing prompt injection, validating outputs, and enforcing policy compliance in LLM-powered applications. As of July 2026.

Tool Developer Type Open Source Pricing Best For GitHub
Llama Guard 3 Meta Safety classifier (8B LLM) Yes (Meta license) Free / ~$0.02/1M tokens (API) Input/output safety classification, 8 languages πŸ”—
NeMo Guardrails NVIDIA Programmable guardrail toolkit (Colang DSL) Yes (Apache 2.0) Free Dialog safety, policy enforcement, LangChain-native πŸ”—
OpenAI Privacy Filter OpenAI PII detection & redaction Yes (Apache 2.0) Free (OSS) Detects & redacts personal info in text πŸ”—
Guardrails AI Guardrails AI Python validator framework Yes Free (OSS) Output validation, PII detection, hallucination guards πŸ”—
Amazon Bedrock Guardrails AWS Managed safety layer No Pay-per-use (AWS) AWS-native, zero-ops compliance and content filtering β€”
ShieldGemma 2 Google Safety classifier (open weights) Yes (open weights) Free Text safety (2B/9B/27B), image safety (4B) β€”
LLM Guard Protect AI Open-source middleware toolkit Yes (MIT) Free PII + toxicity filtering; chains multiple scanners; drop-in middleware πŸ”—
Rebuff Protect AI Prompt injection detector Yes Free Self-hardening anti-injection using vector memory πŸ”—
Lakera Guard Lakera Managed LLM security API No Free tier + Enterprise Runtime LLM security, <50ms latency, PII + injection β€”
Sponsio ⭐ Sponsio Labs Runtime contract enforcement Yes (Apache 2.0) Free (OSS) Deterministic agent safety; trajectory-level policy enforcement; <0.01ms per tool call; OWASP Agentic Top 10 bundles; released 2026-05-06 πŸ”—
ToolShield ⭐ CHATS-lab (ICML 2026) Multi-turn safety defense Yes (MIT) Free Training-free MCP tool guard; 30% attack reduction; plug-and-play for 5 coding agents πŸ”—
AgentGuard ⭐ WhitzardAgent Attribute-based access control Yes (GPL-3.0) Free Tool-call-level access control; supports LangChain/AutoGen/OpenAI Agents SDK πŸ”—
Mirage ⭐ ysham123 Policy gateway Yes (MIT) Free Deterministic policy DSL; same YAML for CI and production; no LLM in decision loop πŸ”—
ToolSafe ⭐ MurrayTom Step-level guardrail Yes Free Proactive monitoring + feedback-driven reasoning; reduces harmful tool executions πŸ”—
SafeHarbor ⭐ ljj-cyber (ICML 2026) Memory-augmented guardrail Yes Free Hierarchical Risk Tree; Safety Projector; no fine-tuning of underlying model πŸ”—
NEXUS ⭐ eliashossain001 Runtime safety monitor Yes Free Structured plan IR; 4 graded interventions (allow/block/confirm/revise); deterministic + learned πŸ”—
mech-gov-framework ⭐ Santander AI Lab Mechanical governance Yes (Apache 2.0) Free Model-agnostic governance regimes (R1/R2/R3); hard gates; entropy commit-reveal πŸ”—
Vermillio ⭐ Vermillio Guardrails-as-a-Service SDK No Paid Real-time prompt/response filtering; copyright and data leakage protection; modular safety filters β€”
Amazon Bedrock Guardrails InvokeGuardrailChecks API ⭐ AWS Per-request safeguard checks No Pay-per-use Granular per-turn safety checks; numeric scores; custom thresholds; detect-only mode; released 2026-06-16 β€”
Google AI Control Roadmap ⭐ Google DeepMind Agent security framework No (internal) Free (published) Defense-in-depth; MITRE ATT&CK-based threat modeling; AI supervisor monitoring; D1-D4/R1-R3 levels; published 2026-06-18 β€”

RAG Frameworks πŸ—‚οΈ

Frameworks and libraries for building Retrieval-Augmented Generation (RAG) pipelines β€” connecting LLMs to external knowledge sources. As of July 2026.

Framework Developer Language Key Features Open Source GitHub
LlamaIndex LlamaIndex Python 160+ data connectors, hybrid search, multi-agent support Yes (MIT) πŸ”—
LangChain LangChain AI Python / JS Chains, agents, memory, 50K+ integrations, LangGraph Yes (MIT) πŸ”—
Dify ⭐ LangGenius Python / JS 131K+ GitHub stars; visual workflow builder, RAG pipelines, MCP client/server, multi-agent, Human Input node, 1M+ apps deployed; $30M Series Pre-A (Mar 2026) Yes (Apache 2.0) πŸ”—
RAGFlow InfiniFlow Python Visual workflow builder, deep document parsing (PDF/tables) Yes (Apache 2.0) πŸ”—
Haystack deepset Python Modular pipelines, enterprise-grade, built-in monitoring Yes (Apache 2.0) πŸ”—
Verba Weaviate Python No-code UI, Weaviate-native vector search Yes πŸ”—
Mem0 Mem0 AI Python / JS Persistent memory layer, graph memory, session recall Yes (Apache 2.0) πŸ”—
txtai NeuML Python All-in-one semantic search + workflow automation Yes (Apache 2.0) πŸ”—
R2R SciPhi Python Lightweight, low-latency, REST API, production-first Yes (MIT) πŸ”—
Ultra RAG ⭐ rblake2320 Python 7-stage ingestion + 10-step query pipeline; KG+PPR, RAPTOR, CRAG, Self-RAG, HyDE; adversarial self-testing Yes (MIT) πŸ”—
ForgeRAG ⭐ deeplethe Python Production-ready; structure-aware reasoning; BM25+vector+KG; pixel-precise citations Yes (MIT) πŸ”—
GRIP (ACL 2026) ⭐ WisdomShell Python Retrieval-as-generation; token-level retrieval control; self-triggered information planning Yes (MIT) πŸ”—
MOTHRAG ⭐ Julian Geymonat Python Deterministic multi-hop RAG; research-SOTA parity on commodity APIs; no GPU; proof tree per answer Yes (Apache 2.0) πŸ”—
AkasicDB / Omni RAG ⭐ KAIST/GraphAI SQL/GQL Unified vector-graph-relational DBMS; Omni RAG improves accuracy 78% vs conventional RAG; 20x faster queries Yes (SIGMOD 2026) β€”
Ennoia ⭐ vunone Python Declarative Document Indexing; typed schemas; hybrid filter+vector search; MCP tool surface Yes (Apache 2.0) πŸ”—
VORTEXRAG ⭐ vignesh2027 Python 7-layer pipeline solving semantic drift + context poisoning; EM=74.8 on multi-hop QA Yes (MIT) πŸ”—
MemGraphRAG ⭐ XMUDeepLIT Python Memory-based multi-agent graph RAG; three-layer memory architecture; KDD 2026 Yes (MIT) πŸ”—
UnWeaver ⭐ Academic Python Entity-based decomposition RAG; better than GraphRAG precision without explicit graph; ICLR 2026 Yes πŸ”—
Google Agentic RAG ⭐ Google Cloud Multi-agent workflow with sufficient context agent; +34% accuracy vs standard RAG; public preview 2026-06-05 No β€”

Fine-tuning Platforms βš™οΈ

Tools and platforms for adapting pre-trained LLMs to specific tasks or domains via supervised fine-tuning, RLHF, LoRA/QLoRA, and related methods. Prices as of July 2026.

Platform Type Supported Models Pricing Best For GitHub
Unsloth OSS library Llama, Mistral, Gemma, Qwen, Phi, DeepSeek, GLM, + more Free 2–5Γ— faster training, 80% VRAM reduction; MoE 12Γ— faster (2026), FP8 RL support (1.4Γ— faster, 60% less VRAM); Unsloth Studio web UI; Windows officially supported πŸ”—
Axolotl OSS framework Most Hugging Face models Free Config-as-code (YAML), reproducibility, multi-GPU training πŸ”—
OpenAI Fine-tuning Managed API GPT-4.1, GPT-4.1-mini (SFT/DPO), o4-mini (RFT) GPT-4.1: ~$3.00/1M training tokens; GPT-4.1-mini: ~$0.80/1M Managed, no infra; note: closed to new users as of 2026 β€” existing users only β€”
Google Vertex AI Managed cloud Gemini 3.1 Pro/Flash, Gemma 4 Gemini 3.1 Pro: $25/1M training tokens GCP-native, Gemini model access β€”
SiliconFlow Managed cloud 100+ open-source models Free tier + pay-per-use Managed fine-tuning + inference; 3-step pipeline (uploadβ†’trainβ†’deploy); 2.3Γ— faster than avg cloud; H100/H200/MI300 πŸ”—
Unsloth Studio ⭐ Unsloth Desktop app 500+ models Free (OSS) / Pro / Enterprise No-code local training and inference; runs 100% offline; GGUF/Safetensors; tool-calling; web search; released 2026-06-18
Arkor ⭐ WlyZhang TypeScript Open-weight LLMs Free (alpha) TypeScript framework for fine-tuning; type-safe configs; local Studio; managed GPUs
Langtrain ⭐ Langtrain Python 20+ open-source models Free (public beta) Sovereign AI platform; fine-tune, align, deploy on own infrastructure; no per-token cost
Tuning Engines ⭐ ShinyLaunch Python 100+ models Free tier + paid Unified AI control and governance layer; single OpenAI-compatible endpoint; fine-tuning + inference + guardrails
Predibase / LoRAX Cloud + OSS server Llama, Mistral, 50+ HF models Free tier + per-GPU pricing Multi-adapter serving: many LoRA adapters on one GPU πŸ”—
PEFT Hugging Face All Hugging Face models Free LoRA, QLoRA, prefix tuning, prompt tuning β€” full HF ecosystem πŸ”—
LLaMA-Factory Community 100+ models Free Web UI, low-code interface, beginner-friendly fine-tuning πŸ”—
torchtune PyTorch Llama, Gemma, Mistral, Phi Free PyTorch-native, composable training recipes πŸ”—
Pioneer (Fastino Labs) ⭐ Managed agentic fine-tuning Qwen, Gemma, Llama, GLiNER API-based First agentic fine-tuning agent; synthetic dataset generation; 10-min fine-tuning; adaptive inference πŸ”—
Langtrain ⭐ Managed platform 20+ open-source models Public beta (free) Sovereign AI platform; LoRA/QLoRA/DPO/GRPO; 100% on-prem; SOC 2 Type II πŸ”—
Fireworks Training ⭐ Managed cloud 100B+ models (Kimi K2.5 1T, etc.) Preview (contact sales) Full-parameter training; custom loss functions; frontier RL; multi-LoRA serving πŸ”—
Together AI Fine-tuning ⭐ Managed cloud 100B+ open-source models Pay-per-use Multi-node orchestration; 100B+ models; fine-tuning + inference in one platform πŸ”—
Tuning Engines ⭐ Unified AI control layer 100+ models Free tier + paid Single OpenAI-compatible endpoint; fine-tuning + routing + guardrails + policy-as-code πŸ”—
NeuralForge ⭐ Local-first platform 1-3B models on consumer GPUs Free (OSS) QLoRA training on consumer GPUs; web UI; GGUF export; single 12GB card sufficient πŸ”—
LLM Fine-Tuner v3.2 ⭐ No-code local tool Most HF models Free (GPL-3.0) No-code web UI; SFT/DPO/RLHF/ORPO; GGUF export; Unsloth 2-5x acceleration πŸ”—

Evaluation & Observability πŸ“Š

Tools for tracing LLM calls, evaluating output quality, debugging RAG pipelines, and monitoring production AI systems. Prices as of July 2026.

Tool Developer Type Open Source Pricing Best For GitHub
LangSmith LangChain AI Tracing + evaluation platform No (enterprise self-host) Free (5K traces/mo), paid plans LangChain apps, chain + agent debugging β€”
Braintrust Braintrust Data Eval-first platform Partial (AI proxy OSS) Free (1M spans), enterprise CI/CD evals, dataset management, LLM-as-judge; Topics active observability (production trace pattern discovery) GA June 2026 β€”
Helicone Helicone Proxy-based observability Yes Free tier, usage-based Cost tracking, request caching, drop-in API proxy πŸ”—
Arize Phoenix Arize AI OSS tracing + evaluation Yes Free (OSS); Arize Cloud paid RAG debugging, LLM-as-judge, local dev πŸ”—
Langfuse Langfuse Tracing + evaluation Yes (MIT) Free / self-host; cloud paid Open-source, 19K+ GitHub stars, OpenTelemetry πŸ”—
MLflow Linux Foundation / Databricks Full AI engineering platform Yes (Apache 2.0) Free (OSS); Databricks managed paid 30M+ monthly downloads; observability, eval, prompt optimization, governance β€” no enterprise paywall πŸ”—
Ragas Ragas RAG evaluation framework Yes Free RAG-specific metrics: faithfulness, recall, precision πŸ”—
DeepEval Confident AI LLM evaluation framework Yes Free (OSS); cloud paid 14+ built-in metrics, pytest-style eval runner, 50+ research-backed metrics, production anomaly detection πŸ”—
Laminar ⭐ Open-source observability Yes (MIT) Free (OSS); managed platform paid Rust-based; ultra-fast; OpenTelemetry-native; traces + evals + AI monitoring; YC S24 πŸ”—
Peekr ⭐ Zero-config observability Yes (MIT) Free Auto-instruments OpenAI/Anthropic/LiteLLM; claim-level hallucination detection; HIPAA/GDPR guardrails πŸ”—
TraceMind ⭐ Open-source eval + observability Yes Free Self-hosted; automatic quality scoring; eval suites; regression alerts; hallucination detection πŸ”—
Agentic CLEAR (ACL 2026) ⭐ IBM Research Multi-level agent eval Yes (OSS) Free Automated multi-level evaluation; dynamic issue discovery; MLflow/Langfuse integration πŸ”—
Styxx ⭐ Cognitive observability Yes (MIT) Free 9 cognometric instruments; hallucination/refusal/tool-call drift/goal-drift detection; per-step localization πŸ”—
Aether ⭐ Local-first cognition debugger Yes Free Chrome DevTools for AI; real-time reasoning tree inspection; VS Code extension; hallucination debugging πŸ”—
Dynatrace dt-evals ⭐ Dynatrace LLM/agent eval from traces Yes (OSS) Free (OSS); cloud paid Eval from real GenAI traces; LLM judge; CI/CD integration; supports OpenAI/Anthropic/Google/AWS/Azure πŸ”—
Currai ⭐ Currai AI observability platform No Paid Prompt tracing; A/B testing; LLM evaluations; cost analytics; OpenTelemetry β€”
tracesage ⭐ Open-source LangGraph observability Yes (MIT) Free Local-first; interactive graph + timeline UI; MCP tool-source attribution; pytest fixture πŸ”—
Opik ⭐ Comet LLM lifecycle platform Yes (Apache 2.0) Free (OSS); cloud paid Evaluation, testing, monitoring, optimization; Opik Guardrails πŸ”—
Orq.ai ⭐ Orq.ai Observability + monitoring No Paid Real-time monitoring; automated evaluations; trace automation; custom dashboards β€”

MCP Ecosystem πŸ”Œ

The Model Context Protocol (MCP) is an open standard originally by Anthropic, now governed by the Agentic AI Foundation (AAIF) under the Linux Foundation (co-founded by Anthropic, Block, and OpenAI; supported by Google, Microsoft, AWS, Cloudflare, Bloomberg). It connects LLMs to external tools and data sources via a unified JSON-RPC 2.0 interface, supporting STDIO and Streamable HTTP transports. The official MCP Registry at registry.modelcontextprotocol.io is currently in preview as the centralized metadata repository for publicly accessible MCP servers, with registry docs at modelcontextprotocol.io/registry. Official registry docs now explicitly position it as metadata infrastructure for downstream aggregators and compatible registries, not the primary direct integration point for host apps. The 2026-07-28 spec release candidate (locked May 21, 2026) introduces stateless core transport, Tasks extension for long-running work, MCP Apps for server-rendered UIs, OAuth 2.1 alignment, a formal deprecation policy, and deprecates legacy roots and sampling features on the new spec line. Docker Custom MCP Catalogs and Profiles reached GA in May 2026. Microsoft Power BI MCP server released June 2026.

MCP Clients: Claude Desktop, Claude Code, Cursor, Windsurf, VS Code (Copilot + ACP Client), Continue.dev, Zed, ChatGPT, Gemini, Microsoft Copilot, LibreChat, and more.

Popular MCP Servers

Tool / Server Developer Category Open Source Best For GitHub
MCP Filesystem Anthropic / Community File I/O Yes (MIT) Read/write local files from any MCP client πŸ”—
MCP GitHub GitHub / Anthropic Code & DevOps Yes Repo management, issues, PRs, code search πŸ”—
MCP Slack Community Messaging Yes Slack workspace read/write interaction πŸ”—
MCP PostgreSQL Community Database Yes Read-only SQL queries against Postgres πŸ”—
MCP Google Drive Community Storage Yes Drive file access and search πŸ”—
MCP Docker Community DevOps Yes Container management and inspection πŸ”—
MCP Brave Search Brave Search Yes Web + local search via Brave API πŸ”—
MCP AWS AWS Labs Cloud Yes (Apache 2.0) AWS service integration πŸ”—
MCP Notion Community Productivity Yes Notion page and database access πŸ”—
FastMCP Community Framework Yes Python framework for building MCP servers fast πŸ”—
Context7 Upstash Dev Tools Yes Up-to-date library docs for AI coding assistants πŸ”—

MCP 2026-07-28 Release Candidate (May 21, 2026): The largest revision since launch. Key changes: stateless core (no sessions, no initialize handshake, plain HTTP load balancing), MCP Apps extension (server-rendered UIs), Tasks extension (long-running work), OAuth 2.0/OpenID Connect alignment, 12-month deprecation policy. Final spec ships July 28, 2026. This is a breaking change -- servers using session state must migrate to explicit handles. Source: πŸ”— MCP SDK beta support (2026-06-29): Python, TypeScript, Go, and C# SDK beta releases now support the 2026-07-28 release candidate, so server maintainers can test stateless transport, Tasks, MCP Apps, and auth changes before final. Source: πŸ”—

Agent Skills & Registries 🎯

Modular capability packages that extend AI agents with specialized knowledge, workflows, and procedural instructions β€” without bloating model context.

skills.sh

skills.sh is the primary registry and package manager for Agent Skills β€” an open standard developed by Anthropic for packaging and distributing reusable agent capabilities. Skills follow a progressive disclosure pattern: agents load only a skill's name and description at startup, then pull full instructions only when a task matches, keeping context overhead minimal.

Vercel's official Agent Resources now publish installable skills for React, Next.js, AI SDK, design/UI, browser automation, deployment, commerce, and workflow use cases. Source: πŸ”—

Feature Detail
Standard Agent Skills (open, SKILL.md format) β€” developed by Anthropic, hosted on GitHub
Registry URL skills.sh
Total installs 90,989+ all-time
Compatible agents Claude Code, Cursor, Windsurf, VS Code Copilot, Continue.dev, Zed, and any MCP-compatible agent
License Open (skills are author-licensed; spec is open standard)

Top Skills by Category

Skill Publisher Category Installs
find-skills vercel-labs/skills Discovery 1.3M
vercel-react-best-practices vercel-labs/agent-skills Frontend 366K
frontend-design anthropics/skills Design 361K
web-design-guidelines vercel-labs/agent-skills Design 291K
microsoft-foundry microsoft/azure-skills Cloud/Azure 286K
azure-ai microsoft/azure-skills AI/Cloud 276K
agent-browser vercel-labs/agent-browser Browser 229K
skill-creator anthropics/skills Meta 180K
browser-use browser-use/browser-use Automation 71.6K
systematic-debugging obra/superpowers Dev 78.5K
test-driven-development obra/superpowers Dev 68.0K
seo-audit coreyhaines31/marketingskills Marketing 95.4K
supabase-postgres-best-practices supabase/agent-skills Database 138K
playwright-best-practices currents-dev/playwright Testing 34.2K

Notable Publisher Ecosystems

Publisher Skills Count Focus
microsoft/azure-skills 19+ Azure cloud, AI, Kubernetes, cost optimization
vercel-labs/agent-skills 15+ React, Next.js, Tailwind, deployment
anthropics/skills 15+ Design, docs, coding, web artifacts
coreyhaines31/marketingskills 20+ SEO, marketing, content, analytics
obra/superpowers 12+ Dev workflows, parallel agents, TDD
firebase/agent-skills 10+ Firebase, Firestore, GenKit
larksuite/cli 13+ Lark workspace automation
pbakaus/impeccable 10+ Design polish, code quality

Model Routers & Load Balancers πŸ”€

Tools for routing LLM requests across multiple providers, models, and deployments β€” optimizing for cost, latency, quality, or reliability. Prices as of July 2026.

Tool Developer Key Features Open Source Pricing GitHub
LiteLLM BerriAI 100+ provider support, proxy server, load balancing, fallbacks, spend tracking Yes (MIT) Free (OSS) / $99/mo cloud πŸ”—
Portkey Portkey 250+ LLMs, AI gateway, guardrails, observability, virtual keys Yes (Apache 2.0) Free tier / $49/mo+ πŸ”—
OpenRouter OpenRouter 200+ model catalog, unified API, pay-per-use credit system No ~5% markup on provider cost β€”
RouteLLM LMSys Open-source router (strong vs. weak model) using classifier or matrix factorization Yes Free πŸ”—
Not Diamond Not Diamond Pre-trained + custom task-specific routers, cost/quality tradeoff No Free tier + enterprise β€”
Unify AI Unify Quality / cost / latency-aware routing across 100+ model deployments No Usage-based β€”
Semantic Router Aurelio AI Embedding-based semantic intent routing for agents and pipelines Yes Free πŸ”—

Small Language Models (SLMs) πŸ“±

Compact models designed for on-device inference, edge deployment, low-latency APIs, and resource-constrained environments. Generally defined as models under ~15B parameters. Specs as of July 2026.

Model Developer Params Context License Best For
Phi-4 Microsoft 14B 16K MIT Reasoning, math, code β€” STEM benchmark leader at class size
Phi-4-mini Microsoft 3.8B 128K MIT On-device STEM reasoning with long context
Phi-4-multimodal Microsoft 5.6B 128K MIT Vision + audio + text multimodal, edge deployment
Gemma 4 31B Google 31B dense 256K Apache 2.0 Multimodal (text+image), hybrid-thinking, open-weight reasoning
Gemma 4 26B A4B Google 26B MoE (4B active) 256K Apache 2.0 Multimodal, efficient MoE, agentic workflows
Gemma 4 E4B Google 4B dense 256K Apache 2.0 On-device, CPU inference, multimodal (text+image+audio)
Gemma 4 E2B Google 2B dense 256K Apache 2.0 Ultra-lightweight, on-device, 5GB RAM (4-bit)
Gemma 4 12B Google 12B dense 256K Apache 2.0 Encoder-free unified multimodal (text+image+audio), hybrid-thinking
Aion 1.0 Instruct ⭐ Microsoft β€” β€” Open weights (July 2026) On-device SLM for Windows; smaller/faster than current Windows SLM; Edge Insider preview; open-source on HuggingFace July 2026
Aion 1.0 Plan ⭐ Microsoft 14B 32K β€” On-device reasoning/tool-calling SLM; agentic workflows locally; shipping in-box Windows "in the coming months"
edge-lm (Gemma 4 E2B/E4B) ⭐ TheStageAI 2B/4B 256K MIT 7x smaller Gemma 4 checkpoints; E2B in 1.44GB; MLX-ready; optimized for Apple Silicon/edge
Atome LM ⭐ TilelliLab 60K β€” Apache 2.0 Ternary zero-heap LM for microcontrollers; runs on ESP32; ~1 tok/s on Cortex-M3
Gemma 3 27B Google 27B 128K Apache 2.0 Top open model, multilingual (140+ languages)
Gemma 3 4B Google 4B 128K Apache 2.0 CPU inference, 140+ languages, mobile-friendly
Gemma 3 1B Google 1B 32K Apache 2.0 On-device, embedded, ultra-lightweight
Qwen3.5-9B Alibaba 9B 256K Apache 2.0 Thinking + non-thinking modes; native multimodal; 201 languages; reasoning closes gap to 30B+ models
Qwen3.5-4B Alibaba 4B 256K Apache 2.0 Multimodal (native VLM), agentic, on-device; edge deployment
Qwen3.5-2B Alibaba 2B 256K Apache 2.0 Runs on recent iPhones (offline); 201 languages; text+image
Qwen3.5-0.8B Alibaba 0.8B 256K Apache 2.0 IoT / rapid prototyping; non-thinking mode only; ultra-lightweight
Mistral Small 4 Mistral AI 119B total / 6B active (MoE, 128 experts) 256K Apache 2.0 Chat + reasoning + coding + vision (multimodal) in one model; released 2026-03-16
SmolLM3 Hugging Face 3B 128K Apache 2.0 Efficient, tool use, multilingual, reasoning
Qwen2.5 3B Alibaba 3B 128K Apache 2.0 Asian and multilingual tasks, coding
Qwen2.5 7B Alibaba 7B 128K Apache 2.0 Strong multilingual baseline, function calling
Llama 3.2 3B Meta 3B 128K Llama 3.2 license General-purpose, on-device, Meta ecosystem
Llama 3.2 1B Meta 1B 128K Llama 3.2 license Lightweight edge inference, distillation target
Granite 3.3 8B IBM 8B 128K Apache 2.0 Enterprise tasks, tool use, business-domain
MiniCPM 3.0 ModelBest / Tsinghua 4B 32K Apache 2.0 Compact yet capable, mobile and edge
Danube 3 500M H2O.ai 500M 8K Apache 2.0 Ultra-lightweight on-device, IoT
MobileMoE ⭐ Academic 0.3B–0.9B active (1.3B–5.3B total) 128K Research First sub-billion MoE family for on-device; 1.8-3.8x faster than dense baselines
SLM-10M (Liodon AI) ⭐ Liodon AI 9.97M 1K Apache 2.0 Leads sub-10M Open SLM Leaderboard (32.38%); beats Pythia-31M with 1/3 params
MiniGPT-16MB ⭐ david-spies 1.46M 256 MIT Sub-16MB model; runs in browser via WebAssembly; 60-120+ tok/s
Micro Language Models (muLMs) ⭐ Sensente 8M–30M β€” Research Instant on-device response initiation; masks cloud latency; collaborative generation
PhoneLM ⭐ UbiquitousLearning 0.5B / 1.5B 32K / 128K Apache 2.0 Smartphone-native SLM; architecture searched for NPU efficiency; Android intent invocation

Apple Foundation Models 3rd Generation (June 8, 2026): Apple announced five foundation models powering Apple Intelligence. AFM 3 Core (3B dense) and AFM 3 Core Advanced (20B sparse, 1-4B active) run on-device. AFM 3 Cloud, ADM 3 Cloud (Image), and AFM 3 Cloud Pro run on Private Cloud Compute. Not available via public API. AFM 3 Core Advanced achieves 4.15 MOS for TTS (vs 3.87 production baseline) and 44.7% preference on dictation quality. Source: πŸ”—

Notable GitHub repos:


Guides πŸ“š

Tutorials, how-tos, and in-depth guides for getting the most out of AI models and tools.

Getting Started πŸš€

A beginner-friendly introduction to AI models and how to start using them effectively.

Understanding LLMs

Concept Description
Parameters Size of model (B = billions). More = more capable
Context Window How much text model can process (128K standard)
Tokens Basic units of text (~0.75 words per token)

Accessing AI Models

Method Best For Setup Difficulty
Web Interfaces Quick experiments Easiest
API Access Building applications Easy
Self-Hosting Privacy, no API costs Medium-Hard
IDE Integration Daily coding Easy

Model Recommendations by Task

Task Free Option Premium Option
Chat Llama 4 (self-hosted) Claude Fable 5, GPT-5.5 Instant
Coding GLM-5.2, Qwen3-Coder (self-hosted) Claude Fable 5, Claude Opus 4.8
Reasoning DeepSeek-R1 Gemini 3 Deep Think, GPT-5.5 Pro
Long docs Llama 4 Scout Gemini 3 Flash
Vision Llama 4 Maverick GPT-5.5, Gemini 3.1 Pro

Free Models & APIs for Vibe Coding πŸ’»

Vibe coding β€” describing what you want in natural language and letting AI generate the code β€” has exploded in 2026. The ecosystem splits into two tracks: free AI APIs you plug into your own editor/agent, and free vibe coding IDEs/platforms that bundle everything together.

Free AI APIs for Coding

These are the raw API endpoints you can use in tools like Cursor (BYOK), Cline, or any agent framework.

Provider Free access Paid baseline if you outgrow free Best For Source
Google Gemini API Gemini Developer API free tier is free of charge Gemini 2.5 Flash-Lite is $0.10 / $0.40 standard; $0.05 / $0.20 in Batch/Flex Long-context and multimodal coding helpers πŸ”—
Groq Cloud GroqCloud Free plan Llama 3.1 8B Instant is $0.05 / $0.08 on paid usage Very fast code loops and agent backends πŸ”—, πŸ”—
Cohere Trial keys are free but limited; Command A+ is free until rate limits are reached Command R is $0.15 / $0.60 Tool use, RAG, and multilingual coding support πŸ”—, πŸ”—, πŸ”—
Mistral AI Free mode is enabled by default with no credit card required Mistral Small 4 is $0.15 / $0.60 on Scale Cheap coding and multimodal experiments πŸ”—, πŸ”—
Anthropic New users receive a small amount of free credits Haiku 4.5 is $1.00 / $5.00 Claude-first agent and prompt testing πŸ”—, πŸ”—
Vercel AI Gateway ⭐ $5/month included on the free tier for eligible models Paid usage follows provider list rates with zero markup Multi-provider prototyping behind one endpoint πŸ”—
OpenAI Developer quickstart includes one free test API request GPT-5.4 nano is $0.20 / $1.25 OpenAI SDK integration checks πŸ”—, πŸ”—

Free Vibe Coding IDEs & Platforms

Tool Type Key Features Best For
Cursor AI IDE Agent Mode, Composer 2, multi-agent workspace Professional development
Cline VS Code Extension Open-source, BYOK/Ollama, MCP tools Self-hosted, unlimited local LLM
Windsurf AI IDE Cascade agent, live browser preview IDE with browser integration
OpenHands Docker Agent Self-hosted, local LLM support, full SDLC Unlimited local development
bolt.diy Browser IDE 19+ LLM providers, Ollama, full-stack apps Free web app building
Open Interpreter CLI Natural language β†’ code, local LLM Simple local automation

Chrome DevTools MCP - Game Changer for Web Dev

Google's Chrome DevTools MCP connects AI agents directly to Chrome for debugging, profiling, and automation:

  • 29 tools across 6 categories (input, navigation, emulation, performance, network, debugging)
  • Run Lighthouse audits, capture performance traces, inspect network requests
  • Works with Claude Code, Cursor, Copilot via MCP
  • Supports BYOK/local LLMs through MCP clients
  • GitHub | 37,783+ stars

Cloudflare Browser Run

Managed browser infrastructure for AI agents:

  • Chrome DevTools Protocol (CDP) direct endpoint
  • MCP client support (Claude, Cursor, OpenCode)
  • Session recordings, Live View, WebMCP, and /json schema extraction with Workers AI or BYOK
  • Free Workers plan / $5+/mo paid
  • Browser Run

Recommendations by Use Case

Free + Local LLM: Cline + Ollama, OpenHands + Qwen3 Coder, bolt.diy + Ollama

Fast API Iteration: Groq (speed) + Cerebras (high limits)

Web Development: Chrome DevTools MCP + Cline (zero-cost debugging)

Many Models: OpenRouter (unified API)

Production Inference: NVIDIA NIM, Cerebras

πŸ’‘ Pro Tip: Combine Chrome DevTools MCP with a local LLM (Ollama) via Cline for completely free, unlimited AI-powered web development and debugging.

A comprehensive guide to running AI models on your own hardware.

Benefits

Benefit Description
Privacy Data never leaves your infrastructure
Cost Control No per-token API costs for unlimited usage
Customization Fine-tune models for specific needs
No Rate Limits Process as much as hardware allows
Offline Access Work without internet

Quick Start with Ollama

For installation and usage instructions, refer to the official Ollama documentation.

Local GPU Quick Guide

Recommended apps (local-first):

  • Ollama - Simple local runtime with a local HTTP API
  • LM Studio - Desktop UI for downloading and running models locally
  • llama.cpp - Fast local inference (CPU/GPU), great for quantized models
  • Open WebUI - Optional local web UI (pairs well with local runtimes)

If you want β€œserver-style” hosting (advanced):

  • vLLM - High-throughput serving for NVIDIA GPUs
  • SGLang - Structured generation and serving workflows

Practical setup tips:

  1. Install the latest NVIDIA drivers (enable GPU acceleration in your chosen app)
  2. Start with smaller quantized models (Q4 is a common β€œbest default”)
  3. Keep context windows realistic for local hardware (lower context = faster, less memory)
  4. Watch VRAM first, then system RAM; reduce model size or quantization if either saturates
  5. Prefer running locally on localhost and only expose to LAN if you understand firewall rules

Example hardware configurations:

Hardware Good starting point Notes
Consumer GPU (24 GB VRAM) 7B–14B quantized e.g., RTX 4090, RTX 3090 β€” great for chat/coding
Pro GPU (48–80 GB VRAM) 14B–70B quantized e.g., A6000, A100 β€” coding agents, longer contexts
Multi-GPU (160+ GB VRAM) 70B+ quantized e.g., 2Γ—A100 β€” larger open-source models
CPU-only (32–64 GB RAM) 7B–14B quantized Slower but viable for offline chat; keep context moderate

Deployment Options

Option Best For Pros Cons
Local Machine Personal use Simple, no latency Limited hardware
Dedicated Server Team use Full control Maintenance
Cloud GPU Rental Experimentation On-demand Hourly costs
Kubernetes Enterprise Scalable Complex

Cost Analysis πŸ’°

Comprehensive pricing comparisons and cost calculations.

Pricing Tiers

Tier Price Range Models
πŸ†“ Free $0 Self-hosted, free tiers
πŸ’Έ Budget $0.025 - $0.50/1M Gemini 3.1 Flash-Lite, GPT-4.1 nano, GLM-4.7-FlashX, GPT-5.4 nano, Grok 4 Fast (aliased to 4.3)
πŸ’° Mid-range $0.60 - $15.00/1M GPT-5.4 mini, Claude Haiku 4.5, Kimi K2.6, Sonar, GLM-5, GPT-5.4, Claude Sonnet, Grok 4.3
πŸ’Ž Premium $15.00 - $600.00/1M Claude Fable 5, GPT-5.5 Pro, Claude Opus 4.8, Claude Opus 4.7, o1-Pro

Subscription Pricing (Monthly, USD)

AI chat apps

Product Plans (USD) Notes Official Source
ChatGPT Go $8, Plus $20, Pro $200, Business $25/seat (annual) or $30/seat (monthly), Enterprise (contact sales) Consumer prices are US-listed; Go is localized in some markets πŸ”—
Claude Pro $20, Max $100 (5Γ—) or $200 (20Γ—), Team/Enterprise (see pricing) Prices shown exclude applicable taxes; availability varies by region πŸ”—
Google AI (Gemini) Free, Plus $4.99, Pro $19.99, Ultra $99.99 (5Γ— limits / 20TB storage) or $200 (20Γ— limits / 20TB storage) US pricing; Plus price dropped from $7.99 on June 8, 2026 with storage doubled to 400GB; Ultra reduced from $249.99 at I/O 2026; $99.99 tier added with Gemini 3.5 Flash integration and Antigravity priority access πŸ”—

Coding assistants

Tool Plans (USD) Notes Official Source
GitHub Copilot Free $0, Pro $10 base + $5 monthly bundled AI Credits, Pro+ $39 base + $31 bundled AI Credits, Business $19/user, Enterprise $39/user Usage-based billing since June 1, 2026; chat, agent mode, code review, cloud agent, CLI, and apps draw from GitHub AI Credits πŸ”—

Model Pricing Comparison

Comparative pricing across the cheapest official public APIs verified in this pass. Only providers and models rechecked against current official pricing pages are included below. As of 2026-07-19 16:54 UTC.

Model Input Output Cached Input Best For Source
Llama 3.1 8B Instant (Groq) $0.05 $0.08 β€” Lowest verified public direct text API price πŸ”—
GPT OSS 20B (Groq) $0.075 $0.30 $0.0375 Cheap open-weight routing and agents πŸ”—
Gemini 2.5 Flash-Lite $0.10 $0.40 $0.01 Cheapest verified multimodal API with an always-on free tier πŸ”—
DeepSeek-V4-Flash $0.14 $0.28 $0.0028 (hit) Strong price/performance for long-context work πŸ”—
Mistral Small 4 $0.15 $0.60 β€” Low-cost coding and multimodal work πŸ”—
Command R (Cohere) $0.15 $0.60 β€” RAG and tool use πŸ”—
GPT-5.4 nano $0.20 $1.25 $0.02 OpenAI lowest-cost standard model πŸ”—
Claude Haiku 4.5 $1.00 $5.00 $0.10 (read) Lowest-cost Claude model πŸ”—
GPT-5.6 Luna $1.00 $6.00 $0.10 Cheapest current GPT-5.6 tier πŸ”—
Kimi K3 ⭐ $3.00 $15.00 $0.30 (hit, 90% cached) New open-weight frontier model (2026-07-16); 2.8T MoE, 1M context, native vision πŸ”—
Meta Muse Spark 1.1 ⭐ API-only API-only β€” Closed-weights frontier multimodal reasoning; agentic tool and computer use; 1M context; released 2026-07-09 πŸ”—
Grok-4.5 $2.00 $6.00 $0.50 (hit) xAI new flagship (2026-07-08); 500K context, text+image; function calling, structured outputs, reasoning πŸ”—
Grok 4.3 $1.25 $2.50 $0.20 xAI flagship at the current public rate card πŸ”—
Claude Sonnet 5 $2.00 intro / $3.00 standard $10.00 intro / $15.00 standard $0.20 intro read Strong coding model if the budget can stretch πŸ”—

Self-Hosting vs API (Monthly)

Usage Level Self-Host (A100) API (GPT-5) Winner
Light (1M tokens) $300 (rental) $10 API
Medium (100M tokens) $300 $1,000 Self-host
Heavy (1B tokens) $300 $10,000 Self-host
Enterprise (10B+ tokens) $2,000 (owned) $100,000+ Self-host

Reference πŸ“–

Reference materials including glossary, comparison tables, and data sources.

Glossary πŸ“–

Definitions of common terms used throughout the documentation.

A-E

Term Definition
Agent AI system that autonomously performs tasks and interacts with environments
API Interface for programmatically accessing AI models
Attention Mechanism Neural network component focusing on relevant input parts
Benchmark Standardized test measuring model performance
Chain-of-Thought (CoT) Prompting technique showing step-by-step reasoning

F-L

Term Definition
Fine-Tuning Adapting pre-trained model to specific tasks
Frontier Model State-of-the-art proprietary model
GPU Hardware accelerator essential for ML
LLM Large Language Model
LoRA Efficient fine-tuning method

M-R

Term Definition
MCP Model Context Protocol for tool interaction
MMLU Massive Multitask Language Understanding benchmark
MoE Mixture of Experts architecture
Multimodal Processing multiple input types
RAG Retrieval-Augmented Generation

S-Z

Term Definition
Self-Hosting Running models on own infrastructure
SLM Small Language Model
SWE-bench Benchmark for real GitHub issue resolution
Token Basic unit of text processing
VRAM GPU memory for model storage

Comparison Tables πŸ“Š

Side-by-side comparisons of AI models sorted by various criteria.

Sort by Latest Update (Default)

🏒 Company πŸ€– Model πŸ“¦ Version πŸ“… Release Date πŸ”„ Latest Updated πŸ’» Coding πŸ“Š Benchmarks πŸ’° Price πŸ–₯️ Self-Host πŸ”— Official Site
πŸ€– Anthropic Claude Fable 5 redeployed 2026-06-09 00:00 UTC 2026-07-01 00:00 UTC ⭐ βœ… SWE-bench 95.0% $10.00 / $50.00 ❌ πŸ”—
πŸ€– Anthropic Claude Sonnet 5 2026-06-30 00:00 UTC 2026-06-30 00:00 UTC ⭐ βœ… SWE-bench 92.4%, SWE-bench Pro 63.2%, HLE 57.4% with tools $2.00 / $10.00 intro ❌ πŸ”—
πŸ€– OpenAI GPT-5.6 Sol / Terra / Luna 2026-07-09 00:00 UTC 2026-07-09 00:00 UTC ⭐ βœ… General availability; Sol $5/$30, Terra $2.50/$15, Luna $1/$6 ❌ πŸ”—
πŸ‡¨πŸ‡³ Moonshot AI Kimi K3 2026-07-16 00:00 UTC 2026-07-16 00:00 UTC ⭐ βœ… 2.8T MoE open-weight frontier; 1M context; $3.00/$15.00; 90% cache-hit discount βœ… πŸ”—
πŸ‡¨πŸ‡³ Meta Meta Muse Spark 1.1 2026-07-09 00:00 UTC 2026-07-09 00:00 UTC ⭐ βœ… Closed-weights frontier multimodal reasoning; agentic tool and computer use; 1M context; API-only ❌ πŸ”—
πŸš€ xAI Grok 4.5 2026-07-08 00:00 UTC 2026-07-08 00:00 UTC ⭐ βœ… New frontier multimodal LLM; 500K context; $2.00/$6.00 ❌ πŸ”—
πŸ€– OpenAI GPT-Live 1 / 1 mini 2026-07-08 00:00 UTC 2026-07-08 00:00 UTC ⭐ β€” Full-duplex voice models powering ChatGPT Voice; GPT-5.5 backend; API pending ❌ πŸ”—
πŸ‡¨πŸ‡³ MiniMax M2 / M2.1 Open-source 2026-07-02 00:00 UTC 2026-07-02 00:00 UTC ⭐ βœ… 230B/10B active MoE, 1M context, Apache 2.0 βœ… πŸ”—
🌐 Google DeepMind Gemini Omni Flash β€” 2026-06-30 00:00 UTC 2026-06-30 00:00 UTC ⭐ βœ… $0.10/s video; any-to-any multimodal ❌ πŸ”—
🌐 Google DeepMind Nano Banana 2 Lite β€” 2026-06-30 00:00 UTC 2026-06-30 00:00 UTC ⭐ β€” $0.034/1K images; 4s generation ❌ πŸ”—
πŸ‡¨πŸ‡³ Zhipu AI GLM-5.2 β€” 2026-06-13 00:00 UTC 2026-06-17 00:00 UTC ⭐ βœ… AIME 2026 99.2%, SWE-bench Pro 62.1%, FrontierSWE 74.4% $1.40 / $4.40 βœ… πŸ”—
🌐 Google DeepMind Gemini 3.5 Flash 2026-05-19 00:00 UTC 2026-05-19 00:00 UTC ⭐ βœ… GPQA ~90.4%, SWE-bench Pro 55.1% $1.50 / $9.00 ❌ πŸ”—
πŸ€– Anthropic Claude Mythos Preview 2026-04-07 00:00 UTC 2026-04-07 00:00 UTC β€” Not disclosed $25.00 / $125.00 ❌ πŸ”—
πŸ”¬ DeepSeek DeepSeek V4 (Flash/Pro) 2026-04-24 00:00 UTC 2026-05-08 00:00 UTC βœ… No public benchmarks From $0.14 / $0.28 (Flash) βœ… πŸ”—
πŸ€– OpenAI GPT-5.5 Pro 2026-04-26 00:00 UTC 2026-05-08 00:00 UTC βœ… GPQA 95.1%, SWE-bench 92.3% $30.00 / $180.00 ❌ πŸ”—
🌐 Google DeepMind Gemini 3 Flash 2026-02-12 00:00 UTC 2026-05-08 00:00 UTC βœ… GPQA 90.4%, SWE-bench 78.0% $0.50 / $3.00 ❌ πŸ”—
πŸš€ xAI Grok 4.3 2026-05-01 00:00 UTC 2026-05-06 00:00 UTC βœ… β€” $1.25 / $2.50 ❌ πŸ”—
πŸ‡¨πŸ‡³ MiniMax Hailuo 02 2026-06-20 00:00 UTC 2026-06-20 00:00 UTC ⭐ β€” #2 Video Arena API ❌ πŸ”—
πŸ‡¨πŸ‡³ MiniMax Speech 2.8 2026-01-23 00:00 UTC 2026-07-03 21:56 UTC ⭐ β€” Native sound tags, high-fidelity cloning, studio-grade clarity API ❌ πŸ”—
πŸ‡¨πŸ‡³ MiniMax Video-01 β€” 2026-06-20 00:00 UTC 2026-06-20 00:00 UTC ⭐ β€” 720p/25fps ~$0.40/video ❌ πŸ”—
πŸ‡«πŸ‡· Mistral AI Voxtral TTS β€” 2026-06-18 00:00 UTC 2026-06-18 00:00 UTC ⭐ β€” 4B, 9 languages, 70ms $0.016/1K chars βœ… πŸ”—
πŸ€– Microsoft VibeVoice Realtime (0.5B) ⭐ β€” 2025-12-03 00:00 UTC 2026-03-06 00:00 UTC ⭐ β€” 0.5B real-time TTS; <300ms TTFA; streaming input Free (MIT) βœ… πŸ”—
πŸ€– Microsoft VibeVoice ASR (7B) ⭐ β€” 2026-01-21 00:00 UTC 2026-03-06 00:00 UTC ⭐ β€” 60-min ASR with diarization + timestamps Free (MIT) βœ… πŸ”—
πŸ€– OpenAI GPT-4.1 ⭐ β€” 2026-04-14 00:00 UTC 2026-04-14 00:00 UTC βœ… 1M context, improved coding/instruction following $2.00 / $8.00 ❌ πŸ”—
πŸ€– OpenAI GPT-4.1 mini ⭐ β€” 2026-04-14 00:00 UTC 2026-04-14 00:00 UTC βœ… 1M context, best value for large-context needs $0.40 / $1.60 ❌ πŸ”—
πŸ‡ΊπŸ‡Έ Zyphra ZONOS2 β€” 2026-06-12 00:00 UTC 2026-06-12 00:00 UTC ⭐ β€” 8B MoE TTS, multilingual Free (Apache 2.0) βœ… πŸ”—
πŸ€– OpenAI GPT-5.5 β€” 2026-04-26 00:00 UTC 2026-04-26 00:00 UTC βœ… GPQA 93.2%, SWE-bench 88.5% $5.00 / $30.00 ❌ πŸ”—
πŸ€– Anthropic Claude Opus 4.7 2026-04-22 00:00 UTC 2026-04-26 00:00 UTC βœ… GPQA 94.2%, SWE-bench 87.6% $5.00 / $25.00 ❌ πŸ”—

Release Windows (Month-level)

🏒 Company πŸ€– Model πŸ“… Release Window Notes πŸ”— Official Site
πŸ€– OpenAI GPT-5.6 2026-07 Sol/Terra/Luna three-tier family; general availability; Sol $5/$30, Terra $2.50/$15, Luna $1/$6 πŸ”—
πŸ‡¨πŸ‡³ Moonshot AI Kimi K3 2026-07 2.8T MoE open-weight frontier; 1M context; $3.00/$15.00; 90% cache-hit discount; open weights promised by 2026-07-27 πŸ”—
πŸ‡¨πŸ‡³ Meta Meta Muse Spark 1.1 2026-07 Closed-weights frontier multimodal reasoning; agentic tool and computer use; 1M context; API-only (Meta Model API) πŸ”—
πŸ‡¨πŸ‡³ MiniMax M2 / M2.1 2026-07 Open-source MoE (230B/10B active); 1M context; Apache 2.0; $0.30/$1.20 API πŸ”—
🌐 Google DeepMind Gemini Omni Flash 2026-06 $0.10/s video generation; any-to-any multimodal; conversational editing πŸ”—
🌐 Google DeepMind Nano Banana 2 Lite 2026-06 Fastest Gemini Image; 4s gen; $0.034/1K images πŸ”—
πŸ€– Anthropic Claude Fable 5 2026-06 First publicly available Mythos-class model; $10/$50; free on subscription through June 22
πŸ€– Anthropic Claude Mythos 5 2026-06 Restricted to Project Glasswing partners; same model as Fable 5 with safeguards lifted
πŸ‡¨πŸ‡³ Alibaba Qwen3.7-Max 2026-05 Proprietary; $2.50 / $7.50 πŸ”—
πŸ‡¨πŸ‡³ Zhipu AI GLM-5.2 2026-06 753B params, 1M context, MIT open-weight; AIME 2026 99.2%; API $1.40/$4.40 πŸ”—
🌐 Google DeepMind Gemini 3.5 Flash 2026-05 GA at I/O 2026; 4x faster than 3.1 Pro
🌐 Google DeepMind Gemini Omni Flash 2026-05 Any-to-any multimodal creation; rolling out to AI Plus/Pro/Ultra πŸ”—
πŸ€– Anthropic Claude Opus 4.8 2026-05 Improved Opus 4.7; SWE-bench 88.6%; fast mode $10/$50 πŸ”—
πŸš€ xAI Grok 4.20 / 4.3 / 4 Fast 2026-05 Major price cuts; Grok 4.20 now $1.25/$2.50; 4 Fast β†’ 4.3 alias πŸ”—
πŸ€– Anthropic Claude Mythos Preview 2026-04 Invitation-only via Project Glasswing πŸ”—
πŸ€– OpenAI GPT-5.5 Instant 2026-05 Default ChatGPT model since 2026-05-05 πŸ”—
πŸš€ xAI Grok 4.3 2026-05 Always-on reasoning; $1.25 / $2.50 πŸ”—
🧠 MiniMax MiniMax M2.5 2026-02 $0.30 / $1.20 πŸ”—
πŸ‡¨πŸ‡³ Alibaba/Qwen Qwen 3.5-Max 2026-02 Open-source release window πŸ”—
🌐 Google DeepMind Gemini 3.1 Flash-Lite 2026-02 Budget Gemini model πŸ”—
🌐 Google DeepMind Gemini 3 Pro 2026-01 Tiered pricing πŸ”—
πŸ€– OpenAI GPT-5.4 family 2026-03 GPT-5.4, GPT-5.4 mini, GPT-5.4 nano πŸ”—
πŸ€– OpenAI GPT-4.1 family ⭐ 2026-04 GPT-4.1 ($2/$8), GPT-4.1 mini ($0.40/$1.60), GPT-4.1 nano ($0.10/$0.40); 1M context; June 2024 cutoff; retired from ChatGPT Feb 13, 2026; API still available πŸ”—
πŸ‡«πŸ‡· Mistral AI Mistral Large 3 2025-11 Apache 2.0 open-source, 123B params πŸ”—

Sort by Price (Cheapest)

Rank Model Input Output License
1 Self-hosted open weights $0 $0 Various
2 Llama 3.1 8B Instant (Groq) $0.05 $0.08 API
3 GPT OSS 20B (Groq) $0.075 $0.30 Open weights via API
4 Gemini 2.5 Flash-Lite $0.10 $0.40 Proprietary
5 DeepSeek-V4-Flash $0.14 $0.28 API
6 Mistral Small 4 $0.15 $0.60 Open weights/API
7 Command R $0.15 $0.60 API
8 GPT-5.4 nano $0.20 $1.25 Proprietary
9 Kimi K3 ⭐ $3.00 $15.00 Open-weight (Modified MIT)
10 Claude Haiku 4.5 $1.00 $5.00 Proprietary
10 GPT-5.6 Luna $1.00 $6.00 Proprietary

Sort by Performance (Coding)

Rank Model SWE-bench Verified Self-Host
1 Claude Mythos 5 95.5% ❌
2 Claude Fable 5 95.0% ❌
3 Claude Sonnet 5 ⭐ 92.4% ❌
4 GPT-5.5 Pro 92.3% ❌
5 Claude Opus 4.8 88.6% ❌
6 Claude Opus 4.7 87.6% ❌
7 GPT-5.3-Codex 85.0% ❌
8 GLM-5.2 ⭐ 62.1% (SWE-bench Pro) βœ…
9 Kimi K3 ⭐ 88.3% (Terminal-Bench 2.1) βœ…
10 Claude Opus 4.6 80.8% ❌
10 Gemini 3.1 Pro 80.6% ❌

Sort by Context Window

Rank Model Context Best For
1 Gemini 3 Flash 10M Entire libraries
2 Llama 4 Scout 10M Long-document RAG
3 Grok 4.20 2M Large codebases with full context
4 Gemini 3 Pro 1M+ Research papers
5 Gemini 3.1 Pro 1M Complex multi-document analysis
6 GLM-5.2 ⭐ 1M Frontier open-weight coding + long context
7 Kimi K3 ⭐ 1M 2.8T MoE open-weight frontier; native vision; $3.00/$15.00
8 Gemini 3.5 Flash 1M Fast multimodal reasoning
8 Claude Opus 4.7 1M Agentic coding with full codebase
9 GPT-5.5 Pro 1.05M Premium reasoning with long context
10 Grok 4.3 1M Reasoning with full project context

Data Sources πŸ“š

Attribution, verification sources, and methodology.

Primary Sources

Company Source URL
OpenAI Official Documentation openai.com
OpenAI ChatGPT agent release notes help.openai.com
OpenAI Model release notes help.openai.com
OpenAI API pricing platform.openai.com
OpenAI API pricing (overview) openai.com/api/pricing
OpenAI GPT-5.5 pro model reference developers.openai.com
OpenAI GPT-5.3-Codex model reference developers.openai.com
OpenAI March 2026 model news openai.com
OpenAI ChatGPT subscriptions (Go/Plus/Pro) openai.com
OpenAI ChatGPT Business pricing help.openai.com
OpenAI GPT-4.5 API deprecation developers.openai.com/api/docs/deprecations
OpenAI GPT-4.1 / GPT-4o retirement from ChatGPT openai.com
OpenAI GPT-Image-2 / ChatGPT Images 2.0 openai.com
OpenAI Codex changelog and CLI docs developers.openai.com
OpenAI Workspace agents in ChatGPT openai.com
OpenAI GPT-Live voice models announcement openai.com
OpenAI Realtime API (gpt-realtime 2.1) docs platform.openai.com
OpenAI Developer quickstart (free test request) developers.openai.com
Anthropic Claude Documentation anthropic.com
Anthropic Claude API pricing docs.anthropic.com
Anthropic Claude rate card and plan pricing claude.com
Anthropic Claude Sonnet 5 announcement anthropic.com
Anthropic Claude Sonnet 5 system card anthropic.com
Anthropic Claude Fable 5 redeployment anthropic.com
Anthropic Claude Opus 4.8 announcement anthropic.com
Anthropic Claude Fable 5 / Mythos 5 announcement anthropic.com
Anthropic Claude Tag announcement anthropic.com
Anthropic Claude Science workbench anthropic.com
Anthropic Claude Haiku 4.5 announcement anthropic.com
Anthropic Claude Pro pricing anthropic.com
Anthropic Max plan pricing anthropic.com
Anthropic Claude model selection and deprecations docs.anthropic.com
Google Gemini Documentation deepmind.google
Google Cloud Vertex AI Gemini pricing cloud.google.com
Google Gemini Developer API pricing ai.google.dev
Google Gemini API models (Flash-Lite pricing) ai.google.dev
Google Project Mariner deepmind.google
Google Google AI plans one.google.com
Google Google AI Plus pricing blog.google
Google Google AI Pro pricing one.google.com
Google Google AI Ultra pricing blog.google
Google Gemini API model catalog ai.google.dev
Google Gemini image model cards deepmind.google
GitHub Copilot plans & pricing github.com
GitHub Copilot changelog github.blog
Kiro IDE changelog kiro.dev
Cursor Pricing and changelog cursor.com, cursor.com
Windsurf Pricing and usage docs windsurf.com, docs.windsurf.com
Zhipu AI (Z.ai) Developer Documentation docs.z.ai
Zhipu AI (Z.ai) GLM-5.2 announcement z.ai
DeepSeek Models & API pricing; corporate site api-docs.deepseek.com, deepseek.com
Mistral AI Mistral Small 4 model card docs.mistral.ai
Mistral AI Free mode API key quickstart docs.mistral.ai
MiniMax Developer Documentation platform.minimax.io
MiniMax Pricing (Pay‑as‑you‑go) platform.minimax.io
MiniMax Speech 2.8 announcement minimax.io
Moonshot AI Developer Documentation platform.moonshot.ai
Moonshot AI Models & Pricing platform.moonshot.ai
Moonshot AI Kimi K3 announcement kimi.com
Meta Llama Documentation llama.meta.com
Meta Muse Spark 1.1 announcement ai.meta.com
Cohere Developer Documentation docs.cohere.com
Cohere Pricing overview and trial usage docs.cohere.com
Cohere Command R and Command A+ pricing docs.cohere.com, docs.cohere.com
Groq On-demand model pricing groq.com
Groq GroqCloud free and developer plans groq.com
Cartesia Sonic 3.5 / Ink-2 launch and docs cartesia.ai, docs.cartesia.ai
AWS Amazon Bedrock AgentCore aws.amazon.com
Cloudflare Browser Run and /json Quick Action developers.cloudflare.com
Browser Use Browser Use pricing, cloud docs, and releases browser-use.com, docs.browser-use.com, github.com
Browser Use Browser Harness github.com
Vercel agent-browser docs agent-browser.dev
Vercel AI SDK 7 vercel.com
Vercel AI Gateway routing rules and pricing vercel.com, vercel.com
Vercel eve Agent Runs in MCP/CLI vercel.com
Vercel Sandbox FUSE support vercel.com
Vercel Agent Resources skills vercel.com
Google ADK 2.0 and ADK Go 2.0 adk.dev, developers.googleblog.com
Google Genkit Agents developers.googleblog.com
Mastra File-based agents mastra.ai
DeepReinforce Ornith-1.0 model family huggingface.co, huggingface.co
Jina AI jina-embeddings-v5-omni jina.ai
Model Context Protocol Official registry and 2026-07-28 SDK betas registry.modelcontextprotocol.io, blog.modelcontextprotocol.io
AI21 Labs Developer Documentation docs.ai21.com
Perplexity Developer Documentation docs.perplexity.ai
ByteDance (Volcengine) Developer Documentation volcengine.com
Tencent (Hunyuan) Cloud Documentation cloud.tencent.com
Baidu (ERNIE) AI Studio Documentation ai.baidu.com
xAI Grok API pricing & models docs.x.ai, docs.x.ai
Apple Foundation Models 3rd Gen announcement machinelearning.apple.com
Meta Llama Documentation llama.meta.com

Benchmark Sources

Benchmark Source Description
GPQA Diamond Google Research Graduate-level science questions (PhD difficulty)
MMLU-Pro TIGER-Lab Extended multi-task language understanding
Arena Elo lmarena.ai Crowdsourced human preference ranking
HLE Scale AI Humanity's Last Exam β€” expert-level questions
SWE-bench Verified Princeton Real GitHub issue resolution (human-verified)
SWE-bench Pro Princeton More challenging subset of SWE-bench
LiveCodeBench LiveCodeBench Live competitive programming problems
AIME 2025 MAA American Invitational Mathematics Examination
ARC-AGI-2 ARC Prize Abstract reasoning challenge (fluid intelligence)
MMMU / MMMU-Pro MMMU Multi-discipline multimodal understanding
IFEval Google Research Instruction-following evaluation
FrontierMath Epoch AI Expert-level research mathematics
HumanEval OpenAI 164 Python programming problems

Verification Methodology

  1. Primary Source Review - Check official documentation
  2. Cross-Validation - Compare multiple sources
  3. Timestamp Verification - All data includes verification date
  4. Update Tracking - Monitor official channels

Supplemental agent & CLI notes πŸ“Ž

Additional context from the June 2026 research pass (supplements earlier tables; nothing below removes or supersedes prior entries).

  • Google I/O 2026 agent announcements - Antigravity 2.0 (desktop + CLI + SDK) GA on 2026-05-19; Gemini Spark (24/7 personal agent) included with AI Ultra; Android CLI 1.0 stable for AI agent Android development; Gemini CLI sunset on 2026-06-18.
  • Microsoft Build 2026 - Windows Agent Framework open-sourced (MIT); Aion 1.0 Instruct and Aion 1.0 Plan on-device SLMs announced; Windows 365 for Agents GA; Surface RTX Spark Dev Box and DGX Station for Windows announced.
  • Claude Sonnet 5 launch - Anthropic launched Sonnet 5 on 2026-06-30 00:00 UTC with 1M context, 128K output, adaptive thinking, $2.00 / $10.00 introductory pricing through 2026-08-31 00:00 UTC, and 92.4% SWE-bench Verified.
  • Claude Fable 5 redeployment - Anthropic restored Fable 5 and Mythos 5 access on 2026-07-01 00:00 UTC after export controls were lifted; Fable 5 is available globally, while Mythos 5 remains limited to approved Project Glasswing partners.
  • Kimi K2.7 Code in GitHub Copilot - Moonshot AI's open-weight coding model became generally available in GitHub Copilot model picker on July 1, 2026; first open-weight model offered in Copilot; rolled out to Pro, Pro+, and Max plans.
  • Kimi K3 - Moonshot AI released Kimi K3 on 2026-07-16: 2.8T-parameter open-weight frontier model with 1M context, native vision, $3.00/$15.00 API pricing, 90% cache-hit discount ($0.30 cached input), and open weights promised by 2026-07-27 under a modified MIT license. Source: πŸ”—
  • Meta Muse Spark 1.1 - Meta released Muse Spark 1.1 on 2026-07-09: closed-weights frontier multimodal reasoning model built for agentic tool and computer use, 1M context, available via Meta Model API in public preview for US developers. Source: πŸ”—
  • MiniMax M2 & M2.1 open-source release - MiniMax open-sourced MiniMax M2 (230B/10B active MoE, 1M context, Apache 2.0) and M2.1 (229B/10B active MoE, 1M context, Apache 2.0) on July 2, 2026; API priced at $0.30/$1.20 per 1M tokens; available on Hugging Face, vLLM, SGLang, and MiniMax Agent.
  • Gemini Omni Flash and Nano Banana 2 Lite - Google released Gemini Omni Flash for video generation ($0.10/s) and Nano Banana 2 Lite for image generation (4s generation, fastest Gemini Image model) on June 30, 2026; available in Google AI Studio, Gemini API, and Gemini Enterprise Agent Platform.
  • GLM-5.2 open-source release - Zhipu AI open-sourced GLM-5.2 (753B MoE, 1M context, MIT license) on 2026-06-17; available via API at $1.40/$4.40 per 1M; leads AIME 2026 at 99.2%.
  • MCP 2026-07-28 spec RC - Largest revision since initial release; six breaking changes including stateless transport, two new required HTTP headers, Tasks lifecycle primitive; SDK maintainers have 10 weeks to ship support.
  • GPT-Live voice models - OpenAI launched GPT-Live (1, 1 mini, Medium, High) on 2026-07-08, replacing Advanced Voice Mode inside ChatGPT; full-duplex consumer voice stack with GPT-5.5 backend delegation. No developer API at launch (Realtime API with gpt-realtime-2.1, shipped 2026-07-06, is the API path). See the TTS Proprietary table entry.
  • Codex CLI 0.144.5 - OpenAI shipped a patch release on 2026-07-16 that expands dangerous-command detection, including more forced rm forms, and returns clearer denial reasons when a command is blocked.
  • Cursor 3.11 and Kiro 1.0.138 - Cursor added side chats, conversation search, and simpler project/repo pickers on 2026-07-10. Kiro's 2026-07-13 IDE release improved session startup, fixed compaction loops on large sessions, added full PowerShell trust on Windows, and improved MCP tool recovery after transient network failures.
  • skills.sh β€” Primary registry for Agent Skills packages (see Agent Skills & Registries 🎯); pairs naturally with CLI-first workflows (Claude Code, Gemini CLI, MCP-capable agents).
  • CrewAI β€” Multi-agent framework with OSS core plus AMP Studio / AMP Cloud paths; vendor positioning emphasizes transparent orchestration, visual workflows, broad tool integrations, and team-scale adoption (see the CrewAI row in the Multi-Agent / Parallel Agent Platforms table in this document).
  • Aiden β€” Listed above under local desktop agents as a local-first, Windows-native autonomous stack for running agents without mandatory cloud offload (taracodlabs/aiden).

Last Updated: 2026-07-19 16:54 UTC Maintained by: ReadyPixels LLC


πŸ“– Additional Resources

Related Projects

Community

Support


πŸ“„ License

CC BY-NC 4.0

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.


Made with ❀️ by ReadyPixels LLC

Star on GitHub

About

Research-based comparison of AI models, development tools, and automation resources. Compare releases, pricing, benchmarks, and deployment options from official sources.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

Watchers

Forks

Releases

Packages

Contributors