Every maker
154 models
- GPT-6.1 Sol2026-09-29
GPT-6.1 Sol is a model from OpenAI for complex coding, computer use, and professional work, offering near-Astra performance at a lower cost.
CurrentProprietaryReasoning1.05M contextfrom $2.0018 hosts1.05M - Claude Sonnet 5.52026-09-28
Claude Sonnet 5.5 is a fast Claude model from Anthropic for everyday coding, agents, and knowledge work.
CurrentProprietaryReasoning1M contextfrom $2.0013 hosts1M - Claude Opus 5.52026-09-22
Claude Opus 5.5 is a Claude model for long-running agentic coding and knowledge work.
CurrentProprietaryReasoning1M contextfrom $4.0015 hosts1M - GPT-6 Luna2026-09-22
GPT-6 Luna is OpenAI's most efficient model for focused, high-volume tasks.
CurrentProprietaryReasoning1.05M contextfrom $0.0822 hosts1.05M - GPT-6 Sol2026-09-22
GPT-6 Sol is an OpenAI model for complex coding and agentic workflows.
ProprietaryReasoning1.05M contextfrom $1.6022 hosts1.05M - MiMo-V2.6-Flash2026-09-22
MiMo-V2.6-Flash is a multimodal model from Xiaomi for coding agents and long-context automation.
CurrentOpen weightsReasoning1.05M contextfrom $0.0419 hosts1.05M - MiMo-V2.6-Pro2026-09-22
MiMo-V2.6-Pro is Xiaomi's Pro model in the MiMo line, built for multimodal coding agents and long-context automation.
Open weightsReasoning1.05M contextfrom $0.4319 hosts1.05M - Grok 4.72026-09-21
Grok 4.7 is xAI's frontier model for long-running agents, coding, knowledge work, and visual projects.
CurrentProprietaryReasoning500K contextfrom $1.6017 hosts500K - GLM-5.3-FlashX2026-09-18
GLM-5.3-FlashX is a high-speed serving option from Z.ai for coding and agent workflows.
CurrentOpen weightsReasoning1M contextfrom $0.158 hosts1M - Qwen3.8 Omni Flash2026-09-17
Qwen3.8 Omni Flash is an omni model from Alibaba's Qwen line for text, vision, audio, and multimodal agent tasks.
CurrentProprietaryReasoning1M contextfrom $0.11269 hosts1M - Step 5 Preview2026-09-16
Step 5 Preview is StepFun's next-generation flagship base model for coding and professional knowledge work.
CurrentProprietaryReasoning1M contextfrom $0.9599 hosts1M - DeepSeek V4 Flash Vision Exp2026-09-10
DeepSeek V4 Flash Vision Exp is a model in the deepseek-flash line, described by the maker as a DeepSeek V4.1 Flash model for reasoning and agentic coding.
Open weightsReasoning1M contextfrom $0.022463 hosts1M - DeepSeek V4.1 Flash2026-09-10
DeepSeek V4.1 Flash is a model for reasoning and agentic coding, part of the deepseek-flash line.
CurrentOpen weightsReasoning1M contextfrom $0.153 hosts1M - Mercury 2.52026-09-08
Mercury 2.5 is a reasoning LLM and diffusion LLM (dLLM) from Inception.
CurrentProprietaryReasoning260K contextfrom $0.045 hosts260K - GPT-6 Astra2026-09-04
GPT-6 Astra is OpenAI's most capable model for complex reasoning, coding, computer use, research, and document creation.
CurrentProprietaryReasoning1.05M contextfrom $8.0024 hosts1.05M - Gemini 3.8 Flash2026-09-02
Gemini 3.8 Flash is Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows.
CurrentProprietaryReasoning1.05M contextfrom $0.7522 hosts1.05M - Muse Spark 1.32026-09-02
Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows.
CurrentProprietaryReasoning1.05M contextfrom $1.2510 hosts1.05M - Claude Fable 5.12026-09-01
Claude Fable 5.1 is a Claude model for demanding reasoning and long-horizon agentic work.
CurrentProprietaryReasoning1M contextfrom $10.0016 hosts1M - GLM-5.3-Flash2026-08-26
GLM-5.3-Flash is a native multimodal GLM model from Z.ai (Zhipu, GLM) designed for efficient coding and long-horizon agent tasks.
Open weightsReasoning1M contextfrom $0.0359 hosts1M - Qwen3.8 Flash2026-08-26
Qwen3.8 Flash is a vision-language model from Alibaba's Qwen line, designed for visual reasoning, documents, and agent tasks.
ProprietaryReasoning1M contextfrom $0.1125 hosts1M - GLM-5.32026-08-14
GLM-5.3 is the flagship model in the GLM line from Z.ai (Zhipu).
CurrentOpen weightsReasoning1M contextfrom $0.4057 hosts1M - Gemini 3.7 Flash2026-08-13
Gemini 3.7 Flash is a high-efficiency model from Google for agentic workflows, coding, and multimodal reasoning.
ProprietaryReasoning1.05M contextfrom $0.7522 hosts1.05M - DeepSeek V4 Pro2026-08-12
DeepSeek V4 Pro is a snapshot in the deepseek-thinking line, released on 2026-08-12.
CurrentOpen weightsReasoning1M contextfrom $0.208866 hosts1M - Grok 4.62026-08-12
Grok 4.6 is xAI's frontier model for long-running agents, coding, knowledge work, and visual projects.
ProprietaryReasoning500K contextfrom $2.0024 hosts500K - Nemotron 3.5 Lightning 30B A3B2026-08-11
Nemotron 3.5 Lightning 30B A3B is a fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads.
CurrentOpen weightsReasoning262K contextfrom $0.003 hosts262K - Solar Pro 42026-08-06
Solar Pro 4 is Upstage's flagship model, specialized for agentic use.
CurrentProprietaryReasoning524K contextfrom $0.035 hosts524K - Muse Spark 1.22026-08-05
Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1 from Meta.
ProprietaryReasoning1.05M contextfrom $1.2513 hosts1.05M - Qwen3.8 Max2026-08-03
Qwen3.8 Max is a 2.4-trillion-parameter mixture-of-experts flagship from Alibaba's Qwen line, aimed at coding, professional work, multimodal understanding, and long-horizon agentic workflows.
ProprietaryReasoning1M contextfrom $0.33831 hosts1M - Sakana Namazu2026-08-03
Sakana Namazu is a Japanese-specialized reasoning model based on Kimi K2.6, tuned for Japanese language, culture, and business workflows.
CurrentProprietaryReasoning262K contextfrom $0.955 hosts262K - Claude Opus 52026-07-24
Claude Opus 5 is Anthropic's strongest Claude Opus model for coding, agents, and professional work.
ProprietaryReasoning1M contextfrom $5.0031 hosts1M - Gemini 3.5 Flash Lite2026-07-21
Gemini 3.5 Flash Lite is a fast model from Google that balances multimodal reasoning, tool use, and cost.
CurrentProprietaryReasoning1.05M contextfrom $0.1524 hosts1.05M - Gemini 3.6 Flash2026-07-21
Gemini 3.6 Flash is a fast Gemini model that balances multimodal reasoning, tool use, and cost.
ProprietaryReasoning1.05M contextfrom $0.37523 hosts1.05M - Laguna S 2.12026-07-21
Laguna S 2.1 is an agentic coding model from Poolside in the XS size class for local deployment.
CurrentOpen weightsReasoning1.05M contextfrom $0.097 hosts1.05M - Kimi K32026-07-16
Kimi K3 is a multimodal model from Moonshot AI (Kimi) with a 1M token context and toggleable max-effort thinking, designed for long-horizon agent work.
CurrentOpen weightsReasoning1.05M contextfrom $0.7268 hosts1.05M - Qwen3.7 Flash2026-07-15
Qwen3.7 Flash is a lightweight multimodal Qwen model for high-throughput text, image, and video tasks.
ProprietaryReasoning1M contextfrom $0.028212 hosts1M - GPT-5.62026-07-09
GPT-5.6 is a frontier model from OpenAI for complex professional work, coding, and agentic workflows.
ProprietaryReasoning1.05M contextfrom $1.502 hosts1.05M - GPT-5.6 Luna2026-07-09
GPT-5.6 Luna is a cost-efficient model from OpenAI for fast, high-volume workloads.
ProprietaryReasoning1.05M contextfrom $0.0632 hosts1.05M - GPT-5.6 Sol2026-07-09
GPT-5.6 Sol is a frontier model from OpenAI for complex professional work, coding, and agentic workflows.
ProprietaryReasoning1.05M contextfrom $1.5033 hosts1.05M - GPT-5.6 Terra2026-07-09
GPT-5.6 Terra is a balanced model from OpenAI for capable, cost-efficient everyday work.
CurrentProprietaryReasoning1.05M contextfrom $1.5031 hosts1.05M - Grok 4.52026-07-08
Grok 4.5 is xAI's model for chat, coding, agentic tools, and lower hallucination risk.
ProprietaryReasoning500K contextfrom $2.0024 hosts500K - Laguna XS 2.12026-07-02
Laguna XS 2.1 is an agentic coding model from Poolside in the XS size class, designed for local deployment.
Open weightsReasoning262K contextfrom $0.065 hosts262K - LongCat-2.02026-06-30
Meituan LongCat-2.0 is a reasoning model with tool calling and a 1M-token context window.
CurrentOpen weightsReasoning1M contextfrom $0.307 hosts1M - Claude Sonnet 52026-06-29
Claude Sonnet 5 is an everyday Claude agent model for coding, planning, browsing, and general work.
ProprietaryReasoning1M contextfrom $1.4434 hosts1M - Seed 2.1 Pro2026-06-23
Seed 2.1 Pro is ByteDance's flagship model in the Seed line, designed for complex multimodal reasoning, coding, and agents.
CurrentProprietaryReasoning256K contextfrom $0.8906256K - Seed 2.1 Turbo2026-06-23
Seed 2.1 Turbo is a faster variant of ByteDance's Seed 2.1 model, designed for multimodal reasoning and latency-sensitive agent workflows.
ProprietaryReasoning256K contextfrom $0.4453256K - Fugu2026-06-15
Fugu is a multi-agent model for routing expert agents across complex analytical tasks.
CurrentProprietaryReasoning1M contextfrom $2.007 hosts1M - Fugu Ultra2026-06-15
Fugu Ultra is a multi-agent model from Sakana AI, designed for hard research, analysis, and competitions.
ProprietaryReasoning1M contextfrom $5.0011 hosts1M - GLM-5.22026-06-13
GLM-5.2 is an open flagship model from Z.ai (Zhipu) for long-horizon coding agents and million-token context work.
Open weightsReasoning1M contextfrom $0.3074 hosts1M - Kimi K2.7 Code HighSpeed2026-06-12
Kimi K2.7 Code HighSpeed is a coding-focused model from Moonshot AI, designed for long-horizon repository work with less overthinking.
CurrentOpen weightsReasoning262K contextfrom $0.27553 hosts262K - Claude Fable 52026-06-07
Claude Fable 5 is a Claude model for creative writing, analysis, and controlled agent workflows.
ProprietaryReasoning1M contextfrom $3.0035 hosts1M - Nemotron 3 Ultra 550B A55B2026-06-04
Nemotron 3 Ultra 550B A55B is the largest model in the Nemotron 3 line, designed for maximum open-weight reasoning and agent accuracy.
Open weightsReasoning1M contextfrom $0.5012 hosts1M - Qwen3.7 Plus2026-06-02
Qwen3.7 Plus is a multimodal model from Alibaba's Qwen line, positioned as a workhorse for long-context agents, visual inputs, and coding.
ProprietaryReasoning1M contextfrom $0.28228 hosts1M - MiniMax-M32026-06-01
MiniMax-M3 is a multimodal model for long-context coding, perception, and agent planning.
CurrentOpen weightsReasoning1M contextfrom $0.2050 hosts1M - Step 3.7 Flash2026-05-29
Step 3.7 Flash is a newer StepFun flash model designed for faster agents, coding, and multimodal prompts.
Open weightsReasoning256K contextfrom $0.18518 hosts256K - Claude Opus 4.82026-05-28
Claude Opus 4.8 is Anthropic's top Claude Opus tier, designed for the hardest reasoning, coding, and long-horizon agents.
ProprietaryReasoning1M contextfrom $0.42528 hosts1M - Qwen3.7 Max2026-05-21
Qwen3.7 Max is the first version of the Qwen line in this registry.
ProprietaryReasoning1M contextfrom $0.82530 hosts1M - Command A Plus2026-05-20
Command A Plus is Cohere's stronger command model for multilingual agents and enterprise workflows.
CurrentOpen weightsReasoning128K contextfrom $0.303 hosts128K - Gemini 3.5 Flash2026-05-19
Gemini 3.5 Flash is a fast model in the gemini-flash line, balancing multimodal reasoning, tool use, and cost.
ProprietaryReasoning1.05M contextfrom $0.185731 hosts1.05M - Mistral Medium2026-04-29
Mistral Medium is a balanced model from Mistral AI for enterprise assistants, multilingual work, and tools.
CurrentOpen weightsReasoning262K contextfrom $0.406 hosts262K - Mistral Medium 3.52026-04-29
Mistral Medium 3.5 is a balanced Mistral model for enterprise assistants, multilingual work, and tools.
Open weightsReasoning262K contextfrom $1.503 hosts262K - Laguna M.12026-04-28
Laguna M.1 is Poolside's open-weight model for agentic coding and long-hizon work.
Open weightsReasoning262K contextfrom $0.002 hosts262K - Nemotron 3 Nano Omni2026-04-28
Nemotron 3 Nano Omni is an open omni model from NVIDIA that combines reasoning with text, vision, and audio.
Open weightsReasoning256K contextfrom $0.106 hosts256K - Qwen3.6 Flash2026-04-27
Qwen3.6 Flash is a vision-language model from Alibaba's Qwen line for visual reasoning, documents, and agent tasks.
CurrentProprietaryReasoning1M contextfrom $0.16518 hosts1M - GPT-5.52026-04-23
GPT-5.5 is OpenAI's default frontier model for coding, computer use, research, and knowledge work.
CurrentProprietaryReasoning1.05M contextfrom $0.187539 hosts1.05M - GPT-5.5 Pro2026-04-23
GPT-5.5 Pro is the highest-accuracy tier in the GPT-5.5 line, designed for slower, precision-heavy reasoning and coding.
CurrentProprietaryReasoning1.05M contextfrom $27.2718 hosts1.05M - MiMo-V2.52026-04-22
MiMo-V2.5 is an open model from Xiaomi for multimodal coding agents and long-context automation.
Open weightsReasoning1.05M contextfrom $0.1422 hosts1.05M - MiMo-V2.5-Pro2026-04-22
MiMo-V2.5-Pro is Xiaomi's Pro tier in the MiMo line, built for multimodal reasoning and coding-agent execution.
Open weightsReasoning1.05M contextfrom $0.4023 hosts1.05M - Kimi K2.62026-04-21
Kimi K2.6 is a multimodal model from Moonshot AI, positioned as a workhorse for agent loops, coding tasks, and visual context.
Open weightsReasoning262K contextfrom $0.27559 hosts262K - Grok 4.32026-04-17
Grok 4.3 is xAI's model for chat, coding, agentic tools, and lower hallucination risk.
ProprietaryReasoning1M contextfrom $1.2524 hosts1M - Grok Build 0.12026-04-16
Grok Build 0.1 is a coding model from xAI, tuned for agentic engineering and iterative edits.
CurrentProprietaryReasoning256K contextfrom $1.0016 hosts256K - Claude Opus 4.72026-04-14
Claude Opus 4.7 is Anthropic's stronger Opus tier for advanced software work and high-stakes reasoning.
ProprietaryReasoning1M contextfrom $5.0027 hosts1M - Muse Spark 1.12026-04-08
Muse Spark 1.1 is a natively multimodal reasoning model from Meta.
ProprietaryReasoning1.05M contextfrom $1.2512 hosts1.05M - GLM-5.12026-04-07
GLM-5.1 is a coding model from Z.ai (Zhipu) for agentic engineering, terminals, and repository generation.
Open weightsReasoning200K contextfrom $0.4548 hosts200K - Gemma 4 26B A4B IT2026-04-02
Gemma 4 26B A4B IT is an open instruction model from Google for efficient chat and self-hosted deployments.
CurrentOpen weightsReasoning262K contextfrom $0.04221 hosts262K - Gemma 4 31B IT2026-04-02
Gemma 4 31B IT is Google's largest instruction-tuned model in the Gemma line, designed for open, self-hosted chat and reasoning.
Open weightsReasoning262K contextfrom $0.0935 hosts262K - Step 3.5 Flash 26032026-04-02
Step 3.5 Flash 2603 is a flash model from StepFun for efficient multimodal reasoning, coding, and tool use, according to the maker.
Open weightsReasoning256K contextfrom $0.105 hosts256K - GLM-5V-Turbo2026-04-01
GLM-5V-Turbo is a vision model from Z.ai (Zhipu, GLM) for screenshots, documents, and multimodal agent tasks.
ProprietaryReasoning200K contextfrom $0.704217 hosts200K - Trinity Large Thinking2026-04-01
Trinity Large Thinking is a reasoning-optimized 398B MoE agent model from Arcee AI, designed for extended thinking in long-horizon and multi-turn tool use.
CurrentOpen weightsReasoning262K contextfrom $0.255 hosts262K - Mercury Edit 22026-03-30
Mercury Edit 2 is a code editing dLLM from Inception, designed for autocomplete (FIM) and next-edit suggestions.
Proprietary32K contextfrom $0.2532K - MiMo-V2-Omni2026-03-18
MiMo-V2-Omni is a legacy model from Xiaomi's mimo line, retained for compatibility with older integrations.
ProprietaryReasoning262K contextfrom $0.142 hosts262K - MiMo-V2-Pro2026-03-18
MiMo-V2-Pro is a model from Xiaomi for multimodal agents, reasoning, and code tasks.
ProprietaryReasoning1.05M contextfrom $0.4357 hosts1.05M - MiniMax-M2.72026-03-18
MiniMax-M2.7 is the open flagship model from MiniMax for coding agents, office automation, and complex environments.
Open weightsReasoning205K contextfrom $0.0838 hosts205K - GPT-5.4 mini2026-03-17
GPT-5.4 mini is a small model from OpenAI for coding subagents, quick tool use, and high-volume work.
CurrentProprietaryReasoning400K contextfrom $0.37532 hosts400K - GPT-5.4 nano2026-03-17
GPT-5.4 nano is OpenAI's cheapest GPT-5.4 lane, designed for simple routing, extraction, and bulk automation.
CurrentProprietaryReasoning400K contextfrom $0.1630 hosts400K - GLM-5-Turbo2026-03-16
GLM-5-Turbo is a faster lane in the GLM-5 line, designed for coding agents that need lower latency.
ProprietaryReasoning200K contextfrom $0.7218 hosts200K - Mistral Small2026-03-16
Mistral Small is an efficient model from Mistral AI for fast chat, extraction, and production assistants.
CurrentOpen weightsReasoning256K contextfrom $0.0756 hosts256K - Mistral Small 42026-03-16
Mistral Small 4 is a production model from Mistral AI for chat, extraction, and cost-sensitive agents.
Open weightsReasoning256K contextfrom $0.1510 hosts256K - Nemotron 3 Super2026-03-11
Nemotron 3 Super is the middle tier in NVIDIA's Nemotron line, designed for collaborative agents and high-volume reasoning workloads.
Open weightsReasoning262K contextfrom $0.059 hosts262K - Grok 4.20 (Non-Reasoning)2026-03-09
Grok 4.20 (Non-Reasoning) is a model from xAI for agentic tool use, reasoning, coding, and live assistance.
Proprietary1M contextfrom $1.256 hosts1M - Grok 4.20 Multi-Agent2026-03-09
Grok 4.20 Multi-Agent is a Grok model from xAI, described for agentic tool use, reasoning, coding, and live assistance.
ProprietaryReasoning1M contextfrom $1.251M - GPT-5.42026-03-05
GPT-5.4 is an agent-ready GPT for coding and computer-use workflows at a lower cost, according to OpenAI.
ProprietaryReasoning1.05M contextfrom $0.7538 hosts1.05M - GPT-5.4 Pro2026-03-05
GPT-5.4 Pro is a model from OpenAI for demanding professional reasoning and agent tasks.
ProprietaryReasoning1.05M contextfrom $24.0018 hosts1.05M - Gemini 3.1 Flash Lite Preview2026-03-03
Gemini 3.1 Flash Lite Preview is a legacy model in the gemini-flash-lite line, retained by Google for compatibility with older integrations.
ProprietaryReasoning1.05M contextfrom $0.12525 hosts1.05M - GPT-5.3 Chat2026-03-03
GPT-5.3 Chat is a chat-tuned GPT model from OpenAI for conversational assistance, writing, and tool workflows.
Proprietary128K contextfrom $1.756 hosts128K - Mercury 22026-02-24
Mercury 2 is a reasoning model from Inception for deliberate analysis, multi-step problem solving, and tool use.
ProprietaryReasoning128K contextfrom $0.256 hosts128K - Gemini 3.1 Pro Preview2026-02-19
Gemini 3.1 Pro Preview is a reasoning-first Gemini preview from Google for agentic coding and complex problem solving.
CurrentProprietaryReasoning1.05M contextfrom $1.0030 hosts1.05M - Sarvam-105B2026-02-18
Sarvam-105B is a flagship Indian-language reasoning model for enterprise multilingual applications.
CurrentOpen weightsReasoning131K contextfrom $0.043 hosts131K - Sarvam-30B2026-02-18
Sarvam-30B is an efficient Indian-language reasoning model for chat, coding, and multilingual work, as described by Sarvam AI.
Open weightsReasoning66K contextfrom $0.022 hosts66K - Claude Sonnet 4.62026-02-17
Claude Sonnet 4.6 is Anthropic's workhorse model for coding agents, careful analysis, and production cost control.
ProprietaryReasoning1M contextfrom $0.9031 hosts1M - Tiny Aya Earth2026-02-17
Tiny Aya Earth is a compact multilingual model from Cohere, specialized for West Asian and African languages.
CurrentOpen weights8K context8K - Tiny Aya Fire2026-02-17
Tiny Aya Fire is a compact multilingual model from Cohere, specialized for South Asian languages.
Open weights8K context8K - Tiny Aya Global2026-02-17
Compact multilingual model balanced across 70 languages and regions.
Open weights8K context8K - Tiny Aya Water2026-02-17
Tiny Aya Water is a compact multilingual model from Cohere, specialized for European and Asia-Pacific languages.
Open weights8K context8K - Seed 2.0 Code2026-02-14
ByteDance Seed 2.0 Code is a coding model for multimodal software engineering and long-running agents, as described by the maker.
ProprietaryReasoning262K contextfrom $0.4752 hosts262K - Seed 2.0 Lite2026-02-14
Seed 2.0 Lite is a cost-efficient model from ByteDance's Seed line, designed for production chat, analysis, and structured generation.
ProprietaryReasoning256K contextfrom $0.082 hosts256K - Seed 2.0 Mini2026-02-14
Seed 2.0 Mini is a lightweight ByteDance Seed 2.0 model for low-latency multimodal reasoning and high-volume tasks.
ProprietaryReasoning256K contextfrom $0.02972 hosts256K - Seed 2.0 Pro2026-02-14
ByteDance Seed 2.0 Pro is the flagship model in the Seed line, designed for complex multimodal reasoning and long-horizon agent workflows.
ProprietaryReasoning256K contextfrom $0.4752 hosts256K - GLM-52026-02-12
GLM-5 is a general GLM flagship for coding, analysis, and tool-heavy engineering workflows.
Open weightsReasoning205K contextfrom $0.5042 hosts205K - MiniMax-M2.52026-02-12
MiniMax-M2.5 is a coding model from MiniMax for agent workflows, office edits, and automation, according to the maker.
Open weightsReasoning205K contextfrom $0.1542 hosts205K - GPT-5.3 Codex2026-02-05
GPT-5.3 Codex is a coding-optimized GPT model for repository edits, reviews, and agentic software work.
CurrentProprietaryReasoning400K contextfrom $1.4025 hosts400K - GPT-5.3 Codex Spark2026-02-05
GPT-5.3 Codex Spark is a coding-optimized GPT model from OpenAI for repository edits, reviews, and agentic software work.
CurrentProprietaryReasoning128K contextfrom $1.753 hosts128K - Claude Opus 4.62026-02-04
Claude Opus 4.6 is Anthropic's high-end model for difficult coding, planning, and slower expert reasoning.
ProprietaryReasoning1M contextfrom $5.0025 hosts1M - Voxtral Mini2026-02-01
Voxtral Mini is a speech transcription model from Mistral AI, built for accurate audio-to-text and captioning workflows.
CurrentProprietary - Step 3.5 Flash2026-01-29
Step 3.5 Flash is StepFun's flash lane for quick multimodal reasoning and coding assistance.
Open weightsReasoning256K contextfrom $0.0914 hosts256K - GLM-4.7-Flash2026-01-19
GLM-4.7-Flash is a budget model in the GLM line for fast coding help, routing, and everyday automation, as described by Z.ai.
Open weightsReasoning200K contextfrom $0.0617 hosts200K - GLM-4.7-FlashX2026-01-19
GLM-4.7-FlashX is an efficient GLM model for fast reasoning, coding, and agent workflows.
Open weightsReasoning200K contextfrom $0.069 hosts200K - Jamba Mini2026-01-01
Jamba Mini is AI21 Labs' efficient, lightweight hybrid SSM-Transformer model for a wide range of tasks.
CurrentOpen weights256K contextfrom $0.20256K - solar-pro32026-01
Solar Pro 3 is the flagship model in the solar-pro line, designed for demanding analysis, coding, and production agent workflows.
ProprietaryReasoning131K contextfrom $0.152 hosts131K - MiniMax-M2.12025-12-23
MiniMax-M2.1 is an agent model from MiniMax for practical coding and productivity tasks.
Open weightsReasoning205K contextfrom $0.2724 hosts205K - Gemini 3 Flash Preview2025-12-17
Gemini 3 Flash Preview is a new model in Google's gemini-flash line.
ProprietaryReasoning1.05M contextfrom $0.0729 hosts1.05M - GPT-5.2 Chat2025-12-11
GPT-5.2 Chat is a chat-tuned GPT model from OpenAI for conversational assistance, writing, and tool workflows.
ProprietaryReasoning128K contextfrom $0.2528 hosts128K - GPT-5.2 Pro2025-12-11
GPT-5.2 Pro is a higher-accuracy variant of GPT-5.2 for tougher reasoning and review workflows.
ProprietaryReasoning400K contextfrom $18.9011 hosts400K - Devstral 22025-12-09
Devstral 2 is Mistral's coding-agent model for repository work, terminal tasks, and software fixes.
CurrentOpen weights262K contextfrom $0.4010 hosts262K - Devstral Small 22025-12-09
Devstral Small 2 is a legacy model from Mistral AI, retained for compatibility with older integrations.
Open weights256K contextfrom $0.00256K - GLM-4.6V-Flash2025-12-08
GLM-4.6V-Flash is a lightweight vision model from Z.ai for visual reasoning, documents, and multimodal agents.
Open weightsReasoning128K contextfrom $0.02185 hosts128K - Nova 2 Pro2025-12-03
Nova 2 Pro is a multimodal reasoning model from Amazon for visual analysis, planning, and tool use.
CurrentProprietaryReasoning1M contextfrom $0.001M - Nova 2 Lite2025-12-01
Nova 2 Lite is a multimodal reasoning model from Amazon for visual analysis, planning, and tool use.
CurrentProprietaryReasoning1M contextfrom $0.304 hosts1M - Claude Opus 4.52025-11-24
Claude Opus 4.5 is Anthropic's flagship model for deep reasoning, coding, and long-horizon agents.
ProprietaryReasoning200K contextfrom $0.7122 hosts200K - GPT-5.12025-11-13
GPT-5.1 is a model in OpenAI's GPT line, described as a sharper GPT-5 generation for coding, product work, and tool-assisted tasks.
ProprietaryReasoning400K contextfrom $1.0028 hosts400K - Nemotron Nano 12B v2 VL2025-10-28
Nemotron Nano 12B v2 VL is a multimodal model from NVIDIA for visual reasoning and agentic AI workflows.
Open weightsReasoning128K contextfrom $0.203 hosts128K - MiniMax-M22025-10-27
MiniMax-M2 is an open model from MiniMax, built for coding agents and tool-heavy workflows.
Open weightsReasoning205K contextfrom $0.1719 hosts205K - Claude Haiku 4.52025-10-15
Claude Haiku 4.5 is Anthropic's fast lane model for lightweight agents, office tasks, and responsive chat.
CurrentProprietaryReasoning200K contextfrom $0.1430 hosts200K - Gemini 2.5 Computer Use Preview 10-20252025-10-07
Gemini 2.5 Computer Use Preview 10-2025 is a specialized Gemini 2.5 model for browser-control agents that automate UI tasks.
ProprietaryReasoning128K contextfrom $1.252 hosts128K - GPT-5 Pro2025-10-06
GPT-5 Pro is a higher-accuracy tier in the GPT-5 line, intended for tough analysis, coding reviews, and planning.
ProprietaryReasoning400K contextfrom $13.5016 hosts400K - Granite-4.0-H-Small2025-10-02
Granite-4.0-H-Small is an open-weight hybrid model from IBM for enterprise chat, coding, retrieval-augmented generation, and tool-calling workloads.
CurrentOpen weights131K contextfrom $0.0636131K - Ling-1T2025-10
Ling-1T is an open-weight instruction model from Ant Group for adaptable chat and self-hosted production workloads.
CurrentOpen weights128K contextfrom $0.562 hosts128K - Ring-1T2025-10
Ring-1T is a reasoning model from Ant Group for deliberate analysis, multi-step problem solving, and tool use.
CurrentOpen weightsReasoning128K contextfrom $0.562 hosts128K - Claude Sonnet 4.52025-09-29
Claude Sonnet 4.5 is a balanced Claude model for coding, analysis, agent workflows, and cost control, per Anthropic.
ProprietaryReasoning1M contextfrom $0.4324 hosts1M - nvidia-nemotron-nano-9b-v22025-08-18
nvidia-nemotron-nano-9b-v2 is a compact Nemotron model from NVIDIA for efficient reasoning and deployable AI agents.
Open weightsReasoning131K contextfrom $0.00131K - Mistral Medium 3.12025-08-12
Mistral Medium 3.1 is a Mistral AI model for multilingual chat, reasoning, and tool-assisted workflows.
Proprietary262K contextfrom $0.40262K - GPT-52025-08-07
GPT-5 is OpenAI's original workhorse model for reasoning, coding, writing, and tool workflows.
ProprietaryReasoning400K contextfrom $1.0031 hosts400K - GPT-5 Mini2025-08-07
GPT-5 Mini is a small model from OpenAI for responsive agents, coding help, and everyday automation.
ProprietaryReasoning400K contextfrom $0.0431 hosts400K - GPT-5 Nano2025-08-07
GPT-5 Nano is a tiny model in the GPT-5 lane for routing, extraction, classification, and bulk jobs.
ProprietaryReasoning400K contextfrom $0.0426 hosts400K - GLM-4.5-Air2025-07-28
GLM-4.5-Air is a lighter variant of GLM-4.5 from Z.ai (Zhipu, GLM).
CurrentOpen weightsReasoning131K contextfrom $0.1019 hosts131K - GLM-4.5-Flash2025-07-28
GLM-4.5-Flash is an efficient GLM model from Z.ai (Zhipu, GLM) for fast reasoning, coding, and agent workflows.
Open weightsReasoning131K contextfrom $0.004 hosts131K - Voxtral Small2025-07-15
Voxtral Small is an instruct model from Mistral AI with native audio input for speech understanding and tool use.
Open weights32K contextfrom $0.104 hosts32K - Devstral Medium2025-07-10
Devstral Medium is a legacy model in the devstral line, retained for compatibility with older integrations.
Open weights128K contextfrom $0.402 hosts128K - Devstral Small2025-07-10
Devstral Small is a legacy model from Mistral AI, retained for compatibility with older integrations.
Open weights128K contextfrom $0.102 hosts128K - Jamba Large2025-07-01
Jamba Large is AI21 Labs' hybrid SSM-Transformer long-context model, built for enterprise agents and grounded generation.
Open weights256K contextfrom $2.00256K - Mistral Small 3.22025-06-20
Mistral Small 3.2 is an efficient Mistral model for fast chat, extraction, and production assistants.
Open weights128K contextfrom $0.103 hosts128K - Gemini 2.5 Flash2025-06-17
Gemini 2.5 Flash is a fast multimodal model for applications where latency and price matter.
ProprietaryReasoning1.05M contextfrom $0.0931 hosts1.05M - Gemini 2.5 Flash-Lite2025-06-17
Gemini 2.5 Flash-Lite is a lean version of Gemini 2.5 for cheap multimodal traffic and quick agents.
ProprietaryReasoning1.05M contextfrom $0.0724 hosts1.05M - Gemini 2.5 Pro2025-06-17
Gemini 2.5 Pro is Google's reasoning model for coding, math, and multimodal analysis.
ProprietaryReasoning1.05M contextfrom $0.8730 hosts1.05M - o3-pro2025-06-10
o3-pro is a high-effort tier for difficult technical reasoning and careful answers.
CurrentProprietaryReasoning200K contextfrom $18.009 hosts200K