Browse all AI Gateway models
Every model available on Vercel AI Gateway, with API access, pricing, and a playground. 393 models · Page 4 of 7.
Search and filter all models →- Thinking MachinesInklingInkling is a multimodal MoE model (975B total, 41B active, 256k context) reasoning over text, image, and audio inputs.
- Thinking MachinesInkling SmallInkling-Small is a lighter-weight model with 12B active parameters, trained with a similar recipe, to Inkling that achieves strong performance with even lower cost and latency.
- InterfazeInterfaze BetaInterfaze is an AI model built on a new architecture that merges specialized DNN/CNN models with LLMs for developer tasks that require deterministic output and high consistency like OCR, scraping, classification, STT and more.
- TypeSafe AIJevJev is TypeSafe AI’s System One decision model for fast, structured decisions in software. It evaluates shared state against typed questions and returns choices, scores, and boolean probabilities, supporting classification, routing, rubric-based assessment, and automated verification. Multiple questions can be evaluated in parallel within a single request.
- Moonshot AIKimi K2 InstructKimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities.
- Moonshot AIKimi K2 ThinkingKimi K2 Thinking is an advanced open-source thinking model by Moonshot AI. It can execute up to 200 – 300 sequential tool calls without human interference, reasoning coherently across hundreds of steps to solve complex problems. Built as a thinking agent, it reasons step by step while using tools, achieving state-of-the-art performance on Humanity's Last Exam (HLE), BrowseComp, and other benchmarks, with major gains in reasoning, agentic search, coding, writing, and general capabilities.
- Moonshot AIKimi K2.5kimi-k2.5 is Kimi's most versatile model to date, featuring a native multimodal architecture that supports both visual and text input, thinking and non-thinking modes, and dialogue and agent tasks.
- Moonshot AIKimi K2.6Kimi K2.6 demonstrates particularly strong performance in long-horizon coding tasks and produces professional-grade design with code and vision.
- Moonshot AIKimi K2.7 CodeKimi-K2.7-Code is a coding model from Moonshot AI. It has improved coding & agent performance over K2.6, more reasoning efficiency with less overthinking, and improved instruction following for long-horizon coding.
- Moonshot AIKimi K2.7 Code High SpeedKimi K2.7 Code HighSpeed is the high-speed version of Kimi K2.7 Code, the same model as Kimi K2.7 Code, but with an output speed of approximately 180 Tokens/s and up to 260 Tokens/s in short context scenarios, delivering a more extreme coding experience.
- Moonshot AIKimi K3Kimi’s flagship model for long-horizon coding and end-to-end knowledge work, with a 1M-token context window.
- Moonshot AIKimi K3 FastFast version of Kimi’s flagship model for long-horizon coding and end-to-end knowledge work, with a 1M-token context window.
- Kling AIKling v2.5 Turbo Image-to-VideoKling 2.5 Turbo is a major update to the AI video generation model focused on significantly improving speed, video quality, temporal stability, and creative control for creators, making professional-grade AI-generated video faster, more coherent, and easier to direct from text prompts.
- Kling AIKling v2.5 Turbo Text-to-VideoKling 2.5 Turbo is a major update to the AI video generation model focused on significantly improving speed, video quality, temporal stability, and creative control for creators, making professional-grade AI-generated video faster, more coherent, and easier to direct from text prompts.
- Kling AIKling v2.6 Image-to-VideoKling 2.6 introduces a groundbreaking "Native Audio" capability, enabling the generation of complete videos in a single go, including natural voice, action sound effects, and environmental ambient sounds, providing an immersive "what you see if what you hear" experience.
- Kling AIKling v2.6 Motion ControlKling 2.6 introduces a groundbreaking "Native Audio" capability, enabling the generation of complete videos in a single go, including natural voice, action sound effects, and environmental ambient sounds, providing an immersive "what you see if what you hear" experience.
- Kling AIKling v2.6 Text-to-VideoKling 2.6 introduces a groundbreaking "Native Audio" capability, enabling the generation of complete videos in a single go, including natural voice, action sound effects, and environmental ambient sounds, providing an immersive "what you see if what you hear" experience.
- Kling AIKling v3.0 Image-to-VideoBuild upon an All-in-One product framework, the Kling 3.0 model series supports full multimodal input and output spanning text, images, audio, and video, bringing the understanding, generation, and editing of video together in one streamlined AI workflow. The models integrate multiple tasks, including text-to-video, image-to-video, reference-to-video, and in-video editing, into a single, native multimodal architecture, enabling the models to follow complex narrative logic, deliver precise shot control, and maintain strong prompt adherence.
- Kling AIKling v3.0 Motion ControlKling 3.0 delivers a major leap in character fidelity for motion-driven generation, with stable facial features across multi-angle and long-duration motion, accurate complex emotions from multi-image face references, identity preservation through partial occlusions (hats, hands, fans), and steady clarity as the camera zooms, pans, or tracks.
- Kling AIKling v3.0 Text-to-VideoBuild upon an All-in-One product framework, the Kling 3.0 model series supports full multimodal input and output spanning text, images, audio, and video, bringing the understanding, generation, and editing of video together in one streamlined AI workflow. The models integrate multiple tasks, including text-to-video, image-to-video, reference-to-video, and in-video editing, into a single, native multimodal architecture, enabling the models to follow complex narrative logic, deliver precise shot control, and maintain strong prompt adherence.
- PoolsideLaguna S 2.1Laguna S 2.1 is Poolside's new open-weight model for agentic coding and long-horizon work.
- PoolsideLaguna S 2.1 FreeLaguna S 2.1 is Poolside's new open-weight model for agentic coding and long-horizon work.
- ConvaiinnovationsLayaLaya is Convai Innovations’ System One evaluation model for fast, structured decisions. It evaluates shared state against choice, score, and yes/no questions and returns typed answers with probabilities in a single forward pass. It supports classification, routing, guardrail checks, and rubric-based assessment without generating free-form text.
- InclusionaiLing 3.0 FlashLing 3.0 Flash is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.
- InclusionaiLing 3.0 Flash FinLing 3.0 Flash Fin is InclusionAI’s finance-enhanced MoE language model, combining 124 billion total parameters with approximately 5.1 billion active parameters for efficient financial reasoning. Its 256K context window, function calling, and support for complex, multi-step investment workflows make it ideal for financial research, analysis, long-horizon planning, and execution, while retaining strong capabilities in coding and mathematics.
- InclusionaiLing 3.0 Flash SanteLing-3.0-Flash-Sante is inclusionAI’s language model specialized for health and medicine, built on a Mixture-of-Experts architecture with 124 billion total parameters and approximately 5.1 billion active per token. With a 256K context window and function calling, it supports medical knowledge reasoning, evidence-based retrieval, and complex medical workflows while retaining general reasoning, coding, and agentic capabilities.
- InclusionaiLing 3.0 Flash VLLing 3.0 Flash VL builds on Ling 3.0 Flash with stronger language capabilities, native visual perception, and visual agent capabilities. It supports text, image, and video inputs with text output, reasoning, and function calling.
- InclusionaiLing 3.1 FlashLing 3.1 Flash is InclusionAI's hybrid reasoning language model for coding, multi-step analysis, and tool-using agents. It has 560B total parameters with 25B activated parameters and supports long-context text workflows.
- Liquid AILiquid d1Liquid d1 is a decision model for classification, routing, and scoring. It evaluates shared state against typed Choice, Score, and yes/no questions and returns structured answers with calibrated probabilities through the TypeSafe-compatible System One API.
- MetaLlama 3.1 70B InstructAn update to Meta Llama 3 70B Instruct that includes an expanded 128K context length, multilinguality and improved reasoning capabilities.
- MetaLlama 3.1 8B InstructAn update to Meta Llama 3 8B Instruct that includes an expanded 128K context length, multilinguality and improved reasoning capabilities.
- MetaLlama 3.3 70B InstructWhere performance meets efficiency. This model supports high-performance conversational AI designed for content creation, enterprise applications, and research, offering advanced language understanding capabilities, including text summarization, classification, sentiment analysis, and code generation.
- MetaLlama 4 Maverick 17B 128E Instruct FP8The Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding. Llama 4 Maverick, a 17 billion parameter model with 128 experts. Served by DeepInfra.
- MetaLlama 4 Scout 17B 16E InstructThe Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding. Llama 4 Scout, a 17 billion parameter model with 16 experts. Served by DeepInfra.
- Meituan LongCatLongCat 2.5 PreviewLongCat-2.5-Preview is Meituan’s multimodal reasoning model for coding and agentic workflows, with a 1M-token context window and support for text and image inputs. It combines image understanding, visual question answering, and content summarization with code generation, code understanding, and automated programming, and supports tool calling and optional thinking for complex tasks.
- Microsoft AIMAI-Transcribe 2MAI-Transcribe 2 is a multilingual speech-to-text model from Microsoft AI, ranked #1 on the FLEURS multilingual benchmark. It supports 60 languages with automatic language identification, code switching for mixed-language speech, speaker diarization, word-level timestamps, keyword biasing for domain-specific terminology, and configurable verbatim or clean transcription styles. It is suited for captions, call transcription, subtitling, accessibility, and other voice-enabled applications, and is faster than MAI-Transcribe-1.5 on long-form audio.
- Microsoft AIMAI-Transcribe-2-StreamingMAI-Transcribe-2-Streaming is a low-latency, "speech in, text out" realtime transcription API. Audio can be sent to the model as a continuous stream over a WebSocket connection and transcripts are returned incrementally as the speaker talks, with no need to wait for the utterance to finish. Intermediate results (aka Partials) update the current transcription; final results confirm a segment.
- Microsoft AIMAI-Voice-2Microsoft AI's highest-fidelity text-to-speech model, with expressive multilingual voices and consistent speakers across long-form narration. Served through Azure Speech (preview).
- Microsoft AIMAI-Voice-2-FlashMicrosoft AI's low-latency text-to-speech model for voice agents and assistants, with the expressive multilingual voices of MAI-Voice-2 at twice the speed.
- Microsoft AIMAI-Voice-2.1MAI-Voice-2.1 is our highest-fidelity, most expressive text-to-speech model, delivering rich, natural speech across 23 languages. It extends the MAI-Voice family with broad multilingual coverage, gated instant voice cloning, and strong long-form generation capabilities. With its detailed prosody, nuanced expressiveness, and studio-grade audio quality, MAI-Voice-2.1 is ideal for experiences where maximum voice quality and fidelity are required - long-form narration, brand-defining audio etc.
- Microsoft AIMAI-Voice-2.1-FlashMAI-Voice-2.1-Flash is a text-to-speech model built for fast, low-latency generation. It produces high- fidelity, natural, and expressive speech across 23 languages and supports gated instant voice cloning, all while being optimized for real-time responsiveness. Its human-like intonation, rhythm, and emotional nuance make it ideal for voice agents, assistants, and other interactive scenarios where latency and cost are critical.
- InceptionMercury 2A diffusion-based reasoning LLM that generates text via parallel refinement (not token-by-token), delivering real-time latency with ~1k tokens/sec plus 128K context and built-in tool/JSON support.
- InceptionMercury 2.5Mercury 2.5 is Inception’s diffusion-based reasoning model for chat, agents, and structured workflows, with tool calling, structured outputs, and a 260K context window.
- InceptionMercury Coder Small BetaMercury Coder Small is ideal for code generation, debugging, and refactoring tasks with minimal latency.
- XiaomiMiMo M2.5A native full-modal model supporting text, image, video, and audio understanding, with powerful Agent capabilities.
- XiaomiMiMo V2.5 ProMiMo V2.5 Pro delivers significant improvements over its predecessor, MiMo-V2-Pro, in general agentic capabilities, complex software engineering, and long-horizon tasks. MiMo-V2.5-Pro is a 1.02T-parameter Mixture-of-Experts model with 42B active parameters, built on a hybrid-attention architecture with a 1M-token context window.
- XiaomiMiMo V2.6 FlashMiMo V2.6 Flash is Xiaomi's efficient multimodal reasoning model for coding, automation, and everyday agent workflows. It accepts text, images, audio, and video within a 1M-token context window, with up to 128K tokens of output. Deep thinking, tool calling, JSON mode, and prompt caching make it suitable for applications that need multimodal understanding at a lower token cost.
- XiaomiMiMo V2.6 ProMiMo V2.6 Pro is Xiaomi's flagship multimodal reasoning model for complex software engineering, long-running agent tasks, and professional workflows. It supports text, image, audio, and video inputs with a 1M-token context window and up to 128K tokens of output. Deep thinking, tool calling, JSON mode, and prompt caching support applications that combine large inputs with multi-step reasoning.
- XiaomiMiMo V2.6 Pro UltraSpeedMiMo V2.6 Pro UltraSpeed is Xiaomi's accelerated inference offering for MiMo V2.6 Pro, designed for interactive agents and workflows where response latency matters. It combines multimodal understanding of text, images, audio, and video with a 1M-token context window and up to 128K tokens of output. It supports deep thinking, tool calling, JSON mode, and prompt caching.
- MiniMaxMiniMax H3H3 is a next-generation open-weights, general-purpose multimodal video model. Rather than being limited to specialized tasks such as generating, editing, or referencing, H3 understands multimodal contexts that bring together text, images, video, and audio. This enables it to interpret creative intent in a unified way and deliver more natural, coherent generation and expression.
- MiniMaxMiniMax H3 MaxMiniMax H3 Max is a next-generation general-purpose multimodal video model ranking #1 for overall quality, prompt understanding, and aesthetics in all evaluations, while generating a 5-second video in under 3 seconds.
- MiniMaxMiniMax M2MiniMax-M2 redefines efficiency for agents. It is a compact, fast, and cost-effective MoE model (230 billion total parameters with 10 billion active parameters) built for elite performance in coding and agentic tasks, all while maintaining powerful general intelligence.
- MiniMaxMiniMax M2.1MiniMax 2.1 is MiniMax's latest model, optimized specifically for robustness in coding, tool use, instruction following, and long-horizon planning.
- MiniMaxMiniMax M2.1 LightningMiniMax-M2.1-lightning is a faster version of MiniMax-M2.1, offering the same performance but with significantly higher throughput (output speed ~100 TPS, MiniMax-M2 output speed ~60 TPS).
- MiniMaxMiniMax M2.5MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. It is capable of handling the entire development process of various complex systems. It covers full-stack projects across multiple platforms including Web, Android, iOS, Windows, and Mac, encompassing server-side APIs, functional logic, and databases.
- MiniMaxMiniMax M2.5 High SpeedM2.5 highspeed: Same performance, faster and more agile (output speed approximately 100 tps)
- MiniMaxMiniMax M2.7M2.7 delivers outstanding performance in real-world software engineering, including end-to-end full project delivery, log analysis and bug troubleshooting, code security, machine learning, and more.
- MiniMaxMiniMax M2.7 High SpeedM2.7 Highspeed: Same performance, faster and more agile (output speed approximately 100 tps)
- MiniMaxMiniMax M3MiniMax-M3 is a frontier-class foundation model that unites the three capabilities defining today's frontier: a 1M-token context window, frontier coding and agentic performance, and native multimodality — the first open-weight model to deliver all three in a single system.
- MistralMinistral 14BMinistral 3 14B is the largest model in the Ministral 3 family, offering state-of-the-art capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. Optimized for local deployment, it delivers high performance across diverse hardware, including local setups.