モデル

OpenAI
gpt-6-luna25%割引
—
入力: $0.075$0.1 / 100万トークン出力: $0.375$0.5 / 100万トークンキャッシュ読取: $0.0075$0.01 / 100万トークン

GPT-6 Luna is a lightweight, hyper-efficient reasoning model released by OpenAI on September 22, 2026. As the most cost-effective member of the GPT-6 family, it leverages Astra’s alignment and reasoning advances to excel in high-volume, focused tasks like summarization, extraction, and rapid coding. Supporting a 1.05M context window and 128k max output, Luna offers adjustable reasoning effort tiers. With API pricing slashed by 50%, it delivers ultra-low-cost, high-frequency intelligence tailored for scaled agentic workflows.

コンテキスト 1.05M
OpenAI
gpt-6-sol25%割引
—
入力: $1.5$2 / 100万トークン出力: $7.5$10 / 100万トークンキャッシュ読取: $0.15$0.2 / 100万トークン

GPT-6 Sol is a mid-tier, fully-featured multimodal reasoning model released by OpenAI on September 22, 2026. Serving as the versatile backbone of the GPT-6 family, it perfectly balances advanced performance, speed, and cost. Sol supports a 1.05M-token context window. Thanks to its native multimodal architecture, it excels not only in multilingual text and complex coding but also in advanced audio/video understanding and multi-step tool reasoning. Positioned as a high-value core engine, it is optimized for enterprise automation, multimodal agents, and heavy data-processing pipelines.

コンテキスト 1.05M
Claude
claude-opus-5-525%割引
—
入力: $3$4 / 100万トークン出力: $15$20 / 100万トークンキャッシュ読取: $0.15$0.2 / 100万トークン

Claude 5.5 Opus is Anthropic's flagship large language model tailored for complex, long-running agentic coding and advanced knowledge work. As the pioneer of the 5.5 family, it features a 1M-token context window and an expanded 128k max output capacity. Powered by always-on adaptive thinking, the model autonomously adjusts its reasoning depth based on task complexity, delivering unparalleled efficiency in multi-step software engineering and extensive document analysis. Remarkably, it achieves this superior performance while reducing API costs by 20% compared to its predecessor.

コンテキスト 1M
Z.ai
glm-5.3-flashx
—
入力: $0.375 / 100万トークン出力: $1.25 / 100万トークンキャッシュ読取: $0.075 / 100万トークン

GLM-5.3-FlashX is a high-speed native multimodal LLM by Z.AI. Using a sparse-linear attention hybrid architecture with 320B parameters (18B active), it supports a 1M context window and speeds up to 200 Tokens/s. It integrates multimodal perception and a visual-coding feedback loop for autonomous agents and engineering.

コンテキスト 1M
Qwen
qwen3.7-flash20%割引
—
入力: $0.08$0.1 / 100万トークン出力: $0.32$0.4 / 100万トークンキャッシュ読取: $0.016$0.02 / 100万トークン

Qwen3.7-Flash is a lightweight, ultra-fast large language model open-sourced by Alibaba. It delivers high speed and low latency while maintaining excellent performance in multilingual understanding, coding, and logical reasoning, making it ideal for real-time chat and high-concurrency enterprise applications.

コンテキスト 1M
ByteDance
doubao-seedream-5-0-pro
—
課金: $0.047

Seedream 5.0 Pro is an image generation and editing model from ByteDance Seed. It is suited for commercial visual-production workflows that require precise editing control, lifelike scenes, and natural rendering.

コンテキスト 1M
ByteDance
seedance-2.515%割引
—
入力: $8.857$10.42 / 100万トークン出力: $8.857$10.42 / 100万トークン

Seedance 2.5 is ByteDance’s flagship AI video model launched in July 2026. It doubles single-pass runtime to a 30-second native long take in 4K without stitching. Supporting up to 50 multimodal references (30 images, 10 videos, 10 audio) with a 20% boost in prompt adherence, it introduces region-level editing and 3D white-model control to lock character and lighting consistency for professional video production.

ByteDance
seedance-2.015%割引
—
入力: $5.865$6.9 / 100万トークン出力: $5.865$6.9 / 100万トークン

Seedance 2.0 is a revolutionary audio-video joint generation model launched by ByteDance in February 2026. Pioneering quad-modal unified input, it seamlessly fuses text, images, video, and audio references. Delivering Hollywood-grade cinematic quality and precise camera control, it achieves industry-first native audio and accurate lip-sync output in a single pass, completely reshaping multi-scene narrative workflows.

DeepSeek
deepseek-v4.1-flash
—
入力: $0.3 / 100万トークン出力: $1.2 / 100万トークンキャッシュ読取: $0.006 / 100万トークン

DeepSeek V4.1 Flash is a next-gen native multimodal MoE model released in September 2026. Featuring an innovative CED asymmetric architecture and CSA2 mechanism, it slashes KV cache storage to 1/4. Offering Flash-level low pricing and speed, it surpasses the flagship V4 Pro in coding and Agent benchmarks, democratizing ultra-long context computing.

コンテキスト 1.05M
Suno
chirp-v6-mini
—
課金: $1

v6-mini is a faster, more efficient version available to everyone. It makes it easier than ever to turn an idea into a song – it delivers better, faster results than any free model on any music creation platform.

Suno
chirp-v6-wild
—
課金: $1

v6-wild is built for exploration. Available to Pro and Premier subscribers, it is less predictable and more varied, producing unexpected, textured and ambitious results. It gives you new ideas to riff on, build from or bring back into v6 for further refinement.

Suno
chirp-v6
—
課金: $1

v6 is reliable, precise and consistently delivers polished music across every genre and style. When you know what you want, v6 helps you get there.

OpenAI
gpt-6-astra25%割引
—
入力: $7.5$10 / 100万トークン出力: $37.5$50 / 100万トークンキャッシュ読取: $0.75$1 / 100万トークン

GPT-6 Astra is OpenAI’s next-generation multimodal AI model, designed for unprecedented reasoning and real-time adaptability. Operating as a highly intuitive digital collaborator, Astra seamlessly integrates advanced text, audio, and visual processing. It excels at breaking down complex, multi-step problems while maintaining a deeply contextual, human-centric tone. With enhanced speed and deep domain expertise, Astra bridges the gap between advanced computation and seamless everyday human interaction.

コンテキスト 1.05M
Claude
claude-fable-5-125%割引
—
入力: $7.5$10 / 100万トークン出力: $37.5$50 / 100万トークンキャッシュ読取: $0.1875$0.25 / 100万トークン

Released in September 2026, Claude Fable 5.1 is Anthropic's flagship model for complex, long-horizon tasks. It doubles scientific benchmark performance, slashes cache read costs by 75%, and introduces Enterprise Frontier Safeguards (EFS) with anti-distillation tech, setting a new standard for secure, autonomous AI.

コンテキスト 1M
Suno
chirp-v5
—
課金: $1

The flagship Chirp model, v5 produces the cleanest, most studio-ready tracks in the family, with the strongest vocals, mixing, and structure. Reach for it when final-quality output matters most.

Suno
chirp-v5-5
—
課金: $1

The flagship Chirp model, v5 produces the cleanest, most studio-ready tracks in the family, with the strongest vocals, mixing, and structure. Reach for it when final-quality output matters most.

Suno
chirp-v4-5
—
課金: $1

Chirp-auk maps to Suno v4.5, known for richer genre blending, stronger prompt adherence, and more expressive vocals than v4. It suits polished tracks when v5 is more than a job needs.

Suno
chirp-v4
—
課金: $1

Version 4 sharpens audio clarity and vocal quality over the v3 line, giving clean, reliable songs at lower cost. A solid default for everyday generation and high volume.

Suno
chirp-v3-5
—
課金: $1

An earlier Chirp release, v3.5 generates full songs quickly and cheaply, useful for drafts, prototyping, and cost-sensitive batches where top fidelity is not required.

Qwen
qwen3.8-flash
—
入力: $0.16 / 100万トークン出力: $0.47 / 100万トークンキャッシュ読取: $0.016 / 100万トークン

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

コンテキスト 1.05M
Qwen
qwen3.8-max
—
入力: $2 / 100万トークン出力: $6 / 100万トークンキャッシュ読取: $0.17 / 100万トークン

Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding, coding, and agentic workflows.

コンテキスト 1.05M
Z.ai
glm-5.3-flash
—
入力: $0.15 / 100万トークン出力: $0.5 / 100万トークンキャッシュ読取: $0.03 / 100万トークン

GLM-5.3-Flash is the first native multimodal model in the GLM-5 series, delivering stronger intelligence than GLM-5.2 while maintaining an exceptionally cost-efficient architecture

コンテキスト 1.05M
DeepSeek
deepseek-v4-flash-073120%割引
—
入力: $0.112$0.14 / 100万トークン出力: $0.224$0.28 / 100万トークンキャッシュ読取: $0.00224$0.0028 / 100万トークン

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash.

コンテキスト 1.05M
DeepSeek
deepseek-v4-pro-081320%割引
—
入力: $1.392$1.74 / 100万トークン出力: $2.784$3.48 / 100万トークンキャッシュ読取: $0.0116$0.0145 / 100万トークン

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

コンテキスト 1.05M
Z.ai
glm-5.3
—
入力: $1.4 / 100万トークン出力: $4.4 / 100万トークンキャッシュ読取: $0.26 / 100万トークン

Released in August 2026, GLM-5.3 is Z.ai's 743B-parameter open-weights flagship model. Relying purely on post-training scaling with a 1M context window, it delivers open-source SOTA in long-horizon agentic tasks and complex software engineering. It boasts breakthrough emergent cybersecurity capabilities, closely rivaling closed frontier systems like Claude Fable 5.

コンテキスト 1M
MoonshotAI
kimi-k3
—
入力: $3 / 100万トークン出力: $15 / 100万トークンキャッシュ読取: $0.3 / 100万トークン

Kimi K3 is a powerful 2.8-trillion-parameter open-weight multimodal reasoning model developed by Moonshot AI. Featuring a massive 1-million-token context window and native vision capabilities, it is specifically engineered for long-horizon coding, complex software engineering, and advanced agentic knowledge work. By utilizing innovative architectures like Kimi Delta Attention, it delivers exceptional scaling efficiency and autonomous problem-solving performance.

コンテキスト 1.05M
Minimax
MiniMax-Hailuo-2.3-Fast20%割引
—
按second課金: $0.032$0.04

Hailuo 2.3-Fast is a speed-optimized, low-latency AI video model developed by MiniMax, tailored specifically for Image-to-Video (I2V) workflows. It cuts generation time and API costs by up to 50%, delivering 768P/1080P clips in around 55 seconds. While preserving strong structural stability and smooth physical motion, it provides creative professionals with an exceptionally cost-effective solution for rapid concept prototyping and high-volume asset production.

Minimax
MiniMax-Hailuo-2.320%割引
—
按second課金: $0.04$0.05

Hailuo 2.3 is the flagship AI video generation model developed by MiniMax, supporting both Text-to-Video and Image-to-Video. It delivers major breakthroughs in complex body motion, adherence to physics, and precise prompt following. The model captures subtle facial micro-expressions with cinematic lighting, while ensuring excellent style stability across anime and e-commerce shots with reduced flicker.

Minimax
MiniMax-M320%割引
—
入力: $0.24$0.3 / 100万トークン出力: $0.96$1.2 / 100万トークンキャッシュ読取: $0.024$0.03 / 100万トークン

MiniMax M3 is the flagship open-source MoE large language model developed by MiniMax. Built on a proprietary Modular Sparse Attention (MSA) architecture, it holds 428 billion total parameters while activating only 23 billion. The model natively supports a 1-million-token context window and delivers top-tier performance in full-stack advanced coding, long-horizon agent collaboration, and cross-modal comprehension, offering global developers exceptional cost-efficiency.

コンテキスト 1.05M
Qwen
qwen3.5-plus20%割引
—
入力: $0.32$0.4 / 100万トークン出力: $1.92$2.4 / 100万トークンキャッシュ読取: $0.032$0.04 / 100万トークン

Qwen3.5-Plus is Alibaba Cloud's flagship MoE large language model, accessible via Model Studio. With 397B total and 17B active parameters, it balances top-tier performance with low cost. It features a 1-million-token context window, excelling in complex reasoning, advanced coding, and agent workflows. Supporting 201 languages, it is built for global enterprise scale.

コンテキスト 1M
Claude
claude-opus-525%割引
—
入力: $3.75$5 / 100万トークン出力: $18.75$25 / 100万トークンキャッシュ読取: $0.375$0.5 / 100万トークン

Claude Opus 5 is Anthropic’s heavy-hitting AI model released in July 2026. It matches flagship performance at half the cost. Built natively for complex, long-horizon tasks, it sets new industry records in coding and reasoning, notably scoring 30.2% on ARC-AGI-3. Featuring bulletproof safety defenses, it powers next-gen developer agents.

コンテキスト 1M
OpenAI
gpt-5.6-luna25%割引
—
入力: $0.15$0.2 / 100万トークン出力: $0.9$1.2 / 100万トークンキャッシュ読取: $0.015$0.02 / 100万トークン

GPT-5.6 Luna is the lightweight, lightning-fast variant optimized for high-volume and cost-sensitive inference workloads. Operating as an agile gateway with the lowest pricing tier, it couples remarkable text-processing speeds with exceptional instruction-following accuracy. It serves as the ideal engine for real-time applications, large-scale content summarization, instant data classification, and high-frequency routing automation.

コンテキスト 1.05M
OpenAI
gpt-5.6-terra25%割引
—
入力: $1.5$2 / 100万トークン出力: $9$12 / 100万トークンキャッシュ読取: $0.15$0.2 / 100万トークン

GPT-5.6 Terra is the perfectly balanced workhorse of the series, optimizing the sweet spot between advanced intelligence and operating cost. Designed as an everyday production engine for businesses and developers, it delivers GPT-5.5-level reasoning capabilities at roughly half the token price. It is highly recommended for scoped code implementations, medium-complexity agentic workflows, content creation, and general-purpose structured data extraction.

コンテキスト 1.05M
OpenAI
gpt-5.6-sol25%割引
—
入力: $3.75$5 / 100万トークン出力: $22.5$30 / 100万トークンキャッシュ読取: $0.375$0.5 / 100万トークン

GPT-5.6 Sol is an advanced large language model tailored for high-efficiency, long-context reasoning. Integrating next-generation RAG with explicit and implicit context caching, it drastically cuts computation costs for multi-turn dialogues while maintaining ultra-low latency. It delivers secure, privacy-focused, and cost-effective dual-track AI power across local and cloud environments, driving enterprise agent automation, deep research, and massive document analysis.

コンテキスト 1.05M
Gemini
gemini-3.1-flash-lite-image25%割引
—
入力: $0.1875$0.25 / 100万トークン出力: $22.5$30 / 100万トークンキャッシュ読取: $0.0188$0.025 / 100万トークン

Gemini 3.1 Flash-Lite Image (also known as Nano Banana 2 Lite) is a highly efficient multimodal AI model developed by Google DeepMind. Tailored for ultra-low latency, it delivers sub-2 second text-to-image generation and multi-turn interactive local editing at a fraction of standard computing costs. Optimized for 1K resolution and 14 native aspect ratios, it seamlessly balances rapid execution with studio-quality precision. This makes it the ideal choice for high-volume creator workflows and real-time enterprise applications.

コンテキスト 65.5K
Claude
claude-sonnet-525%割引
—
入力: $1.5$2 / 100万トークン出力: $7.5$10 / 100万トークンキャッシュ読取: $0.15$0.2 / 100万トークン

Claude 5 Sonnet builds on Sonnet 4.7 with improvements across benchmarks, delivering a more effective and collaborative experience. It's available today. Claude 5 Sonnet launches alongside several powerful new features. Users on claude.ai now have control over the amount of effort Claude dedicates to each task. Claude Code introduces a new 'dynamic workflows' capability that enables it to tackle very large-scale, complex problems. Additionally, fast mode for Claude 5 Sonnet allows the model to work at 2.5× the standard speed, significantly enhancing productivity for time-sensitive workflows.

コンテキスト 1M
Qwen
qwen3.7-max20%割引
—
入力: $2$2.5 / 100万トークン出力: $6$7.5 / 100万トークンキャッシュ読取: $0.4$0.5 / 100万トークン

Qwen3.7-Max is Alibaba's closed-source, flagship AI model engineered for the agentic era. Supporting a 1,000,000-token context window, it excels at complex coding, advanced office automation, and deep scientific reasoning. Designed as a versatile foundation for multi-agent workflows, it delivers elite long-horizon autonomy, sustaining multi-hour autonomous execution across thousands of consecutive steps. It sets SOTA benchmarks in complex logic without requiring framework-specific tuning. [1, 2, 3, 4, 5, 6, 7]

コンテキスト 1M
Qwen
qwen3.7-plus20%割引
—
入力: $0.32$0.4 / 100万トークン出力: $1.28$1.6 / 100万トークンキャッシュ読取: $0.064$0.08 / 100万トークン

Qwen3.7-Plus is Alibaba's proprietary multi-modal Hybrid-Agent model powered by All-field Thinking architecture. Supporting a 1,000,000-token context window and 65,536-token output, it seamlessly processes text, image, and video inputs. By dynamically toggling between deep reasoning and standard conversation, it tackles complex logical tasks with high accuracy. It delivers near-flagship intelligence matching Qwen3.7-Max but operates at a fraction of the cost, making it ideal for deploying scalable, low-latency AI agents.

コンテキスト 1M
Z.ai
glm-5.220%割引
—
入力: $1.12$1.4 / 100万トークン出力: $3.52$4.4 / 100万トークンキャッシュ読取: $0.208$0.26 / 100万トークン

GLM-5.2 is Z.ai's open-weight, flagship 744B-parameter Mixture-of-Experts (MoE) model engineered for long-horizon software engineering and complex agentic workflows. It features a robust 1,000,000-token context window alongside adjustable "Thinking Effort" modes to optimally balance reasoning performance against execution speed. Boasting SOTA benchmarks on Terminal-Bench and SWE-bench Pro, it delivers near-proprietary, frontier-level coding intelligence under an open-source MIT license.

コンテキスト 1M
MoonshotAI
kimi-k2.7-code20%割引
—
入力: $0.76$0.95 / 100万トークン出力: $3.168$3.96 / 100万トークンキャッシュ読取: $0.128$0.16 / 100万トークン

Kimi-K2.7-Code is Moonshot AI's open-weight, 1-trillion parameter Mixture-of-Experts (MoE) model meticulously optimized for software engineering and agentic workflows. Supporting a 256K context window with native vision-text processing, it operates exclusively in long-thinking reasoning mode. It significantly minimizes overthinking—cutting thinking-token usage by 30%—while dramatically boosting multi-turn tool calling, instruction compliance, and success rates in complex, long-horizon coding tasks.

コンテキスト 262.1K
Minimax
MiniMax-M2.520%割引
—
入力: $0.24$0.3 / 100万トークン出力: $0.96$0.3 / 100万トークンキャッシュ読取: $0.024$0.03 / 100万トークン

MiniMax-M2.5 is an open-weight, 230-billion parameter Mixture-of-Experts (MoE) model that activates only 10 billion parameters per token. Designed natively for agentic workflows, it features a 200K context window and achieves an impressive 80.2% on SWE-bench Verified. Driven by large-scale reinforcement learning, it delivers elite performance in full-stack coding, multi-turn tool calling, and office document generation while maintaining superior cost-efficiency.

コンテキスト 204.8K
Minimax
MiniMax-M2.720%割引
—
入力: $0.24$0.3 / 100万トークン出力: $0.96$1.2 / 100万トークンキャッシュ読取: $0.048$0.06 / 100万トークン

MiniMax-M2.7 is an open-weight, 230-billion parameter Mixture-of-Experts (MoE) model featuring a 205K context window, specifically engineered for autonomous productivity. Driven by a breakthrough recursive self-evolution pipeline, it achieved a 30% performance boost through autonomous optimization. It excels at multi-agent collaboration, multi-round complex office suite editing, and system-level software engineering, scoring a SOTA 1495 ELO on GDPval-AA and an impressive 56.22% on SWE-Pro.

コンテキスト 204.8K
Claude
claude-fable-525%割引
—
入力: $7.5$10 / 100万トークン出力: $37.5$50 / 100万トークンキャッシュ読取: $0.75$1 / 100万トークン

Claude Fable 5 is a Mythos-class model from Anthropic designed for autonomous knowledge work, advanced coding, and complex multi-step reasoning. With support for text, image, and file inputs, text output, reasoning capabilities, and a 1M-token context window, it is well suited for long-running tasks that require deep understanding, planning, and sustained execution. Fable 5 performs especially well on ambiguous or highly involved workflows—such as large codebase analysis, end-to-end software development, research synthesis, document review, and asynchronous project work. It can break down complex goals, verify intermediate results, self-correct through iterative reasoning loops, and reduce the need for frequent human check-ins. Built with robust safeguards, Claude Fable 5 offers a strong balance of capability, reliability, and safety for teams handling demanding knowledge and engineering workloads.

コンテキスト 1M
Claude
claude-opus-4-825%割引
—
入力: $3.75$5 / 100万トークン出力: $18.75$25 / 100万トークンキャッシュ読取: $0.375$0.5 / 100万トークン

Claude Opus 4.8 builds on Opus 4.7 with improvements across benchmarks, delivering a more effective and collaborative experience. It's available today. Opus 4.8 launches alongside several powerful new features. Users on claude.ai now have control over the amount of effort Claude dedicates to each task. Claude Code introduces a new ""dynamic workflows"" capability that enables it to tackle very large-scale, complex problems. Additionally, fast mode for Opus 4.8 allows the model to work at 2.5× the standard speed, significantly enhancing productivity for time-sensitive workflows.

コンテキスト 1M
Gemini
gemini-3.5-flash25%割引
—
入力: $1.125$1.5 / 100万トークン出力: $6.75$9 / 100万トークンキャッシュ読取: $0.1125$0.15 / 100万トークン

Gemini 3.5 Flash is a next-generation lightweight, highly cost-effective large language model introduced by Google. It not only retains the "Flash" series' defining features of low latency, high concurrency, and high cost-effectiveness, but also possesses reasoning and multitasking capabilities that rival those of large flagship models.

コンテキスト 1.05M
Gemini
gemini-3-pro-preview25%割引
—
入力: $1.5$2 / 100万トークン出力: $9$12 / 100万トークンキャッシュ読取: $0.15$0.2 / 100万トークン

Gemini 3 Pro Preview is the flagship reasoning model in the Gemini 3 generation, designed for demanding agentic, analytical, and multimodal tasks. It targets scenarios where getting every step right matters more than getting the answer quickly. The model can comprehend vast datasets and complex problems across diverse information sources—including text, audio, images, video, PDFs, and entire code repositories—leveraging its massive 1-million-token context window. Compared to its predecessors, it introduces significant improvements in multi-step function calling, structured planning, reasoning over complex images and long documents, and rigid instruction following. As a reasoning-first model, it supports exposing its internal thought processes (thinking mechanisms), making it the ideal choice for developers building high-stakes AI pipelines, reliable autonomous agents, and deep technical research tools.

コンテキスト 1.05M
Gemini
gemini-3-flash-preview25%割引
—
入力: $0.375$0.5 / 100万トークン出力: $2.25$3 / 100万トークンキャッシュ読取: $0.0375$0.05 / 100万トークン

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool use performance with substantially lower latency than larger Gemini variants, making it well suited for interactive development, long running agent loops, and collaborative coding tasks. Compared to Gemini 2.5 Flash, it provides broad quality improvements across reasoning, multimodal understanding, and reliability. The model supports a 1M token context window and multimodal inputs including text, images, audio, video, and PDFs, with text output. It includes configurable reasoning via thinking levels (minimal, low, medium, high), structured output, tool use, and automatic context caching. Gemini 3 Flash Preview is optimized for users who want strong reasoning and agentic behavior without the cost or latency of full scale frontier models.

コンテキスト 1.05M
Gemini
gemini-3.1-flash-lite-preview25%割引
—
入力: $0.1875$0.25 / 100万トークン出力: $1.125$1.5 / 100万トークンキャッシュ読取: $0.0188$0.025 / 100万トークン

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across key capabilities. Improvements span audio input/ASR, RAG snippet ranking, translation, data extraction, and code completion. Supports full thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs. Priced at half the cost of Gemini 3 Flash.

コンテキスト 1.05M
DeepSeek
deepseek-v4-flash20%割引
—
入力: $0.112$0.14 / 100万トークン出力: $0.224$0.28 / 100万トークンキャッシュ読取: $0.00224$0.0028 / 100万トークン

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance. The model includes hybrid attention for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.

コンテキスト 1.05M
DeepSeek
deepseek-v4-pro20%割引
—
入力: $1.392$1.74 / 100万トークン出力: $2.784$3.48 / 100万トークンキャッシュ読取: $0.0116$0.0145 / 100万トークン

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks. Built on the same architecture as DeepSeek V4 Flash, it introduces a hybrid attention system for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for complex workloads such as full-codebase analysis, multi-step automation, and large-scale information synthesis, where both capability and efficiency are critical.

コンテキスト 1.05M
Gemini
gemini-3.1-flash-image-preview25%割引
—
入力: $0.375$0.5 / 100万トークン出力: $45$60 / 100万トークン

Gemini 3.1 flash image preview (Nano Banana 2) is a lightweight multimodal model from Google DeepMind. It delivers Pro-quality image generation and multi-turn editing at Flash speed. Key capabilities include multi-step reasoning (Thinking), real-time search grounding, 4K resolution, and video-to-image conversion. It supports 14 reference images, ensures character consistency, and runs on Google AI Studio.

コンテキスト 65.5K
Qwen
qwen3.6-plus20%割引
—
入力: $0.4$0.5 / 100万トークン出力: $2.4$3 / 100万トークンキャッシュ読取: $0.04$0.05 / 100万トークン

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers major gains in agentic coding, front-end development, and overall reasoning, with a significantly improved “vibe coding” experience. The model excels at complex tasks such as 3D scenes, games, and repository-level problem solving, achieving a 78.8 score on SWE-bench Verified. It represents a substantial leap in both pure-text and multimodal capabilities, performing at the level of leading state-of-the-art models.

コンテキスト 1M
Z.ai
glm-5.120%割引
—
入力: $0.784$0.98 / 100万トークン出力: $2.464$3.08 / 100万トークンキャッシュ読取: $0.1456$0.182 / 100万トークン

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on a single task for more than 8 hours, autonomously planning, executing, and improving itself throughout the process, ultimately delivering complete, engineering-grade results.

コンテキスト 202.8K
Gemini
gemini-3.1-pro-preview25%割引
—
入力: $1.5$2 / 100万トークン出力: $9$12 / 100万トークンキャッシュ読取: $0.15$0.2 / 100万トークン

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation of the Gemini 3 series, it combines high-precision reasoning across text, image, video, audio, and code with a 1M-token context window. Reasoning Details must be preserved when using multi-turn tool calling. The 3.1 update introduces measurable gains in SWE benchmarks and real-world coding environments, along with stronger autonomous task execution in structured domains such as finance and spreadsheet-based workflows. Designed for advanced development and agentic systems, Gemini 3.1 Pro Preview improves long-horizon stability and tool orchestration while increasing token efficiency. It introduces a new medium thinking level to better balance cost, speed, and performance. The model excels in agentic coding, structured planning, multimodal analysis, and workflow automation, making it well-suited for autonomous agents, financial modeling, spreadsheet automation, and high-context enterprise tasks.

コンテキスト 1.05M
Gemini
gemini-3-pro-image-preview25%割引
—
入力: $1.5$2 / 100万トークン出力: $90$120 / 100万トークンキャッシュ読取: $0.15$0.2 / 100万トークン

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and high-fidelity visual synthesis. The model generates context-rich graphics, from infographics and diagrams to cinematic composites, and can incorporate real-time information via Search grounding. It offers industry-leading text rendering in images (including long passages and multilingual layouts), consistent multi-image blending, and accurate identity preservation across up to five subjects. Nano Banana Pro adds fine-grained creative controls such as localized edits, lighting and focus adjustments, camera transformations, and support for 2K/4K outputs and flexible aspect ratios. It is designed for professional-grade design, product visualization, storyboarding, and complex multi-element compositions while remaining efficient for general image creation workflows.

コンテキスト 65.5K
OpenAI
gpt-image-225%割引
—
入力: $6$8 / 100万トークン出力: $22.5$30 / 100万トークンキャッシュ読取: $1.5$2 / 100万トークン

GPT-5.4 Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and visual generation within the same interaction.

コンテキスト 278.5K
Claude
claude-opus-4-725%割引
—
入力: $3.75$5 / 100万トークン出力: $18.75$25 / 100万トークンキャッシュ読取: $0.375$0.5 / 100万トークン

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on complex, multi-step tasks and more reliable agentic execution across extended workflows. It is especially effective for asynchronous agent pipelines where tasks unfold over time - large codebases, multi-stage debugging, and end-to-end project orchestration. Beyond coding, Opus 4.7 brings improved knowledge work capabilities - from drafting documents and building presentations to analyzing data. It maintains coherence across very long outputs and extended sessions, making it a strong default for tasks that require persistence, judgment, and follow-through.

コンテキスト 1M
OpenAI
gpt-5.525%割引
—
入力: $3.75$5 / 100万トークン出力: $22.5$30 / 100万トークンキャッシュ読取: $0.375$0.5 / 100万トークン

GPT-5.5 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for text and image inputs, enabling high-context reasoning, coding, and multimodal analysis within the same workflow. The model delivers improved performance in coding, document understanding, tool use, and instruction following. It is designed as a strong default for both general-purpose tasks and software engineering, capable of generating production-quality code, synthesizing information across multiple sources, and executing complex multi-step workflows with fewer iterations and greater token efficiency.

コンテキスト 1.05M
Claude
claude-haiku-4-5-2025100125%割引
—
入力: $0.75$1 / 100万トークン出力: $3.75$5 / 100万トークンキャッシュ読取: $0.075$0.1 / 100万トークン

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance across reasoning, coding, and computer-use tasks, Haiku 4.5 brings frontier-level capability to real-time and high-volume applications. It introduces extended thinking to the Haiku line; enabling controllable reasoning depth, summarized or interleaved thought output, and tool-assisted workflows with full support for coding, bash, web search, and computer-use tools. Scoring >73% on SWE-bench Verified, Haiku 4.5 ranks among the world’s best coding models while maintaining exceptional responsiveness for sub-agents, parallelized execution, and scaled deployment.

コンテキスト 200K
Claude
claude-sonnet-4-5-20250929
—
入力: $3 / 100万トークン出力: $15 / 100万トークンキャッシュ読取: $0.3 / 100万トークン

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with improvements across system design, code security, and specification adherence. The model is designed for extended autonomous operation, maintaining task continuity across sessions and providing fact-based progress tracking. Sonnet 4.5 also introduces stronger agentic capabilities, including improved tool orchestration, speculative parallel execution, and more efficient context and memory management. With enhanced context tracking and awareness of token usage across tool calls, it is particularly well-suited for multi-context and long-running workflows. Use cases span software engineering, cybersecurity, financial analysis, research agents, and other domains requiring sustained reasoning and tool use.

コンテキスト 200K
Claude
claude-sonnet-4-625%割引
—
入力: $2.25$3 / 100万トークン出力: $11.25$15 / 100万トークンキャッシュ読取: $0.225$0.3 / 100万トークン

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with memory, polished document creation, and confident computer use for web QA and workflow automation.

コンテキスト 1M
Claude
claude-opus-4-5-20251101
—
入力: $5 / 100万トークン出力: $25 / 100万トークンキャッシュ読取: $0.5 / 100万トークン

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and reasoning benchmarks, and improved robustness to prompt injection. The model is designed to operate efficiently across varied effort levels, enabling developers to trade off speed, depth, and token usage depending on task requirements. It comes with a new parameter to control token efficiency, which can be accessed using the OpenRouter Verbosity parameter with low, medium, or high. Opus 4.5 supports advanced tool use, extended context management, and coordinated multi-agent setups, making it well-suited for autonomous research, debugging, multi-step planning, and spreadsheet/browser manipulation. It delivers substantial gains in structured reasoning, execution reliability, and alignment compared to prior Opus generations, while reducing token overhead and improving performance on long-running tasks.

コンテキスト 200K
Claude
claude-opus-4-625%割引
—
入力: $3.75$5 / 100万トークン出力: $18.75$25 / 100万トークンキャッシュ読取: $0.375$0.5 / 100万トークン

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective for large codebases, complex refactors, and multi-step debugging that unfolds over time. The model shows deeper contextual understanding, stronger problem decomposition, and greater reliability on hard engineering tasks than prior generations. Beyond coding, Opus 4.6 excels at sustained knowledge work. It produces near-production-ready documents, plans, and analyses in a single pass, and maintains coherence across very long outputs and extended sessions. This makes it a strong default for tasks that require persistence, judgment, and follow-through, such as technical design, migration planning, and end-to-end project execution.

コンテキスト 1M
OpenAI
gpt-5.4-mini25%割引
—
入力: $0.5625$0.75 / 100万トークン出力: $3.375$4.5 / 100万トークンキャッシュ読取: $0.0563$0.075 / 100万トークン

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding, and tool use, while reducing latency and cost for large-scale deployments. The model is designed for production environments that require a balance of capability and efficiency, making it well suited for chat applications, coding assistants, and agent workflows that operate at scale. GPT-5.4 mini delivers reliable instruction following, solid multi-step reasoning, and consistent performance across diverse tasks with improved cost efficiency.

コンテキスト 409.6K
OpenAI
gpt-5.4-nano25%割引
—
入力: $0.15$0.2 / 100万トークン出力: $0.9375$1.25 / 100万トークンキャッシュ読取: $0.015$0.02 / 100万トークン

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency use cases such as classification, data extraction, ranking, and sub-agent execution. The model prioritizes responsiveness and efficiency over deep reasoning, making it ideal for pipelines that require fast, reliable outputs at scale. GPT-5.4 nano is well suited for background tasks, real-time systems, and distributed agent architectures where minimizing cost and latency is essential.

コンテキスト 400K
MoonshotAI
kimi-k2.520%割引
—
入力: $0.48$0.6 / 100万トークン出力: $2.4$3 / 100万トークンキャッシュ読取: $0.0813$0.1016 / 100万トークン

Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed visual and text tokens, it delivers strong performance in general reasoning, visual coding, and agentic tool-calling.

コンテキスト 262.1K
OpenAI
gpt-5.425%割引
—
入力: $1.875$2.5 / 100万トークン出力: $11.25$15 / 100万トークンキャッシュ読取: $0.1875$0.25 / 100万トークン

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for text and image inputs, enabling high-context reasoning, coding, and multimodal analysis within the same workflow. The model delivers improved performance in coding, document understanding, tool use, and instruction following. It is designed as a strong default for both general-purpose tasks and software engineering, capable of generating production-quality code, synthesizing information across multiple sources, and executing complex multi-step workflows with fewer iterations and greater token efficiency.

コンテキスト 1.05M