Last updated: 2026-09-30
This is a list of free llm providers and their rate usage limits
Updates from time to time.
免费LLM api平台(包括cn平台)
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
You may also want to read my other posts:
- How to Choose Your LLM For Translation?
- Model & Performance FAQ
- Local LLMs Collection For Translation
- LLM Timeline
Easy guide to deploy online LLM api(luna):
- Sign up/register in certain platform;
- Get this platform's Api Key and Endpoint Address
- Check if there's an available date for the key
- Put it in the softwares that support it.
https://ai.google.dev/gemini-api/docs/rate-limits#free-tier
Gemma3 had been removed from api.
Last updated: 2026-09-23 UTC
RPM: Requests per minute
TPM: Tokens per minute
RPD Requests per day
- No free api for Gemini 3 Pro
https://ai.google.dev/gemini-api/docs/gemini-3?thinking=high#faq
Endpoint: https://generativelanguage.googleapis.com
| Model | Requests/minute (RPM) | Tokens/minute (TPM) | Requests/day (RPD) |
|---|---|---|---|
| Gemini 2.5~3.8 Flash | 5 | 250k | 20 |
| Gemini 2.5 Flash Lite | 10 | 250k | 20 |
| Gemini 3.1/3.5 Flash Lite | 15 | 250k | 500 |
| Gemma 4 26B/31B | 30 | 16k | 14.4K |
Google has set gemini limit to low rate.
★Other solutions:
Google AI Studio to API Adapter
Last updated: 2026-09-23
https://build.nvidia.com/explore/discover
Endpoint: https://integrate.api.nvidia.com
Rate limit: Usually up to 40 Requests per minute(RPM)
Nvidia: Maximum API requests accepted in a given timeframe.
Rate limits may vary by model and traffic from other users may cause throttling.
For dedicated availability, deploy models as a dedicated endpoint with NVIDIA NIM.
- 55 Models list
- OpenAI/Google/Meta/Microsoft/NVIDIA/Mistral AI/DeepSeek/Moonshot AI/Z-ai/Minimax etc.
deepseek-ai/deepseek-v4-flash
deepseek-ai/deepseek-v4-flash-0731
deepseek-ai/deepseek-v4-pro
google/codegemma-7b
google/gemma-7b
meta/llama2-70b
meta/llama-3.1-8b-instruct
meta/llama-3.1-70b-instruct
meta/llama-3.2-1b-instruct
meta/llama-3.2-3b-instruct
meta/llama-3.3-70b-instruct
microsoft/phi-4-mini-instruct
microsoft/phi-4-mini-flash-reasoning
minimaxai/minimax-m2.5
minimaxai/minimax-m2.7
mistralai/mistral-nemotron
mistralai/mixtral-8x7b-instruct
mistralai/mixtral-8x22b-instruct
moonshotai/kimi-k2-instruct
moonshotai/kimi-k2-thinking
moonshotai/kimi-k3
nvidia/gliner-pii
nvidia/llama-3.1-nemoguard-8b-content-safety
nvidia/llama-3.1-nemoguard-8b-topic-control
nvidia/llama-3.1-nemotron-safety-guard-8b-v3
nvidia/llama-3.3-nemotron-super-49b-v1
nvidia/llama-3.3-nemotron-super-49b-v1.5
nvidia/llama-3.1-nemotron-ultra-253b-v1
nvidia/nemotron-3-ultra-550b-a55b
nvidia/nemotron-3.5-lightning-30b-a3b
nvidia/nemoguard-jailbreak-detect
nvidia/nemotron-3-nano-30b-a3b
nvidia/nemotron-3-super-120b-a12b
nvidia/nemotron-content-safety-reasoning-4b
nvidia/nvidia-nemotron-nano-9b-v2
nvidia/riva-translate-4b-instruct-v1.1
nvidia/riva-translate-4b-instruct-v2
nvidia/usdcode
openai/gpt-oss-20b
openai/gpt-oss-120b
poolside/laguna-xs-2-1
qwen/qwen2.5-coder-32b-instruct
qwen/qwen3-next-80b-a3b-instruct
qwen/qwen3-next-80b-a3b-thinking
qwen/qwq-32b
sarvamai/sarvam-m
stepfun-ai/step-3.5-flash
stockmark/stockmark-2-100b-instruct
thinking-machines/inkling
upstage/solar-10.7b-instruct
z-ai/glm4.7
z-ai/glm5.1
z-ai/glm-5.2
z-ai/glm-5.3
z-ai/glm-5.3-flash
Last Check: 2026-09-30
https://ollama.com/cloud
https://ollama.com/search?c=cloud
- OpenAI — GPT-OSS
- MiniMax — M2 / M3
- Moonshot AI — Kimi K2-K3
- NVIDIA — Nemotron 3
- DeepSeek — DeepSeek V4
- Z.ai — GLM 5
- Alibaba Qwen — Qwen 3.5
- Google — Gemma 4
- Mistral AI — Mistral Large 3 / Devstral Small 2
Free model list
- Gemma4: 31b
- gpt-oss 20b/120b
- nemotron-3-nano 30B/super 120B-A12B/ultra 550B-A55B
Session usage reset in 3hours
| Company | Models |
|---|---|
| OpenAI—GPT-OSS | GPT-OSS 20B GPT-OSS 120B |
| MiniMax—M2 / M3 | MiniMax M2.7(229B/A10B) MiniMax M3 (428B/A22B) |
| Moonshot AI—Kimi K2 / K3 | Kimi K2.6(1T/A32B) Kimi K2.7 Code(1T/A32B) Kimi K3(2.8T/A104B) |
| NVIDIA—Nemotron 3 | Nemotron 3 Nano(30B/A3B) Nemotron 3 Super(120B/A12B) Nemotron 3 Ultra(550B/A55B) |
| DeepSeek—DeepSeek V4 | DeepSeek V4 Flash(284B/A13B) DeepSeek V4 Flash 0731(284B/A13B) DeepSeek V4 Pro(1.6T/A49B) DeepSeek V4 Pro 0813(1.6T/A49B) DeepSeek v4.1 Flash(552B) |
| Z.ai—GLM 5 | GLM-5.1(754B/A40B) GLM-5.2(753B/A40B) GLM-5.3-Flash(320B/A18B) |
| Alibaba Qwen—Qwen 3.5 | Qwen3.5-397B-A17B |
| Google—Gemma 4 | Gemma 4 31B |
| Mistral AI—Mistral Large 3 / Devstral Small 2 | Mistral Large 3 675B Devstral Small 2 24B |
| Plan | Price | Included Usage Credits | Concurrent Requests | Model Access | Main Features |
|---|---|---|---|---|---|
| Free | $0 | Starter usage credits | 1 | Starter models by default; add credits to unlock all models | Run models locally; no service fees |
| Pro | $20/month or $200/year | $60/month | 3 | Access to larger Pro models | Everything in Free; multiple models concurrently; Fast Mode (coming soon) |
| Max | $100/month | $300/month | 10 | All Pro access + early access to newest models | Everything in Pro; designed for power users running multiple agents simultaneously |
| Team | $500/month | $1,000/month shared across the team | 10 | Pro-level model access | Unlimited users; centralized billing and administration; priority support; shared projects, skills and instructions (coming soon) |
| Enterprise | Custom pricing | Custom / volume-based | Custom | Configurable model access | Everything in Team; model access controls; user/API-key cost budgets; private Slack support channel; custom security questionnaires |
- Pro annual billing: $200/year, equivalent to about $16.67/month.
- Local model usage is unlimited on your own hardware regardless of plan.
- All plans, including Free, can purchase additional usage credits.
- Purchased/additional credits can unlock all cloud models even on the Free plan.
- Included monthly credits are used first, followed by the extra usage balance.
- Unused included monthly credits do not roll over.
- Pro, Max and Team included usage resets monthly on the subscription anniversary date.
- Free usage resets monthly from the account signup date.
- Requests exceeding the concurrency limit are queued until a slot becomes available; requests may be rejected if the queue is full.
- Team credits and additional usage balances are shared across the organization.
- Ollama states that prompt and response data is not logged or used for training.
Prices are in USD per 1 million tokens. "Input + Output (1:1)" assumes 1M input tokens + 1M output tokens.
| Model | Input | Output | Input + Output (1:1) | Note |
|---|---|---|---|---|
| deepseek-v4.1-flash | $0.15 | $0.60 | $0.75 | unpeak |
| deepseek-v4-flash | $0.44 | $1.32 | $1.76 | peak |
| deepseek-v4-pro | $1.32 | $3.96 | $5.28 | |
| gemma4 | $0.14 | $0.40 | $0.54 | |
| glm-5.3 | $1.40 | $4.40 | $5.80 | |
| glm-5.3-flash | $0.15 | $0.50 | $0.65 | |
| glm-5.2 | $1.40 | $4.40 | $5.80 | |
| glm-5.1 | $1.00 | $3.20 | $4.20 | |
| gpt-oss:120b | $0.15 | $0.60 | $0.75 | |
| gpt-oss:20b | $0.07 | $0.30 | $0.37 | |
| kimi-k3 | $3.00 | $15.00 | $18.00 | |
| kimi-k2.7-code | $0.95 | $4.00 | $4.95 | |
| kimi-k2.6 | $0.95 | $4.00 | $4.95 | |
| minimax-m3 | $0.60 | $2.40 | $3.00 | |
| minimax-m2.7 | $0.30 | $1.20 | $1.50 | |
| mistral-large-3 | $0.50 | $1.50 | $2.00 | |
| nemotron-3-nano | $0.06 | $0.24 | $0.30 | |
| nemotron-3-super | $0.015 | $0.60 | $0.615 | |
| nemotron-3-ultra | $0.10 | $3.00 | $3.10 | |
| qwen3.5:397b | $0.60 | $3.60 | $4.20 |
Last Check: 2026-09-23
*Current only GPT-OSS 20b/120b & Qwen3.8-27b
https://console.groq.com/docs/rate-limits
Endpoint: https://api.groq.com/openai
| Model | Request/Minute | Request/Day | Token/Minute | Token/Day |
|---|---|---|---|---|
| groq/compound | 30 | 250 | 70K | - |
| groq/compound-mini | 30 | 250 | 70K | - |
| openai/gpt-oss-120b | 30 | 1K | 8K | 200K |
| openai/gpt-oss-20b | 30 | 1K | 8K | 200K |
| qwen/qwen3.8-27b | 30 | 1K | 8K | 2M |
> Last Check: 2026-09-23
Unusable rate limit 5$ per account
https://inference-docs.cerebras.ai/support/rate-limits
Endpoint: https://api.cerebras.ai
| Model | Requests/Minute | Tokens/Minute | Tokens/Hour | Tokens/Day |
|---|---|---|---|---|
| gpt-oss-120b | 5 | 30k | 1M | 1M |
| gemma-4-31b | 5 | 30k | 1M | 1M |
Last Check: 2026-09-30
https://openrouter.ai/models?q=free
https://openrouter.ai/pricing
Endpoint: https://openrouter.ai/api
Models: Based on what OpenRouter (the platform) provide as Free
- Free usage limits: If you’re using a free model variant (with an ID ending in
:free/(free))
you can make up to20requests/minute.
- If you have purchased less than
$10 credits, you’re limited to50freemodel requests/Day. - If you purchase at least
$10 credits, your daily limit is increased to1000freemodel requests/Day.
- If you have purchased less than
| Model variant (ID) | Credits purchased | Rate limit (requests/min) | Daily limit (requests/day) |
|---|---|---|---|
| *:free | < $10 credits | 20 | 50 |
| *:free | ≥ $10 credits | 20 | 1000 |
Last web updated: 2026-08-18
Last Check: 2026-09-23
https://developers.cloudflare.com/workers-ai/platform/pricing/#llm-model-pricing
https://developers.cloudflare.com/workers/platform/pricing/
- Workers Free
10,000 Neurons/0.11$ per day
"Neurons are our way of measuring AI outputs across different models, representing the GPU compute needed to perform your request. Our serverless model allows you to pay only for what you use without having to worry about renting, managing, or scaling GPUs."
Models list
- Llama
- Llama2-7b
- Llama3.1-8b/70b
- Llama3.2-1b/3b/11b(vision)
- Llama4-scout-17b-16e
- Qwen
- qwq-32b
- qwen2.5-coder-32b
- qwen3-30b-a3b
- qwen3.8-27b
- Mistral
- Mistral-7b-intruct-v0.1
- Mistral-small-3.1b-24b
- deepseek-r1-distill-qwen-32b
- deepseek v4 flash/pro
- Gemma
- gemma-3-12b
- gemma-4-26B-A4B
- gemma-sea-lion-v4-27b-it
- granite-4.0-h-micro
- glm-4.7-flash
- glm-5.2
- nemotron-3-120b-a12b
- kimi-k2.5/k2.6/k2.7-code
Full list with token cost
Last official page update: August 18, 2026
Unit: USD per 1M tokens
Ranking method:
Combined = Input + OutputFree allocation: 10,000 Neurons per day
Free Tokens/day assumes an Input : Output token ratio of
1 : 1.3.Free-token estimates are calculated directly from Cloudflare's Neurons-per-token rates rather than rounded USD prices.
Cached input is excluded from the free-token calculation.
Identical models from the same family are merged only when their pricing and relevant notes are identical.
| Rank | Provider | Model | Input | Output | Combined | 10K Neurons ≈ Tokens/day (I:O = 1:1.3) | Notes |
|---|---|---|---|---|---|---|---|
| 1 | IBM | Granite 4.0 H Micro | $0.017 | $0.112 | $0.129 | ~1.560M | — |
| 2 | Meta | Llama 3.2 1B Instruct | $0.027 | $0.201 | $0.228 | ~878K | — |
| 3 | Mistral | Mistral 7B Instruct v0.1 | $0.110 | $0.190 | $0.300 | ~708K | — |
| 4 | Meta | Llama 3.2 3B Instruct | $0.051 | $0.335 | $0.386 | ~520K | — |
| 5 | Qwen | Qwen3 30B-A3B FP8 | $0.051 | $0.335 | $0.386 | ~520K | — |
| 6 | Meta | Llama 3/3.1 8B Instruct AWQ | $0.123 | $0.266 | $0.389 | ~539K | Identical pricing; merged |
| 7 | Gemma 4 26B-A4B IT | $0.100 | $0.300 | $0.400 | ~516K | — | |
| 8 | Meta | Llama 3.1 8B Instruct FP8 Fast | $0.045 | $0.384 | $0.429 | ~465K | — |
| 9 | Meta | Llama 3.1 8B Instruct FP8 | $0.152 | $0.287 | $0.439 | ~482K | — |
| 10 | Z.AI | GLM-4.7 Flash | $0.060 | $0.400 | $0.460 | ~435K | — |
| 11 | OpenAI | GPT-OSS 20B | $0.200 | $0.300 | $0.500 | ~429K | — |
| 12 | Meta | Llama Guard 3 8B | $0.484 | $0.030 | $0.514 | ~484K | Safety classifier; not a general-purpose chat LLM |
| 13 | Meta | Llama 3.2 11B Vision Instruct | $0.049 | $0.676 | $0.725 | ~273K | Vision-capable |
| 14 | Gemma 3 12B IT | $0.345 | $0.556 | $0.901 | ~237K | — | |
| 15 | AI Singapore | Gemma SEA-LION v4 27B IT | $0.351 | $0.555 | $0.906 | ~236K | — |
| 16 | Mistral | Mistral Small 3.1 24B Instruct | $0.351 | $0.555 | $0.906 | ~236K | — |
| 17 | OpenAI | GPT-OSS 120B | $0.350 | $0.750 | $1.100 | ~191K | — |
| 18 | Meta | Llama 3/3.1 8B Instruct | $0.282 | $0.827 | $1.109 | ~187K | Identical pricing; merged |
| 19 | Meta | Llama 4 Scout 17B-16E Instruct | $0.270 | $0.850 | $1.120 | ~184K | — |
| 20 | Qwen | QwQ 32B | $0.660 | $1.000 | $1.660 | ~129K | — |
| 21 | Qwen | Qwen2.5 Coder 32B Instruct | $0.660 | $1.000 | $1.660 | ~129K | — |
| 22 | DeepSeek | DeepSeek V4 Flash 0731 | $0.440 | $1.320 | $1.760 | ~117K | Cached input $0.014; paid billing method required |
| 23 | NVIDIA | Nemotron 3 120B-A12B | $0.500 | $1.500 | $2.000 | ~103K | — |
| 24 | Meta | Llama 3.1/3.3 70B Instruct FP8 Fast | $0.293 | $2.253 | $2.546 | ~78.5K | Identical pricing; merged |
| 25 | Moonshot AI | Kimi K2.5 | $0.600 | $3.000 | $3.600 | ~56.2K | Cached input $0.100 |
| 26 | Qwen | Qwen3.8 27B | $0.450 | $3.200 | $3.650 | ~54.9K | — |
| 27 | Moonshot AI | Kimi K2.6 | $0.950 | $4.000 | $4.950 | ~41.1K | Cached input $0.160; paid billing method required |
| 28 | Moonshot AI | Kimi K2.7 Code | $0.950 | $4.000 | $4.950 | ~41.1K | Cached input $0.190; paid billing method required |
| 29 | DeepSeek | DeepSeek V4 Pro 0813 | $1.320 | $3.960 | $5.280 | ~39.1K | Cached input $0.044; paid billing method required |
| 30 | DeepSeek | DeepSeek R1 Distill Qwen 32B | $0.497 | $4.881 | $5.378 | ~37.0K | — |
| 31 | Z.AI | GLM-5.2 | $1.400 | $4.400 | $5.800 | ~35.5K | Cached input $0.260; paid billing method required |
| 32 | Meta | Llama 2 7B Chat FP16 | $0.556 | $6.667 | $7.223 | ~27.4K | — |
Last web updated: 2026-06-09
Last Check: 2026-09-23
https://docs.cohere.com/docs/rate-limits
Endpoint: https://api.cohere.ai/compatibility
| Endpoint | Trial rate limit (requests/min) | Trial monthly cap (calls/month) |
|---|---|---|
| Chat API | 20/min (per model) | 1000/month |
- Chat model includes
- Command A+
- Command A Reasoning
- Command A Translate
- Command A Vision
- Command A
- Command R+
- Command R
- Command R7B
- North Mini Code
All endpoints are limited to 1,000 calls per month with a trial key
Last Check: 2026-08-29
https://docs.z.ai/guides/overview/pricing
Endpoint: https://api.z.ai/api/paas/v4
Free Models :
- GLM-4.5-Flash
- GLM-4.7-Flash
- GLM-4.6V-Flash
No offical rate usage limits
https://docs.github.com/en/github-models/use-github-models/prototyping-with-ai-models#rate-limits
GitHub Models has been retired.
As of July 30, 2026, GitHub Models has been fully retired. The playground, model catalog, inference API, and bring your own key (BYOK) are no longer available to any customer.
GitHub Models was a separate service from GitHub Copilot and is unrelated to GitHub Copilot services.
https://docs.mistral.ai/deployment/laplateforme/tier/
(need to login)
Endpoint: https://api.mistral.ai
- From community&reports; No offcial list
| Plan / Tier | Requests/second (RPS) | Tokens/minute (TPM) | Tokens/month |
|---|---|---|---|
| Mistral API Free | 1/sec | 500k | 1 billion |
Last Check: 2026-09-23
https://docs.sambanova.ai/docs/en/models/rate-limits#free-tier
| Model | Requests/Minute | Requests/Day | Tokens/Day |
|---|---|---|---|
| DeepSeek-V3.1 | 20 | 20 | 200k |
| Meta-Llama-3.3-70B-Instruct | 20 | 20 | 200k |
| gpt-oss-120b | 20 | 20 | 200k |
| DeepSeek-V3.2 | 20 | 20 | 200k |
| gemma-4-31B-it | 20 | 20 | 200k |
Last Check: 2026-09-23
https://www.aionlabs.ai/docs/rate-limits/
https://www.aionlabs.ai/docs/models/
Endpoint: https://api.aionlabs.ai/v1
| Model | Requests/Minute | Tokens/Day |
|---|---|---|
| Aion-2.0 (DeepSeek V3.2 finetune) | 15 | 20k |
| Aion-2.5 (Improve of Aion-2.0) | 15 | 20k |
| Aion-3.0 (GLM) | 15 | 20k |
| Aion-3.0 Mini (Deepseek) | 15 | 20k |
| Aion-RP 1.0 (8B) (llama-3.1-8b fineune/openweight) | 15 | 20k |
Free Api for A.X 4.0 (7B/72B, based on Qwen2.5)
Korean⇌EN
https://github.com/SKT-AI/A.X-4.0/blob/main/apis/README.md
Last Check: 2026-09-23
https://www.ibm.com/products/watsonx-ai/pricing
| Plan / Tier | Requests/second (RPS) | Tokens/month |
|---|---|---|
| watsonx.ai Free tier (Foundation Models) |
2/sec | 300k |
https://www.scaleway.com/en/docs/generative-apis/faq/#how-does-the-free-tier-work
> Last updated: 2025-11-17
https://www.together.ai/models
- Llama3-70b
https://www.together.ai/models/llama-3-3-70b-free
https://www.together.ai/models/deepseek-r1-distilled-llama-70b-free
NO FREE TIER Now
https://support.together.ai/articles/1862638756-changes-to-free-tier-and-billing-july-2025
https://modelscope.cn/docs/model-service/API-Inference/limits
- 需要须首先绑定阿里云账号。对应云账号需已通过实名认证后,才可正常使用API-Inference
- 每位魔搭注册用户,当前每天允许进行总数为2000次的API-Inference调用,其中每单个模型上限不超过500次,具体每个模型的限制可能随时动态调整。
- 在每个模型每天不超过500次调用的基础上,平台可能对于部分模型再进行单独的限制
- 例如,deepseek-ai/DeepSeek-R1-0528,deepseek-ai/DeepSeek-V3.2-Exp等规格较大模型,当前限制单模型每天100次调用额度。其他模型的API调用,也可能会有类似的限制并进行动态调整
部分小模型免费;新cn手机用户注册送20M token
https://siliconflow.cn/
(Hunyuan-lite free; 1 M free token for other models/per account)
- 首次开通腾讯混元大模型服务后,混元生文将发放一定量级的免费调用额度(100M tokens)
- 资源包有效期为1年,自开通服务之日起1年内若免费资源包次数未使用完,则过期作废
- Hunyuan-lite 为免费模型
https://cloud.tencent.com/document/product/1729/97731
Tongyi Qwen free (no point cost; 100 time/day)
https://www.volcengine.com/docs/84458/1585102
https://www.volcengine.com/docs/84458/1585097
- 个人免费版为
500资源点/天 - 目前(2025.11.03),扣子模型中仅豆包模型、DeepSeek 模型和 Kimi-K2 模型收费。使用 Kimi(8K)等其他扣子提供的模型暂不收取费用,但每日调用次数有一定限制
(100次/天)
模型表
| 模型名称 | 条件 (千tokens) |
输入单价 (资源点/ktok) |
输出单价 (资源点/ktok) |
合计单价 (资源点/ktok) |
500资源点可用总tokens (ktok) |
|---|---|---|---|---|---|
| 豆包·1.6·视觉理解·250815(Doubao-Seed-1.6-vision) | [0,32] | 0.8 | 8 | 8.8 | 56.82 |
| 豆包·1.6·视觉理解·250815(Doubao-Seed-1.6-vision) | (32,128] | 1.2 | 16 | 17.2 | 29.07 |
| 豆包·1.6·深度思考 / 豆包·1.6·深度思考·250715(Doubao-Seed-1.6-thinking) | [0,32] | 0.8 | 8 | 8.8 | 56.82 |
| 豆包·1.6·深度思考 / 豆包·1.6·深度思考·250715(Doubao-Seed-1.6-thinking) | (32,128] | 1.2 | 16 | 17.2 | 29.07 |
| 豆包·1.6·自动深度思考(Doubao-Seed-1.6) | [0,32]&[0,0.2] | 0.8 | 2 | 2.8 | 178.57 |
| 豆包·1.6·自动深度思考(Doubao-Seed-1.6) | [0,32]&(0.2,+∞] | 0.8 | 8 | 8.8 | 56.82 |
| 豆包·1.6·自动深度思考(Doubao-Seed-1.6) | (32,128] | 1.2 | 16 | 17.2 | 29.07 |
| 豆包·1.6·极致速度 / 豆包·1.6·极致速度·250828 / 豆包·1.6·极致速度·250715(Doubao-seed-1.6-flash) | [0,32] | 0.15 | 1.5 | 1.65 | 303.03 |
| 豆包·1.6·极致速度 / 豆包·1.6·极致速度·250828 / 豆包·1.6·极致速度·250715(Doubao-seed-1.6-flash) | (32,128] | 0.3 | 3 | 3.3 | 151.52 |
| 豆包·1.5·Pro·视觉深度思考(Doubao-1.5-thinking-vision-pro) | 默认 | 3 | 9 | 12 | 41.67 |
| 豆包·1.5·Pro·视觉理解-250328(Doubao-1.5-vision-pro) | 默认 | 3 | 9 | 12 | 41.67 |
| 豆包·1.5·Pro·视觉理解(Doubao-1.5-vision-pro-32k) | 默认 | 3 | 9 | 12 | 41.67 |
| 豆包·1.5·Pro·深度思考·128K / 豆包·1.5·Pro·视觉推理·128K(Doubao-1.5-thinking-pro) | 默认 | 4 | 16 | 20 | 25.00 |
| 豆包·1.5·Pro·角色扮演 / 豆包·1.5·Pro·角色扮演·250715(Doubao-1.5-pro-32k) | 默认 | 0.8 | 2 | 2.8 | 178.57 |
| 豆包·1.5·Pro·32k(Doubao-1.5-pro-32k) | 默认 | 0.8 | 2 | 2.8 | 178.57 |
| 豆包·1.5·Pro·256k(Doubao-1.5-pro-256k) | 默认 | 5 | 9 | 14 | 35.71 |
| 豆包·1.5·Lite·32k(Doubao-1.5-lite-32k) | 默认 | 0.3 | 0.6 | 0.9 | 555.56 |
| 豆包·通用模型·Lite(Doubao-lite-32k) | 默认 | 0.3 | 0.6 | 0.9 | 555.56 |
| 豆包·工具调用 / 豆包·角色扮演·Pro(Doubao-pro-32k) | 默认 | 0.8 | 2 | 2.8 | 178.57 |
| DeepSeek-V3.1 | 默认 | 4 | 12 | 16 | 31.25 |
| DeepSeek-V3 / DeepSeek-V3 工具调用 / DeepSeek-V3-0324 | 默认 | 2 | 8 | 10 | 50.00 |
| DeepSeek-R1 / DeepSeek-R1 工具调用 / DeepSeek-R1-250528 | 默认 | 4 | 16 | 20 | 25.00 |
| Kimi-K2 | 默认 | 4 | 16 | 20 | 25.00 |
- 使用豆包 1.6 模型时,输入 token 单价和输出 token 单价均由输入长度决定。例如调用豆包·1.6·自动深度思考模型时,当 1 个请求的输入长度为 200 千tokens,输出长度为 14 千token 时,满足条件输入长度 (128, 256],将采用计费项 **Doubao-Seed-1.6-256k(输入)**和 Doubao-Seed-1.6-256k(输出)。
- Doubao-Seedance-1.0-lite、Doubao-Seedance-1.0-pro 模型各自为每个扣子账号(主账号+子账号)提供累计 100 万tokens 免费额度。免费额度耗尽后如需继续使用,会从账号中扣减资源点。
https://platform.iflow.cn/docs
阿里云服务器
https://www.streamlake.com/document/WANQING/mdsor5767ob7s796sp6
Spark-lite free
- 首次开通后,免费包(个人)有200k免费额度(所有模型),有效期为一年
https://www.xfyun.cn/doc/spark/HTTP调用文档.html
https://xinghuo.xfyun.cn/sparkapi?scr=price
Unit: USD per 1M tokens
Primary ranking:
Combined = Input + OutputCombined represents the cost of 1M uncached input tokens + 1M output tokens.
Tie-breaker: when Combined prices are identical, the model with the lower directly comparable Cached Input price is ranked first.
Standard real-time API pricing is used.
Temporary promotional prices and normal/list prices are shown separately when the promotion is still active.
Expired promotional prices are not included in the current ranking.
Batch, Flex, Priority/Fast and other alternative processing modes are excluded.
For context-tiered models, the lowest/default context tier is used for ranking.
Same-family models are merged only when their main pricing and relevant pricing conditions are sufficiently equivalent.
Qwen prices use Alibaba Cloud Model Studio Singapore / International pricing.
| Rank | Provider | Model | Input | Cached Input | Output | Combined | Notes |
|---|---|---|---|---|---|---|---|
| 1 | Qwen | Qwen3.7 Flash | $0.030 | — | $0.130 | $0.160 | ≤32K; 32K–256K: $0.10/$0.40; 256K–1M: $0.20/$0.80 |
| 2 | Meta / OpenRouter | Muse Spark 1.2/1.3 Contributor | $0.10 | $0.002 | $0.20 | $0.30 | 1M context; prompts/outputs may be used to improve Meta products |
| 3 | Qwen | Qwen3.5 Flash | $0.10 | — | $0.40 | $0.50 | International |
| 4 | OpenAI | GPT-6 Luna | $0.10 | $0.01 | $0.50 | $0.60 | 1.05M context; >272K: $0.20/$0.75 |
| 5 | Qwen | Qwen3.8 Flash | $0.15 | — | $0.47 | $0.62 | International; 1M context |
| 6 | GLM / Z.AI | GLM-5.3-Flash | $0.15 | $0.03 | $0.50 | $0.65 | Standard price |
| 7 | DeepSeek | DeepSeek V4.1 Flash — Off-peak | $0.15 | $0.003 | $0.60 | $0.75 | 1M context; off-peak pricing |
| 8 | Mistral | Mistral Small 4 | $0.15 | $0.015 | $0.60 | $0.75 | 256K context |
| 9 | OpenAI | GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | $1.40 | >272K: $0.40/$1.80 |
| 10 | Meta / OpenRouter | Muse Glimmer 30B | $0.30 | $0.04 | $1.10 | $1.40 | 131K context |
| 11 | OpenAI | GPT-5.4 nano | $0.20 | $0.02 | $1.25 | $1.45 | — |
| 12 | DeepSeek | DeepSeek V4.1 Flash — Peak | $0.30 | $0.006 | $1.20 | $1.50 | 1M context; peak pricing |
| 13 | Qwen | Qwen3.7 Plus — Promo | $0.32 | — | $1.28 | $1.60 | International alias; 20% off; list $0.40/$1.60 |
| 14 | Gemini 3.1 Flash-Lite | $0.25 | $0.025 | $1.50 | $1.75 | — | |
| 15 | Qwen | Qwen3.6 Flash | $0.25 | — | $1.50 | $1.75 | ≤256K; International |
| 16 | Qwen | Qwen3.7 Plus — List Price | $0.40 | — | $1.60 | $2.00 | Normal International price |
| 17 | DeepSeek | DeepSeek V4 Pro — Off-peak | $0.66 | $0.022 | $1.98 | $2.64 | Off-peak pricing |
| 18 | Gemini 3.5 Flash-Lite | $0.30 | $0.03 | $2.50 | $2.80 | — | |
| 19 | Qwen | Qwen3.5 Plus | $0.40 | — | $2.40 | $2.80 | ≤256K; International |
| 20 | xAI | Grok Build 0.1 | $1.00 | $0.20 | $2.00 | $3.00 | <200K; ≥200K: $2/$4 |
| 21 | Qwen | Qwen3.6 Plus | $0.50 | — | $3.00 | $3.50 | ≤256K; International |
| 22 | xAI | Grok 4.20/4.3 | $1.25 | $0.20 | $2.50 | $3.75 | <200K; ≥200K: $2.50/$5 |
| 23 | GLM / Z.AI | GLM-5 | $1.00 | $0.20 | $3.20 | $4.20 | — |
| 24 | Gemini 3.6/3.7/3.8 Flash — Promo | $0.75 | $0.075 | $3.75 | $4.50 | Introductory price through Dec 31, 2026 | |
| 25 | Kimi | Kimi K2.6 | $0.95 | $0.16 | $4.00 | $4.95 | ~262K context |
| 26 | Kimi | Kimi K2.7 Code | $0.95 | $0.19 | $4.00 | $4.95 | ~262K context |
| 27 | GLM / Z.AI | GLM-5-Turbo | $1.20 | $0.24 | $4.00 | $5.20 | 200K context |
| 28 | OpenAI | GPT-5.4 mini | $0.75 | $0.075 | $4.50 | $5.25 | — |
| 29 | DeepSeek | DeepSeek V4 Pro — Peak | $1.32 | $0.044 | $3.96 | $5.28 | Peak pricing |
| 30 | Meta / OpenRouter | Muse Spark 1.1/1.2/1.3 | $1.25 | $0.15 | $4.25 | $5.50 | 1M context; standard data policy |
| 31 | GLM / Z.AI | GLM-5.1 | $1.40 | $0.26 | $4.40 | $5.80 | 200K context |
| 32 | GLM / Z.AI | GLM-5.2/5.3 | $1.40 | $0.26 | $4.40 | $5.80 | 1M context |
| 33 | xAI | Grok 4.5/4.6 | $2.00 | $0.30 / $0.50 | $6.00 | $8.00 | <200K; ≥200K: $4/$12 |
| 34 | xAI | Grok 4.7 | $2.00 | $0.50 | $6.00 | $8.00 | 500K context; ≥200K: $4/$12 |
| 35 | Qwen | Qwen3.8 Max | $2.00 | — | $6.00 | $8.00 | International; 1M context |
| 36 | Gemini 3.6/3.7/3.8 Flash — List Price | $1.50 | $0.15 | $7.50 | $9.00 | Standard price from Jan 1, 2027 | |
| 37 | Mistral | Mistral Medium 3.5 | $1.50 | $0.15 | $7.50 | $9.00 | 256K context |
| 38 | Kimi | Kimi K2.7 Code Highspeed | $1.90 | $0.38 | $8.00 | $9.90 | ~262K context |
| 39 | Qwen | Qwen3.7 Max | $2.50 | — | $7.50 | $10.00 | International; 1M context |
| 40 | Gemini 3.5 Flash | $1.50 | $0.15 | $9.00 | $10.50 | Thinking tokens billed as output | |
| 41 | OpenAI | GPT-6.1 Sol | $2.00 | $0.10 | $10.00 | $12.00 | New; lower cache price than GPT-6 Sol |
| 42 | OpenAI | GPT-6 Sol | $2.00 | $0.20 | $10.00 | $12.00 | 1.05M context; >272K: $4/$15 |
| 43 | Anthropic | Claude Sonnet 5/5.5 | $2.00 | $0.20 | $10.00 | $12.00 | Same current token pricing |
| 44 | OpenAI | GPT-5.6 Terra | $2.00 | $0.20 | $12.00 | $14.00 | >272K: $4/$18 |
| 45 | Gemini 3.1 Pro Preview | $2.00 | $0.20 | $12.00 | $14.00 | ≤200K; >200K: $4/$18 | |
| 46 | OpenAI | GPT-5.4 | $2.50 | $0.25 | $15.00 | $17.50 | >272K long-context surcharge |
| 47 | Anthropic | Claude Sonnet 4.6 | $3.00 | $0.30 | $15.00 | $18.00 | — |
| 48 | Kimi | Kimi K3 | $3.00 | $0.30 | $15.00 | $18.00 | 1M context |
| 49 | Anthropic | Claude Opus 5.5 | $4.00 | $0.20 | $20.00 | $24.00 | 1M context |
| 50 | OpenAI | GPT-5.6 Sol — Promo | $4.00 | $0.40 | $20.00 | $24.00 | Promo available at least through Nov 21, 2026; list $5/$30 |
| 51 | Anthropic | Claude Opus 4.7/4.8/5 | $5.00 | $0.50 | $25.00 | $30.00 | — |
| 52 | OpenAI | GPT-5.5 | $5.00 | $0.50 | $30.00 | $35.00 | — |
| 53 | OpenAI | GPT-5.6 Sol — List Price | $5.00 | $0.50 | $30.00 | $35.00 | Original/list price; currently $4/$20 promo |
| 54 | Anthropic | Claude Fable 5.1 | $10.00 | $0.25 | $50.00 | $60.00 | 1M context; 0.025× cache-read multiplier |
| 55 | Anthropic | Claude Fable 5 | $10.00 | $1.00 | $50.00 | $60.00 | — |
| 56 | OpenAI | GPT-6 Astra | $10.00 | $1.00 | $50.00 | $60.00 | 1.05M context; >272K: $20/$75 |
| 57 | OpenAI | GPT-5.4/5.5 Pro | $30.00 | — | $180.00 | $210.00 | Maximum-compute tier |
DeepSeek introduced Peak / Off-peak pricing on August 16, 2026.
Peak hours:
- 01:00–04:00 UTC
- 06:00–10:00 UTC
All other hours are Off-peak.
DeepSeek V4 Flash:
- Off-peak: $0.22 Input / $0.66 Output
- Peak: $0.44 Input / $1.32 Output
- Cached Input: $0.007 Off-peak / $0.014 Peak
DeepSeek V4 Pro:
- Off-peak: $0.66 Input / $1.98 Output
- Peak: $1.32 Input / $3.96 Output
- Cached Input: $0.022 Off-peak / $0.044 Peak











