qwen3.8-max — 價格、上下文與效能
qwen3.8-max 是 Alibaba 的模型,上下文 1M,支援 Reasoning, Tools, Files, Vision。輸入 $2.00,輸出 $6.00,皆為每百萬 tokens。
總覽
| 廠商 | Alibaba |
|---|---|
| 上下文 | 1M |
| 能力 | Reasoning, Tools, Files, Vision |
| API 格式 | openai, openai-response, openai-response-compact, anthropic, gemini, openai-alpha-search |
| 計費方式 | p * 2 + cr * 0.25 + cc * 2.5 + c * 6 |
2.4-trillion-parameter MoE flagship for coding, professional work, multimodal understanding, and long-horizon agentic workflows價格
| 方案 | 倍率 | 輸入 | 輸出 | 快取輸入 | 說明 |
|---|---|---|---|---|---|
| 基準價 | 1× | $2.00 | $6.00 | $0.25 | |
| default | 0.6× | $1.20 | $3.60 | $0.15 | Standard — standard reasoning, moderate stability. |
| Plus | 1× | $2.00 | $6.00 | $0.25 | High — strong reasoning, stable. |
| Pro | 2× | $4.00 | $12.00 | $0.50 | Enterprise — full reasoning, high stability. |
| Max | 4× | $8.00 | $24.00 | $1.00 | Premium — max reasoning & stability. |
僅列出該模型已開放的套餐。
效能實測
| API 端點 | https://api.airai.cc/v1 |
|---|---|
| OpenAI 相容 | OpenAI-compatible |
常見問題
qwen3.8-max 的價格是多少?
輸入 $2.00,輸出 $6.00,皆為每百萬 tokens。
qwen3.8-max 的上下文視窗有多大?
qwen3.8-max 的上下文視窗為 1M,即單次請求中輸入與輸出合計的 token 上限。
如何在程式碼中呼叫 qwen3.8-max?
把 base_url 指向 https://api.airai.cc/v1,model 填 qwen3.8-max 即可,介面與 OpenAI 完全相容。
qwen3.8-max 為什麼會回傳 429?
429 是限流回應。降低並發、加上退避重試,或切換到配額更高的方案即可。
模型比較
| Model | 廠商 | 輸入 | 輸出 | 上下文 |
|---|---|---|---|---|
| qwen3.8-max | Alibaba | $2.00 | $6.00 | 1M |
| qwen3.6-plus | Alibaba | $0.50 | $3.00 | 1M |
| qwen3.6-35b-a3b | Alibaba | $0.25 | $1.49 | 262.1K |
| qwen3-vl-plus | Alibaba | $0.20 | $4.80 | 262.1K |
| qwen3.7-plus | Alibaba | $0.40 | $1.60 | 1M |
| wan2.7-image-pro | Alibaba | - | — | 8.2K |
| qvq-max | Alibaba | $1.20 | $4.80 | 131.1K |
| qwen-vl-max | Alibaba | $0.80 | $3.20 | 131.1K |
| qwen3.5-27b | Alibaba | $0.30 | $2.40 | 262.1K |
| qwen-flash | Alibaba | $0.05 | $0.40 | 1M |
| qwq-plus | Alibaba | $0.80 | $2.40 | 131.1K |
| qwen3.5-397b-a17b | Alibaba | $0.60 | $3.60 | 262.1K |
| qwen3-max | Alibaba | $1.20 | $6.00 | 262.1K |
錯誤排查
429 是限流回應。降低並發、加上退避重試,或切換到配額更高的方案即可。
常見原因
- 並發請求過多
- 目前方案配額已用完
- 失敗後立即重試沒有退避
解決辦法
- 限制並發並以指數退避重試
- 切換到配額更高的方案
- 對重複請求做快取
如何接入
curl https://api.airai.cc/v1/chat/completions \
-H "Authorization: Bearer <YOUR_API_KEY>" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-max",
"messages": [{"role": "user", "content": "Hello"}]
}'from openai import OpenAI
client = OpenAI(
api_key="<YOUR_API_KEY>",
base_url="https://api.airai.cc/v1"
)
response = client.chat.completions.create(
model="qwen3.8-max",
messages=[{"role": "user", "content": "Hello"}]
)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
apiKey: "<YOUR_API_KEY>",
baseURL: "https://api.airai.cc/v1"
});
const response = await client.chat.completions.create({
model: "qwen3.8-max",
messages: [{ role: "user", content: "Hello" }]
});
console.log(response.choices[0].message.content);const OpenAI = require("openai");
const client = new OpenAI({
apiKey: "<YOUR_API_KEY>",
baseURL: "https://api.airai.cc/v1"
});
client.chat.completions.create({
model: "qwen3.8-max",
messages: [{ role: "user", content: "Hello" }]
}).then((r) => console.log(r.choices[0].message.content));將 <YOUR_API_KEY> 替換為令牌設定中的 API Key。
相容端點
openaiopenai-responseopenai-response-compactanthropicgeminiopenai-alpha-search
Gemini SDK · TypeScript
import { GoogleGenerativeAI } from "@google/generative-ai";
const genAI = new GoogleGenerativeAI("<YOUR_API_KEY>");
const model = genAI.getGenerativeModel({ model: "qwen3.8-max" });
const result = await model.generateContent(
"Explain quantum entanglement in one paragraph."
);
console.log(result.response.text());身分驗證
所有請求必須攜帶 Authorization: Bearer <TOKEN> 請求頭;Anthropic 格式端點也接受 x-api-key 請求頭。在「令牌」頁面產生 API Key,可依模型、分組、IP、速率等維度精細化授權。
支援的參數
| 參數 | 類型 | 預設值 | 範圍 | 說明 |
|---|---|---|---|---|
temperature | float | 1 | 0–2 | 取樣溫度;越低越穩定 |
top_p | float | 1 | 0–1 | 核取樣累積機率 |
max_tokens | integer | — | — | 回應中最大 token 數 |
frequency_penalty | float | 0 | −2–2 | 懲罰高頻 token 的重複出現 |
presence_penalty | float | 0 | −2–2 | 鼓勵引入新話題 |
stop | string[] | null | ≤ 4 | 最多 4 個停止生成的字串 |
seed | integer | null | — | 盡量保證可復現的取樣種子 |
n | integer | 1 | — | 產生的候選條數 |
stream | boolean | false | — | 透過 SSE 串流回傳 token |
response_format | object | — | — | 強制輸出 JSON 物件或符合 Schema 的結果 |
tools | array | — | — | 模型可呼叫的工具 / 函式宣告 |
tool_choice | string | object | auto | — | 工具選擇策略或具體工具名 |
logprobs | boolean | false | — | 回傳每個 token 的對數機率 |
top_logprobs | integer | null | 0–20 | 每個 token 回傳的 top 機率數量 |
logit_bias | map | — | — | 依 token 的 logit 偏置對應 |
user | string | — | — | 用於風險審計的終端使用者識別 |
速率限制
| 分組 | RPM | TPM | RPD |
|---|---|---|---|
| Eco | 990 | 398K | 20K |
| Test | 980 | 393K | 20K |
資料更新於: 2026-10-10 18:35