131K context · 72B params · Apache 2.0
alibaba/qwen3-72b-instruct
Alibaba's Qwen3 72B Instruct is the latest open-weight flagship in the Qwen family, with strong multilingual (especially Chinese) and reasoning capabilities. Apache 2.0 licensed; available on DashScope (official) and major open-weight aggregators.
Reachability & verification
Live reachability
Target (Alibaba Cloud DashScope): https://api.together.xyz/v1
Overseas reference latency (est.)
Live reachability probe
Measured in real time from the ModelHub gateway egress to this model's API endpoint. Green means reachable right now. Overseas reference latencies are estimates; the live probe is the proof of current reachability.
Official Path Verified by ModelHub
Pricing across providers
2Prices are shown in each provider's billing currency — ¥ (CNY) for Chinese providers, $ (USD) for overseas gateways.
| Provider | Input /1M | Output /1M | Blended | Ctx | p50 latency | API | Verified |
|---|---|---|---|---|---|---|---|
| Together.aiQwen/Qwen3-72B-Instruct | $0.9 | $0.9 | $0.9 | 131K | 220 ms | OAI | 140d ago |
| Alibaba Cloud DashScopeqwen-72b-instruct | ¥0.75 | ¥1.5 | ¥0.9375 | 131K | 420 ms | OAI | 140d ago |
Works with
Any OpenAI-compatible client works — point the base URL at the provider endpoint.
Capabilities
Languages: en, zh, ja, ko, ar, es, fr, de
Benchmarks
| Benchmark | Setting | Verification | Score |
|---|---|---|---|
| MMLU | 5-shot | official | 85.9% |
| HumanEval | 0-shot pass@1 | official | 82.1% |
Code samples
from openai import OpenAI
client = OpenAI(
base_url="https://api.together.xyz/v1",
api_key="YOUR_API_KEY",
)
resp = client.chat.completions.create(
model="Qwen/Qwen3-72B-Instruct",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)curl https://api.together.xyz/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3-72B-Instruct",
"messages": [{"role": "user", "content": "Hello!"}]
}'import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.together.xyz/v1",
apiKey: "YOUR_API_KEY",
});
const resp = await client.chat.completions.create({
model: "Qwen/Qwen3-72B-Instruct",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(resp.choices[0].message.content);Technical specs
Context
131K
Max output
8K
Parameters
72B
Release
2025-01-28
Training cutoff
2024-09-01
License
Apache 2.0
Similar models
Frequently asked questions
How much does Qwen3 72B Instruct cost?
The cheapest tracked blended price is $0.9 per 1M tokens on Together.ai. See the pricing matrix above for input/output splits per provider.
Can I use Qwen3 72B Instruct from outside China?
Availability depends on the hosting provider. Use our Path Planner on the Alibaba Cloud DashScope page to map a verified official-vs-brokered access path, including payment and KYC constraints.
Is Qwen3 72B Instruct open source?
Yes — Qwen3 72B Instruct is open-weight under the Apache 2.0 license. You can self-host it or use any listed inference provider.
Is Qwen3 72B Instruct OpenAI-compatible?
Yes — at least one tracked provider exposes an OpenAI-compatible endpoint, so Cursor, Cline, Aider and similar clients work with just a base-URL change.
What is the maximum context window?
Qwen3 72B Instruct supports up to 131K tokens of context with a maximum output of 8K tokens.
Need verified access to Qwen3 72B Instruct?
We broker vetted intros to Alibaba Cloud DashScope — payment, KYC and endpoint verification handled.