modelhub.help

R1 Distill Llama 70B

DeepSeekOpen-weight

131K context · — params · —

deepseek/deepseek-r1-distill-llama-70b

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...

Cheapest blended: ¥0.725 / 1M tokens on DeepSeekGet custom quote

Reachability & verification

Live reachability

Probing…

Target (DeepSeek): https://api.deepseek.com/v1

Live reachability probe

Measured in real time from the ModelHub gateway egress to this model's API endpoint. Green means reachable right now. Overseas reference latencies are estimates; the live probe is the proof of current reachability.

Pricing across providers

1

Prices are shown in each provider's billing currency — ¥ (CNY) for Chinese providers, $ (USD) for overseas gateways.

ProviderInput /1MOutput /1MBlendedCtxp50 latencyAPIVerified
DeepSeekdeepseek-r1-distill-llama-70b¥0.7¥0.8¥0.725131K OAI119d ago

Works with

CursorClineAiderContinueOpenCodeOpen WebUI

Any OpenAI-compatible client works — point the base URL at the provider endpoint.

Capabilities

Code samplesExample usingDeepSeek— the cheapest hosting for this model as of last verification. Swapbase_urlandmodelto use a different provider from the matrix above.PythoncURLNode.jsCopyfrom openai import OpenAI client = OpenAI( api_key="YOUR_API_KEY", base_url="https://api.deepseek.com/v1", ) response = client.chat.completions.create( model="deepseek-r1-distill-llama-70b", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content)code

Code samples

Python
from openai import OpenAI

client = OpenAI(
    base_url="https://api.deepseek.com/v1",
    api_key="YOUR_API_KEY",
)

resp = client.chat.completions.create(
    model="deepseek-r1-distill-llama-70b",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
cURL
curl https://api.deepseek.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-r1-distill-llama-70b",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
Node.js
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.deepseek.com/v1",
  apiKey: "YOUR_API_KEY",
});

const resp = await client.chat.completions.create({
  model: "deepseek-r1-distill-llama-70b",
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(resp.choices[0].message.content);

Technical specs

Context

131K

Max output

16K

Parameters

Release

Training cutoff

License

Similar models

Frequently asked questions

How much does R1 Distill Llama 70B cost?

The cheapest tracked blended price is ¥0.725 per 1M tokens on DeepSeek. See the pricing matrix above for input/output splits per provider.

Can I use R1 Distill Llama 70B from outside China?

Availability depends on the hosting provider. Use our Path Planner on the DeepSeek page to map a verified official-vs-brokered access path, including payment and KYC constraints.

Is R1 Distill Llama 70B open source?

Yes — R1 Distill Llama 70B is open-weight under the — license. You can self-host it or use any listed inference provider.

Is R1 Distill Llama 70B OpenAI-compatible?

Yes — at least one tracked provider exposes an OpenAI-compatible endpoint, so Cursor, Cline, Aider and similar clients work with just a base-URL change.

What is the maximum context window?

R1 Distill Llama 70B supports up to 131K tokens of context with a maximum output of 16K tokens.

Need verified access to R1 Distill Llama 70B?

We broker vetted intros to DeepSeek — payment, KYC and endpoint verification handled.

Request a brokerage intro