Q
ALIBABA QWEN / qwen3-max-2026-01-23

Qwen3‑Max Thinking

Multimodal reasoning, coding, and complex problems.

01 / SUMMARY

The essential

ORIENTATIONLong-running reasoning

Primary use stated or inferred cautiously from official documentation.

CONTEXT262.000 tokens

Maximum input capacity when published by the source.

EXIT65.000 tokens

Documented maximum generation limit.

ACCESSQwen API · Alibaba Cloud Model Studio

Verified availability channels.

MODEL TYPEText and reasoning

Primary functional family and verified specialties.

LICENSEProprietary

Declared terms for API access, model weights, or self-hosting.

02 / TECHNICAL SHEET

Limits and integration

API IDqwen3-max-2026-01-23
Model typeText and reasoning
Access modelPaid proprietary
LicenseProprietary
DeploymentHosted API
Release25 JAN 2026
Knowledge cutoffNot published
EntranceText
ExitText
Context window262.000 tokens
maximum output65.000 tokens
ReasoningModes with and without reasoning
Published toolsFeatures · Web search · Web extraction · Code interpreter
Structured OutingsSupported
Batch processingSupported
Prompt cacheSupported
Fine-tuningNot published
Verified platformsQwen API · Qwen Chat · Alibaba Cloud
BEST SUITED FOR

Multimodal reasoning, coding, and complex problems.

WORTH MONITORING

Execution conditions vary across benchmarks; do not compare the number alone.

03 / CAPABILITIES

What it can do

01

Orientation

Multimodal reasoning, coding, and complex problems.

02

Context

262.000 tokens · 65.000 tokens

03

Tools and integration

Features · Web search · Web extraction · Code interpreter

04

Access

Qwen API · Alibaba Cloud Model Studio

Input modalities
Text
04 / PRICES

Documented cost

ConceptWorthUnit/condition
Standard input1,20 USD / 1 M tokens (≤32k)Standard API rate
Cached input0,24 USD / 1 M tokensReading reused prefixes
Cache write or storageNot publishedThe condition varies by provider
Standard output6,00 USD / 1 M tokens (≤32k)May include reasoning tokens
Batch input0,60 USD / 1 M tokens (≤32k)Asynchronous processing
Batch output3,00 USD / 1 M tokens (≤32k)Asynchronous processing

Higher pricing tiers apply above 32,000 tokens.

i Prices change and may depend on level, region or context length. Check the source before making a decision.

05 / EVALUATIONS

How to read the results

“

A public benchmark provides guidance, but does not replace an evaluation with your data, tools, budget, and error tolerance.

BenchmarkResultMetricSource
GPQA Diamond92,8 %AccuracyView source ↗
LiveCodeBench v691,4 %AccuracyView source ↗
Humanity's Last Exam58,3 %With toolsView source ↗

Inferama only highlights a “best result” when the metric, test set, configuration, and date allow for an equivalent comparison. The supplier's figures are presented as claims from its own source.

06 / VERSIONS

Chronology

Qwen3‑Max Thinking

Version added to Inferama's verified catalog.

07 / SOURCES

Traceability