Qwen3.8 Max
Qwen logo

Qwen3.8 Max

qwen3.8-maxllms.txt
Qwen
Qwen3.8-Max is Alibaba Cloud Tongyi Qianwen's next-generation flagship large language model, featuring a mixture-of-experts (MoE) architecture with 2.4 trillion parameters. It achieves another breakthrough in encoding depth, enabling it to handle more complex engineering-grade projects and long-term autonomous development; collaborative agent capabilities are significantly enhanced, performing more confidently in multi-tool orchestration and end-to-end delivery; visual understanding is comprehensively improved, with more sensitive and accurate chart reasoning, document parsing, and multimodal perception. Continuing the 1-million-context window, reasoning modes, and a complete tool ecosystem, it continues to evolve at a higher level of intelligence.

Pricing

PricingWeb SearchCache WriteCache Read
$1.69$5.07
$0.00055/request$2.1125/M tokens$0.169/M tokens

Input Modalities

  • Text
  • Vision
  • Video

Output Modalities

  • Text

Context length

  • 1M tokens

Max output

  • 131K tokens

Capabilities

  • Thinking
  • Streaming
  • Tool calling
  • Web search
  • URL context
  • Code interpreter
  • Computer use
  • File search
  • Memory tool
  • Structured outputs
  • Citations
  • Prompt caching
  • Background mode
  • Server-side sessions

Providers

Alibaba Cloud alicloud-qwen3.8-max
Pricing$1.69$5.07
Web Search$0.00055/request
Cache Write$2.1125/M tokens
Cache Read$0.169/M tokens
Context991K
Max output128K
Latency2.5S
Throughput39.5TPS
Uptime
99.90% uptime 2 days ago
99.89% uptime yesterday
99.45% uptime today

Performance for qwen3.8-max

Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).

Uptime
Loading...
Latency
Loading...
Throughput
Loading...

Try this model

Python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com/v1",
)

response = client.chat.completions.create(
    model="qwen3.8-max",
    messages=[
      {
        "role": "user",
        "content": "Hello, how are you?"
      }
    ],
    max_tokens=1024,
    stream=False,
)

print(response.choices[0].message.content)

Frequently asked questions

What is Qwen3.8 Max?

Qwen3.8-Max is Alibaba Cloud Tongyi Qianwen's next-generation flagship large language model, featuring a mixture-of-experts (MoE) architecture with 2.4 trillion parameters. It achieves another breakthrough in encoding depth, enabling it to handle more complex engineering-grade projects and long-term autonomous development; collaborative agent capabilities are significantly enhanced, performing more confidently in multi-tool orchestration and end-to-end delivery; visual understanding is comprehensively improved, with more sensitive and accurate chart reasoning, document parsing, and multimodal perception. Continuing the 1-million-context window, reasoning modes, and a complete tool ecosystem, it continues to evolve at a higher level of intelligence.