Skip to content
huggingface.co·

🤖Qwen3.8-Max Released With Vision Input and 1M Context

New Qwen Version Adds Vision Capabilities and Massive Context

TL;DR

Qwen Cloud launches Qwen3.8-Max with vision input, non-thinking support, and a default context length of 1 million tokens. The model supports vLLM, SGLang, TokenSpeed, and more.

Qwen Cloud has just released Qwen3.8-Max, the latest version of their powerful language model with significant enhancements. This new release includes vision input capabilities, non-thinking support, and a default context length of 1 million tokens. If you're working on complex multi-step tasks or need robust coding assistance, this update is crucial. The model boasts over 2 trillion parameters in total and has an activated hidden dimension of 95 billion with a token embedding dimension of 248,320 (Padded). It's designed to be deployed using popular inference frameworks like SGLang, vLLM, or TokenSpeed.

Qwen3.8-Max Released With Vision Input and 1M Context — huggingface.co

Key Points

1

Qwen3.8-Max features a default context length of 1 million tokens, up from previous versions.

2

The model supports popular inference frameworks such as SGLang, vLLM, and TokenSpeed.

3

Vision input capabilities allow the model to handle image-based tasks more effectively.

4

Coding Agent Terminal Bench score for Qwen3.8-Max is 88.8, outperforming previous versions.

5

Qwen Cloud's official API service provides managed, scalable inference without infrastructure maintenance.

Why It Matters

If you're working on complex multi-step tasks or need robust coding assistance, the vision input and non-thinking support in Qwen3.8-Max are game-changers. The default context length of 1 million tokens is a massive improvement for long-horizon agentic tasks. However, smaller teams might find the premium pricing for managed services less appealing unless they hit high IOPS.

Qwen3.8-Maxvision-inputcontext-lengthinference-frameworks

Frequently Asked Questions

Why does this matter?

If you're working on complex multi-step tasks or need robust coding assistance, the vision input and non-thinking support in Qwen3.8-Max are game-changers. The default context length of 1 million tokens is a massive improvement for long-horizon agentic tasks. However, smaller teams might find the premium pricing for managed services less appealing unless they hit high IOPS.

What happened?

Qwen Cloud launches Qwen3.8-Max with vision input, non-thinking support, and a default context length of 1 million tokens. The model supports vLLM, SGLang, TokenSpeed, and more.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 2,950 builders reading daily.

Also get