Alibaba Qwen logo

Launch

Qwen3 reasoning flagship
By Alibaba Qwenqwen3:235b235B parameters32K tokens of contextApache 2.0

The Qwen3 flagship with 22B active parameters competes with leading reasoning models in code and maths. Requires at least two 80GB GPUs.

GPU configurations for Qwen3 235B-A22B

Other models in the private loop

FAQ

Frequently asked about renting Qwen3 235B-A22B

Which GPU does Qwen3 235B-A22B need?

The entry configuration is 2x A100, which delivers around 32 tokens per second. For more parallel requests and headroom on speed, go with 2x H200.

How much does Qwen3 235B-A22B cost?

An hour of Qwen3 235B-A22B runs from $4.30 to $18.02 depending on the GPU configuration. Billing is hourly: you pay for server time, not for tokens.

How do I call Qwen3 235B-A22B from my code?

Once deployed you get an OpenAI-compatible endpoint: point any OpenAI client at the base_url you receive and pass qwen3:235b as the model name. Code written against the OpenAI API needs no changes.

Does the data stay private?

Yes. Qwen3 235B-A22B runs on a dedicated GPU server inside a private loop: requests and responses never reach the model vendor and are not used for training.

What context window and license does Qwen3 235B-A22B have?

Qwen3 235B-A22B holds 32K tokens of context per request and ships under the Apache 2.0 license, so you can use it commercially on your own hosting terms.