Alibaba Qwen logo

Launch

Compact vision model
By Alibaba Qwenqwen3-vl:4b4B parameters32K tokens of contextApache 2.0

A compact vision-language model for OCR, screenshots, diagrams and documents. Suited to bulk image processing without an expensive GPU.

GPU configurations for Qwen3-VL 4B

Other models in the private loop

FAQ

Frequently asked about renting Qwen3-VL 4B

Which GPU does Qwen3-VL 4B need?

The entry configuration is 1x RTX3080TI, which delivers around 70 tokens per second. For more parallel requests and headroom on speed, go with 1x H200.

How much does Qwen3-VL 4B cost?

An hour of Qwen3-VL 4B runs from $0.24 to $8.26 depending on the GPU configuration. Billing is hourly: you pay for server time, not for tokens.

How do I call Qwen3-VL 4B from my code?

Once deployed you get an OpenAI-compatible endpoint: point any OpenAI client at the base_url you receive and pass qwen3-vl:4b as the model name. Code written against the OpenAI API needs no changes.

Does the data stay private?

Yes. Qwen3-VL 4B runs on a dedicated GPU server inside a private loop: requests and responses never reach the model vendor and are not used for training.

What context window and license does Qwen3-VL 4B have?

Qwen3-VL 4B holds 32K tokens of context per request and ships under the Apache 2.0 license, so you can use it commercially on your own hosting terms.