Google logo

Launch

Compact Gemma with vision
By Googlegemma3:4b4B parameters32K tokens of contextGemma Terms of Use

A compact multimodal Gemma for images and short conversations. Supports over 140 languages and runs well on inexpensive GPUs.

GPU configurations for Gemma 3 4B

1x RTX3080TI

Best value
~70 tok/s
$0.19per hour

1x RTX4070TI

~84 tok/s
$0.18per hour

1x RTX4060TI

~76 tok/s
$0.16per hour

1x RTX4070TISUPER

~92 tok/s
$0.25per hour

1x RTX5070TI

~108 tok/s
$0.24per hour

1x RTX3090

~88 tok/s
$0.90per hour

1x RTX3090TI

~94 tok/s
$0.53per hour

1x RTX4090

~120 tok/s
$1.33per hour

1x RTX5090

~155 tok/s
$1.78per hour

1x A100

~115 tok/s
$2.50per hour

1x H100

~185 tok/s
$5.92per hour

1x H200

~208 tok/s
$8.26per hour
Pick a GPU
VRAM
RAM
Storage
CPU

Gemma 3 4B benchmarks

as published by the model vendor
Gemma 3 4B scores in public benchmarks
MMLU59.6
GSM8K38.4
HumanEval36

Other models in the private loop

FAQ

Frequently asked about renting Gemma 3 4B

Which GPU does Gemma 3 4B need?

The entry configuration is 1x RTX3080TI, which delivers around 70 tokens per second. For more parallel requests and headroom on speed, go with 1x H200.

How much does Gemma 3 4B cost?

An hour of Gemma 3 4B runs from $0.16 to $8.26 depending on the GPU configuration. Billing is hourly: you pay for server time, not for tokens.

How do I call Gemma 3 4B from my code?

Once deployed you get an OpenAI-compatible endpoint: point any OpenAI client at the base_url you receive and pass gemma3:4b as the model name. Code written against the OpenAI API needs no changes.

Does the data stay private?

Yes. Gemma 3 4B runs on a dedicated GPU server inside a private loop: requests and responses never reach the model vendor and are not used for training.

What context window and license does Gemma 3 4B have?

Gemma 3 4B holds 32K tokens of context per request and ships under the Gemma Terms of Use license, so you can use it commercially on your own hosting terms.