Google logo

Launch

Lightweight text model
By Googlegemma3:1b1B parameters32K tokens of contextGemma Terms of Use

A small text-only Gemma for fast predictable tasks. Useful as a classifier, filter or inexpensive LLM pipeline stage.

GPU configurations for Gemma 3 1B

Gemma 3 1B benchmarks

as published by the model vendor
Gemma 3 1B scores in public benchmarks
MMLU26.5
GSM8K1.4
HumanEval6.1

Other models in the private loop

FAQ

Frequently asked about renting Gemma 3 1B

Which GPU does Gemma 3 1B need?

The entry configuration is 1x RTX3080TI, which delivers around 105 tokens per second. For more parallel requests and headroom on speed, go with 1x H200.

How much does Gemma 3 1B cost?

An hour of Gemma 3 1B runs from $0.24 to $8.26 depending on the GPU configuration. Billing is hourly: you pay for server time, not for tokens.

How do I call Gemma 3 1B from my code?

Once deployed you get an OpenAI-compatible endpoint: point any OpenAI client at the base_url you receive and pass gemma3:1b as the model name. Code written against the OpenAI API needs no changes.

Does the data stay private?

Yes. Gemma 3 1B runs on a dedicated GPU server inside a private loop: requests and responses never reach the model vendor and are not used for training.

What context window and license does Gemma 3 1B have?

Gemma 3 1B holds 32K tokens of context per request and ships under the Gemma Terms of Use license, so you can use it commercially on your own hosting terms.