Google logo

Launch

Compact model for text
By Googlegemma2:9b9B parameters8K tokens of contextGemma Terms of Use

A Google model with a tidy, well-mannered answer style — it drifts into improvisation and rudeness far less often. Context is only 8K, so it is for short tasks: support replies, emails, rewriting.

GPU configurations for Gemma 2 9B

1x RTX4060TI

Best value
~34 tok/s
$0.13per hour

1x RTX4070TISUPER

~48 tok/s
$0.23per hour

1x RTX5070TI

~58 tok/s
$0.24per hour

1x RTX3090

~55 tok/s
$0.22per hour

1x RTX3090TI

~60 tok/s
$0.53per hour

1x RTX4090

~82 tok/s
$0.54per hour

1x RTX5090

~108 tok/s
$0.64per hour

1x A100

~72 tok/s
$1.52per hour

1x H100

~118 tok/s
$3.37per hour

1x H200

~132 tok/s
$7.63per hour
Pick a GPU
VRAM
RAM
Storage
CPU

Gemma 2 9B benchmarks

as published by the model vendor
Gemma 2 9B scores in public benchmarks
MMLU71.3
GSM8K68.6
HumanEval40.2

Other models in the private loop

FAQ

Frequently asked about renting Gemma 2 9B

Which GPU does Gemma 2 9B need?

The entry configuration is 1x RTX4060TI, which delivers around 34 tokens per second. For more parallel requests and headroom on speed, go with 1x H200.

How much does Gemma 2 9B cost?

An hour of Gemma 2 9B runs from $0.13 to $7.63 depending on the GPU configuration. Billing is hourly: you pay for server time, not for tokens.

How do I call Gemma 2 9B from my code?

Once deployed you get an OpenAI-compatible endpoint: point any OpenAI client at the base_url you receive and pass gemma2:9b as the model name. Code written against the OpenAI API needs no changes.

Does the data stay private?

Yes. Gemma 2 9B runs on a dedicated GPU server inside a private loop: requests and responses never reach the model vendor and are not used for training.

What context window and license does Gemma 2 9B have?

Gemma 2 9B holds 8K tokens of context per request and ships under the Gemma Terms of Use license, so you can use it commercially on your own hosting terms.