Meta logo

Launch

Flagship multimodal Llama
By Metallama4:maverick400B parameters32K tokens of contextLlama 4 Community License

A large multimodal Llama 4 with 400B total and 17B active parameters. It offers flagship quality but requires a multi-GPU server.

GPU configurations for Llama 4 Maverick

Llama 4 Maverick benchmarks

as published by the model vendor
Llama 4 Maverick scores in public benchmarks
MMLU Pro80.5
GPQA Diamond69.8
LiveCodeBench43.4

Other models in the private loop

FAQ

Frequently asked about renting Llama 4 Maverick

Which GPU does Llama 4 Maverick need?

The entry configuration is 2x H200, which delivers around 27 tokens per second. For more parallel requests and headroom on speed, go with 4x H200.

How much does Llama 4 Maverick cost?

An hour of Llama 4 Maverick runs from $14.20 to $18.20 depending on the GPU configuration. Billing is hourly: you pay for server time, not for tokens.

How do I call Llama 4 Maverick from my code?

Once deployed you get an OpenAI-compatible endpoint: point any OpenAI client at the base_url you receive and pass llama4:maverick as the model name. Code written against the OpenAI API needs no changes.

Does the data stay private?

Yes. Llama 4 Maverick runs on a dedicated GPU server inside a private loop: requests and responses never reach the model vendor and are not used for training.

What context window and license does Llama 4 Maverick have?

Llama 4 Maverick holds 32K tokens of context per request and ships under the Llama 4 Community License license, so you can use it commercially on your own hosting terms.