Microsoft logo

Launch Phi-4 Mini 3.8B

Compact model for reasoning
By Microsoftphi4-mini:3.8b3.8B parameters32K tokens of contextMIT

A compact Phi for instructions, maths and function calling. Useful when stronger reasoning than typical small models is needed at high speed.

GPU configurations for Phi-4 Mini 3.8B

Other models in the private loop

FAQ

Frequently asked about renting Phi-4 Mini 3.8B

Which GPU does Phi-4 Mini 3.8B need?

The entry configuration is 1x RTX3080TI, which delivers around 70 tokens per second. For more parallel requests and headroom on speed, go with 1x H200.

How much does Phi-4 Mini 3.8B cost?

An hour of Phi-4 Mini 3.8B runs from $0.24 to $8.26 depending on the GPU configuration. Billing is hourly: you pay for server time, not for tokens.

How do I call Phi-4 Mini 3.8B from my code?

Once deployed you get an OpenAI-compatible endpoint: point any OpenAI client at the base_url you receive and pass phi4-mini:3.8b as the model name. Code written against the OpenAI API needs no changes.

Does the data stay private?

Yes. Phi-4 Mini 3.8B runs on a dedicated GPU server inside a private loop: requests and responses never reach the model vendor and are not used for training.

What context window and license does Phi-4 Mini 3.8B have?

Phi-4 Mini 3.8B holds 32K tokens of context per request and ships under the MIT license, so you can use it commercially on your own hosting terms.