LLM rental with hourly GPU billing

Run open-weight models on dedicated Qudata GPUs instead of paying per token.

Pick a model and we bring it up on a matching GPU behind an OpenAI-compatible API. Prompts and data never leave the rented server, there is no cap on requests, and you pay only for GPU hours.

Alibaba Qwen

Qwen

By Alibaba Qwen

Mainstream multimodal all-rounder

Runs on 1x RTX4060TI~34 tok/s
OpenAI

GPT-OSS

By OpenAI

Open reasoning model for agents

Runs on 1x RTX3090~27 tok/s
Z.ai

GLM

By Z.ai

Fast GLM for code and agents

Runs on 1x RTX5090~52 tok/s
DeepSeek

DeepSeek

By DeepSeek

Powerful reasoning model

Runs on 1x RTX5090~32 tok/s
Google

Gemma

By Google

Compact Gemma with vision

Runs on 1x RTX3080TI~70 tok/s
Moonshot AI

Kimi

By Moonshot AI

Flagship MoE for agentic coding

Runs on 3x H200~19 tok/s
Meta

Llama

By Meta

Long-context multimodal MoE

Runs on 1x A100~28 tok/s
Mistral AI

Mistral

By Mistral AI

Precise multimodal all-rounder

Runs on 1x RTX3090~24 tok/s
Microsoft

Phi

By Microsoft

Enhanced reasoning model

Runs on 1x RTX3090~36 tok/s
Agentica & Together AI

DeepCoder

By Agentica & Together AI

Reasoning model for programming

Runs on 1x RTX3090~36 tok/s