GPU Compute Rental and Hosted Inference

Rent GPU infrastructure for model training, fine-tuning, and other demanding workloads. Choose from the systems below and pay by the hour. We manage the physical hardware and hosting so your team can focus on building and running its models.

We are preparing hosted inference endpoints for open models. The table below shows the planned models, token pricing, and expected performance. These endpoints are not active yet, but you can contact us about early access or a custom deployment.

NODES

GPUsCPUsDRAMVRAMNVMe/HR
8x B300 SXM2x Xeon 128c @ 2.3 GHz2 TB @ 6400 MT/s2.30 TB HBM3e8x 3.84 TB @ 14 GB/s$59.12/HRRENT
8x B200 SXM2x Xeon 112c @ 2.1 GHz2 TB @ 5600 MT/s1.44 TB HBM3e8x 3.84 TB @ 7 GB/s$47.12/HR
4x H200 NVL2x EPYC 96c @ 2.75 GHz1 TB @ 4800 MT/s564 GB HBM3e4x 3.84 TB @ 7 GB/s$17.56/HR
4x Instinct MI455X1x EPYC 96c @ 5.0 GHz1 TB @ 8533 MT/s1.73 TB HBM44x 7.68 TB @ 14 GB/s$39.84/HR

ENDPOINTS

MODELINPUT / 1MOUTPUT / 1MCACHE / 1MLATENCYTOK/S
KIMI K3$3.00$15.00$0.30463 MS73 TOK/SPAUSED
GLM 5.2$0.475$1.493$0.095192 MS98 TOK/SPAUSED

CONTACT

PHONE
UNAVAILABLE

Have a question about availability, pricing, or a custom configuration? Send us a message and we’ll get back to you.