RunPod
Listed at https://www.runpod.io
Overall Rank #21 ⭐ Consider
⚠️ Partially available (31 global regions; CN direct access unstable — proxy recommended) | 🌍 International
💰 Token Pricing
| Type | Price | Note |
|---|---|---|
| Input | Per-second GPU: H100 SXM $2.99/hr, A100 SXM $1.49/hr, B200 $5.89/hr, L40S $0.99/hr, RTX 4090 $0.69/hr | per million tokens |
| Output | Same as input (per-GPU-second — no idle fees, no egress fees) | per million tokens |
| Cache Read | N/A (billed on actual GPU runtime, no prompt cache concept) | Discounted |
💡 Free Credits: $5 starter credit on signup, no credit card required; per-second billing means spiky workloads never pay for idle
🤖 Supported Models (22)
B300 288GB HBM3eB200 180GBH200 141GBH100 SXM 80GBH100 PCIe 80GBH100 NVL 94GBA100 SXM 80GBA100 PCIe 80GBRTX Pro 6000 96GBL40S 48GBRTX 6000 Ada 48GBL40 48GBRTX 5090 32GBRTX 4090 24GBL4 24GBA40 48GBRTX 3090 24GBSelf-hosted LLM (vLLM / SGLang / TensorRT-LLM)Stable Diffusion XL / FLUX / SD3 image generationWhisper / XTTS speech inferenceEmbeddings (BGE, GTE, sentence-transformers)Custom PyTorch / TensorFlow / JAX containers
✨ Pros
- ✓Per-second GPU billing, zero idle fees, no egress — cost-leader for bursty/inference workloads
- ✓13+ GPU tiers from RTX 4090 ($0.69/hr) to B300 288GB ($7.39/hr) — widest catalog among serverless GPU clouds
- ✓Serverless FlashBoot sub-200ms cold start; Pods deploy in under 30 seconds across 31 global regions
- ✓Community Cloud (low cost, peer-supplied GPUs) + Secure Cloud (Tier-3 datacenter) tier choice
- ✓Hub ships 100+ ready-to-deploy OSS templates (Llama, Qwen, DeepSeek, ComfyUI, Whisper)
- ✓Referral program: lifetime 5% credit rebate on referred developers' spending
⚠️ Cons
- ×Not a model API provider — you bring your own model + container image
- ×Community Cloud GPUs are interruptible (~15% preemption rate/hour) — production should use Secure Cloud
- ×Volume reservations require 1-3 month commit; pay-as-you-go is 20-40% more expensive than reserved
- ×No native prompt cache — long-context repeated-prefix workloads cost more than Anthropic/OpenAI
- ×China access unstable — needs proxy; no China-direct edge
- ×Documentation is sparser than AWS/GCP; expect to read template JSON for debugging
🎯 Best For
Open-source LLM hosting with low idle-cost (ComfyUI / Llama / Qwen / DeepSeek); bursty inference workloads; developers who need sub-30s GPU deployment without Kubernetes overhead; budget-conscious AI startups