RunPod (RunPod)

收录于 https://www.runpod.io

综合排名 #21 ⭐ 可考虑
⚠️ 部分可用(基础设施在全球 31 区域,国内直连不稳定,需稳定代理) | 🌍 国际

💰 Token 价格

类型 价格 备注
输入 (Input) Per-second GPU: H100 SXM $2.99/hr, A100 SXM $1.49/hr, B200 $5.89/hr, L40S $0.99/hr, RTX 4090 $0.69/hr 每百万 tokens
输出 (Output) Same as input (per-GPU-second — no idle fees, no egress fees) 每百万 tokens
缓存读取 N/A (billed on actual GPU runtime, no prompt cache concept) 享折扣
💡 免费额度

$5 starter credit on signup, no credit card required; per-second billing means spiky workloads never pay for idle

🤖 支持模型(共 22 个)

B300 288GB HBM3e B200 180GB H200 141GB H100 SXM 80GB H100 PCIe 80GB H100 NVL 94GB A100 SXM 80GB A100 PCIe 80GB RTX Pro 6000 96GB L40S 48GB RTX 6000 Ada 48GB L40 48GB RTX 5090 32GB RTX 4090 24GB L4 24GB A40 48GB RTX 3090 24GB Self-hosted LLM (vLLM / SGLang / TensorRT-LLM) Stable Diffusion XL / FLUX / SD3 image generation Whisper / XTTS speech inference Embeddings (BGE, GTE, sentence-transformers) Custom PyTorch / TensorFlow / JAX containers

✨ 优势

  • Per-second GPU billing, zero idle fees, no egress — cost-leader for bursty/inference workloads
  • 13+ GPU tiers from RTX 4090 ($0.69/hr) through B300 288GB ($7.39/hr) — widest catalog among serverless GPU clouds
  • Serverless FlashBoot sub-200ms cold start; Pods deploy in <30s across 31 global regions
  • Community Cloud (lower cost, peer-supplied GPUs) + Secure Cloud (Tier-3 datacenter) tier choice
  • Hub ships 100+ ready-to-deploy OSS templates (Llama, Qwen, DeepSeek, ComfyUI, Whisper)
  • Referral program: lifetime 5% credit rebate on referred developers' spending

⚠️ 不足

  • × Not a model API provider — you bring your own model + container image
  • × Community Cloud GPUs are interruptible (~15% preemption rate/hour) — production should use Secure Cloud
  • × Volume reservations require 1-3 month commit; pay-as-you-go is 20-40% more expensive than reserved
  • × No native prompt cache — long-context repeated-prefix workloads cost more than Anthropic/OpenAI
  • × China access unstable — needs proxy; no China-direct edge
  • × Documentation is sparser than AWS/GCP; expect to read template JSON for debugging

🎯 适合场景

Open-source LLM hosting with low idle-cost (ComfyUI / Llama / Qwen / DeepSeek); bursty inference workloads; developers who need sub-30s GPU deployment without Kubernetes overhead; budget-conscious AI startups