Hyperbolic

Listed at https://hyperbolic.xyz

Overall Rank #15 ⭐ Consider
❌ Proxy required | 🌍 International

💰 Token Pricing

TypePriceNote
Input Per-token inference scaled from market GPU rates; H100 from $2.89/GPU-hour, H200 $3.49/GPU-hour per million tokens
Output Pay-as-you-go GPU compute billed per hour; reserved discounts for committed capacity per million tokens
💡 Free Credits: No permanent free tier and no long-term lock-in. On-demand GPUs are pay-as-you-go billed per hour, with no quota limits and no minimum commitment to start.

🤖 Supported Models (30)

DeepSeek R1DeepSeek R1-0528Llama 3.3 70BQwen 2.5 72BNVIDIA NemotronMixtral 8x22B

✨ Pros

  • OpenAI-compatible inference API for open-weights models
  • On-demand H100/H200/B200 GPUs launched in minutes, zero quota limits
  • H200 at $3.49/GPU-hour, below RunPod/AWS/Azure mainstream rates
  • Forge infra layer with only ~2% virtualization overhead
  • Dedicated Model Hosting (single-tenant) with HIPAA/SOC2/GDPR readiness

⚠️ Cons

  • ×No mainland China endpoint — proxy or aggregator required
  • ×Per-token inference pricing not always listed publicly; driven by live GPU market rates
  • ×Marketplace GPU prices fluctuate in real time with supply and demand
  • ×Docs are JS-rendered, developer-oriented, English only
  • ×No first-party frontier proprietary models; focused on open-weights ecosystem

🎯 Best For

Teams needing on-demand GPU for training/fine-tuning plus an OpenAI-compatible inference API; developers chasing the lowest H100/H200 hourly rates; international users who can use a proxy

💰 Pricing & Plans

GPU / PlanPriceMeteringNotes
H100 SXM~$2.89 / GPU-hourPer hour, on-demandMarketplace rate; cheapest mainstream Hopper option
H200 SXM$3.49 / GPU-hourPer hour, on-demand141GB HBM3e; below RunPod $4.39, AWS ~$5.97+, Azure ~$13.78
B200Market ratePer hour, on-demandBlackwell for larger models and FP4 inference
ReservedDiscounted prepaidCommitment <1 yearDedicated capacity with predictable availability
Private CloudCustom quoteLong-termSupplier-network dedicated infra, lowest cost at scale
Inference APIPer-tokenUsage-basedOpenAI-compatible, scaled from live GPU market rates

🔧 API & Developer Experience

  • API Surface: Hyperbolic exposes an OpenAI-compatible inference API, so the standard OpenAI SDK works by swapping base_url. The same account also provisions on-demand GPU instances over SSH for training, fine-tuning, batch compute and agent automation.
  • GPU Marketplace: Launch H100, H200 or B200 capacity in minutes from a dashboard with zero quota limits. Pay-as-you-go hourly billing with no long-term commitment, plus a live supply-and-demand-driven spot price.
  • Forge: Hyperbolic's infrastructure layer (launched June 10, 2026) manages the full machine lifecycle across distributed GPU suppliers — provisioning, security hardening, image management, monitoring and post-run sanitization — at roughly 2% virtualization overhead.
  • Serving Stack: Model hosting runs on Hyperbolic's proprietary inference engine, vLLM or SGLang. Dedicated Model Hosting exposes a private, customer-only API endpoint on single-tenant reserved GPUs.
  • Compliance: Dedicated and Private Cloud tiers support HIPAA, SOC2 and GDPR requirements via single-tenant compute, isolated networking and no prompt logging, with SLA-backed uptime targets.

🖥️ GPU Marketplace & Forge

Hyperbolic's core differentiator is the on-demand GPU marketplace. In July 2026 its H200 listed at $3.49 per GPU-hour (141GB HBM3e, 4.8TB/s bandwidth) — the lowest of the providers it compares against, undercutting RunPod's $4.39, AWS Capacity Blocks at roughly $5.97-6.87, Oracle at $10 and Azure at about $13.78 per GPU-hour. Behind the marketplace sits Forge, an infrastructure layer that standardizes provisioning, security hardening and sanitization across dozens of distributed GPU suppliers, with virtualization overhead of only about 2% versus bare metal. New users get instant capacity with no quota games, and can scale from a single on-demand GPU into Reserved clusters or Private Cloud as workloads grow.

🌐 Regional Availability & Latency

Hyperbolic is a US-based open-access AI cloud that sources GPU capacity from a global network of compute providers, so latency depends on which supplier and region you provision. There is no mainland China regional endpoint and no official China access program, so production use from China requires a stable overseas proxy or an aggregator that fronts the Hyperbolic API. For teams outside China the on-demand model means you can provision capacity close to your users or your training data, and the OpenAI-compatible inference surface keeps integration latency low. The distributed supplier model trades single-region consistency for broad, flexible capacity across North America, Europe and Asia-Pacific regions.