Hyperbolic
Listed at https://hyperbolic.xyz
💰 Token Pricing
| Type | Price | Note |
|---|---|---|
| Input | Per-token inference scaled from market GPU rates; H100 from $2.89/GPU-hour, H200 $3.49/GPU-hour | per million tokens |
| Output | Pay-as-you-go GPU compute billed per hour; reserved discounts for committed capacity | per million tokens |
🤖 Supported Models (30)
✨ Pros
- ✓OpenAI-compatible inference API for open-weights models
- ✓On-demand H100/H200/B200 GPUs launched in minutes, zero quota limits
- ✓H200 at $3.49/GPU-hour, below RunPod/AWS/Azure mainstream rates
- ✓Forge infra layer with only ~2% virtualization overhead
- ✓Dedicated Model Hosting (single-tenant) with HIPAA/SOC2/GDPR readiness
⚠️ Cons
- ×No mainland China endpoint — proxy or aggregator required
- ×Per-token inference pricing not always listed publicly; driven by live GPU market rates
- ×Marketplace GPU prices fluctuate in real time with supply and demand
- ×Docs are JS-rendered, developer-oriented, English only
- ×No first-party frontier proprietary models; focused on open-weights ecosystem
🎯 Best For
Teams needing on-demand GPU for training/fine-tuning plus an OpenAI-compatible inference API; developers chasing the lowest H100/H200 hourly rates; international users who can use a proxy
💰 Pricing & Plans
| GPU / Plan | Price | Metering | Notes |
|---|---|---|---|
| H100 SXM | ~$2.89 / GPU-hour | Per hour, on-demand | Marketplace rate; cheapest mainstream Hopper option |
| H200 SXM | $3.49 / GPU-hour | Per hour, on-demand | 141GB HBM3e; below RunPod $4.39, AWS ~$5.97+, Azure ~$13.78 |
| B200 | Market rate | Per hour, on-demand | Blackwell for larger models and FP4 inference |
| Reserved | Discounted prepaid | Commitment <1 year | Dedicated capacity with predictable availability |
| Private Cloud | Custom quote | Long-term | Supplier-network dedicated infra, lowest cost at scale |
| Inference API | Per-token | Usage-based | OpenAI-compatible, scaled from live GPU market rates |
🔧 API & Developer Experience
- •API Surface: Hyperbolic exposes an OpenAI-compatible inference API, so the standard OpenAI SDK works by swapping base_url. The same account also provisions on-demand GPU instances over SSH for training, fine-tuning, batch compute and agent automation.
- •GPU Marketplace: Launch H100, H200 or B200 capacity in minutes from a dashboard with zero quota limits. Pay-as-you-go hourly billing with no long-term commitment, plus a live supply-and-demand-driven spot price.
- •Forge: Hyperbolic's infrastructure layer (launched June 10, 2026) manages the full machine lifecycle across distributed GPU suppliers — provisioning, security hardening, image management, monitoring and post-run sanitization — at roughly 2% virtualization overhead.
- •Serving Stack: Model hosting runs on Hyperbolic's proprietary inference engine, vLLM or SGLang. Dedicated Model Hosting exposes a private, customer-only API endpoint on single-tenant reserved GPUs.
- •Compliance: Dedicated and Private Cloud tiers support HIPAA, SOC2 and GDPR requirements via single-tenant compute, isolated networking and no prompt logging, with SLA-backed uptime targets.
🖥️ GPU Marketplace & Forge
Hyperbolic's core differentiator is the on-demand GPU marketplace. In July 2026 its H200 listed at $3.49 per GPU-hour (141GB HBM3e, 4.8TB/s bandwidth) — the lowest of the providers it compares against, undercutting RunPod's $4.39, AWS Capacity Blocks at roughly $5.97-6.87, Oracle at $10 and Azure at about $13.78 per GPU-hour. Behind the marketplace sits Forge, an infrastructure layer that standardizes provisioning, security hardening and sanitization across dozens of distributed GPU suppliers, with virtualization overhead of only about 2% versus bare metal. New users get instant capacity with no quota games, and can scale from a single on-demand GPU into Reserved clusters or Private Cloud as workloads grow.
🌐 Regional Availability & Latency
Hyperbolic is a US-based open-access AI cloud that sources GPU capacity from a global network of compute providers, so latency depends on which supplier and region you provision. There is no mainland China regional endpoint and no official China access program, so production use from China requires a stable overseas proxy or an aggregator that fronts the Hyperbolic API. For teams outside China the on-demand model means you can provision capacity close to your users or your training data, and the OpenAI-compatible inference surface keeps integration latency low. The distributed supplier model trades single-region consistency for broad, flexible capacity across North America, Europe and Asia-Pacific regions.