CoreWeave
Listed at https://www.coreweave.com
💰 Token Pricing
| Type | Price | Note |
|---|---|---|
| Input | GPU On-Demand hourly: NVIDIA HGX H100 $49.24, HGX H200 $50.44, HGX B200 $68.80, A100 80GB $21.60, L40S $18.00, L40 $10.00, GB200 NVL72 $42.00. Spot discounts: H100 $19.71, H200 $20.93, B200 $34.11, A100 $9.65. Serverless Inference is pay-per-token via W&B Inference; Dedicated Inference is GPU-hour / node-based. | per million tokens |
| Output | per million tokens |
🤖 Supported Models (8)
✨ Pros
- ✓Nasdaq-listed (CRWV), NVIDIA owns ~11% — a leading 'neocloud' provider
- ✓All three inference products are OpenAI-compatible: Serverless (W&B Inference, per-token), Dedicated (bring-your-own model weights, GPU-hour), and Inference on CKS (full control over the serving stack)
- ✓Notable customers: OpenAI (additional $6.5B), Anthropic (multi-year), Meta ($21B expansion), Perplexity, Runway, Jane Street ($6B), MasterClass (W&B Weave to trace AI teaching agents)
- ✓MLPerf training and inference records; trained DeepSeek V3 in ~2 minutes; first to bring up NVIDIA Vera Rubin NVL72 and GB300 NVL72
- ✓0EM (Zero Egress Migration) program + AI Object Storage + Sandboxes (RL training) lower the cost of migrating and running AI workloads
⚠️ Cons
- ×No mainland China direct endpoint; China production deployments require proxy; first APAC node only landed in Indonesia in 2026
- ×GPU On-Demand hourly prices run high (H100 $49.24/hr is a premium over many self-run endpoints); Serverless per-token rates are scattered across the W&B catalog with no unified pricing table
- ×Console/docs lean toward Kubernetes and professional AI-engineering teams; a pure-API newcomer faces a higher learning curve than OpenAI
- ×Capacity/demand fluctuates — high-demand GPUs (e.g. GB200) may need waitlisting or Capacity Plans/reservations on On-Demand
- ×Serverless inference ecosystem is delivered through W&B (acquired by CoreWeave) and sits partly outside the CoreWeave account system, adding a modest integration-learning cost
🎯 Best For
AI teams that want OpenAI-compatible inference API plus underlying GPU compute in one place (Serverless per-token + Dedicated bring-your-own-model + CKS full control); teams that train or fine-tune their own models on CoreWeave GPUs and then serve them via OpenAI-compatible endpoints; enterprises that value MLPerf records and NVIDIA's latest platforms (Vera Rubin NVL72, GB300); engineering orgs with large usage that want 0EM zero-egress migration and native object storage.
💰 Pricing & Plans
| Resource | Billing Unit | Price | Notes |
|---|---|---|---|
| NVIDIA HGX H100 (On-Demand) | per GPU-hour | $49.24 | Spot $19.71 (~60% off); Inference single-GPU $6.16/hr |
| NVIDIA HGX H200 (On-Demand) | per GPU-hour | $50.44 | Spot $20.93; Inference single-GPU $6.31/hr |
| NVIDIA HGX B200 (On-Demand) | per GPU-hour | $68.80 | Spot $34.11; Inference single-GPU $8.60/hr |
| NVIDIA A100 80GB (On-Demand) | per GPU-hour | $21.60 | Spot $9.65; Inference single-GPU $2.70/hr |
| NVIDIA L40S (On-Demand) | per GPU-hour | $18.00 | Spot $7.88; Inference single-GPU $2.25/hr |
| NVIDIA L40 (On-Demand) | per GPU-hour | $10.00 | Spot $6.27; Inference single-GPU $1.25/hr |
| NVIDIA GB200 NVL72 (On-Demand) | per GPU-hour | $42.00 | Spot N/A; 4-GPU Blackwell superchip system |
| Serverless Inference (W&B) | per token | W&B catalog | Pay-per-token for catalog models (GLM 5.2, Kimi K2.6/K2.7, DeepSeek R1/V3, Llama, Qwen) |
| Dedicated Inference | GPU-hour / node | Contract | Bring-your-own model weights on dedicated GPUs; full gateway + scaling control |
| Inference on CKS | GPU-hour / reserved node | CKS pricing | Full control over serving stack on CoreWeave Kubernetes Service |
| 0EM Zero Egress Migration | free during migration | $0 egress | No egress fees, no lock-in; expert-led transfer to AI object storage |
🔧 API & Developer Experience
- •API Style: All three inference products expose OpenAI API-compatible endpoints. Existing OpenAI client libraries, agents, and tooling connect with minimal changes. Dedicated Inference is configured through the CoreWeave Inference management API (REST/JSON at api.coreweave.com, gRPC over HTTP/2, and a Terraform provider).
- •Deployment Options: Three ways to serve models: Serverless Inference (CoreWeave-managed catalog, auto-scaling, no infra to manage), Dedicated Inference (bring your own model weights on dedicated GPU infra with full control over gateways, scaling, and capacity reservations), and Inference on CKS (complete control over your serving stack, runtimes, and networking on CoreWeave Kubernetes Service).
- •NVIDIA GPU Access: H100, H200, B200, A100, L40S, L40, and GB200 NVL72 on demand with spot discounts of 40-60%. First cloud to deploy NVIDIA Vera Rubin NVL72 and GB300 NVL72 — early access to the newest NVIDIA platforms for those who need it.
- •Model Catalog (Serverless): CoreWeave-managed catalog of popular models — GLM 5.2, Kimi K2.6/K2.7, DeepSeek R1/V3, Llama 3.x, Qwen — plus the option to deploy any catalog model or bring your own weights via Dedicated. MLPerf records: CoreWeave trained DeepSeek V3 in ~2 minutes and leads inference speed/price-performance for Kimi K2.6.
- •Training + Inference Loop: Unlike pure inference relays, CoreWeave lets you train or fine-tune on the same GPU fleet (CKS, SUNK, Mission Control observability) and then serve via OpenAI-compatible endpoints — closing the training-to-inference gap for autonomous agent improvement.
- •Sandboxes / RL: CoreWeave Sandboxes (launched 2026) accelerate reinforcement-learning agent tool-use and model evaluation; a serverless RL capability was launched for building reliable RL reward systems. Good for teams that iterate agents, not just call static endpoints.
- •Storage & Data: AI Object Storage (2 GB/s per GPU throughput), distributed file storage, and dedicated VAST storage. The 0EM (Zero Egress Migration) program lets you migrate into CoreWeave with no egress fees and no lock-in.
- •Observability: Mission Control provides fleet lifecycle controller, node lifecycle controller, observability, security, and audit visibility — the operating standard for AI at scale. Tensorizer accelerates model loading for inference.
- •Rate Limits & Enterprise: Capacity through On-Demand, Spot, Capacity Plans, or reserved nodes. High-demand GPUs (e.g. GB200) may require waitlisting or reservations. Enterprise agreements with OpenAI, Anthropic, Meta, and hedge funds like Jane Street indicate strong capacity scale.
🧠 GPU-Native Training-to-Inference Capability
CoreWeave's defining capability is that it is a GPU-native cloud rather than a pure inference API relay. A team can rent an NVIDIA H100/H200/B200 instance on demand (or reserve capacity), train or fine-tune a model on CoreWeave Kubernetes Service (CKS) with Mission Control observability, then serve that same model through an OpenAI-compatible inference endpoint — Serverless (managed catalog) or Dedicated (bring-your-own weights). This closes the training-to-inference loop that pure relays like OpenRouter or Portkey cannot offer, because the serving happens on the same GPU fleet where the weights were built. CoreWeave leads MLPerf both on training (DeepSeek V3 in ~2 minutes) and inference price-performance (Kimi K2.6), and was first to deploy NVIDIA Vera Rubin NVL72 and GB300 NVL72. For API consumers, the OpenAI-compatible surface means you can prototype with a hosted catalog model and later swap to a Dedicated endpoint running your own fine-tune with zero client changes. The tradeoff: this power is aimed at engineers who can operate Kubernetes and reason about GPU capacity, not at a one-line drop-in replacement for a frontier API.
🌐 Regional Availability & Latency
CoreWeave is a US-headquartered AI cloud (Roseland, New Jersey) with data centers concentrated in the United States. In 2026 it expanded internationally: two UK data centers are operational, it announced expansion into Indonesia as its first Asia-Pacific node, and Sweden capacity via a Conapto partnership. There is no mainland China direct endpoint; the Inference management API and console are served from US-fronted regions. For China-based production deployments, expect to use a proxy or a regional provider with local presence. Latency wise, CoreWeave is optimized for North American and European user bases: cross-Pacific requests from Asia to US regions typically add 150-250ms of first-byte latency versus ~50-80ms from a regional provider like Alibaba Bailian, ByteDance Volcano Engine, or Cloudflare Workers AI. For teams serving APAC users, the Indonesia expansion (2026) is the first step toward reducing that distance, but mainland China latency remains a factor. Choose CoreWeave for training-to-inference workloads anchored in the US/EU; pick a regional provider for latency-sensitive Asia-facing traffic.