Lambda GPU Cloud Review 2026: H100 Pricing & 1-Click Clusters
Lambda (lambda.ai) describes itself as "The Superintelligence Cloud" — a GPU cloud and AI infrastructure company (formerly branded Lambda Labs / Lambda Cloud) that provides NVIDIA-accelerated compute for training and inference at every scale. It is not a per-token model API, and it is not AWS Lambda (the serverless-functions product); the two are unrelated. Lambda's credibility signal is unmatched among dedicated GPU clouds: OpenAI's GPT-5 and Google DeepMind's Gemini are among the frontier models with documented training runs on Lambda GPU clusters. Its product stack scales three ways — per-second on-demand Instances (single-node B200/H100), self-serve 1-Click Clusters (16x-512x GPUs), and single-tenant Superclusters / Private Cloud (4,000+ GPUs) on NVIDIA Quantum-2 InfiniBand — spanning NVIDIA Rubin-generation (VR200 NVL72), Blackwell (GB300 NVL72, HGX B300, HGX B200), and Hopper (H200, H100) systems. Read the full Lambda provider page for the 4-section breakdown, or compare with RunPod, Crusoe, and Hyperbolic on the providers index.
What is Lambda, and why does the "Superintelligence Cloud" matter in 2026?
Lambda began as a GPU-compute and deep-learning-hardware company, then pivoted to a fully managed NVIDIA GPU cloud ("GPU cloud" or "neocloud") as AI demand exploded. Its positioning, "The Superintelligence Cloud," reflects a deliberate focus: instead of competing with hyperscalers on general-purpose cloud services, Lambda concentrates on dense NVIDIA fleets plus the orchestration layer (Managed Kubernetes, Slurm), the one-line Lambda Stack installer (CUDA, cuDNN, PyTorch, TensorFlow, NVIDIA drivers), and enterprise compliance (SOC 2 Type II, ISO 27001/27017/27701/22301, live Trust Portal). The public GPU pricing page and the llms.txt overview are the primary sources for the data in this review.
Lambda GPU pricing: on-demand instances, 1-Click Clusters, Superclusters
Lambda bills GPU instances per GPU-hour (per-second granularity) with no egress fees. On-demand list prices and 1-Click Cluster reserved rates (verified from the public pricing page, August 2026):
| GPU / Service | On-Demand ($/GPU/hr) | Notes |
|---|---|---|
| NVIDIA B200 SXM6 | $6.69 | 180GB VRAM, 8-GPU host |
| NVIDIA H100 SXM | $3.99 | 80GB VRAM — the most rentable Hopper GPU |
| NVIDIA A100 SXM 80GB / 40GB | $2.79 / $1.99 | Ampere workhorse |
| NVIDIA GH200 | $2.29 | 96GB Grace-Hopper superchip |
| NVIDIA A100 PCIe / A6000 / A10 / V100 | $1.99 / $1.09 / $1.29 / $0.79 | Entry data-center tiers |
| 1-Click Cluster — HGX B200 (16 GPU) | $9.86 | Reserved 2 wks-1 yr; $8.87 at 256+ GPUs |
| 1-Click Cluster — HGX H100 (16 GPU) | $6.16 | $5.54 at 256+ GPUs; prepaid term required |
| Superclusters / Private Cloud | Contact sales | VR200 NVL72 (Rubin) / GB300 NVL72 (Blackwell), 4000+ GPUs, Quantum-2 InfiniBand |
A few notes on pricing. On-demand H100 at $3.99/GPU-hr is competitive with RunPod and Hyperbolic's burst rates, and Lambda's "no egress fees" policy is a real cost lever for training runs that move large checkpoints and datasets. 1-Click Cluster and Supercluster rates carry a 2-week-to-1-year prepaid commitment, which unlocks the volume discounts; if you only need bursty capacity, per-second on-demand instances are the cheaper fit. Lambda offers no consumer GPUs (no RTX 4090) — the lineup is data-center-grade only, starting at A10/A6000/V100.
Lambda's API surface: it is an infrastructure cloud, not a model API
The single most important thing to understand about Lambda for an API consumer: it is not a pay-per-token model endpoint. Lambda's hosted Inference API is winding down (2026), and the official guidance is to "continue deploying and scaling models seamlessly on NVIDIA GPU instances." That means the developer experience is operations-first: you launch a GPU instance or cluster from the Cloud console or REST API, then deploy open weights yourself with vLLM, SGLang, or TensorRT-LLM and expose your own OpenAI-compatible endpoint. The trade-off is control: you own quantization, batching, KV-cache policy, and, crucially, cost and latency. The upside is the widest possible model freedom — DeepSeek V4, Qwen 3.x, Llama, GLM, Kimi, or a fine-tune of any of them — with none of the per-token markups of a managed API.
- Lambda Stack: one-line installer for CUDA, cuDNN, PyTorch, TensorFlow, and NVIDIA drivers — images are preinstalled on managed instances, so you go from empty GPU to working ML environment in minutes.
- Managed Kubernetes + Slurm: both orchestration layers ship on 1-Click Clusters with no per-node fee — teams running PyTorch DDP/FSDP or Slurm job arrays get the same scheduler they use locally.
- 1-Click Cluster launch: self-serve 16x-512x GPU provisioning in minutes from the console; no sales call below the Supercluster tier.
- By-the-second billing: on-demand instances bill per second and terminate anytime; clusters carry the 2-week-to-1-year prepaid term.
- No egress fees: outbound data transfer is not metered the way hyperscalers price it.
Why Lambda matters for API consumers: the frontier-training credential and infrastructure durability
For an engineer choosing AI infrastructure in 2026, Lambda matters for two reasons. First, credibility: OpenAI's GPT-5 and Google DeepMind's Gemini are documented to have trained on Lambda GPU clusters — a signal few dedicated GPU clouds can match, and a strong hint that if you need guaranteed NVIDIA capacity for large scale, Lambda has the fleet. Second, durability: Lambda raised $1.5 billion in November 2025 in a round anchored by Microsoft and NVIDIA, and is reportedly pursuing an IPO (a $350M pre-IPO round was discussed in January 2026). That is the capital profile of infrastructure you can trust with production workloads. The compliance stack (SOC 2 Type II, ISO 27001/27017/27701/22301, live Trust Portal) further de-risks the enterprise security review. The caveat is that Lambda's value is realized almost entirely through self-managed GPU usage — you trade a managed API's convenience for raw control. Teams that want a managed OpenAI-compatible endpoint should pair Lambda GPUs with vLLM self-hosting or choose a per-token provider instead.
Lambda vs. RunPod, Crusoe, and Hyperbolic
These three competitors bracket the GPU-cloud spectrum differently, and each maps to a different Lambda use case. RunPod is the developer-friendly serverless option: per-second GPU pods, a large marketplace, and consumer-grade cards (RTX 4090) that make it the cheapest choice for bursty inference and experimentation. Crusoe and Nebius sit closest to "GPU cloud with a managed inference API": they rent you GPU compute AND expose OpenAI-compatible per-token endpoints for open models, so you get Lambda-class compute without managing vLLM yourself. Hyperbolic is a decentralized GPU marketplace that aggregates idle GPUs at the lowest entry prices, trading predictable capacity for cost. Lambda's niche is the enterprise training + self-hosted production inference tier: the most blue-chip customer list, $1.5B+ funding, strict compliance, and a clean (if commitment-heavy) path to 4000+ GPU single-tenant clusters.
Who should choose Lambda — and who should not
Choose Lambda if you are (a) training or fine-tuning large models and want guaranteed NVIDIA capacity on a well-capitalized, compliant platform; (b) running self-hosted production inference at scale with vLLM/SGLang and want per-GPU-hour honesty with no egress fees; or (c) an enterprise that needs SOC 2/ISO compliance and the option to scale to 4000+ GPU single-tenant clusters. Skip it if you (a) just want to call an open or frontier model with minimal setup — a managed per-token API (OpenAI, an aggregator, or Crusoe/Nebius) is simpler and cheaper for pure consumption; (b) need cheap, bursty, per-second GPU without a multi-week commitment — RunPod or Hyperbolic fit better and carry consumer GPUs like the RTX 4090; or (c) serve latency-sensitive traffic to mainland China — there is no China endpoint and cross-Pacific latency runs 150-250ms.
Regional availability and latency
Lambda operates data centers in the United States (California and other US regions), serving North American and, increasingly, European teams; Europe and the Middle East are the current international expansion focus. There is no mainland China direct endpoint and no announced China region as of August 2026. For China-based production, you would need a proxy or a regional provider with local presence. For latency-sensitive Asia-facing traffic, cross-Pacific requests to US regions typically add 150-250ms of first-byte delay versus ~50-80ms from a regional provider like Alibaba Bailian, ByteDance Volcano Engine, or Cloudflare Workers AI. For training workloads — where GPUs are rented by the hour and data moves in bulk — this network cost matters far less than for synchronous inference. Lambda is best suited to US/EU-anchored training and self-hosted inference; pick a regional GPU cloud for latency-critical Asia traffic.
Lambda GPU Cloud FAQ
A: Not directly. Lambda has no managed pay-per-token inference endpoint (its Inference API is winding down). You rent a GPU instance and deploy open weights with vLLM/SGLang yourself, then call your own OpenAI-compatible endpoint.
A: On-demand instances by the second (H100 SXM $3.99/GPU-hr, A10 $1.29/GPU-hr) with vLLM self-hosted are the cheapest path for bursty work. 1-Click Cluster prepaid terms (2 weeks-1 year) lower per-GPU cost only if you sustain high utilization.
A: No. Lambda publishes no-egress-fees pricing, which matters for training runs that move large checkpoints and datasets out of the cloud.
A: On demand: B200 SXM6, H100 SXM, A100 SXM 80GB/40GB, GH200, A100 PCIe, A6000, A10, V100. Clusters: HGX B200 and HGX H100 up to 512 GPUs. Superclusters: VR200 NVL72 (Rubin) and GB300 NVL72 (Blackwell), 4000+ GPUs. No consumer cards (RTX 4090 absent).
A: Yes. Rent H100/B200 instances, fine-tune with your framework of choice, and either serve on the same instance or snapshot the image and launch a serving node — full control, no per-token fine-tuning markup.
A: No. Lambda (lambda.ai) is a GPU cloud / AI infrastructure company formerly branded Lambda Labs; AWS Lambda is Amazon's serverless-functions product. They are unrelated companies and products.