Lambda
Listed at https://lambda.ai
💰 Token Pricing
| Type | Price | Note |
|---|---|---|
| Input | GPU instance on-demand per GPU/hr: NVIDIA B200 SXM6 $6.69, H100 SXM $3.99, A100 SXM 80GB $2.79, A100 SXM 40GB $1.99, GH200 $2.29, A100 PCIe $1.99, A6000 $1.09, A10 $1.29, V100 $0.79. 1-Click Cluster reserved per GPU/hr: B200 $9.86 (16 GPU) to $8.87 (256+), H100 $6.16 to $5.54. Superclusters / Private Cloud (4000+ GPUs) via sales. No egress fees. | per million tokens |
| Output | Billed by the second per GPU instance (no idle GPU cost); clusters and reserved capacity carry a minimum commitment (2 weeks to 1 year). Managed orchestration: Managed Kubernetes and Slurm included at no extra per-node fee; Lambda Stack one-line install. No per-token inference pricing — Lambda's Inference API is winding down; self-hosted inference billed as GPU runtime. | per million tokens |
| Cache Read | N/A (GPU cloud — no per-token prompt-cache concept; KV-cache cost is absorbed into GPU-hour billing) | Discounted |
🤖 Supported Models (8)
✨ Pros
- ✓Frontier model training infrastructure: OpenAI's GPT-5 and Google DeepMind's Gemini were trained on Lambda GPU clusters (official customer cases)
- ✓Per-second on-demand GPU instances with no idle GPU cost and no egress fees — H100 SXM $3.99/GPU-hr, B200 SXM6 $6.69/GPU-hr
- ✓1-Click Clusters from 16x to 512x GPUs, self-serve launch in minutes; prepaid 2-week to 1-year terms carry volume discounts (H100 16-GPU $6.16 down to 256+ GPU $5.54/GPU-hr)
- ✓Superclusters / Private Cloud: single-tenant VR200 NVL72 (Rubin) and GB300 NVL72 (Blackwell) clusters, 4000+ GPUs with NVIDIA Quantum-2 InfiniBand
- ✓SOC 2 Type II + ISO 27001/27017/27701/22301 certified with a live Trust Portal; raised $1.5B in Nov 2025 led by Microsoft + NVIDIA, pursuing IPO
⚠️ Cons
- ×Not a per-token model API — you deploy inference yourself on GPU instances with vLLM/SGLang; the official Inference API is winding down (2026)
- ×No mainland China direct endpoint; China production deployments require proxy
- ×1-Click Clusters and Superclusters require a 2-week to 1-year prepaid commitment, unsuitable for bursty / short workloads
- ×Compared to RunPod / Hyperbolic etc., consumer-grade entry GPUs (RTX 4090) are absent — only data-center GPUs (starting at A10 / A6000)
- ×No free tier; per-second billing still requires a credit card + billing-address verification
🎯 Best For
AI labs and enterprises training / fine-tuning large models (OpenAI, DeepMind have used Lambda); teams self-hosting open-model inference on GPU instances (vLLM/SGLang); teams needing 16x-512x GPU clusters with a 2-week to 1-year commitment; enterprises that value compliance (SOC 2/ISO) and no egress fees
💰 Pricing & Plans
| Service | Billing Unit | Price | Notes |
|---|---|---|---|
| On-demand instance — NVIDIA B200 SXM6 | per GPU / hr | $6.69 | 180GB VRAM, 8-GPU host; no egress fees |
| On-demand instance — NVIDIA H100 SXM | per GPU / hr | $3.99 | 80GB VRAM; the most rentable Hopper GPU |
| On-demand instance — NVIDIA A100 SXM 80GB / 40GB | per GPU / hr | $2.79 / $1.99 | Ampere workhorse, 80GB or 40GB variants |
| On-demand instance — NVIDIA GH200 | per GPU / hr | $2.29 | 96GB Grace-Hopper superchip |
| On-demand instance — NVIDIA A100 PCIe / A6000 / A10 / V100 | per GPU / hr | $1.99 / $1.09 / $1.29 / $0.79 | Entry data-center tiers |
| 1-Click Cluster — NVIDIA HGX B200 (16 GPU) | per GPU / hr | $9.86 | Reserved 2 wks-1 yr; volume drops to $8.87 at 256+ |
| 1-Click Cluster — NVIDIA HGX H100 (16 GPU) | per GPU / hr | $6.16 | Drops to $5.54 at 256+ GPUs; prepaid term required |
| Superclusters / Private Cloud | contract | Contact sales | Single-tenant VR200 NVL72 (Rubin) / GB300 NVL72 (Blackwell), 4000+ GPUs, Quantum-2 InfiniBand |
| Managed Kubernetes / Slurm | included | $0 extra | Cluster orchestration at no per-node fee |
| Self-hosted inference (vLLM / SGLang) | per GPU / hr | Same as instance rate | Inference API winding down — deploy open models on GPU instances |
🔧 API & Developer Experience
- •API Style: Lambda is a GPU cloud, not a per-token LLM API. The primary interface is the Cloud console + REST API for launching instances, 1-Click Clusters, and Kubernetes. For model inference you deploy open weights yourself on a GPU instance with vLLM, SGLang, or TensorRT-LLM and expose your own OpenAI-compatible endpoint.
- •Lambda Stack: One-line installer that sets up CUDA, cuDNN, PyTorch, TensorFlow, and NVIDIA drivers on any Lambda system — the fastest path from empty GPU to a working ML environment (images are preinstalled on managed instances).
- •Managed Kubernetes + Slurm: Both orchestration layers are available on 1-Click Clusters with no per-node fee. Teams that run PyTorch DDP / FSDP or Slurm job arrays get the same scheduler they use locally, without standing up the cluster themselves.
- •1-Click Cluster Launch: Self-serve multi-node provisioning from 16x to 512x GPUs in minutes — choose GPU type, node count, storage, and orchestration (Managed K8s, Preinstalled K8s, or Slurm) from the console. No sales call for clusters under the Supercluster tier.
- •By-the-second Billing: On-demand GPU instances bill by the second and can be terminated anytime — no idle-GPU waste. Clusters require a 2-week to 1-year prepaid term, which unlocks the volume pricing.
- •No Egress Fees: Lambda publishes no-egress-fees pricing; data transfer out of the cloud is not metered the way hyperscalers price it. This matters for training runs that move large checkpoints and datasets.
- •Inference API Status: Lambda's hosted Inference API is officially winding down (2026) — the recommended path for inference is to deploy models on GPU instances. Teams that need a managed pay-per-token endpoint should pair Lambda GPUs with vLLM self-hosted or use a separate per-token provider.
🧠 Frontier Training & Self-Hosted Inference Capability
Lambda's core capability is raw NVIDIA GPU capacity for frontier training and self-hosted inference. OpenAI's GPT-5 and DeepMind's Gemini have documented training runs on Lambda clusters — a credibility signal few GPU clouds match. The stack spans Rubin VR200 NVL72, Blackwell GB300 NVL72 / HGX B300 / B200, and Hopper H200/H100, delivered three ways: per-second on-demand instances for prototyping, self-serve 1-Click Clusters (16x-512x GPUs) for distributed training, and single-tenant Superclusters (4000+ GPUs) on Quantum-2 InfiniBand. Because Lambda is infrastructure-first, model choice is unrestricted — deploy any open weight (DeepSeek V4, Qwen 3.x, Llama, GLM, Kimi) with vLLM or SGLang and control quantization, batching, and KV-cache policy yourself, giving full control over inference cost and latency.
🌐 Regional Availability & Latency
Lambda operates data centers in the United States (California and other US regions), serving North American and, increasingly, European teams. There is no mainland China direct endpoint and no announced China region as of August 2026; China-based users need a proxy or relay, and cross-Pacific latency to US regions typically runs 150-250ms first byte versus ~50-80ms for regional providers (Alibaba Bailian, ByteDance Volcano). For training jobs this network cost matters less because GPUs are rented by the hour and data moves in bulk; but for latency-sensitive self-hosted inference serving Chinese end-users, a regional GPU cloud is the better fit. Europe and the Middle East are the current expansion focus. Compliance (SOC 2 Type II, ISO 27001/27017/27701/22301) with a live Trust Portal eases enterprise security review.