PRODUCT • October 5, 2026 • 4 min read

Iterate.ai Lifeboat GA: 2–6× Agent Sessions per GPU

Thumbnail for: Iterate.ai Lifeboat GA: 2-6x Agent Sessions per GPU

Iterate.ai made Lifeboat generally available — a self-hosted LLM inference engine with confidential computing built in that the company says runs 2–6× more concurrent AI agent sessions per GPU, according to its product page and October 5, 2026 SiliconANGLE coverage of the launch.

The numbers on the wire

  • Company claim: 2–6× more concurrent agent sessions per GPU; product page also cites roughly 2× concurrent capacity without weight quantization.
  • Company benchmark on one NVIDIA RTX PRO 6000 Blackwell with Qwen 30B-A3B: 2,048 concurrent sessions at 100% success vs 1,024 baseline; 8,714 tok/s vs 4,965 at 2,048 concurrency; KV cache 568K vs 284K tokens (2.0×).
  • Under 128 sessions × 18K-token requests: p99 TTFT 1.5s vs baseline 189s (company figures).
  • Pricing: free Developer (≤2 servers, single node, non-commercial); Standard $49.99/mo; Confidential Computing $499.99/mo; paid tiers include 7-day trial (licensing docs).
  • Confidential Computing attests Intel TDX or AMD SEV-SNP CPUs plus NVIDIA CC mode on H100/H200/B100/B200/GB200/GB300 / RTX PRO 6000 Blackwell Server Edition.

What Lifeboat does

Iterate frames agent workloads as dozens of model calls with growing KV cache — enough to stall standard engines at a handful of long-context sessions. Lifeboat adds fair scheduling and admission control, KV-cache compression (TurboQuant / H2O-style) while keeping weights full BF16/precision, MoE expert loading, and per-session “security capsules” with filtering and token budgets. An OpenAI-compatible API, web control plane, RBAC, audit logs and multi-node orchestration are included; ROCm builds cover AMD Instinct alongside CUDA.

Confidential Computing edition refuses to serve until hardware attestation passes; model weights stay sealed in the TEE. CEO Jon Nordmark and CTO Brian Sathianathan are quoted in SiliconANGLE on keeping agent data inside enterprise walls before buying more GPUs.

What the launch leaves out

No independent third-party benchmark. The 2–6× range and all RTX PRO 6000 figures are company-run. Early-access “thousands of downloads” is company-reported without a public telemetry dump. GPU confidential computing has no AMD GPU equivalent per Iterate’s own docs.

Why it matters

Agent concurrency is becoming a GPU-capacity tax. A generally available engine that sells confidential compute plus measured 2× session density on mid-range enterprise GPUs is a direct product answer to that tax — if the company benchmarks hold up outside Iterate’s lab.

This article was ultrathought.

Stay ahead of AI

Get breaking news, funding rounds, and analysis delivered to your inbox. Free forever.

Related stories