Clockwork.io Raises $31M to Keep AI Jobs Running Through GPU Failures
Clockwork.io raised $31 million for software that keeps large AI training, reinforcement learning and inference jobs running when GPUs, links or servers fail, bringing its total funding to $73 million, according to the company’s October 5, 2026 release. LinkedIn says the software prevents tens of thousands of GPU-hours of downtime a month across its AI fleet.
The numbers on the wire
- $31 million new round; $73 million raised to date. No valuation or series label disclosed.
- Co-leads: Premji Invest, Wing Venture Capital and Seligman Ventures; existing investors NEA and e& Capital participated.
- LinkedIn runs Clockwork’s LinkPass network fault tolerance across its AI infrastructure fleet. Before that, LinkedIn says, one InfiniBand NIC flap could pull an 8-GPU server out of service, and a switch-port flap could drain a second, for 16 GPUs idled.
- SemiAnalysis founder Dylan Patel is quoted saying that in its ClusterMAX ratings, TorchPass cut training goodput loss from 14% to under 3% for a gold-rated neocloud.
- Clockwork says its new multi-node platform snapshots capture a running distributed training job in under 20 seconds with no training-code changes (company site).
- The problem it is selling against: Meta reported unexpected interruptions about once every three hours over a 54-day Llama 3 run on 16,384 GPUs. Clockwork says reloading a checkpoint can take up to 90 minutes while healthy GPUs sit idle.
What the product does
Clockwork sits as a layer between the hardware and the workload. LinkPass reroutes traffic around a failed network link so the job never sees the fault. TorchPass moves work off a failing GPU to a healthy one so training continues instead of rolling back to a checkpoint. Both are in production.
The two additions announced today extend TorchPass. Platform snapshots save a whole distributed job across every node so infrastructure teams can restore it after a failure too big to migrate around, without waiting for application owners to add checkpointing. Fast asynchronous application checkpoints run in the background and push updated weights to the inference replicas that generate rollouts in reinforcement learning, so those replicas spend less time idle or working from a stale model.
CEO Suresh Vasudevan called fault tolerance “a goodput multiplier,” meaning more of each paid GPU-hour goes to useful work rather than waiting on recovery or repeating finished steps.
Who is using it
- LinkedIn: LinkPass deployed fleet-wide (SVP and CTO Infrastructure Raghu Hiremagalur quoted).
- Together AI: bringing TorchPass to market as a service on its GPU Clusters; the two plan a live demo at the PyTorch Conference of a multi-node training job running through injected network and GPU failures.
- WhiteFiber (NASDAQ: WYFI): existing customer expanding Clockwork across its GPU-as-a-service clusters, including automated fleet audits before customer acceptance.
What the release leaves out
No valuation, revenue, pricing or customer count. The 14%-to-under-3% goodput figure comes from a SemiAnalysis quote in the company’s own release, not a published benchmark report. LinkedIn’s savings are stated as “tens of thousands” of GPU-hours a month, without an exact figure or fleet size. Some coverage calls the snapshot feature “TorchSnap”; the release itself describes it as TorchPass platform snapshots.
Why it matters
As clusters grow to tens of thousands of GPUs, a single bad optic or NIC can stall an entire job, and every restart burns paid GPU time. Clockwork’s pitch is that resilience belongs in the infrastructure layer that neoclouds and enterprise platform teams control, not in each team’s training code. The Together AI resale deal is the clearest test of whether GPU clouds will package that as a product customers pay for.
This article was ultrathought.
Get breaking news, funding rounds, and analysis delivered to your inbox. Free forever.