Clockwork Systems Inc. has raised $31 million in a new funding round and introduced TorchSnap, a feature aimed at reducing wasted compute in artificial intelligence clusters. The data center infrastructure startup said the round was co-led by Seligman Ventures, Wing Ventures and Premji Invest, with existing backers New Enterprise Associates and e& Capital returning. The new capital brings Clockwork's total raised to date to $73 million.
The funding arrives as organizations running large distributed AI workloads shift their attention from acquiring raw graphics processing unit capacity to getting more out of the clusters they already operate. Clockwork points to the experience of Meta Platforms Inc., which reported hardware issues every three hours on average during the 54-day training run of its Llama 3 model across a cluster of 16,384 GPUs. When such failures occur, teams typically reload saved progress from a snapshot, a recovery process the startup said can take as long as 90 minutes, leaving thousands of healthy GPUs idle and forcing clusters to repeat completed work.
Clockwork addresses that problem with a programmable software layer positioned between GPUs and running AI workloads. The layer synchronizes GPU clusters and provides nanosecond-accurate telemetry intended to identify failures before they trigger a full cluster restart. Chief Executive Suresh Vasudevan said GPU failures are unavoidable at this scale, but teams should not have to lose hours of work. He described fault tolerance as a goodput multiplier that keeps GPUs doing useful work rather than waiting for recovery or repeating prior work, and said the software was built alongside enterprises and cloud providers operating some of the largest GPU fleets.
TorchSnap adds a third protective layer alongside Clockwork's existing LinkPass network failover tool and TorchPass GPU migration software. According to Vasudevan, it captures multinode snapshots of distributed AI inference workloads across each node in a cluster without requiring developer code modifications, allowing running jobs to restart where they left off. Teams can also add checkpointing logic at the application level to reduce lost progress and repeated computation after a failure.
SemiAnalysis analyst Dylan Patel said fault tolerance is now urgently needed for AI inference workloads, not just training. He said Clockwork keeps replicas serving through link flaps and network failures, and that its fast checkpoints speed weight transfer back into the rollout fleet so neither direction stalls the run.
Adoption has grown across public cloud infrastructure providers, neoclouds and enterprise fleets since Clockwork's previous round just over a year ago. Microsoft Corp.'s LinkedIn has deployed LinkPass across its entire GPU fleet, eliminating thousands of GPU-hours of downtime monthly by rerouting AI traffic around optical link and switch failures. Together AI Inc. offers TorchPass as a service on its GPU clusters, and neocloud provider WhiteFiber Corp. uses Clockwork's technology to audit and validate cluster reliability before new AI workloads enter production.
NEA Venture Partner Greg Papadopoulos said GPUs are defined by their unreliability, and that stopping an entire job to reload a checkpoint made sense for supercomputers thirty years ago but does not at this scale. He said the firm backed the team early and invested again because this layer is becoming part of what an AI cluster is.
More company and startup news from TechManNews.





