The company’s funding round, co-led by Premji Invest, Wing Venture Capital, and Seligman Ventures, brings its total capital to $73 million. Clockwork.io focuses on a critical pain point: the high rate of failure in massive GPU clusters, where a single malfunctioning component can trigger a cascade of downtime. By acting as a software layer between hardware and AI workloads, the system ensures training and inference tasks continue even when individual GPUs or network links fail.
New updates to the company's TorchPass solution allow for platform-level snapshots and fast background checkpoints. These features enable teams to save the state of a distributed job without requiring modifications to the original training code. LinkedIn, which has already implemented the software’s LinkPass technology, reports saving tens of thousands of GPU-hours each month by eliminating the need to restart jobs after minor network disruptions. This shift toward treating infrastructure failure as an expected condition rather than an exception is becoming a core requirement for hyperscalers and cloud providers managing intensive AI deployments.





Comments (0)
No comments yet. Be the first!