IBM Research and CoreWeave have moved beyond a straightforward capacity arrangement into joint engineering of identity management and workload controls, according to Brian Belgodere, a senior technical staff member at IBM. The two companies refined the implementation through several rounds of feedback, with CoreWeave regularly bringing proposals to IBM for review. Speaking with theCUBE Research's Dave Vellante and John Furrier at the Fully Connected event, Belgodere described how IBM's research infrastructure requirements helped shape that collaboration.
The work reflects a shift in what research systems must handle. Reinforcement learning introduces a task-execution stage in the middle of model training, where a checkpoint is loaded into inference, given a task and measured. Belgodere characterized that stage as a testing phase. Systems built for intensive computation now have to accommodate workloads that interact with tools, storage and other services, making agent workload isolation a practical infrastructure challenge.
IBM's work on its Granite family of models demanded substantial computing resources. Belgodere said the cooling and power requirements of a subsequent hardware generation drove the company's decision to work with CoreWeave. IBM built a large H100 cluster on its own, securing the space and handling the project from start to finish, which he described as a huge task.
Much of the IBM Research cluster is single-tenant, with IBM storage deployed inside CoreWeave and additional capacity available within cost and security parameters. The collaboration also includes CoreWeave Sandboxes, which support isolated execution on dedicated infrastructure or through a managed serverless runtime. Those options let researchers decide where agent code runs and which resources it can reach.
Belgodere warned that architecture choices made early are costly to reverse. He said many teams underestimate that cost, and that a poor decision can lead either to buying far more networking infrastructure than needed or to refitting everything later.
IBM measures how security controls affect performance against benchmark results, then uses those findings in discussions with security teams about tradeoffs. Workload isolation sits alongside enterprise identity integration within that broader security architecture. Belgodere described the overall challenge as a supply chain problem extending from hardware and firmware through kernel levels and code to data provenance, and on into agent images, which he called an absolute provenance problem.
More software news from TechManNews.



