The neocloud market is moving past its origins as a stopgap for scarce graphics processing units. AI-native startups now choose their infrastructure on latency, burst capacity and openness, not just chip availability.
That shift is playing out at CoreWeave Inc., which is expanding beyond GPU compute into networking, storage and software as inference demand grows. At the same time, LlamaIndex Inc. has evolved from an open-ilder that rents its compute rather than owning it, according to Jerry Liu (pictured, right), co-founder and chief executive officer of LlamaIndex
“We’re effectively a specialized AI lab right now that’s purely focused on building models for document parsing and extraction. We post-train open-weight models, we gather our own datasets and we make it really, really good at analyzing and reading documents to basically extract that data,” Liu said. “We care a lot about making sure that we can actually tailor everything we’re doing at the Pareto frontier of performance, cost, and latency for our customers.”
Liu and Lukas Biewald (left), senior vice president of AI initiatives at CoreWeave, spoke with theCUBE’s John Furrier and Dave Vellante at Fully Connected, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed long-running agents, governance and why AI-native startups are turning to specialized clouds for inference-heavy workloads. (* Disclosure below.)
Why bursty AI workloads favor the neocloud model
LlamaIndex’s compute footprint barely existed a year ago. Today its workload runs about 75% inference and 25% training, and it processes millions of document pages per day for finance, legal and insurance customers whose paperwork arrives in bursts so guaranteed capacity matters more than hardware ownership
“We serve a lot of different customers at extremely persistent and also spiky workloads,” ,” Liu said. “We really, really need to make sure that we have the right capacity to serve our customers without getting throttled.”
CoreWeave is betting that capacity alone is not the differentiator. Biewald joined the company through its acquisition of Weights & Biases, the AI observability startup he co-founded, and CoreWeave used the event to launch CoreWeave Forge, a development layer that runs training, inference, evaluation and agent development in one connected environment. Coming from software, he initially questioned how much chip configuration could really matter, Biewald noted.
“I’ll tell you, the answer is ‘massive difference,’” Biewald said. “I’m talking orders of magnitude difference depending on how you do the networking for the chips [and] how you do the power distribution.”
Openness also separates CoreWeave from the hyperscalers Where providers such as Amazon Web Services Inc. lean on proprietary application programming interfaces that make workloads hard to move, CoreWeave follows the standard networking protocols recommended by Nvidia Corp., which brings broader open- its portfolio much as AWS did in its early days, even as the company bristles at the neocloud label
“CoreWeave knows that everyone is coming from a different cloud,” Biewald said. “Everyone’s going to host their web service on AWS or GCP, not on CoreWeave. CoreWeave is okay with that, so CoreWeave plays much more nicely with the other clouds.”
That ecosystem points to a larger change in who gets to build intelligence, Liu noted. Post-training a small open-weight model remains a skill limited to a narrow group of specialists today. Abundant neocloud capacity, combined with fast-improving coding agents, could open that work to far more people.
“Everyone is starting to get really good at defining observability and evals and the right metrics to focus on,” Liu said. “I think there’s going to be a world where we’re basically just going to automate this entire loop and make it accessible to everybody.”
Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of Fully Connected:
(* Disclosure: TheCUBE is a paid media partner for the Fully Connected 2026 event. Neither CoreWeave, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)
Photo: SiliconANGLE
A message from John Furrier, co-founder of SiliconANGLE:
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
- 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
- 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network
Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from ourMarketplace portal page and links:https://siliconangle.com/aws-marketplace/
About SiliconANGLE Media
SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.