How does Databricks power the governed, multicloud data platform behind ChatGPT?
OpenAI’s data platform powers one of the fastest-scaling products in tech history. ChatGPT, the OpenAI API, and the internal tools used by <a href="https://bitcomme.com/thousands-of-north-korean-it-workers-are-infiltrating-corporate-america/” title=”Thousands of North Korean IT workers are infiltrating corporate America”>thousands of OpenAI employees require a foundation that can withstand sudden growth, run governed analytics across multiple business units, and operate across multiple clouds simultaneously. That foundation has grown up around Databricks.
Since 2024, OpenAI has leveraged Databricks in Azure and AWS, with usage data landing in Delta tables governed by Unity Catalog. Databricks Jobs and Data Warehouses transform that data through layered pipelines to serve product analytics, finance, security, and trust, and safety teams. Workloads have ranged from abuse-detection pipelines to privacy-sensitive analytics to company-wide dashboards, with thousands of internal users accessing the platform for ad hoc analysis. The result is one of Databricks’ largest and most sophisticated deployments. “We’re very happy to be a Databricks customer,” said Sam Altman, CEO of OpenAI, in aprior interviewwith Ali, celebrating the launch of our partnership.
