Add to Google Preferred Sources
AI data company Mercor has begun purchasing internal datasets from failed or acquired startups, including Slack chat logs, Google Drive files, emails, and Zoom meeting recordings, with individual offers reaching as high as $300,000. AI agent startup Warmly received an acquisition offer from Mercor just eight days after announcing its acquisition by HubSpot, but its founder declined to sell. Bobby Samuels, CEO of data company Protege, said AI labs have essentially exhausted publicly available internet content and now need real human-to-human interaction data. Protege has processed at least $100 million in data transactions this year, a sharp increase from roughly $30 million in the same period last year. This data commands high prices because AI agents need to learn how humans actually complete work, not just the answers themselves. However, privacy and intellectual property concerns make this business fraught with challenges—data must be anonymized before it can be legally traded.
Key Elements
When a startup shuts down or gets acquired, its most valuable asset may no longer be its code, patents, or customer lists—but rather the communication records its employees used every day. As artificial intelligence (AI) evolves from chatbots into autonomous digital workers, AI data companies are aggressively acquiring internal datasets from defunct companies, including Slack chat logs, Google Drive files, emails, and Zoom meeting recordings, with individual offers reaching as high as $300,000 (approximately NT$9.6 million).
AI data company Mercor has been actively targeting failed or acquired startups, seeking to purchase their internal datasets. Mercor is the data supplier behind AI giants such as OpenAI and Anthropic, and its reach has now extended from publicly available internet content into the most private communication trails within enterprises.
In late June, AI agent startup Warmly announced it would be acquired by HubSpot. Just eight days later, Warmly founder and CEO Max Greenwald received an email from Mercor seeking to purchase or license Warmly’s internal data—including Slack chat logs, GitHub records, Asana project management content, Google Drive files, and employee meeting transcripts—with offers reaching up to $300,000.
In the days that followed, Warmly received multiple similar inquiries from Mercor and other AI data companies, but Greenwald ultimately turned them all down.
From the Open Web to Enterprise Internals: The Data Arms Race
AI data companies buying books and startup code is hardly news, but as AI giants accelerate development of autonomous AI agents, the nature of data demand is undergoing a fundamental shift. Public internet content, books, and code are no longer sufficient to support the training needs of next-generation AI models.
Bobby Samuels, CEO of data company Protege, put it bluntly: AI labs have essentially “scraped the entire internet,” and what they need now is real human-to-human interaction. Protege primarily helps companies assess whether their internal data qualifies for sale or licensing. This year alone, the company has processed at least $100 million (approximately NT$3 billion) in data transactions—a more than threefold increase from roughly $30 million (approximately NT$1 billion) in the same period last year.
The following table shows the change in data transaction volume handled by Protege:
| Period | Total Data Transaction Volume |
|---|---|
| Same period in 2025 | Approximately $30 million (about NT$1 billion) |
| 2026 year-to-date | At least $100 million (about NT$3 billion) |
Note: Figures disclosed by Protege CEO Bobby Samuels.
Why Did Chat Logs Suddenly Become Valuable?
When AI transitions from a chatbot that simply answers questions to a digital worker that can genuinely share the workload, it needs far more than just “what the answer is”—it needs to understand “how humans actually get work done.”
For example: how a human employee responds when a customer raises an issue; how an engineer troubleshoots step by step when an IT system fails; how a finance staffer handles an invoice discrepancy; how a salesperson updates customer records in a CRM system. Even the specific clicks and inputs an employee makes when software suddenly throws an error. These real-world work experiences are nearly impossible to find on the public internet.
An IT ticket might record “employee unable to log into website,” followed by Slack chat logs showing engineers discussing the root cause, emails containing the final resolution, and a GitHub commit documenting a related code fix. To a human, this is just routine troubleshooting. But to an AI agent, it is an entire set of highly valuable “on-the-job training”—precisely the material needed to train AI systems capable of autonomous task execution.
The Gray Zone of Privacy and Intellectual Property
This emerging business is far more complicated than it appears on the surface. Companies involved must navigate thorny issues around privacy and intellectual property. Slack conversations may contain customer information, emails may include employee names, meeting transcripts may reveal trade secrets, and data from healthcare companies may even involve patient privacy. All such data must undergo anonymization before it can be legally sold or licensed to AI companies.
Warmly CEO Greenwald’s decision to decline selling internal data reflects the reservations some entrepreneurs hold about such transactions. Even when a company is about to be acquired or has already shut down, internal communication records may still implicate the privacy of employees, customers, and partners.
Yet the rapid expansion of the data trading market signals that the AI industry’s demand for “real work data” is forming an irreversible wave. Protege’s data transaction volume surging from $30 million to $100 million in a single year is the most direct quantitative proof of this trend.
As the AI agent race heats up, the digital legacies of failed companies—those traces of human work scattered across Slack, email, and meeting recordings—are becoming the scarcest raw material for AI training. And how to strike a balance between data demand and privacy protection will be the core challenge this emerging industry must confront going forward.
Once added, BigGo Finance appears first in Google Search Top Stories, so you get the broadest, most up-to-the-minute, and most comprehensive global financial news first.
