Synthetic datais used to replace poor-quality, fragmented, sensitive or nonexistent data for analytics or to train AI. Some organizations also use it to createsynthetic panelsto generate feedback. Using synthetically generated data can be faster and more cost-effective than real enterprise data. However, there are concerns over bias and inaccuracies.
Kathy Langeis an IDC research director for AI, Data and automation software. Her main focus is on unified AI platform tools and technology. Lange previously worked at SAS, where she led analytics consulting, analytics and text analytics presales and competitive intelligence teams. In this conversation, Lange discusses synthetic data with No Jitter. Lange said she views synthetic data as a supplementary tool for real-world data rather than a replacement.
No Jitter: How mature is synthetic data as an alternative for fine-tuning communications and routing?
Lange:If by fine-tuning communications and routing you mean using synthetic data to optimize customer messaging and train customer-service workflows, the technology is becoming better and more practical. However, if possible, I would still view it as complementary to real customer data, since real customer conversations provide behaviors and actions that would be difficult to generate with synthetic data.
No Jitter: How mature are synthetic customer-service models when organizations cannot use real customer or employee data?
Lange:Synthetic customer-service models are still relatively immature when they rely entirely on synthetic data. Synthetic customer-service models are increasingly viable for training, testing, and simulation, especially where privacy or regulatory requirements limit access to real data. That said, they are generally better at representing common scenarios than the emotional, ambiguous, or unusual interactions that often come within real customer examples.
No Jitter: Which AI use cases are best suited to synthetic data?
Lange:In my view, the best use cases for synthetic data are those where the cost, privacy risks, or practical challenges of obtaining large amounts of real data outweigh the need for perfect accuracy and when the outcomes are driving directional insights, rather than high-stakes business decisions.
No Jitter: Where does synthetic data still fail to reproduce the ambiguity, edge cases, behavioral variation, and bias found in real customer interactions?
Lange:Synthetic data is generally strongest at reproducing common patterns and averages. It is less effective at capturing unexpected behavior, emerging trends, or emotional responses.
No Jitter: What validation practices should organizations have before trusting an AI model that is trained on synthetic data?
Lange:Models trained on synthetic data should be benchmarked against actual customer interactions or other observed outcomes, such as survey responses, click-through rates, support interactions, or purchasing behavior.
