AI workflows inherit the weaknesses in your data. Learn how to find gaps across your data pipeline and fix them at the source.
Table of Contents
Table of Contents
Spy on Any Website
Get traffic data and keyword intel on competitors instantly.
Most marketing teams I talk to use AI on their marketing data to build campaigns, write copy, segment audiences, or spot customers who might leave.
When AI works well, it feels like the future is here. When it doesn’t, it can be frustrating. If AI writes copy, you can often tell when the tone is off. On the other hand, if you’re asking AI to choose your audiences, mistakes might slip by until a customer complains or someone checks the results. Sometimes, you only notice when someone asks why the numbers don’t add up.
Anyone who started in data-driven marketing learned early on that your results are only as good as the systems behind them. AI makes it riskier to ignore this rule, since models give answers quickly and confidently whether they’re right or not.
Almost across the board — particularly in consumer-focused marketing — we need better data systems. To learn how to get there, I got in touch with an old colleague who has more experience than anyone else I know in building data systems: Subu Desaraju.
He and I worked together at Tempur-Pedic, building data warehouses and CRM strategies. Before that, he helped create one of the first people-based marketing platforms at Digitas, worked at WPP, and led global data and analytics at MRM.
Now, Desaraju leads commercial and operations at iceDQ, a data reliability platform. I don’t know many people who have spent as much time bridging the gap between what marketers think their data says and what it really means.
We had an hour-long conversation that covered:
- The industries most in need of better data systems.
- The impact poor data has on your customers.
- Red flags that hint at poor data.
- The four-stage process from data collection to consumption.
- Two frameworks to build a better data system.
Desaraju summed up the concern this way: “While everyone’s rushing into the AI game, I really fear for the output of that process without the right foundations in terms of reliable data.”
10X your SEO with Semrush for Enterprise.
The world’s most powerful SEO platform, purpose-built for Enterprise.
Request demo
Which industries most critically need systems?
I started my chat with Desaraju by asking him directly: Which industries have reliable data and actually get this right?
His answer focused more on consequences than on technical details. Regulated industries are ahead, especially those that face financial penalties for getting data wrong. Financial services pay close attention because authorities like FINRA will step in if a bank’s trade and position data don’t match. Healthcare faces similar risks. Privacy rules add another layer for both. When the cost of mistakes is clear and steep, companies become more disciplined.
Perhaps because consumer and retail marketing don’t face FINRA or HIPAA consequences, the approach to data systems in those industries tends to be more lax. If our segments are wrong, no one fines us — but there are still consequences in the form of misdirection. If a campaign underperforms, we question the creative. If a segment acts strangely, we blame the algorithm. If a model gives a weird result, we think we need a better prompt. Or a dashboard shows a misleading performance result.
Desaraju is worried that we’re making this problem worse. As more people rush into AI, he’s concerned about what happens when the basics aren’t solid. Many organizations care about data quality, but not many realize how hard it is to create data that’s truly clean, consistent, and ready for business use.
What your customer experiences when your data isn’t clean
Data hygiene might sound like an IT issue, but your customers feel it in a much more personal way.
They notice it when a welcome offer goes to someone who’s been a member for 10 years. They see it when they get the same email three times, or when a brand recommends something they just bought. To them, these aren’t just technical glitches. They’re signs that you don’t really know them, and that’s a hard impression to fix.
This should be on the CMO’s agenda, not buried in a technology roadmap. Every claim we make about knowing our customers depends on it. Every AI workflow we build on top of customer data inherits it, often without us realizing.
Data red flags: What to look for
When I ask Desaraju about the complaint he hears that most strongly represents a data red flag, his first answer is unpredictability.
He told me about a client last year who couldn’t match its campaign execution reports with the audiences it had chosen. The team thought they had emailed three million customers, but they couldn’t confirm which message went to which person, or whether the delivery report matched the original segment definitions.
If you run CRM, this should catch your attention. The segmentation and strategy were solid, but the team couldn’t prove that what they sent matched what they planned. That’s where personalization breaks down, and where a model trained on those records can learn the wrong lessons and repeat them at scale.
Two other questions come up a lot in his conversations: How do I connect and coordinate work across all these different systems? And how do I show the return on what I’ve already invested? Both questions get tougher as more platforms are added.
The four stages of data systems and how to address each
According to Desaraju, most companies check their data at the wrong stage.
Organizations often measure quality at what Desaraju calls the consumption layer. This is where analytics models use the data, BI reports read it, and business teams finally interact with it. Checking the output at this point is too late. By the time a problem shows up on a dashboard, the data has moved through every system, accumulating faults along the way.
Desaraju suggests focusing on thethe system? As it moves, gets transformed, joined, and segmented, is each step working as it should?
He points out that this problem involves people, processes, and tools. But more importantly, he sees it as a shift in behavior. Checks and controls only get built when an organization changes how it thinks about its data.
Think of a water system, not a warehouse. Picture data as water flowing through a city, with four steps from start to consumption.
- The lake is your reservoir, holding almost everything. This is the point where data enters your system.
- From there, it moves into a distribution system to get filtered and treated. This step is where you segment the data, clean it up, add rules, etc.
- Next, it travels through pipelines to households. In this analogy, the households are your business users and their applications. The data changes as it moves from platform to platform.
- Last: The data is consumed through reports and dashboards, which are usually where marketers in consumer-based industries decide to do QA.
That’s the problem: if you want clean water at the tap, you don’t just test the tap. You monitor the entire system, checking pressure and volume at every step.
“If we want to consume clean, consistent, accurate data, we have to look at the value chain of the data and ensure that there are checks and controls in place at each point in that pipeline and in that distribution mechanism. Is the pressure right? Is the volume right?” Desaraju said.
For nearly 20 years, our industry has tried to create a single view of the customer. His view on why customer data platforms haven’t solved this is important. Technology wasn’t the real barrier. The changes needed in people, processes, and platforms never fully happened in most companies.
A true CDP depends on gathering customer data, setting clear standards for how data should be received, and using an engineering approach to combine identities and attributes from different business units, channels, and data sources. Without that discipline, you just end up with an expensive system that still has the same old problems.
Two frameworks to improve your data
Desaraju laid out two frameworks that he’s used sequentially to help companies clean up their data.
Framework 1: Trace a campaign backward
Most marketing directors, VPs, and CMOs I know don’t control the data organization. But you can still find out what shape your data is in, and you can do it in a single working session with your current team.
Pick your last major campaign. Walk through it in reverse, from what the customer received back to where the data started. Ask these five questions in order.
- What did we actually send? Pull the execution report. How many people got each version of each message?
- Does that match the audience file we gave to the platform? Compare the counts and makeup. Is it close enough? Is it exact?
- Does that file match the segment definition in the brief? This is a question that tends to highlight issues. If your answer is “probably,” you’ve found your first gap.
- What happened to the data between the source and that file? List every transformation in between. Joins, deduplication, suppression rules, identity matching, exclusions. For each one, name who validated it and how.
- Where did the data come from, and how do we know it arrived? Which systems fed this campaign? When did they last deliver? Who would notice if one of them didn’t?
Then fill in three columns for every step you just reviewed: what check exists today, who owns that check, and what happens when it fails. Most teams find that the middle column is mostly empty and the third column is completely empty.
The blanks are your plan. You don’t need a budget or permission to do this exercise. You just need a conference room and two hours.
Framework 2: Build your solution across people, process, and tools
Desaraju uses this framework to turn the gaps you just identified into a plan an organization can fund and follow. It focuses on three areas: people, process, and tools.
Start with a question that often gets missed in big companies: Who owns data quality, and how do we check it?
Having a chief data and analytics officer helps, but silos still exist between application, development, data, and business teams. Developers build, but quality assurance isn’t always thorough or ongoing. The business side assumes the data is accurate and integrated, but requirements are often not documented clearly enough for anyone to verify.
Two steps help close that gap.
- Write your requirements in a way that someone can test them, and update them as things change.
- Invest in people who can bridge the gap between business and engineering.
Desaraju looks for problem solvers who go beyond their job descriptions, and he encourages training that helps data people understand the business and business people understand the systems. The old model where someone is only a developer, only a QA specialist, or only an analyst doesn’t work well anymore.
Here, Desaraju borrows from manufacturing, and the sequence is the most useful part of the whole conversation.
No one starts an assembly line without testing the equipment first. No one runs a plant without monitoring each unit as it moves down the line. No one ships without checking the finished product. Three stages need to be addressed, in that order.
Test before you deploy. Nothing should go into production untested. For a marketing team, that means new audience logic, a new suppression rule, a new integration between your CRM and an ad platform, a new data feed, and any AI workflow that reads customer records. The test isn’t just whether it ran — it’s whether it produced what the requirement said it would.
Monitor while it runs. Set up continuous checks on the data as it moves through the pipeline, so problems surface there rather than on a dashboard. Volume, timing, and completeness are good places to start. If a feed fails to arrive on Tuesday, someone should know on Tuesday.
Observe the output. Check the finished data products in the consumption layer for accuracy, completeness, and timeliness, and look for anomalies at scale. This stage still matters. But as Desaraju says, if the first two stages are working, about 99% of defects never reach this point.
Toyota proved this thinking decades ago, and automotive quality in the 1980s is a fair picture of where much enterprise data sits today. He cites Gartner research estimating that poor data quality costs organizations an average of $15 million per year.
“You never switch on the assembly line without testing that all of the instrumentation works correctly,” Desaraju said.
We’ve all seen those crowded slides full of martech logos. Desaraju believes that fragmentation comes at a real cost: the more scattered your tools, the more complex things get and the less insight you gain. Business units built their own stacks for good reasons, but now the industry is trying to simplify.
His advice is to reframe the question. Stop asking which tool you need. Instead, ask these three questions in order.
- What business outcome am I responsible for? Is it retention, acquisition efficiency, lifetime value, or something more specific to this quarter?
- Where does my data value chain break down on the way to that outcome? You already made this list in the first framework.
- Does this tool close those specific gaps? Evaluate it based on your gaps, not on a feature list or someone else’s use case.
Choosing tools in that order usually works even after future reorganizations, because the reason was never just the tool.
“The more fragmented the ecosystem becomes, the higher the complexity and lower the intelligence,” Desaraju said.
Fix the foundation before scaling AI
In two years, the marketing teams getting real value from AI will have one thing in common: someone made the practical choice to fix the data pipelines early, before it was popular. Their models learned the real story about their customers, not just a likely version.
Contributing authors are invited to create content for MarTech and are chosen for their expertise and contribution to the martech community. Our contributors work under the oversight of the editorial staff and contributions are checked for quality and relevance to our readers. MarTech is owned by Semrush. Contributor was not asked to make any direct or indirect mentions of Semrush. The opinions they express are their own.
Google’s “preferred sources” feature allows users to customize their search results by selecting news outlets they want to see more often in the “Top Stories” section.
