Devops tools are making it increasingly easy to monitor data pipeline failures. But in many cases, those alerts are already too late. Take marketing campaigns, for example. Every mistake in data for marketing and sales has a trickle-down effect on revenue.
William Flaiz, founder of CleanSmartLabs, shares one example. One of his B2B clients increased their marketing spend by 40%, yet revenue remained flat. On digging further, he found there was an issue in the data used for the campaign. That glaring mistake cost the company $2 million in losses.
There is a need to catch key data errors before they reach downstream, and AI makes it possible.
In this article, I’ll cover the gaps in the current data pipeline monitoring system and how using AI can catch failures before the damage is done.
The current gaps in data pipeline monitoring
Today’s data pipeline monitoring systems need an overhaul. There are three important ways that current systems fall short.
Slow
Data teams are drowning in alerts, often up to 20 alerts per day. There is no clear visibility of impact of each alert to better prioritize them. The ones that get escalated get immediate attention. And when that happens, everything else takes a back seat. So even if an alert exists, it’s sometimes waiting way too long in the queue. By the time the issue gets noticed, it has already caused harm that could be avoided.
Static
Currently, alerts are set with static thresholds or conditions. Keeping these thresholds or conditions up to date requires regular input from business stakeholders and involvement of data teams. That’s an additional maintenance activity. On top of it, the threshold needs to be temporarily tweaked for some business use cases. For instance, a retail company can expect traffic spikes during the holiday season. This static nature of alerts can lead to false alarms that waste the support team’s time or to missed issues that impact the business.
Scattered
The current system notifies about failures, but without any context. The on-call person then has to check logs, understand the issue, and inform the respective team. With complex workflows, logs are scattered across servers, cloud platforms, log management systems, and scheduling tools. Manually browsing through each takes time. Take Sourcegraph, for example. Their support team was spending from 45 to 90 minutes just to check logs for each issue before they started using AI. The scattered nature of pipeline logs and telemetry data increases the on-call person’s work and delays the resolution of issues.
How AI will shift data pipeline monitoring from reactive to proactive
AI models bring together your scattered historical logs and metrics across platforms and summarize only key information, such as:
- Failures and root causes
- Traffic patterns
- Usage trends
- Recurring issues
- Unusual patterns
The models run on this summarized information and send predictive alerts with possible root cause context on:
- Unexpected data volume, schema changes, and distribution
- Potential resource and CPU utilization issues
- Possible memory leaks
- Increased capacity needs for upcoming traffic spikes
- Unusual upstream delays or failures
You’re no longer boxed in by rigid thresholds or a fixed set of static alerts.
How does early detection of data pipeline failures impact business?
Early detection prevents failures or shortens the time frame between failure and resolution, limiting the spread and impact of bad data. Early detection can:
- Prevent data outages by predictive alerts before failure happens.
- Improve customer experience by reducing unexpected downtime.
- Reduce revenue leakages due to poor quality of data.
The cost saved on each prevented issue depends on the business and the type of data. A New Relic study found that some high-impact outages can cost as much as $2 million per hour. Not all issues can be prevented, but the losses can be minimized with early detection.
Is early detection required for every data pipeline?
Early detection is not required for every pipeline. In fact, it will make matters worse if implemented for every pipeline with rising alert fatigue. Data teams don’t need more alerts; they need timely and critical ones. Increasing the volume of alerts would mean a lot more issues waiting in the queue for attention and delayed resolutions.
A hybrid approach is much better, where you implement early detection for datasets with a high financial impact and rely on regular alerts for datasets whose issues can be fixed as per the SLAs. That’s why it’s necessary to understand SLAs for every dataset before you plan early detection monitoring.
Solve this one bottleneck to start AI observability
Finding AI models to implement and gathering historical data to train them is easy. Even ready-to-use data observability platforms with AI features are available in the market. Where things actually get stuck is understanding the domain data better and using models in right place. That means understanding:
- Which datasets are highly time-sensitive and require predictive monitoring?
- What are the business expectations from the dataset in terms of quality?
- What are the expected SLAs to resolve these issues?
- How to stay up to date with evolving business rules and processes for every dataset?
You can’t build reliable monitoring on data you don’t fully understand. Build appropriate documentation for your datasets and pipelines first. Then implementation becomes much easier.
