Skip to main content

Reliable output needs a reliable upstream

Pipelines can only be as reliable as their inputs
·3 mins

Hello 👋

This week I write about how, when our data output needs to be reliable, so does our input.

There’s also links to posts on using data contracts to build a shared language, federated data architecture with DuckDB, and Bitol (ODCS) becoming a graduate project at LF AI & Data.


Reliable output needs a reliable upstream

Years ago at GoCardless, we started feeding our data into ML models to power features for our customers.

And we quickly realised the data wasn’t reliable enough to do that.

That realisation is where my work on data contracts started. We needed to improve the reliability of the data we were consuming, and we decided the best way to achieve that would be to create an interface between the upstream services producing the data and the pipelines consuming it.

I think a lot more teams are having that same realisation now.

For a long time, it was accepted that if the data team’s pipelines broke for a day, or a few days, or even a week, it didn’t really matter. The numbers were for finance and the business, and they could wait.

I’m not sure that was ever really true! Certainly as someone on that team, being stuck permanently firefighting is not much fun… But I can see why, from the outside, it looked like an acceptable trade-off.

But it isn’t any more.

More and more, the data we produce doesn’t just feed a dashboard someone checks once a week. It powers a product feature, either directly or via an AI agent.

So, we need to invest in the reliability of the data.

That includes us upskilling and building more reliable pipelines, improving observability, and so on.

But there’s no point doing that if the data we depend on is also unreliable, which is why we have to improve data reliability at the source.

And we do that with data contracts.

I talk about all this on The Data Engineering Show. We also get into:

  • Moving from a centralised data team to a self-serve platform
  • Designing for different user personas
  • Building the new world while still supporting the old
  • Using agents to automate support requests
  • Rethinking data governance in an AI-driven world

Check it out and let me know what you think!


Data Contracts: Bringing Data and the Business Closer Together by Ian Røpke (on LinkedIn)

Nice post on using data contracts to build a shared language and understanding.

Why we rebuilt our data warehouse and how it unlocks self-driving products by Eric Duong (Posthog)

Interesting federated architecture using DuckDB. Interesting site too…

Three Years, 59 RFCs, One Vote: Bitol Graduates by Jean-Georges Perrin

Bitol, the the Open Data Contract Standard, is now a graduated project of the LF AI & Data Foundation.


Being punny 😅

Roman numeral puns are great and I for one will continue to make them.


Upcoming events


Thanks! If you’d like to support my work…

Thanks for reading this weeks newsletter — always appreciated!

If you’d like to support my work consider buying my book, Driving Data Quality with Data Contracts, or if you have it already please leave a review on Amazon.

Enjoy your weekend.

Andrew


Want great, practical advice on implementing data mesh, data products and data contracts?

In my weekly newsletter I share with you an original post and links to what's new and cool in the world of data mesh, data products, and data contracts.

I also include a little pun, because why not? 😅

    Newsletter

    (Don’t worry—I hate spam, too, and I’ll NEVER share your email address with anyone!)


    Andrew Jones
    Author
    Andrew Jones
    I build data platforms that reduce risk and drive revenue.