Reliable output needs a reliable upstream
Hello 👋
This week I write about how, when our data output needs to be reliable, so does our input.
There’s also links to posts on using data contracts to build a shared language, federated data architecture with DuckDB, and Bitol (ODCS) becoming a graduate project at LF AI & Data.
Reliable output needs a reliable upstream
Years ago at GoCardless, we started feeding our data into ML models to power features for our customers.
And we quickly realised the data wasn’t reliable enough to do that.
That realisation is where my work on data contracts started. We needed to improve the reliability of the data we were consuming, and we decided the best way to achieve that would be to create an interface between the upstream services producing the data and the pipelines consuming it.
I think a lot more teams are having that same realisation now.
For a long time, it was accepted that if the data team’s pipelines broke for a day, or a few days, or even a week, it didn’t really matter. The numbers were for finance and the business, and they could wait.
I’m not sure that was ever really true! Certainly as someone on that team, being stuck permanently firefighting is not much fun… But I can see why, from the outside, it looked like an acceptable trade-off.
But it isn’t any more.
More and more, the data we produce doesn’t just feed a dashboard someone checks once a week. It powers a product feature, either directly or via an AI agent.
So, we need to invest in the reliability of the data.
That includes us upskilling and building more reliable pipelines, improving observability, and so on.
But there’s no point doing that if the data we depend on is also unreliable, which is why we have to improve data reliability at the source.
And we do that with data contracts.
I talk about all this on The Data Engineering Show. We also get into:
- Moving from a centralised data team to a self-serve platform
- Designing for different user personas
- Building the new world while still supporting the old
- Using agents to automate support requests
- Rethinking data governance in an AI-driven world
Check it out and let me know what you think!
Interesting links
Data Contracts: Bringing Data and the Business Closer Together by Ian Røpke (on LinkedIn)
Nice post on using data contracts to build a shared language and understanding.
Why we rebuilt our data warehouse and how it unlocks self-driving products by Eric Duong (Posthog)
Interesting federated architecture using DuckDB. Interesting site too…
Three Years, 59 RFCs, One Vote: Bitol Graduates by Jean-Georges Perrin
Bitol, the the Open Data Contract Standard, is now a graduated project of the LF AI & Data Foundation.
Being punny 😅
Roman numeral puns are great and I for one will continue to make them.
Upcoming events
- Data Community Conference, September, Online
Thanks! If you’d like to support my work…
Thanks for reading this weeks newsletter — always appreciated!
If you’d like to support my work consider buying my book, Driving Data Quality with Data Contracts, or if you have it already please leave a review on Amazon.
Enjoy your weekend.
Andrew