Your customer discovered the poor data and emailed you about it at 9 am, copying your boss. The night before, all the jobs in the pipeline finished successfully, and the one person who checked them thought the row counts looked okay. The data quality assessment you arranged last year is stored on the shared drive, is 80 pages long, includes a maturity score on page 4, and has a roadmap no one has funded.
We’ve experienced this ourselves, and so have the people we work with. Over the last few years, we have tried to document the differences made by teams that avoided it, and as a result, we’ve developed 24 principles plus one more. We called it the DataOps Quality Manifesto, and I want to explain why we created it, what it contains, and how it connects to the two manifestos we wrote earlier.
Why we wrote it
That’s a fair question; we’ve already produced two. We published the DataOps Manifesto in 2017, stating that the data team should operate like the most effective software teams, using version control, automated tests, small changes, a fast cycle time, and avoiding reliance on any individual. The Data Journey Manifesto, released in 2023, claimed that data teams should monitor production like a factory monitors its production line: anticipate problems at every stage, alert when a failure occurs, and never let the customer discover the problem first.
They both state that you should test your data, but neither explains how. “Quality is paramount” is principle 15 of the DataOps Manifesto; that is correct, but it’s useless at 2 am. The Data Journey Manifesto claims it’s unacceptable for customers to find problems, and that is true too, but it doesn’t say where to run the first test.
As to who cares, here is the honest answer: the individual who nodded along when the first two manifestos were presented returned to their desk and didn’t know what to do on Monday morning. This document is for Monday morning; it explains where the tests originate, where they go, who is responsible for them, how to score them, and how the results can be turned into the budget you have been asking for since the assessment.
Where the ideas came from
It isn’t anything new; that’s the whole idea behind a manifesto. By putting forward what already works, you help people avoid relearning it every time there’s an outage.
Begin with the two manifestos that came before it. The DataOps Manifesto provided the structure for this one by introducing the approach of stating the values first, followed by numbered principles and an additional point at the end. It also included principle 15, which states that quality is paramount, and an extra point recommending that you start with data testing. The Data Journey Manifesto contributed the factory idea: production must stop when a defect is found, and it is unacceptable for customers to discover problems. About half of the Quality Manifesto consists of these two ideas, made specific so they can be put into action on Monday.
The four testing points were identified by observing teams as they developed medallion architectures and added tests to one layer, typically the one they owned. The same null check serves a different function at the source, at ingestion, in production, and in development, and if any of these checks is skipped, someone else will report a failure to you. We documented this in a blog post on the four points in your medallion architecture where data testing matters, and this became the second part of the manifesto.
The influence principles came from the data quality leaders that we speak to every week. They do not have control over the systems that produce the data, and they can’t impose any requirements. What they can do is present a table profile together with the forty faults identified in it, and that approach always outperforms what a steering committee could achieve. This argument forms the core of our white paper on data quality from the DataOps perspective.
The idea of using tests as a shared resource came from a problem we kept running into. Engineers want the rules stored in Git, but stewards don’t have a Git account and never will, so analysts modify the rules in a spreadsheet, and the spreadsheet ends up winning. The solution was a single database as the official record for tests and results, with a separate door for each kind of user, and Git tracking all changes. That idea, first written up in Your Best Data Quality Rules Live in Someone Else’s Head, became principle eight.
Score what matters, keep the trend line, and start with one person came from our book on data quality and from observing which data quality initiatives made it through their second budget cycle. The ones that obtained a low score for the completeness of the fax number field did not survive.
What it says
It begins in the same way as the DataOps Manifesto by highlighting what we have come to appreciate: a preference for automated tests over relying on hope and carrying out manual checks, measurement over opinion, generated tests over manually written rules, test coverage over lineage diagrams, influence and evidence over mandates, and having a working test at 70% now rather than waiting for a perfect one in the future.
The 19 principles are grouped into five categories.
- Evidence: You can only know what you have tested. When checking the data, keep a close eye on the tools, since a perfectly formatted table is useless if the refresh process never runs. Manual checks are pointless because no one will remember at 2 am. Set your standards only after you have measured.
- People: Begin by selecting one person and using a free tool. Influence is the job, since you will never come to own the source system.
- Tests: Let the profile be responsible for writing the tests. Consider the tests to be a shared resource with a single authoritative version. Release a version today at 70%.
- Scores and action: Score the things your customers care about, not the fax number column. Each score must be linked to a test. Simplify action by assigning a named person and providing them with a ticket in the system that they already use. Continue to display the trend line since six months’ worth of scores tied to a KPI that your boss is already monitoring represents a budget.
- Practice: Embrace your mistakes, and assign fault to the step, not to the person. Include the delivery terms in the tests. Verify the migration. Reduce the noise, and downgrade the false alarms that occur every night rather than suppressing them. Take responsibility for it in production. Test the agent’s output.
Testing happens at four stages: at the source, at ingestion, in production, and during development. That is one null check and four tasks, four locations. If you omit one of them, the customer will let you know about the failure.
And a plus one: if you’re starting DataOps from scratch, start with tests, since they are the cheapest step and the ones your customers experience first.
Three manifestos, one line
The three manifestos fit together like this. The DataOps Manifesto explains how you build: use version control, maintain environments, employ orchestration, and run automated tests with every change. The Data Journey Manifesto explains how you monitor by setting expectations at each production step, which you observe but don’t execute. The Quality Manifesto explains how you prove it by running tests at four points: generating them from the profile, scoring them for the customer, assigning ownership, and feeding the results back into the other two.
They share the same spine: all three recommend automating testing. The DataOps Manifesto concludes with a plus one: start with data testing, while the Quality Manifesto ends with a plus one stating that DataOps should start with quality. This is no coincidence. Of the 64 principles, if you do one thing, write a test and run it on every load. Then everything else becomes much easier.
What to do with it
Read it; it’s just one page long. Then keep a list of every error your team has, data errors and process errors alike. After a few weeks, convene a meeting of some people. Look at those errors, and instead of blame, try to find some automated fixes so they never happen again. We call this a quality circle. It’s the cheapest and easiest first way to see that these principles make sense. Then start with some automation. Pick a table, write some tests, watch which tests are failing, and put them in the right place.
TestGen is the open-source tool we created to do exactly this: it provides profiling, generates tests, runs hygiene screening, produces scores, includes dashboards, and offers an MCP server so your agent can read the results. It’s free, it can be run using Docker, and you’ll get your first profile by the same afternoon.
Take a walk, read the manifesto, then start using it with TestGen.
FAQ
What are the key points in this blog?
DataKitchen wrote the DataOps Quality Manifesto because its first two manifestos say test your data but not how. The new one is 24 principles and a plus one: 19 principles in five groups and four points where tests go. The DataOps Manifesto is how you build, the Data Journey Manifesto is how you watch, and the Quality Manifesto is how you prove.
What is the DataOps Quality Manifesto?
The DataOps Quality Manifesto is a one-page statement of how to do data quality the DataOps way, published by DataKitchen in 2026. It opens with values such as automated tests over hope and manual checks, then sets out 19 principles in five groups, the four points where tests go, and a plus one: start DataOps with the tests.
Why did DataKitchen write a third manifesto?
The DataOps Manifesto (2017) and the Data Journey Manifesto (2023) both say test your data, and neither says how. Quality is paramount is principle 15 of the DataOps Manifesto, which is true and useless at 2 am. The Quality Manifesto says where tests come from, where they go, who owns them, and how to score them.
How do the three DataOps manifestos fit together?
The DataOps Manifesto is how you build: version control, environments, orchestration, and automated tests on every change. The Data Journey Manifesto is how you watch: an expectation at every step in production. The Quality Manifesto is how you prove: tests at four points, generated from the profile, scored for the customer, and owned by a named person.
What are the four points where data tests go?
Data tests go at the source, at ingestion, in production, and in development. The same null check does a different job in each place. Teams building medallion architectures often test only the layer they own, and skipping any one of the four leaves a failure you hear about from someone else, usually the customer.
What are the five groups of principles in the DataOps Quality Manifesto?
The 19 principles fall into five groups. Evidence: you only know what you tested. People: start with one person and a free tool. Tests: let the profile write the tests. Scores and action: score what your customer cares about and keep the trend line. Practice: love your errors, prove the migration, and test what the agent wrote.
Why should data quality tests be a shared resource?
Engineers want data quality rules in Git, stewards will never have a Git account, and analysts fix the rule in a spreadsheet. The manifesto answers with one database of record for tests and results, a door for each kind of user, and Git capturing every change. That idea became principle eight of the DataOps Quality Manifesto.
How many principles are in the three DataOps manifestos?
There are 64 principles across the three: 18 in the DataOps Manifesto, 22 in the Data Journey Manifesto, and 24 in the DataOps Quality Manifesto. All three say automate the testing, and two end with a plus one about starting with tests. If you do one thing, write a test and run it on every load.
How do I start using the DataOps Quality Manifesto?
Read it, since it is one page. Then keep a list of every error your team has, data and process alike. After a few weeks, hold a quality circle: look at those errors and, instead of blame, find automated fixes so they never happen again. Then start automating: pick a table, write some tests, and watch which ones fail. DataOps TestGen, an open-source tool, does the profiling, test generation, and scoring.
