Continuous Governance with DataGovOps

DataGovOps turns manual, error-prone governance into automated, repeatable orchestrations that run continuously as part of DataOps development workflows.

Written by DataKitchen Marketing Team on September 3, 2020

Data Governance
Continuous Governance with DataGovOps

Key points

  • DataGovOps is governance automation: turning the inefficient, time-consuming, error-prone manual processes around governance into code and repeatable orchestrations.
  • Governance framed as enforcement puts it in conflict with analytics productivity. Positive incentives and enablement get further than handing out the equivalent of speeding tickets.
  • Automated orchestrations fold glossary and catalog updates into the change management process, so governance deploys with everything else rather than landing on an already-busy analyst.
  • Process lineage records all metadata related to data, including the code that acts on it — test results, timing data, quality assessments — with everything stored in version control.
  • A self-service sandbox is a purpose-built race track rather than a speed trap: it enforces where you can go while letting you move fast, and raises an automated alert when a policy is violated.

Data teams using inefficient, manual processes often find themselves working frantically to keep up with the endless stream of analytics updates and the exponential growth of data. If the organization also expects busy data scientists and analysts to implement data governance, the work may be treated as an afterthought, if not forgotten altogether. Enterprises using manual procedures need to carefully rethink their approach to governance.

With DataOps automation, governance can execute continuously as part of development and operations workflows. Governance automation is called DataGovOps, and it is a part of the DataOps movement.

DataGovOps in Data Governance

Governance is, first and foremost, concerned with policies and compliance. Some governance initiatives focus on enforcement – somewhat akin to policing traffic by handing out speeding tickets. Focusing on violations positions governance in conflict with analytics development productivity. Data governance advocates can get much farther with positive incentives and enablement rather than punishments.

DataGovOps looks to turn all of the inefficient, time-consuming and error-prone manual processes associated with governance into code or scripts. DataGovOps reimagines governance workflows as repeatable, verifiable automated orchestrations. DataGovOps strengthens the pillars of governance through governance-as-code, automation, and on-demand enablement in the following ways:

Figure 1: The orchestrations that implement continuous deployment incorporate DataGovOps updates into the change management process.

Figure 2: All artifacts that relate to data pipelines are stored in version control so that you have as complete a picture of your data journey as possible.

Conclusion

The concept of governance as a policing function that restricts development activity is out-moded and places governance at odds with freedom and innovation. DataGovOps provides a better approach that actively promotes safe use of data with automation that improves governance while freeing data analysts and scientists from manual tasks. DataGovOps is a prime example of how DataOps can optimize the execution of workflows without burdening the team. DataGovOps transforms governance into a robust, repeatable process that executes alongside development and data operations.


FAQ

What are the key points in this blog?

DataGovOps applies DataOps automation to governance, turning manual policy work into repeatable orchestrations. It folds glossary and catalog updates into change management, records process lineage in version control, runs data testing continuously rather than periodically, and provides self-service sandboxes that enforce policy while letting analysts move quickly.

What is DataGovOps?

DataGovOps is governance automation and a part of the DataOps movement. It reimagines the inefficient and error-prone manual processes associated with governance as code, scripts, and repeatable verifiable orchestrations, so governance executes continuously as part of development and operations workflows instead of as a separate periodic exercise.

Why does enforcement-based governance fail?

Because focusing on violations positions governance in conflict with analytics development productivity, somewhat like policing traffic by handing out speeding tickets. Teams route around it. Governance advocates get considerably farther with positive incentives and enablement, which is what makes the automated approach work where the policing approach stalls.

What is process lineage?

Process lineage is the record of all metadata related to data, including the code that acts on that data. Test results, timing data, data quality assessments, and every other artifact generated by executing the pipeline document how the data came to be, and all of it is stored in version control.

What is a self-service sandbox in governance terms?

It is an environment containing everything an analyst or data scientist needs to create analytics, created on demand with background processes monitoring governance. If manual governance is handing out speeding tickets, a sandbox is a purpose-built race track: it enforces where you can go and is built so you can go fast.

Install Open Source TestGen Free, no vendor lock-in Request a Demo See TestGen Enterprise in action

DataKitchen Marketing Team

The DataKitchen marketing team curates industry news, resources, and thought leadership on DataOps, data quality, and data observability.