You have a governance program. There's a committee, an inventory of approved use cases, an acceptable use policy, a review gate before anything touches customer data, and a named owner for model risk.
It's more than most organizations have, and it will hold up well in a board conversation.
Then someone asks a harder question, about whether the answers coming out of these systems can be defended. Where did that number come from. Which definition of revenue did the assistant use when it produced that summary. Show me.
That question travels down through the policies and lands somewhere the program was never designed to look.
Every control describes data. Few of them describe the description
Governance controls govern what happens to data: who can see it, where it can move, how long it's retained, which use cases are approved.
All of those controls bind to an inventory. They assume something upstream has produced a complete, current, accurate account of what exists in the estate, what each field means, and where each number comes from.
That account is your metadata. And in most organizations it has never been measured.
The result is a program that is well designed on paper and untested at its foundation. The policies are sound. The question is whether the description they operate on is complete enough for them to attach to anything.
Three ways this surfaces
The confident wrong answer. An assistant returns a revenue figure that differs from the board deck. Both came from the same warehouse the same morning. There were four definitions of revenue written down across the estate and no record of which one the finance close actually runs on. The system picked one. It had no way to know it mattered.
The audit request that becomes a project. A regulator or an internal auditor asks how a reported figure was produced. If the estate is well described, that's a query. If it isn't, it's six weeks of engineering time, a spreadsheet, and a set of answers assembled by hand that will be hard to reproduce next quarter.
The reconciliation tax. Two teams present different numbers for the same metric. Someone spends a week finding the discrepancy. This happens often enough that it's stopped being remarkable, and the cumulative cost has never been counted because it never appears as a line item. Ask your analytics team how many hours a quarter go to reconciling definitions across tools. The number is usually available and usually worse than expected.
Each of those has been treated as a separate problem with a separate fix. They share one cause.
The dependency is measurable
Metadata completeness can be measured. It scores on five dimensions, and the score can be tracked over time like any other operational number.
Five questions, and your team can answer all five this quarter:
- How much of the estate is described at all? Including the older systems that still feed regulated reports.
- How deep does the description go? Knowing a report touches a system is a different thing from knowing which specific field produced the number.
- Where did the description come from? Some of it was written by people, some observed by tools, some derived from the code that defines the transformations. These carry very different levels of confidence.
- How current is it? A description that reflects the estate as it stood eighteen months ago will confirm things that are no longer true.
- Who owns each definition? A named person who would answer if you asked them today, beyond an entry in a document.
Score those five across your three most-reported metrics and you have a number. Run it again in two quarters and you have a direction.
These five describe the metadata, never the values in your data. Whether last night's load was correct is a separate discipline with separate tooling. This is about whether the estate is described well enough for your controls, your reporting, and your AI systems to rest on it.
What to ask, and when
Bring these to your next data leadership review:
What is our metadata completeness score, and what's the denominator? A team that owns this can answer with a number. A team that answers qualitatively has given you the finding.
Which systems are dark? The list is usually short, specific, and older than the current team. It's also usually where the regulated reporting lives.
Who owns this? In most organizations the answer is currently four teams, which functions as no one. Some organizations have started naming a single role for it.
What would it cost to close the top two gaps? Almost always less than one AI initiative that stalls waiting for a definition.
Where this leaves the AI conversation
Every AI governance claim your organization makes rests on the same foundation, and it's the one part of the stack that rarely gets a budget line, because it produces nothing visible when it's working.
The organizations that will be able to answer the harder board question in twelve months are the ones measuring this now. The measurement is cheap. It takes a quarter, it produces a number, and the number gives you something to manage.
Start with three metrics and five questions. The gap list you end up with is the beginning of the plan.