Whitepaper

What Should a Lineage POC Prove on Your Own Estate?

A practical guide to running a data lineage proof of concept against your own estate. Test six production workflows, from impact analysis and BI migration to metric consistency, audit evidence, and AI context, with measurable pass criteria.

On this page
Text Link

You've shortlisted. The vendors that made the cut passed the connector match, the second harvest, and the parsing test in Which Lineage Holds Up Outside the Demo?, and now one or two of them get your estate for a few weeks.

This guide covers that POC. It walks through six workflows where lineage carries production weight in a large estate, what the data team does in each one, the number to record, and what sits underneath all six.

Trusted BI. Reliable AI. Defensible governance.

Those are the outcomes on the table, and every one of them depends on a trace that holds when someone asks for it.

Before the first harvest

A POC that starts without scope turns into a feature tour. Settle four things in writing before any connector runs.

Scope. Name the systems in the path of one board-deck figure, end to end, including the oldest one. Then add the specific artifacts the six workflows need: a schema change already sitting in a pull request, a metric two dashboards disagree on, a BI model slated for migration, and an AI use case in pilot. Check every system against the connector explorer before the kickoff call.

People. You need a data engineer to run the traces, a steward who owns the definitions in dispute, someone from the BI platform team, and the person who signs off on the result. The steward is the one most often left off the invite, and the metric consistency workflow stalls without them.

Baseline. Record today's numbers before MetaKarta touches anything: the consumer count your current tool returns for the schema change, the hours your last root cause investigation took, the sync jobs running between your metadata tools, and how long the last audit request took to answer. Every workflow below ends with a metric, and a metric needs a before.

Exit criteria. Agree on the timeline and on what passing means for each of the six checks at the end of this guide, then get the sign-off owner to initial it. That page is what the business case gets built from.

How the lineage gets computed

MetaKarta reverse-engineers the estate by reading the code and design artifacts that move data: SQL scripts, stored procedures, ETL mappings, BI semantic models, notebooks. Transformation language parsers cover SQL DML across on-premises databases (Microsoft SQL Server Transact-SQL, Oracle PL/SQL, Teradata BTEQ and Fastload) and cloud databases (Snowflake SQL), plus Python and Scala as used in Databricks. Data flow emulators cover the hundreds of transformations inside platforms like Informatica PowerCenter, IBM DataStage, Qlik Talend, dbt, and Matillion.

Nothing has to run first. Lineage by design, not by observation.

Two flow types render in the same diagram. Data flow follows connection definitions and physical transformation rules, drawn left to right as solid lines. Semantic flow follows definition and usage relationships from a term or logical model down to its physical representation, drawn top to bottom as dashed blue lines.

Semantic flow is what makes the metric consistency and AI context workflows below possible. Everything computes at the column level first, and table, schema, and model views are unions of the underlying column traces. Hold onto that detail for the next section.

Impact and root cause analysis

The trigger. A schema change is sitting in a pull request, or a dashboard number moved overnight and three teams are on a call about it.

The workflow. Open the column, go to the Lineage tab, and set Direction to Impact for the forward view or Lineage for sources. Depth runs from one adjacent step out to nine, or Any for the complete trace. The Tree view returns the same trace as text, scoped to Adjacent, Ultimate, or All, and downloads to CSV, which is the artifact that goes into the change ticket.

The Processes panel at the bottom carries the answer for root cause. It splits into processes that make the selected object and processes that use it. Click through to lineage details and the actual script opens, with two-way highlighting: select a step in the diagram and the matching code highlights, select the code and the diagram object selects.

Control flow is the part most impact analysis misses. A column can govern what moves without moving itself, through a WHERE clause, a filter, or a lookup. Set Control Flow to Limited for adjacent control dependencies or Complete for the full set, and the filter conditions surface in the Processes panel with the expression visible.

Why column level is the whole argument. Take a payment table in a staging warehouse. Traced at the table level, the union of all its column traces implicates nine reporting technologies downstream. Select one column, CheckNumber, and two of those reporting models drop out of the trace entirely.

A table-level tool reports those two as impacted anyway. The team that spends the afternoon regression-testing them spent it for nothing.

Know the blast radius before the first ticket lands.

What the repository adds. The catalog entry, the owner, the glossary term, and the sensitivity label hang off the same object you're tracing. Impact analysis returns a list of consumers with the names of the people accountable for each one.

The number to record. Coverage delta: the consumers MetaKarta returns for your schema change, against the count from your baseline. Then replay your last real incident and record mean time to root cause against the hours it took the first time. The Impact & root cause analysis brief covers the business case behind both.

Data and BI platform migration

The trigger. A BI consolidation or a platform move lands on the schedule, scoped from a report inventory that was already incomplete.

The workflow. Start with the inventory the migration actually needs. Trace impact from the source tables to Ultimate in the Tree view and the terminal nodes are report fields, which is the real dependency list. Filters keep it readable: exclude model types, exclude specific models, hide temporary objects, and show or hide internal transformation steps and external source objects.

Lineage overview at the model level answers the second question, which is what happens inside a given job. Internal data flow shows design-level movement inside the model, external data flow extends to stitched models, transformation data flow includes every intermediate transformation, and summary data flow strips those out for the readable version.

Then the semantic assets move. Semantic Hub imports existing BI definitions from Power BI and Tableau, and from legacy platforms like MicroStrategy, SAP BusinessObjects, IBM Cognos Framework Manager, and Oracle OBIEE, as governed starting points. Governed definitions then compile into native artifacts per target: Snowflake Semantic Views, Databricks Metric Views, Power BI semantic models, and Tableau data sources.

Write once, compile anywhere.

What the repository adds. The pre-migration lineage and the post-migration compiled artifacts sit in the same repository under the same version history. Proving the new environment produces the old number becomes a side-by-side comparison.

The number to record. Rebuild scope: how many of the models in scope import cleanly as governed definitions, each one a rebuild that comes off the migration plan. Record dependency coverage beside it, the report fields the trace found against the ones in your original inventory. The Data & BI platform migration brief covers both.

BI metric consistency

The trigger. The CFO's deck shows one revenue figure, the sales dashboard shows another, and someone has asked which one is right.

The workflow. Start with the two measures, one in a Power BI semantic model and one in a Tableau data source, and set Direction to Lineage with Depth at Any. Each trace runs back through the BI model, the ETL, and the SQL to the source columns. The Processes panel shows the expression at every step, so the divergence is on screen: one version drops refunds in a WHERE clause, and the other joins to a different customer dimension.

Set Control Flow to Complete for this one. The filter that separates two revenue numbers usually governs the data without moving it, which is the dependency a table-level trace leaves out.

Then turn on Show Term Definitions. Definition lookup shows which glossary term, if any, each version claims to implement. Semantic usage shows every other report and model that depends on the same term, which tells you how far the disagreement already reaches.

A definition you can't trace is a definition you can't trust.

What the repository adds. Semantic Hub imports both versions as governed starting points, and a steward picks the canonical one in Data Governance with the two traces as the evidence for the decision. The governed definition then compiles back out to the Power BI semantic model and the Tableau data source, so both dashboards read the same definition. Documentation describes what should be true. Compilation makes it true.

The number to record. Time from the question to the divergence on screen, and the count of reports and models that semantic usage shows depending on the disputed term.

Audit and compliance

The trigger. A regulator or an internal auditor names a figure and a date. The date is rarely today.

The workflow. Version and configuration management runs across any scope of lineage model components, called multi-models. A multi-model combines what belongs together: a data modeling import with its logical and physical models, an RDBMS source organized by catalog, a file system import covering a directory hierarchy, and the ETL jobs that read from both. Configurations carry named versions, so a Published configuration and a Development configuration coexist and get compared.

Compare Metadata returns the side-by-side view of additions, changes, and removals between two versions, color-coded. That comparison is the audit answer to "what changed," and the parser-derived dependency data behind it is the answer to "what did it affect."

Conditional labels travel with the trace. Turn on PII or a confidentiality level in Display Options and the diagram shows every place sensitive data lands downstream of its source. Object audit logs record every change with an attribution and a timestamp.

The audit answer is a query.

Multi-model versions also pay for themselves operationally. Only changed databases, schemas, packages, reports, or files need re-collection, so incremental harvesting stays affordable through the 200th refresh.

What the repository adds. Governance policy, stewardship history, and semantic definitions version on the same clock as the lineage. Audit preparation becomes a set of saved queries, and the staffing plan goes away.

The number to record. Hours to answer one audit-style request end to end, against your baseline from the last real one.

AI context and governance

The trigger. A text-to-SQL interface or an agent works in the demo and stalls at review, because the team can't produce the derivation behind an answer.

The workflow. Semantic flow carries this one. Definition lookup traces from a physical column back to the term that defines it, and semantic usage traces the other direction, from a term to every physical asset it governs. Turn on Show Term Definitions in the lineage diagram and the glossary definitions render alongside the data flow, on the same canvas.

That structure is what gets served. MetaKarta MCP exposes the repository to external agents through two toolsets: Metadata Management Tools deliver governed lineage, catalog, and governance metadata, and Semantic Hub Tools serve compiled context, the compiled semantic definitions. Structured AI consumption through Snowflake Cortex Analyst and Databricks Genie reads the same governed definitions, so an agent and a BI report resolve the same metric against the same source.

Ontology Modeling binds business concepts to physical assets from the start, and Context Sandbox validates the grounding against live data before an agent sees it. Lineage then traces from the governed definition the agent consumed back to the columns and the code that produced them, which is the evidence a model risk review asks for.

What the repository adds. The governed context an agent reads is the same context the catalog shows a person and the same context governance applies policy to. One definition and one lineage serve all three.

The number to record. For the pilot's test questions, the share of answers that come back with the governed definition and its lineage attached.

What sits underneath all six

The trigger. This one usually surfaces during procurement, when someone adds up what the current stack costs to keep in sync.

The workflow. Data Lineage, Data Catalog, Data Governance, and Semantic Hub read from and write to one metamodel, the Meta Integration Repository. A governance decision shows up in lineage immediately, and a catalog update reaches the semantic definitions with no sync job between them.

Compare that to the common production architecture: a lineage tool, a catalog, a governance workflow tool, and a standalone semantic product, each with its own metadata store, its own connection schedule, and its own model of what a column is. Reconciling them is standing work that grows with every source added, and API maturity in any one tool doesn't retire it.

Every workflow above depends on this. The owner in the impact report, the steward's decision in the metric dispute, and the policy version in the audit answer all come from the same record as the trace.

The number to record. Integration hours: the sync jobs, custom connectors, and reconciliation scripts from your baseline that the POC estate ran without. Record the drift interval beside it, how far behind any one tool runs between syncs today. The Metadata tool consolidation brief covers both.

The POC scorecard

Six checks, one per workflow. Run them on your estate, with your worst systems included, and fill in the last column from your baseline.

Workflow Check What passing looks like Number to record
Impact and root cause Run impact on a proposed schema change with Control Flow set to Complete, then replay one recent incident A consumer list including WHERE-clause dependencies, exported to CSV; the incident's root cause traced to a line of code Coverage delta; mean time to root cause
Platform migration Import an existing Power BI or Tableau semantic model, compile it to the target, and open the compiled artifact in the target tool The compiled artifact matches the original definition Rebuild scope; dependency coverage
Metric consistency Trace two conflicting versions of one metric to source The filter or join that makes them differ, with the expression on screen Time to the divergence; reports depending on the term
Audit and compliance Compare a model against its state six months ago, then trace one filed figure to the expression that produced it Side-by-side additions, changes, and removals; the join, the filter, and the line of code Hours to answer the request
AI context Trace one agent answer from the governed definition it consumed back to source An unbroken chain from definition to code Share of answers with derivation attached
Shared repository Classify a dataset in Data Governance, then check lineage, the catalog entry, and a compiled definition that depends on it The classification visible in all three, with no sync step Integration hours; drift interval

Weight the first check most heavily. Impact analysis that runs before deployment changes how a team schedules work, and it's the hardest of the six for an assembled stack to produce.

What one trace is worth in production

Every workflow in this guide runs on the same trace: column level, computed from the code, and stored in the repository that also holds the catalog, the policies, and the governed definitions. That's why one trace answers the change ticket, the migration plan, the revenue dispute, the auditor, and the model risk review.

MetaKarta is the Metadata Management Platform built on that repository. Meta Integration Technology, Inc. has spent nearly 30 years building the metadata connectors that other vendors ship inside their own products, embedded as the OEM engine in Microsoft Purview, Informatica from Salesforce, IBM, Oracle, and Qlik Talend.

The industry has long called that engineering the Switzerland of Metadata.

Data Lineage computes column-level lineage from the code itself across 400+ connectors, on the same shared metadata repository as Data Catalog, Data Governance, and Semantic Hub. Get in touch to scope these six workflows on your own estate.