- A data catalog is the searchable inventory of an organization's data: what exists, what it means, who owns it, where it comes from and whether it can be trusted.
- It combines technical metadata (schemas, types), business metadata (definitions, glossary) and operational metadata (freshness, usage, quality) in one place.
- Active metadata is the shift that made catalogs useful: metadata harvested automatically and pushed back into the tools people work in, instead of a manual registry that rots.
- The catalog is where governance becomes visible. Ownership, lineage and quality only help consumers if they surface at the moment someone is choosing which dataset to use.
In most enterprises, finding the right dataset is tribal knowledge. You ask a colleague, who asks another, until someone remembers which table has the numbers and which of its three copies is current. That search tax is paid on every analysis, every model, every report. A data catalog replaces it with a lookup. This article covers what a catalog actually contains, why active metadata changed the game, and how to choose between buying one and running an open-source one.
What a data catalog is, and is not
A data catalog is a searchable inventory of an organization's data assets, enriched with the context needed to understand and trust them. It answers the questions that otherwise require a human: what data exists, what each field means, who owns it, where it came from, how fresh it is and whether anyone considers it reliable.
It is not a storage system and not a governance policy. It is the discovery and context layer that sits on top of your existing warehouses, lakes and databases. The distinction matters because a catalog does not move or own data; it describes it. Its value is entirely in making the data estate legible to the people who need to use it.
Three kinds of metadata
A useful catalog brings together three layers of metadata that usually live in different places.
- Technical metadata: schemas, column names, data types, table sizes, partitioning. This comes from the systems themselves and describes the physical shape of the data.
- Business metadata: what a field actually means in business terms, the definitions, the glossary, the mapping between "MRR" on a slide and the column that computes it. This is the layer humans add, and it is where most of the value and most of the effort lives.
- Operational metadata: freshness, update frequency, usage statistics, quality scores, lineage. This is dynamic, harvested from pipelines and query logs, and it is what lets a consumer judge whether a dataset is safe to use right now.
A catalog that has only technical metadata is a schema browser. The business and operational layers are what turn it into a decision tool.
Discovery, glossary and the human side
The everyday job of a catalog is discovery: someone needs revenue by region, searches the catalog, and finds the certified dataset with its definition, owner and freshness, instead of guessing among five candidates. Good search, with business terms and not just table names, is what makes this work.
Underneath sits the business glossary: the agreed definitions of the terms the organization uses. When "active customer" has one documented definition linked to the datasets that implement it, arguments about whose number is right largely disappear. Building and maintaining that glossary is human work that no tool automates, and it is the part that pays off most, because it aligns the language of the business with the structure of the data.
Active metadata and automation
Early data catalogs failed for a predictable reason: they were manual registries. Someone had to document each dataset by hand, the documentation went stale the moment anything changed, and consumers learned not to trust it.
Active metadata is the shift that fixed this. Modern catalogs harvest metadata automatically by crawling source systems, parsing query logs to infer lineage and popularity, and integrating quality checks. Just as important, they push that metadata back out: into the BI tool, the query editor, the notebook, so context appears where people already work rather than in a catalog they have to remember to visit. A catalog that only pulls metadata in is a library; one that pushes context back out is part of the workflow.
Buy versus open source
There are two credible paths, and the choice follows team maturity and control requirements.
Commercial platforms (Collibra, Alation, and the catalog features of the major cloud data platforms) offer polished governance workflows, glossary management and enterprise support. They suit organizations that want governance capability quickly and are willing to pay for it.
Open-source platforms (DataHub, OpenMetadata) offer strong metadata models, active-metadata ingestion and full control, at the cost of running them yourself. They suit teams with the engineering capacity to operate the platform and a preference for avoiding vendor lock-in, which in a European context often aligns with sovereignty requirements too.
Neither is universally right. The wrong move is buying a heavy governance platform and using it as a schema browser, or standing up an open-source catalog with no one to curate the business glossary. The tool is a fraction of the work; the curation is the rest.
Failure patterns
From the field, the patterns that leave a catalog unused:
1. Manual and stale. A catalog documented by hand, out of date within weeks, distrusted within months.
2. Technical metadata only. A schema browser with no business definitions, so it answers "what columns exist" but not "which dataset should I use."
3. Catalog no one opens. Metadata that lives in a separate tool instead of surfacing in the BI and query environments people already work in.
4. No glossary ownership. Definitions that nobody maintains, so the business language and the data drift apart again.
5. Bought for the logo. A heavyweight platform acquired for compliance optics and never curated, delivering none of the discovery value that justified it.
Talk through your data catalog
DNA Solutions helps European enterprises make their data findable and trustworthy: a catalog with active metadata, a business glossary that reflects how the organization actually talks, and ownership, lineage and quality surfaced where people work. Whether you are choosing between a commercial and an open-source platform or reviving a catalog no one uses, we focus on the curation that turns an inventory into a decision tool. Talk to us.
Related services: Data & Analytics



