Case study

Building a Governed Data Lakehouse for HAPPY END

A governed BigQuery foundation for reporting, automation, and operational systems.

On this page

Overview

HAPPY END EHS Solutions provides products and services for safe, sustainable operations across Central Europe. Its commercial and operational systems manage customers, products, suppliers, orders, invoices, and sales activity across several markets.

Those systems contained years of useful data, but their schemas reflected application history rather than analytical contracts. Abbreviated Czech names, country-specific identifiers, historical conventions, and undocumented exceptions made correct interpretation depend heavily on experience.

Copying tables into the cloud would have preserved the ambiguity. HAPPY END needed to protect production systems, retain source evidence, define stable business entities, document what the data meant, control access, and publish outputs other teams could reuse.

The result is a separate analytical platform where core concepts such as customers, products, transactions, revenue, and margin are defined once and consumed by reporting, departmental data products, commercial analysis, operational alerting, and automation.

Platform scale and reuse

Curated business models
208
Reusable entities, transactions, facts, and marts built from legacy operational data.
Documented source fields
7,374
Maintained alongside descriptions for 489 source tables through catalog-as-code.
Governed BigQuery objects
886
Tables and views maintained across 74 production datasets.

How legacy data becomes a governed data product

Operational sources

Protect production systems while moving data through controlled ingestion paths.

  • Source

    Operational systems remain authoritative while analytical workloads run against a separate platform.

  • Ingest

    Transactional replication, managed ingestion, and controlled pipelines move operational and reference data into BigQuery.

Governed data foundation

Turn source-shaped records into documented, reusable business definitions.

  • Model

    Dataform turns source-shaped records into reusable entities, transactions, facts, and decision-specific marts.

  • Govern

    Catalog-as-code, documentation, IAM, and row-level policies keep meaning, lineage, ownership, and access reviewable.

Reusable data products

Publish trusted outputs for reporting, alerts, automation, and future systems.

  • Serve

    Curated data products become shared inputs for reporting, analysis, alerts, automation, and downstream systems.

Why the legacy environment needed an analytical boundary

The legacy environment was built to run the business. Complex analytical queries risked competing with operational work, while repeated extracts and locally maintained calculations created several versions of the same business concept.

Revenue, margin, customer, product, and transaction logic varied by consumer. Experienced users also relied on memory to interpret historical fields, market-specific identifiers, and source behavior that had never been documented as a shared contract.

HAPPY END separated transaction processing from analysis. Operational systems continued to own daily work, while the lakehouse became the controlled source for analytical models, reporting, alerting, and automation.

Safe ingestion and explicit modeling

A one-way transactional replica isolated analytical extraction from production. CDC-based managed ingestion moved operational changes into BigQuery, while external reference data followed controlled pipelines through Cloud Run, Workflows, and Cloud Scheduler. Dataform then resolved source structures into business entities and outputs that reports, alerts, and automated workflows could consume.

Operational foundation

Isolate data movement from production and give ingestion a recoverable operating path.

  • Decouple change capture

    The replica and CDC pipeline let source systems remain authoritative without turning analytical ingestion into part of their production workload.

  • Make ingestion recoverable

    Source-shaped records, extraction metadata, and explicit pipeline boundaries make failed loads easier to diagnose, replay, and validate.

Reusable data products

Replace source-specific structures with stable contracts for downstream consumers.

  • Stabilize business contracts

    Consistent entities, identifiers, relationships, history, and measures insulate consumers from legacy schemas and source conventions.

  • Separate logic from consumption

    Reports, marts, alerts, and workflows start from curated outputs instead of repeating source joins and transformation rules.

Meaning and ownership stay visible

Each data product has explicit definitions for grain, lineage, ownership, access, and lifecycle. Documentation, permissions, models, and infrastructure are version-controlled so the platform carries its own explanation as it changes.

  • What does this field mean?

    Version-controlled field and table descriptions preserve shared definitions while recording historical names, source exceptions, and operational caveats.

  • Where did this data come from?

    Governed datasets and Dataplex zones keep source-shaped records, transformations, and published outputs traceable across the platform.

  • Who can use it?

    Dataset IAM bindings and row-level policies define access by department and use case instead of granting broad visibility by default.

  • Can this environment be reproduced?

    Pulumi defines datasets, ingestion assets, schedules, permissions, Dataplex resources, and supporting services through reviewed infrastructure changes.

Built once, reused across the business

  • Consistent business meaning

    Revenue, margin, customers, products, and transactions retain the same meaning across reporting and automation.

  • Controlled access by role and scope

    Departmental permissions and row-level policies keep analytical access deliberate, reviewable, and proportionate.

  • Build once, reuse across workflows

    Curated models provide starting points for dashboards, analysis, operational alerts, automation, and future systems.

Core customers, products, invoices, orders, and quotes are modeled once and reused across reporting, departmental data products, commercial analysis, inventory history, and operational alerting.

The transaction models cover more than a decade of commercial history, including thousands of invoices and millions of invoice lines. They also provide the input contracts for HAPPY END's data-driven alerting platform.

Source systems produce operational records. The lakehouse turns them into governed business information. Downstream systems use that information to drive action, and the resulting outcomes return as additional analytical context.

Outcome

HAPPY END gained an analytical platform that protects operational systems, preserves source knowledge, and gives reporting, data products, alerts, and automation the same reviewable definitions and access controls.

  • Source knowledge becomes maintainable

    Legacy operational knowledge is documented, normalized, and exposed through reusable analytical models.

  • Analytical workloads stay separate

    Reporting and analysis no longer compete directly with production systems and can scale independently.

  • Trusted models are reused faster

    Reports, alerts, automation, and downstream systems can start from trusted data products instead of rediscovering source rules.

Give shared data logic one production home.

Turn operational data into governed models that reporting, alerts, automation, and future systems can reuse.

0 / 2,000