Case study

Building a Governed Data Lakehouse for HAPPY END

A governed analytical foundation for reporting, automation, and operational systems.

On this page

Overview

HAPPY END EHS Solutions provides products and services for safe, sustainable operations across Central Europe. Its operational and commercial systems support a broad portfolio of customers, products, suppliers, orders, invoices, and sales activity.

Those systems held years of valuable business data, but their structures had grown around operational needs rather than analytical use. Tables and fields contained abbreviated Czech names, historical conventions, application-specific behavior, and exceptions that were understood mainly through experience.

Moving the data into the cloud was therefore only part of the work. HAPPY END needed to replicate it without burdening production systems, reconstruct reliable business concepts, preserve source knowledge, control access, and publish models that other teams and systems could reuse.

The result is a governed analytical foundation for reporting, departmental data products, operational alerting, automation, and future applications. Clients, products, transactions, revenue, margins, and other core concepts are defined once instead of being rebuilt for every new use case.

Platform scale and reuse

Curated business models
208

Reusable entities, transactions, facts, and marts built from legacy operational data.

Documented source fields
7,374

Maintained alongside descriptions for 489 source tables through catalog-as-code.

Governed BigQuery objects
886

Tables and views maintained across 74 production datasets.

How legacy data becomes a governed data product

Operational sources

Protect production systems while moving data through controlled ingestion paths.

  • Source

    Operational systems remain authoritative while analytical workloads run against a separate data platform.

  • Ingest

    Transactional replication, managed ingestion, and controlled pipelines bring operational and reference data into BigQuery.

Governed data foundation

Turn source-shaped records into documented, reusable business definitions.

  • Model

    Dataform converts legacy structures into reusable entities, transactions, facts, and marts.

  • Govern

    Catalog-as-code, documentation, IAM, and row-level policies keep data traceable, understandable, and controlled.

Reusable data products

Publish trusted outputs for reporting, alerts, automation, and future systems.

  • Serve

    Curated data products become shared inputs for reporting, analysis, alerts, automation, and downstream systems.

Why the legacy environment needed a separate analytical layer

The legacy environment was built to run the business, not to support unrestricted analytical workloads. Complex reporting queries risked competing with operational activity, while repeated extracts and locally maintained logic could produce different versions of the same figures.

Definitions of revenue, margin, clients, products, and transactions could vary depending on who prepared the data. Even experienced users relied on memory to interpret historical fields, country-specific identifiers, and source behavior that had never been documented formally.

HAPPY END therefore separated transaction processing from analysis. Operational systems continued serving daily work, while the lakehouse became the governed source for reporting, decision-making, alerting, and automation.

Safe ingestion and governed modeling

A one-way transactional replica separated analytical extraction from production, while CDC-based managed ingestion moved operational changes into BigQuery without placing heavy workloads on primary systems. External reference data followed controlled pipelines through Cloud Run, Workflows, and Cloud Scheduler. Dataform then converted source-shaped records into governed business entities that reports, alerts, and automated workflows could reuse.

Operational foundation

Isolate data movement from production and make ingestion failures recoverable.

  • Decouple change capture

    The replica and CDC pipeline let source systems remain authoritative without making analytical ingestion part of their operational workload.

  • Make ingestion recoverable

    Source-shaped records and explicit pipeline boundaries make failed loads easier to diagnose, replay, and validate without reconstructing history.

Reusable data products

Replace source-specific structures with stable contracts for downstream consumers.

  • Stabilize business contracts

    Consistent entities, identifiers, relationships, and measures insulate downstream systems from legacy schemas and source-specific conventions.

  • Separate logic from consumption

    New reports, marts, alerts, and workflows can start from curated outputs instead of repeating source joins and transformation rules.

The platform makes ownership and meaning explicit

Each data product has explicit, reviewable definitions for meaning, lineage, ownership, and access. Documentation, permissions, and infrastructure are version-controlled alongside the platform, so those decisions remain visible as the system changes.

  • What does this field mean?

    Version-controlled field and table descriptions preserve shared definitions while documenting historical names, source-specific exceptions, and operational caveats.

  • Where did this data come from?

    Governed datasets and Dataplex zones keep source-shaped records, transformations, and published outputs traceable across the platform.

  • Who can use it?

    Dataset IAM bindings and row-level policies define access by department and use case instead of granting broad visibility by default.

  • Can this environment be reproduced?

    Pulumi defines datasets, ingestion assets, schedules, permissions, Dataplex resources, and supporting services through reviewable infrastructure changes.

Built once, reused across the business

  • Consistent business meaning

    Revenue, margin, clients, products, and transactions retain the same meaning across reporting and automation.

  • Controlled access by role and scope

    Departmental permissions and row-level policies keep analytical access deliberate, reviewable, and proportionate.

  • Build once, reuse across workflows

    Curated models provide starting points for dashboards, analysis, operational alerts, automation, and future systems.

Core clients, products, invoices, orders, and quotes are modeled once and reused across reporting, departmental data products, commercial analysis, inventory history, and operational alerting.

The transaction models cover more than a decade of commercial history, including thousands of invoices and millions of invoice lines. They also provide the foundation for HAPPY END's data-driven alerting platform, which identifies relevant commercial conditions from governed data.

This creates a broader operational data loop. Source systems produce data, the lakehouse turns it into governed business information, downstream systems use it to drive action, and the resulting outcomes return as additional analytical context.

Outcome

Reporting, departmental data products, analysis, and operational alerting now reuse the same governed definitions, while access remains controlled by department, role, and appropriate data scope.

  • Source knowledge becomes maintainable

    Legacy operational knowledge is documented, normalized, and exposed through reusable analytical models.

  • Analytical workloads stay separate

    Reporting and analysis no longer compete directly with production systems and can scale independently.

  • Trusted models are reused faster

    Reports, alerts, automation, and downstream systems can start from trusted data products instead of rediscovering source rules.

Stop rebuilding the same data logic.

Turn legacy operational data into governed models that reporting, alerts, automation, and future systems can reuse.