Case study
Building a Governed Data Lakehouse for HAPPY END
A governed BigQuery foundation for reporting, automation, and operational systems.
Case study
A governed BigQuery foundation for reporting, automation, and operational systems.
HAPPY END EHS Solutions provides products and services for safe, sustainable operations across Central Europe. Its commercial and operational systems manage customers, products, suppliers, orders, invoices, and sales activity across several markets.
Those systems contained years of useful data, but their schemas reflected application history rather than analytical contracts. Abbreviated Czech names, country-specific identifiers, historical conventions, and undocumented exceptions made correct interpretation depend heavily on experience.
Copying tables into the cloud would have preserved the ambiguity. HAPPY END needed to protect production systems, retain source evidence, define stable business entities, document what the data meant, control access, and publish outputs other teams could reuse.
The result is a separate analytical platform where core concepts such as customers, products, transactions, revenue, and margin are defined once and consumed by reporting, departmental data products, commercial analysis, operational alerting, and automation.
Protect production systems while moving data through controlled ingestion paths.
Operational systems remain authoritative while analytical workloads run against a separate platform.
Transactional replication, managed ingestion, and controlled pipelines move operational and reference data into BigQuery.
Turn source-shaped records into documented, reusable business definitions.
Dataform turns source-shaped records into reusable entities, transactions, facts, and decision-specific marts.
Catalog-as-code, documentation, IAM, and row-level policies keep meaning, lineage, ownership, and access reviewable.
Publish trusted outputs for reporting, alerts, automation, and future systems.
Curated data products become shared inputs for reporting, analysis, alerts, automation, and downstream systems.
The legacy environment was built to run the business. Complex analytical queries risked competing with operational work, while repeated extracts and locally maintained calculations created several versions of the same business concept.
Revenue, margin, customer, product, and transaction logic varied by consumer. Experienced users also relied on memory to interpret historical fields, market-specific identifiers, and source behavior that had never been documented as a shared contract.
HAPPY END separated transaction processing from analysis. Operational systems continued to own daily work, while the lakehouse became the controlled source for analytical models, reporting, alerting, and automation.
A one-way transactional replica isolated analytical extraction from production. CDC-based managed ingestion moved operational changes into BigQuery, while external reference data followed controlled pipelines through Cloud Run, Workflows, and Cloud Scheduler. Dataform then resolved source structures into business entities and outputs that reports, alerts, and automated workflows could consume.
Isolate data movement from production and give ingestion a recoverable operating path.
The replica and CDC pipeline let source systems remain authoritative without turning analytical ingestion into part of their production workload.
Source-shaped records, extraction metadata, and explicit pipeline boundaries make failed loads easier to diagnose, replay, and validate.
Replace source-specific structures with stable contracts for downstream consumers.
Consistent entities, identifiers, relationships, history, and measures insulate consumers from legacy schemas and source conventions.
Reports, marts, alerts, and workflows start from curated outputs instead of repeating source joins and transformation rules.
Each data product has explicit definitions for grain, lineage, ownership, access, and lifecycle. Documentation, permissions, models, and infrastructure are version-controlled so the platform carries its own explanation as it changes.
Version-controlled field and table descriptions preserve shared definitions while recording historical names, source exceptions, and operational caveats.
Governed datasets and Dataplex zones keep source-shaped records, transformations, and published outputs traceable across the platform.
Dataset IAM bindings and row-level policies define access by department and use case instead of granting broad visibility by default.
Pulumi defines datasets, ingestion assets, schedules, permissions, Dataplex resources, and supporting services through reviewed infrastructure changes.
Revenue, margin, customers, products, and transactions retain the same meaning across reporting and automation.
Departmental permissions and row-level policies keep analytical access deliberate, reviewable, and proportionate.
Curated models provide starting points for dashboards, analysis, operational alerts, automation, and future systems.
Core customers, products, invoices, orders, and quotes are modeled once and reused across reporting, departmental data products, commercial analysis, inventory history, and operational alerting.
The transaction models cover more than a decade of commercial history, including thousands of invoices and millions of invoice lines. They also provide the input contracts for HAPPY END's data-driven alerting platform.
Source systems produce operational records. The lakehouse turns them into governed business information. Downstream systems use that information to drive action, and the resulting outcomes return as additional analytical context.
HAPPY END gained an analytical platform that protects operational systems, preserves source knowledge, and gives reporting, data products, alerts, and automation the same reviewable definitions and access controls.
Legacy operational knowledge is documented, normalized, and exposed through reusable analytical models.
Reporting and analysis no longer compete directly with production systems and can scale independently.
Reports, alerts, automation, and downstream systems can start from trusted data products instead of rediscovering source rules.
Turn operational data into governed models that reporting, alerts, automation, and future systems can reuse.