Case study
Building a Governed Data Lakehouse for HAPPY END
A governed analytical foundation for reporting, automation, and operational systems.
Case study
A governed analytical foundation for reporting, automation, and operational systems.
HAPPY END EHS Solutions provides products and services for safe, sustainable operations across Central Europe. Its operational and commercial systems support a broad portfolio of customers, products, suppliers, orders, invoices, and sales activity.
Those systems held years of valuable business data, but their structures had grown around operational needs rather than analytical use. Tables and fields contained abbreviated Czech names, historical conventions, application-specific behavior, and exceptions that were understood mainly through experience.
Moving the data into the cloud was therefore only part of the work. HAPPY END needed to replicate it without burdening production systems, reconstruct reliable business concepts, preserve source knowledge, control access, and publish models that other teams and systems could reuse.
The result is a governed analytical foundation for reporting, departmental data products, operational alerting, automation, and future applications. Clients, products, transactions, revenue, margins, and other core concepts are defined once instead of being rebuilt for every new use case.
Reusable entities, transactions, facts, and marts built from legacy operational data.
Maintained alongside descriptions for 489 source tables through catalog-as-code.
Tables and views maintained across 74 production datasets.
Protect production systems while moving data through controlled ingestion paths.
Operational systems remain authoritative while analytical workloads run against a separate data platform.
Transactional replication, managed ingestion, and controlled pipelines bring operational and reference data into BigQuery.
Turn source-shaped records into documented, reusable business definitions.
Dataform converts legacy structures into reusable entities, transactions, facts, and marts.
Catalog-as-code, documentation, IAM, and row-level policies keep data traceable, understandable, and controlled.
Publish trusted outputs for reporting, alerts, automation, and future systems.
Curated data products become shared inputs for reporting, analysis, alerts, automation, and downstream systems.
The legacy environment was built to run the business, not to support unrestricted analytical workloads. Complex reporting queries risked competing with operational activity, while repeated extracts and locally maintained logic could produce different versions of the same figures.
Definitions of revenue, margin, clients, products, and transactions could vary depending on who prepared the data. Even experienced users relied on memory to interpret historical fields, country-specific identifiers, and source behavior that had never been documented formally.
HAPPY END therefore separated transaction processing from analysis. Operational systems continued serving daily work, while the lakehouse became the governed source for reporting, decision-making, alerting, and automation.
A one-way transactional replica separated analytical extraction from production, while CDC-based managed ingestion moved operational changes into BigQuery without placing heavy workloads on primary systems. External reference data followed controlled pipelines through Cloud Run, Workflows, and Cloud Scheduler. Dataform then converted source-shaped records into governed business entities that reports, alerts, and automated workflows could reuse.
Isolate data movement from production and make ingestion failures recoverable.
The replica and CDC pipeline let source systems remain authoritative without making analytical ingestion part of their operational workload.
Source-shaped records and explicit pipeline boundaries make failed loads easier to diagnose, replay, and validate without reconstructing history.
Replace source-specific structures with stable contracts for downstream consumers.
Consistent entities, identifiers, relationships, and measures insulate downstream systems from legacy schemas and source-specific conventions.
New reports, marts, alerts, and workflows can start from curated outputs instead of repeating source joins and transformation rules.
Each data product has explicit, reviewable definitions for meaning, lineage, ownership, and access. Documentation, permissions, and infrastructure are version-controlled alongside the platform, so those decisions remain visible as the system changes.
Version-controlled field and table descriptions preserve shared definitions while documenting historical names, source-specific exceptions, and operational caveats.
Governed datasets and Dataplex zones keep source-shaped records, transformations, and published outputs traceable across the platform.
Dataset IAM bindings and row-level policies define access by department and use case instead of granting broad visibility by default.
Pulumi defines datasets, ingestion assets, schedules, permissions, Dataplex resources, and supporting services through reviewable infrastructure changes.
Revenue, margin, clients, products, and transactions retain the same meaning across reporting and automation.
Departmental permissions and row-level policies keep analytical access deliberate, reviewable, and proportionate.
Curated models provide starting points for dashboards, analysis, operational alerts, automation, and future systems.
Core clients, products, invoices, orders, and quotes are modeled once and reused across reporting, departmental data products, commercial analysis, inventory history, and operational alerting.
The transaction models cover more than a decade of commercial history, including thousands of invoices and millions of invoice lines. They also provide the foundation for HAPPY END's data-driven alerting platform, which identifies relevant commercial conditions from governed data.
This creates a broader operational data loop. Source systems produce data, the lakehouse turns it into governed business information, downstream systems use it to drive action, and the resulting outcomes return as additional analytical context.
Reporting, departmental data products, analysis, and operational alerting now reuse the same governed definitions, while access remains controlled by department, role, and appropriate data scope.
Legacy operational knowledge is documented, normalized, and exposed through reusable analytical models.
Reporting and analysis no longer compete directly with production systems and can scale independently.
Reports, alerts, automation, and downstream systems can start from trusted data products instead of rediscovering source rules.
Turn legacy operational data into governed models that reporting, alerts, automation, and future systems can reuse.