Skip to content
CBN data localisation deadline, 1 Jan 2027: 99d 00h 25m left.Talk to us
NuxFamily

Insurance · A leading European insurer

A governed lakehouse for an insurer's claims and pricing data

The insurer's analytics ran on an ageing warehouse and a set of unmanaged data lakes with no lineage. NuxFamily delivered a single on-premises lakehouse with a catalogue, lineage and access control that satisfied the internal audit function.

  • 100 %

    Of regulated reports with end-to-end lineage in the catalogue

  • Faster actuarial model runs compared with the legacy warehouse

  • 0

    Reports lost or rewritten during the transition

The challenge

The challenge

Claims, policy and pricing data lived in a legacy data warehouse that could no longer scale, plus several departmental data lakes built on object storage without a catalogue. Actuarial models depended on manual extracts, and nobody could answer with confidence where a given figure had come from.

Internal audit and the data protection office required lineage, role-based access at column level and a retention policy that could actually be enforced. The insurer also wanted the platform on-premises for cost and data-residency reasons.

The migration had to preserve existing reports and models during the transition. Analysts would not accept a period without their dashboards.

The architecture

The architecture

The warehouse layer is Greenplum, later moved to Apache Cloudberry, deployed as a multi-segment cluster on bare metal with mirrored segments for availability. It holds the curated, modelled data used by actuarial and finance teams and answers the existing reporting workload with no query rewrites.

Raw and semi-structured data lands in S3-compatible object storage as Apache Iceberg tables. Iceberg gives schema evolution, time travel and snapshot isolation, and it lets the warehouse and Spark-based pipelines read the same tables without copies.

OpenMetadata provides the catalogue, column-level lineage from source systems through pipelines to dashboards, data quality tests and the ownership model the audit function asked for. Access policies are defined once in the catalogue and enforced in the query engines.

OpenSearch indexes claims documents and pipeline logs for search and operational monitoring. The whole platform runs on Kubernetes for the service components and on dedicated nodes for the database engines, with backups and disaster recovery tested quarterly.

Technologies

  • Greenplum
  • Apache Cloudberry
  • Apache Iceberg
  • OpenMetadata
  • OpenSearch
  • Apache Spark
  • Kubernetes
For the first time we could show the auditors where a number came from, step by step, without a spreadsheet.
Head of Data, European insurer

Your institution could be the next case

Tell us about your platform and regulatory timeline; we will show you which of these patterns applies.

No mailing lists, no automated follow-ups. We reply personally within one working day.