Most early-stage 'data platforms' are a Postgres replica and a Notion doc. We build the real thing — Fabric on Azure, PySpark for transformation, a medallion lakehouse, and a semantic layer your product and AI features can both query against.
What the work involves
Ingestion
Batch and streaming ingestion into Fabric / OneLake
CDC from operational stores (Postgres, MongoDB, SAP)
Third-party SaaS and webhook integrations
Transformation
PySpark notebooks and Spark jobs for bronze → silver → gold
Python + dbt-style modelling on the gold layer
Data quality contracts and idempotent pipelines
Semantic layer & governance
Semantic models for BI and AI consumption
Lineage, catalog, and access control (Purview)
DataOps: CI for notebooks, versioning, environment promotion
What you get
Production lakehouse on Fabric (bronze / silver / gold)
PySpark pipelines under CI with data quality checks