Senior data engineer with 7 years building lakehouse platforms on Spark for mobility and banking. Makes data reliable, cheap and fast enough that analysts stop waiting, and treats pipelines like products with owners, contracts and tests. Databricks-certified.
- Migrated a 400 TB warehouse to a Delta Lake lakehouse, cutting compute costs 38% and query times from minutes to seconds.
- Built streaming pipelines in Spark Structured Streaming delivering trip data within 2 minutes instead of next day.
- Introduced data contracts and tests for 120 tables; broken dashboards fell from weekly to rare.
- Designed a self-serve ingestion framework that let analysts add new sources in hours, not weeks.
- Led a team of 4 engineers and ran the data platform's quarterly roadmap.
- Cut pipeline failures 70% by adding retries, idempotent writes and clear ownership for every job.
- Documented lineage for 200 critical tables so analysts could trace every metric to its source.
- Designed the first Airflow platform, now running 300 daily jobs.
- Cut the nightly ETL window from 7 to 2 hours with partitioning and incremental loads.
- Built GDPR deletion pipelines that processed 15,000 customer requests a year automatically.
- Partnered with risk analysts to deliver a credit data mart used in 4 regulatory reports.
- Reduced cloud storage costs €60k a year by archiving cold data to cheaper tiers.
- Mentored 2 analysts who moved into data engineering roles.
- Built 25 SQL Server reports for branch managers across 120 branches.
- Optimised slow stored procedures, cutting report runtimes by an average of 70%.
- Documented the data model for 80 core tables, onboarding new analysts twice as fast.
- Automated data-quality alerts that caught 30 broken loads before users noticed.