Azure End-to-End Data Engineering Pipeline
Designed and built a cloud-native ETL pipeline on Azure using Medallion Architecture, ingesting eight datasets from GitHub REST API, MySQL, and MongoDB into ADLS Gen2 through a metadata-driven Azure Data Factory pipeline. Performed PySpark transformations in Azure Databricks, implemented star schema joins and MongoDB enrichment, secured access with OAuth2 Service Principal, and served the Gold layer through Azure Synapse Analytics external tables.
Microsoft Fabric End-to-End Lakehouse Pipeline
Built a Medallion Lakehouse on Microsoft Fabric, ingesting Parquet data into OneLake through Fabric Data Factory and transforming it into partitioned Delta tables with PySpark. Implemented Delta Lake MERGE-based incremental loading and built a Direct Lake star-schema Semantic Model for Power BI.
Wide World Importers — Microsoft Fabric End-to-End Project
An end-to-end data engineering and analytics solution built on Microsoft Fabric, implementing a full Medallion Lakehouse architecture (Bronze → Silver → Gold) using PySpark notebooks, Data Factory pipelines, a Direct Lake Semantic Model, and a Power BI report — all within a single unified OneLake environment.
Olist Brazilian E-Commerce — Azure End-to-End Data Engineering Pipeline
A production-style cloud data pipeline built on Azure, ingesting raw e-commerce data from multiple heterogeneous sources, transforming it with distributed compute, and serving a clean analytical dataset using a Medallion (Bronze → Silver → Gold) Lakehouse architecture.