Alan Almeida

Alan Almeida

Analyst

Borivali (W), Mumbai - 400103, Maharashtra, IndiaData Engineering
1roles
46skills
6education
7credentials

About

Data Engineer with hands-on experience designing and building end-to-end ETL and ELT data pipelines on Azure and Microsoft Fabric. Proficient in PySpark, Apache Spark, Azure Data Factory, Azure Databricks, ADLS Gen2, Synapse Analytics, Delta Lake, and Fabric Lakehouse.

Experience

Analyst

Capgemini

Sep 2025 - Apr 2026 · Mumbai

Completed enterprise software training covering Java, Selenium test automation, SQL, QA testing methodologies, test case design, and software quality assurance in an SDLC context. Independently pursued and obtained four cloud certifications: AZ-900, DP-900, DP-700, and GCP ACE.

Education

St. Francis Institute of Technology

B.E., Information Technology

2021 - 2025

66.59%

T.P. Bhatia College of Science

12th HSC

2019 - 2021

80.33%

St. Lawrence High School

10th SSC

2019

66.40%

St. Francis Institute Of Technology

Information Technology, Engineering

2021 - 2025

PrepInsta

Aug 2023

Coursera

Skills

PythonSQLApache SparkPySparkDelta LakeParquetSpark Structured StreamingAzure Data Factory (ADF)ADLS Gen2Azure DatabricksSynapse AnalyticsApp RegistrationOAuth2Service PrincipalMicrosoft FabricOneLakeFabric LakehouseFabric Data FactoryDirect LakeSemantic ModelPower BIMySQLMongoDBSQL ServerETLELTData PipelineData LakeData WarehouseMedallion ArchitectureIncremental LoadStar SchemaData ModelingGitGitHubGoogle ColabREST APIDatabricks NotebooksMicrosoft Power BIAzure Data LakeAzure Data FactoryApache KafkaApache AirflowData EngineeringPython (Programming Language)Cloud Computing

Projects

Azure End-to-End Data Engineering Pipeline

Designed and built a cloud-native ETL pipeline on Azure using Medallion Architecture, ingesting eight datasets from GitHub REST API, MySQL, and MongoDB into ADLS Gen2 through a metadata-driven Azure Data Factory pipeline. Performed PySpark transformations in Azure Databricks, implemented star schema joins and MongoDB enrichment, secured access with OAuth2 Service Principal, and served the Gold layer through Azure Synapse Analytics external tables.

Microsoft Fabric End-to-End Lakehouse Pipeline

Built a Medallion Lakehouse on Microsoft Fabric, ingesting Parquet data into OneLake through Fabric Data Factory and transforming it into partitioned Delta tables with PySpark. Implemented Delta Lake MERGE-based incremental loading and built a Direct Lake star-schema Semantic Model for Power BI.

Wide World Importers — Microsoft Fabric End-to-End Project

An end-to-end data engineering and analytics solution built on Microsoft Fabric, implementing a full Medallion Lakehouse architecture (Bronze → Silver → Gold) using PySpark notebooks, Data Factory pipelines, a Direct Lake Semantic Model, and a Power BI report — all within a single unified OneLake environment.

Olist Brazilian E-Commerce — Azure End-to-End Data Engineering Pipeline

A production-style cloud data pipeline built on Azure, ingesting raw e-commerce data from multiple heterogeneous sources, transforming it with distributed compute, and serving a clean analytical dataset using a Medallion (Bronze → Silver → Gold) Lakehouse architecture.