Diwakar Vashisht
Data Engineer (Azure + Databricks) | 4+ Years Building Scalable Data Platforms
4 yrs experience · Gurugram, Haryana · <10 hrs/week
About
Data Engineer with 4+ years of experience designing and implementing scalable data platforms using Azure (Data Factory, Synapse Analytics, Fabric) and Databricks. Expertise in end-to-end data pipelines following Medallion architecture, reliable ingestion/transformation, analytics enablement, pipeline performance optimization, data quality improvements, and cost-efficient cloud data solutions. Holds 6 active Microsoft and Databricks certifications; has delivered large-scale ETL/streaming pipelines and enterprise GenAI RAG solutions, including team leadership and measurable latency/cost reductions.
Skills
Experience
- Data Engineer · Infosys Limited01-01-2022
Large-Scale Gaming Analytics Platform Designed 10+ ETL pipelines for a global gaming service, processing large-scale telemetry data for near real-time insights. ∗ Engineered end-to-end data pipelines leveraging Azure Data Factory and Azure Databricks to ingest, process, and transform 500+ GB of gaming telemetry data from diverse formats (CSV, JSON, XML, Parquet) sourced from on-premise and cloud storage (ADLS Gen2), enabling near real-time analytics for a global user base. ∗ Refined complex transformation logics in Spark SQL, resolving data quality issues through rigorous validation checks and automated monitoring, ensuring high reliability of downstream analytics. ∗ Spearheaded the implementation of near-real-time streaming ETL pipelines using Azure Event Hubs, reducing data latency from hours to minutes and supporting live dashboards for business stakeholders. ∗ Partnered to drive agile sessions with client architects to refine data requirements and system design for high-throughput scenarios, including peak sale events such as Black Friday and holiday seasons. Enterprise GenAI Retrieval Platform (RAG) Delivered 3 RAG approaches(Streamlit-LangChain RAG workflows, QnA completion pipelines, and Azure OpenAI Playground) with one adopted for enterprise deployment. ∗ Envisioned a Retrieval-Augmented Generation (RAG) platform using Streamlit, LangChain and Azure OpenAI, enabling semantic search over enterprise documents; collaborated with senior leadership and pre-sales stakeholders to evaluate deployment trade-offs, present client demos, and influence architectural decisions–reducing document retrieval time by 40% during internal testing. ∗ Deployed compliant GenAI infra using ARM-based IaC, enabling secure client demos and accelerating enterprise adoption (received Infosys Insta Award). ∗ Led a team of two data engineers, coordinating tasks and ensuring timely delivery of optimizations; architected automated cost-control measures‘ and query optimizations, delivering annual savings of $50k+. ∗ Identified and proposed adopting Microsoft Fabric to the client and Infosys pre-sales, initiating a new business conversation grounded in architectural best practices. Retail Distribution Data Modernization Processed Silver transformations reducing runtime by 33% across 10+ daily pipelines. ∗ Owned Silver layer by authoring Pyspark transformation logics using on Databricks, processing tables with 10+ billion records and 150+ columns, applying complex business logic across 30+ table joins, and implementing partitioning strategies, join optimizations, and skew mitigation techniques to improve performance and stability. ∗ Built dimensions and fact tables in Azure Synapse Analytics using star-schema modeling, enabling self-service BI for 10+ business users and reducing query response time by 30%. ∗ Implemented incremental load, upsert, and deduplication frameworks in Databricks to handle primary key constraints, improving data freshness by 25%. ∗ Optimized Spark jobs using broadcast joins, partitioning, and UDFs, reducing compute costs by 30% and cutting average pipeline runtime from 6 hours to 4 hours. ∗ Translated business and source system requirements into analytical data models using mapping documents and close collaboration with data architects, ensuring accuracy and consistency across curated layers.
Education
- Rajasthan Technical UniversityBachelor of Technology (B.Tech), Computer Science and Engineering
Similar talent on Pangea
Hire Diwakar through Pangea
Describe your project to the Pangea agent — see if Diwakar is a fit, with transparent pricing and interviews booked straight onto your calendar. No contact details change hands until you hire.
See if Diwakar is a fit