Sri Omkar D
Open to Senior Data Engineer & Lakehouse Roles

Architecting Resilient Cloud Data Lakehouses.

I'm Sri Omkar D, a Senior Data Engineer with 7+ years of experience designing high-throughput batch and streaming architectures across AWS & Azure using PySpark, Databricks, and Python.

View System Architectures LinkedIn Email Me

7+ Years

Enterprise Experience

25–40%

Runtime Latency Reduction

Multi-TB

Daily Pipeline Scale

100%

Reconciled SLA Integrity

01. Architectural Case Studies

Production-tested designs focusing on throughput, reliability, and cost-efficiency.

STREAMING LAKEHOUSE Near-Real-Time

IoT Fleet Telemetry & Asset Engine

Ingested high-frequency JSON/XML telemetry payloads from 50,000+ rental units via AWS Kinesis. Orchestrated structured streaming via Databricks to feed S3 Delta Lake curated layers for operational Qlik Sense dashboards.

Pipeline: Kinesis → Databricks Streaming → S3 Curated → Qlik Sense
AWS Kinesis Databricks PySpark AWS S3
40% Faster executive dashboard refreshes
SPARK TUNING Performance Study

Spark Catalyst & Skew Mitigation

Benchmarked high-volume banking compliance jobs suffering from partition skew and shuffle overhead. Resolved bottlenecks via Spark UI analysis, salting keys, Adaptive Query Execution (AQE), and broadcast hash joins.

Optimization: Anti-Join Appends + Data Salting + AQE
Spark UI AQE Data Salting Redshift
35–40% Batch runtime reduction
INGESTION FRAMEWORK Enterprise ERP

Asynchronous ERP Ingestion Engine

Architected an API-driven extraction framework using Python, FastAPI, and SQLAlchemy to ingest multi-source enterprise client ERP files, enforcing strict schema validation and staging to Azure Data Lake Storage Gen2.

Stack: FastAPI → ADLS Gen2 → Azure Databricks → Snowflake
FastAPI ADLS Gen2 Snowflake Terraform
99.8% Schema reconciliation compliance

02. Professional Experience

Senior Data Engineer @ Nortek Consulting INC (Herc Rentals Inc.)

Nov 2025 - Present

Florida, USA • Fleet Operations & Logistics

  • Developed Python ingestion pipelines routing high-volume operational data from RentalMan, AS400, Oracle OLTP, and Teradata into AWS S3 data lakes.
  • Built near-real-time telemetry processing pipelines using AWS Kinesis and Databricks Streaming for GPS, engine metrics, and fleet utilization analytics.
  • Reduced daily batch runtimes by 25–30% by tuning PySpark partition strategies, optimizing Spark SQL joins, and shifting to incremental CDC load patterns.
  • Decreased pipeline rerun effort by 35% during failures by decoupling long Databricks transformation chains into checkpointed, modular stages.

Data Engineer @ Quess Corp (Blue Yonder India Pvt Ltd.)

June 2022 - July 2023

India • Supply Chain Optimization

  • Engineered high-throughput Azure Databricks pipelines, reducing batch transformation runtimes by 25–30% through unified PySpark transformations.
  • Built custom Python, FastAPI, and SQLAlchemy ingestion frameworks to ingest complex ERP payloads safely into ADLS Gen2.
  • Optimized Snowflake and Exasol analytical layers via advanced query profiling, partition tuning, and SQL performance tuning.
  • Managed orchestration workflows using Airflow alongside secure Terraform-deployed Azure environments (Key Vault, RBAC).

Data Engineer @ Trigent Software (Accenture)

Jan 2020 - June 2022

India • Banking & Regulatory Compliance

  • Cut runtime by 35–40% on large-scale financial pipelines by diagnosing Spark UI bottlenecks and applying data salting, AQE, and Broadcast Joins.
  • Modernized legacy Ab Initio workflows into Python-based S3 data lakes cataloged via AWS Glue Data Catalog.
  • Designed audit-ready Amazon Redshift dimensional models and structured rigorous source-to-target reconciliation checks.

ETL Developer @ Experis IT (Thomson Reuters)

June 2016 - Sep 2019

India • Sales & Revenue Analytics

  • Extracted enterprise SQL Server legacy data into S3 data lakes using Apache Sqoop on AWS EMR.
  • Built AWS Glue and PySpark pipelines to process sales, subscription, and revenue data for Redshift and Athena analytics.

03. Technical Competencies

Cloud & Distributed Compute

AWS (S3, EMR, Glue) Azure (ADLS Gen2) Apache Spark PySpark Databricks AWS Kinesis Athena HiveQL

Warehousing & Modeling

Amazon Redshift Snowflake PostgreSQL Teradata Microsoft SQL Server Star Schema SCD Type 2 CDC

DevOps & Tooling

Python SQL Apache Airflow FastAPI Docker Terraform GitHub Actions Spark UI Profiling

04. Education

GRADUATED 2024

Masters in Information Technology Management

California Baptist University • Riverside, CA

GRADUATED 2017

B. Tech in Computer Science and Engineering

JNTU-K University • India