Sign up to access all features of our service
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Data Engineer — Supply Chain Lakehouse & Streaming Pipelines

HyrEzy Talent Solutions

Senior Data Engineer — Supply Chain Lakehouse & Streaming Pipelines

  • Location: Bengaluru / Hyderabad / NCR (Hybrid / Remote Flexibility)

  • Employment Type: Full-Time

  • Department: Data Engineering & Analytics Infrastructure

  • Experience Range: 6–10 Years

About the Company & Data Engineering Culture

At the heart of modern supply chain orchestration lies a massive volume of heterogeneous data—ranging from real-time GPS fleet telemetry and warehouse IoT sensor streams to asynchronous purchase orders and ERP ledger entries. As a high-growth Supply Chain Technology SaaS enterprise, our platform processes petabytes of operational data daily.

Our data engineering culture revolves around building immutable, fault-tolerant data pipelines, real-time stream processing architectures, and high-performance data lakehouse foundations. We empower our data science, operations research, and enterprise client reporting squads with pristine, low-latency data models. If you are passionate about architecting massive-scale data systems that drive real-world logistics optimization, this is your arena.

Position Overview

We are looking for an expert, hands-on Senior Data Engineer to design, build, and scale our next-generation data lakehouse and real-time streaming infrastructure. In this role, you will own the end-to-end data lifecycle—from ingestion and transformation to serving layers that power predictive analytics, machine learning feature stores, and client-facing supply chain visibility dashboards.

You will work closely with software architects, backend engineers, and data scientists to ensure high data quality, lineage tracking, and sub-second query performance across complex, multi-tenant enterprise datasets.

Key Responsibilities & Technical Ownership1. Real-Time Streaming & Ingestion Architecture

  • High-Throughput Pipelines: Design, build, and optimize real-time streaming data pipelines using Apache Kafka, Apache Flink, and Spark Streaming to process high-frequency logistics events and IoT telemetry.

  • Asynchronous Integration: Build scalable ingestion connectors for disparate external data sources, including third-party carrier APIs, EDI feeds, and enterprise ERP/WMS systems (SAP, Oracle, Blue Yonder).

  • Fault Tolerance & Resilience: Implement robust error-handling mechanisms, dead-letter queues, and automatic recovery protocols for streaming workloads to guarantee zero data loss during peak operational surges.

2. Lakehouse Infrastructure & Data Modeling

  • Modern Lakehouse Implementation: Architect and manage our enterprise data lakehouse using open table formats (Apache Iceberg / Delta Lake) built on top of cloud object storage (AWS S3 / GCP Cloud Storage).

  • Dimensional Modeling: Design scalable, optimized data warehouses and data marts using dimensional modeling principles (Kimball methodology) to support multi-tenant operational reporting.

  • Data Quality & Governance: Establish automated data quality frameworks, schema validation protocols, anomaly detection alerts, and end-to-end data lineage tracking.

3. Performance Tuning & Cost Optimization

  • Query Optimization: Tune expensive distributed queries, optimize shuffle operations, manage partitioning/clustering strategies, and configure caching layers in engines like Trino, Presto, or Snowflake.

  • FinOps Governance: Continuously monitor and optimize cloud compute and storage expenditures associated with heavy batch processing and continuous stream consumers.

Comprehensive Tech Stack & Technical RequirementsCore Technical Stack:

  • Languages: Advanced proficiency in Python, Scala, and SQL (ANSI SQL, complex window functions, performance tuning).

  • Stream Processing & Messaging: Expert hands-on experience with Apache Kafka, Apache Flink, Spark Streaming, or Kafka Connect.

  • Data Processing Frameworks: Apache Spark (PySpark), Ray, or distributed data processing engines.

  • Data Lakehouse & Warehousing: Open table formats (Apache Iceberg, Delta Lake), Snowflake, AWS Redshift, or Google BigQuery.

  • Orchestration & Infrastructure: Apache Airflow, Dagster, Prefect; Docker, Kubernetes, Terraform; AWS / GCP cloud environments.

Experience & Educational Qualifications:

  • Experience: 6 to 10 years of professional software engineering experience, with at least 4+ years dedicated exclusively to building large-scale data engineering architectures, data lakes, or real-time streaming platforms.

  • Education: Bachelor’s or Master’s degree in Computer Science, Data Science, Information Technology, or a related engineering discipline from a premier institution (IITs, NITs, IISc, or top-tier universities).

  • Domain Exposure: Prior experience in supply chain analytics, logistics tech, e-commerce fulfillment, or fintech transactional data processing is strongly preferred.

Competencies & Behavioral Traits

  • Obsession with Data Integrity: Uncompromising standards regarding data accuracy, consistency, and correctness across complex distributed pipelines.

  • Systems Mindset: Ability to view data systems holistically—understanding how source microservices changes propagate downstream into analytics models.

  • Collaborative Problem Solver: Eagerness to partner with data scientists, backend developers, and product teams to unblock data bottlenecks and deliver scalable solutions.

What We Offer

  • Scale & Impact: Direct ownership of data pipelines processing millions of tracking events daily, powering automated supply chain optimization across global markets.

  • Technical Autonomy: Freedom to evaluate, test, and implement modern open-source data technologies and cloud architectures.

  • Competitive Remuneration: Industry-leading salary packages, comprehensive medical benefits, learning stipends, and generous ESOP options.

Professional Interview Process

  • Initial Technical Screening: Deep-dive discussion into data architecture patterns, stream processing challenges, and past project scale.

  • Data Systems Design Round: Collaborative design exercise solving a large-scale data ingestion and lakehouse modeling scenario.

  • Coding & SQL Deep-Dive: Live coding assessment focusing on distributed data manipulation, performance optimization, and advanced SQL.

  • Leadership & Culture Fit Interview: Interaction with engineering leadership focusing on execution velocity, code quality standards, and cross-functional collaboration.

Vacancy posted more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Data Engineer — Supply Chain Lakehouse & Streaming Pipelines. Be the first to apply!