EXL - Databricks Engineer
EXL Services.com ( I ) Pvt. Ltd.
Role Overview :We are looking for a skilled and passionate Databricks Engineer to design, build, and optimize enterprise-scale data lakehouse solutions on the Databricks platform. The successful candidate will be responsible for creating Databricks pipeline delivering Financial Crime platforms covering Anti-Money Laundering (AML), Know Your Customer (KYC), Customer Risk Assessment (CRA), Sanctions Screening, Transaction Monitoring, Fraud Detection, and Regulatory ReportingResponsibilities :- Design, build, and maintain Databricks workspaces, clusters, and compute pools across development, testing, and production environments.- Configure and manage Unity Catalog for data governance, fine-grained access control, permissions, metadata management, and data lineage.- Optimize Databricks cluster configurations, including instance types, auto-scaling, spot/preemptible nodes, and compute pools to improve performance and reduce costs.- Implement workspace best practices, including folder structures, access controls, secret management using Databricks Secrets, Azure Key Vault, or AWS Secrets Manager.- Create, schedule, and manage Databricks Jobs, Workflows, and multi-task job orchestration with dependency management.- Design and implement Delta Lake tables using partitioning, Z-Ordering, OPTIMIZE, VACUUM, and file compaction techniques.- Build and maintain Medallion Architecture (Bronze, Silver, and Gold layers) for scalable and governed data lakehouse solutions.- Develop Delta Live Tables (DLT) pipelines with built-in data quality expectations for reliable ETL/ELT processing.- Manage schema evolution, table versioning, Time Travel, and Change Data Feed (CDF) to support incremental data processing.- Design and implement lakehouse architectures integrating Delta Lake with cloud storage and external systems such as Azure Data Lake Storage (ADLS), Kafka, Event Hubs, and Kinesis.- Develop scalable batch and real-time data pipelines using PySpark, Spark SQL, Structured Streaming, and Delta Lake.- Build streaming ingestion pipelines from Kafka, Azure Event Hubs, and other streaming platforms into Delta tables.- Optimize PySpark applications using broadcast joins, Adaptive Query Execution (AQE), dynamic partition pruning, caching, and Photon Engine.- Develop reusable transformation frameworks, utility libraries, and pipeline templates to improve engineering productivity and standardization.- Implement robust error handling, retry mechanisms, logging, monitoring, and dead-letter queue (DLQ) patterns for production-grade pipelines.- Set up and manage MLflow experiment tracking, model registry, and model lifecycle management.- Support machine learning workloads by enabling scalable model training, inference, and GPU-based compute environments.- Develop feature engineering pipelines using Databricks Feature Store to create reusable and versioned machine learning features.- Enable Generative AI solutions, including Retrieval-Augmented Generation (RAG), vector search, LLM fine-tuning, and Mosaic AI capabilities.- Implement MLOps best practices, including model versioning, model deployment, A/B testing, and Databricks Model Serving.- Integrate Databricks with Azure Data Lake Storage (ADLS) and other cloud-native services.- Develop and maintain CI/CD pipelines using Azure DevOps, GitHub Actions, or GitLab CI for Databricks notebooks, jobs, and workflows.- Automate Databricks infrastructure deployment using Databricks Asset Bundles (DABs), Terraform, and Infrastructure-as-Code (IaC) practices.- Build and manage data ingestion frameworks using Auto Loader, COPY INTO, and third-party integration tools such as Fivetran, dbt, and Airbyte.- Monitor pipeline execution, cluster utilization, system performance, and cloud costs using Databricks system tables and cloud monitoring tools.- Implement row-level security, column-level masking, dynamic views, and governance policies using Unity Catalog.- Enforce data quality through Delta Live Tables expectations and Great Expectations frameworks.- Perform query optimization, execution plan analysis, caching strategies, and performance tuning to improve workload efficiency.- Maintain enterprise data cataloging, metadata management, and end-to-end data lineage.- Prepare technical documentation, architecture diagrams, operational runbooks, and standard operating procedures for Databricks platform and data engineering solutions.Qualifications :- Bachelor's or master's degree in computer science, Information Technology, Data Engineering, or related field.- 6+ years of total experience in data engineering or software engineering.- 3+ years of dedicated hands-on experience with the Databricks platform in production environments.- Strong background in big data engineering, cloud data platforms, and distributed computing.- Deep expertise in Databricks Workspaces, Clusters, Jobs, Workflows, and Repos.- Proficiency with Unity Catalog - metastore setup, catalog/schema/table management, access controls, and data lineage.- Hands-on experience with Delta Live Tables (DLT) - pipeline development, expectations, and monitoring.- Strong command of Delta Lake internals - transaction log, ACID guarantees, file layout, and optimization techniques.- Experience with Databricks SQL Warehouses, SQL Analytics, and dashboard creation.- Knowledge of Databricks Photon engine, serverless compute, and cost optimization strategies.- 4+ years of PySpark development - Dataframe, Datasets, Spark SQL, RDD operations.- Expert-level SQL - window functions, lateral joins, CTEs, recursive queries, and analytical functions.- Experience with Spark performance tuning - AQE, query plans (EXPLAIN), partitioning, and caching.- Proficiency with Python for pipeline development, utilities, and automation.- Hands-on experience with at least one: Azure (ADLS Gen2, ADF, Azure Databricks), AWS (S3, EMR, Glue, AWS Databricks), or GCP (GCS, BigQuery, Dataproc).- Experience with cloud networking for Databricks: VNet/VPC injection, private endpoints, and firewall configurations.- Familiarity with IAM roles, managed identities, and service principal authentication for Databricks.MLflow & ML Engineering (Nice to Have) :- Working knowledge of MLflow - experiment tracking, model registry, and deployment.- Experience supporting ML pipelines on Databricks for training, evaluation, and serving.- Exposure to Databricks Feature Store and Mosaic AI / GenAI capabilities. (ref:hirist.tech)
- Job Title : Azure Databricks & Microsoft Dynamics 365 (D365) EngineerLocation : PuneExperience : 46 YearsJob Summary :We are looking for a skilled Azure Databricks & Microsoft Dynamics 365 (D365) Engineer with strong technical expertise in Azure Data Engineering and Microsoft...Suggested
- Job Description : We are seeking an innovative, forward-thinking Data Engineer with a strong footprint in Azure Databricks and Artificial Intelligence (AI) solutions. This position is a full-time, permanent remote role based out of Pune, India, within our IT Services & Consulting...SuggestedPermanent employmentFull timeRemote job
- Job Summary:We are seeking an experienced Databricks ETL Testing Engineer with 6+ years of experience in validating enterprise data platforms and ETL pipelines. The ideal candidate should have hands-on expertise in Databricks, ETL testing, data validation, SQL, and cloud-based...Suggested
- ...YearsAre you passionate about building modern data platforms on Microsoft Azure? We're looking for experienced Microsoft Azure Databricks Engineers to join a high-performing team delivering enterprise-scale data engineering and analytics solutions.If you have expertise in Azure...Suggested
- About the Role : We are seeking an analytical and highly skilled Data Engineer with a strong background in Databricks and deep domain experience in Manufacturing or Supply Chain Planning. In this role, you will be part of our software engineering division, designing and executing...Suggested
- Key Responsibilities :- Design, develop, and deploy data engineering solutions using Azure Databricks and Azure cloud services.- Build, optimize, and maintain scalable ETL/ELT pipelines using PySpark, SQL, and Azure Data Factory.- Develop and maintain high-performance data pipelines...
- ...PRIVATE LIMITED is a technology services firm specializing in data engineering, cloud transformation, and advanced analytics. We partner with... ...the finance, retail, and healthcare sectors.Role Overview:As a Databricks Developer, you will be responsible for designing and...Hybrid work
- Job Title : Databricks on AWS and PySpark EngineerJob SummaryWe're seeking an experienced Databricks on AWS and PySpark Engineer to join our team. The ideal candidate will have a strong background in designing, building, and maintaining large-scale data pipelines and architectures...
- ...with enterprise clients.- Solution Design: Architect end-to-end Databricks Lakehouse solutions covering ingestion, transformation,... ...Core Skills & Requirements:- Experience: 8 to 14 years in data engineering, architecture, or consulting roles.- Big Data Expertise: Apache...
- Position : Databricks Architect.Purpose of the Position :To lead the design, development, and implementation of scalable data solutions using... ...Spark jobs.- Stakeholder Collaboration: Work closely with data engineers, analysts, and business stakeholders to translate requirements...
- ...Overview :NTT DATA Americas is seeking a highly skilled Senior Data Engineer to support a Teradata Utilization Analysis engagement, which... ...candidate combines deep Teradata DBA expertise with hands-on Databricks engineering capability and applied AI/ML skills. The analysis will...
- Databricks Admin - ArchitectLocation : Only PuneExperience : 11-18 YearsMandatory Skills : Databricks Administration, Unity Catalog, TerraformCore Platform Skills :- Databricks workspace and account administration- Cluster management, policies, and SQL warehouses- Strong Unity...
- Job Title : DataOps Technical Lead Databricks/AzureWork Location : Bangalore / Chennai / Gurgaon/ Pune/ Kolkata / HyderabadDivision/Department... ...pipelines. Should understand complex data pipelines and data engineering concepts. Mentoring and guiding junior team members.- Data...Full timeUS shiftShift work
- Job Title: Data Engineer Azure DatabricksLocation: Pan India / HybridExperience: 4-10 YearsMandatory Skills: Azure, Azure Databricks & BI Tool (Power BI/Tableau)Job Summary:We are looking for a Data Engineer with strong hands-on experience in Azure data services and Databricks...
Rs 15 - 20 lakhs p.a.
? Hiring Alert | Senior Platform Engineer – Python / Airflow / Databricks / AWS ? ? Location: Pune ? Experience: 8–10 Years ? Budget: ₹19lpa ? Work Mode: Work from Office (5 Days Mandatory) ⏰ Shift Timing: 12:00 PM – 9:00 PM IST ? Contract Duration: 6...Full timeContract workWork at officeShift workRs 15 - 18 lakhs p.a.
...Hiring Alert | Senior Platform Engineer – Python / Airflow / Databricks / AWS ? ? Location: Pune ? Role: Senior Platform Engineer ? Experience: 8–10 Years ? Work Mode: Work from Office (5 Days Mandatory) ⏰ Shift Timing: 12:00 PM – 9:00 PM IST ? Contract...Full timeContract workWork at officeShift work- ...and written communication skills in English to join our team.Job Description : We are looking for a Senior Data Engineer with strong expertise in Azure Databricks, PySpark, and distributed computing to develop and optimize scalable ETL pipelines for manufacturing analytics.The...Full time
- ...-time data pipelines.- Build and optimize data solutions using Databricks, PySpark, Delta Lake, and Apache Kafka.- Develop ETL/ELT workflows... ..., and cost efficiency.- Collaborate with Data Scientists, ML Engineers, Product teams, and business stakeholders to deliver reliable...
- Experience:- 5 to 10 years of experience in Data Engineering.Key Skills Required:- Strong hands-on experience with Databricks + Streaming.- PySpark.- SQL.- Delta Lake.- Unity Catalog.- Databricks Workflows / Jobs orchestration.- Experience with Azure Data Factory (ADF).- Experience...
- ...build, and operate scalable and reliable data pipelines on the Databricks platform.- Develop end-to-end data workflows from ingestion through... ...to minimize disruption during code transitions.Data Engineering Excellence : - Implement data quality checks and validation frameworks...
- ...solve complex business challenges at scale.Role Overview:As a Data Engineer based in Pune, you will be responsible for designing, building,... ...and maintain high-performance data processing solutions using Databricks and Delta Lake to support large-scale analytical workloads.-...Hybrid work
- About the Role:We are seeking a highly capable, hands-on Lead Data Engineer (Manager) with deep expertise in Databricks to lead the design and delivery of scalable data solutions.In this role, you will:- Lead engineering teams across distributed environments- Own end-to-end...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to EXL - Databricks Engineer. Be the first to apply!
