Senior Site Reliability Engineer (SRE) Engineer
Rs 20 - 25 lakhs p.a.Umanist Staffing LLC
Senior Site Reliability Engineer (SRE) Engineer
Location: Viman Nagar, Pune – Work From Office
Experience Overall(must have): 8 Years
CTC: Up to ₹25 LPA
Notice Period: Immediate Joiners Only within 15d or (if serving max 30days)
Working Hours: 3:00 PM – 12:00 AM, Monday to Friday
On-Call: 24/7 Production Support – On-Call Rotation Required
Employment Type: Full-Time
About the Role
We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to manage and improve the reliability, scalability, performance, security, and observability of mission-critical production environments.
The role requires strong hands-on expertise in Cloud, Kubernetes, DevOps automation, Monitoring & Observability, Incident Management, and SRE practices . The ideal candidate should be comfortable handling production incidents while also driving long-term initiatives around reliability, automation, scalability, and reduction of operational toil.
Must-Have Skills & Experience1. SRE & Production Operations
Relevant 7+ years of relevant experience in SRE / DevOps / Cloud Infrastructure / Production Engineering .
Hands-on experience with 24/7 production support and on-call operations .
Strong experience in incident management, troubleshooting, RCA, and post-mortems .
Good understanding of SLI, SLO, SLA, Error Budgets, MTTR, and reliability engineering .
Experience with toil reduction, capacity planning, high availability, disaster recovery, and failover strategies .
Ability to improve system availability, performance, scalability, and operational reliability.
2. Cloud & Infrastructure
Strong hands-on experience with Microsoft Azure, AWS, and/or GCP .
Strong understanding of cloud infrastructure, networking, IAM, storage, compute, and cloud-native services.
Hands-on experience with at least one major cloud platform and good exposure to multi-cloud environments.
Experience with:
Azure: VMs, Networking, Storage, IAM, Azure Monitor, AKS
AWS: EC2, S3, RDS, IAM, VPC, CloudWatch, EKS
GCP: Compute Engine, Cloud Storage, IAM, VPC, GKE, Cloud Monitoring
3. Kubernetes & Containerization
Strong hands-on experience with Kubernetes and containerized workloads.
Experience with AKS / EKS / GKE or equivalent Kubernetes environments.
Hands-on experience with Helm deployments.
Understanding of Kubernetes troubleshooting, scaling, networking, and workload management.
4. Infrastructure as Code & DevOps
Hands-on experience with Terraform / Infrastructure as Code (IaC) .
Experience with Git-based workflows using GitHub, GitLab, or Azure Repos .
Strong DevOps automation and CI/CD understanding.
Strong scripting skills in Python and/or Bash .
5. Monitoring & Observability
Strong hands-on experience with OpenTelemetry .
Experience with monitoring and observability tools such as:
Prometheus
Grafana
Datadog
Azure Monitor
AWS CloudWatch
GCP Cloud Monitoring
Strong understanding of metrics, logs, distributed tracing, and alerting .
Experience implementing monitoring based on Golden Signals :
Latency
Traffic
Errors
Saturation
Ability to develop symptom-based, user-impact-focused alerting.
6. Linux & Networking
Strong knowledge of Linux system administration .
Strong understanding of:
DNS
TCP/IP
Load Balancing
SSL/TLS
Networking fundamentals
Experience supporting highly available production environments.
7. Incident & Reliability Engineering
Ability to rapidly diagnose and resolve high-severity production incidents .
Experience driving MTTR reduction .
Strong debugging and analytical problem-solving skills.
Ability to identify recurring issues and implement permanent corrective/preventive solutions.
Good-to-Have Skills
Experience working across Azure + AWS + GCP in a multi-cloud environment.
Knowledge of Go (Golang) .
Experience with OpenSearch / ELK Stack .
Experience supporting AI/ML workloads in production.
Exposure to Azure AI Services and Azure AI Foundry .
Experience supporting RAG (Retrieval-Augmented Generation) workloads.
Experience designing infrastructure for AI/ML platforms.
Experience building enterprise-wide OpenTelemetry observability frameworks .
Strong understanding of distributed systems architecture.
Exposure to advanced cloud-native architectures and reliability patterns.
Experience with security, compliance, vulnerability remediation, secrets management, and network segmentation.
Key ResponsibilitiesProduction & Incident Management
Participate in the 24/7 on-call rotation .
Diagnose, mitigate, and resolve production incidents.
Lead RCA and post-incident reviews.
Implement corrective and preventive actions.
Continuously improve MTTR and production stability.
Reliability Engineering
Define and improve SLIs, SLOs, SLAs, and Error Budgets .
Identify and eliminate operational toil.
Conduct reliability and capacity reviews.
Improve redundancy, failover, disaster recovery, and system resilience.
Cloud & Infrastructure
Manage and optimize cloud infrastructure across Azure, AWS, and/or GCP.
Manage Kubernetes clusters and containerized applications.
Implement and maintain Infrastructure as Code using Terraform.
Support CI/CD and Git-based development workflows.
Observability & Performance
Build and improve monitoring, logging, metrics, and tracing.
Implement OpenTelemetry and distributed tracing .
Establish Golden Signals-based monitoring and alerting.
Identify and resolve infrastructure and application performance bottlenecks.
Security
Implement cloud security best practices around IAM, network segmentation, and secrets management .
Support vulnerability remediation and compliance initiatives.
Collaborate with Development, Security, and Infrastructure teams.
Ideal Candidate
We are looking for someone with:
Strong SRE mindset and production ownership .
Excellent troubleshooting and incident-management skills.
Hands-on expertise in Cloud + Kubernetes + Terraform + Observability .
Strong understanding of OpenTelemetry and Golden Signals .
Experience working in highly available, production-critical environments.
Ability to remain calm and make effective decisions during critical incidents.
Strong communication and cross-functional collaboration skills.
Passion for automation, scalability, reliability, and continuous improvement .
Important Hiring Criteria
Must be:
7+ years relevant experience
Immediate joiner
Willing to work from office in Viman Nagar, Pune
Comfortable with 3:00 PM – 12:00 AM shift
Comfortable with 24/7 on-call rotation
Strong hands-on SRE/DevOps experience
Strong Cloud + Kubernetes + Observability experience
Strong production incident management experience
Good to have:
Multi-cloud: Azure + AWS + GCP
OpenTelemetry
AI/ML or RAG production workloads
Azure AI / AI Foundry
Go
OpenSearch / ELK
Distributed systems
- Responsibilities Design and develop scalable software functionality for SAS products using React and TypeScript/JavaScript. Collaborate with managers, developers, user interface and visual designers, product managers, and distributed teams to implement business requirements...SuggestedFull timeFlexible hours
- ...implementation direction with product and engineering leaders. Lead design reviews and make... ...delivery risks, and recommendations to senior stakeholders. Requirements... ...region operations. Strong knowledge of reliability engineering practices, including service...SuggestedFull time
Rs 12 - 18 lakhs p.a.
.... Preferred (Experience 4) - Microsoft certifications in .NET technologies are preferred. Preferred (Experience 5) - Candidates from product companies, SaaS organizations, enterprise technology companies, IT services, or cloud-first engineering teams are preferred....SuggestedRs 12 - 26 lakhs p.a.
...Senior Engineer – Salesforce Marketing Cloud (SFMC) Location: Pune Experience: 5–8 Years Notice Period: Immediate to 30 Days preferred... ...with Salesforce and marketing technology teams to ensure reliable data flow and platform operations. Release, Change & Operational...SeniorFull timeImmediate start- ...Senior Manager/ Assistant Vice President/ Vice President - Sales, Health & Benefits ~202601105 ~Mumbai, Maharashtra, India ~Pune, Maharashtra, India ~Gurugram, Haryana, India ~Full time View favourites Description About the team Our Health & Benefits...SeniorFull time
- ...Job Description: Role Summary We are looking for a Senior Machine Learning Engineer who combines deep machine learning expertise with strong software... ...validated but architected and deployed as scalable, reliable products. Whether it's advancing NLP, optimisation, simulation...SeniorLong term contractFull timeHybrid workRelocation packageWork at officeLocal areaRemote jobRelocationFlexible hours
- ...deployment pipelines and production operations Set and uphold engineering standards for code quality, reviews, testing, observability... ...party integrations (payments, loyalty, identity, content) Drive reliability, performance and cost efficiency of the platform Mentor...Long term contractFull timeHybrid workRelocation packageWork at officeLocal areaRemote jobRelocationFlexible hours
- ...standalone deployment (monolithic) Event sourced data architecture / CQRS Oauth2 + OpenID Connect Azure experience of as a dev-ops engineer, including: AKS CosmosDB API for MongoDB Azure Service Bus Azure Storage Azure Application Gateway Azure Cache for...SeniorFull timeHybrid workRelocation packageWork at officeLocal areaRemote jobRelocationFlexible hoursShift work
- ...digital solutions and agile ways of working. Helpdesk Execution Senior Analyst is expected to act as technical SME for SAP/Coupa and... ...sophisticated problem or situation. Challenges assumptions and reliability of acquired information Decision Making - Makes decisions affecting...SeniorLong term contractFull timeTemporary workHybrid workRelocation packageWork at officeLocal areaRemote jobRelocationUS shiftFlexible hoursNight shift
- ...our outstanding team? Join our team, and develop your career in an encouraging, forward-thinking environment! Role: People Data Senior Specialist Seeking a highly analytical and detail-oriented People Data Senior Specialist to join our team at BP Business Solutions...SeniorFull timeHybrid workRelocation packageWork at officeLocal areaRemote jobRelocation
- ...relationships, grasps interdependencies, and reviews trends within a sophisticated problem or situation. Challenges assumptions and reliability of acquired information Decision Making - Makes decisions affecting both own tasks and those of others. Combines a variety of...SeniorFull timeTemporary workHybrid workRelocation packageWork at officeLocal areaRemote jobRelocationShift work
- ...Apart: Experience working in a manufacturing, industrial, engineering, or global supply chain environment. Knowledge of transportation... ...sustainably while improving productivity, energy security and reliability. With global operations and a comprehensive portfolio...SeniorFull timeNo agencyWorldwide
- ...Job Description Role Summary The R2R Senior Analyst – Lease is accountable for technical judgement and policy guidance across lease accounting under IFRS 16 / ASC 842, acting as a quality checkpoint between lease preparation and tower-level close governance. This...SeniorFull timeContract workLocal areaFlexible hours
- ...We are looking for an enthusiastic and inquisitive Senior Operations Support Engineer (Integration Platform) who will be responsible for operational... ...improvement of enterprise integration platforms. Ensures reliable and secure operation of more than 2,000 integrations and...Senior
- ...professional to join our team in the role of Associate Director, Software Engineering Principal responsibilities The Performance Testing... ...and stability. The role will work closely with engineering, SRE/DevOps, and product teams to identify bottlenecks early, drive performance...Permanent employmentFlexible hours
- ...seeking an experienced professional to join our team in the role of Senior Consultant Specialist In this role, you will: Lead... ...path. Remove impediments proactively by coordinating across engineering, QA, architecture, environments, release management, and third parties...SeniorPermanent employmentFlexible hours
- ...ambitions. We are currently seeking an experienced professional to join our team in the role of an Associate Director, Software Engineering In this role, you will Analyse Android/iOS mobile malware (static/dynamic), extract IOCs, and explain risks to technical and...Permanent employmentFlexible hours
- ...Job Description Bloom Engineering improves the environmental sustainability of industrial... ...must be able to identify and provide reliable solutions for all technical issues to assure... ...future projects • Travel to customer sites to provide assistance with field service,...Full timeWorldwide
- ...currently seeking a Business Consulting-DevOps Engineer to join our team in Pune, Madhya Pradesh... ...• Monitor system performance and ensure reliability • Manage cloud infrastructure and... ...hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective...Hybrid workWork at officeRemote jobFlexible hours
- ...opportunities and alternative courses of action to Senior Procurement/Sourcing leadership.... ...issue resolution projects. Support of engineering releases and other change management processes... ...; assimilates cross-functional, cross-site performance data across geographies; ensures...SeniorLong term contractPermanent employmentContract workTemporary workRelocation package
- Job Description We are looking for a skilled Angular Developer with 1-2 years of hands-on experience in Angular 16+ (preferably Angular 19) , Angular Signals , RxJS, and modern frontend development practices. You will be responsible for building high‑quality, scalable...Full timeLocal areaFlexible hours
- ...significantly enhance customer service and operational efficiency while supporting seamless omnichannel interactions. We seek a Senior Test Automation Engineer to join our Software Quality Engineering team. This role involves providing automation and test support for software...SeniorFull timeInternshipRemote jobRelocation
- ...ChampionAI is QAD | Redzone's agentic platform, purpose-built for manufacturing and used across all business units. As a junior-level Engineer, you will join the core platform team to build new platform capabilities, help Applied AI teams ship Champion agents correctly, and...Full timeLocal area
- ...Position Overview Job Title: Senior Data Platform Engineer Corporate Title: Assistant Vice President Location: Pune, India Role Description ~ This position sits within Deutsche Bank's eDiscovery and Archiving Technology group and will play a leading engineering...SeniorFull timeContract workFlexible hours
- ...capture flows. Embed and configure Wistia video assets while ensuring page performance. Implement landing pages, content updates, and site redesigns in collaboration with designers and product marketing managers. Support A/B testing by implementing front-end changes,...SeniorFull time
- ...Software Engineer, Business Systems Function: Business Technology & Security — Business Systems Location: Pune, India · Hybrid Reports to: Product Manager, Business Systems Experience: 2–4 years Short description Build and run the integrations, pipelines...Hybrid work
- ...Role Overview We are transitioning our organization to an AI-first operational model. We are seeking an AI Product Engineer to architect and build the company’s central intelligence engine. In this role, you will be the primary builder of automated intelligence systems...Full time
- ...exceptional experience for yourself, and a better working world for all. Job Title: ERP Analytics - Oracle EPM Reporting – PBCS - Senior Experience: 4 to 7 years Employment Type: Full-Time Job Summary We are looking for an experienced Oracle PBCS/EPBCS...SeniorFull time
- Design, develop, modify, and implement software programming for products (both internal and external) with focus on surpassing customer expectations, on achieving high quality and on-time delivery. Responsible for ensuring the overall functional quality of the released product...
- ...are automated, and solutions are built. Our solutions are mainly focused on Supply Chain and Commercial. We are looking for (senior) data engineers to join our high performing team. We aim to attract and further develop the best Data Science & Supply Chain talent. Key Activities...SeniorPermanent employment
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer (SRE) Engineer. Be the first to apply!
