Sign up to access all features of our service
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer - Imunify Reliability Platform

India
  • Remote job

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Site Reliability Engineer - Imunify Reliability Platform based in India.

This is a greenfield SRE leadership opportunity within a large-scale, security-focused product environment. You will define what “healthy” means across roughly 70 components spanning cloud services and customer-hosted agents. You will establish SLIs, SLOs, error budgets, monitoring standards, alerting, and escalation practices from the ground up. Your work will directly improve the ability to detect silent security-control degradation before it becomes a widespread customer-impacting issue. You will collaborate closely with engineering leads and senior engineers while building the telemetry and reliability platform yourself. The environment is remote-first, async, technically demanding, and focused on measurable outcomes rather than dashboards for their own sake. This is an opportunity to shape the reliability culture and foundations of a major security product.

Accountabilities
  • Define and establish meaningful SLIs for approximately 70 product components, working with squad leads and senior engineers to agree on ownership, measurement, tiering, SLOs, and error budgets.
  • Develop a reliability taxonomy covering service availability and latency, fleet reachability and configuration convergence, security-control efficacy, artifact delivery, and telemetry pipeline health.
  • Ensure reliability indicators are independently measurable and cannot be disabled by the same failure they are intended to detect.
  • Design and build the telemetry collection pipeline for customer-hosted agents and cloud services, balancing push-based collection, sampling, privacy constraints, data quality, and cardinality.
  • Extend instrumentation across Python, Go, and Rust components in collaboration with product engineering teams.
  • Consolidate existing dashboards, queries, and reporting mechanisms into a smaller, more reliable observability platform, retiring tooling that does not provide meaningful operational value.
  • Implement symptom-based, SLO-driven alerting with multi-window burn-rate principles and clear page, ticket, and dashboard classifications.
  • Ensure every production alert has a defined owner, documented failure mode, and actionable runbook.
  • Establish ongoing alert-quality practices, including periodic reviews, measurable actionable-alert rates, and deliberate removal of unnecessary alerts.
  • Build a machine-readable ownership and escalation model that routes incidents to the appropriate engineering squads.
  • Establish severity definitions, acknowledgement expectations, follow-the-sun escalation practices, and clean handoff procedures across multiple time zones.
  • Strengthen incident command and blameless postmortem practices, including reliable timelines, ownership, and follow-through on corrective actions.
  • Coach engineering squads to own their own operational responsibilities and paging rather than becoming a centralized buffer for other teams' alerts.
  • Deliver measurable reliability outcomes over the first year, including complete SLI ownership, production telemetry, tiered alerting, squad on-call adoption, and a significant reduction in the time required to detect silent security-control degradation.
  • Requirements

    • Substantial production engineering or SRE experience, including experience defining and implementing an SLO framework rather than simply operating within an existing one.
    • Strong Python skills and the ability to read and modify Go or Rust code when implementing instrumentation and reliability improvements.
    • Strong hands-on experience with time-series and event telemetry at scale, including Prometheus/OpenMetrics, Grafana, Alertmanager-class routing systems, and columnar or high-cardinality data stores such as ClickHouse or equivalent technologies.
    • Experience debugging distributed systems running on bare metal and long-lived hosts; this role requires more than Kubernetes-centric operational experience.
    • Practical experience with production-scale configuration management and CI/CD tooling such as Ansible, GitLab CI, Jenkins, or comparable technologies.
    • Strong understanding of telemetry for systems that cannot be directly scraped or fully controlled, including push-based collection, sampling, clock skew, partial reporting, and privacy considerations on customer-managed infrastructure.
    • Excellent written and asynchronous communication skills, with the ability to align multiple engineering teams around measurable definitions of system health.
    • Strong engineering judgment and a pragmatic approach to observability, reliability, alerting, and operational ownership.
    • Experience with security products such as WAF, EDR, antivirus, or vulnerability-management platforms is a strong advantage, particularly an understanding that security-control reliability must measure effective enforcement rather than simple uptime.
    • Familiarity with monitoring and continuous-monitoring requirements related to SOC 2, ISO 27001, NIST SP 800-137, or similar frameworks is valuable.
    • Exposure to OpenTelemetry, eBPF, Sentry, cost-aware telemetry, or cardinality-management techniques is beneficial.
    • Experience working effectively with AI-assisted development tools and modern agentic engineering workflows is a plus.
    • Kubernetes experience is useful for supporting the smaller portion of the platform that runs in Kubernetes.
    • This role is focused on SRE and reliability engineering rather than DevOps ticket management, build-system ownership, cloud cost management, or acting as the on-call team for other engineering squads.
    • Benefits

      • Fully remote work with flexible working hours, enabling you to work from anywhere worldwide.
      • 24 paid vacation days per year.
      • 10 paid national holidays.
      • Unlimited sick leave.
      • Compensation toward private medical insurance.
      • Co-working space reimbursement.
      • Gym and sports reimbursement.
      • Professional development opportunities through challenging technical projects, learning opportunities, mentoring, and knowledge-sharing programs.
      • Opportunity to receive a reward for an innovative idea that can be patented.
      • Remote-first, asynchronous working environment spanning multiple time zones.
      • Opportunity to define an SRE function, reliability standards, and operational culture from the ground up.
Vacancy posted 6 hours ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer - Imunify Reliability Platform in India vacancy
  •  ...countries. Founded in 1976, CGI is a leading IT and business process...  .... Develop secure, reliable, and high-performance RESTful...  ...management. Leverage cloud platforms and containerization technologies...  ...reviews to ensure adherence to engineering standards. Mentor... 
    Suggested
    Permanent employment
    Full time
    No agency
    Bangalore
    1 day ago
  •  ...GIB). Learn more at cgi.com. Job Title: Full Stack Developer / Delivery Lead Position: Full Stack Developer / Delivery Lead Experience: 8+ Years Category: Software Development / Engineering Shift: To be confirmed Main location: India, Telangana, Hyderabad /... 
    Suggested
    Full time
    Local area
    Shift work
    Andhra Pradesh
    1 day ago
  • DEVELOPER (Senior Software Engineer) - Data Analyst Position Description Company Profile: Founded in 1976, CGI is among the largest...  ...can inform business decisions. 7. Technical Leadership § Lead code reviews, ensuring adherence to coding standards, best practices... 
    Suggested
    Full time
    Local area
    Mumbai
    1 day ago
  •  ...scalable RESTful APIs and microservices for enterprise applications. . Deploy, configure, and support APIs on Red Hat OpenShift Container Platform (OCP). . Strong experience with Oracle and SQL Server database development and integration. . Proficient in SQL, PL/SQL, and T... 
    Suggested
    Full time
    Andhra Pradesh
    1 day ago
  •  ...Developer Experience: 5+ Category: Software Development/ Engineering Shift: Timing/rotation etc. details Main location: Bangalore...  ...patterns. Ensure solutions are scalable, secure, reliable, and production-ready Must-Have Skills: • Python and/or C#... 
    Suggested
    Full time
    Local area
    Shift work
    Bangalore
    1 day ago
  • SAP PLM Consultant Position Description Job Title: SAP PLM Consultant Position: Lead Analyst Experience: 10-15yrs Category: Software Development/ Engineering Shift: Timing/rotation etc. details Main location: India, Karnataka, Bangalore Position ID:... 
    Full time
    Shift work
    Bangalore
    1 day ago
  •  ...NYSE (GIB). Learn more at cgi.com. Job Title: Software Engineering Position: Lead Consultant Experience: 7 to 10 years Category: Software...  ...(Kubernetes/Kustomize/FluxCD/Helm) and contributing to platform decisions Owning and evolving the GraphQL API server (Go... 
    Permanent employment
    Full time
    Hybrid work
    Local area
    Flexible hours
    Shift work
    Bangalore
    1 day ago
  • Azure Data Engineer Position Description About CGI Founded in 1976, CGI is among the world's largest independent IT and business consulting...  ...more at cgi.com. Job Title: Azure Data Engineer Position: Lead Data Engineer Experience: 7+ years Category: Emerging... 
    Permanent employment
    Full time
    Hybrid work
    Local area
    Shift work
    Bangalore
    1 day ago
  •  ...We are seeking a talented and motivated best-in-class Senior Site Reliability Engineer. This role presents an exciting opportunity to thrive in a...  ...patterns. Help drive the infrastructure team’s roadmap, leading us to higher levels of reliability, recoverability, and scalability... 
    Remote job
    Full time
    Flexible hours
    India
    3 hours ago
  •  ...boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software...  ...professionals from software engineering, DevOps, SRE, cloud, platform, or infrastructure engineering backgrounds to test and evaluate... 
    Remote job
    Hourly pay
    Contract work
    For contractors
    India
    3 hours ago
  • Tesco India •  Bengaluru, Karnataka, India •  Full-Time •  Working hours 45•  Apply by 02-Sep-2026
    Full time
    Bangalore
    4 hours ago
  •  ...ArchitectLocation : Pune (preferred)Experience : 11-18 YearsMandatory Skills : Databricks Administration, Unity Catalog, TerraformCore Platform Skills :- Databricks workspace and account administration- Cluster management, policies, and SQL warehouses- Strong Unity Catalog... 

    Wipro Technologies

    Bangalore
    4 hours ago
  • Skill: SAP DatasphereLocation: Bengaluru (1-2 Customer rounds)Experience: 4-6 Years (Relevant minimum 2 years in SAP)Job Description :We are looking for a skilled SAP Datasphere Developer with hands-on experience in data modeling, transformation and ETL within SAP Datasphere...
    Side job
    Remote job

    Wipro Technologies

    Bangalore
    4 hours ago
  • Data Engineer GCPLocation : PuneExperience : 6+ YearsJoining : Immediate to 15 DaysJob Summary...  ...data pipelines for performance, reliability, and scalability.- Collaborate with cross...  ...contribute to continuous improvement of data platforms.Mandatory Skills :- 6+ years of... 
    Immediate start

    FUNIC TECH PRIVATE LIMITED

    Pune
    4 hours ago
  • Experienced C/C++ Linux Engineer with strong expertise in OpenWrt, networking, and 5G technologies. The candidate will be responsible for...  ...software for embedded Linux and next-generation networking/telecom platforms.Key Responsibilities :- Design, develop, and maintain software... 

    Akshaya Business IT Solutions

    Chennai
    4 hours ago
  •  ...dog-friendly offices. We hope to meet you soon!About the role : Heady is looking for an experienced, well-rounded Software Development Engineer In Test (SDET) to join our growing team. In this position, you will own responsibility for the quality of our work and work closely... 
    Flexible hours

    Heady Technologies Consultancy Pvt. Ltd.

    Mumbai
    4 hours ago
  • Role Overview :As a Senior Data Engineer, you will serve as a cornerstone of our data infrastructure...  ...the efficiency of our analytics platforms and drive data-backed outcomes that...  ...effectively while maintaining high standards of code quality and system reliability. (ref:hirist.tech)
    Remote job

    EiCTechsys

    India
    4 hours ago
  •  ...experienced Snowflake Architect / Senior Data Engineer to design, develop, and deliver scalable...  ..., along with good exposure to cloud platforms and data architecture.Key Responsibilities...  ...Snowflake Architect, Snowflake Data Engineer, Lead Data Engineer, Cloud Data Architect, or... 
    Work at office

    Durus Consulting Private Limited

    Chennai
    4 hours ago
  •  ...transformation, data analytics, and cloud engineering. The company partners with global...  ...data ecosystems, leveraging advanced cloud platforms to drive actionable business intelligence...  ...delivery, ensuring that our clients have reliable, real-time access to critical business insights... 
    Remote job

    KPI PARTNERS INDIA PRIVATE LIMITED

    Pune
    4 hours ago
  • Job Description :We are seeking an experienced SAP Security & Access Governance Consultant to join our Global IT team.In this role, you will be responsible for SAP user administration, role design, access governance, and security compliance across the SAP landscape.As a subject...

    Gyansys Infotech Pvt ltd.

    Bangalore
    4 hours ago
  •  ...experienced Agentic AI Manager to lead the development, deployment,...  ...AI models, ensuring their reliability, safety, and compliance with...  ...Science, Artificial Intelligence, Engineering, or a related field. PhD is a...  ...:- Experience with cloud platforms (AWS, GCP, Azure) and MLOps... 

    BLJ TECH GEEKS

    Delhi
    4 hours ago
  •  ...databases.- Develop LLM-powered applications leveraging prompt engineering, AI agents, and multi-agent workflows.- Fine-tune, evaluate,...  ...Engineering, DevOps, and Product teams.- Ensure scalability, reliability, and performance of AI applications in production environments... 

    Infinium Associates

    Gurgaon
    4 hours ago
  •  ...application design, development, testing, deployment, and production support.- Collaborate with product managers, architects, QA, and other engineering teams.- Write clean, maintainable, and well-tested code.- Troubleshoot performance, integration, and production issues.-... 

    Talent Toppers

    Noida
    4 hours ago
  •  ...Immediate to 45 Days only.About the Position :The Principal Data Platform Engineer is a member of the Data & Analytics (D&A) Platform and...  ...platform, collaborating closely with HQ architects and engineering leads.- Lead the design, development, and maintenance of scalable data... 
    Immediate start

    Wipro Technologies

    Pune
    4 hours ago
  • Role : Cloud Operations and Engineering role for a USA Based companyLocation : RemoteExperience...  ...: 6+ years in infrastructure, DevOps, platform engineering, cloud operations, or SRE; ~...  ...alerts, incident response, performance, reliability, and cost optimisation.- Scripting... 
    Remote job

    Hunastreet Technology Pvt Ltd

    work from home
    4 hours ago
  •  ...:- 5+ years of experience in Salesforce development.- Salesforce Platform Developer II certification.- Strong proficiency in Apex, Lightning...  ...integrity and real-time synchronization across the ecosystem.- Lead the end-to-end development lifecycle, utilizing Salesforce DX to... 
    Long term contract

    THOMPSONS HR CONSULTING PRIVATE LIMITED

    Hyderabad
    4 hours ago
  •  ...C++Job Responsibilities : - Collaborate with Managers and other Engineers to help define, scope, and implement high-quality features that...  ...Bangalore-based product company providing a research and deal sourcing platform for Venture Capital, Private Equity, Corp Dev, and professionals... 

    Tracxn Technologies

    Bangalore
    4 hours ago
  •  ...Python, along with an understanding of modern cloud-based data platforms. While this is not primarily a Power BI development role, familiarity...  ...- Understanding of cloud-based data architectures4. Data Engineering & Analytics : - SQL (Advanced)- Python- Data extraction, cleansing... 
    Work at office

    CodeChavo

    Pune
    4 hours ago
  • SAP ABAP Consultant L2Location : PuneExperience : 6 - 8 YearsLevel : L2Work Mode : As per business requirementJob Summary :We are looking for an experienced SAP ABAP Consultant L2 with 6 - 8 years of hands-on experience in SAP ABAP and S/4HANA development. The ideal candidate...

    Tech Grow Global

    Pune
    4 hours ago
  •  ...the global rollout, adoption and continuous improvement of its CRM platform based on SAP Sales Cloud Version 2.In this role, you will work...  ...a strong understanding of B2B Sales and CRM processes, including lead management, opportunity management, customer engagement and customer... 

    Abeyaantrix solutions

    Bangalore
    4 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer - Imunify Reliability Platform. Be the first to apply!