Site Reliability Engineer - Cloud Infrastructure
Sysmind
Job Summary : We are seeking a proactive and experienced Site Reliability Engineer (SRE) with 46 years of hands-on experience in managing cloud-native infrastructure, ensuring high availability, and improving the reliability of enterprise applications. The ideal candidate should possess strong expertise in Kubernetes, Python, Linux, and modern observability tools such as Dynatrace and Prometheus. You will work closely with development, DevOps, and cloud teams to automate operations, optimize system performance, and build resilient production environments.Key Responsibilities : - Design, deploy, and manage highly available applications on Kubernetes clusters.- Develop automation scripts and operational tools using Python.- Administer, monitor, and troubleshoot Linux/Unix production servers.- Implement and maintain monitoring, alerting, and observability solutions using Dynatrace, Prometheus, and related tools.- Analyze production incidents, perform root cause analysis (RCA), and implement preventive measures.- Monitor infrastructure performance, application health, and system capacity to ensure maximum uptime.- Collaborate with DevOps and development teams to improve deployment strategies and operational efficiency.- Support CI/CD pipelines and automate operational workflows wherever possible.- Maintain system documentation, operational runbooks, and incident response procedures.- Participate in on-call rotations and provide production support for critical enterprise applications.- Drive continuous improvements in system reliability, scalability, security, and operational excellence.Required Skills : - 46 years of experience in Site Reliability Engineering (SRE), DevOps, or Platform Engineering.- Strong hands-on expertise in Kubernetes administration and container orchestration.- Proficiency in Python scripting for automation and operational tasks.- Strong experience with Linux/Unix system administration and troubleshooting.- Hands-on experience with monitoring and observability tools such as Dynatrace and Prometheus.- Experience supporting production environments with high availability requirements.- Knowledge of incident management, root cause analysis, and system performance tuning.- Familiarity with CI/CD pipelines and DevOps methodologies.- Strong analytical, troubleshooting, and communication skills.Mandatory Skills : - Kubernetes- Site Reliability Engineering (SRE)- Python- Linux / Unix Administration- Dynatrace- Prometheus- Monitoring & Observability- Production Support- Incident Management- Automation & Scripting (ref:hirist.tech)
- ...advantage.Your Future Employer :A leading global AI and Analytics organisation helping businesses leverage data, advanced analytics, data engineering and AI-powered solutions to solve complex business challenges.Responsibilities :- Engage with business stakeholders to understand,...Suggested
- ...specializing in building scalable digital infrastructure and high-performance web... ...Operating at the intersection of cloud computing and modern software engineering, the company provides robust solutions... ...of code quality and system reliability.- Troubleshoot and resolve production...SuggestedHybrid work
- Principal Data Engineer - 9+ Years - Bengaluru/Hyderabad/GurugramWe are looking for an experienced Principal - Data Engineer to lead the design... ...organization helping global enterprises leverage data, cloud, and AI to drive business transformation.Responsibilities : - Partner...Suggested
- Job Description : PayPay India is looking for a QA Automation Engineer to enhance our payment systems and deliver the best customer experience.Main Responsibilities : - Utilizing the Testing Methodology, analyzes testing requirements as the basis for developing testing scenarios...Suggested
- Description :Job Title : Voice AI Engineer (Machine Learning & NLP)Locations : Gurugram (Sector 54) / Chennai (T. Nagar) Hybrid (4 Days In-Office) Experience : 3+ Years Role Summary :We are an agentic AI platform for financial services, backed by Y Combinator and other world...SuggestedHybrid workWork at office
- Daily Responsibilities : - Manage L2 incident tickets and service requests within defined SLAs.- Monitor M365 health dashboards and resolve service availability alerts daily.- Administer Microsoft Defender policies and investigate security alerts for endpoints.- Manage user ...Work at office
- Job Description :We are looking for a Data Run Support Engineer with strong experience in AWS Data Engineering, Snowflake, Python/PySpark... ...data quality, monitoring platforms, and improving operational reliability.The ideal candidate should be proactive, have strong...
- ...future regressions.- Maintain version control best practices using Git to ensure seamless collaboration and code integrity across the engineering team.Preferred Skills : - JMeter / Locust- Experience with distributed systems, microservices, IoT, robotics, or supply-chain...
- ...C#, .NET technologies, object-oriented programming, REST APIs, and database development.About Us : CodeVyasa is a mid-sized product engineering company that works with top-tier product and solutions organizations such as McKinsey, Walmart, RazorPay, Swiggy, and others. We...Long term contractFlexible hours
- ...Description : We are seeking a skilled NOC engineer DevOps in Gurgaon having 4 - 8 yrs of... ...solutions.Key Responsibilities : - Monitor infrastructure, applications, services, dashboards, and... ...to DevOps, application, infrastructure, cloud, or network teams.- Track incidents...Flexible hoursShift work
- Role : Software Engineer - Full Stack" in a bank.Responsibilities : - Develop and maintain web and mobile applications for the company.- Work on both the front end and back end.- Convert business/technical requirements into working software.- Write clean, reusable and secure...
- Consultant Data Engineer Databricks | Azure | PySpark | Spark SQL Job Description :Looking... ...architecture while ensuring data quality and reliability.- Monitor production data pipelines,... ...you :- Opportunity to work on large-scale cloud data engineering projects.- Exposure to modern...Worldwide
- ...Responsibilities Define the data validation, automation, and quality roadmap with Data Engineering and Product leadership. Design and execute complex tests for backend data systems, transformations, distributed-system logic, and batch or stream processing. Develop...Full timeWork at officeRemote jobFlexible hours
- ...Elastic/OTEL, and Site24x7. Steer cloud migration and improve the reliability and performance of services.... ...and monitoring standards. Mentor engineers and help mature operational practices... ...or similar languages. ~ Strong infrastructure-as-code experience with Terraform,...Full timeWork at officeRemote job
- ...performance. Implement technical SEO best practices, including site structure, metadata, sitemaps, structured data, and other on-... ...strong focus on reducing recurring issues and improving website reliability. Security & Maintenance Maintain website security and ensure...Temporary workImmediate start
- Inside Sales Specialist-Cloud Start Date Starts Immediately... ...Understand customer cloud strategy and map ZeaCloud solutions (Infrastructure-as-a-Service {IaaS}, Platform-as-a-Service {PaaS}, Backup-as-a-...Hybrid workImmediate startRemote job
- E-learning Developer Start Date Starts Immediately CTC (ANNUAL) Competitive salary Competitive salary Apply By ...Contract workImmediate start
Rs 10000 - Rs 12000 per month
.... Mastery of AI developer tools: Deep daily experience with tools like Cursor, Claude, GitHub Copilot, v0, Bolt.new, or similar AI engineering environments. 3. Production debugging: Proven ability to trace errors in an existing legacy or live codebase rather...Immediate start- ...Role: Application Reliability Engineer Location: Gurgaon Who we are Graviton Research Capital is a privately funded quantitative... ...operation and minimal downtime. Monitor trading systems and infrastructure., Triage issues across trading support services,...Long term contract
- ...skilled Senior Python Data Scraping Engineers to join the Tendem project and... ...to deliver accurate, reliable, and high-quality results. This... ...rendered content and changing site behavior. Enforce data quality... ...at scale Experience with cloud infrastructure (AWS or equivalent...Hourly payPart timeFreelanceRemote job
- ...system integration. - Collaborate with engineering and business stakeholders to deliver... ...operations. - Improve the performance, reliability, and maintainability of backend and data... ...NICE TO HAVES - Familiarity with cloud platforms, CI/CD, and data pipeline operations...Full timeLocal areaRemote jobFlexible hours
- ...THE ROLE We are looking for an Agentic AI Engineer to build generative AI workflows that accelerate cloud configuration and remediation activities. This person... ...and remediation activities - Policy & Infrastructure as Code (IaC): generate and validate compliant Terraform...Full timeLocal areaRemote jobFlexible hours
- ...ABOUT THE ROLE We are looking for a Lead AI Engineer to architect generative AI workflows that accelerate cloud configuration and remediation. Using RAG,... ...response and remediation activities - Policy & Infrastructure as Code (IaC): accelerate the generation and...Full timeLocal areaRemote jobFlexible hours
- About the Role:We are looking for a skilled and proactive Site Reliability Engineer who will be responsible for ensuring the stability, availability... ...applications.The role will involve production support, cloud infrastructure management, monitoring and observability, risk...Work at office3 days week
- ...achieved with a high level of quality. You are a high-performance engineer expected to work in a product squad and deliver solutions for... ...webservices and consuming webservices. - Hands-on experience with spring Cloud/Spring Boot.- Hands-on experience with any of the logging...
- Role Overview :As a Full Stack Software Engineer, you will take ownership of the end-to-... ...platforms, ensuring that our technical infrastructure remains resilient as we expand our... ...high standards of code quality and system reliability.- Troubleshoot and resolve complex technical...Remote jobFlexible hours
- ...of concurrency and distributed computing.- Degree in Computer Engineering or Computer Science or 5+ years equivalent experience in SaaS... ...our high throughput applications.- Understand how to leverage infrastructure for solving such large scale problems.- Develop tools and contribute...
- ...Server administration, licensing, and deployment models including cloud, on-premise, and hybrid environments.- Must have hands-on... ...Bachelor's degree in Computer Science, Information Systems, Statistics, Engineering, or related field (Master's preferred). (ref:hirist.tech)Hybrid work
- Description :AuxoAI is hiring a Senior Applied AI Engineer to design and deploy production-grade AI agents capable of structured reasoning... ...candidate will develop robust agent architectures that operate reliably in real-world environments with constraints around latency, cost...
- ...are looking for a Principal AI Engineer to build AI-native products... ..., model governance, and data infrastructure roadmap.Mandatory Requirements... ...on scalability, latency, reliability, and high concurrency.- Strong... ...AI deployment platforms, and cloud-native architectures.- Experience...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer - Cloud Infrastructure. Be the first to apply!
