Skip to main content
Sails Software Solutions logo

Senior Site Reliability Engineer

Sails Software Solutions
Full Timesenior
Visakhapatnam, Andhra Pradesh, INPosted April 15, 2026

Resume Keywords to Include

Make sure these keywords appear in your resume to improve ATS scoring

PythonGoBashAWSDockerKubernetesTerraformLinuxCI/CDDevOpsMicroservices

Sign up free to auto-tailor your resume with all these keywords and get a higher ATS score

Job Description

Experience: 7+ Years

About the Role

We are looking for a highly experienced Senior SRE with strong expertise in AWS to help design, operate, and scale the infrastructure powering our product platforms. This is a mission-critical role in a fast-moving product development environment, where system reliability, automation, and performance are core business drivers.

Key Responsibilities

Reliability & Operations

Own reliability, availability, and performance of large-scale production systems.

Establish SLOs, SLAs, and error budgets for mission-critical services.

Lead incident response, root cause analysis, and continuous improvement initiatives.

Design fault-tolerant architectures and disaster recovery strategies.

Cloud & Infrastructure Engineering

Architect, deploy, and manage infrastructure on AWS using IaC (Terraform / CloudFormation).

Optimize cloud costs while maintaining performance and reliability.

Implement multi-region, highly available architectures.

Manage container platforms (Docker, Kubernetes, EKS).

Automation & DevOps

Build automation pipelines for infrastructure provisioning, deployment, and scaling.

Improve CI/CD pipelines and release engineering processes.

Develop tools and scripts to reduce operational toil.

Observability & Performance

Implement comprehensive monitoring, logging, and alerting systems.

Drive performance tuning and capacity planning.

Lead chaos engineering and resilience testing practices.

Leadership & Mentorship

Mentor SREs and DevOps engineers.

Partner with Engineering and Product teams to embed reliability into product design.

Skills

Required Skills & Experience

7+ years in Site Reliability Engineering / DevOps / Infrastructure roles.

Deep hands-on experience with AWS services (EC2, EKS, RDS, S3, Lambda, VPC, IAM, etc.).

Expertise in infrastructure as code: Terraform, CloudFormation.

Strong experience with Linux systems, networking, and distributed systems.

Experience with Kubernetes, container orchestration, and microservices environments.

Strong scripting skills (Python, Bash, Go).

Knowledge of security best practices and compliance requirements.

Soft Skills

Strong problem-solving and decision-making ability under pressure.

Excellent communication and stakeholder collaboration.

High ownership and accountability mindset.

Ability to thrive in an aggressively-paced product development culture.

Education

Bachelors degree in Computer Science, Engineering, or related field (preferred).

Want AI-powered job matching?

Upload your resume and get every job scored, your resume tailored, and hiring manager emails found - automatically.

Get Started Free