Job Purpose
We are seeking an experienced Site Reliability Engineer (SRE) to ensure the reliability, availability, performance, and scalability of critical applications and infrastructure. The ideal candidate will have strong experience in monitoring, automation, cloud technologies, incident management, and DevOps practices.
Key Responsibilities
- Monitor and maintain the availability, performance, and reliability of applications and infrastructure.
- Design and implement effective monitoring, logging, and alerting solutions.
- Manage and troubleshoot production incidents and perform root cause analysis.
- Automate operational and deployment processes to improve system reliability and efficiency.
- Work closely with development, infrastructure, and DevOps teams to improve application performance and resilience.
- Implement and maintain CI/CD pipelines and Infrastructure as Code (IaC).
- Support containerized environments and cloud-based infrastructure.
- Develop scripts and automation tools to reduce manual operational activities.
- Implement observability solutions, including monitoring, logging, and distributed tracing.
- Ensure proper documentation of operational procedures, incidents, and system configurations.
Required Technical Skills
- Monitoring: Prometheus, Grafana, Zabbix
- Logging: ELK/Elastic Stack, Splunk
- APM: Dynatrace, AppDynamics, New Relic
- Cloud: AWS, Microsoft Azure, or GCP
- Containers: Docker, Kubernetes, OpenShift
- CI/CD: Jenkins, GitLab CI/CD, Azure DevOps
- Infrastructure as Code: Terraform, Ansible
- Version Control: Git, GitHub, GitLab
- Incident Management: ServiceNow, PagerDuty
- Distributed Tracing: OpenTelemetry, Jaeger
- Scripting: Bash, Python, PowerShell
- Databases: PostgreSQL, Oracle, SQL Server
- Web & APIs: IIS, Nginx, Apache, REST APIs
Qualifications & Experience
- Bachelor's degree in Computer Science, Information Technology, or a related field.
- Proven experience as a Site Reliability Engineer, DevOps Engineer, or Production Support Engineer.
- Strong experience in cloud infrastructure, automation, monitoring, and incident management.
- Hands-on experience with Kubernetes and containerized environments.
- Strong troubleshooting and root cause analysis skills.
- Experience working in highly available, large-scale production environments.
- Excellent communication and collaboration skills.
#J-18808-Ljbffr
This listing is sourced from a third-party job board. Applying will redirect you to the original posting.