Role Summary
We are seeking an experienced DevOps Engineer (AWS) to lead observability and monitoring initiatives for a cloud-native AWS platform. In this role, you will design and implement enterprise-grade monitoring solutions, define SRE standards, establish SLIs/SLOs, and improve platform reliability through proactive monitoring, alerting, and incident management. You will work closely with engineering and operations teams to ensure high availability, performance, and operational excellence across AWS environments.
Key Responsibilities
- Design and implement observability solutions for AWS-based platforms, including EKS, Lambda, and SNS.
- Define and implement SLIs, SLOs, monitoring standards, and alerting strategies.
- Build, configure, and maintain centralized monitoring solutions using AWS CloudWatch.
- Develop dashboards and alerts for engineering, operations, and business stakeholders.
- Lead incident monitoring, troubleshooting, and platform reliability improvement initiatives.
- Implement Infrastructure as Code (IaC) for monitoring and observability deployments.
- Perform root cause analysis and continuously optimize platform performance.
- Collaborate with cross-functional teams to improve cloud operations, reliability, and automation.
- Promote DevOps and Site Reliability Engineering (SRE) best practices across the organization.
Required Skills & Experience
- Strong hands-on experience with AWS CloudWatch.
- Experience working with Amazon EKS, AWS Lambda, and Amazon SNS.
- Strong knowledge of Observability, Monitoring, and Logging practices.
- Experience implementing SRE principles, including SLIs, SLOs, and incident management.
- Experience creating dashboards, alerts, and monitoring strategies.
- Hands-on experience with Infrastructure as Code (Terraform or AWS CloudFormation).
- Strong understanding of DevOps practices and AWS Cloud Operations.
- Experience performing root cause analysis, performance monitoring, and system optimization.
- Excellent troubleshooting, communication, and stakeholder collaboration skills.
Desirable Skills
- Experience with enterprise monitoring tools such as Grafana, Datadog, Splunk, Dynatrace, or similar platforms.
- AWS and/or DevOps certifications.
- Experience working in Agile and cloud-native environments.
Pay: From $120,000.00 per year
Work Location: Hybrid remote in Sydney NSW