- Build reliable, scalable cloud platforms and automation solutions
- Implement observability, monitoring and incident response practices
- Drive CI/CD, Kubernetes and service reliability improvements
Company Description
Showtime Consulting is a leading provider of Shielded Cloud and Digital Solutions across Australia and New Zealand. We specialise in delivering secure, enterprise-scale technology solutions across cloud, data, DevSecOps, platform engineering, and digital transformation programs. We partner with major organisations to build high-performing technology teams and deliver innovative cloud-native solutions that improve reliability, scalability, security, and operational efficiency.
The Role
We are seeking an experienced Site Reliability Engineer (SRE) to join a high-performing engineering team supporting large-scale, mission-critical telecommunications platforms.
This is a hands-on engineering role where operational challenges are approached as software engineering problems. You will work across cloud platforms, microservices environments, CI/CD pipelines, observability tooling, and modern infrastructure platforms to build resilient, scalable, and highly automated systems.
You will play a key role in driving reliability, automation, monitoring, and incident response initiatives while championing a culture of Observability-as-Code and continuous improvement. The ideal candidate will bring strong SRE and DevOps experience, a passion for automation, and a proven ability to improve platform reliability through engineering and operational excellence.
Key Responsibilities
- Design, build, and maintain reliable, scalable, and highly available cloud-native platforms and applications.
- Implement and manage CI/CD pipelines to streamline software delivery and deployment processes.
- Automate operational tasks and eliminate manual effort through scripting and platform engineering practices.
- Develop and support self-healing systems to improve platform resilience and reduce operational risk.
- Implement monitoring, logging, alerting, and observability solutions across complex distributed systems.
- Build and maintain dashboards, metrics, and health-check frameworks to improve operational visibility.
- Lead incident response activities, root cause analysis, and post-incident reviews.
- Collaborate with software engineering, infrastructure, networking, and platform teams to improve system reliability and performance.
- Support containerised workloads and Kubernetes-based environments.
- Drive best practices for reliability engineering, operational excellence, and service ownership.
- Contribute to architectural discussions and platform strategy initiatives.
- Implement security, compliance, and governance controls across cloud and application environments.
- Support performance tuning, capacity planning, and continuous service improvement activities.
What You'll Need
- Strong experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, or similar roles.
- Hands-on experience with Java 11 and modern software engineering principles.
- Experience designing and supporting microservices-based architectures.
- Strong experience building and maintaining CI/CD pipelines using Jenkins, GitLab CI, or similar technologies.
- Experience with containerisation technologies including Docker and Kubernetes.
- Strong understanding of observability, monitoring, and logging platforms.
- Hands-on experience with Prometheus, Grafana, ELK Stack, or similar monitoring solutions.
- Experience supporting cloud environments across AWS, Azure, or Google Cloud Platform.
- Experience with scripting and automation using Python, Bash, or similar technologies.
- Strong understanding of incident management, root cause analysis, and service reliability practices.
- Excellent troubleshooting, analytical, and problem-solving skills.
- Strong communication skills and the ability to work effectively across technical and business teams.
Why Join Us?
- Work on large-scale, mission-critical telecommunications platforms supporting millions of users.
- Drive reliability, automation, observability, and cloud engineering initiatives.
- Work with modern cloud-native technologies, microservices, Kubernetes, and enterprise-scale platforms.
- Collaborate with experienced engineering, cloud, platform, and operations teams.
- Join a consulting culture focused on innovation, technical excellence, continuous learning, and engineering best practices.