Senior Site Reliability Engineer (SRE)
SpotMe
SeniorAbout the position
Join a leading B2B event platform to ensure the reliability and scalability of a 24/7 SaaS platform. You'll optimize infrastructure and automate processes while collaborating with engineering and product teams.
Tech stack
- aws
- terraform
- python
- docker
- jenkins
- datadog
Requirements
Required:
- Around five or more years in a site reliability role, built on earlier experience as a system administrator or software developer
- Strong, production-grade experience with AWS, and knowledge of Azure is a bonus
- Automate and manage infrastructure with Terraform, and build and maintain CI/CD pipelines with Jenkins
- You diagnose and resolve complex system issues under pressure, and design for resilience before incidents happen
Nice to have:
- Experience with JavaScript, Node.js, or Go is an asset
- Hands-on with cloud-native architectures, distributed systems, and high-availability platforms
- Comfortable across both document-oriented and relational databases
Responsibilities
- Develop and deploy scalable infrastructure using Terraform and cloud-native AWS services
- Optimize the platform's cloud infrastructure for high availability and cost efficiency
- Take your turn in the on-call infrastructure rotation, responding to incidents and resolving them quickly
- Strengthen the platform’s monitoring and observability to catch issues before they reach end users