We’re looking for a DevOps Engineer / Site Reliability Engineer (SRE) to build, automate, operate, and improve highly available and scalable production systems.
The role involves managing infrastructure, CI/CD pipelines, containers, monitoring, observability, databases, messaging platforms, and distributed systems. The engineer will also troubleshoot production issues, lead root-cause analysis, support incident response, improve reliability and fault tolerance, and participate in an on-call rotation.
Candidates should have strong experience with Linux, Docker, Kubernetes, Git, Python, Bash, Infrastructure as Code, CI/CD, networking, databases, Kafka or RabbitMQ, and tools such as Prometheus, Grafana, and the Elastic Stack.
Experience in large-scale or data-intensive environments, strong communication skills, and familiarity with DevOps/SRE best practices are preferred.
Why Join Us?
This is an opportunity to work on challenging infrastructure and reliability problems while helping build robust systems that operate at scale. If you’re passionate about automation, distributed systems, and improving the reliability of production systems, we’d love to hear from you.