Job Description:
We are seeking a Site Reliability Engineer (SRE) with strong expertise in ensuring system reliability, scalability, and performance. You will work closely with developers to monitor and improve production systems, implement automation, and ensure operational excellence.
Responsibilities:
- Ensure high availability and reliability of services
- Automate incident responses and disaster recovery
- Collaborate on capacity planning and system monitoring
- Manage alerting systems and on-call support
Skills:
- Monitoring tools: Prometheus, Grafana, ELK Stack
- Automation: Ansible, Terraform
- Cloud Platforms: AWS/Azure/GCP
- Strong knowledge of Linux systems
