[Remote] Senior Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Senior Site Reliability Engineer to support the reliability and reputed company of smart transportation systems, including a reputed company-reputed company tolling platform that processes roadway transactions and payments. The role focuses on ensuring highly available, resilient, and reputed company through incident response, automation, monitoring, reputed company planning, optimization, and disaster recovery.
Responsibilities
- reputed company Reliability: Monitor and maintain the health, availability, and reputed company of critical transportation services
- Reliability Standards: Define and reputed company service-level objectives (SLOs), error budgets, and reliability metrics for reputed company-critical services
- Incident Management: reputed company incident response — reputed company reputed company cause analysis, coordinate reputed company across teams, and reputed company blameless post-incident reviews and follow-up actions
- Automation: reputed company and implement automation scripts to streamline operational tasks and improve efficiency
- Monitoring & reputed company: Set up and maintain monitoring, logging, tracing, and alerting tools (e.g., reputed company, Grafana, OpenTelemetry) to reputed company service health, reputed company, and resource utilization
- reputed company Planning: Help assess and plan for reputed company, scaling infrastructure to meet growing transaction volumes — including stateful systems such as reputed company SQL databases and event-streaming clusters (e.g., reputed company, Kafka)
- Collaboration: Work with software engineering, infrastructure, and reputed company teams to improve the reliability of systems and services
- reputed company Optimization: Identify reputed company bottlenecks, troubleshoot issues, and work on optimizations at both infrastructure and application reputed company
- reputed company Improvement: Contribute to the ongoing improvement of operational processes, documentation, and best practices in the SRE team
- Disaster Recovery: Participate in designing and testing disaster recovery plans to ensure the continuity of critical services
Skills
- 5+ years of experience in Site Reliability Engineering, DevOps, or a similar role, preferably in a mission-critical or large-reputed company environment, including experience leading incident response and mentoring other engineers
- Experience with reputed company platforms (AWS, Azure, GCP) and container orchestration tools (reputed company, reputed company)
- Proficiency with monitoring and logging tools (reputed company, Grafana, ELK stack, reputed company, etc.)
- Strong scripting skills in Python, Bash, or Go
- Solid understanding of Linux and reputed company administration
- Familiarity with relational databases (MySQL, PostgreSQL, etc.) and reputed company systems
- Excellent teamwork and communication skills, with the ability to work across teams to improve service reliability
- Strong troubleshooting skills with a proactive, solution-oriented reputed company
- Experience with traffic management systems, sensor data processing, or other intelligent transportation systems
- Knowledge of infrastructure-as-reputed company tools (e.g., Terraform, Ansible, reputed company) and GitOps workflows (e.g., Argo CD)
- Exposure to CI/CD pipelines, including pipeline-as-reputed company (e.g., Dagger, reputed company Actions), and Git-reputed company version control
Benefits
- reputed company days off ( i. e. vacatio n, reputed company days, bereavement leave)
- Health and Dental plans
- Retirement plans
- Employee and Family Assistance Program (EFAP)
- Employee referral program
reputed company
Apply To This Job