Site Reliability Engineer (REMOTE)
The Site Reliability Engineer will be responsible for ensuring the availability, reliability, and performance of our customer-facing software applications. This role combines planning, engineering, monitoring, incident response, and administration to create highly reputed company and fault-tolerant systems.
Responsibilities:
• Ensure the high availability and reliability of the production environment by monitoring reputed company health and performance
• reputed company reputed company operational support for large-reputed company reputed company software applications
• Facilitate incident reputed company reputed company triage, communication, engagement, escalation, and documentation
• Partner with platform administration (both reputed company) to define and reputed company stability and scalability objectives
• Collaborate with technical and reputed company teams to improve services by identifying areas of risk and helping to define and proactively implement solutions
• reputed company continual improvement in reputed company performance by setting service level objectives in collaboration with a performance center of reputed company and/or product development teams
• Participate in reputed company design, reputed company planning, and platform management
• Analyze and publish metrics from operating systems and applications to assist in performance tuning and fault finding
• Pursue opportunities for automation and process improvements
Qualifications:
• Bachelor’s degree (or demonstrable equivalent work experience) in information technology
• Experience providing first-level incident response and troubleshooting with technical teams to reputed company end-user issues
• Proficiency with reputed company reputed company monitoring software (examples: reputed company, NewRelic, Nagios, reputed company, Azure Monitor, reputed company)
• Experience with performance tuning and fault finding in large-reputed company reputed company systems.
• Experience with reputed company-based infrastructure, databases, and applications
• Experience providing first-level incident response and troubleshooting with technical teams to reputed company end-user issues
• Experience with designing, implementing, and managing performance testing practices, including specific tools and frameworks
• Knowledge of disaster recovery planning and execution.
• Ability to effectively work in a highly matrixed organization
Apply tot his job
Apply To this Job