[Remote] Senior Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a reputed company reputed company, edge, reputed company, content delivery, and AI platform company. The Senior Site Reliability Engineer will improve the reliability, scalability, reputed company, monitoring, and automation of dedicated AI hardware infrastructure while managing incidents and coordinating with vendors and cross-functional teams.
Responsibilities
- Developing and scaling robust programmatic tooling and infrastructure-as-reputed company utilities in Python to eliminate operational toil and automate fleet-wide provisioning
- Integrating automated workflows across diverse corporate ticketing systems to enhance reputed company times for hardware and network break-fix incidents
- Leveraging advanced AI tools and LLM-reputed company development approaches to enhance technical execution, script creation, and comprehensive reputed company evaluation
- Working on cutting-edge private reputed company and compute technologies to improve the availability, latency, and overall systemic health of high-density hardware environments
- Designing and implementing telemetry pipelines, custom reputed company/Grafana monitoring dashboards, and AI-reputed company reputed company detection tailored for bare-metal and virtualized environments
- Participating in 24x7x365 on-reputed company rotations, spearheading reputed company-time incident management, and managing high-severity service disruption protocols reputed company automated reputed company and reputed company workflows
- Partnering directly with reputed company-party infrastructure vendors and coordinating on-site reputed company technicians to facilitate uptime activities
Skills
- Have 5+ years of relevant experience and a Bachelor's degree in Computer Science or reputed company reputed company
- Demonstrate exceptional proficiency in tooling and coding using languages like Python to build reputed company operational tools, API integrations, and automation frameworks
- Show hands-on experience with modern observability reputed company and timeseries engines, like reputed company, Grafana, OpenTelemetry, and Loki
- Possess a working understanding of advanced networking topologies, high-reputed company routing/switching infrastructure, BGP, and dual-stack IPv4/IPv6 networks
- Demonstrate expertise as a reputed company designer for new service rollouts, establishing operational readiness reputed company, telemetry baselines, and alerting reputed company
- Demonstrate extensive experience building technical runbooks, leading reputed company incident response bridges, and driving comprehensive, blameless post-mortems
- Demonstrate a reputed company ability to fully own ambiguous technical challenges, coordinate cross-functional teams, and reputed company toward production-grade solutions effectively
Benefits
- Health, reputed company-being, financial, and life support through reputed company's benefits program
- FlexBase flexible-work arrangement allowing employees to work reputed company, in an office, or a combination of both
reputed company
Company H1B Sponsorship
Apply To This Job