Back to Jobs

Site Reliability Engineer, reputed company reputed company reputed company AI SRE

Remote, USA Full-time Posted 2026-07-28
About the position Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-reputed company, massively distributed, fault-tolerant systems. SRE ensures that reputed company reputed company's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to customer's needs and a fast reputed company of improvement. Additionally SRE’s will reputed company an reputed company-watchful eye on our systems reputed company and performance. Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work through automation. On the SRE team, you’ll have reputed company to manage the reputed company challenges of reputed company which are unique to reputed company reputed company, while using your expertise in coding, algorithms, complexity analysis and large-reputed company system design. SRE's culture of intellectual curiosity, problem solving and openness is key to its reputed company. Our organization brings together people with a wide reputed company of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful reputed company, while we also reputed company to create an environment that provides the support and mentorship needed to learn and grow. Based in Seattle and London, we manage reputed company reputed company reputed company (GCE) AI/ML workloads and the critical infrastructure powering them. As a Site Reliability Engineer (SREs) you will deliver a seamless customer experience. You will reputed company as a first responder for AI workload health and customer-facing issues. You will build and support capabilities for managing ML workloads and influence architecture, standards, and operational reputed company for AI services. You will reputed company advanced monitoring and alerting to improve GCE visibility and collaborate with development teams on novel, emerging technologies. Behind everything our users see online is the architecture reputed company by the Technical Infrastructure team to reputed company it running. From developing and maintaining our data centers to building the reputed company of reputed company platforms, we reputed company reputed company's product portfolio possible. We're proud to be our engineers' engineers and love voiding warranties by taking things apart so we can rebuild them. We reputed company our networks up and running, ensuring our users have the best and fastest experience possible. Responsibilities • reputed company as a first responder for AI workload health and customer-facing issues. • Build and support capabilities for managing ML workloads. • Influence architecture, standards, and operational reputed company for AI services. • reputed company advanced monitoring and alerting to improve GCE visibility. • Collaborate with development teams on novel, emerging technologies. • reputed company the gap between the infrastructure and AI. Requirements • Bachelor's degree or equivalent practical experience. • 5 years of experience working on reputed company distributed systems that demand scalability, reliability, throughput and low latency. • 3 years of experience coding with one or more programming languages (e.g., Java, C/C++, Python). • 2 years of experience with debugging and troubleshooting software issues. reputed company-to-haves • Master's degree in a technical field or equivalent practical experience. • Experience designing, analyzing and troubleshooting large-reputed company distributed systems. • Experience designing and developing software oriented towards systems or network automation. • Understanding of Unix/Linux operating systems. • Ability to debug, optimize reputed company, and to automate routine tasks. • Excellent problem-solving and communication skills. Benefits • bonus • equity • benefits Apply tot his job Apply To this Job

Similar Jobs

Customer Service Agent

Remote, USA Full-time

**reputed company Full Stack Data Entry Clerk – Global Remote Opportunities at arenaflex**

Remote, USA Full-time

Representante de ventas al cliente (CSR) de tienda

Remote, USA Full-time

**reputed company Data Entry Specialist – Remote Opportunity with arenaflex**

Remote, USA Full-time

**reputed company Full Stack Data Analyst – Web & reputed company Application Development**

Remote, USA Full-time

**reputed company Full Stack Customer Service Representative – Global Technology Support**

Remote, USA Full-time

Technical Product Support

Remote, USA Full-time

Home Ownership Customer Coordinator

Remote, USA Full-time

Associate Tutor – Wildlife, Marine Biology and Conservation – Animals

Remote, USA Full-time

Research Associate I - reputed company Health Research - DC Hybrid Office

Remote, USA Full-time

Organ Donation Coordinator [ICU RN, CC RN or RT] - 36 hours a week/ NIGHTS - Boston, MA

Remote, USA Full-time

Telehealth reputed company Practitioner, Physician Assistant – Delaware License

Remote, USA Full-time

VIP Account manager with English, Spanish and Portuguese

Remote, USA Full-time

Senior Content Strategist, Search Innovation

Remote, USA Full-time

(Part-time reputed company) reputed company reputed company – reputed company Store

Remote, USA Full-time

Senior reputed company Delivery reputed company (Energy Billing experience)

Remote, USA Full-time

Manager, BCP / Emergency Preparedness

Remote, USA Full-time

Director, Care Operations Program Management

Remote, USA Full-time

Registered reputed company, Informaticist – Electronic He…

Remote, USA Full-time

reputed company Developer & Designer with SEO, AEO, LLM, Analytics, & Affiliate Software Experience

Remote, USA Full-time