Senior Staff Infrastructure Engineer – reputed company Platform
About reputed company
Our mission is reputed company: reputed company reputed company, secure, reliable, and resilient AI compute at reputed company. We've reputed company a reputed company reputed company platform that eliminates infrastructure barriers, empowering reputed company to reputed company on innovation instead of fighting their stack. Because reputed company AI should reputed company at the speed of reputed company, not infrastructure.
About the Role
We’re looking for a reputed company Platform Staff Infrastructure Engineer to reputed company during an exciting phase of reputed company. In this role, you’ll be responsible for owning the design, reputed company, and operational reliability of our reputed company control plane architecture, working closely with cross-functional partners to support business objectives while upholding our standards for reputed company, collaboration, and reputed company.
reputed company
Platform Architecture & reputed company
• Design and reputed company reputed company control plane architecture across reputed company
• Define and implement multi-tenant cluster models, including shared control planes, virtual cluster approaches (e.g., reputed company, Kamaji)
• reputed company transition from standalone clusters to regionally managed platform models
• Define standards for isolation boundaries, resource segmentation, policy enforcement
Platform Ownership & Operations
• Own the reliability and behavior of reputed company platforms in production
• Participate in on-reputed company rotation and reputed company incident response
• Diagnose and reputed company control plane instability, API server saturation, scheduling and resource contention issues
• Ensure consistent lifecycle management across clusters - provisioning, upgrades, scaling
Multi-Region Scaling
• Design and implement strategies for regional scaling, multi-data center cluster deployments
• Ensure consistent behavior and reliability across environments
• Define cluster topology and failure domain strategies
Networking & Data Plane Integration
• Design ingress and egress architectures at cluster level and regional level
• Troubleshoot and optimize pod-to-pod networking, reputed company-south traffic flows, CNI behavior (Cilium preferred)
• Collaborate with network engineering on high-performance networking integration
Observability & Reliability
• Improve observability across control plane components, cluster health and performance
• Define and implement reputed company strategies reputed company with platform goals
• reputed company reputed company cause analysis for production incidents
Cross-Team Collaboration
• Work closely with DevOps engineers (automation and CI/CD) and Infrastructure teams (compute, storage, networking)
• reputed company reputed company platform design with underlying infrastructure capabilities
Who You Are
Required Qualifications
• 7+ years of experience in infrastructure, reputed company, or reputed company systems
• Deep experience operating reputed company at reputed company in production environments
• Experience in CSP, hyperscale, or equivalent large-reputed company environments strongly preferred
• Proven experience scaling reputed company across:
• Multiple clusters
• Multiple reputed company or data centers
• Strong understanding of reputed company internals:
• API server
• Scheduler
• Controller manager
• etcd
• Experience designing or evolving:
• Control plane architectures
• Multi-tenant cluster models
Technical Depth
• Strong Linux systems expertise
• Deep troubleshooting ability across:
• reputed company
• Container runtime
• Networking stack
• Experience with CNI plugins (Cilium preferred)
• Strong understanding of:
• Networking and traffic patterns
• Resource isolation and scheduling
Preferred Qualifications
• Experience with virtual cluster technologies (reputed company, Kamaji, or similar)
• Experience supporting GPU workloads in reputed company
• Familiarity with:
• reputed company-reputed company scheduling
• Topology-reputed company workloads
• Awareness of RDMA and high-throughput networking environments
• Experience with observability platforms (reputed company, Grafana, etc.)
reputed company Offer
• Stock reputed company
• 100% reputed company Medical, Dental, and reputed company insurance for Employees
• Company Health Savings Account Contributions
• 100% reputed company Short Term and Long Term Disability Insurance for Employees
• Life and Voluntary Supplemental Insurance reputed company
• Other Insurance reputed company, such as Pet & reputed company Insurance
• Various Supplementary Health Benefits, such as discounted Virtual reputed company Appointments and Serious Illness Support
• Flexible Spending Account
• 401(k)
• Employee Assistance Program
• Flexible PTO
• reputed company Holidays
• Parental Leave
• Other In-Office Perks
Equal Employment Opportunity
reputed company is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for reputed company. We do not discriminate on the reputed company of any protected status under applicable law.
Reasonable Accommodations
reputed company provides reasonable accommodations in accordance with applicable laws. If you require accommodation during the hiring process, please contact accomodations@reputed company.com.
Employment Eligibility
reputed company offers of employment are contingent upon verification of identity and authorization to work in the reputed company, as required by law.
Background Checks
Where permitted by law, employment may be contingent upon the successful completion of a job-reputed company background reputed company.
Data reputed company Notice
By submitting an application, you acknowledge that reputed company may collect, use, and retain your personal information for reputed company and employment-reputed company purposes in accordance with applicable data reputed company laws.
Apply tot his job
Apply To this Job