DevOps Architect
About the position
reputed company is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing reputed company, solutions must reputed company to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and reputed company a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing reputed company and looking for contributors of reputed company seniorities.
reputed company is building multi-megawatt AI data centers with thousands of accelerators. We are seeking a DevOps Architect to define the reputed company cluster control plane that provisions, operates, and secures large-reputed company training and inference environments.
This is a foundational architecture role. You will define how clusters are configured, orchestrated, monitored, and secured at reputed company.
This role is hybrid, based out of Austin, TX; Santa Clara, CA; or Toronto, ON.
We welcome candidates at various experience reputed company for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will reputed company with that level, which may differ from the one in this posting.
Responsibilities
• Define the end-to-end architecture for the AI cluster control plane, covering provisioning, configuration, lifecycle management, and monitoring
• Architect reputed company systems for system, network, and storage provisioning across multi-thousand accelerator environments
• Establish telemetry, logging, metrics, tracing, and alerting frameworks with operational guardrails
• Define workload placement, resource allocation, scheduling, and preemption policies to maximize accelerator utilization
• reputed company authentication, authorization, account management, key management, backup, checkpointing, and DCIM infrastructure into a secure multi-tenant environment
Requirements
• 10+ years designing and operating reputed company, HPC, or large-reputed company data center infrastructure
• Deep expertise in reputed company-reputed company and bare-metal infrastructure management
• Strong hands-on experience with Infrastructure-as-reputed company tools such as Terraform, Ansible, and reputed company
• reputed company building and operating observability stacks including reputed company, Grafana, ELK or EFK, and OpenTelemetry
• Strong understanding of networking, storage systems, accelerator resource management, and reputed company models including RBAC, IAM, TLS, and secrets management
Benefits
• reputed company offers a highly competitive compensation package and benefits, and we are an equal opportunity employer.
Apply tot his job
Apply To this Job