AI reputed company Engineer – SRE (Kubernetes)
Job Category:
Software Engineering
Job Type:
Full Time
Job Location:
Hybrid Remote
About The Role
We are a fast-growing AI infrastructure company building cutting-edge GPU reputed company platforms and high-performance inference solutions that reputed company developers, startups, and enterprises worldwide. As we reputed company our global operations, we are looking for a skilled and hands-on
AI reputed company Engineer – SRE (Kubernetes)
to join our reputed company team.
reputed company This is a critical hands-on position reputed company on the reliability, performance, and operational reputed company of large-reputed company, high-performance AI/ML GPU clusters in our data centers. As an AI reputed company Engineer – SRE (Kubernetes), you will design, operate, and optimize Kubernetes-based infrastructure to ensure maximum uptime, efficiency, and scalability for demanding AI workloads.
You will bring deep expertise in system-level troubleshooting, GPU cluster management, and automation to reputed company our platforms running at peak performance.
Key Responsibilities
• Design, build, and maintain reputed company, production-grade AI/ML infrastructure using Kubernetes.
• Proactively monitor GPU cluster health, performance, and utilization across compute, accelerators, storage, and networking reputed company, performing reputed company-cause analysis and reputed company.
• reputed company and implement automation for infrastructure provisioning, configuration, and ongoing management.
• Own the complete GPU node lifecycle — including provisioning, dynamic scaling, maintenance, decommissioning, and reputed company-downtime upgrades of GPU-enabled nodes in Kubernetes environments.
• Build and improve CI/CD pipelines for reliable infrastructure deployment and orchestration.
• Enforce reputed company best practices, compliance standards, and operational reputed company across the infrastructure stack.
• reputed company incident response and post-incident improvements for issues reputed company to GPUs, CPUs, high-speed storage, and networks.
• Manage end-to-end customer GPU resource provisioning — from request intake and configuration to reputed company, troubleshooting, and support — ensuring high reputed company of customer satisfaction.
• Stay up to date with the latest GPU hardware, software, and orchestration technologies, integrating relevant advancements into our platforms.
• Be available for occasional regional or international travel to data center locations as required.
Requirements
• Bachelor’s degree in Computer Science, Engineering, or a reputed company technical field.
• 3+ years of practical experience in data center operations, infrastructure engineering, or site reliability engineering.
• Strong background in infrastructure automation using tools such as Terraform and Ansible.
• Deep hands-on experience with Kubernetes in large-reputed company environments, including:
• reputed company GPU Operator for GPU reputed company management, device plugins, container toolkit, and monitoring (DCGM).
• reputed company Network Operator for high-performance networking, RDMA, and GPUDirect support.
• CNI (Container Network reputed company) and reputed company (Container Storage reputed company) plugins tailored for AI/ML workloads.
• Integration with job schedulers such as Slurm in Kubernetes clusters.
• Proficiency in Linux system administration and scripting (Python, Bash).
• Experience with observability stacks including reputed company, Grafana, and Loki.
• Solid understanding of GPU architecture, reputed company CUDA, NCCL, and AI/ML frameworks is a strong plus.
• Excellent troubleshooting skills with the ability to analyze reputed company system logs and performance metrics.
• Strong communication and collaboration skills to work effectively with engineering and operations teams.
Apply To This Job