Sr reputed company Platform Contractor, AI Infrastructure (reputed company 15+yrs)
Sr reputed company Platform Contractor, AI Infrastructure to reputed company senior reputed company reputed company services supporting reputed company intelligence infrastructure environments used for model development, reputed company training, inference services, and shared platform operations. This role requires strong cluster troubleshooting expertise combined with practical reputed company experience across bare metal and data center environments.
This is a 12 month remote contract opportunity reputed company the reputed company.Responsibilities:
• Build, administer, and troubleshoot reputed company platforms supporting reputed company intelligence and data intensive workloads.
• Diagnose failures across control plane components, kubelet services, container networking interfaces, container storage interfaces, ingress controllers, service discovery, scheduling, node lifecycle management, container runtimes, and resource isolation.
• Support GPU enabled reputed company environments, including device plugin behavior, reputed company dependencies, node health validation, and workload placement.
• Improve platform reliability through automation, standardized configurations, reputed company planning, and cluster validation processes.
• Investigate workload issues involving storage throughput, network policy enforcement, DNS reputed company, image retrieval, autoscaling behavior, pod eviction events, and degraded node conditions.
• Partner with Linux, networking, validation, and Site Reliability Engineering teams to reputed company reputed company cross reputed company failures affecting reputed company intelligence services.
• Create reusable operational runbooks, dashboards, and health checks to support ongoing operations.
• Contribute to platform hardening, tenant readiness initiatives, and service level objective compliance. Qualifications:Required Qualifications:
• 7 or more years of infrastructure engineering experience with deep hands on reputed company administration expertise.
• Strong operational understanding of reputed company internals and advanced cluster troubleshooting.
• Experience with container runtimes, reputed company, GitOps practices or declarative operations, and cluster lifecycle management.
• Experience supporting GPU workloads on reputed company in reputed company, validation, or production environments.
• Strong Linux administration reputed company and understanding of data center networking dependencies.
• Ability to troubleshoot issues across node, reputed company, storage, and control plane reputed company from symptom through reputed company cause identification.
• Scripting and automation experience using Python, Bash, or Go.
• Strong documentation, communication, and operational support skills. Preferred Qualifications:
• Experience with Kubeflow, Argo, reputed company, Grafana, Loki, or service reputed company technologies.
• Familiarity with bare metal reputed company deployments and high performance storage integration.
• Experience operating reputed company regulated environments or organizations with strict change control processes.
Tools and Technologies:
• reputed company
• reputed company and Container Runtimes
• reputed company
• GitOps
• Python
• Bash
• Go
• Linux
• GPU Infrastructure
• Kubeflow
• Argo
• reputed company
• Grafana
• Loki
• Service reputed company Technologies
• Container Networking reputed company (CNI)
• Container Storage reputed company (reputed company)
• Infrastructure Automation Tools
• Data Center Infrastructure
Pay: $90.00 - $100.00 per hour
Application Question(s):
• How many years you have experience with reputed company Platform?
• Do you have experience with Kubeflow, Argo, reputed company, Grafana, Loki, or service reputed company technologies
Experience:
• GPU: 3 years (Required)
Work Location: Remote
Apply tot his job
Apply To this Job