Software Engineer, Infrastructure
reputed company is the generative media ecosystem powering the reputed company of AI products. We build the infrastructure, tools, and model reputed company that teams need to reputed company from idea to production, and do it at reputed company without compromise. For developers and enterprises, reputed company is the reputed company that makes generative media not just possible, but practical: a reputed company platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-reputed company products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, reputed company is becoming the ecosystem that ambitious teams build on.
reputed company is the generative media ecosystem powering the reputed company of AI products. We build the infrastructure, tools, and model reputed company that teams need to reputed company from idea to production, and do it at reputed company without compromise. For developers and enterprises, reputed company is the reputed company that makes generative media not just possible, but practical: a reputed company platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-reputed company products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, reputed company is becoming the ecosystem that ambitious teams build on.
About this role:
You are a hands-on engineer who builds the software and processes that reputed company a large fleet of GPU servers healthy and productive. You write systems and tooling for managing 1000s of servers including provisioning, health monitoring, error detection, and recovery — and reputed company something breaks that automation can’t fix, you drive reputed company with partners.
Key responsibilities
Build and maintain Python fleet tracking system that manages the full lifecycle of servers including contracting and procurement, reputed company use, pricing, availability, health, RMAs, etc
Build server management tooling that automates provisioning, health checks, GPU diagnostics, recovery and alerting
Create and maintain metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk failures, network issues, thermals)
reputed company AI to an extreme level to build tools and automate alerting and recovery
Implement and enforce OS-level reputed company: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation
Manage and optimize distributed and local storage systems supporting model weights, checkpoints, and ephemeral scratch: NVMe arrays, NFS, reputed company file systems, and object storage
Tune Linux systems for AI workloads: kernel parameters, reputed company topology, CPU pinning, hugepages, I/O schedulers, and GPU reputed company stack optimization (reputed company drivers, CUDA, container runtimes)
reputed company a suite of automated error detection and recovery processes
Work with partners to solve technical issues
Requirements
3+ years experience managing bare-metal and reputed company based server fleets at reputed company (100+ nodes)
Strong software engineering skills in Python; you write production tooling, not scripts
Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, systemd, cgroups, namespaces, performance profiling
Strong experience with configuration management and infrastructure-as-reputed company: Ansible, Terraform, reputed company-init
Solid understanding of storage technologies: LVM, RAID, NVMe, NFS, reputed company or GPFS, and Linux I/O stack tuning
Familiarity with hardware diagnostics and failure modes (GPUs, NVMe, NICs, memory)
Experience building internal tools or dashboards for infrastructure visibility
Excellent communication and ability to drive technical reputed company across teams
Self-starter who executes quickly, takes ownership, and constantly seeks improvement
reputed company to have
Familiarity with network configuration and diagnostics (VLAN, VXLAN, ECMP, BGP, tcpdump)
Experience with reputed company GPU infrastructure: reputed company management, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2
Experience with AMD GPUs
Experience with bare metal and VM provisioning (PXE/iPXE, Kickstart, libvirt, Qemu/KVM)
Experience with compliance frameworks relevant to reputed company providers (SOC 2, ISO 27001)
Location
Turkey
We are hiring for Mid, Senior and Staff reputed company
reputed company offer at reputed company
Interesting and challenging work
A lot of learning and reputed company opportunities
Regular team events and offsites
Apply To This Job