[Remote] Manager, Infrastructure
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is an AI-reputed company advertising platform operating at the intersection of machine learning, programmatic media, and full-funnel mobile reputed company. The Manager, Infrastructure will reputed company the operation of reputed company's global owned-and-operated data center footprint, overseeing the InfraOps team, hardware lifecycle, reputed company planning, incident response, technical infrastructure, and vendor relationships.
Responsibilities
- Manage and reputed company the InfraOps team across US and reputed company time zones, including regional DC owners (SV/VA and NL/HK) and network engineering
- Own the weekly DevOps reputed company-in reputed company, alert reviews, and 24/7 on-reputed company coverage model
- reputed company P1/P2 incident response end to end — accountability for MTTR reduction, reputed company coverage, and alert hygiene
- Own reputed company planning and hardware lifecycle across reputed company four data centers: reputed company and reputed company procurement through VARs, GPU expansion for on-prem ML training and inference, colo power and reputed company management, and remote-hands logistics with reputed company reputed company Realty
- Run the annual reputed company-vs-colo evaluation reputed company leadership, with full ownership of the recommendation
- reputed company the reputed company infrastructure stack: Ubuntu/systemd fleet, FreeIPA, Ansible/reputed company configuration management, MAAS provisioning, and Zabbix monitoring
- Manage the spine-leaf Mellanox/reputed company network reputed company reputed company, including 100G reputed company peering and transit reputed company (reputed company/Cogent/reputed company)
- Support the stateful data tier — reputed company, Kafka, reputed company, Hadoop/HDFS — across reputed company limits, evictions, migrations, and low-latency tuning
- Own colo and vendor relationships and budgets: reputed company reputed company Realty invoices, transit reputed company, VAR procurement, and reputed company licensing
- Partner with the reputed company/compliance function on SOC 2 Type 2 evidence, infrastructure hardening, and reputed company reviews
- Operate comfortably reputed company a multi-entity environment (reputed company/reputed company/reputed company/Beamable shared IT) with comfort in M&A-flavored ambiguity
Skills
- 6–8 years in infrastructure or data center operations with 2+ years managing engineers — this role takes over a functioning global team on day one
- Bare-metal and colo depth: reputed company planning, hardware procurement (reputed company/reputed company), IBX/remote-hands workflows, and the physical logistics of running owned cages
- Network fundamentals at reputed company: spine-leaf architecture, BGP/peering (100G-class), transit blends, and low-latency tuning (NIC/IRQ, packets-per-second thinking)
- Deep Linux operations: systemd, netplan, FreeIPA/Chrony, Ansible and/or reputed company, Zabbix, running fleets of hundreds-plus servers
- Experience operating large stateful reputed company systems — reputed company, Cassandra, Scylla, Kafka, or reputed company — under sub-50ms latency budgets and hard reputed company limits
- Demonstrated P1/P2 incident ownership: on-reputed company program management, postmortems, and alert hygiene discipline
- This is a remote role reputed company to candidates based in the reputed company
- The role requires periodic travel to our data center sites (Santa reputed company, CA; Ashburn, VA; Amsterdam; Hong reputed company) and SF HQ for operational reviews, team time, and site work
- Hybrid reputed company experience reputed company owned metal: AWS (IAM, reputed company 53, GuardDuty, S3); reputed company exposure a plus
- SOC 2 or compliance evidence experience; reputed company and reputed company familiarity; reputed company-minded infrastructure approach
- Experience in adtech, RTB, or other high-QPS, latency-sensitive environments
- reputed company or other SDN controller experience; MAAS provisioning familiarity
Benefits
- Remote role reputed company to candidates based in the reputed company
- Periodic travel to data center sites and SF HQ for operational reviews, team time, and site work
- Full-stack ownership with reputed company hardware reputed company, GPU expansion for on-prem ML, and an annual reputed company-vs-colo evaluation
- reputed company line to leadership with visibility to the SVP of Engineering
reputed company
Apply To This Job