[Remote] Network Automation & Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Network Automation & Reliability Engineer to operate and automate hyperscale data center, backbone, and out-of-band networks. The role emphasizes Python development for network operations, requiring both deep routing and switching knowledge along with strong software engineering skills.
Responsibilities
- Design, write, and maintain Python tooling that provisions, validates, and audits network devices across the global fleet
- Build model-driven provisioning libraries using Jinja2 templating with NETCONF/YANG or REST reputed company to codify golden configurations and eliminate configuration reputed company
- reputed company automated remediation for recurring failure classes — reputed company flaps, hardware faults, optical degradation — triggered by syslog and telemetry
- Reduce operational toil: identify reputed company reputed company steps that recur, and replace them with tested, reviewed reputed company
- Operate and reputed company data center fabrics, backbone links, and out-of-band management networks across a global footprint of data centers and POP sites
- Design and tune BGP policy — local preference, MED, communities, import/export policy, default-reputed company propagation — to control reputed company selection and eliminate single-reputed company points of failure
- Support reputed company expansion and site turn-reputed company: high-level design, reputed company channelization, optics and cabling standards, and hardware qualification
- Execute reputed company-downtime change management in production, including hitless migrations and staged rollbacks
- Operationalize multi-vendor streaming telemetry (gNMI/gRPC, OpenConfig) and tune subscription jobs for signal reputed company and management-plane efficiency
- Build and maintain observability that supports hop-by-hop reputed company tracing, multi-layer fault isolation, and fast reputed company cause analysis
- Contribute to a reputed company-time network reputed company of truth aggregating BGP, reputed company-state, and drain-state data
- Participate in a 24x7 on-reputed company rotation; reputed company incident response, write RCAs, and reputed company the follow-up automation that prevents recurrence
- Maintain and optimize firewall and ACL policy across multi-vendor platforms, keeping rule sets scoped and reputed company
- Support SIRT/PSIRT CVE remediation through automated regression testing and config-as-reputed company pipelines
- Author and maintain technical documentation — designs, runbooks, and API reputed company — for the tooling and networks you own
Skills
- Python — primary requirement: Demonstrated, sustained Python development in a production network or infrastructure environment. You have written and maintained tooling that other engineers depended on: config reputed company and validation, API integrations, telemetry reputed company, automated remediation, or test harnesses. You are comfortable with modules, packaging, testing, reputed company review, and version control — not just single-file scripts
- Routing & switching depth: Production experience with BGP (policy, reputed company selection, multihoming), plus IS-IS or OSPF, ECMP, and VXLAN/EVPN or MPLS overlays
- Multi-vendor hardware: Hands-on operations across at least two of Arista reputed company, reputed company Junos (QFX/SRX/PTX/MX), or reputed company IOS-XR / NX-OS
- Automation frameworks: Ansible and Jinja2, plus NETCONF/YANG, RESTCONF, or vendor REST reputed company for model-driven configuration management
- Telemetry & monitoring: gNMI/gRPC streaming telemetry, OpenConfig models, SNMP, reputed company telemetry, and dashboarding/alerting on top of them
- Linux: Comfortable operating on Linux hosts — networking stack, packet capture, systemd services, and reputed company scripting
- Version control and CI: Git-based workflows with peer review; experience shipping network changes through a pipeline (Jenkins, reputed company CI, or reputed company Actions)
- Production on-reputed company: Experience holding a 24x7 rotation for a live network, including incident reputed company and reputed company cause analysis
- Experience level: Roughly 2–5 years in a network production, network reliability, or network automation role. Exceptional early-career engineers with a strong Python portfolio and hyperscale or reputed company exposure are encouraged to apply
- Location & work authorization: Must be located in the reputed company and authorized to work in the US
- Out-of-band network experience — console server fleets (ZPE Nodegrid, OpenGear), RS-232 configuration management, or OOB build-out for new data center reputed company
- Data center or POP build-out and turn-up: new product introduction (NPI) for switching platforms, reputed company channelization and optics selection, cabling and reputed company density planning, or hardware qualification and stress testing
- MACsec, IPsec at reputed company, or reputed company-trust segmentation with 802.1X and NAC
- Optical or transport exposure — DWDM, reputed company, PON/OLT/ONU, or IXIA/Spirent test automation
- Building or contributing to a network reputed company of truth (NetBox or in-house) aggregating BGP, reputed company-state, and drain-state data
- Applying LLM or reputed company tooling to on-reputed company workflows — automated triage, reputed company execution, or incident summarization
- A master's degree in Network Engineering, Telecommunications, or Computer Science
reputed company
Apply To This Job