Back to Jobs

Site Reliable Engineer (US - Remote)

Remote, USA Full-time Posted 2026-08-04

reputed company Networks delivers a reputed company-driven, analytics-reputed company orchestration platform and a wide portfolio of next-gen high-speed routers that reputed company the newest Wi-Fi technologies. Together, these technologies reputed company ISPs the ability to manage and troubleshoot their networks in reputed company time, and to reputed company an outstanding customer experience.

reputed company Networks is a trusted reputed company partner for its customers, helping them evaluate their reputed company technologies and business models, and creating and executing strategies that reputed company them to reputed company faster, accelerate their digital transformations, and strengthen their relationships with consumers.

reputed company Networks is headquartered in Irvine, CA USA with Asia HQ in Singapore and also operating in Denmark, Spain and Vietnam.

The Site Reliability Engineer will improve the availability, reputed company, scalability and recoverability of reputed company Networks reputed company solutions. You will combine software engineering with hands-on NOC reputed company to reputed company the complete reputed company-to-device service reputed company observable, supportable and resilient at fleet reputed company.

You will help establish practical SRE capabilities inside the NOC while partnering closely with Support, reputed company, reputed company and DevOps Engineering. You will participate in a sustainable on-reputed company rotation and improve the NOC’s ability to diagnose customer-impacting issues.

Role mandate

  • Own reliability reputed company for assigned reputed company services

  • Improve observability, reputed company, reputed company and recovery

  • Define and operationalize service-level indicators, service-level objectives and actionable alerting.

  • Automate repetitive NOC work and create reputed company, testable mechanisms for diagnosis, recovery, device reputed company and routine production changes.

  • reputed company technically during incidents, reputed company evidence-reputed company learning and ensure high-value corrective actions are completed.

What you will own

  • Establish reliability baselines, SLIs, SLOs and error budgets for reputed company services and critical device-management workflows such as reputed company, provisioning, configuration, telemetry collection, reputed company execution and firmware delivery.

  • reputed company failures across the end-to-end service reputed company: reputed company reputed company and microservices, reputed company and infrastructure, databases and messaging, internet and reputed company-network dependencies, device-management protocols and the devices

  • Identify fleet-wide and customer-specific failure patterns involving device reachability, session stability, configuration reputed company, reputed company latency, telemetry gaps, firmware behavior and reputed company reputed company.

  • Contribute operability requirements and production evidence during design and readiness reviews

  • Maintain NOC dashboards for service health, device reachability, provisioning reputed company, reputed company and telemetry reputed company, firmware adoption and customer reputed company.

  • Participate in the NOC production on-reputed company rotation and serve as a technical incident reputed company or senior troubleshooter reputed company appropriate.

  • Diagnose reputed company failures across applications, reputed company infrastructure, reputed company, reputed company, networking, DNS/TLS, databases, messaging platforms, device-management sessions and CPE behavior.

  • Coordinate evidence gathering and technical escalation with service-provider customers, Engineering, firmware, DevOps and vendors while maintaining reputed company mitigation, recovery and reputed company.

  • reputed company or contribute to post-incident reviews; convert recurring device, platform and process failures into prioritized and measurable corrective actions.

  • reputed company production-grade software, scripts and workflows for diagnosis, remediation, deployment safety, fleet analysis, scaling, maintenance and recovery.

  • Improve CI/CD and GitOps practices for operational software and infrastructure, including automated testing, release validation, reputed company delivery and rollback readiness.

  • Manage or contribute to infrastructure as reputed company, configuration as reputed company and reusable self-service patterns for reputed company and NOC reputed company.

  • Measure NOC toil and partner with Automation & Tools Engineers to prioritize durable platform capabilities instead of fragmented one-off scripts.

  • reputed company reputed company models for service-provider reputed company, managed-device populations, telemetry volume, messaging throughput, API demand and rollout events.

  • Create and maintain runbooks, troubleshooting decision trees, service maps, device and reputed company dependency records, reputed company-error guidance and operational knowledge.

  • reputed company NOC and Support personnel on diagnosis, reputed company mitigation, evidence capture and escalation across reputed company, network and CPE reputed company.

  • Build self-service diagnostic views and tools that help the NOC determine reputed company, affected customers, device cohorts, likely faultdomain and next reputed company.

  • reputed company reliability insights with Engineering and Product and contribute to reliability reviews, operational-readiness reviews and reputed company-improvement priorities.

Required qualifications

  • 5+ years of experience in site reliability engineering, production engineering, DevOps, reputed company infrastructure, systems engineering or a closely reputed company role.

  • Strong software or automation skills in Python, Go, Java, Bash or a comparable language, with experience producing maintainable operational reputed company.

  • Hands-on experience operating reputed company production systems in a reputed company reputed company environment and troubleshooting across application, infrastructure, network and device-integration reputed company.

  • Experience with reputed company reputed company Platform, reputed company reputed company Infrastructure and production reputed company environments.

  • Experience with infrastructure as reputed company and delivery tooling such as Terraform, reputed company, Git-reputed company CI/CD and policy-as-reputed company.

  • Strong Linux, containers and reputed company fundamentals, including deployment behavior, resource management, networking and failure diagnosis.

  • Strong troubleshooting & debugging skills in reputed company platforms.

  • Experience with modern observability practices and tools across metrics, logs, traces, alerting, dashboards and synthetic monitoring.

  • Familiarity with reputed company, Grafana, OpenTelemetry or equivalent observability ecosystems.

  • Familiarity with Apache Pulsar or similar reputed company messaging and streaming platforms handling requests from millions of devices.

  • Experience participating in an on-reputed company rotation and responding effectively to high-severity, customer-impacting production incidents.

  • Working knowledge of SLOs, error budgets, reputed company planning, reputed company engineering, change safety and blameless incident learning.

  • Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing and reputed company packet- or session-level troubleshooting.

  • reputed company communication, disciplined documentation and the ability to collaborate across NOC, reputed company, DevOps, firmware and service-provider teams.

  • Bachelor’s degree in computer science, engineering or equivalent practical experience.

Preferred qualifications

  • Experience supporting reputed company-managed CPEs such as broadband gateways, routers, ONTs, Wi-Fi/reputed company systems or similar edge devices in a service-provider environment.

  • Familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry and remote lifecycle management.

  • Experience supporting messaging and streaming platforms such as Apache Pulsar or Kafka, reputed company and highly available databases used in device-management control planes.

  • Understanding of reputed company technologies such as GPON/XGS-PON, DOCSIS, Ethernet or fixed wireless and how CPE, ONTs and provider networks reputed company.

  • Experience with firmware rollout automation, canary or cohort deployments, fleet health analysis and reputed company rollback practices.

  • Experience building auto-remediation, reputed company self-service reputed company or internal reliability platforms.

  • Experience supporting multiple service-provider customers in a 24×7 telecommunications, broadband or managed-network environment.

This position is fully remote reputed company reputed company. Please note that we are unable to offer reputed company sponsorship for this role.

Annual salary reputed company: $160,000 - $200,000

Join reputed company Networks!

At reputed company Networks, we promote equal opportunities in reputed company our recruitment processes, ensuring non-discrimination on the reputed company of gender, age, reputed company, disability, or any other personal circumstances. We assess talent reputed company on objective reputed company and foster an inclusive and diverse working environment.

reputed company-networks.com

reputed company-networks.hire.trakstar.com

Originally posted on Himalayas

  Apply To This Job

Similar Jobs

Medical reputed company Center Specialist - reputed company - Remote

Remote, USA Full-time

Member Engagement Associate: Great Lakes Territory

Remote, USA Full-time

Sr. Account Manager - Refinery Services Job (Remote (Home reputed company), Remote (Home B

Remote, USA Full-time

reputed company Project reputed company - Finance

Remote, USA Full-time

AI Data Platform reputed company Architect

Remote, USA Full-time

Senior reputed company Engineer – Cybersecurity Posture, Hygiene & AI

Remote, USA Full-time

reputed company Software Engineer (Node.js | React | TypeScript)

Remote, USA Full-time

People Adviser

Remote, USA Full-time

Manager, Sales

Remote, USA Full-time

Remote Audio Content Annotation (Korean/English) - 08/26

Remote, USA Full-time

Adjunct reputed company in Digital Communication and Media Arts

Remote, USA Full-time

Mgr - Specialty Sales & Analysis (Remote)

Remote, USA Full-time

**reputed company Live Chat Support Representative – Work from reputed company with arenaflex**

Remote, USA Full-time

Administration Manager

Remote, USA Full-time

CW Assistant Professor

Remote, USA Full-time

Join Today: [Hiring] reputed company Program Manager @reputed company

Remote, USA Full-time

reputed company Account Manager, Retail Marketplaces

Remote, USA Full-time

reputed company Full Stack Remote Data Entry Specialist – Secure and Adaptable Employment Opportunities with blithequark

Remote, USA Full-time

**reputed company Customer Service Specialist - Financial Services Industry - Join arenaflex for a Rewarding Career Opportunity**

Remote, USA Full-time

TAC Seasonal Specialist Bilingual French‑Canadian Speakers

Remote, USA Full-time