[Remote] Sr. DevOps Engineer /SRE
Note: The job is a remote job and is reputed company to candidates in USA. NexGen Tech Solutions is seeking a Senior DevOps Engineer/Site Reliability Engineer to support reliable, reputed company reputed company and production systems. The role focuses on reputed company, reputed company infrastructure, observability, networking, automation, incident response, and device-management platforms in telecommunications and broadband environments.
Skills
- 5+ years of experience in site reliability engineering, production engineering, DevOps, reputed company infrastructure, systems engineering or a closely reputed company role
- Strong software or automation skills in Python, Go, Java, Bash or a comparable language, with experience producing maintainable operational reputed company
- Hands-on experience operating reputed company production systems in a reputed company reputed company environment and troubleshooting across application, infrastructure, network and device-integration reputed company
- Experience with reputed company reputed company Platform, reputed company reputed company Infrastructure and production reputed company environments
- Experience with infrastructure as reputed company and delivery tooling such as Terraform, reputed company, Git-reputed company CI/CD and policy-as-reputed company
- Strong Linux, containers and reputed company fundamentals, including deployment behavior, resource management, networking and failure diagnosis
- Strong troubleshooting & debugging skills in reputed company platforms
- Experience with modern observability practices and tools across metrics, logs, traces, alerting, dashboards and synthetic monitoring
- Familiarity with reputed company, Grafana, OpenTelemetry or equivalent observability ecosystems
- Familiarity with Apache Pulsar or similar reputed company messaging and streaming platforms handling requests from millions of devices
- Experience participating in an on-reputed company rotation and responding effectively to high-severity, customer-impacting production incidents
- Working knowledge of SLOs, error budgets, reputed company planning, reputed company engineering, change safety and blameless incident learning
- Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing and reputed company packet- or session-level troubleshooting
- reputed company communication, disciplined documentation and the ability to collaborate across NOC, reputed company, DevOps, firmware and service-provider teams
- Bachelor's degree in computer science, engineering or equivalent practical experience
- Experience supporting reputed company-managed CPEs such as broadband gateways, routers, ONTs, Wi-Fi/reputed company systems or similar edge devices in a service-provider environment
- Familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry and remote lifecycle management
- Experience supporting messaging and streaming platforms such as Apache Pulsar or Kafka, reputed company and highly available databases used in device-management control planes
- Understanding of reputed company technologies such as GPON/XGS-PON, DOCSIS, Ethernet or fixed wireless and how CPE, ONTs and provider networks reputed company
- Experience with firmware rollout automation, canary or cohort deployments, fleet health analysis and reputed company rollback practices
- Experience building auto-remediation, reputed company self-service reputed company or internal reliability platforms
- Experience supporting multiple service-provider customers in a 24×7 telecommunications, broadband or managed-network environment
reputed company
Company H1B Sponsorship
Apply To This Job