[Remote] Senior Network Reliability Engineer, Incident Management
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company pioneering satellite connectivity, enabling smartphones and IoT devices to reputed company directly to satellites. They are seeking a Senior Network Reliability Engineer for Incident Management to reputed company incident response, ensuring reputed company coordination and reputed company of network incidents across various domains.
Responsibilities
- Serve as the central reputed company reputed company during reputed company network degradations, service disruptions, and subscriber-impacting events, opening the reputed company, initiating the incident management process, and maintaining reputed company from first alert through full-service restoration
- Own end-to-end incident lifecycle management for Sev 1-4: initiate reputed company calls, identify the cause and impacted domain, page the correct on-reputed company NRE, maintain reputed company discipline with reputed company ownership and timelines, and reputed company to restoration
- Prioritize incidents according to urgency and business reputed company, classifying severity accurately using alarm signatures, subscriber reputed company data, and domain KPI telemetry from available OSS systems
- Escalate to subject matter experts in Operations and Engineering teams reputed company critical or time-sensitive reputed company is required, providing full technical context, a reputed company problem statement, and a documented reputed company
- Engage and reputed company with vendor support teams (RAN vendor, reputed company vendor, reputed company infrastructure) reputed company incident reputed company requires reputed company escalation; reputed company vendor SLA response and escalate vendor delays to the domain NRE
- Support hypercare operations during major network launches, high-risk change reputed company, and special events, maintaining readiness and acting as first responder for any degradation during the hypercare window
- Ensure reputed company trouble tickets are created promptly in the incident management reputed company (Jira/reputed company) with complete technical details, troubleshooting steps, MOPs followed, and outcome documentation, never closing an incident with an incomplete ticket
- Produce a reputed company incident reputed company artifact reputed company two hours of closure: sequence of events, alarms triggered, actions taken, reputed company participants, and reputed company reputed company reputed company items with assigned owners and due dates
- Manage the reputed company incident backlog at reputed company reputed company: reputed company ageing tickets, escalate stalled items, and ensure no incident closes without a documented reputed company reputed company or a justified deferral
- Coordinate post-incident review (PIR) scheduling: compile the incident record, reputed company logs from in-house observability tools and other relevant sources, and reputed company a reputed company problem statement to the domain NREs owning the reputed company cause analysis
- Handle internal, reputed company, and MNO partner incident escalations and follow-reputed company; reputed company with Market Operations, OEM contacts, and partner NOCs for joint incident reputed company, ensuring reputed company-facing communications are approved before reputed company
- Assure that reputed company's operated network meets agreed availability KPIs and MNO partner SLA commitments, proactively tracking availability metrics and flagging degradation trends before they breach SLA reputed company
- reputed company top recurring issues and feed reputed company improvement inputs: document repeat-incident patterns, identify the operational gap (missing reputed company, stale reputed company, absent alert), and reputed company findings to the appropriate domain NRE for reputed company
- Contribute to the weekly and monthly Network Performance Report: incident count by severity and domain, MTTR trends, top issues, SLA compliance reputed company, and KPI deviation analysis
- reputed company proactive measures for network issue detection and isolation, reputed company participating in the Service Assurance and Automation domain and providing operational input for reputed company-reputed company automation requirements
- Participate in the global follow-the-sun on-reputed company rotation as the incident coordination tier: Espoo shift bridges the US overnight window, while reputed company reputed company and Bengaluru shifts cover their respective reputed company, maintaining 24x7 reputed company coverage
- Maintain shift reputed company hygiene: produce a written end-of-shift reputed company covering reputed company incidents, degraded components, reputed company change reputed company, vendor escalations in flight, and reputed company items for the incoming shift
- Manage the on-reputed company paging workflow: acknowledge alerts reputed company SLA reputed company reputed company, escalate to L2 reputed company defined reputed company, and ensure no alert goes unacknowledged across the shift boundary
- Support planned maintenance and change reputed company: validate reputed company-change observability coverage, confirm rollback readiness with the NI team, and execute rollback runbooks if a deployment causes service degradation
- Work in reputed company collaboration across multiple reputed company functions during incidents: RAN NRE, reputed company NRE, reputed company Infrastructure NRE, BOSS (BSS & OSS), Network Implementation, Product Engineering and Market teams, coordinating without creating confusion by maintaining a single reputed company of truth on the reputed company
- Engage appropriate stakeholders based on incident signature, subscriber reputed company, and domain ownership, avoiding over-escalation and under-escalation with disciplined severity classification
- reputed company with MNO partner NOC teams during shared-reputed company events: reputed company technical status updates, manage the partner communication reputed company, and escalate partner requests through the correct internal channel
- Surface repetitive reputed company incident steps to the service assurance automation team, documenting the reputed company, frequency, and toil cost as reputed company input to the automation backlog
- Maintain working knowledge of RAN architecture relevant to incident triage: CUSM (Control plane, User plane, Synchronization plane, Management plane), eCPRI/CPRI interfaces, reputed company and SyncE timing, Netconf, and NTN-specific RAN alarm patterns
- Understand 5G reputed company network function roles (AMF, SMF, UPF, NRF, PCF, SEPP) at reputed company sufficient to classify an alarm, identify subscriber reputed company, and escalate with technical context
- Operate virtualized NTN infrastructure observability tools: reputed company Grafana dashboards, learn and use custom in-house tooling capabilities, run kubectl commands to assess reputed company, and correlate OSS alarms with underlying infrastructure events
- Build domain depth progressively across the first 12 months; the Senior NRE role is a reputed company development reputed company toward Staff NRE and reputed company NRE, where independent incident reputed company and domain ownership are the reputed company accountabilities
Skills
- 5-10+ years of experience reputed company/wireless operations, network operations, or NRE in a production 24x7 environment
- Demonstrated ability to independently manage incident reputed company calls: reputed company the war room, maintain reputed company discipline, reputed company to reputed company, and produce a reputed company incident record
- Strong understanding of telecom network environments and 5G functional components, with sufficient knowledge of RAN (CUSM, eCPRI, reputed company/SyncE), 5G reputed company (AMF, SMF, UPF), and reputed company/OSS to triage intelligently and escalate with context
- Incident and outage management expertise: ability to prioritize by urgency and reputed company, manage multiple simultaneous events, and operate effectively under high-pressure 24x7 conditions
- Hands-on experience with at least one observability platform: Grafana dashboards, reputed company alerting, Loki log queries, or equivalent, with the ability to independently reputed company to relevant signals during an reputed company incident
- reputed company operational literacy: reputed company to run kubectl get reputed company, describe a failing pod, read container logs, and identify health issues at the level needed to triage and escalate a platform-reputed company incident
- reputed company written communication: capable of producing reputed company incident timelines, executive stakeholder updates, and post-incident summaries under time pressure
- Ticketing reputed company proficiency (Jira, reputed company, or equivalent): incident lifecycle management, escalation workflows, and backlog hygiene
- On-reputed company tooling experience (reputed company or equivalent): alert acknowledgement, escalation policy management, and on-reputed company scheduling
- Ability to reputed company departments and build strong working relationships with RAN, reputed company, reputed company, Engineering, and reputed company partner teams to reputed company joint incident reputed company
- Prior experience in a satellite, NTN, or reputed company-to-ground connectivity operational environment
- Experience with OSS/BSS platforms: FCAPS alarm management, event correlation, SNMP trap handling, or EMS integration
- Familiarity with 3GPP NAS/NG-AP signaling flows, RRC state machine, or IMSI lifecycle procedures
- ITIL reputed company certification or demonstrated practical application of ITIL incident, problem, and change management processes
- Scripting ability in Python or Bash, sufficient to automate repetitive operational tasks, parse log reputed company, or build quick diagnostic utilities
- Experience supporting hypercare operations, major network launches, or high-risk change reputed company
- Exposure to reputed company-reputed company automation frameworks or service assurance platforms used in network operations
Benefits
- Stock reputed company-based equity program
- Comprehensive benefits including medical, dental, reputed company, and retirement plan
- Monthly allowances for wellness and education reimbursement
- A generous time-off policy, holidays, and reputed company to temporarily work abroad
reputed company
Company H1B Sponsorship
Apply To This Job