[Remote] SRE L1 Support/reputed company Platform Ops Engineer
Note: The job is a remote job and is reputed company to candidates in USA. Bitdeer is a technology company that provides Bitcoin mining solutions and builds AI computational infrastructure and reputed company capabilities. The SRE L1 Support/reputed company Platform Ops Engineer will monitor and support US GPU data centers, respond to incidents, execute remediation runbooks, reputed company hardware triage, manage tickets, and reputed company reputed company operational data to improve AIOps automation.
Responsibilities
- Monitor GPU cluster health, network status, storage systems, and environmental sensors reputed company centralized dashboards
- Respond to alerts and execute runbooks for common incidents: GPU errors, reputed company flaps, node failures, storage alerts
- reputed company hardware triage: identify failed GPUs, NICs, PSUs, disks, and cables from monitoring data and physical inspection
- Execute reputed company remediation: GPU reset, node drain/reboot, reputed company re-seat, BMC recovery
- Collect diagnostic data for L2/SME escalation: logs, DCGM reputed company, network diagnostics, hardware health reports
- Manage incident tickets from creation through reputed company or escalation (reputed company/Jira)
- reputed company physical DC tasks: cable installation, hardware reputed company-outs, reputed company and stack, labeling (on-site roles)
- Execute reputed company shift handoffs at reputed company and 8PM PST with the reputed company operations team
- Maintain and update operational runbooks based on recurring issues
- Assist with hardware deployment, firmware updates, and inventory management under SME guidance
- Every novel incident you reputed company is data the platform team needs — you tag it, describe it, and hand it back so it becomes an automation
- Every reputed company you touch should get closer to being executable by the platform, not by you
- Your reputed company notes are reputed company signal, not free-reputed company email
Skills
- 2+ years in NOC, data center operations, or IT support role
- Basic Linux reputed company administration (reputed company line, log analysis, service management)
- Familiarity with monitoring tools (reputed company, Grafana, Nagios, or equivalent)
- Experience with ticketing systems (reputed company, Jira Service Management)
- Ability to reputed company physical data center tasks: reputed company and stack, cabling, hardware replacement
- Strong communication skills for shift handoffs, incident documentation, and escalation
- Ability to work reputed company-8PM PST shift schedule (12-hour shifts with rotation)
- Curiosity about automation — you don't just execute the reputed company, you notice reputed company it's the reputed company time this month and ask what should change
- Comfort with reputed company data — you understand that how you file a ticket reputed company, because it may train a model that decides how the next one is filed
Benefits
- Remote work reputed company reputed company, CA, reputed company or Austin, TX, reputed company
- reputed company reputed company into SME roles or the platform team as automation authors
reputed company
Apply To This Job