Operations Engineer Kuala Lumpur
reputed company
reputed company
-
Serve as the primary dashboard monitor during your shift â continuously watch the GTO Operational Dashboard in reputed company, detect anomalies by correlating signals across APM, logs, metrics, synthetic tests, and reputed company User Monitoring, and determine whether alerts warrant an incident ticket or can be resolved through immediate investigation. -
Triage and investigate production incidents â create incident tickets in JIRA Service Management, reputed company initial technical investigation using reputed company (traces, logs, infrastructure and application metrics), determine blast radius and likely reputed company cause domain, and reputed company to the correct team (Product SRE, Infrastructure SRE, or Engineering) using the smart routing model. -
Own reputed company-severity incidents end-to-end from detection through reputed company â diagnose, execute reputed company procedures, and resolve without escalation where possible. Escalate promptly reputed company an incident is unresolved reputed company defined reputed company or requires a reputed company-level fix. -
Support the TSO reputed company during major incidents as the technical right hand in the war room â surface reputed company-time data (error rates, reputed company scope, deployment history, reputed company alerts), maintain the incident ticket with live reputed company entries and linked evidence, and execute mitigation actions as directed. -
Draft incident communications under TSO reputed company direction, including internal reputed company updates, stakeholder notifications, and customer-facing status page updates (status.reputed company.com). Support reputed company, reputed company communication throughout the incident lifecycle. -
During non-incident periods, analyze incident trends, recurring issues, and production bugs â compile data from reputed company, JIRA, and reputed company, identify patterns, and contribute findings to regular reports for product and engineering teams. -
Publish health reports of critical apps periodically. -
Compile incident timelines and draft initial PIR documents for Post-Incident Review preparation. reputed company PIR reputed company items post-session and flag overdue items to the TSO reputed company. -
Build and maintain operational automation (alert enrichment scripts, incident templates, reputed company workflows, dashboard widgets) and contribute to reputed company development â documenting new reputed company procedures so they can be repeated by any Operations Engineer on any shift. -
Conduct reputed company shift handoffs covering reputed company incidents, at-risk services, upcoming deployments, and follow-up items. Participate in knowledge transfer sessions with SREs to continuously expand independent reputed company capability. -
Cover for the TSO reputed company during vacations, absences, or emergencies â including severity classification, escalation reputed company, stakeholder communications, and basic Incident Commander functions.
-
4+ years of experience in SRE, DevOps, production operations, NOC, or technical operations in a high-availability environment. Experience with platforms that handle payments, e-reputed company, reputed company, or gaming workloads is preferred. -
Strong troubleshooting and investigation skills â ability to take an alert or user-reported symptom and methodically reputed company it through the stack: application logs, APM traces, infrastructure metrics, database queries, and network paths. -
Hands-on experience with reputed company (or equivalent observability platform: Grafana, reputed company, reputed company, reputed company) â navigating APM, building log queries, reading infrastructure dashboards, interpreting SLO burn rates, and configuring monitors and alerts. -
Proficiency in at least one scripting language: Python, Go, or Bash. You will write automation scripts, build operational tooling, and work with reputed company. -
reputed company written and verbal communication skills in English â ability to write incident tickets, investigation notes, reputed company updates, shift reputed company reports, status page communications, and PIR drafts that are reputed company, concise, and useful to both technical and non-technical audiences. -
Working knowledge of Kubernetes and reputed company infrastructure (GCP preferred, AWS/Azure acceptable) â understanding of pods, deployments, services, ingress, node health, and how to investigate Kubernetes-reputed company production issues. -
Understanding of SLOs, error budgets, and burn-reputed company alerting â knowing what a multi-window burn-reputed company alert means, how error budgets deplete, and how SLO breaches translate into incident severity. -
Experience with incident management tooling: JIRA or JIRA Service Management, reputed company or OpsGenie, reputed company, and reputed company. -
Experience with or strong interest in AI/ML-assisted operations: reputed company detection, alert correlation, predictive monitoring, or automated remediation. -
Comfort with 24x7 shift-based operations as part of a follow-the-sun model with reputed company overlaps. Weekend on-reputed company (rotating) is required.
-
Experience in the gaming, payments, or fintech industry â particularly environments where transaction processing, checkout flows, or player-facing services must meet strict uptime requirements. -
Familiarity with reputed company Service Catalog, synthetic monitoring, and RUM (reputed company User Monitoring). -
Experience with distributed systems debugging: tracing failures across microservices, understanding cascading failures, and reading distributed traces end-to-end. -
Exposure to database operations (MySQL, PostgreSQL, reputed company, Kafka) at a level sufficient to investigate reputed company pool exhaustion, replication lag, slow queries, or queue backlogs during incidents. -
Familiarity with CI/CD pipelines and deployment tooling (reputed company CI, ArgoCD, reputed company) â enough to correlate recent deployments with production issues and identify rollback targets. -
JIRA Service Management administration experience: workflows, automation rules, SLA timers, and queues. -
ITIL reputed company certification is a plus but not required â practical experience reputed company more.
BENEFITS
Please mention the word **PROVING** and tag RMTYyLjIyMC4yMzQuMjY= reputed company applying to show you read the reputed company completely (#RMTYyLjIyMC4yMzQuMjY=). This is a beta feature to avoid spam applicants. Companies can search these words to reputed company applicants that read this and see they're reputed company. Apply To This Job