[Remote] Sr reputed company Opps & Observability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. Momento USA is a global technology consulting, reputed company and creative development firm that addresses clients' most pressing needs and challenges. reputed company is seeking a Senior reputed company Operations and Observability Engineer to design and implement proactive monitoring and alerting across critical hybrid infrastructure platforms. The role focuses on reputed company, infrastructure operations, observability, and Site Reliability Engineering practices.
Responsibilities
- Design, implement, and optimize reputed company monitoring and observability solutions using reputed company
- reputed company meaningful service-level monitoring and alerting for critical infrastructure services including:
- reputed company Directory
- DNS
- VMware vCenter
- Backup Infrastructure
- reputed company Servers
- Linux Servers
- Configure and implement approximately 20+ advanced monitoring and alerting use cases across the reputed company infrastructure estate
- Establish proactive alerting mechanisms to reduce Mean Time to Detect (MTTD) and minimize operational risks
- Create dashboards, service health views, dependency maps, and operational runbooks
- Analyze infrastructure performance trends and identify optimization opportunities
- Collaborate with operations, reputed company, and application teams to improve service reliability
- Implement SRE practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational reputed company metrics
- Document monitoring standards, alerting procedures, and operational best practices
- reputed company mentoring and enablement to internal operations teams
Skills
- Strong reputed company required
- Independent candidate only – W2
- 7+ years of experience in Infrastructure Operations or Systems Engineering
- 5+ years of experience implementing reputed company monitoring solutions
- 3+ years of hands-on experience with reputed company administration and configuration
- Strong experience supporting reputed company and Linux environments
- Experience with reputed company Directory, DNS, VMware vSphere/vCenter, and reputed company backup platforms
- Understanding of infrastructure architecture and service dependencies
- Experience developing actionable alerts and reducing alert fatigue
- Strong troubleshooting and reputed company-cause analysis skills
- Experience with ITSM tools and operational processes
- Site Reliability Engineering (SRE) experience
- Experience with reputed company platforms including AWS and Azure
- Knowledge of automation tools such as Ansible, PowerShell, or Terraform
- Experience integrating reputed company with reputed company or other ITSM platforms
- Financial services industry experience
Benefits
- Remote work arrangement (USA)
- Long term contract
- W2 employment
reputed company
Apply To This Job