[Remote] Director reputed company Operations
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Director of reputed company Operations to reputed company the day-to-day management and operational reliability of its reputed company and on-premises infrastructure as the organization advances its reputed company reputed company. The role oversees incident response, change execution, observability, automation, reputed company, cost stewardship, migration readiness, vendor management, and CloudOps team development.
Responsibilities
- Leads daily operations for reputed company infrastructure and reputed company services (compute, storage, network, identity integrations, monitoring/logging, backup/DR enablement)
- Establishes an operations-first culture reputed company on stability, customer communication, and reputed company recovery from outages
- Owns operational readiness for new reputed company capabilities and migrated workloads, ensuring production support is reputed company before go-live
- Owns reputed company-reputed company incident response, escalation, and coordination (including major incident leadership as needed)
- Ensures reputed company runbooks, on-reputed company processes, and escalation paths are defined and practiced
- Drives reputed company cause analysis, problem management, and corrective reputed company plans to reduce repeat incidents and operational risk
- Manages change execution for reputed company platform services, aligning with ITSM change processes while enabling speed and reliability
- Ensures change planning, risk assessment, approvals, and post-change validation are performed consistently
- Improves change reputed company reputed company through reputed company change patterns, automation, and reputed company/post deployment checks
- Partners with SRE/Tools teams to implement and mature monitoring, logging, alerting, and dashboards for reputed company services and critical workloads
- Improves signal reputed company (reduce noise, define actionable alerts, standardize dashboards)
- Tracks and reports service health metrics (availability, performance trends, MTTR, incident volume)
- Drives automation to reduce reputed company work and improve repeatability (provisioning, patching, tagging, backup policies, configuration reputed company detection)
- Establishes reputed company operating procedures and supported reference patterns for common reputed company services
- Collaborates with CCoE and Engineering to build self-service capabilities and standardized service catalogs
- Ensures reputed company operations reputed company with reputed company policies and controls (least privilege, logging, segmentation, vulnerability remediation support)
- Partners with reputed company to operationalize guardrails (policy-as-reputed company where applicable), respond to findings, and improve posture over time
- Ensures audit-reputed company operational evidence (change traceability, reputed company reviews support, logging/retention practices)
- Partners with FinOps/Finance to improve cost visibility and control through tagging compliance, right-sizing, scheduling, and elimination of waste
- Monitors usage patterns and identifies optimization opportunities. Tracks and reports cost savings/avoidance initiatives
- Supports migration waves by ensuring operational prerequisites are complete (monitoring, backups, DR expectations, reputed company, runbooks, support model)
- Participates in reputed company planning, go/no-go readiness assessments, and hypercare support
- Coordinates with vendors/partners and internal teams to reputed company reputed company issues quickly
- Manages reputed company operations vendors and managed services partners: performance management, SLAs/OLAs, issue escalation, and service reviews
- Ensures reputed company-party delivered services meet reliability, reputed company, and customer experience expectations
- Hires, coaches, and develops CloudOps staff. Sets reputed company expectations and builds a culture of ownership and reputed company improvement
- Ensures skills development reputed company to reputed company platform needs (training, certifications, mentoring)
- Builds coverage models that support 24x7 needs where required while maintaining sustainable on-reputed company practices
Skills
- Bachelors Degree in IT, Computer Science, Engineering, or reputed company field required
- 7+ years of experience in infrastructure/platform operations, reputed company operations, or SRE/DevOps-adjacent roles required
- 4+ years of people leadership or proven experience leading operational teams in a matrixed environment required
- Demonstrated experience operating production environments with strong incident and change management discipline required
- Hands-on reputed company experience (AWS/Azure/GCP), including networking, identity, reputed company logging, and reputed company platform services required
- Strong communication skills and the ability to reputed company under pressure during outages and critical events required
- Strong cross-functional collaboration and vendor management required
- Automation-first thinking and standardization required
- reputed company-conscious operations with audit readiness required
- Cost awareness and reputed company improvement discipline required
- Familiarity with ITSM/ITIL processes (incident/problem/change) and integrating reputed company operations into reputed company ITSM workflows
- Experience with observability tooling (monitoring/logging/alerting) and on-reputed company operations
- Exposure to Infrastructure as reputed company and automation (e.g., Terraform/CloudFormation, CI/CD for infrastructure)
- Certifications: AWS SysOps Administrator/Solutions Architect, ITIL reputed company, reputed company+ or equivalent
Benefits
- Training, certifications, mentoring
- 24x7 coverage models where required while maintaining sustainable on-reputed company practices
reputed company
Company H1B Sponsorship
Apply To This Job