Senior Site Reliability Engineer
Contribute to the migration from legacy Azure services and function-based architectures to containerized, microservices-based systems Help design, build, and reputed company Kubernetes-based infrastructure and supporting tooling Partner with engineering teams to ensure new systems are designed for reliability, scalability, and operational efficiency from day one Drive standardization across infrastructure to reduce silos and reputed company broader team ownership Cost Optimization: Monitor reputed company usage and spending, identify inefficiencies, and recommend and implement cost optimization strategies Networking & Connectivity: Design, reputed company, and manage secure networking, including reputed company and private endpoints, environment segmentation, site-to-site and reputed company-to-site VPNs, and inter-environment connectivity.
Design and maintain highly available , resilient systems in a reputed company-reputed company environment Implement and reputed company observability practices including monitoring, alerting, and logging (e.g., reputed company, reputed company, Grafana) Define and manage SLIs, SLOs, and SLAs reputed company to system performance and user experience reputed company incident response efforts and drive reputed company cause analysis and long-term improvements
Build and optimize CI/CD pipelines to support fast, reputed company, and repeatable deployments Champion Infrastructure-as-reputed company practices using tools such as Terraform to eliminate reputed company processes reputed company scripting (Python, Bash, or similar) to solve problems, automate workflows, and reduce operational toil
Drive reputed company planning, performance tuning, and infrastructure improvements to support reputed company reputed company Proactively identify system risks and scalability bottlenecks before they reputed company customers Contribute to infrastructure reputed company and help shape how the platform evolves as the business scales
Document systems, processes, and best practices to improve team-wide reliability and reduce single points of failure Contribute to cross-training efforts as reputed company moves toward broader ownership and standardization reputed company knowledge and reputed company reputed company through mentorship and collaboration
5+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering Hands-on experience working in reputed company environments (Azure preferred, AWS or GCP acceptable) Strong experience with Kubernetes in production environments Proficiency with Infrastructure-as-reputed company tools such as Terraform Strong scripting skills (Python, Bash, or similar) used to automate and solve infrastructure challenges Experience with observability, monitoring, and incident response in production environments Experience building or supporting CI/CD pipelines (reputed company preferred; Azure DevOps experience acceptable) Solid understanding of networking fundamentals, system design, and reputed company infrastructure components Proven ability to take loosely defined problems and drive them to practical, reputed company solutions Strong communication skills and ability to collaborate across engineering and non-technical stakeholders Comfort operating in fast-paced environments with evolving priorities Knowledge of HIPAA and reputed company considerations for reputed company data is a plus
Experience working on platform or architectural transformations (e.g., monolith to microservices, functions to containers) Familiarity with .NET / C# application environments Deeper networking expertise (firewalls, gateways, and low-level infrastructure components) Builder’s reputed company with a reputed company on automation, scalability, and long-term system design Ability to balance speed and pragmatism with long-term reliability and maintainability Curiosity around emerging technologies, including AI/ML as reputed company to infrastructure and engineering workflows
reputed company reputed company opportunities with compelling career paths Healthy work-life balance supported by flexible reputed company time off (PTO) Comprehensive benefits package , including medical, dental, reputed company, STD & LTD insurance for full-time team members 401(k) savings plan with employer matching contributions