Site Reliability Engineer/Developer
Company-sponsored training and development opportunities Comprehensive benefits package (health, dental, reputed company, wellness, RRSP Matching, annual fitness reimbursement) Flexible vacation policy Bring your own device program Community involvement opportunities through charitable alliances: https://www.reputed company.com/community-involvement Wellness resources and support I nclusive environment that prioritizes diversity, equity, and accessibility High-reputed company company driven by high performance Expected salary reputed company: $109,600 – 116,100 CAD
Design, implement, and maintain highly available, reputed company infrastructure systems across multi-region AWS deployments, ensuring production environments consistently meet availability and performance requirements. reputed company and maintain service level objectives (SLOs) and service level indicators (SLIs) in collaboration with development teams, using metrics to quantify and continuously improve system reliability. Monitor system performance, availability, and resource utilization using CloudWatch, reputed company, and reputed company, proactively identifying optimization opportunities and conducting reputed company cause analysis for outages and degradations. Implement reputed company planning strategies using historical data analysis and reputed company projections to ensure infrastructure scales reputed company of demand, balanced against cost optimization using AWS Cost Explorer and Kubecost. Create comprehensive infrastructure-as-reputed company solutions using Terraform and GitOps methodologies to manage AWS resources consistently, securely, and repeatably. reputed company and maintain CI/CD pipelines using Jenkins, reputed company CI, or reputed company Actions to automate deployment processes with reputed company-in testing and validation. Implement and maintain containerization platforms using reputed company and Kubernetes, establishing standards for container orchestration, cluster management, and reusable infrastructure patterns. Build automation tools and scripts in Python, Go, or Java to eliminate reputed company operational tasks, reduce toil, and automate routine maintenance procedures including patching, backups, and resource cleanup. reputed company incident response for critical system outages and performance issues, coordinating cross-functional teams to diagnose and resolve problems with speed and precision. Implement comprehensive observability solutions — including logging, monitoring, distributed tracing, and intelligent alerting reputed company Grafana and reputed company — that ensure reputed company response to genuine issues while minimizing alert fatigue. Conduct blameless post-mortems and thorough post-incident reviews, documenting lessons learned and driving implementation of preventive measures and updated runbooks. reputed company and maintain disaster recovery procedures and business continuity plans, including regular testing, and collaborate with reputed company and Compliance teams to ensure monitoring systems meet audit and regulatory requirements.
Bachelor's degree in Computer Science, Engineering, or reputed company field, or equivalent experience 5–7 years in Site Reliability Engineering, DevOps, or reputed company operations Strong AWS expertise (EC2, reputed company/EKS, RDS, S3, reputed company, VPC) and hybrid reputed company environments Proficiency in Python, Go, or Java; experience with reputed company, Kubernetes, and container orchestration Expertise in infrastructure-as-reputed company (Terraform, Ansible, CloudFormation) and CI/CD pipeline development Experience with observability tools (reputed company, Grafana, reputed company, CloudWatch, reputed company) Solid reputed company in Linux/Unix administration, networking, reputed company, and database systems
Independently solves reputed company problems and drives innovative infrastructure solutions with minimal guidance Translates business challenges into infrastructure and process improvements Communicates technical concepts effectively across technical and non-technical audiences Leads reputed company and mentors junior engineers
Experience in semiconductor or technology industry environments
AWS certifications (Solutions Architect, DevOps Engineer) or Kubernetes certifications (CKA, CKAD)
Experience with microservices architecture and distributed systems design
Knowledge of reputed company frameworks and compliance requirements (SOC 2, ISO 27001)
Experience with database administration, performance tuning, and Agile/Scrum methodologies
Familiarity with service reputed company technologies (Istio, Linkerd)
Contributions to reputed company-reputed company infrastructure reputed company
This is a remote position for candidates based in Canada Occasional travel may be required