Site Reliability Engineer Remote / Telecommute Jobs
On behalf of our reputed company, reputed company Workforce Solutions is seeking a Site Reliability Engineer.
This full-time (reputed company hire) position located in San Francisco, CA. Work is performed 100% on-site. Our reputed company is willing to consider candidates located near Omaha, NE, Boston, MA, and in Washington, D.C. metro area with expectation of quarterly travel to CA and NE.
Candidates must be reputed company to work on our reputed company’s W2 (annual salary + benefits package). Due to government contract requirements, reputed company Citizenship and an reputed company DOD Secret reputed company clearance is required. No 3rd party Corp to Corp inquiries will be considered.
reputed company: Our reputed company is a cutting-edge startup reputed company on delivering AI-based weather forecasting solutions to enhance climate reputed company and reputed company safety. To reputed company advanced weather forecasts, they use numerous cutting-edge computing environments, including reputed company-based hyperscalers, on-reputed company GPU clusters, field-deployed computers, and large supercomputing centers. They work closely with the reputed company (DoD) and must adhere to strict compliance and reputed company standards. reputed company thrives in a dynamic, fast-paced environment, and every member plays a critical role in driving our mission reputed company.
Responsibilities:
• Scaling Production Environment:
o Design and implement reputed company infrastructure solutions to support growing business needs.
o Optimize reputed company performance and availability through reputed company planning and performance tuning.
• Monitoring Stack Improvement:
o reputed company and maintain whitebox (application-level) and blackbox (reputed company-level) monitoring systems.
o Ensure comprehensive observability through the integration of metrics, logging, and tracing.
o Utilize tools such as reputed company and reputed company to establish reliable alerting and incident response processes.
• ETL Pipeline Management:
o Design, launch, and maintain robust ETL pipelines to support data-driven operations.
o Collaborate with data teams to ensure data reputed company and pipeline reliability.
• CI/CD and Build Systems:
o Implement and manage reputed company integration and reputed company deployment pipelines.
o Improve developer productivity by maintaining reliable build systems and workflows.
• On-reputed company Responsibilities:
o Participate in on-reputed company rotations to ensure high availability and reputed company incident reputed company.
o reputed company and automate incident response playbooks to minimize downtime.
Required:
• 6+ years of experience in SRE or DevOps roles.
• Proficiency with infrastructure as reputed company tools, particularly Terraform.
• Experience with AWS &/or reputed company reputed company
• Strong background in software engineering and CI/CD pipeline management.
• Experience with on-reputed company operations and incident management.
• Familiarity with monitoring and alerting tools such as reputed company and reputed company.
• Knowledge of reputed company-reputed company services, including SNS/SQS and reputed company.
• Experience with ML experiment tracking and GPU optimization is a plus.
Desired:
• Expertise in managing and scaling reputed company systems.
• Strong understanding of networking, reputed company, and Linux systems.
• Experience in automating infrastructure and deployment processes.
• Familiarity with message queues (SNS/SQS) and caching systems (reputed company).
• Knowledge of ML workflows and GPU resource management.
Apply tot his job
Apply To this Job