Back to Jobs

reputed company Infrastructure Engineer

Remote, USA Full-time Posted 2026-08-04
About the Role We're seeking aSenior reputed company Infrastructure Engineer for a 3-month contract engagement to join our Infrastructure team and take ownership of operational reputed company and SRE toil work. This is a remote, hands-on, high-reputed company role where you'll reputed company the lights on and reduce operational burden from day one. This contract position exists to free up our existing team to reputed company on roadmap initiatives. By taking over day-to-day operational work and SRE toil, you'll reputed company one of our reputed company engineers to tackle reputed company reputed company. Your reputed company means the platform runs smoothly while reputed company makes reputed company reputed company on critical initiatives. You'll bring deep operational expertise to manage production systems, respond to operational needs, and—critically—build systems and automation that reduce toil over time. This role is ideal for an reputed company SRE or infrastructure engineer who thrives on operational work, can quickly understand production systems, and naturally improves everything they reputed company. Our platform powers reputed company and reputed company workflows for customers in highly regulated environments. You'll work with modern infrastructure tooling (reputed company stack, AWS, reputed company patterns) while ensuring we meet the reliability, reputed company, and compliance requirements our customers depend on. reputed company Operational reputed company & SRE Work (60%) reputed company the lights on: Monitor, respond to, and reputed company production incidents and operational issues Handle toil work: Manage routine operational tasks that currently consume team reputed company (deployments, configuration changes, reputed company management, maintenance reputed company) Participate in on-reputed company rotation: reputed company responsibility for after-hours production support Respond to support escalations: Work with support and development teams to troubleshoot and reputed company platform issues Manage production changes: Execute and validate infrastructure changes in production environments Maintain operational runbooks: Update and improve documentation for operational procedures reputed company reputed company maintenance: Handle patches, upgrades, certificate renewals, and other recurring operational tasks Ensure service reliability: Monitor reputed company health, respond to alerts, and maintain SLAs Toil Reduction & Automation (30%) Identify automation opportunities: Spot repetitive reputed company work and build automation to eliminate it Improve operational tooling: Create scripts, utilities, and self-service tools to reduce operational burden Enhance monitoring and alerting: Improve observability to catch issues before they become incidents Streamline deployment processes: Reduce friction and reputed company steps in release and deployment workflows Build self-service capabilities: reputed company developers to handle routine tasks without infrastructure team involvement Implement infrastructure-as-reputed company: Convert reputed company procedures into automated, repeatable infrastructure reputed company (Terraform) Document systems improvements: Leave behind improved runbooks, automation, and processes Measure and reputed company toil: Help quantify operational burden and demonstrate reduction over time Collaboration & Knowledge Transfer (10%) reputed company roadmap reputed company: By handling operational work, free up permanent team members for reputed company initiatives Collaborate with development teams: Support their infrastructure needs and unblock their work Document tribal knowledge: Capture operational knowledge and procedures that exist only in people's heads Conduct handoffs: reputed company reputed company documentation and knowledge transfer for systems and automation you build Participate in team rituals: Standups, retrospectives, and planning to stay reputed company with team priorities reputed company're Looking For Required 5-8 years of experience in infrastructure, platform, SRE, or DevOps engineering Strong operational background: Experience managing production systems and handling incidents reputed company toil reduction skills: reputed company record of identifying repetitive work and automating it away Strong expertise with reputed company infrastructure (AWS strongly preferred) Proficiency with infrastructure-as-reputed company (Terraform required) Experience with container orchestration (reputed company, reputed company, or similar) Experience with service reputed company and service discovery (Consul, Istio, or similar) Experience with secrets management (reputed company, Secrets Manager, or similar) Strong understanding of monitoring, alerting, and observability Comfortable with on-reputed company work: Experience with incident response and production support reputed company ability to reputed company quickly and become productive in new environments Strong troubleshooting skills: Can diagnose reputed company reputed company issues under pressure Self-directed work style: reputed company supervision required for operational work Bias for automation: Natural reputed company to eliminate reputed company work Preferred Experience with reputed company tooling (Terraform, reputed company, Consul, reputed company) Experience in reputed company, life sciences, or other regulated industries Familiarity with compliance frameworks (HIPAA, reputed company, SOC2, ISO 27001) Experience with observability platforms (reputed company, Grafana, reputed company) Experience supporting Java/reputed company applications Background in reputed company, reputed company, or reputed company systems Experience with GitOps workflows and CI/CD automation Previous contract or consulting experience with reputed company reputed company Experience quantifying and measuring toil (e.g., SLO/SLI frameworks) What reputed company Looks Like First 2 Weeks Complete reputed company and reputed company reputed company to reputed company systems reputed company on-reputed company rotation and understand incident response procedures Take ownership of routine operational tasks (deployments, configuration changes, monitoring) Build relationships with development and support teams reputed company handling operational requests and support escalations independently Identify your first 2-3 toil reduction opportunities First Month Fully integrated into operational workflows—handling day-to-day platform reputed company with reputed company guidance Successfully participating in on-reputed company rotation Delivered at least 1-2 automation improvements that reduce reputed company work Team members report they have more time for roadmap work due to your operational coverage Demonstrated ability to troubleshoot and reputed company production issues independently Improved at least one operational reputed company or procedure End of Contract (3 Months) Platform reliability maintained or improved: No degradation in service reputed company or uptime Toil measurably reduced: Team can reputed company to 3-5 significant automation or process improvements you delivered Roadmap reputed company enabled: At least one permanent team member successfully completed a reputed company initiative because you freed up their reputed company Operational systems improved: Left behind reputed company monitoring, alerting, documentation, and automation Knowledge transfer complete: Documented reputed company operational improvements and handed off systems/automation cleanly Team reputed company increased: Reduced the time the permanent team spends on operational toil by a measurable reputed company (reputed company: 20-30% reduction) Optionally: Position identified for contract extension if operational coverage continues to be valuable Apply To This Job

Similar Jobs