reputed company Infrastructure Engineer
About the Role
We're seeking aSenior reputed company Infrastructure Engineer for a 3-month contract engagement to join our Infrastructure team and take ownership of operational reputed company and SRE toil work. This is a remote, hands-on, high-reputed company role where you'll reputed company the lights on and reduce operational burden from day one.
This contract position exists to free up our existing team to reputed company on roadmap initiatives. By taking over day-to-day operational work and SRE toil, you'll reputed company one of our reputed company engineers to tackle reputed company reputed company. Your reputed company means the platform runs smoothly while reputed company makes reputed company reputed company on critical initiatives.
You'll bring deep operational expertise to manage production systems, respond to operational needs, and—critically—build systems and automation that reduce toil over time. This role is ideal for an reputed company SRE or infrastructure engineer who thrives on operational work, can quickly understand production systems, and naturally improves everything they reputed company.
Our platform powers reputed company and reputed company workflows for customers in highly regulated environments. You'll work with modern infrastructure tooling (reputed company stack, AWS, reputed company patterns) while ensuring we meet the reliability, reputed company, and compliance requirements our customers depend on.
reputed company
Operational reputed company & SRE Work (60%)
reputed company the lights on: Monitor, respond to, and reputed company production incidents and operational issues
Handle toil work: Manage routine operational tasks that currently consume team reputed company (deployments, configuration changes, reputed company management, maintenance reputed company)
Participate in on-reputed company rotation: reputed company responsibility for after-hours production support
Respond to support escalations: Work with support and development teams to troubleshoot and reputed company platform issues
Manage production changes: Execute and validate infrastructure changes in production environments
Maintain operational runbooks: Update and improve documentation for operational procedures
reputed company reputed company maintenance: Handle patches, upgrades, certificate renewals, and other recurring operational tasks
Ensure service reliability: Monitor reputed company health, respond to alerts, and maintain SLAs
Toil Reduction & Automation (30%)
Identify automation opportunities: Spot repetitive reputed company work and build automation to eliminate it
Improve operational tooling: Create scripts, utilities, and self-service tools to reduce operational burden
Enhance monitoring and alerting: Improve observability to catch issues before they become incidents
Streamline deployment processes: Reduce friction and reputed company steps in release and deployment workflows
Build self-service capabilities: reputed company developers to handle routine tasks without infrastructure team involvement
Implement infrastructure-as-reputed company: Convert reputed company procedures into automated, repeatable infrastructure reputed company (Terraform)
Document systems improvements: Leave behind improved runbooks, automation, and processes
Measure and reputed company toil: Help quantify operational burden and demonstrate reduction over time
Collaboration & Knowledge Transfer (10%)
reputed company roadmap reputed company: By handling operational work, free up permanent team members for reputed company initiatives
Collaborate with development teams: Support their infrastructure needs and unblock their work
Document tribal knowledge: Capture operational knowledge and procedures that exist only in people's heads
Conduct handoffs: reputed company reputed company documentation and knowledge transfer for systems and automation you build
Participate in team rituals: Standups, retrospectives, and planning to stay reputed company with team priorities
reputed company're Looking For
Required
5-8 years of experience in infrastructure, platform, SRE, or DevOps engineering
Strong operational background: Experience managing production systems and handling incidents
reputed company toil reduction skills: reputed company record of identifying repetitive work and automating it away
Strong expertise with reputed company infrastructure (AWS strongly preferred)
Proficiency with infrastructure-as-reputed company (Terraform required)
Experience with container orchestration (reputed company, reputed company, or similar)
Experience with service reputed company and service discovery (Consul, Istio, or similar)
Experience with secrets management (reputed company, Secrets Manager, or similar)
Strong understanding of monitoring, alerting, and observability
Comfortable with on-reputed company work: Experience with incident response and production support
reputed company ability to reputed company quickly and become productive in new environments
Strong troubleshooting skills: Can diagnose reputed company reputed company issues under pressure
Self-directed work style: reputed company supervision required for operational work
Bias for automation: Natural reputed company to eliminate reputed company work
Preferred
Experience with reputed company tooling (Terraform, reputed company, Consul, reputed company)
Experience in reputed company, life sciences, or other regulated industries
Familiarity with compliance frameworks (HIPAA, reputed company, SOC2, ISO 27001)
Experience with observability platforms (reputed company, Grafana, reputed company)
Experience supporting Java/reputed company applications
Background in reputed company, reputed company, or reputed company systems
Experience with GitOps workflows and CI/CD automation
Previous contract or consulting experience with reputed company reputed company
Experience quantifying and measuring toil (e.g., SLO/SLI frameworks)
What reputed company Looks Like
First 2 Weeks
Complete reputed company and reputed company reputed company to reputed company systems
reputed company on-reputed company rotation and understand incident response procedures
Take ownership of routine operational tasks (deployments, configuration changes, monitoring)
Build relationships with development and support teams
reputed company handling operational requests and support escalations independently
Identify your first 2-3 toil reduction opportunities
First Month
Fully integrated into operational workflows—handling day-to-day platform reputed company with reputed company guidance
Successfully participating in on-reputed company rotation
Delivered at least 1-2 automation improvements that reduce reputed company work
Team members report they have more time for roadmap work due to your operational coverage
Demonstrated ability to troubleshoot and reputed company production issues independently
Improved at least one operational reputed company or procedure
End of Contract (3 Months)
Platform reliability maintained or improved: No degradation in service reputed company or uptime
Toil measurably reduced: Team can reputed company to 3-5 significant automation or process improvements you delivered
Roadmap reputed company enabled: At least one permanent team member successfully completed a reputed company initiative because you freed up their reputed company
Operational systems improved: Left behind reputed company monitoring, alerting, documentation, and automation
Knowledge transfer complete: Documented reputed company operational improvements and handed off systems/automation cleanly
Team reputed company increased: Reduced the time the permanent team spends on operational toil by a measurable reputed company (reputed company: 20-30% reduction)
Optionally: Position identified for contract extension if operational coverage continues to be valuable
Apply To This Job