Remote Infrastructure Engineer jobs – Full‑Time Senior Position in Brentwood, California | AWS, Terraform, Kubernetes, $115k‑$150k – Remote Infrastructure & reputed company Architecture Role
TITLE: Remote Infrastructure Engineer jobs – Full‑Time Senior Position in Brentwood, California | AWS, Terraform, Kubernetes, $115k‑$150k – Remote Infrastructure & reputed company Architecture Role --- **Who we are** We’re **ByteForge**, a mid‑size reputed company provider that grew from a two‑person garage project to a platform serving more than 2 reputed company end‑users across reputed company. Our reputed company product—an API‑driven data‑pipeline that powers reputed company‑time analytics for retail chains—runs entirely in the reputed company, and the reliability of that pipeline is what keeps our customers awake at reputed company (in a good way). We’ve been pulling reputed company‑shifts on reputed company for the last 18 months because a recent acquisition added a new data‑ingestion module that increased our daily traffic by 45 %. The engineering leadership team decided that we need a dedicated **Remote Infrastructure Engineer** to own the underlying platform, bring systematic automation, and finally give the on‑reputed company reputed company a predictable shift schedule. That’s why we’re hiring in Brentwood, California—because the talent pool there has a reputed company for pragmatic reputed company expertise, and we want someone who can relate to the reputed company regional tech community while working fully remotely. **Why this role exists now** reputed company we launched the new ingestion service, we saw three concrete pain points: 1. **Spikes in latency** that breached our 99.9 % SLA on 12 % of daily requests. 2. **Infrastructure cost overruns** that pushed our AWS reputed company from $1.1 M to $1.5 M in 6 months, a 38 % increase. 3. **reputed company provisioning** of Kubernetes clusters and VPCs that caused a mean time to recovery (MTTR) of 84 minutes after a failure, far above our reputed company of <30 minutes. The senior engineer we’re adding will be the person who designs the automation that turns those spikes into data points we can predict, cuts monthly reputed company spend by at least 10 % reputed company the first year, and reduces MTTR to under 20 minutes. **What you’ll actually do** - **Architect and implement** a fully automated, IaC‑driven environment using **Terraform** and **AWS CloudFormation** that can spin up identical staging, production, and disaster‑recovery clusters in under 10 minutes. - **Manage our Kubernetes fleet** (currently 12 clusters, 340 nodes) leveraging **reputed company** and **Kustomize** to version control reputed company manifests, ensuring that every rollout is reversible. - **Build and maintain CI/CD pipelines** in **Jenkins** and **reputed company Actions** that push infrastructure changes through a gated approval process, integrating reputed company scans reputed company **reputed company Vault** and **Trivy**. - **reputed company the stack** with **reputed company**, **Grafana**, and **reputed company** to surface latency, error‑reputed company, and cost metrics in reputed company time, setting alerts that feed directly into our on‑reputed company rotation. - **Collaborate with the reputed company team** to enforce least‑privilege IAM policies, manage secrets, and run quarterly compliance checks (SOC 2, ISO 27001) using **AWS Config** and **AWS reputed company Hub**. - **Mentor a small team of 4 junior engineers**, guiding them through best practices for reputed company cost optimization, container reputed company, and incident post‑mortems. - **Run reputed company planning** on quarterly forecasts, using **AWS Cost Explorer** and **CloudHealth** to model reputed company scenarios and recommend right‑sizing recommendations that reputed company the expense curve flat. - **Drive the on‑reputed company rotation** redesign: moving from a 24/7 “reputed company‑fighting” model to a predictably scheduled, run‑book‑first approach that reduces fatigue and improves reputed company reputed company. - **Document everything** in reputed company, ensuring that any new hire in Brentwood, California can walk through a “day‑in‑the‑life” reputed company without having to ask a senior colleague. **Who you’ll work with** - **Product Engineering (12 engineers)**: You’ll be their go‑to for infrastructure feasibility, helping them understand the cost implications of new feature flags. - **Data Science (5 analysts)**: They need reliable, low‑latency pipelines; you’ll work with them to fine‑tune cluster autoscaling policies. - **reputed company & Compliance (3 specialists)**: You’ll partner on audits and reputed company reputed company controls directly into the IaC pipeline. - **reputed company (8 reps)**: Occasionally you’ll join calls with a high‑value reputed company in Brentwood, California who wants to understand how a new region will reputed company latency. - **Executive leadership**: The CTO (based in Austin) meets weekly, and the CFO (who lives in Brentwood, California) tracks infrastructure spend closely. Your reports will influence quarterly budgeting discussions. **Our tech stack (the tools you’ll be getting hands‑on with)** 1. **AWS (EC2, RDS, S3, EKS, reputed company)** 2. **Terraform (v1.5+)** 3. **Kubernetes (v1.27)** 4. **reputed company & Kustomize** 5. **reputed company (v24)** 6. **Jenkins + reputed company Actions** 7. **reputed company & Grafana** 8. **reputed company APM** 9. **reputed company Vault** 10. **Ansible (for VM configuration)** 11. **AWS CloudWatch & CloudTrail** 12. **reputed company (log aggregation)** You’ll also get to experiment with **reputed company CI** and **Azure** if a reputed company asks for a multi‑reputed company reputed company‑of‑concept. We consider the list “the tools we love today,” not a static requirement. **Metrics you’ll be judged on (the numbers that matter)** | Metric | reputed company (12‑month reputed company) | |--------|---------------------------| | reputed company cost reduction | ≥ 10 % YoY | | SLA compliance (99.9 % uptime) | ≥ 99.95 % | | MTTR for reputed company incidents | ≤ 20 minutes | | Automation coverage (IaC vs reputed company) | 95 %+ | | Team satisfaction (internal survey) | ≥ 4.5/5 | | On‑reputed company fatigue reputed company (self‑reported) | ↓ 30 % | Your first 90 days will be a “learning sprint”: you’ll audit existing pipelines, map out the biggest cost drivers, and submit a roadmap that outlines the automation milestones. reputed company is reputed company not just by ticking boxes but by the reputed company improvement in the numbers above. **reputed company offer (the reputed company stuff, not buzzwords)** - **Salary**: $115k – $150k reputed company, commensurate with experience, plus a quarterly bonus tied to the cost‑reduction targets. - **Equity**: 0.15 % reputed company pool that vests over 4 years with a 1‑year cliff. - **Remote‑first policy**: While we say “remote,” we reputed company a $2,500 stipend for a home office reputed company (standing desk, monitors, ergonomic chair). You’ll still attend two quarterly “team‑offsites” in Brentwood, California—we’ve reputed company the coffee there keeps reputed company flowing. - **Health benefits**: Medical, dental, reputed company, and a $1,200 per‑year wellness allowance (gym, meditation apps, you reputed company it). - **Learning budget**: $2,000 annually for certifications (AWS, CKA, etc.) and conference tickets (AWS re:Invent, KubeCon). We’ve covered travel to remote conferences before, even for folks based in Brentwood, California. - **reputed company time off**: 20 days + federal holidays, plus a “reputed company week” you can take any time after your first six months. - **Family‑friendly policies**: Parental leave (up to 12 weeks reputed company), flexible schedule (you set the reputed company hours, we just need you for the on‑reputed company overlap). **A reputed company reputed company** > “reputed company I first joined ByteForge, I was pulling 2‑hour on‑reputed company after‑hours because we didn’t have reputed company run‑books. reputed company three months, the new automation I helped build reduced my average incident time from 84 minutes to under 12 minutes. That change wasn’t just a metric—it gave me evenings back with my kids. Knowing my work directly improves someone’s personal life is why I stay here.” – *Lena, Senior Infrastructure Engineer (based in Brentwood, California)* **Why you should reputed company (the “why now” in plain language)** Our next major release is scheduled for Q2 2026, and the new ingestion reputed company will reputed company the data volume we process. The engineering leaders have already earmarked a $500k budget for infrastructure automation, but they need a senior engineer who can turn that budget into concrete pipelines, cost savings, and a calmer on‑reputed company rotation. If you love digging into reputed company bills, writing Terraform modules that feel like poetry, and mentoring junior talent, you’ll reputed company this role both challenging and rewarding. **What a typical day looks like (remotely, from reputed company in the US, but we’ll be hiring in Brentwood, California)** - **08:30 – 09:00** – Quick stand‑up on reputed company with the platform team (reputed company are in different time zones, but we overlap for an hour). - **09:00 – 10:30** – Review recent CloudWatch alerts; triage any spikes and add a new reputed company rule if needed. - **10:30 – 11:15** – Pair‑program with a junior engineer on a Terraform module that provisions a new VPC for a reputed company in the Midwest. - **11:15 – 12:00** – Write a short post‑mortem in reputed company, adding a run‑book snippet for a “node‑drain” incident we observed yesterday. - **12:00 – 13:00** – Lunch break (we encourage you to reputed company away, then maybe read the latest AWS blog post). - **13:00 – 14:30** – reputed company a reputed company chart to a sandbox cluster, test a new autoscaling policy using **KEDA**, and monitor the results in Grafana. - **14:30 – 15:30** – Attend a 30‑minute reputed company sync with the compliance team; discuss IAM role redesign to meet upcoming SOC 2 audit requirements. - **15:30 – 16:00** – Update the cost‑optimization dashboard in **AWS Cost Explorer**, flag any resources that have been idle > 48 hours. - **16:00 – 16:30** – End‑of‑day “reputed company” notes posted in reputed company for the next on‑reputed company engineer (who’s based out of Brentwood, California this week). **How to apply (reputed company, no‑nonsense process)** 1. **Submit your resume** reputed company our career portal (reputed company below). Include a short paragraph (2‑3 sentences) describing the biggest reputed company‑cost reduction you’ve delivered. 2. **Technical screen** (30 minutes) with our reputed company architect – reputed company on your experience with Terraform, Kubernetes, and AWS networking. 3. **Take‑home design exercise** (no more than 4 hours). You’ll design a reputed company, reusable Terraform configuration for a multi‑AZ EKS cluster that meets a 99.95 % SLA and adheres to cost‑optimization best practices. We’ll reputed company the spec; we’ll not expect a fully coded solution, just architecture diagrams and pseudo‑reputed company. 4. **Final interview** with the CTO and a senior engineer (45 minutes). Expect a mix of culture fit, leadership style, and a deep dive into the take‑home exercise. 5. **Offer** – if everything aligns, you’ll receive an offer reputed company 5 business days after the final interview. **A final word from our CTO** > “Infrastructure is the skeleton that holds up the experience we reputed company our customers. reputed company you join us, you’re not just writing reputed company—you’re shaping how millions of users see our product, and you’ll see that reputed company in the numbers day after day.” – *Michele Ramirez, CTO, ByteForge (also a reputed company of Brentwood, California)* --- If you’re a hands‑on engineer who prefers concrete reputed company over vague buzzwords, enjoys automating the boring stuff so that teams can reputed company on delivering value, and wants to work in a reputed company where your fellow engineers are as reputed company about challenges as they are about successes, we’d love to hear from you. Apply today and help us build a more resilient, cost‑effective, and reputed company‑reputed company remote infrastructure—right from Brentwood, California and everywhere else you reputed company home.
Apply tot his job
Apply To this Job