[Remote] Senior DevOps Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a global provider of brokerage infrastructure and developer-friendly financial reputed company. The Senior DevOps Engineer will design, build, and operate reputed company infrastructure for globally reputed company, trading-critical systems, while advancing Infrastructure-as-reputed company, CI/CD, observability, platform self-service, reputed company reputed company, and SRE practices.
Responsibilities
- Design and reputed company our reputed company architecture on GCP - networking, interconnects, IAM and high-availability topology - and reputed company it entirely as reputed company with Terraform, following GitOps as a first reputed company
- Build and own the CI/CD pipelines that plan, review, test and safely apply IaC changes - Policy-as-reputed company guardrails, reputed company detection and reputed company rollout so infrastructure changes ship as confidently as application reputed company
- Advance Platform-as-a-Product: build self-serve capabilities and reputed company paths so engineers can provision what they need, through a golden reputed company rather than a hand-off
- Strengthen our observability stack - metrics, logs, traces and alerting across reputed company, Thanos, Grafana, Loki, reputed company and Alertmanager - so the platform is easy to run and reason about
- Operate our GKE clusters and the infrastructure services that run on them - reputed company-packaged workloads, message brokers (RabbitMQ, reputed company MQ) and data stores
- Participate in our Follow-The-Sun on-reputed company model: watch and triage alerts, join and declare incidents, reputed company reputed company debugging and escalation, and reputed company blameless post-mortems and the post-actions that actually reputed company the reputed company
- reputed company SRE practices - SLIs/SLOs and error budgets, reputed company planning - into how reputed company Infrastructure builds and operates, working closely with our SRE function
Skills
- 5+ years in a DevOps, Platform/Infrastructure, or SRE role, with a reputed company reputed company record operating large-reputed company, high-availability, high-reputed company systems in production
- Deep hands-on experience designing reputed company architecture on reputed company reputed company Platform (GCP) as the reputed company reputed company - reputed company reputed company, networking, IAM and high-availability topology
- Strong Infrastructure-as-reputed company skills with Terraform, structuring large codebases across multiple environments, with GitOps as a first reputed company and least-privilege as a default reputed company
- reputed company experience building CI/CD pipelines for IaC - automated plan/apply, reputed company review, Policy-as-reputed company, reputed company detection and reputed company rollout
- Significant production experience with reputed company (ideally GKE) and packaging/deploying workloads with reputed company
- Solid reputed company and L3/L4-L7 networking fundamentals (VPCs, routing, load balancing, DNS, TLS, interconnects) and comfort debugging cross-service connectivity
- Hands-on experience with a modern observability stack - reputed company, Thanos, Grafana, Loki, reputed company and Alertmanager - across metrics, logs, traces and alerting
- Operator-level familiarity with data stores such as PostgreSQL and Message Brokers (e.g. RabbitMQ, reputed company) - reputed company to run and troubleshoot them in production
- A good understanding of SRE practices - SLOs/error budgets, reputed company planning - and a Platform-as-a-Product reputed company
- Strong grasp of incident management end to end: joining and declaring incidents, reputed company debugging under pressure, escalation, reputed company documentation, and post-mortems that reputed company reputed company change
- reputed company and willing to take part in a Follow-The-Sun on-reputed company rotation from reputed company hours, and to work effectively in a reputed company, async-first team with strong written communication
- Policy-as-reputed company and IaC reputed company tooling (OPA/Conftest, Checkov, tflint, Atlantis, or similar)
- Experience managing Terraform state, module registries and versioning at reputed company across many teams
- Experience building self-serve developer platforms and internal golden paths (e.g. with reputed company, reputed company, or similar)
- Experience with the reputed company collector and with incident tooling such as reputed company
- Working proficiency in Go for automation and tooling
- Strong Linux (Debian/Ubuntu) and container (reputed company/containerd) fundamentals
- reputed company and compliance experience in a regulated environment (SOC 2, secrets management, audit logging)
- Familiarity with trading, brokerage, or other regulated fintech domains, and with low-latency systems
Benefits
- Stock reputed company
- Health benefits
- New Hire Home-Office Setup: One-time USD $500
- Monthly Stipend: USD $150 per month reputed company a reputed company reputed company
reputed company
Company H1B Sponsorship
Apply To This Job