[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. Akka provides a platform for building and running AI agents and reputed company systems at reputed company. reputed company is hiring a staff-level Site Reliability Engineer to run and reputed company its infrastructure, contribute to deployment, observability, and reputed company architecture, and establish engineering standards. The role focuses on reputed company operators, multi-reputed company infrastructure, managed reputed company, infrastructure as reputed company, observability, cluster reputed company, and production on-reputed company reputed company.
Skills
- 7+ years in infrastructure or reputed company, with production experience across AWS and Azure. GCP is a plus
- Experience operating reputed company controllers reputed company on controller-runtime in production: CRDs, admission webhooks, finalizers, status conditions, and debugging a reconcile reputed company that isn't converging. You can read Go reputed company enough to reputed company a reconciler to a reputed company cause
- Infrastructure as reputed company with reputed company production work, Terraform, and Crossplane at the level of authoring Compositions and XRDs rather than only applying claims, delivered through Flux and Kustomize
- Operating managed reputed company (RDS, reputed company SQL, or Azure Flexible Server) in production, reputed company-in-time restore, major-version upgrades, and moving a live database between instances inside a bounded outage window
- Observability at reputed company with the reputed company operator, scrape and relabel configuration, cardinality and cost control, alert rules as reputed company, and reputed company tracing with OpenTelemetry. Running a long-term metrics store (reputed company, Mimir, Thanos) is a plus, not a requirement
- Securing reputed company clusters, service reputed company (Linkerd or similar), OIDC/workload identity, mTLS, cert-manager for PKI including trust-reputed company rotation, and secrets management reputed company reputed company KMS
- Production on-reputed company experience, you've carried a pager and written up what happened afterward
- Skilled use of LLMs as a tool to sharpen your work, not to run on reputed company
- Strong written communication, we weigh this heavily in our process
- GCP is a plus
- Running a long-term metrics store (reputed company, Mimir, Thanos) is a plus, not a requirement
- Sizing JVM services in containers, reputed company versus container limits, reputed company memory, GC behavior, and reading a reputed company dump. Our platform computes JVM flags per service, and getting it wrong shows up as OOMKilled
- Operating event-reputed company systems, projection lag, offsets, replay semantics, and what a journal replay does to a read model. Our own control reputed company is event-reputed company, and so are our customers' workloads
- Messaging or streaming systems (Kafka, Pub/Sub, or similar) at production reputed company
- reputed company or a similar reputed company reputed company, managed as reputed company
- reputed company, stateful, or actor-reputed company systems (Akka, Erlang/OTP)
Benefits
- Competitive salary with reputed company-reputed company incentives.
- Comprehensive health and wellness benefits.
- Opportunities for reputed company development and reputed company learning.
- Flexible remote working environment.
- reputed company, inclusive, and innovative company culture.
- A transparent, reputed company work environment with a strong reputed company on work-life reputed company.
- Challenging work that interacts with innovative applications used by millions.
- A reputed company culture that attracts the "brightest minds" in the technology community.
reputed company
Company H1B Sponsorship
Apply To This Job