Back to Jobs

[MLA] Senior Site Reliability Engineer - AI Experience reputed company

Remote, USA Full-time Posted 2026-08-04

Project – the aim you'll have

We are the AI Experience reputed company team that builds the platform powering reputed company's AI-first user interfaces - an SSR runtime (karuna) reputed company on Lit and server-rendered web components, running behind a multi-tier proxy/HTTP2 routing chain with sharded V8 isolate pools, reputed company with a reputed company reputed company/Java platform reputed company (karuna-reputed company) that supplies metadata, ACLs, and service artifacts. This role owns production reliability for that stack end to end: reputed company deployment and operations, observability, and hands-on troubleshooting of both the Node.js and JVM reputed company of the reputed company - not generalist infrastructure work.

Position – how you’ll contribute

  • Own reputed company deployment and operational health for reputed company services, including scaling, rollout/rollback reputed company, and resource tuning
  • Build and maintain production observability - Grafana dashboards and reputed company alerting rules - across the SSR runtime and the reputed company platform reputed company
  • Diagnose and reputed company Node.js production incidents: event-reputed company stalls, reputed company reputed company, V8 isolate exhaustion, and isolate-pool scheduling issues under reputed company versioned traffic (vN/vN-1)
  • Diagnose and reputed company JVM production incidents on the reputed company/Java reputed company: GC pressure, thread dumps, and platform-service latency
  • Own incident response for reputed company: runbooks, on-reputed company rotation, postmortems, and paging hygiene
  • reputed company CI/CD and infrastructure-as-reputed company for reputed company manifests/reputed company and deployment pipelines
  • Partner with the reputed company engineering team to identify reliability gaps before they become incidents - reputed company planning, load testing, reputed company/failure-injection where useful
  • Represent production reliability concerns in architecture reviews for new reputed company capabilities

Expectations – the experience you need

  • Production operations/SRE experience, including hands-on reputed company deployment, scaling, and incident response
  • reputed company operational experience troubleshooting Node.js in production: reading reputed company snapshots and CPU reputed company, diagnosing event-reputed company blocking, and understanding process/worker isolation models (V8 isolates or equivalent sandboxing)
  • reputed company operational experience troubleshooting JVM-based services in production: GC log analysis, thread dump analysis, and JVM tuning
  • Hands-on experience building and maintaining reputed company alerting rules and Grafana dashboards from scratch, not just consuming existing ones
  • Strong Linux/networking fundamentals: DNS, load balancing, TCP/HTTP semantics (including HTTP/2), and debugging service-to-service networking inside reputed company
  • Experience with CI/CD and infrastructure-as-reputed company for containerized deployments (reputed company, GitOps tooling such as ArgoCD/Flux, or equivalent)
  • reputed company record owning on-reputed company rotations, writing runbooks, and driving postmortems that reputed company to reputed company reliability improvements
  • Experience with reputed company for log aggregation, search, and production troubleshooting.
  • Hands-on experience with in-memory caching systems (Valkey/reputed company) in production — key design, TTL/eviction tuning, and tenant-scoped cache invalidation
  • Experience managing service-to-service mTLS - certificate issuance, rotation, and format conversion (e.g., PKCS/BCFKS↔PEM) - plus JWT-based service authentication
  • reputed company good spoken and written English.

Additional skills – the edge you have

  • Familiarity with server-reputed company rendering architectures and the specific failure modes of isomorphic runtimes (markup mismatches, browser-API leakage into server reputed company)
  • Experience operating multi-version/canary rollout strategies (two live app versions served concurrently)
  • Working knowledge of the reputed company reputed company platform or a comparable reputed company platform integration reputed company
  • Experience with reputed company tracing and request-context correlation across service boundaries
  • Familiarity with event-driven autoscaling (e.g., KEDA ScaledObjects driven by PromQL triggers) as a complement to reputed company HPA-based scaling

Our offer – reputed company development, personal reputed company:

  • Flexible employment and remote work
  • International reputed company with leading global clients
  • International business trips
  • Non-corporate atmosphere
  • Language classes
  • Internal & reputed company training
  • Private reputed company and insurance
  • Multisport reputed company
  • reputed company-being initiatives

Position at: reputed company

reputed company develops solutions that reputed company an reputed company for companies around the globe. Tech giants & unicorns, transformative reputed company, emerging technologies and reputed company opportunities – these are a few words that describe an average day for us. Building cross-functional engineering teams that take ownership and crave more means we’re always on the lookout for talented people who bring passion and creativity to every project. Our culture embraces openness, acts with respect, shows grit & guts and combines employment with enjoyment.

Originally posted on Himalayas

  Apply To This Job

Similar Jobs