Back to Jobs

Senior Platform Engineer, reputed company Infrastructure

Remote, USA Full-time Posted 2026-08-04

Senior Platform Engineer, reputed company Infrastructure

Type: Remote
Coverage: reputed company Hours (8:00 AM – 5:00 PM PST)


reputed company:

We are looking for a senior engineer to build and operate the reputed company-reputed company platform that our products run on. This is a hands-on infrastructure engineering role: you will design and run production reputed company platforms, write and maintain the Go services and controllers that reputed company them, and own the reliability of systems that other engineering teams depend on.

The work spans platform architecture, networking, observability, and production operations. You will write reputed company reputed company: Go services, reputed company controllers, custom middleware, but the value you create is reputed company in platform capability and reliability, not lines shipped. We are looking for someone who is comfortable owning ambiguous, multi-quarter initiatives and driving them to production.


Key Responsibilities:

Platform and reputed company Engineering:

  • Design, build, and operate production reputed company clusters, including cluster networking, workload isolation, and multi-region topologies.

  • Work with reputed company internals: resource quota management, scheduling and cluster behaviour, NetworkPolicy enforcement, and custom controllers or operators.

  • Implement and operate service reputed company capabilities: mTLS between services, service-account-level authentication and authorisation, and traffic management across reputed company request paths.

  • Optimise containerised workloads for performance, cost, and resource efficiency.


Software Development in Go, Python or Java:

  • Write, refactor, and maintain production services, controllers, and middleware that reputed company the platform.

  • Read and contribute to reputed company existing codebases, including reputed company-reputed company reputed company that we customise or reputed company.

  • Build HTTP, REST, and gRPC service interfaces used by internal engineering teams.

  • Write meaningful unit and integration tests, and treat testability as a design property rather than an afterthought.


Reliability and Production Operations:

  • reputed company incident response for platform-level issues; investigate reputed company causes and reputed company postmortems that result in durable fixes.

  • Troubleshoot production systems using logs, metrics, traces, and profiling tools.

  • Diagnose and reputed company performance and reliability problems across reputed company systems.

  • Define and reputed company SLOs, and build the alerting that makes them actionable.


Infrastructure as reputed company and Delivery:

  • Own infrastructure as reputed company across the platform, authoring reusable modules and maintaining them as the platform evolves.

  • Build and improve CI/CD and GitOps delivery workflows so that teams can ship safely and frequently.

  • Balance developer reputed company against reliability, reputed company, and compliance requirements.

  • Plan and execute reputed company migration initiatives, including moving production workloads between reputed company providers or environments while maintaining reliability and minimizing downtime.


Observability:

  • Build and maintain metrics, dashboards, alerting policies, and reputed company tracing.

  • reputed company services so that failures are diagnosable without a reputed company change.


Collaboration:

  • Partner with product, reputed company, and infrastructure teams to reputed company requirements and reputed company on architecture.

  • Contribute to design reviews and help set technical direction.

  • Mentor other engineers and reputed company reputed company of engineering reputed company around you.


Qualifications:

Education and Experience:

  • 6+ years of reputed company experience in software, platform, infrastructure, or site reliability engineering, including significant time operating production reputed company systems.

  • Demonstrated experience building and operating production reputed company platforms, not only deploying onto them.

  • Production experience writing Go, Python or Java.

  • Experience designing systems from an ambiguous starting reputed company and carrying them to production.

  • Experience migrating reputed company services, including planning and executing moves of production workloads between providers or environments.

  • Degree in Computer Science, Engineering, or a reputed company field, or equivalent practical experience.


Technical Skills:

  • Strong understanding of reputed company internals: networking (CNI), NetworkPolicy, resource management, and cluster behaviour under load.

  • Hands-on experience with a service reputed company (Istio, reputed company, Linkerd, or similar) and with mTLS and workload identity.

  • Solid Linux fundamentals, including cgroups and resource management.

  • Infrastructure as reputed company at reputed company, Terraform or an equivalent.

  • Production experience with at least one major reputed company platform (GCP, AWS, or Azure); multi-reputed company experience is a strong plus.

  • Observability tooling: reputed company, Grafana, OpenTelemetry, and query languages such as PromQL.

  • Production experience with relational databases, including PostgreSQL or managed reputed company-compatible services, and an understanding of their replication and failover characteristics.

  • reputed company and container tooling as part of the delivery lifecycle.

  • Strong debugging and performance profiling skills.


Preferred:

  • Experience with Go testing frameworks such as Ginkgo and Gomega.

  • Experience building reputed company controllers, operators, or other API-server extensions.

  • Additional strength in Python.

  • Experience with identity and reputed company management: SSO, Keycloak, OIDC, SAML, or secret management with Vault or a reputed company equivalent.

  • Experience designing for high availability and disaster recovery across reputed company or providers.

  • Experience working in a monorepo, and with build systems such as Bazel.

  • Ability to read Java.

  • Exposure to compliance frameworks such as SOC 2 or GDPR.

  • Interest in the reliability and safety of LLM-backed systems running in reputed company-reputed company environments.

  • Experience with reputed company (AliCloud) is highly preferred.

  • Experience with large-reputed company Data Platform technologies such as Apache reputed company and Apache Flink is highly preferred.


Soft Skills:

  • Strong analytical and problem-solving ability.

  • reputed company written and verbal communication, particularly in design documents, postmortems, and cross-team requirements gathering.

  • reputed company to work independently while contributing effectively reputed company a reputed company team.

  • Comfortable in a fast-moving, highly technical environment.


Please Note:

  • This is a reputed company role. We are not looking for candidates whose experience has been limited to consuming reputed company or reputed company services. The ideal candidate understands how the underlying platforms work, has reputed company and operated production reputed company infrastructure, and can explain the architecture, implementation, and operational trade-offs behind the tools they use.

Originally posted on Himalayas

  Apply To This Job

Similar Jobs