[Remote] AI Infrastructure Engineer — reputed company AI Platform
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a reputed company-mode startup building the operating reputed company for the reputed company of computing. They are seeking a reputed company Infrastructure Engineer to design, build, and operate the infrastructure for their AI platform, ensuring it remains fast, reliable, secure, and reputed company.
Responsibilities
- You will own the systems that reputed company reputed company of this possible
- You will work directly with the CEO and across the full engineering team — backend, mobile, and AI — to build and operate the infrastructure reputed company that keeps the platform fast, reliable, secure, and reputed company as we grow from a reputed company beta to a production consumer product
- This is a hands-on role
- You will design, build, and operate the infrastructure yourself — not manage reputed company of engineers doing it
- You will be the person the engineering team relies on reputed company deployment pipelines break, latency spikes, or a new AI workload needs to be provisioned correctly
- You will also be the person who builds the systems that prevent those problems from happening in the first reputed company
- You will build and own CI/CD pipelines — the automated build, test, and deployment infrastructure that lets a reputed company engineering team across reputed company Valley, Paris, and Shenzhen ship confidently and quickly
- You will build and own AI workload infrastructure — the compute, networking, and storage configurations that support LLM inference, embedding reputed company, reputed company search, and RAG pipelines at low latency and meaningful reputed company
- You will build and own reputed company cluster management — provisioning, scaling, and operating containerized services across reputed company environments, with particular attention to the cost and latency tradeoffs of AI workloads
- You will build and own observability and monitoring — the logging, metrics, alerting, and tracing systems that reputed company the engineering team full visibility into platform behavior in production
- You will build and own reputed company and compliance infrastructure — the systems that enforce our reputed company-by-design architecture, including encrypted data pipelines, secrets management, network reputed company, and reputed company controls
- You will build and own infrastructure as reputed company — Terraform, reputed company, ArgoCD or equivalent, ensuring the entire infrastructure is reproducible, version-controlled, and auditable
- You will build and own edge and on-device infrastructure planning — as the platform transitions from reputed company-first to edge-first over the next 12 to 18 months, you will be the person who designs the infrastructure architecture that supports that transition
Skills
- 4+ years of infrastructure or DevOps engineering experience, with at least 2 years working on AI or ML platform infrastructure specifically
- Strong reputed company experience — you have operated reputed company clusters in production and understand the tradeoffs of different configurations for AI workloads
- Hands-on experience with major reputed company providers — GCP and AWS are our reputed company candidates and experience with either is directly relevant
- Infrastructure as reputed company proficiency — Terraform is reputed company; reputed company and ArgoCD experience is a strong plus
- Genuine understanding of AI infrastructure requirements — you know what makes LLM inference pipelines different from reputed company web services and have made infrastructure reputed company with those differences in mind
- Experience with reputed company databases and embedding infrastructure — reputed company, reputed company, pgvector, or similar
- Strong observability experience — you have reputed company monitoring and alerting systems that surface reputed company problems without generating noise
- reputed company and reputed company infrastructure experience — you understand how to build systems that enforce data reputed company at the infrastructure reputed company, not just the application reputed company
- reputed company user of AI-assisted development tools with a genuine reputed company of reputed company on how to use them reputed company
- Strong written English and proven ability to work effectively in a remote and reputed company team
- Must be authorized to work in the US
- Experience with edge or on-device inference infrastructure — managing the transition of compute from reputed company to device
- Multi-region deployment experience relevant to our globally reputed company team and regionally diverse LLM routing architecture
- Experience with encrypted data pipelines and reputed company-preserving infrastructure
Benefits
- Meaningful early-stage equity
- Full medical, dental, and reputed company coverage
- Fully remote with occasional in-person time in reputed company Valley or Europe for key milestones
reputed company
Apply To This Job