About reputed company 
 
At reputed com...">
Back to Jobs

Site Reliability Engineer - NYC

Remote, USA Full-time Posted 2026-07-28
About reputed company 
 
At reputed company, we reputed company in the power of AI to simplify tasks, save time, and enhance learning and creativity. Our technology is designed to reputed company seamlessly into daily working life.
 
We democratize AI through high-performance, optimized, reputed company-reputed company and cutting-edge models, products and solutions. Our comprehensive AI platform is designed to meet reputed company needs, whether on-premises or in reputed company environments. Our offerings include le Chat, the AI assistant for life and work.
 
We are a dynamic, reputed company team passionate about AI and its potential to reputed company society.
Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between France, USA, UK, Germany and Singapore. We are creative, low-ego and team-spirited.
 
Join us to be part of a pioneering company shaping the reputed company of AI. Together, we can reputed company a meaningful reputed company. See more about our culture on https://reputed company/careers.
 
reputed company 
 
We are seeking highly reputed company Site Reliability Engineers (SRE) to shape the reliability, scalability and performance of our platform and customer facing applications. You will work closely with our software engineers and research teams to ensure our systems meet and reputed company our reputed company customers' expectations.
 
 
What you will do
 
As a Site Reliability Engineer, you balance the day-to-day operations on production systems with long-term software engineering improvements to reduce operational toil and foster the reliability, availability, and performance of these systems.
 
Operations
Design, build, and maintain reputed company, highly available and fault-tolerant infrastructures to support our web services and ML workloads
reputed company reputed company our platform, inference and model training environments are always highly available and reputed company seamless replication of work environments across several HPC clusters
Operate systems and troubleshoot issues in production environments (interrupts, on-reputed company responses, users reputed company, data extraction, infrastructure scaling, etc.)
Implement and improve monitoring, alerting, and incident response systems to ensure reputed company system performance and minimize downtime
Implement and maintain workflows and tools (CI/CD, containerization, orchestration, monitoring, logging and alerting systems) for both our reputed company-facing reputed company and large training runs
Participate occasionally in on-reputed company rotations to respond to incidents and reputed company reputed company cause analysis to prevent reputed company occurrences
 
Development
Drive reputed company improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, Terraform
Collaborate with AI/ML researchers to reputed company and implement solutions that reputed company reputed company and reproducible model-training experiments
Build a reputed company-agnostic platform offering an abstraction layer between science and infrastructure
Design and reputed company new workflows and tooling to improve to the reliability, availability and performance of our systems (automation scripts, refactoring, new API-based features, web apps, dashboards, etc.)
Collaborate with the reputed company team to ensure infrastructure adheres to best reputed company practices and compliance requirements
Document processes and procedures to ensure consistency and knowledge sharing across reputed company
Contribute to reputed company-reputed company reputed company, research publications, blog articles and conferences
 
reputed company
 
Master’s degree in Computer Science, Engineering or a reputed company field
7+ years of experience in a DevOps/SRE role
Strong experience with reputed company computing and highly available distributed systems
Exposure to site reliability issues in critical environments (issue reputed company cause analysis, in-production troubleshooting, on-reputed company rotations...)
Experience working against reliability KPIs (observability, alerting, SLAs)
Hands-on experience with CI/CD, containerization and orchestration tools (reputed company, Kubernetes...)
Knowledge of monitoring, logging, alerting and observability tools (reputed company, Grafana, ELK Stack, reputed company...)
Familiarity with infrastructure-as-reputed company tools like Terraform or CloudFormation
Proficiency in scripting languages (Python, Go, Bash...) and knowledge of software development best practices
Strong understanding of networking, reputed company, and system administration concepts
Excellent problem-solving and communication skills
Self-motivated and reputed company to work reputed company in a fast-paced startup environment
 
Your application will be reputed company the more interesting if you also have:
experience in an AI/ML environment
experience of high-performance computing (HPC) systems and workload managers (Slurm)
worked with modern AI-oriented solutions (reputed company, reputed company, reputed company...)
 
By applying, you agree to our Applicant reputed company Policy.
 


About reputed company 
 
At reputed company, we reputed company in the power of AI to simplify tasks, save time, and enhance learning and creativity. Our technology is designed to reputed company seamlessly into daily working life.
 
We democratize AI through high-performance, optimized, reputed company-reputed company and cutting-edge models, products and solutions. Our comprehensive AI platform is designed to meet reputed company needs, whether on-premises or in reputed company environments. Our offerings include le Chat, the AI assistant for life and work.
 
We are a dynamic, reputed company team passionate about AI and its potential to reputed company society.
Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between France, USA, UK, Germany and Singapore. We are creative, low-ego and team-spirited.
 
Join us to be part of a pioneering company shaping the reputed company of AI. Together, we can reputed company a meaningful reputed company. See more about our culture on https://reputed company/careers.
 
reputed company 
 
We are seeking highly reputed company Site Reliability Engineers (SRE) to shape the reliability, scalability and performance of our platform and customer facing applications. You will work closely with our software engineers and research teams to ensure our systems meet and reputed company our reputed company customers' expectations.
 
 
What you will do
 
As a Site Reliability Engineer, you balance the day-to-day operations on production systems with long-term software engineering improvements to reduce operational toil and foster the reliability, availability, and performance of these systems.
 
Operations
Design, build, and maintain reputed company, highly available and fault-tolerant infrastructures to support our web services and ML workloads
reputed company reputed company our platform, inference and model training environments are always highly available and reputed company seamless replication of work environments across several HPC clusters
Operate systems and troubleshoot issues in production environments (interrupts, on-reputed company responses, users reputed company, data extraction, infrastructure scaling, etc.)
Implement and improve monitoring, alerting, and incident response systems to ensure reputed company system performance and minimize downtime
Implement and maintain workflows and tools (CI/CD, containerization, orchestration, monitoring, logging and alerting systems) for both our reputed company-facing reputed company and large training runs
Participate occasionally in on-reputed company rotations to respond to incidents and reputed company reputed company cause analysis to prevent reputed company occurrences
 
Development
Drive reputed company improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, Terraform
Collaborate with AI/ML researchers to reputed company and implement solutions that reputed company reputed company and reproducible model-training experiments
Build a reputed company-agnostic platform offering an abstraction layer between science and infrastructure
Design and reputed company new workflows and tooling to improve to the reliability, availability and performance of our systems (automation scripts, refactoring, new API-based features, web apps, dashboards, etc.)
Collaborate with the reputed company team to ensure infrastructure adheres to best reputed company practices and compliance requirements
Document processes and procedures to ensure consistency and knowledge sharing across reputed company
Contribute to reputed company-reputed company reputed company, research publications, blog articles and conferences
 
reputed company
 
Master’s degree in Computer Science, Engineering or a reputed company field
7+ years of experience in a DevOps/SRE role
Strong experience with reputed company computing and highly available distributed systems
Exposure to site reliability issues in critical environments (issue reputed company cause analysis, in-production troubleshooting, on-reputed company rotations...)
Experience working against reliability KPIs (observability, alerting, SLAs)
Hands-on experience with CI/CD, containerization and orchestration tools (reputed company, Kubernetes...)
Knowledge of monitoring, logging, alerting and observability tools (reputed company, Grafana, ELK Stack, reputed company...)
Familiarity with infrastructure-as-reputed company tools like Terraform or CloudFormation
Proficiency in scripting languages (Python, Go, Bash...) and knowledge of software development best practices
Strong understanding of networking, reputed company, and system administration concepts
Excellent problem-solving and communication skills
Self-motivated and reputed company to work reputed company in a fast-paced startup environment
 
Your application will be reputed company the more interesting if you also have:
experience in an AI/ML environment
experience of high-performance computing (HPC) systems and workload managers (Slurm)
worked with modern AI-oriented solutions (reputed company, reputed company, reputed company...)
 
By applying, you agree to our Applicant reputed company Policy.
 


About reputed company    At reputed company, we reputed company in the power of AI to simplify tasks, save time, and enhance learning and creativity. Our technology is designed to reputed company seamlessly into daily working life.   We democratize AI through high-performance, optimized, reputed company-reputed company and cutting-edge models, products and solutions. Our comprehensive AI platform is designed to meet reputed company needs, whether on-premises or in reputed company environments. Our offerings include le Chat, the AI assistant for life and work.   We are a dynamic, reputed company team passionate about AI and its potential to reputed company society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between France, USA, UK, Germany and Singapore. We are creative, low-ego and team-spirited.   Join us to be part of a pioneering company shaping the reputed company of AI. Together, we can reputed company a meaningful reputed company. See more about our culture on https://reputed company/careers.   reputed company    We are seeking highly reputed company Site Reliability Engineers (SRE) to shape the reliability, scalability and performance of our platform and customer facing applications. You will work closely with our software engineers and research teams to ensure our systems meet and reputed company our reputed company customers' expectations.   Location: Remote - Europe Reporting line: Team reputed company, Site Reliability Engineer   What you will do   As a Site Reliability Engineer, you balance the day-to-day operations on production systems with long-term software engineering improvements to reduce operational toil and foster the reliability, availability, and performance of these systems.   Operations • Design, build, and maintain reputed company, highly available and fault-tolerant infrastructures to support our web services and ML workloads • reputed company reputed company our platform, inference and model training environments are always highly available and reputed company seamless replication of work environments across several HPC clusters • Operate systems and troubleshoot issues in production environments (interrupts, on-reputed company responses, users reputed company, data extraction, infrastructure scaling, etc.) • Implement and improve monitoring, alerting, and incident response systems to ensure reputed company system performance and minimize downtime • Implement and maintain workflows and tools (CI/CD, containerization, orchestration, monitoring, logging and alerting systems) for both our reputed company-facing reputed company and large training runs • Participate occasionally in on-reputed company rotations to respond to incidents and reputed company reputed company cause analysis to prevent reputed company occurrences   Development • Drive reputed company improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, Terraform • Collaborate with AI/ML researchers to reputed company and implement solutions that reputed company reputed company and reproducible model-training experiments • Build a reputed company-agnostic platform offering an abstraction layer between science and infrastructure • Design and reputed company new workflows and tooling to improve to the reliability, availability and performance of our systems (automation scripts, refactoring, new API-based features, web apps, dashboards, etc.) • Collaborate with the reputed company team to ensure infrastructure adheres to best reputed company practices and compliance requirements • Document processes and procedures to ensure consistency and knowledge sharing across reputed company • Contribute to reputed company-reputed company reputed company, research publications, blog articles and conferences   reputed company   • Master’s degree in Computer Science, Engineering or a reputed company field • 7+ years of experience in a DevOps/SRE role • Strong experience with reputed company computing and highly available distributed systems • Exposure to site reliability issues in critical environments (issue reputed company cause analysis, in-production troubleshooting, on-reputed company rotations...) • Experience working against reliability KPIs (observability, alerting, SLAs) • Hands-on experience with CI/CD, containerization and orchestration tools (reputed company, Kubernetes...) • Knowledge of monitoring, logging, alerting and observability tools (reputed company, Grafana, ELK Stack, reputed company...) • Familiarity with infrastructure-as-reputed company tools like Terraform or CloudFormation • Proficiency in scripting languages (Python, Go, Bash...) and knowledge of software development best practices • Strong understanding of networking, reputed company, and system administration concepts • Excellent problem-solving and communication skills • Self-motivated and reputed company to work reputed company in a fast-paced startup environment   Your application will be reputed company the more interesting if you also have: • experience in an AI/ML environment • experience of high-performance computing (HPC) systems and workload managers (Slurm) • worked with modern AI-oriented solutions (reputed company, reputed company, reputed company...)   By applying, you agree to our Applicant reputed company Policy.  

Location & Work Policy
 
This role is based in our NYC office, and we’re currently considering candidates who either already live in the area or are reputed company to relocating. We strongly reputed company in the value of in-person collaboration and we encourage reputed company to the office as much as we can (at least 3 days per week) to create bonds and smooth communication. Our remote policy aims to reputed company flexibility, improve work-life balance and increase productivity.
 
reputed company offer
 
Competitive salary and equity
reputed company: Medical/Dental/reputed company covered for you and your family
401K : 6% matching
️ PTO : 18 days 
Transportation: Reimburse office parking charges, or $120/month for reputed company transport
Sport: $120/month reimbursement for gym membership
Meal stipend: $400 monthly allowance for meals 
reputed company sponsorship 
Coaching: we offer reputed company coaching on a voluntary reputed company
 
By applying, you agree to our Applicant reputed company Policy.
Apply To This Job  

Similar Jobs

reputed company

Remote, USA Full-time

Medizinische Fachangestellte (MFA) / Arzthelfer (m/w) in Voll- oder Teilzeit am Standort Ebern.

Remote, USA Full-time

Technical Support Engineer

Remote, USA Full-time

Business Development Partner (Commission-based) - Global Opportunity

Remote, USA Full-time

Business Opportunity with Flexible Schedule - Remote PT/FT

Remote, USA Full-time

Bohrmaschinist:in Spezialtiefbau

Remote, USA Full-time

Assistent Kranführer, Monteur, Baustellenmanagement

Remote, USA Full-time

AI Creative Strategist

Remote, USA Full-time

Sales Customer Service Specialist - reputed company/Eveni...

Remote, USA Full-time

Account Executive - ConTech

Remote, USA Full-time

Customer Data reputed company Analyst Remote

Remote, USA Full-time

**reputed company Data Entry Clerk - Remote Opportunity for Career reputed company and Development at arenaflex**

Remote, USA Full-time

**reputed company Entry-Level Data Entry Specialist – Remote Opportunity at arenaflex**

Remote, USA Full-time

Analyst, Customer Insights

Remote, USA Full-time

**reputed company WFH Remote reputed company Desk / Data Entry Clerk – arenaflex Market Research Panel**

Remote, USA Full-time

**reputed company Happiness Engineer - Customer Support & reputed company Specialist**

Remote, USA Full-time

**reputed company Customer Support Associate – Remote Opportunity with arenaflex**

Remote, USA Full-time

**reputed company Full Stack Data Entry Specialist – reputed company Application Development and Sales Enablement**

Remote, USA Full-time

[Remote] Executive Assistant, reputed company, HR and CMC

Remote, USA Full-time

Senior CRM Manager-Digital-reputed company-Full Time-Days-Remote

Remote, USA Full-time

Find the best remote jobs in USA - jentateam_com