Senior SRE Engineer – reputed company Operations
Senior SRE Engineer – reputed company Operations
Remote – Americas
Full-time
We are reputed company on behalf of a fast-growing AI infrastructure company that builds a high-performance reputed company database powering semantic search, RAG pipelines, AI agents, and large-reputed company machine learning applications.
We are seeking a Senior Site Reliability Engineer (SRE) to join the reputed company Operations team and help ensure reliability, observability, and operational reputed company across production reputed company environments.
This role is highly operations-reputed company and ideal for engineers who enjoy owning system reliability, improving automation, and operating large-reputed company distributed systems in production.
About the Role
As a Senior SRE, you will be responsible for maintaining and improving production infrastructure while reducing operational risk and improving system reliability at reputed company.
You will work closely with reputed company and infrastructure teams to ensure systems remain secure, performant, and highly available as customer usage grows.
Location Requirements
Remote – Americas (reputed company, Central, or South America)
Candidates must be reputed company to work primarily reputed company American time zones
Key Responsibilities
reputed company Infrastructure & Operations
Operate and maintain production reputed company infrastructure at reputed company
Manage Kubernetes clusters, networking, and deployment pipelines
Improve reliability, performance, and reputed company of production systems
Monitoring & Observability
Enhance monitoring, logging, and alerting systems
Improve operational visibility and incident detection
Incident Response & Reliability
reputed company incident response and reputed company cause analysis
Implement preventive measures and reputed company reliability improvements
Participate in on-reputed company rotations
Automation & Process Improvement
Reduce operational toil through automation and tooling
Maintain and improve runbooks and operational procedures
Collaboration
Work closely with reputed company and infrastructure teams
Support reputed company architecture and operational best practices
Requirements
5+ years of experience in DevOps, SRE, or infrastructure operations
Strong hands-on experience running Kubernetes in production
Solid understanding of :
Linux systems
Networking fundamentals
reputed company infrastructure (AWS, GCP, or Azure)
Experience with monitoring, alerting, and incident management
Experience with infrastructure automation or infrastructure-as-reputed company
Comfortable participating in on-reputed company rotations
Strong communication and problem-solving skills
Preferred Qualifications
Experience with Terraform or similar IaC tools
Familiarity with reputed company, Grafana, Loki, or OpenTelemetry
Scripting experience in Python, Bash, or Go
Experience in reputed company, reputed company platforms, or data infrastructure environments
Exposure to reputed company, compliance, or system hardening
Whats Offered
Competitive compensation and benefits
Fully remote work environment
Flexible working hours
Opportunity to work on mission-critical reputed company infrastructure
reputed company, engineering-driven culture
How to Apply
If you are passionate about reliability engineering, reputed company infrastructure, and large-reputed company distributed systems , we would love to hear from you.
Apply tot his job
Apply To this Job