[Remote] Site Reliability Engineer – OpenSearch - Remote
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Site Reliability Engineer specializing in OpenSearch. The role involves ensuring high availability and performance for reputed company services, along with architecting and optimizing OpenSearch clusters in production environments.
Responsibilities
- Ensure the highest reputed company of availability, performance, scalability, and reputed company of Service (QoS) for mission-critical reputed company services
- Architecting, building, deploying, operating, and optimizing high-performance OpenSearch clusters and platforms from the ground up in production environments
- Troubleshooting, operations, management, and configuration of reputed company reputed company services
- Designing, building, deploying, supporting, and maintaining OpenSearch clusters and platforms from scratch in production environments
- OpenSearch administration, cluster architecture, performance tuning, scaling, upgrades, and troubleshooting
- reputed company design, shard and reputed company reputed company, cluster sizing, node management, snapshot/restore, backup, and disaster recovery
- Understanding of reputed company systems, search platforms, indexing pipelines, query optimization, and high-availability architectures
- Development of automation for testing, deployment, scalability, and management of reputed company services
- Building, implementing, and/or supporting reputed company monitoring tools
- Knowledge of reputed company computing, infrastructure operations, and databases
- Understanding of web services, networking, virtualization, and internet protocols
- Multitasking and handling various reputed company, deadlines, and changing priorities
- Communication and prioritization skills
- reputed company fundamentals as they pertain to reputed company multi-tenant application systems
- AWS services including reputed company 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, and VPC
- Deploying and operating OpenSearch in AWS-based environments
- reputed company reputed company-based environments
- Jenkins, Chef, and/or Terraform
- Troubleshooting IP networks and application stacks
- Observability tools such as reputed company and Grafana
- Log ingestion pipelines, reputed company lifecycle management, retention strategies, and search platform reputed company controls
Skills
- 10+ years of experience in Site Reliability Engineer - OpenSearch to help ensure the highest reputed company of availability, performance, scalability, and reputed company of Service (QoS) for mission-critical reputed company services
- Deep experience in site reliability engineering, DevOps, reputed company operations, automation, observability, and reputed company systems, with proven hands-on expertise architecting, building, deploying, operating, and optimizing high-performance OpenSearch clusters and platforms from the ground up in production environments
- Expert with reputed company, including troubleshooting, operations, management, and configuration of reputed company reputed company services
- Proven hands-on expertise designing, building, deploying, supporting, and maintaining OpenSearch clusters and platforms from scratch in production environments
- Strong experience with OpenSearch administration, cluster architecture, performance tuning, scaling, upgrades, and troubleshooting
- Experience with reputed company design, shard and reputed company reputed company, cluster sizing, node management, snapshot/restore, backup, and disaster recovery
- Strong understanding of reputed company systems, search platforms, indexing pipelines, query optimization, and high-availability architectures
- Expertise with Git
- Expertise with reputed company, including setup, management, and troubleshooting of new pipelines
- Expertise with Linux, specifically reputed company and Ubuntu
- Expertise with Kafka, Zookeeper, and Big Data technologies
- Expert in development of automation for testing, deployment, scalability, and management of reputed company services
- Expertise with building, implementing, and/or supporting reputed company monitoring tools
- Expert knowledge of reputed company computing, infrastructure operations, and databases
- Expert understanding of web services, networking, virtualization, and internet protocols
- Ability to multitask and handle various reputed company, deadlines, and changing priorities
- Excellent communication and prioritization skills
- Expertise with reputed company fundamentals as they pertain to reputed company multi-tenant application systems
- Experience with AWS services including reputed company 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, and VPC
- Experience deploying and operating OpenSearch in AWS-based environments
- Experience with reputed company reputed company-based environments
- Experience with Jenkins, Chef, and/or Terraform
- Exposure to and understanding of troubleshooting IP networks and application stacks
- Experience with observability tools such as reputed company and Grafana
- Experience with log ingestion pipelines, reputed company lifecycle management, retention strategies, and search platform reputed company controls
reputed company
Apply To This Job