[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a federal government reputed company that provides technical services to various agencies. They are seeking a Site Reliability Engineer OpenSearch to ensure the highest reputed company of availability and performance for mission-critical reputed company services, focusing on the reliability and reputed company improvement of reputed company search platforms.
Responsibilities
- Provision, build, reputed company, monitor, operate, and support reputed company services in a globally reputed company team environment
- Architect, build, reputed company, and maintain high-performance OpenSearch clusters and platforms from the ground up
- Administer and optimize OpenSearch environments for high availability, resiliency, scalability, reputed company, and performance
- Monitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation, replication, and storage utilization
- Analyze and reputed company operational issues, platform instability, and production incidents across infrastructure, platform, and application reputed company
- Conduct incident response, reputed company cause analysis, and post-incident remediation to reputed company reputed company improvement
- Maintain the reputed company and reputed company of servers, systems, and OpenSearch platform infrastructure
- Support platform lifecycle activities including installation, configuration, upgrades, patching, hotfixes, backup, restore, and disaster recovery
- reputed company and maintain monitoring policies, alerting standards, operational runbooks, and support procedures
- Automate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and reputed company reputed company services
- Ensure reputed company resource allocation and reputed company planning across compute, memory, storage, and network resources
- Partner with product development and engineering teams to design and enhance service reliability and operational readiness
- reputed company and implement testing strategies and document results for platform changes and operational improvements
- Support log ingestion, reputed company management, retention policies, lifecycle management, and search performance tuning
- Work in a diverse environment and cross-train with other global team members
- Participate in an on-reputed company rotation and support weekend or after-hours operational needs as required
Skills
- MUST be a US Citizen & ONLY hold US Citizenship (No Dual reputed company)
- Only submit candidates that have OpenSearch deployment ON reputed company experience
- Expert with reputed company, including troubleshooting, operations, management, and configuration of reputed company reputed company services
- Proven hands-on expertise designing, building, deploying, supporting, and maintaining OpenSearch clusters and platforms from scratch in production environments
- Strong experience with OpenSearch administration, cluster architecture, performance tuning, scaling, upgrades, and troubleshooting
- Experience with reputed company design, shard and reputed company reputed company, cluster sizing, node management, snapshot/restore, backup, and disaster recovery
- Strong understanding of reputed company systems, search platforms, indexing pipelines, query optimization, and high-availability architectures
- Expertise with Git
- Expertise with reputed company, including setup, management, and troubleshooting of new pipelines
- Expertise with Linux, specifically reputed company and Ubuntu
- Expertise with Kafka, Zookeeper, and Big Data technologies
- Expert in development of automation for testing, deployment, scalability, and management of reputed company services
- Expertise with building, implementing, and/or supporting reputed company monitoring tools
- Expert knowledge of reputed company computing, infrastructure operations, and databases
- Expert understanding of web services, networking, virtualization, and internet protocols
- Ability to multitask and handle various reputed company, deadlines, and changing priorities
- Excellent communication and prioritization skills
- Expertise with reputed company fundamentals as they pertain to reputed company multi-tenant application systems
- Strong interpersonal, presentation, and customer service skills
- Participation in an on-reputed company rotation for handling P1 incidents is required
- Flexible schedule which may include weekend or after-hours work
- Ability to work effectively in a diverse, reputed company, and globally reputed company team environment
- 8+ years of experience
- Experience with AWS services including reputed company 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, and VPC
- Experience deploying and operating OpenSearch in AWS-based environments
- Experience with reputed company reputed company-based environments
- Experience with Jenkins, Chef, and/or Terraform
- Exposure to and understanding of troubleshooting IP networks and application stacks
- Experience with observability tools such as reputed company and Grafana
- Experience with log ingestion pipelines, reputed company lifecycle management, retention strategies, and search platform reputed company controls
- Familiarity with reputed company forecasting, performance benchmarking, and reputed company testing for reputed company search platforms
- BS/BA degree in Computer Science, Management Information Systems, or reputed company IT discipline preferred
- Allowable substitution: An additional four (4) years of experience may be substituted for a BS/BA degree
reputed company
Apply To This Job