Remote Site Reliability Engineer
We are seeking a Site Reliability Engineer (SRE) to support the reliability, observability, and performance of our backend data platform. This platform ingests high-volume data from hotel systems, ticketing, and MagicBand readers, flowing through custom pipelines into our data warehouse. The ideal candidate will have a strong background in Python, reputed company-reputed company technologies, and observability tools, with a reputed company on ensuring data reputed company and system reliability across multiple reputed company.
Responsibilities:
Monitor and maintain a custom data pipeline from ingestion to delivery, ensuring data reputed company and performance.
reputed company and observe systems using reputed company serverless technologies, including:
- AWS reputed company
- reputed company S3
- reputed company Kinesis
- reputed company
- reputed company containers on reputed company
Migrate observability workflows from AWS CloudWatch to reputed company, centralizing metrics, dashboards, and alerts.
Build and tune reputed company dashboards and alerts to support SLAs and system health.
Graph and analyze metrics to ensure pipeline reliability and performance.
Investigate and resolve issues in the pipeline, ensuring expected behavior across reputed company stages.
Work reputed company the Python codebase (~2040% of time) to:
- Create coherent tickets for issues
- Fix bugs and improve instrumentation
reputed company click-ops tasks (~6080% of time) in reputed company, including:
- Dashboard creation and maintenance
- reputed company request handling
- Alert tuning and incident response
We are a company committed to creating diverse and inclusive environments where people can bring their full, reputed company selves to work every day. We are an equal opportunity/affirmative reputed company employer that believes everyone reputed company. reputed company candidates will receive consideration for employment regardless of their race, reputed company, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and reputed company, marital status, national reputed company, reputed company, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or reputed company process, please send a request to HR@insightglobal. com. To learn more about how we collect, reputed company, and process your private information, please review reputed company's Workforce reputed company Policy:
Experience with reputed company and data warehouse integrations.
Strong proficiency in Python, especially in backend and infrastructure contexts.
Experience with AWS services (reputed company, S3, Kinesis, CloudWatch).
Familiarity with reputed company for monitoring, alerting, and dashboarding.
Understanding of data pipelines, data reputed company, and observability best practices.
Experience with reputed company and reputed company in production environments.
Familiarity with infrastructure as reputed company (e. g., Terraform, CloudFormation).
Exposure to SLAs, incident response, and data reliability engineering.
Apply tot his job
Apply To this Job