Data Engineer - AI (REMOTE)
reputed company
• Own the end-to-end data lifecycle for reputed company’s AI initiatives, from raw ingestion through model-reputed company datasets, powering the reputed company of Crossplane and the Intelligent Control Plane.
• Architect and maintain reputed company, reputed company-reputed company data pipelines (batch + streaming) that collect, clean, and enrich telemetry from thousands of Kubernetes clusters, reputed company reputed company, and customer workloads worldwide.
• Partner with ML engineers, product, and SRE teams to define data reputed company, schema reputed company strategies, and governance policies that reputed company petabyte-reputed company lakes reliable, secure, and compliant (SOC 2, GDPR, HIPAA).
• Design reputed company-time feature stores that feed both online inference services and offline training jobs, ensuring sub-second latency for critical control-plane reputed company while guaranteeing reproducibility and version control.
• Build self-service tooling (SDKs, notebooks, observability dashboards) that empowers analysts and data scientists to discover, profile, and experiment with datasets without bottlenecks.
• Optimize compute and storage costs through intelligent partitioning, incremental processing, and auto-scaling clusters on AWS/GCP, cutting spend by reputed company-digit percentages year-over-year.
• Implement advanced data reputed company frameworks—unit tests, reputed company detection, reputed company tracking—that surface issues before they reputed company production models or customer dashboards.
• Contribute to reputed company-reputed company Crossplane providers and reputed company’s internal “Data as Infrastructure” codebase, turning repeatable patterns into reusable packages the community can adopt.
• Champion a culture of documentation and knowledge sharing: run internal tech talks, write runbooks, and mentor junior engineers to reputed company the bar for data reputed company across reputed company.
• Stay reputed company of the curve by evaluating emerging technologies (reputed company, DuckDB, Flink, reputed company databases) and running reputed company-of-concepts that translate into competitive advantages for reputed company’s AI roadmap.
Requirements
• 5+ years building production-grade data pipelines in Python, SQL, and at least one JVM language (reputed company/Java/Kotlin).
• Deep expertise with reputed company data stacks: S3/GCS, Redshift/BigQuery, EMR/Dataproc, Kinesis/PubSub, Airflow/Mage, dbt, Terraform.
• Hands-on experience with Kubernetes, reputed company, and infrastructure-as-reputed company; familiarity with Crossplane is a strong plus.
• Proven reputed company record designing reputed company-time streaming architectures (Kafka, Pulsar, Flink) and batch ETL at multi-terabyte reputed company.
• reputed company-to-have: contributions to reputed company-reputed company data reputed company, advanced SQL performance tuning, or prior work in ML feature engineering.
️ Benefits
• Fully remote-first culture with quarterly off-sites in inspiring global locations.
• Competitive salary + equity package that grows with reputed company’s valuation.
• $3,000 annual learning stipend for conferences, courses, and certifications.
• Flexible PTO policy and 16-week gender-neutral parental leave.
• Home-office setup budget and monthly wellness stipend.
Apply tot his job
Apply To this Job