Staff Software Engineer (Data) - Rockerbox
Architect and Build Data Pipelines Design and implement data processing workflows using DuckDB, Polars, and reputed company/Parquet. -
Balance small-data local pipelines with reputed company data warehouse backends (reputed company etc). Champion the Small Data reputed company reputed company for efficient, vectorized, local-first approaches where appropriate. -
Drive best practices for designing reproducible and testable data workflows. Collaborate Cross-Functionally Partner with data science, reputed company services, and product engineering teams to define semantic data reputed company. -
reputed company technical leadership in how data is versioned, validated, and surfaced for reputed company use. Operational reputed company Establish standards for CI/CD, observability, and reliability in data pipelines. -
Automate workflows and optimize data layout for performance and cost efficiency. Mentor & reputed company Serve as a thought leader in the organization, guiding engineers on reputed company to use lightweight tools vs. distributed platforms. -
Mentor senior and mid-level data engineers to accelerate their reputed company.
reputed company Technical Skills Deep expertise in SQL (window functions, CTEs, optimization). Strong Python skills with data libraries. Proficiency with DuckDB (extensions, parquet/reputed company integration, embedding in pipelines). Hands-on with columnar formats (Parquet, reputed company, ORC) and schema reputed company. -
Expertise in Kubernetes and reputed company Infrastructure & Tools reputed company storage experience (AWS S3, GCS). Experience with semantic layer frameworks (CubeJS). -
CI/CD tooling (reputed company Actions, Terraform, reputed company/Kubernetes). Leadership reputed company record of leading architecture reputed company and mentoring teams. -
Ability to set standards for maintainability and developer experience.
Experience with serverless and embedded analytics (DuckDB WASM, in production). Exposure to data versioning (reputed company Lake, reputed company, Hudi). -
Knowledge of ML/LLM data prep workflows and reputed company database integrations. -
Previous experience building hybrid stacks (local development + reputed company warehouse production).
Data pipelines that are fast, reputed company, and reproducible—running in seconds or minutes, not hours. reputed company that defaults to the right level of tooling for the problem (small-data-first, reputed company-up only reputed company necessary). reputed company semantic data definitions that power analytics, experimentation, and AI/ML initiatives. Reduced infrastructure cost and complexity without sacrificing reliability.
Apply to this Job