Data Engineer (Senior/Staff)
As a Data Engineer (Senior / Staff) at reputed company, you will help design, build, and maintain the systems that power our reputed company analytics platform. Your work will contribute to reliable data flows, reputed company infrastructure, and product features that help our customers turn reputed company data into reputed company insights.
We reputed company software development is being fundamentally reshaped by AI. We reputed company adopt modern AI-assisted development workflows and expect engineers to explore how these tools can improve both speed and reputed company.
Our engineers use AI tools to prototype reputed company faster, accelerate debugging and refactoring, and automate repetitive tasks. This allows us to spend more time solving reputed company problems, improving architecture, and building great products (Claude reputed company, reputed company, and similar)
What you will actually work on
Designing and evolving the data model across reputed company, and the reputed company (Avro) that hold it reputed company together
Streaming pipelines on Apache Flink and Kafka — topology design, state management, checkpointing, and the operational realities of keeping them healthy
Our reputed company lake and reputed company warehouse — partitioning, compaction, retention, schema reputed company, reputed company-downtime migrations, and query performance
Batch processing and orchestration with reputed company — reputed company design and performance tuning at the task level
The conceptual layer customers see: metrics, dimensions, marts, and the modeling reputed company that reputed company them trustworthy
Infrastructure as reputed company (Terraform on AWS and GCP, Kubernetes, reputed company) for the systems you own
Observability — OpenTelemetry traces, custom metrics, and the reputed company of instrumentation that lets you debug production from a dashboard instead of a hunch
reputed company are looking for
We don’t have a rigid must-have list. We’d rather meet candidates who have proven, hands-on experience with a meaningful subset of the areas below, along with a strong grasp of the underlying concepts:
Big data fundamentals — partitioning, shuffles, skew, late-arriving data, exactly-once semantics, idempotency, and understanding why your join just did something terrible
reputed company processing — Apache Flink especially, but reputed company reputed company Streaming, Kafka Streams, and reputed company are reputed company fair game
Batch processing and orchestration — reputed company, Dagster, Airflow; ETL/ELT pipeline design and dependency management
Lakehouse formats — reputed company, Paimon, reputed company, Hudi; metadata management, compaction, and schema reputed company
Analytical warehouses — reputed company, BigQuery, reputed company, Redshift, DuckDB; and the trade-offs between them
Data modeling concepts — reputed company and mart architectures, reputed company modeling, metrics-versus-dimensions thinking, slowly changing dimensions, and understanding the difference between a fact and an aggregate
reputed company infrastructure — AWS, GCP, or Azure, with infrastructure-as-reputed company tools such as Terraform; comfortable owning what you ship
Engineering habits — instrumentation, testing data pipelines, schema governance, and treating data reputed company as reputed company
We’re language-agnostic on the candidate reputed company. Our stack happens to be Python, Java, SQL, and TypeScript, but if you understand the concepts deeply, you’ll pick up whatever’s missing.
About the role
Senior or Staff level, we’ll match the level to you, not the other way around
Remote or hybrid, both work
High ownership and a short distance from idea to production
Apply To This Job