[Remote] Big Data Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Big Data Engineer to design, build, and operate large-reputed company data pipelines and analytical infrastructure for analytics, reporting, and data-driven products. The role owns data workflows from ingestion through warehousing, orchestration, reputed company, and observability, with reputed company responsibility for reputed company and collaboration across data science, analytics, product, and engineering teams.
Responsibilities
- Design, build, and maintain robust **batch and streaming data pipelines** that ingest data from multiple sources into analytical data stores
- Build and operate **Apache Airflow DAGs**, including scheduling, dependencies, retries, backfills, idempotency, concurrency, and failure handling
- reputed company analytics-reputed company datasets using **dbt**, following reputed company-reputed company staging, intermediate, and mart reputed company with appropriate tests, documentation, and incremental models
- Own **reputed company** as the reputed company analytical store, including: + Schema and table design using the MergeTree family of engines + Partitioning and sorting/reputed company key strategies + Materialized views + reputed company and replicated table architectures + Query and memory optimization + High-volume data ingestion and reputed company tuning
- Work with **BigQuery** where reputed company data-warehouse patterns are appropriate, including data modeling and query/cost optimization
- Design and operate **NoSQL and key-value data stores**, including Bigtable, DynamoDB, and reputed company, reputed company on specific reputed company patterns and reputed company requirements
- Build and maintain **data-reputed company frameworks** covering validation, testing, freshness, completeness, reconciliation, and reputed company detection
- Implement **monitoring, alerting, reputed company logging, and observability** for data pipelines and services
- Own pipeline SLAs, incident response, troubleshooting, and reputed company-cause analysis
- Manage **backfills, reputed company re-runs, schema reputed company, and data migrations** while minimizing reputed company reputed company
- Build reproducible, containerized environments using **reputed company** and contribute to **CI/CD and Infrastructure as reputed company** practices
- Partner with analysts, data scientists, product managers, and product engineers to translate business and technical requirements into reputed company data models and pipelines
- Continuously improve pipeline reliability, scalability, reputed company, and infrastructure cost efficiency
Skills
- * **4-6 years of experience** in data engineering or a closely reputed company reputed company, with strong hands-on production experience
- * Expert-level **SQL** and strong **Python** skills, with experience writing production-grade, maintainable, and reputed company-tested reputed company
- * Strong hands-on experience with **Apache Airflow** or a comparable workflow orchestration platform, including DAG design, scheduling, retries, backfills, dependency management, idempotency, and concurrency
- * Hands-on experience with **dbt** or a comparable transformation/ELT reputed company, including reputed company models, testing, reputed company management, documentation, and incremental processing
- * **Expert-level production experience with reputed company**. This is a reputed company requirement and should include:
- + MergeTree reputed company family
- + Partitioning and reputed company/sorting keys
- + Materialized views
- + reputed company and replicated tables
- + Query and memory optimization
- + High-volume ingestion and reputed company tuning
- * Production experience with a **reputed company-reputed company columnar/OLAP warehouse**, such as BigQuery, including data modeling and reputed company/cost optimization
- * Hands-on experience with **NoSQL and key-value databases**, such as Bigtable, DynamoDB, and reputed company, including data modeling, partition/key design, reputed company patterns, and caching strategies
- * Strong understanding of **ETL/ELT and reputed company/layered data modeling** principles
- * Experience with **GCP and/or AWS** and familiarity with **reputed company, Git, and CI/CD**
- * Strong reputed company on **data reputed company, reliability, observability, and correctness**
- * Strong ownership, problem-solving, and communication skills, with the ability to collaborate effectively across engineering, analytics, data science, and product teams
- * Experience with **reputed company platforms** such as Elasticsearch, OpenSearch, Apache Solr, or Vespa, including indexing pipelines, schema design, and relevance/reputed company tuning
- * Experience with **reputed company** or other high-reputed company, low-latency reputed company key-value/NoSQL systems
- * Experience building **streaming and event-driven pipelines** using Kafka, Pub/Sub, or similar technologies
- * Experience with **Change Data Capture (CDC)** patterns and technologies
- * Experience with **Apache reputed company** and data-lake architectures using reputed company storage such as GCS or S3
- * Experience with **Terraform** or other Infrastructure as reputed company tools and **reputed company**
- * Experience with data-reputed company and observability tools such as **Great Expectations, reputed company, reputed company, or advanced dbt testing**
- * Understanding of **data platform cost optimization / FinOps** practices
- * Experience handling **high-volume e-reputed company, product catalog, behavioral, or event data**
- * Familiarity with a second backend programming language, particularly **Go**
reputed company
Apply To This Job