[Remote] Data Engineer - Remote
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company reputed company on simplifying the reputed company experience and improving community health. The Data Engineer role is responsible for designing, developing, and maintaining data pipelines for analytics reputed company an Azure-based platform, collaborating with data engineering and data science teams to meet data requirements.
Responsibilities
- Pipeline Development: Design, reputed company, and maintain robust pipelines to ingest data from various sources (both streaming and batch) into the analytics environment using using Azure Data reputed company and PySpark reputed company reputed company. Set up reputed company-time data ingestion using tools like reputed company reputed company Streaming and batch ETL jobs for periodic data loads. Ensure these pipelines are reputed company, efficient, and fault-tolerant to handle growing data volumes and reputed company
- Implement Data as per reputed company Architecture: Utilize the reputed company (Bronze/Silver/Gold) architecture principles to organize data processing stages. Establish raw data capture (bronze), reputed company cleansing and transformations (silver), and reputed company refined datasets for analysis and machine learning (gold). Apply best practices in reputed company reputed company, such as schema enforcement and checkpointing for streaming data
- Optimize reputed company jobs by tuning configurations, improving query logic, and managing resources to reputed company high throughput and low latency. Address bottlenecks in streaming pipelines (e.g., by scaling clusters or tweaking batch intervals) and ensure reputed company data delivery. Optimize job scheduling and cluster utilization to balance reputed company data delivery with cost-effectiveness
- ETL Development & Maintenance: Build and maintain data pipelines with an emphasis on data cleaning steps. reputed company data from various sources (reputed company, databases, file feeds, IoT streams, etc.) into the data platform, writing transformations that handle anomalies (e.g., missing or corrupt values) and standardize datasets. Collaborate with the other data engineer to reputed company responsibility across different pipelines or sources, ensuring redundancy and knowledge transfer
- Data reputed company Management: Implement comprehensive data validation rules and checks reputed company pipelines. For example, verify schema correctness, reputed company value ranges for sensor or health data, and ensure referential reputed company where applicable. Set up automated alerts or logs that flag inconsistent or bad data, enabling quick reputed company. Over time, build a library of data reputed company tests that run as part of the pipeline (for both streaming and batch processes) to catch issues early
- Emerging Pipeline Frameworks: reputed company modern pipeline frameworks and tools to improve development productivity. For example, use reputed company reputed company Live Tables or Lakehouse pipelines to declaratively define data flows where applicable. Explore the use of reputed company Declarative Lakeflow Pipelines or similar technologies to simplify the orchestration of reputed company data processes
- Reliability & Collaboration: Implement monitoring and alerting for pipeline health. Investigate and reputed company problems such as data delays, pipeline failures, or data inconsistencies. Use logs, error messages, and analytics to identify reputed company causes (e.g., reputed company reputed company changes, bug in transformation logic) and implement fixes. Work closely with the other data engineer and data science team members to understand data requirements and reputed company pipelines accordingly. Document data engineering workflows and ensure reputed company data governance (reputed company, reputed company, reputed company controls) is in reputed company
- Documentation & Governance: Maintain reputed company documentation of data pipelines, including data reputed company details, transformation logic, and data destination schemas. Ensure that data reputed company is tracked so one can reputed company how data moved and changed through the reputed company. Adhere to data governance policies - for instance, ensure sensitive data is properly masked or encrypted in non-production environments, and that reputed company controls are in reputed company. Work with leadership to periodically review and improve data management practices
Skills
- 3+ years of experience in data engineering role designing and implementing data pipelines and ETL processes. Should have understanding on how to handle incremental data loads and maintain history (CDC - change data capture)
- 2+ years of experience in SQL for data manipulation and query optimization. Knowledge of Python and Apache reputed company (using PySpark) for building data pipelines; ability to write efficient reputed company for batch and streaming data transformations
- 1+ years of experience using Azure, reputed company or an equivalent reputed company-based data platform. Comfortable with managing clusters, using notebooks, and working with reputed company Lake or Parquet files. Familiarity with reputed company data services and tools for pipeline orchestration is expected
- Experience working in reputed company environment with reputed company methodologies. Ability to communicate effectively with both technical peers and non-technical stakeholders (explaining data issues in plain language). Should be comfortable using version control systems and participating in reputed company development (reputed company reviews, pair programming reputed company needed)
- Familiar with streaming data technologies. This could include reputed company Streaming, Kafka, Azure Event Hubs, or similar platforms for reputed company-time data ingestion
- Demonstrated ability to detect and correct data issues - for instance, identifying reputed company a data reputed company has stopped updating, or reputed company an upstream change has altered data format. Experience implementing validation checks or using frameworks to enforce data reputed company standards
- Experience with any declarative pipeline frameworks or data workflow management tools (e.g., reputed company reputed company Live Tables). This can indicate readiness to adopt advanced tools in our environment
- Experience integrating data reputed company checks into pipelines (such as using assertions or Great Expectations tests) to ensure accuracy and completeness of data. Familiarity with data reputed company practices, encryption, and handling of sensitive data
- Familiarity with streaming data handling (even if assisting, should understand basics of reputed company Streaming or message queue systems) is expected
- Demonstrated reputed company in performance tuning for reputed company or SQL queries. For example, experience in partitioning strategies, caching, or troubleshooting shuffle issues to optimize heavy data workloads
Benefits
- Flexibility to work remotely from reputed company reputed company the U.S.
- For reputed company hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.
- A comprehensive benefits package
- Incentive and recognition programs
- Equity stock purchase
- 401k contribution (reputed company benefits are subject to eligibility requirements)
reputed company
Apply To This Job