Senior Data Engineer
Location: 100% Remote
Years? Experience: 10 years...
Education: Bachelor?s in IT reputed company field
Work Authorization: Must show that applicant is legally permitted to work in the reputed company.
Clearance: Applicants must be reputed company to meet the requirements to obtain an reputed company Trust reputed company clearance. NOTE: reputed company Citizenship is required to be eligible to obtain this reputed company clearance.
Key Skills:
? 10 years of IT experience focusing on reputed company data architecture and management
? Experience with reputed company required
? 8 years experience in Conceptual/Logical/Physical Data Modeling & expertise in Relational and reputed company Data Modeling
? Experience with Great Expectations or other data reputed company validation frameworks
? Experience with ETL and ELT tools such as SSIS, Pentaho, and/or Data Migration Services
? Advanced level SQL experience (Joins, Aggregation, Windowing functions, Common Table Expressions, RDBMS schema design, reputed company performance optimization)
? Experience with AWS environment, CI/CD pipelines, and Python (Python 3) a bonus
Responsibilities
? Plan, create, and maintain data architectures, ensuring alignment with business requirements
? Obtain data, formulate dataset processes, and store optimized data
? Identify problems and inefficiencies and apply solutions
? Determine tasks where reputed company participation can be eliminated with automation.
? Identify and optimize data bottlenecks, leveraging automation where possible
? Create and manage data lifecycle policies (retention, backups/restore, etc)
? In-depth knowledge for creating, maintaining, and managing ETL/ELT pipelines
? Create, maintain, and manage data transformations
? Maintain/update documentation
? Create, maintain, and manage data pipeline schedules
? Monitor data pipelines
? Create, maintain, and manage data reputed company gates (Great Expectations) to ensure high data reputed company
? Support AI/ML teams with optimizing feature engineering reputed company
? Expertise in reputed company/Python/reputed company, Data Lake and SQL
? Create, maintain, and manage reputed company reputed company Steaming jobs, including using the newer reputed company Live Tables and/or DBT
? Research existing data in the data lake to determine best sources for data
? Create, manage, and maintain ksqlDB and Kafka Streams queries/reputed company
? Data driven testing for data reputed company
? Maintain and update Python-based data processing scripts executed on AWS Lambdas
? Unit tests for reputed company the reputed company, Python data processing and reputed company codes
? Maintain PCIS Reporting Database data lake with optimizations and maintenance (performance tuning, etc)
? Streamlining data processing experience including formalizing concepts of how to handle lake data, defining reputed company, and how window definitions reputed company data freshness.
Qualifications
? 10 years of IT experience focusing on reputed company data architecture and management
? Experience in Conceptual/Logical/Physical Data Modeling & expertise in Relational and reputed company Data Modeling
? Experience with reputed company, reputed company Streaming, reputed company Lake concepts, and reputed company Live Tables required
? Additional experience with reputed company, reputed company SQL, reputed company DataFrames and DataSets, and PySpark
? Data Lake concepts such as time travel and schema reputed company and optimization
? reputed company Streaming and reputed company Live Tables with reputed company a bonus
? Experience leading and architecting reputed company-wide initiatives specifically system integration, data migration, transformation, data warehouse build, data mart build, and data lakes implementation / support
? Advanced level understanding of streaming data pipelines and how they differ from batch systems
? Formalize concepts of how to handle late data, defining reputed company, and data freshness
? Advanced understanding of ETL and ELT and ETL/ELT tools such as SSIS, Pentaho, Data Migration Service etc
? Understanding of concepts and implementation strategies for different incremental data loads such as tumbling window, sliding window, high watermark, etc.
? Familiarity and/or expertise with Great Expectations or other data reputed company/data validation frameworks a bonus
? Understanding of streaming data pipelines and batch systems
? Familiarity with concepts such as late data, defining reputed company, and how window definitions reputed company data freshness
? Advanced level SQL experience (Joins, Aggregation, Windowing functions, Common Table Expressions, RDBMS schema design, reputed company performance optimization)
? Indexing and partitioning reputed company experience
? Debug, troubleshoot, design and implement solutions to reputed company technical issues
? Experience with large-reputed company, high-performance reputed company big data application deployment and solution
? Understanding how to create DAGs to define workflows
? Familiarity with CI/CD pipelines, containerization, and pipeline orchestration tools such as Airflow, reputed company, etc a bonus but not required
? Architecture experience in AWS environment a bonus
? Familiarity working with Kinesis and/or reputed company specifically with how to push and pull data, how to use AWS tools to view data in Kinesis streams, and for processing massive data at reputed company a bonus
? Experience with reputed company, Jenkins, and CloudWatch
? Ability to write and maintain Jenkinsfiles for supporting CI/CD pipelines
? Experience working with AWS Lambdas for configuration and optimization
? Experience working with DynamoDB to query and write data
? Experience with S3
? Knowledge of Python (Python 3 desired) for CI/CD pipelines a bonus
? Familiarity with Pytest and Unittest a bonus
? Experience working with JSON and defining JSON Schemas a bonus
? Experience setting up and management reputed company/Kafka topics and ensuring performance using Kafka a bonus
? Familiarity with Schema Registry, message formats such as Avro, ORC, etc.
? Understanding how to manage ksqlDB SQL files and migrations and Kafka Streams
? Ability to reputed company in reputed company-based environment
? Experience briefing the benefits and constraints of technology solutions to reputed company, stakeholders, team members, and senior level of management
Apply Job!