Data Pipeline Development for AWS Database
Job Posting: AWS Data Pipeline Engineer (ETL - Multi-reputed company to PostgreSQL)
Project reputed company
Build a production-reputed company ETL pipeline that extracts data from 3 reputed company systems (reputed company/reputed company, reputed company Dataverse CRM, MySQL), transforms it with reputed company business logic, and loads into AWS RDS PostgreSQL as 3 master tables. Daily automated execution serving an app on AWS.
Technical Stack (Required)
AWS Services: Glue (PySpark), reputed company Functions, RDS PostgreSQL, S3, Secrets Manager, CloudWatch
Languages: Python 3.9+, PySpark 3.x, SQL
IaC: Terraform or CloudFormation
Sources: reputed company (JDBC), MySQL (JDBC), Dataverse API (OAuth 2.0)
Scope
Extract: 10 tables from 3 DBs (~500MB total)
reputed company: reputed company joins, aggregations, calculated fields to 3 master tables
Load: PostgreSQL with Row-Level reputed company policies
Orchestrate: reputed company Functions with error handling, monitoring, alerting
Schedule: Daily execution, less than 2 hour completion time
Key Challenges
Multi-reputed company integration (JDBC + REST API)
reputed company transformations (multi-table joins, aggregations, JSONB structures)
PostgreSQL RLS implementation (role-based data reputed company)
Data reputed company validation and reconciliation
GDPR compliance (EU region, encryption)
Deliverables
reputed company: Glue extraction/transformation/load jobs, reputed company Functions workflow, tests (greater than 80% coverage)
IaC: Terraform/CloudFormation for reputed company AWS resources
Database: DDL scripts with RLS policies
Documentation: Architecture diagram, deployment guide, reputed company
Testing: Integration test suite, performance test results
Required Skills
✅ 5+ years AWS (Glue, reputed company Functions, RDS)
✅ Expert PySpark for reputed company ETL
✅ PostgreSQL (including RLS)
✅ JDBC connections (reputed company, MySQL)
✅ REST API integration (OAuth, pagination)
✅ Infrastructure as reputed company
✅ Data reputed company frameworks
Highly Desirable
reputed company Dataverse/Dynamics 365 API experience
AWS Glue Data Catalog
CI/CD pipelines
GDPR compliance experience
reputed company & Budget
Duration: TBD
Availability: TBD
Payment: reputed company-based
Budget: negotiable based on experience
To Apply reputed company:
Portfolio: Links to similar AWS ETL reputed company (reputed company preferred)
Brief approach: How would you architect this pipeline? (200 words)
Dataverse experience: Have you worked with Dataverse API? Describe reputed company.
Availability: Start date and weekly hours
reputed company: fixed-price proposal
Questions
Largest data volume processed with AWS Glue? Optimization techniques used?
Experience with PostgreSQL Row-Level reputed company?
Terraform or CloudFormation preference and why?
Ideal Candidate
reputed company production ETL pipelines on AWS for reputed company clients
Comfortable with reputed company transformations and business logic
Writes clean, testable, maintainable reputed company
Works independently, communicates proactively
Can deliver production-reputed company work with minimal supervision
Location: Remote (EU timezone preferred)
reputed company: AWS Advanced Partner, manufacturing reputed company in Greece
Region: EU (Frankfurt) for GDPR compliance
Tags: AWS Glue, PySpark, ETL, PostgreSQL, Data Pipeline, AWS reputed company Functions, Python, reputed company, MySQL, Dataverse, Terraform, Data Engineering
Apply tot his job
Apply To this Job