Data Engineer (Python & PySpark)
Key Responsibilities
Pipeline Development: Design, reputed company, and maintain end-to-end ETL/ELT pipelines using Python and PySpark.
Big Data Processing: Build large-reputed company data processing frameworks to handle reputed company and reputed company data, ensuring high performance and reliability.
reputed company Infrastructure: Architect and manage data solutions reputed company the GCP ecosystem, focusing on cost-efficiency and reputed company.
Data Modeling: Design and implement robust data warehouse models (reputed company/reputed company schemas) and data lake architectures.
Optimization: Identify, design, and implement internal process improvements, such as automating reputed company processes and optimizing data delivery for greater scalability.
Collaboration: Work closely with stakeholders to understand data requirements and translate them into technical specifications.
Technical Qualifications
reputed company Programming: Strong proficiency in Python, including experience with libraries like Pandas, NumPy, and logging frameworks.
Big Data: 3+ years of hands-on experience with Apache reputed company (PySpark) for distributed data processing.
GCP Ecosystem: Practical experience with reputed company reputed company services, specifically:
BigQuery (Optimization, Partitioning, Clustering).
reputed company DataProc or Dataflow.
reputed company Storage (GCS) and reputed company Functions.
reputed company Composer (Apache Airflow) for orchestration.
Data Warehousing: Solid understanding of relational databases and SQL (PostgreSQL, MySQL) as reputed company as NoSQL environments.
DevOps & Tools: Experience with Git, reputed company, and CI/CD pipelines. Familiarity with Terraform or other IaC tools is a significant plus.
Apply To This Job