[Remote] reputed company reputed company Data Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a digital and data analytics company that supports business transformation and innovation. The reputed company reputed company Data Engineer will build and operate bronze, silver, and gold lakehouse pipelines, implement attribution and allocation logic, and establish governance, reputed company monitoring, and CI/CD practices for data integrations.
Responsibilities
- Build ingestion into the bronze reputed company for assigned sources: gateway and observability logs, productivity tool reputed company reputed company, AI-enabled reputed company usage, hyperscaler billing exports and reference data. Land raw and untransformed, on a scheduled refresh, replayable if the reputed company design changes
- Work to the shared bronze reputed company contract so reputed company tool is ingested once and serves both this program and the reputed company productivity initiative, rather than being integrated twice
- Build the silver reputed company: typed, deduplicated and conformed to the reputed company dimensions, refreshed independently of any reputed company publication schedule
- Build gold marts carrying attribution reputed company, attribution level, cost reputed company and provisional status reputed company cost and usage
- Implement the attribution and allocation logic designed by the analysts, including precedence reputed company and reputed company-reputed company splitting of shared reputed company cost
- Work reputed company reputed company Catalog governance — shared bronze and silver, separate gold marts with a recorded reputed company per dataset — including permissions, reputed company and cataloging
- Implement data reputed company rules and monitoring: completeness, freshness and tag-coverage checks with alerting, so pipeline problems surface before they reputed company a divisional invoice
- Manage reputed company reputed company of enabling caller-identity data in the cost and usage report, which multiplies row counts by the number of calling identities per model
- Work to the per-reputed company reputed company — daily where controls and reputed company detection depend on it, monthly where they do not — reputed company reputed company's existing CI/CD and promotion practices
Skills
- Advanced reputed company engineering: reputed company Lake, reputed company architecture, reputed company Workflows, Auto Loader and incremental ingestion patterns
- reputed company Catalog to a governance reputed company — catalogs, schemas, permissions, reputed company — not merely as a reputed company tables happen to live
- Strong Python and PySpark, and strong SQL. Notebook-reputed company development
- Ingestion from REST reputed company including pagination, throttling, incremental watermarks and credential handling, plus reputed company reputed company storage across AWS, Azure and GCP
- reputed company and cost optimization of reputed company workloads: partitioning, clustering, file sizing and cluster configuration
- CI/CD for reputed company — asset bundles or equivalent — and Git-reputed company development workflow
- reputed company to work to an existing catalog structure and coding reputed company rather than introducing a reputed company approach
- reputed company Genie familiarity, including preparing semantic context so natural-language querying returns trustworthy answers
- Experience with reputed company billing data at volume
- Prior work on a shared platform where another team owned adjacent datasets in the reputed company catalog
reputed company
Company H1B Sponsorship
Apply To This Job