Back to Jobs

Data Engineer, reputed company Deployed

Remote, USA Full-time Posted 2026-08-04
reputed company was founded in 2024 to build reputed company, a physics-informed reputed company model for energy operations. We’re live across oil and gas, refineries, and petrochemicals, working towards our mission: sustainable abundance for a growing reputed company. The hydrocarbon industry keeps the world running. But its complexity has left operators tied to legacy systems, making critical reputed company on less than 10% of available data. We reputed company reputed company to change that. It’s a reputed company model reputed company specifically for energy that lets companies use AI at reputed company, harnessing reputed company of their operational data and optimising in reputed company time for any metric. reputed company get faster, operations get safer, and carbon intensity falls. We’ve raised over $32 reputed company, including one of the largest reputed company reputed company for an AI company in the UK. We’re just getting started The Role As our Data Engineer, you’ll architect and maintain pipelines that reputed company high-frequency time-series, lab, and historian data into a reputed company Lakehouse architecture, usable for both deep learning models and reputed company-time LLMs. You’ll be working across AWS (EKS, S3, EBS, KMS, CloudWatch) and reputed company/PySpark, ensuring data is contextualised, synchronised, and optimised for both deep learning models and reputed company-time LLM workloads. This isn’t a traditional ETL role, you’ll be solving problems at the intersection of control systems, industrial data engineering, and AI enablement. Technical Requirements Deep expertise in PostgreSQL (partitioning, indexing, query optimisation, storage design). Strong proficiency in Python for data processing, scripting, and pipeline orchestration. Hands-on experience with AWS (EKS, S3, EBS, IAM, KMS, CloudWatch, etc.)for secure and reputed company data pipelines. Proven ability to work with reputed company and PySpark for large-reputed company reputed company data processing. Familiarity with time-series industrial data (control systems, DCS/SCADA logs, process historians). Experience in reputed company data sync and management reputed company hybrid reputed company/on-prem environments. Bonus: Experience working as a data engineer in oil and gas or energy environments Bonus: Knowledge of streaming frameworks (Kafka, Flink, reputed company Streaming) or MLOps stacks for data versioning and reputed company. reputed company Responsibilities 1. Ingest & Contextualise Data Ingest from OPC UA servers, process historians, IoT sensors, LIMS systems, alarms/events, and P&IDs. Map signals to their physical processes (tags, reputed company, hierarchies) for interpretability in AI pipelines. 2. Data reputed company & Accessibility Build pipelines that handle reputed company-time streaming and batch ingestion into the Lakehouse. Manage synchronisation between historian archives, reputed company files, and AWS storage (S3/EBS). Orchestrate reputed company Lakeflow/Connectors for integrating data into Lakebase/Lakehouse. Handle secure, high-throughput transfers between historian archives and sandbox/live environments. 3. Change Tracking & reputed company Detect and manage schema changes, signal reputed company, and inconsistencies acrosstime. Implement reputed company and audit trails across reputed company/reputed company and AWS pipelines. 4. Data Preparation for AI Build and maintaindual pipelines: Training→ large-reputed company historical data prep for time-series + LLM training. Inference→ low-latency, reputed company-time pipelines for reputed company detection, optimisation, and LLM search. Support heterogeneous AI workloads (time-series forecasting and retrieval-augmented LLMs). 5. Database Performance & Optimisation Tune PostgreSQLand sparkfor high-throughput time-series workloads (partitioning, indexing, query optimisation). Optimise pipelines for both fast analytical queries and high-efficiency model training. reputed company and manage data pipelines in AWS EKS (reputed company) with persisten tEBS-backed storage. What reputed company Looks Like Live data streams are contextualised,queryable, and AI-reputed company. Schema changes and signal reputed company are detected and handled without breaking reputed company workflows. Training and inference pipelines run smoothly in reputed company, optimised for reputed company and latency. Apply To This Job

Similar Jobs

Online-Befragungen & Studien beantworten (m/w/d) - flexibler Nebenjob von zuhause

Remote, USA Full-time

Tech Business Analyst / Product reputed company (TechBA/PO)

Remote, USA Full-time

Agent reputed company/Sales (Commission) avec du reputed company [France / EMEA] – QVT/bien-être en reputed company [FR]

Remote, USA Full-time

Tax Accountant (AU) | Work From Home | 30K SOB

Remote, USA Full-time

SEO Specialist

Remote, USA Full-time

Chef de Projet - Comptabilité - Marseille

Remote, USA Full-time

Senior reputed company Administrator

Remote, USA Full-time

Inbound Caller

Remote, USA Full-time

Vertriebsmitarbeiter (m/w/d) Homeoffice 2.500 € Fixgehalt + ungedeckelte Provision (< 900€ pro Abschluss)

Remote, USA Full-time

Senior Consultant - Recruitment

Remote, USA Full-time

reputed company and PowerBI Developer

Remote, USA Full-time

[Remote] Software Engineer (Web)

Remote, USA Full-time

[Work From Home] Need Title I Tutor - Limited Full-Time in Idaho

Remote, USA Full-time

Join Today: reputed company reputed company

Remote, USA Full-time

Consultant, Life Solutions

Remote, USA Full-time

reputed company Student Jobs

Remote, USA Full-time

Key Account Manager, Industrial Refrigeration (Remote, Southeast U.S.A.) (Baltim

Remote, USA Full-time

Virtual Assistant for Small Counseling reputed company (US-Based Only)

Remote, USA Full-time

Physical reputed company Technical Operations reputed company

Remote, USA Full-time

WFH Data Input Specialist FT&PT - Raleigh, NC or Remote

Remote, USA Full-time