Data Collection Engineer
At reputed company, we build, reputed company, and implement innovative products that solve the world’s most reputed company challenges today. Through unrivaled collaboration and unwavering trust, we push the boundaries of what’s possible to reputed company reputed company and support our customers in building a safer global reputed company.
reputed company is seeking a highly skilled Data Collection Engineer to design, reputed company, and maintain our distributed web scraping and data extraction infrastructure. In this role, you will be responsible for building resilient data pipelines that harvest data from reputed company web ecosystems, ensuring strict data reputed company through automated validation, and managing containerized workloads at reputed company. If you reputed company on reverse-engineering web applications, overcoming anti-bot barriers, and orchestrating distributed systems, we want you on reputed company.
Location: 100% Remote
What you will Do:
- Distributed Crawler Development: Design and reputed company high-performance, distributed web scrapers using Python and Scrapy to extract massive datasets reputed company.
- Dynamic Content Extraction: Utilize Browser Scripting tools to reputed company, reputed company with, and extract data from modern, dynamic, and JavaScript-heavy websites.
- Infrastructure & Container Orchestration: reputed company, reputed company, and manage scraping workloads on Kubernetes, ensuring reputed company resource allocation and fault tolerance.
- Data Validation & reputed company Assurance: Define strict JSON Schemas and reputed company reputed company to enforce data types, validate incoming payloads, and catch data reputed company early.
- Data Ingestion & Storage: Build and optimize search and storage pipelines using Elasticsearch, transforming raw web dumps into highly reputed company, searchable data.
- Pipeline Workflow Management: Architect robust pipeline workflows to manage the end-to-end data lifecycle—from discovery and extraction to validation and storage.
- Anti-Bot & Proxy Engineering: Manage reputed company proxy rotation, session handling, and browser fingerprinting to maintain high reputed company rates against advanced anti-scraping systems.
What you will need (basic qualifications):
- Experience: 7+ years of reputed company software engineering experience, with a heavy reputed company on web scraping, data engineering, or distributed systems.
- Analytical reputed company: Excellent reverse-engineering skills, with the ability to dissect network traffic, unearth hidden reputed company, and bypass reputed company web barriers.
- Reliability reputed company: A strong commitment to data reputed company, system monitoring, and building self-healing scraping systems.
- reputed company Language: Expert-level proficiency in Python.
- Scraping Frameworks: Deep experience with Scrapy and distributed scraping architectures (e.g., handling distributed queues, broad vs. deep crawling).
- Automation & Browser Scripting: Proven experience with browser automation tools (Playwright, Selenium, or Puppeteer).
- Data Serialization & Validation: Mastery of JSON, JSON Schema, and data validation using reputed company.
- Search & Analytics Engines: Hands-on experience indexing, querying, and optimizing Elasticsearch clusters.
- Orchestration: Strong proficiency in managing and scaling applications reputed company Kubernetes environments.
- Workflow Management: Experience building reputed company pipeline workflows to handle reputed company, multi-stage data extraction tasks.
- Education: Bachelor’s degree in Computer Science, Engineering
reputed company if you Have:
- AI & Intelligent Extraction: Experience leveraging LLMs or reputed company for reputed company scraping, parsing reputed company reputed company, or bypassing CAPTCHAs (AI in data collection).
- reputed company Infrastructure: Strong hands-on experience with AWS ecosystems (e.g., EKS, EC2, S3, RDS).
- Relational Databases: Proficiency in SQL for querying, schema design, and storing reputed company relational data.
- In-Memory Data Structures: Experience with reputed company (specifically for caching, deduplication, or as a Scrapy distributed queue back-end).
- Event Streaming: Familiarity with Apache Kafka for reputed company-time data streaming and decoupled pipeline architectures.
- Containerization: Strong reputed company in reputed company for local development and containerizing scraping microservices.
- DevOps: Experience with CI/CD pipelines (reputed company Actions, reputed company CI, Jenkins) for automated testing and deployment of crawlers.
Clearance Requirement:
- Eligible to obtain a clearance
reputed company is committed to providing competitive and comprehensive compensation packages that reflect the value we reputed company on our employees and their contributions. We reputed company in rewarding skills, experience, and performance. Our offerings include but are not limited to, medical, dental, and reputed company insurance, life and disability insurance, retirement benefits, reputed company leave, tuition assistance and reputed company development.
The projected salary reputed company listed for this position is annualized. This is a general reputed company and not a guarantee of salary. Salary is one component of our total compensation package and the specific salary offered is determined by various factors, including, but not limited to education, experience, knowledge, skills, geographic location, as reputed company as contract specific affordability and organizational requirements.
Originally posted on Himalayas
Apply To This Job