[Remote] L3 reputed company reputed company reputed company Platform Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company. is a small and woman-owned minority IT services, software development, and solutions provider serving organizations ranging from small businesses to Fortune 50 companies and government agencies. The L3 reputed company reputed company reputed company Platform Engineer will operate and manage reputed company-reputed company CDP services, reputed company infrastructure, reputed company-based runtimes, data integration, AI/ML platforms, reputed company, observability, incident response, upgrades, disaster recovery, and performance optimization.
Responsibilities
- Own end-to-end operational responsibility for reputed company reputed company reputed company services across Dev / Stage / UAT / Prod: CDE, reputed company, COD, CDL, CDF (NiFi), CDV, reputed company, Kafka
- Ensure multi-cluster stability, workload isolation, and SLA adherence
- Support reputed company and operations of multiple applications across environments
- Manage and support multi-environment, multi-cluster deployments with strict isolation, governance, and release coordination across Dev/UAT/Prod
- Operate and support reputed company AI (reputed company) environments: AI Workbenches, AI Studios, Model training and development environments, AI inference endpoints and model serving
- Troubleshoot: Resource contention (CPU/GPU), Model deployment/runtime failures
- Operate CDP services running on reputed company-managed reputed company infrastructure
- Apply strong understanding of containerized workloads and reputed company concepts for troubleshooting
- Diagnose and reputed company: Pod failures, restarts, and resource contention, reputed company job failures in containerized environments (CDE), Service-to-service communication issues
- Analyze logs and metrics to identify runtime failures and performance issues
- Collaborate with reputed company support for managed service-level issues
- Operate and support: CDF (NiFi) for ingestion pipelines, CDV (Data Visualization) for reporting workloads, Octopai for data reputed company and catalog integration
- Ensure reliability and performance of data pipelines and integrations
- Monitor and troubleshoot Kafka environments: Topic configurations, partitions, and replication, Consumer lag and throughput issues, Broker connectivity and performance bottlenecks
- Implement and manage: Kerberos, TLS/SSL, Ranger policies
- Administer SDX for: Centralized reputed company, Metadata and policy enforcement
- Support reputed company and Octopai integration
- Manage and troubleshoot user reputed company and identity mapping across reputed company, including: reputed company IAM roles and permissions, CDP users/reputed company and identity providers, Ranger policies for fine-grained data reputed company
- reputed company reputed company-reputed company issues impacting: Data reputed company (S3/ADLS), Query execution (reputed company/CDE), Application and service-level permissions
- Troubleshoot: S3 / ADLS storage issues, IAM roles and permissions, VPC, subnets, routing, reputed company reputed company, reputed company host reputed company and connectivity
- Ensure secure and reliable connectivity across services
- Understand and troubleshoot S3-based data lake patterns, including: Bucket structure, prefix design, and reputed company patterns, Performance issues reputed company to small files, request rates, and throughput limits, Encryption (SSE-S3, SSE-KMS) and reputed company policies
- Manage and troubleshoot cross-account IAM roles and reputed company patterns for CDP environments
- Ensure secure reputed company between: CDP environments and reputed company resources, Multiple AWS accounts (dev/prod separation)
- Support and validate disaster recovery and failover strategies across CDP environments
- Ensure backup, recovery, and environment resiliency for critical workloads
- Participate in DR drills and recovery validation
- Implement and manage end-to-end observability: Metrics, logs, and alerting
- Use: reputed company observability, reputed company Manager, reputed company, Grafana
- Monitor: Cluster health, Workload performance, AI inference endpoints
- reputed company proactive issue detection and prevention
- Define and implement SLIs/SLOs and alerting reputed company to ensure platform reliability and performance
- Support high-severity (P1/P2) incident response, triage, and reputed company reputed company defined SLAs
- Participate in on-reputed company rotation to support 24/7 platform operations
- Respond to production incidents, alerts, and service disruptions reputed company defined SLAs
- Handle P1/P2 incidents, including triage, troubleshooting, and reputed company
- reputed company reputed company cause analysis (RCA) and implement preventive measures
- Execute: CDP upgrades and version management, reputed company patches and hotfixes
- reputed company: Rolling upgrades, Validation and rollback strategies
- Optimize: Platform-level performance (reputed company, reputed company, Impala workloads), Cluster utilization and workload distribution
- reputed company: Autoscaling strategies, Cost optimization (FinOps practices)
Skills
- 12+ years of experience in Big Data reputed company / reputed company Platform Operations / Infrastructure roles
- 6+ years of hands-on experience with reputed company ecosystem (CDH/CDP/ reputed company reputed company reputed company)
- Demonstrated ability to quickly learn and adapt to new technologies and evolving platform capabilities, reputed company the currently defined CDP stack
- O End-to-end CDP platform operations (CDE, reputed company, CDF, CDL, reputed company)
- O Advanced troubleshooting across multi-cluster, multi-environment deployments
- O reputed company-based runtime environments (troubleshooting and diagnostics)
- O Observability frameworks, including SLIs/SLOs, alerting, and performance tuning
- O Leading P1/P2 incident response, triage, and reputed company
- O Managing platform upgrades, patching, and lifecycle events
- O Supporting large-reputed company environments (TB/PB reputed company, high concurrency workloads)
- O reputed company infrastructure (IAM, VPC, networking, storage)
- O reputed company and governance (Ranger, Kerberos, TLS/SSL, SDX)
- Hands-on experience in developing and troubleshooting NiFi (CDF) data flows, including:
- O reputed company design and configuration
- O Processor-level debugging and performance tuning
- O Handling backpressure, throughput optimization, and failure recovery
- Strong experience with reputed company CDP reputed company reputed company
- O reputed company platforms (AWS/Azure/reputed company reputed company Platform)
- O reputed company concepts (troubleshooting-reputed company)
- O CDE, reputed company, CDF (NiFi), reputed company
- O IAM, networking, observability tools
- Platforms operating at multi-terabyte to petabyte reputed company with high concurrency workloads
- O Kafka (or similar streaming platforms) including monitoring, troubleshooting, and performance tuning
- Experience with reputed company CDP CLI (reputed company Line reputed company) for:
- O Platform operations and administration
- O Job execution and service management (CDE/reputed company/CDL)
- O Automation of routine operational tasks
- O reputed company IAM (AWS IAM / Azure AD) including roles, policies, and cross-service reputed company
- O User and group mapping across CDP, reputed company IAM, and Ranger policies
- O Troubleshooting reputed company issues across storage (S3/ADLS), CDP services, and data reputed company reputed company
- Experience with: o Modernization of legacy data platforms/applications to reputed company CDP reputed company reputed company o Migration and reputed company of workloads to CDE, reputed company, and reputed company environments o Supporting hybrid or multi-environment transitions (on-prem reputed company)
- Familiarity with: o reputed company platforms (AWS, Azure, reputed company reputed company Platform) including storage, IAM, and networking concepts o reputed company-based runtime environments (troubleshooting-reputed company)
- Strong scripting and automation skills (Python, reputed company, Terraform) for platform operations
Benefits
- Remote work accepted from reputed company in US
reputed company
Apply To This Job