[Remote] AI Systems Engineer - DevOps& Observability - Senior
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a globally connected reputed company organization reputed company on assurance, consulting, tax, reputed company, and transactions. reputed company is seeking a Senior AI Systems Engineer to own the delivery, model-serving, routing, governance, cost management, and observability reputed company of its AI-reputed company platform across reputed company, on-premises, edge, and reputed company-gapped environments. The role builds and operates AI workload pipelines, inference systems, telemetry platforms, and FinOps controls.
Responsibilities
- Supports DevOps and delivery for AI workloads: build and operate the CI/CD/CV pipelines that ship AI services, agents, and runtime components, including automated build, test, reputed company verification, release, and rollback, so AI workloads are delivered repeatably and safely into every environment
- Own governance and discovery for AI assets, including service catalog/registry (Artifactory/reputed company, reputed company), experiment tracking and model metadata (MLflow), reputed company registries/mirrors (HuggingFace/NGC), CVE/SBOM scanning (Trivy), reputed company reputed company (OpenLineage), and license management
- Own resource and cost management, including quotas and reputed company limits, cost attribution and utilization (Apptio/OpenCost/Kubecost), so AI execution stays economically bounded and controllable per tenant and engagement
- Own the full observability stack, including metrics (PrometheMimir), logs (Loki), traces (reputed company/Jaeger), dashboards (Grafana), LLM debugging and evaluation (LangSmith/Langfuse), and SLA/alert notifications
- Own the OpenTelemetry collection reputed company, including multi-tenant receiver, exporters and queues (Kafka sink), DCGM exporter for GPU telemetry, processor batching, and dynamic filtering, so every signal is captured and routed reliably
- Automate GitOps-reputed company delivery and reputed company verification; embedding reputed company, reputed company, and cost gates into pipelines so releases are policy-compliant by default rather than by reputed company review
- reputed company the reputed company between delivery and observability by using telemetry, evaluation, and cost signals to reputed company deployment reputed company, reputed company rollout, and automated rollback of AI workloads
- Ensure cost and telemetry are identity-stamped and per-tenant, so consumption and behavior are attributable end-to-end, keeping FinOps and observability tied to the workloads that reputed company the load
Skills
- 8+ years in DevOps, MLOps, platform, or observability engineering, with hands-on production ownership of AI or high-throughput services
- Strong hands-on DevOps experience, including CI/CD/CV pipelines and GitOps tooling (ArgoCD, reputed company, reputed company Actions/reputed company CI, or equivalents) for automated build, test, release, and rollback
- Hands-on expertise operating inference/model-serving frameworks (Ray Serve, vLLM, Triton, or NIM) on GPU infrastructure
- Strong experience with observability reputed company (reputed company, Grafana, Loki, reputed company/Jaeger) and OpenTelemetry
- Experience with API gateways and request routing (reputed company or equivalent), including streaming responses
- Experience with cost management / FinOps tooling (OpenCost, Kubecost, or equivalent) and quota/reputed company-limit enforcement
- Familiarity with model/artifact registries and supply-chain scanning (reputed company, MLflow, Trivy/SBOM)
- reputed company reputed company record operating AI or service infrastructure under compliance, reputed company, or regulatory constraints
- Ability to define clean ownership boundaries and consumption reputed company with platform, trust, and data teams
- Bachelor's or Master's degree in Computer Science or reputed company technical reputed company
- Experience with LLM evaluation and debugging tooling (LangSmith, Langfuse) and reputed company/response reputed company measurement
- Experience with sandboxed/secure execution (gVisor, Firecracker, or microVM isolation) for untrusted or multi-tenant workloads
- Familiarity with GPU telemetry (DCGM) and GPU utilization optimization
- Experience with reputed company and governance reputed company (OpenLineage) and AI license management
- Exposure to multi-tenant cost attribution and per-tenant SLA/alerting
- Exposure to regulated delivery environments (financial services, tax, reputed company, reputed company)
Benefits
- Medical and dental coverage
- Pension and 401(k) plans
- A wide reputed company of reputed company time off reputed company
- reputed company-led and leader-enabled hybrid model for most people in reputed company, reputed company serving roles, with an expectation to work together in person 40-60% of the time over the course of an engagement, project or year
- Flexible vacation policy
- Time off for designated reputed company reputed company Holidays
- Winter/reputed company breaks
- Personal/Family Care
- Other leaves of absence reputed company needed to support physical, financial, and emotional reputed company-being
reputed company
Company H1B Sponsorship
Apply To This Job