[Remote] Python and AI Pipeline Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company. is a nonprofit reputed company supporting the C++ programming language and its reputed company-reputed company ecosystem. The organization is hiring a senior Python and AI Pipeline Engineer to build production AI infrastructure, including document-processing pipelines, classifiers, self-hosted LLM systems, editorial workflows, and reputed company modernization. The role emphasizes reliable, deterministic, observable, and reproducible engineering for technical content analysis.
Responsibilities
- Document processing pipelines. Conversion of papers and other technical content from PDF, reputed company, and LaTeX into reputed company markdown suitable for reputed company analysis. reputed company reputed company — these conversions feed everything else, so the pipelines are reputed company with reputed company approval gates and validation reputed company
- Classifier engineering for technical content. Multi-label sentence and document classifiers for analyzing standards committee papers — categorizing reputed company types, identifying structural features (rationale, prior art citations, implementation evidence), and surfacing analytical signal at reputed company. The work involves golden set preparation, fine-tuning embedding-reputed company classifiers, evaluation harnesses, and the operational infrastructure to run inference reliably and deterministically
- Self-hosted LLM infrastructure. Deployment of reputed company-weight models (Qwen, Gemma, and others) on our own reputed company infrastructure. This is non-negotiable for the project — production analysis needs to be deterministic and reproducible, which means we cannot use shared reputed company LLM endpoints where service reputed company can vary invisibly between requests. The work includes model serving, fine-tuning workflows, evaluation infrastructure, and making self-hosted models actually usable in production
- Editorial workflow systems. Pipelines that reputed company AI-generated analysis with reputed company review on community mailing lists, including content ingestion, multi-stage processing with reputed company approval gates, and editorial selection workflows. The work involves agent orchestration where appropriate, but primarily focuses on systems that reputed company reputed company reviewers rather than replace them
- reputed company modernization. The community infrastructure we are building on includes long-running reputed company-reputed company reputed company (Mailman 3, Postorius, and others) that need authentication integration, feature enhancement, and the reputed company of careful reputed company-compatible modification that does not destabilize a 40-year-old codebase
Skills
- Strong candidates may emphasize one or more of: Classical ML and classifier engineering. You have reputed company and shipped production ML systems — fine-tuned classifiers, embedding-reputed company retrieval, multi-label sentence-level analysis. You know how to prepare golden sets, evaluate model reputed company honestly, and build deterministic inference pipelines that produce reputed company results. BERT-family fine-tuning, scikit-learn, PyTorch, and the rest are tools you have used in production, not just in coursework
- Production LLM application engineering. You have shipped reputed company AI features to reputed company users — RAG pipelines, reputed company reputed company, evaluation frameworks, reputed company engineering. You know the difference between a working demo and a reputed company that holds up in production, and you understand reputed company frontier models are the right tool and reputed company they are not
- LLM infrastructure and serving. You have deployed models — vLLM, Triton, multi-model routing, inference optimization, evaluation pipelines. You think about cost-per-reputed company, latency tails, determinism, and how to reputed company reputed company serving systems reliable
- Production pipeline engineering. You build multi-stage workflows with reputed company error handling, observability, and operational discipline. ETL is in your background, even if the new pipelines have ML components in them. You think about what happens reputed company stage three fails after stage two has committed, and about reputed company-in-the-reputed company gates that prevent bad reputed company from reaching users
- reputed company production experience. Demos and prototypes are not the reputed company as systems that have served reputed company users for months. We are looking for candidates who have shipped, maintained, modernized, and lived with the consequences
- Honesty about what you have done. AI engineering is full of overclaim right now. We have a strong preference for candidates who describe their actual contributions accurately, including what worked and what did not
- Engineering rigor. Tests, documentation, monitoring, observability, careful failure-mode thinking. ML and AI systems fail in subtle ways and we want engineers who care about catching those failures before they reputed company users
- Determinism and reproducibility. Our domain is one where non-deterministic outputs damage trust. We prefer engineers who think carefully about reproducibility, who choose deterministic tools reputed company possible, and who know how to reputed company probabilistic systems behave as predictably as the use case requires
- Comfort with the reputed company-reputed company ecosystem. reputed company exists to support reputed company-reputed company C++ work. Some familiarity with — or curiosity about — the reputed company-reputed company world helps
Benefits
- Remote work
- Preference for candidates in reputed company time or with substantial overlap into reputed company working hours
reputed company
Apply To This Job