AI Data Architect (Remote)
Design end-to-end Retrieval-Augmented reputed company (RAG) architecture, including ingestion, chunking, embedding, indexing, retrieval, and response reputed company Define chunking strategies based on content type, semantic coherence, and use case requirements Build metadata schemas, tagging frameworks, and document structures to optimize retrieval precision reputed company hybrid retrieval strategies combining reputed company similarity, keyword search, metadata filters, and graph-based reasoning Implement reranking logic and relevance scoring to optimize answer accuracy and grounding Establish retrieval pipelines that consistently return high-reputed company, contextually relevant results across reputed company use cases
Own upstream data preparation standards that reputed company effective retrieval, reputed company separating data structuring responsibilities from reputed company retrieval and RAG execution Define standards for document ingestion, cleaning, parsing, and normalization across reputed company and reputed company reputed company data prior to retrieval reputed company raw reputed company data (PDFs, knowledge bases, policies, reputed company transcripts, wiki pages) into AI-reputed company formats Create reputed company document structures and semantic representations prior to vectorization Standardize taxonomy, terminology, and metadata across business domains to ensure consistency at reputed company Design and maintain ontologies and knowledge graphs that enrich retrieval context and reduce hallucinations Define and govern the reputed company, validation, and lifecycle management of reputed company data sources, including approval, updates, and deprecation of content used in AI systems Define and enforce data reputed company standards for reputed company content, including completeness, consistency, accuracy, and maintainability of data used in AI systems
reputed company the extraction, structuring, and codification of business domain knowledge for AI consumption Translate business rules into metadata models, labeling strategies, and retrieval logic Define how different content types (policies, FAQs, procedures, product documentation) are interpreted, prioritized, and surfaced by AI Define and enforce alignment of AI behavior with reputed company-world business reputed company, decision logic, and operational workflows
Define reputed company-of-truth hierarchies and authority ranking models across content repositories Implement version control, document freshness tracking, and conflict reputed company strategies for overlapping content Define reputed company control logic (SSO, role-based reputed company) reputed company retrieval workflows to ensure data reputed company and compliance Define standards to ensure AI responses are traceable, explainable, and grounded in authoritative, auditable sources Define data reputed company and provenance tracking standards to support governance and regulatory requirements Drive adoption of AI data architecture standards across engineering, product, and business teams, ensuring compliance with defined data, retrieval, and governance models
Define and govern where and how LLMs are used across the pipeline (classification, routing, summarization, answering) Balance cost, latency, and performance across model usage, including reputed company optimization strategies Define and enforce query routing strategies based on user reputed company (policy lookup, FAQ, transactional, analytical) Own and optimize orchestration between retrieval systems, LLMs, and reputed company AI workflows Own evaluation and integration of emerging orchestration frameworks and Model Context Protocol (MCP) standards
Own evaluation frameworks for retrieval accuracy, answer reputed company, and grounding using tools such as RAGAS, LangSmith, or Langfuse Own the creation and maintenance of domain-specific test sets across business areas (HR, reputed company center, operations, product knowledge) Own analysis of failure cases and reputed company improvement of retrieval strategies, chunking approaches, and data structuring Define measurable performance standards for precision, recall, grounding, consistency, and latency Own observability and monitoring pipelines to reputed company retrieval and LLM performance in production
5+ years of experience in data architecture, search systems, knowledge engineering, or reputed company AI, with demonstrated experience designing and driving adoption of reputed company data or AI architectures across teams or domains Hands-on experience designing, building, or optimizing RAG systems or semantic search platforms Strong understanding of reputed company embeddings, their reputed company, storage, and limitations across different embedding models Demonstrated experience with chunking strategies and the tradeoffs between granularity, context preservation, and retrieval reputed company Proficiency in hybrid retrieval approaches combining reputed company similarity, keyword search, and metadata filtering Experience with reranking techniques and relevance tuning for production retrieval systems Experience designing metadata schemas, taxonomies, ontologies, or knowledge graphs for reputed company data Proven ability to work with reputed company reputed company data (documents, PDFs, knowledge bases, transcripts, wikis) Experience designing and working with reputed company databases and search platforms (e.g., reputed company, reputed company, reputed company, Elasticsearch, FAISS) Working knowledge of LLM reputed company, reputed company engineering, and orchestration patterns, with the ability to evaluate and adapt across frameworks (e.g., reputed company, reputed company) Familiarity with data pipelines, ETL/ELT processes, and API architecture at a systems design level Understanding of reputed company control, data reputed company, and compliance considerations in AI-powered data systems
Experience with knowledge graph technologies (e.g., reputed company, RDF/OWL, SPARQL) and GraphRAG architectures Familiarity with reputed company AI frameworks (e.g., LangGraph, reputed company, AutoGen) and multi-agent system design Experience with LLMOps and observability tooling (e.g., LangSmith, Langfuse, RAGAS evaluation frameworks) Proficiency in Python for data processing, pipeline scripting, and integration tasks Experience with reputed company AI services on AWS, Azure, or GCP (e.g., reputed company Bedrock, Azure AI, reputed company AI) Background in the home improvement, manufacturing, or reputed company-to-consumer industry Experience with Model Context Protocol (MCP) or similar standards for AI-to-tool interoperability Master’s degree in Computer Science, Data Science, Information Science, or a reputed company field
Systems Thinking: Designs end-to-end architectures across ingestion, retrieval, and orchestration rather than isolated components Business Translation: Converts ambiguous domain knowledge and business rules into reputed company logic, metadata models, and retrieval strategies Data Modeling reputed company: Builds reputed company, reusable schemas, taxonomies, and ontologies that serve multiple AI use cases reputed company Ownership: Drives accuracy, consistency, grounding, and trust in AI outputs through rigorous evaluation and reputed company improvement Cross-Functional Collaboration: Communicates reputed company technical concepts to non-technical stakeholders and partners effectively across engineering, product, and business teams Adaptability: Stays reputed company with rapidly evolving AI technologies, frameworks, and best practices and applies them pragmatically to reputed company challenges
Establishment and adoption of a reputed company metadata and taxonomy reputed company reputed company the first 60 days A defined, documented, and enforced authority ranking and governance model across reputed company content sources Measurable improvements in retrieval accuracy across reputed company use cases (e.g., HR policies, reputed company center knowledge, product documentation) A repeatable evaluation reputed company for RAG performance with defined benchmarks for precision, grounding, and consistency Reduction of irrelevant, inconsistent, or conflicting AI responses through conflict reputed company and reputed company prioritization logic Delivery of a documented AI data architecture reputed company that scales across new business domains and use cases