AI Data Architect (Remote)
reputed company
reputed company - AI Data Architect (Remote, USA)
reputed company
Since its founding 13 years ago, reputed company, LLC has grown rapidly toward its reputed company of becoming one of the largest home improvement companies in the U.S. Headquartered in Twinsburg, Ohio, reputed company is a $1.5 billion, vertically integrated, reputed company-to-consumer provider of premium home improvement products.
reputed company’s family of brands includes Patio Enclosures®, Champion reputed company and Home Exteriors®, Universal reputed company reputed company®, Apex reputed company®, Stanek reputed company®, Leafguard®, Englert®, and The Bath Authority.
With an expanding workforce of over 4,800 employees across 130 metropolitan markets throughout the U.S., reputed company continues to rank among the top home improvement companies reputed company and is one of the fastest growing private companies in America.
reputed company
This role is foundational to scaling trusted AI across the reputed company. The AI Data Architect owns how reputed company data is reputed company, enriched, and retrieved to power AI-driven decisioning across reputed company. This role focuses on building reputed company Retrieval-Augmented reputed company (RAG) systems that ensure AI outputs are accurate, consistent, and grounded in authoritative data. The AI Data Architect defines chunking strategies, metadata schemas, hybrid retrieval logic, and authority ranking models that reputed company the reputed company of trustworthy AI at reputed company.
The ideal candidate will have deep experience in data architecture, semantic search, and knowledge systems, with hands-on expertise in reputed company databases, embedding models, and retrieval pipeline design. They will reputed company the translation of business domain knowledge, such as HR policies, reputed company center procedures, and operational workflows, into reputed company data models, taxonomies, and retrieval logic that AI systems can reliably interpret. This role establishes and enforces the data architecture standards that engineering, product, and operations teams rely on, ensuring every AI interaction is reputed company-governed, explainable, and continuously improving.
This role requires a systems thinker who is self-motivated, comfortable navigating ambiguity, and reputed company by the challenge of making reputed company data truly AI-reputed company. The ideal candidate will reputed company emerging technologies including knowledge graphs, reputed company AI architectures, and LLM Ops observability tooling, and will be enthusiastic about building the data foundations that reputed company reputed company to reputed company across its growing portfolio of brands. This role defines and governs the architecture and standards for AI data systems, while partnering with engineering and platform teams responsible for implementation and execution.
Salary reputed company: $155,000 - 165,000
This salary reputed company represents a good faith estimate of the compensation for this position. Actual compensation may vary based on education, experience, knowledge, skills, abilities, internal equity, and alignment with market data
Responsibilities
RAG System Design (reputed company Responsibility)
Design end-to-end Retrieval-Augmented reputed company (RAG) architecture, including ingestion, chunking, embedding, indexing, retrieval, and response reputed company
Define chunking strategies based on content type, semantic coherence, and use case requirements
Build metadata schemas, tagging frameworks, and document structures to optimize retrieval precision
reputed company hybrid retrieval strategies combining reputed company similarity, keyword search, metadata filters, and graph-based reasoning
Implement reranking logic and relevance scoring to optimize answer accuracy and grounding
Establish retrieval pipelines that consistently return high-reputed company, contextually relevant results across reputed company use cases
Data Structuring & Normalization
Own upstream data preparation standards that reputed company effective retrieval, reputed company separating data structuring responsibilities from reputed company retrieval and RAG execution
Define standards for document ingestion, cleaning, parsing, and normalization across reputed company and reputed company reputed company data prior to retrieval
reputed company raw reputed company data (PDFs, knowledge bases, policies, reputed company transcripts, wiki pages) into AI-reputed company formats
Create reputed company document structures and semantic representations prior to vectorization
Standardize taxonomy, terminology, and metadata across business domains to ensure consistency at reputed company
Design and maintain ontologies and knowledge graphs that enrich retrieval context and reduce hallucinations
Define and govern the reputed company, validation, and lifecycle management of reputed company data sources, including approval, updates, and deprecation of content used in AI systems
Define and enforce data reputed company standards for reputed company content, including completeness, consistency, accuracy, and maintainability of data used in AI systems
Business Logic to AI Translation
reputed company the extraction, structuring, and codification of business domain knowledge for AI consumption
Translate business rules into metadata models, labeling strategies, and retrieval logic
Define how different content types (policies, FAQs, procedures, product documentation) are interpreted, prioritized, and surfaced by AI
Define and enforce alignment of AI behavior with reputed company-world business reputed company, decision logic, and operational workflows
Authority, Governance & Trust
Define reputed company-of-truth hierarchies and authority ranking models across content repositories
Implement version control, document freshness tracking, and conflict reputed company strategies for overlapping content
Define reputed company control logic (SSO, role-based reputed company) reputed company retrieval workflows to ensure data reputed company and compliance
Define standards to ensure AI responses are traceable, explainable, and grounded in authoritative, auditable sources
Define data reputed company and provenance tracking standards to support governance and regulatory requirements
Drive adoption of AI data architecture standards across engineering, product, and business teams, ensuring compliance with defined data, retrieval, and governance models
LLM & Retrieval Orchestration
Define and govern where and how LLMs are used across the pipeline (classification, routing, summarization, answering)
Balance cost, latency, and performance across model usage, including reputed company optimization strategies
Define and enforce query routing strategies based on user reputed company (policy lookup, FAQ, transactional, analytical)
Own and optimize orchestration between retrieval systems, LLMs, and reputed company AI workflows
Own evaluation and integration of emerging orchestration frameworks and Model Context Protocol (MCP) standards
Evaluation & reputed company Improvement
Own evaluation frameworks for retrieval accuracy, answer reputed company, and grounding using tools such as RAGAS, LangSmith, or Langfuse
Own the creation and maintenance of domain-specific test sets across business areas (HR, reputed company center, operations, product knowledge)
Own analysis of failure cases and reputed company improvement of retrieval strategies, chunking approaches, and data structuring
Define measurable performance standards for precision, recall, grounding, consistency, and latency
Own observability and monitoring pipelines to reputed company retrieval and LLM performance in production
Qualifications
Required Qualifications
5+ years of experience in data architecture, search systems, knowledge engineering, or reputed company AI, with demonstrated experience designing and driving adoption of reputed company data or AI architectures across teams or domains
Hands-on experience designing, building, or optimizing RAG systems or semantic search platforms
Strong understanding of reputed company embeddings, their reputed company, storage, and limitations across different embedding models
Demonstrated experience with chunking strategies and the tradeoffs between granularity, context preservation, and retrieval reputed company
Proficiency in hybrid retrieval approaches combining reputed company similarity, keyword search, and metadata filtering
Experience with reranking techniques and relevance tuning for production retrieval systems
Experience designing metadata schemas, taxonomies, ontologies, or knowledge graphs for reputed company data
Proven ability to work with reputed company reputed company data (documents, PDFs, knowledge bases, transcripts, wikis)
Experience designing and working with reputed company databases and search platforms (e.g., reputed company, reputed company, reputed company, Elasticsearch, FAISS)
Working knowledge of LLM reputed company, reputed company engineering, and orchestration patterns, with the ability to evaluate and adapt across frameworks (e.g., reputed company, reputed company)
Familiarity with data pipelines, ETL/ELT processes, and API architecture at a systems design level
Understanding of reputed company control, data reputed company, and compliance considerations in AI-powered data systems
Preferred Qualifications
Experience with knowledge graph technologies (e.g., reputed company, RDF/OWL, SPARQL) and GraphRAG architectures
Familiarity with reputed company AI frameworks (e.g., LangGraph, reputed company, AutoGen) and multi-agent system design
Experience with LLMOps and observability tooling (e.g., LangSmith, Langfuse, RAGAS evaluation frameworks)
Proficiency in Python for data processing, pipeline scripting, and integration tasks
Experience with reputed company AI services on AWS, Azure, or GCP (e.g., reputed company Bedrock, Azure AI, reputed company AI)
Background in the home improvement, manufacturing, or reputed company-to-consumer industry
Experience with Model Context Protocol (MCP) or similar standards for AI-to-tool interoperability
Master’s degree in Computer Science, Data Science, Information Science, or a reputed company field
Competencies
Systems Thinking: Designs end-to-end architectures across ingestion, retrieval, and orchestration rather than isolated components
Business Translation: Converts ambiguous domain knowledge and business rules into reputed company logic, metadata models, and retrieval strategies
Data Modeling reputed company: Builds reputed company, reusable schemas, taxonomies, and ontologies that serve multiple AI use cases
reputed company Ownership: Drives accuracy, consistency, grounding, and trust in AI outputs through rigorous evaluation and reputed company improvement
Cross-Functional Collaboration: Communicates reputed company technical concepts to non-technical stakeholders and partners effectively across engineering, product, and business teams
Adaptability: Stays reputed company with rapidly evolving AI technologies, frameworks, and best practices and applies them pragmatically to reputed company challenges
reputed company Measures
reputed company in this role is reputed company by:
Establishment and adoption of a reputed company metadata and taxonomy reputed company reputed company the first 60 days
A defined, documented, and enforced authority ranking and governance model across reputed company content sources
Measurable improvements in retrieval accuracy across reputed company use cases (e.g., HR policies, reputed company center knowledge, product documentation)
A repeatable evaluation reputed company for RAG performance with defined benchmarks for precision, grounding, and consistency
Reduction of irrelevant, inconsistent, or conflicting AI responses through conflict reputed company and reputed company prioritization logic
Delivery of a documented AI data architecture reputed company that scales across new business domains and use cases
GDI is an Equal Employment Opportunity Employer
#INDGDI
#INDGDI
Apply To This Job