Generative AI Architect
- Los Angeles, United States
Posted 15 hours ago
About the role
Generative AI Architect
C2CContractW2
Los Angeles, CA (Remote)
Visa: US CitizenGreen CardGC-EADH1BH4-EADTNEAD
Job Summary: We are seeking an experienced Generative AI Architect to define and drive the organization's enterprise Generative AI strategy and architecture. The role will focus on designing scalable RAG platforms, multi-agent AI systems, cost-efficient LLM inference infrastructure, and enterprise AI governance frameworks. The ideal candidate will have strong experience in Generative AI, LLM architecture, RAG, vector databases, model serving, cloud AI platforms, and enterprise AI governance, along with the ability to provide technical leadership across AI initiatives. Responsibilities Define the client's multi-year Generative AI blueprint, including technology strategy, architecture roadmap, and platform capabilities. Develop model selection strategies for open-source and proprietary models based on performance, cost, scalability, security, and business requirements. Establish governance standards for GenAI frameworks, models, development practices, and enterprise adoption. Design enterprise-grade RAG architectures using high-throughput vector databases such as Pinecone, Milvus, and pgvector, integrating enterprise data sources including video telemetry, content metadata libraries, and customer data platforms. Architect multi-agent orchestration systems using LangGraph, AutoGen, and Semantic Kernel for complex, multi-step workflows. Define scalable architecture patterns for LLM applications, retrieval systems, agents, tools, and AI workflows. Design cost-efficient and low-latency LLM inference pipelines using AWS Bedrock, Azure OpenAI, and Google Cloud Vertex AI. Define scalable model deployment and inference strategies, including GPU utilization and resource optimization. Implement model optimization techniques including quantization, caching layers, token usage optimization, efficient inference, and workload optimization. Establish strategies to manage and optimize enterprise AI compute and operational costs and evaluate model-serving technologies such as NVIDIA Triton and vLLM. Establish GenAI guardrails and responsible AI frameworks covering hallucination detection, toxicity filtering, DLP, output validation, and secure LLM interactions. Define governance standards for model usage, data access, AI application development, and production deployment. Ensure AI solutions comply with DRM, CCPA/user privacy, and copyright requirements, including appropriate controls for video scripts, captions, user logs, and other enterprise content and data. Partner with Product Managers, Data Engineering teams, and Content Operations to translate business requirements into viable GenAI initiatives. Evaluate AI use cases, define technical architectures and implementation approaches, mentor GenAI Engineers, and conduct architecture and code reviews. Create and maintain Architecture Decision Records (ADRs) and promote engineering best practices across GenAI development and deployment teams. Required Skills Master's or Ph.D. in Computer Science, Artificial Intelligence, Machine Learning, or a related technical field. Equivalent practical experience with demonstrated expertise in Generative AI architecture may also be considered. 10+ years of experience in software engineering, data architecture, machine learning, or related technical disciplines, including at least 3+ years specifically leading Generative AI architecture, LLM deployment, and RAG systems at scale. Proven experience designing enterprise-grade AI platforms and production LLM solutions, providing technical leadership and mentoring AI engineering teams. Deep expertise in Python and strong experience with PyTorch and/or TensorFlow. Hands-on experience with LangChain and/or LlamaIndex. Strong knowledge of enterprise RAG architectures and vector databases, including Pinecone, Milvus, and pgvector. Experience with LLM model serving, including NVIDIA Triton and/or vLLM. Strong experience with AWS Bedrock, Azure OpenAI, and Google Cloud Vertex AI. Strong understanding of LLM inference, embeddings, vector search, prompt engineering, model optimization, and AI application architecture. Experience designing scalable and reliable AI systems for enterprise workloads. Strong capabilities in Generative AI Architecture, LLM Architecture & Deployment, RAG Architecture, Multi-Agent Systems, Vector Search & Retrieval, Cloud AI Platforms, Model Serving & Inference Optimization, GPU & Cost Optimization, AI Security & Governance, Responsible AI, and Technical Leadership & Mentoring. Desired Skills Experience in recommendation engines, conversational search, multi-modal AI systems, text-to-video applications, audio-to-text applications, content intelligence, AI-powered media workflows, or large-scale enterprise AI platforms.
Refer
