Air-Gapped AI RAG Engine: A Responsible AI (RAI) Architecture
A secure, locally hosted Retrieval-Augmented Generation (RAG) pipeline built to bridge commercial AI agility with rigorous enterprise and federal governance standards.
Executive Summary
Designed and deployed a fully air-gapped Large Language Model (LLM) architecture utilizing agentic workflows to synthesize proprietary documents without transmitting sensitive data to external APIs. Drawing on experience managing complex, high-stakes deployments in disconnected and secure environments, this project demonstrates how to rapidly field commercial-grade AI while maintaining zero-trust data privacy. It serves as a blueprint for organizations—from defense-tech startups to commercial hyperscalers—seeking to deploy AI responsibly and securely.
Technical Architecture & Data Foundation
- ▷
LLM Engine: Llama 3.2 running locally via Ollama for zero-latency, private inference, aligning with modern decentralized computing needs.
- ▷
Vector Database: ChromaDB utilized for generating, storing, and querying document embeddings locally, ensuring high-quality foundational data management.
- ▷
Orchestration: LangChain framework employed for multi-document vector retrieval and recursive character text splitting.
- ▷
Embeddings (Nomic): Implemented local Nomic embeddings to transform unstructured text into high-dimensional vectors efficiently on consumer hardware.
TPM Impact: Governance, Trust, & Continuous Deployment
As a Technical Program Manager, successfully scaling AI requires moving beyond basic implementation and establishing a "journey to trust." This pipeline was engineered with strict adherence to established AI governance frameworks, including the NIST AI RMF and the DoD's Responsible AI (RAI) guidelines.
- ▷
Zero-Trust Security & Compliance: Eliminated data leakage risks by processing all queries and documents locally, meeting strict defense-grade security protocols.
- ▷
NIST AI RMF Alignment: Mapped system architecture and data handling processes directly to the NIST AI Risk Management Framework to establish baseline governance.
- ▷
Cloud Cost Avoidance: Shifted inference and embedding workloads away from metered cloud APIs, resulting in a 100% reduction in per-token operational cloud costs.
- ▷
Rapid Prototyping & Delivery: Accelerated the discovery-to-deployment lifecycle by utilizing open-source models, allowing for rapid iteration without procurement blockers.
Source Code & CI/CD Pipeline
The full Python architecture and continuous deployment configuration are available for review.
View Repository on GitHub →