←Back to Projects

Air-Gapped AI RAG Engine: A Responsible AI (RAI) Architecture

A secure, locally hosted Retrieval-Augmented Generation (RAG) pipeline built to bridge commercial AI agility with rigorous enterprise and federal governance standards.

Executive Summary

Designed and deployed a fully air-gapped Large Language Model (LLM) architecture utilizing agentic workflows to synthesize proprietary documents without transmitting sensitive data to external APIs. Drawing on experience managing complex, high-stakes deployments in disconnected and secure environments, this project demonstrates how to rapidly field commercial-grade AI while maintaining zero-trust data privacy. It serves as a blueprint for organizations—from defense-tech startups to commercial hyperscalers—seeking to deploy AI responsibly and securely.

Technical Architecture & Data Foundation

  • ▷

    LLM Engine: Llama 3.2 running locally via Ollama for zero-latency, private inference, aligning with modern decentralized computing needs.

  • ▷

    Vector Database: ChromaDB utilized for generating, storing, and querying document embeddings locally, ensuring high-quality foundational data management.

  • ▷

    Orchestration: LangChain framework employed for multi-document vector retrieval and recursive character text splitting.

  • ▷

    Embeddings (Nomic): Implemented local Nomic embeddings to transform unstructured text into high-dimensional vectors efficiently on consumer hardware.

TPM Impact: Governance, Trust, & Continuous Deployment

As a Technical Program Manager, successfully scaling AI requires moving beyond basic implementation and establishing a "journey to trust." This pipeline was engineered with strict adherence to established AI governance frameworks, including the NIST AI RMF and the DoD's Responsible AI (RAI) guidelines.

  • ▷

    Zero-Trust Security & Compliance: Eliminated data leakage risks by processing all queries and documents locally, meeting strict defense-grade security protocols.

  • ▷

    NIST AI RMF Alignment: Mapped system architecture and data handling processes directly to the NIST AI Risk Management Framework to establish baseline governance.

  • ▷

    Cloud Cost Avoidance: Shifted inference and embedding workloads away from metered cloud APIs, resulting in a 100% reduction in per-token operational cloud costs.

  • ▷

    Rapid Prototyping & Delivery: Accelerated the discovery-to-deployment lifecycle by utilizing open-source models, allowing for rapid iteration without procurement blockers.

Source Code & CI/CD Pipeline

The full Python architecture and continuous deployment configuration are available for review.

View Repository on GitHub →