# Owl Alpha — AI Infrastructure Architecture

The vision: Z840 as the AI nerve centre. Claude as the persistent orchestrator. All models, all agents, all tools centralised on one server — accessible from any device, always on, always context-aware.

## The Goal

- Claude has persistent memory across every session from any device (phone, laptop, desktop) - Every new chat bootstraps in <30 seconds with full context via CLAUDE.md + wiki - Specialised agents handle domains — SEO, Divi, infrastructure, home automation, design, customer service - Local Mixture of Experts routes tasks to the right model — no unnecessary API costs - Everything runs offline — client data never leaves the network (GDPR-friendly) - The wiki grows as the living brain — updated after every session

## Hardware Layer

### Z840 — AI Nerve Centre - Dual Xeon - 2x RTX 3060 arriving Saturday (24GB VRAM combined) - Optional 3rd RTX 3060 = 36GB VRAM total - 36GB VRAM unlocks: 70B models fully in VRAM, no CPU offload, full inference speed

## Model Layer

### Ollama — Local Model Serving Serves all local LLMs via unified API. Handles VRAM management, model switching, quantisation.

### Model Router Claude decides, n8n or config executes. Task → optimal model:

Task Type Model
———–——-
Code, debugging, scripting DeepSeek
Reasoning, planning, long-context Mixtral / Llama 3.3 70B
Quick, high-volume, lightweight Phi-4 / Gemma 3
Image generation Flux / SDXL
Semantic search / embeddings nomic-embed

### Models to install - [ ] DeepSeek (code/logic) - [ ] Mixtral 8x7B or Llama 3.3 70B (reasoning) - [ ] Phi-4 or Gemma 3 (lightweight) - [ ] nomic-embed-text (embeddings) - [ ] Flux or SDXL (image gen)

## Orchestration Layer

### Claude — Orchestrator & Router - Reads CLAUDE.md + wiki on every session boot - Routes tasks to the right model or agent - Interprets results, makes decisions - Writes findings back to wiki before session ends - Accessible from any device via Tailscale + Open WebUI

## Memory Layer

### Wiki + CLAUDE.md — Persistent Memory - CLAUDE.md = session bootstrap (who, what, where, current state) - Wiki = long-term brain (infrastructure, decisions, client info, project state) - Embeddings via nomic-embed = semantic search across all wiki pages - Lives on Hetzner VPS, accessible to Z840 and all agents

## Access Layer

### Tailscale — Secure Mesh Phone, laptop, desktop, all servers on one private network. No port forwarding.

### Open WebUI Already installed. Front door to all models and agents from any browser on Tailscale.

## Automation Layer

### n8n — Agent Workflows Already running. Triggers agents on events, schedules, webhooks. Routes tasks between agents.

### MCP Servers One tool layer shared by all agents: - OPNsense (firewall) - CyberPanel (hosting) - Cloudflare (DNS) - Google Drive - Home Assistant - MainWP (WordPress) - Hetzner (VPS)

## Agent Layer

Agent Domain Primary Models
——-——–—————-
Hermes Infrastructure monitoring Phi/Gemma (lightweight, always on)
SEO Agent Audits, content, rankings DeepSeek (technical), Mixtral (strategy)
Divi Agent Design & site builds DeepSeek (code), Flux (images)
HA Agent Home Assistant automations Phi/Gemma (speed)
Design Agent Brand assets, imagery Flux / SDXL
CS Agents Customer service (hireable) Client-trained, local, GDPR-friendly

## Build Order

### Phase 1 — Saturday (GPU day) - [ ] Install 2x RTX 3060 in Z840 - [ ] Install Ollama - [ ] Pull first models: DeepSeek, Phi-4, nomic-embed - [ ] Test inference speed - [ ] Decide on 3rd GPU

### Phase 2 — Model layer - [ ] Pull Mixtral or Llama 70B - [ ] Install Flux/SDXL - [ ] Configure model router config - [ ] Wire Ollama into Open WebUI

### Phase 3 — Memory & bootstrap - [ ] Formalise CLAUDE.md structure - [ ] Set up embedding pipeline for wiki semantic search - [ ] Test full session bootstrap from phone

### Phase 4 — Agents - [ ] Migrate Hermes to Z840 (already on kaburuaibox) - [ ] Build SEO agent - [ ] Build Divi agent - [ ] Build HA agent - [ ] Build Design agent

### Phase 5 — Customer service agents - [ ] Design white-label CS agent framework - [ ] First client pilot

## Diagram

See owl-alpha-architecture.jsx (interactive React diagram)

```mermaid graph TD

  Z840[Z840 — AI Nerve Centre<br/>2-3x RTX 3060] --> GPU[GPU Layer<br/>24-36GB VRAM]
  Z840 --> Ollama[Ollama<br/>Model Serving]
  GPU --> Ollama
  GPU --> ImgGen[Flux / SDXL<br/>Image Gen]
  Ollama --> DeepSeek[DeepSeek<br/>Code & Logic]
  Ollama --> Mixtral[Mixtral / Llama<br/>Reasoning]
  Ollama --> Phi[Phi-4 / Gemma<br/>Lightweight]
  Ollama --> Embed[Embeddings<br/>nomic-embed]
  Claude[Claude<br/>Orchestrator] --> Router[Model Router]
  Router --> Ollama
  Router --> ImgGen
  Claude --> Wiki[Wiki + CLAUDE.md<br/>Persistent Memory]
  Wiki --> Embed
  Claude --> Tailscale[Tailscale<br/>Secure Mesh]
  Tailscale --> OpenWebUI[Open WebUI]
  Claude --> Hermes & SEO & Divi & HA & Design & CS
  n8n --> Hermes & SEO & Divi & HA & Design & CS
  MCP --> Hermes & SEO & HA

```

## Notes

- The name Hermes running on Hetzner is unintentional mythology — messenger god as monitoring agent - Owl Alpha = the overarching AI stack name - All client data stays on-premises — strong GDPR position for selling CS agents - Z840 is powerful enough to run everything; 3rd GPU is insurance not necessity