# Owl Alpha — AI Infrastructure Architecture
The vision: Z840 as the AI nerve centre. Claude as the persistent orchestrator. All models, all agents, all tools centralised on one server — accessible from any device, always on, always context-aware.
## The Goal
- Claude has persistent memory across every session from any device (phone, laptop, desktop) - Every new chat bootstraps in <30 seconds with full context via CLAUDE.md + wiki - Specialised agents handle domains — SEO, Divi, infrastructure, home automation, design, customer service - Local Mixture of Experts routes tasks to the right model — no unnecessary API costs - Everything runs offline — client data never leaves the network (GDPR-friendly) - The wiki grows as the living brain — updated after every session
—
## Hardware Layer
### Z840 — AI Nerve Centre - Dual Xeon - 2x RTX 3060 arriving Saturday (24GB VRAM combined) - Optional 3rd RTX 3060 = 36GB VRAM total - 36GB VRAM unlocks: 70B models fully in VRAM, no CPU offload, full inference speed
—
## Model Layer
### Ollama — Local Model Serving Serves all local LLMs via unified API. Handles VRAM management, model switching, quantisation.
### Model Router Claude decides, n8n or config executes. Task → optimal model:
| Task Type | Model |
| ———– | ——- |
| Code, debugging, scripting | DeepSeek |
| Reasoning, planning, long-context | Mixtral / Llama 3.3 70B |
| Quick, high-volume, lightweight | Phi-4 / Gemma 3 |
| Image generation | Flux / SDXL |
| Semantic search / embeddings | nomic-embed |
### Models to install - [ ] DeepSeek (code/logic) - [ ] Mixtral 8x7B or Llama 3.3 70B (reasoning) - [ ] Phi-4 or Gemma 3 (lightweight) - [ ] nomic-embed-text (embeddings) - [ ] Flux or SDXL (image gen)
—
## Orchestration Layer
### Claude — Orchestrator & Router - Reads CLAUDE.md + wiki on every session boot - Routes tasks to the right model or agent - Interprets results, makes decisions - Writes findings back to wiki before session ends - Accessible from any device via Tailscale + Open WebUI
—
## Memory Layer
### Wiki + CLAUDE.md — Persistent Memory - CLAUDE.md = session bootstrap (who, what, where, current state) - Wiki = long-term brain (infrastructure, decisions, client info, project state) - Embeddings via nomic-embed = semantic search across all wiki pages - Lives on Hetzner VPS, accessible to Z840 and all agents
—
## Access Layer
### Tailscale — Secure Mesh Phone, laptop, desktop, all servers on one private network. No port forwarding.
### Open WebUI Already installed. Front door to all models and agents from any browser on Tailscale.
—
## Automation Layer
### n8n — Agent Workflows Already running. Triggers agents on events, schedules, webhooks. Routes tasks between agents.
### MCP Servers One tool layer shared by all agents: - OPNsense (firewall) - CyberPanel (hosting) - Cloudflare (DNS) - Google Drive - Home Assistant - MainWP (WordPress) - Hetzner (VPS)
—
## Agent Layer
| Agent | Domain | Primary Models |
| ——- | ——– | —————- |
| Hermes | Infrastructure monitoring | Phi/Gemma (lightweight, always on) |
| SEO Agent | Audits, content, rankings | DeepSeek (technical), Mixtral (strategy) |
| Divi Agent | Design & site builds | DeepSeek (code), Flux (images) |
| HA Agent | Home Assistant automations | Phi/Gemma (speed) |
| Design Agent | Brand assets, imagery | Flux / SDXL |
| CS Agents | Customer service (hireable) | Client-trained, local, GDPR-friendly |
—
## Build Order
### Phase 1 — Saturday (GPU day) - [ ] Install 2x RTX 3060 in Z840 - [ ] Install Ollama - [ ] Pull first models: DeepSeek, Phi-4, nomic-embed - [ ] Test inference speed - [ ] Decide on 3rd GPU
### Phase 2 — Model layer - [ ] Pull Mixtral or Llama 70B - [ ] Install Flux/SDXL - [ ] Configure model router config - [ ] Wire Ollama into Open WebUI
### Phase 3 — Memory & bootstrap - [ ] Formalise CLAUDE.md structure - [ ] Set up embedding pipeline for wiki semantic search - [ ] Test full session bootstrap from phone
### Phase 4 — Agents - [ ] Migrate Hermes to Z840 (already on kaburuaibox) - [ ] Build SEO agent - [ ] Build Divi agent - [ ] Build HA agent - [ ] Build Design agent
### Phase 5 — Customer service agents - [ ] Design white-label CS agent framework - [ ] First client pilot
—
## Diagram
See owl-alpha-architecture.jsx (interactive React diagram)
```mermaid graph TD
Z840[Z840 — AI Nerve Centre<br/>2-3x RTX 3060] --> GPU[GPU Layer<br/>24-36GB VRAM] Z840 --> Ollama[Ollama<br/>Model Serving] GPU --> Ollama GPU --> ImgGen[Flux / SDXL<br/>Image Gen] Ollama --> DeepSeek[DeepSeek<br/>Code & Logic] Ollama --> Mixtral[Mixtral / Llama<br/>Reasoning] Ollama --> Phi[Phi-4 / Gemma<br/>Lightweight] Ollama --> Embed[Embeddings<br/>nomic-embed] Claude[Claude<br/>Orchestrator] --> Router[Model Router] Router --> Ollama Router --> ImgGen Claude --> Wiki[Wiki + CLAUDE.md<br/>Persistent Memory] Wiki --> Embed Claude --> Tailscale[Tailscale<br/>Secure Mesh] Tailscale --> OpenWebUI[Open WebUI] Claude --> Hermes & SEO & Divi & HA & Design & CS n8n --> Hermes & SEO & Divi & HA & Design & CS MCP --> Hermes & SEO & HA
```
—
## Notes
- The name Hermes running on Hetzner is unintentional mythology — messenger god as monitoring agent - Owl Alpha = the overarching AI stack name - All client data stays on-premises — strong GDPR position for selling CS agents - Z840 is powerful enough to run everything; 3rd GPU is insurance not necessity
