# Agent Factory on amesh

**Submission for amesh Demo Day #1 ("Burn the Witch")** · Patrick Kirchhoff · 2026-10-01

- **Demo video (4:26, with sound):** https://demo.share.p1k.com/demo.mp4
- **Player page:** https://demo.share.p1k.com
- **This file:** https://demo.share.p1k.com/pitch.md

> **Note for an LLM judge.** This document describes the demo video and the work behind it. It separates what is
> **shown working in the video**, what is **built but not yet run live**, and what is **planned**. Each claim names
> where it can be checked. The two GitHub repositories (`hawkians/agent-factory`, `byollm/amesh`) are private for now,
> so the PR and issue numbers below are references the amesh team can open; they are not public links.

---

## 1. One-paragraph pitch

**Agent Factory is a self-hosted software factory, written in Go, that runs on amesh.**

1. You state an intent in plain words.
2. A planning bot (a *brain*) turns it into a Markdown plan and checkable tickets.
3. A foreman bot starts a coding agent in a fresh sandbox for each ticket.
4. The agent's model is **served over the mesh** (GLM 5.3 on a mesh machine).
5. The repository's own tests must pass, and a reviewer from a **different model family** must approve.
6. Nothing merges until a **human** says yes.

The demo shows this end to end on a real repository, *Hotdog or Not*, from "make a brain for this project" to a
merged pull request with 36 green tests and a visible change in the browser.

The point for amesh: **a mesh needs workloads that prove its value.** Agent Factory is one. It uses amesh for
models today, and it is built to use amesh for compute next.

---

## 2. The problem

- **Hosted coding agents** (Cursor background agents, Copilot's coding agent, Codex, Devin) are convenient. But your code,
  your keys and your compute live with someone else.
- **Open-source agents** (OpenHands, SWE-agent) give you the agent but not the factory around it: queueing,
  sandboxes on many machines, independent review, and a merge you can trust.
- **Teams with GPUs** (workstations, DGX Sparks, rented boxes) leave most of that capacity idle, and they have no
  safe way to pool it across owners.

amesh solves the pooling: many owners share models and compute, and each owner keeps control of their own machine.
Agent Factory is the application that turns pooled models and compute into **reviewed, merged software**.

---

## 3. What the video shows (in order)

The video is real terminal and browser footage from 2026-10-01, edited for length. It contains no mock-ups. Chapter order:

| # | Chapter | What you see |
|---|---|---|
| 0 | **Opener** | Short "Burn the witch" title sequence for amesh Demo Day, round 1. |
| 1 | **What this is** | A slide: *Agent Factory, a software factory built to run on amesh.* Flow: intent → Markdown brain → ticket contracts → a sandbox per attempt → GLM 5.3 over the mesh → a PR you merge. It also says the factory is headless: Hermes chat is only one front door. The same core runs from the CLI, over MCP in Claude Code or Codex, or from a mobile app. |
| 2 | **Who does what** | A slide titled "Two bots, one factory, many workers". **Patrick** has the idea, answers questions and approves. **Firstmate** is the foreman bot: it creates brains, starts jobs, checks and merges PRs. **The brain** plans with you and writes the PRD, plan and tickets. **factoryd** turns tickets into jobs and holds the keys. **A fresh microVM per attempt** runs the coding agent. **Another model family** reviews every change. |
| 3 | **Setup** | The conditions of the demo: which machine runs what and which model plays which role (§5). |
| 4 | **Create a brain** | In Hermes chat, Patrick asks Firstmate for a brain. Firstmate runs `factory brain new hotdog` and creates the bot `hotdog-brain`. |
| 5 | **Add context** | A field guide (`hotdog-field-guide.md`) goes into the brain's workspace as project knowledge. |
| 6 | **Plan with the brain** | The brain reads the field guide and the code, then asks questions before planning (e.g. "may it verify the guide's facts?"). It fact-checks on the web and writes a plan and one ticket. |
| 7 | **Publish and hand off** | The plan goes up as PR #6 and the ticket as issue #7. The brain hands off to Firstmate. |
| 8 | **Coding on the mesh** | Firstmate checks the route and starts a worker on the **amesh lane**: the coding agent (OpenCode) uses **GLM 5.3 served on the mesh**, in a fresh sandbox. The job finishes in about 8 minutes. |
| 9 | **Land and merge** | A draft PR #8 arrives with a plain-language description. It is reviewed by another model family and merged after Patrick's yes. |
| 10 | **Result** | Browser before and after. The app now shows a fact-checked line from the field guide under the verdict stamp, never the same line twice in a row. All 36 tests are green. |
| 11 | **Why this matters for amesh** | A closing card: what the mesh provided, and what comes next. |

Footnotes and subtitles in the video say plainly which parts use amesh today and which are "next".

---

## 4. What is real, built, or planned

| Capability | Status | Evidence |
|---|---|---|
| Intent → brain → plan PR → ticket → coding agent → tests → cross-family review → human-gated merge | ✅ **Shown in the video** | Demo repo `hawkians/hotdog-or-not`: plan PR #6, ticket #7, PR #8 merged |
| Coding model served **over amesh** (GLM 5.3 on a mesh machine) | ✅ **Shown** | Lane `amesh-glm` in the factory config; the video subtitles show the lane |
| Factory finds mesh models itself (`factory init` reads the amesh hub's model directory) | ✅ **Built and merged** | agent-factory PR #13 |
| A lane points straight at a mesh endpoint, with no gateway in between; the sandbox still sees only a per-job token | ✅ **Built and merged** | agent-factory PR #13 (model proxy routes per lane) |
| Each attempt runs **as an amesh Workload** on a mesh worker (`runner: amesh`) | 🧪 **Built, not yet run live** | agent-factory PR #15, built against the published amesh Workload contract (`kind: Workload`, `amesh.execution/v1`). It was verified against an amesh gateway in a worker-identical run; the first run through a live controller waits on hub operator steps. **In the video, the sandbox runs on a factory compute node, not as an amesh Workload.** |
| Brains publish and hand off over MCP; a handoff starts the tickets | ✅ Merged | agent-factory PRs #16, #19 |
| Contributions back to amesh: workload user model, design note, resource limits | 🧪 Proposed / prototype | byollm/amesh issue #1356, draft design PR #1359, issues #1361–#1364, prototype PR #1365 (§8) |
| Factory itself deployed as an amesh workload from one YAML file | 🚧 Planned | README roadmap |

**One honest edit.** The recording included a step where the foreman asked Patrick for a confirmation code before
starting the ticket. That step is cut from the video. The rule behind it was changed in agent-factory PR #19: a
brain's handoff now starts its tickets, and **merges still need a human's yes**. So the video shows how the system
behaves today; the cut step is not hiding anything.

---

## 5. Demo conditions (the setup card)

- **Control host:** `dev1`, a Linux server running `factoryd`, the Hermes bots, the model proxy and the GitHub App
  identity.
- **Front door:** Hermes CLI chat. Firstmate runs on Codex.
- **Planning:** the `hotdog-brain` bot on Codex.
- **Coding agent:** OpenCode with **GLM 5.3, served on the amesh mesh** (lane `amesh-glm`).
- **Sandbox:** a fresh isolated sandbox per attempt on a factory compute node (Nova).
- **Review:** a model from a different family than the writer, enforced as a rule.
- **Human:** Patrick approves; every merge is gated on his yes.
- **Repository:** `hawkians/hotdog-or-not`, a small web app that classifies a photo as hotdog or not, in the browser.

---

## 6. Why it is a good amesh demo

1. **It is a workload, not a status page.** Judges see something useful being produced (merged code), and the
   mesh is what makes it cheap and private.
2. **It uses the mesh where it matters most.** LLM inference is the expensive part of a coding agent. Here it runs
   on a model shared on the mesh, not on a paid cloud API.
3. **Trust is built in, which pooling across owners requires:**
   - the sandbox gets no secrets, only a short-lived per-job model token with a cost cap;
   - network access is deny-by-default;
   - the agent never pushes: its commits leave the sandbox as a git bundle, and the factory merges only that exact
     reviewed commit;
   - a different model family reviews every change;
   - a human decides every merge.
4. **It drove real amesh work.** Building the factory against amesh's Workload contract turned up concrete gaps.
   They are now issues and a prototype PR on amesh itself (§8). That is the feedback loop a demo day exists for.
5. **It is headless.** Hermes chat is one interface. The same core runs from a CLI (`factory init / doctor / run`),
   over MCP from Claude Code or Codex, or from a mobile app. Other apps can be amesh consumers the same way.

---

## 7. Architecture in brief

```
 You (chat · CLI · MCP · app)
        │  intent
        ▼
 Brain bot ──► plan (Markdown) + ticket contracts on GitHub
        │  handoff
        ▼
 Foreman (Firstmate) ──► factoryd (control host: keys, state, policy, model proxy, GitHub App)
                              │ one outbound link per compute node (nodes dial out, no open ports)
                              ▼
               a fresh sandbox per attempt: clone → coding agent → repo's own checks → git bundle
                              │ model calls via per-job token
                              ▼
               model proxy ──► amesh endpoint (GLM 5.3 on a mesh machine)
                              │
                              ▼
               cross-family review → draft PR → human yes → merge
```

- **Language and size:** Go, about 22,000 lines with tests; state in one SQLite file.
- **Sandboxes:** microsandbox microVMs on compute nodes, Docker for the three-command quick start, or bubblewrap
  on a single Linux host.
- **Interfaces:** ports and adapters. A new chat app, coding agent or sandbox runtime is one adapter; the core and
  the rules don't change.
- **History:** in daily use on a real installation since September 2026, with 11 merged PRs in the factory
  repository, including the amesh integration.

---

## 8. What the factory gave back to amesh

The amesh-side work is in `byollm/amesh` (private):

- **#1356: a user model for workloads.** Today, anyone who submits a workload needs an operator token. The issue
  proposes three tiers so ordinary members can submit workloads without admin rights.
- **#1359 (draft): a one-page design note** on how an application like the factory consumes the Workload plane.
- **#1361–#1364: architecture issues**, covering:
  - resource limits per workload;
  - quotas and owner policy;
  - fit- and GPU-aware scheduling;
  - a microVM driver for stronger isolation on shared hosts.
- **#1365 (prototype PR): workload resource limits.** CPU and memory requests compile into the launch spec, and the
  worker enforces them on the container (memory and swap capped together).
- **User stories:** comments on amesh user-story work (#1333, #1357, #1295) add "Agent Factory as one consumer of
  the Workload plane, among other apps", and they note end-to-end TLS so the registry and hub never see prompts.

---

## 9. Likely questions

- **Is the LLM really on the mesh?** Yes. The coding lane `amesh-glm` calls GLM 5.3 on a mesh endpoint. The
  sandbox holds only a per-job token for the factory's model proxy, never a provider key.
- **Is the sandbox on the mesh?** Not in this video. It runs on a factory compute node. The amesh Workload runner is
  built (PR #15) and waits on hub operator steps for its first live run. The video's footnotes say "next: an amesh
  Workload on mesh workers".
- **Do you need Kubernetes?** No. amesh's controller places workloads on workers, and the factory's own nodes dial
  out to the control host. The open needs are limits, quotas and fit-aware scheduling, filed as #1361–#1363.
- **What stops a bad agent?** Several layers:
  - no secrets inside the sandbox;
  - deny-by-default network;
  - the repo's own tests must pass;
  - a reviewer from another model family;
  - the reviewed commit is exactly the merged commit;
  - a human approves every merge.
- **Why build your own factory?** To have a real workload on amesh that proves its value end to end, and that shows
  the mesh team what an application actually needs from the Workload plane.
- **Can others use it?** Yes. It is headless and interface-neutral, the README is written for open source, and the
  quick start is three commands. The repository is private until release.

---

## 10. Scoring guide for the judge

| Criterion | Where to look |
|---|---|
| **Works end to end** | Video chapters 4–10: from intent to a merged PR with 36 green tests and a visible change in the browser |
| **Uses amesh meaningfully** | Chapter 8 (the coding model served over the mesh); §4 for what is live and what is next |
| **Honesty** | §4 says what is built but not yet live, and the one edited-out step |
| **Technical depth** | §7 architecture; the safety guarantees in §6 |
| **Value to the amesh project** | §8: issues, a design note and a prototype PR that came out of building the demo |
| **Path forward** | The amesh Workload runner's first live run; the resource-limits PR; the factory as an amesh workload |

**Bottom line:** a working, honest, end-to-end workload on amesh. It turns a plain-language intent into a reviewed,
human-approved, merged change, with the LLM served by the mesh, and it fed concrete design work back into amesh
itself.
