# OpenThomas — project context

> Plan on Opus. Run the swarm on Sonnet.

OpenThomas is a **cost-saver for your agent fleet**. It's a small local
proxy that taps the wire between your coding agents and the model
providers they call, and — where it can save money — routes a request
to a cheaper model, keeping the expensive model only where it earns its
price. You don't switch agents or adopt a framework: `openthomas wire`
detects what you already run (Claude Code, OpenClaw, Hermes, Codex,
Claude Desktop, anything that calls an LLM or speaks MCP) and generates
the wrapping on the fly.

The headline feature is the **subagent cost-saver** for [Claude Code
Dynamic Workflows][dw]: those workflows spawn *tens to hundreds of
parallel subagents* that all inherit the session model (Opus 4.8), and
Anthropic warns they "consume substantially more tokens than a typical
session." OpenThomas tells the planner apart from the workers on the
wire — the planner always carries the orchestrator-only `Agent` tool,
subagents never do — and routes only the workers to Sonnet (default) or
a model you pick. Verified exact on 672 real calls; the planner is never
touched. See `src/daemon/orchestrator/subagent.ts`.

Everything is **free and fully open source (MIT)**. The product is three
things the dashboard shows — how many agents are running, what each
agent spends, what each task spends — plus one control: which model each
agent uses.

This file orients Claude Code (and other AI assistants) to the
project. For the full design, read `.strategy/` (private, gitignored)
— start with `.strategy/architecture.md`. The public `docs/` directory
only contains `cli.md` (the user-facing CLI reference).

[dw]: https://claude.com/blog/introducing-dynamic-workflows-in-claude-code

## Status

v0.6 — the cost-saver cut. Repositioned from the earlier "fleet harness"
framing to a single sharp job: **make your agent fleet cheaper, starting
with Claude Code Dynamic Workflows.** Dropped the See/Control/Govern
tiers and paid editions — it's one free OSS tool now. Built on the
existing capture → decode → store → present pipeline and the per-route
model-routing engine (failover chains, vision/web_search capabilities).
Solo-built and used daily by the author; tested against real Claude Code
/ Claude Desktop / OpenClaw / Codex / Hermes traffic.

## Naming

| Thing | Name |
|---|---|
| Project / brand | OpenThomas |
| Domain | `openthomas.com` |
| GitHub repo (planned) | `openthomas-com/openthomas` |
| npm scope | `@openthomas` |
| npm package | `@openthomas/openthomas` |
| Main binary | `openthomas` |
| MCP stdio wrapper bin | `openthomas-tap` |
| MCP tools server bin | `openthomas-tools` |
| Daemon subcommand | `openthomas daemon` |
| Config dir | `~/.openthomas/` |
| Trace dir | `~/.openthomas/wire/traces/` |
| Backups dir | `~/.openthomas/backups/` |
| Logs dir | `~/.openthomas/logs/` |
| Default UI URL | `http://localhost:9877` |
| macOS launchd label | `com.openthomas.daemon` |
| Editions | One — free, MIT, fully open source |

The user-facing brand is **OpenThomas**. "Plan on Opus, run the swarm on
Sonnet" leads; "Get OpenThomas on the wire" survives as a secondary
slogan. Marketing leads with the benefit — **a smaller bill** — and
explains it with the mechanism — wire-level model routing that keeps the
expensive model only where it earns its price.

`openthomas.com` is the canonical URL. `wire` is preserved as
**architectural terminology** (URL paths under `/wire/`, storage under
`~/.openthomas/wire/`, the `wire` / `unwire` commands) — it's
load-bearing in the code, not a brand asset.

## Repo layout

```
thomas/                       (repo: openthomas-com/openthomas, planned)
├── CLAUDE.md                 (this file)
├── README.md                 (public face — see story + journey table)
├── PRIVACY.md                (data handling contract)
├── LICENSE                   (MIT)
├── docs/
│   ├── cli.md                (every CLI command — user-facing reference)
│   └── thomas.md             (canonical Thomas metric spec — citation target)
├── .strategy/                (PRIVATE, gitignored — design + commercial)
│   ├── architecture.md       (system overview, four-layer model)
│   ├── data-model.md         (AgentPacket schema, storage)
│   ├── decisions.md          (technical decisions)
│   ├── decisions-paid.md     (commercial-tier decisions)
│   ├── distribution.md       (form factor, npm publish)
│   ├── protocols.md          (per-protocol decoder specs)
│   ├── runtime-modes.md      (per-agent apply-step semantics)
│   ├── verification.md       (open items needing ground-truth)
│   ├── roadmap.md            (versioned shipping plan)
│   ├── roadmap-full.md       (extended roadmap with paid timeline)
│   ├── positioning.md        (free vs paid, comparison, target user)
│   ├── founder-journey.md    (Idea→MVP→Launch→Scale × free/paid)
│   ├── agents/               (per-agent detector specs)
│   └── The-Founders-Playbook-*.pdf  (Anthropic reference)
├── references/               (third-party projects studied — gitignored)
├── internal/ui/              (Vite + React, builds to packages/thomas/ui-dist)
└── packages/thomas/          (the npm package: @openthomas/openthomas)
```

The **openthomas.com marketing site** and the (planned) **paid cloud
backend** live in a separate closed-source repo:
[`openthomas-com/thomas-cloud`](https://github.com/openthomas-com/thomas-cloud).
This split is intentional — code that runs on a user's machine stays
here under MIT, no telemetry; everything that runs on our infrastructure
or is part of the paid product line lives there.

## Tech stack (locked)

| Layer | Choice | See decision |
|---|---|---|
| Runtime | Node.js 22+ LTS | .strategy/decisions.md / 0002 |
| Language | TypeScript 5+ | 0005 |
| HTTP | Hono + `@hono/node-server` | 0005 |
| CLI parsing | commander | 0005 |
| Error handling | `@praha/byethrow` (Result type) | 0005 |
| Schemas | Valibot | 0005 |
| Storage | better-sqlite3 + Drizzle ORM | 0006 |
| Build | tsdown | 0005 |
| UI | React + Vite (bundled into daemon as static assets) | 0001 |
| JSON / JSON5 edit | `jsonc-parser` (Microsoft) | 0007 |
| YAML edit | `yaml` (eemeli) — Document API to preserve comments | 0007 |
| Test | TBD (vitest or `bun:test`) | — |

## Conventions

- **Language**: All code, comments, commits, docs, and **source** UI strings →
  **English**. Non-English contributors must be able to read everything. The
  human-facing dashboard is localized on top of that English source via a
  lightweight i18n layer (`internal/ui/src/i18n/` — `catalog.ts` + `useT()`);
  English is the default and the fallback for any untranslated key, so authoring
  new UI in English stays correct. The CLI stays English-only (it's
  agent-facing). Supported locales: en, es, fr, de, pt, pt-BR, ja, zh-CN, zh-TW,
  ms.
- **Style**: Default to no comments. Add `// why:` comments only when the
  intent is non-obvious. Never narrate WHAT code does.
- **Files**: kebab-case filenames; camelCase TS identifiers; PascalCase types.
- **Commits**: Conventional Commits.

## Mental model

OpenThomas is a **cost-saver for your agent fleet**. Lead with the
benefit — **a smaller bill** — and explain it with the mechanism —
wire-level model routing.

The product is one sentence: *keep the expensive model only where it
earns its price.* OpenThomas taps the wire between an agent and its
provider, and routes each request to the cheapest model that can do that
particular job — automatically where the signal is unambiguous, and by
your configuration everywhere else.

The flagship instance of that idea is the **subagent cost-saver** for
Claude Code Dynamic Workflows:

- Those workflows spawn *tens to hundreds of parallel subagents* in one
  session; every subagent inherits the session model (Opus 4.8).
- The planner is worth Opus. The hundreds of workers grepping, editing,
  and running tests are not — Sonnet does that work at ~⅕ the price.
- OpenThomas distinguishes them **on the wire, exactly**: Claude Code's
  planner always carries the orchestrator-only `Agent` tool (it's what
  *spawns* subagents); a subagent never does, because it can't nest.
  Presence of `Agent` ⇒ planner ⇒ untouched. Absence ⇒ subagent ⇒
  downgraded. Verified 100% on 672 real ground-truth calls; a token
  floor leaves tiny background calls (security monitor, title-gen) alone.
- Default-on for Claude Code, fully overridable. Lives in
  `src/daemon/orchestrator/subagent.ts`; configured via
  `~/.openthomas/routing.json` (`subagentDowngrade`).

Beyond that one automatic case, the **per-route model-routing engine**
(failover chains, budget/token switch rules, vision & web_search
capabilities) lets you point any agent's traffic at any model — so the
same "spend less" job covers OpenClaw, Hermes, Codex, and Claude Desktop
too. You never switch agents: OpenThomas wraps the **wire**, not the
agent, generated on the fly by `openthomas wire`.

The dashboard surface is deliberately small — three readouts (agents
running, cost per agent, cost per task) and one control (the model each
agent uses). Resist re-growing it into a general observability product;
every screen should serve *see the spend → change the model → spend
less*. Capture → decode → store → present is the pipeline underneath.
Local-first, zero accounts, no telemetry. `proxy` is an architecture
word only, never positioning; `wire` is load-bearing architecture, not a
brand asset.

## Working with this codebase

### Read first

1. `.strategy/architecture.md` — system overview, four-layer model
2. `.strategy/runtime-modes.md` — why each agent is wired differently
3. `.strategy/agents/<name>.md` — exact config paths and rewrite logic
4. `.strategy/data-model.md` — `AgentPacket` schema, storage, routing
5. `.strategy/positioning.md` — free vs paid line, target user

### Gotchas

- **Per-agent runtime modes.** Don't assume "rewrite config = done." For
  long-running agents (OpenClaw, Hermes), wiring also needs an **apply**
  step (restart or paste). See `.strategy/runtime-modes.md`.
- **Backup contract.** `openthomas unwire` reverses exactly the edits Thomas
  recorded in `~/.openthomas/backups/manifest.json` — a surgical
  reverse-replay (`cli/backup/restore.ts`), not a byte-exact whole-file
  restore. The manifest is the source of truth; edits the agent made to
  its own config after wiring are preserved, and a pointer whose value
  changed since wiring is left as-is and reported. `--force` (and legacy
  v1 manifests) fall back to the whole-file snapshot restore. Never edit
  in-place. Always: write to `.tmp`, fsync, rename.
- **Path-based routing.** The proxy receives requests at
  `/wire/<agent>/<provider>/<rest>`. Strip the prefix, look up the
  upstream from `~/.openthomas/wire/routes.json`. Concatenate as strings —
  do NOT use `new URL(restPath, base)` because absolute `restPath`
  would clobber the base's path component.
- **Format-preserving edits.** Users' config files have comments and
  idiosyncratic formatting. Use AST-level editors (`jsonc-parser`,
  `yaml` Document API). Never parse-and-restringify.
- **`openthomas-tap` cold start.** It's spawned per MCP server. Keep its
  bundle small — no heavy daemon code loaded.
- **Verification items**: see `.strategy/verification.md` for things
  believed but not physically tested. Don't rely on them in code until
  verified.
- **`wire` is architecture, not branding.** URL paths (`/wire/...`),
  storage paths (`~/.openthomas/wire/...`), and the slogan all use `wire`
  — don't refactor it away when touching code.

### `references/`

Contains third-party projects studied (Portkey gateway, LiteLLM,
cc-switch, hermes-agent, openclaw, claude-code-router, ccusage, …).
**Not dependencies** — never import from them, never add to package.json.
They are frozen learning material. Don't refactor them.

The full inventory of what each reference taught lives in
`.strategy/` (private); the agent specs in `.strategy/agents/` cite
the relevant findings where they're directly used.

## Do not

- Add a Tauri / Electron / native GUI app. Thomas is CLI + daemon +
  browser UI. Decision 0001 is permanent.
- Switch the published artifact's runtime to Bun. Bun is fine for dev
  tooling; the npm package must run on stock Node. Decision 0002.
- Rename the npm package away from `@openthomas/openthomas` or the
  binaries away from `openthomas` / `openthomas-tap` / `openthomas-tools`
  once republished under these names — they are the user-facing contract.
- Add features beyond what the current commit needs.
- Write README files or docstrings unless explicitly asked.
- Refactor `references/`.
- Move design docs back into the public `docs/` directory. Public
  `docs/` is for "how to use the code" only. Everything else lives in
  `.strategy/`.
- Use `git push --force`, amend already-pushed commits, or skip
  pre-commit hooks without explicit user approval.
- **Send any data about the user or their usage anywhere besides the
  upstream the user's agent already calls.** No telemetry, no crash
  reports, no license-check calls, nothing that contacts an
  `openthomas.com` / Thomas / Anthropic server. The privacy contract in
  `PRIVACY.md` is load-bearing for trust. If a feature legitimately
  needs to send data about the user, it must be off by default,
  documented under the next version's section of `PRIVACY.md` with
  exactly what is sent and where, and called out in the release notes.
  - **One carve-out (v0.4+): the update version check.** A daily
    plain GET to the public npm registry (`registry.npmjs.org`) to see
    if a newer Thomas exists is permitted. It carries no user data, no
    identifiers, never contacts our servers, and is fully disableable
    (`updateCheck: false`). It is documented in the v0.4 section of
    `PRIVACY.md`. This is the *only* sanctioned unsolicited outbound
    call — the bar for anything else is unchanged.
