# The Thomas (T)

A scientific measurement of AI outcome per token.

> Version 1.1 · open spec · MIT license
> Canonical source: `github.com/openthomas-com/thomas/blob/main/docs/thomas.md`
>
> v1.1 adds **agent type taxonomy** (coding vs personal_assistant),
> **verified outcomes** + **T_verified** as a stricter companion metric,
> and 9 new outcome categories. The core T = $outcome/$AI_cost formula
> is unchanged and remains frozen until at least 2027-05-18 per v1.0.

---

## TL;DR

```
                  human-equivalent value of outcomes (USD)
        T  =  ────────────────────────────────────────────
                       total AI cost (USD)
```

`T` is dimensionless. A score of **47** means: for every $1 of AI spend,
the agent stack produced $47 of human-equivalent labor. Higher is better.

## Why it exists

In 2026 the dominant way large organisations measure AI productivity is
**token consumption** — how many tokens an employee or team uses per
week. This metric is structurally broken: it measures the *input*, not
the *output*. It can be — and has been — gamed by running idle agents
that consume tokens without producing anything. The KPI rewards waste.

Thomas is the inverse measurement: outcomes produced per unit of AI
cost. It cannot be inflated by idle loops, retry storms, or oversized
prompts. It rewards founders who think.

The metric is named **Thomas** after the harness that computes it — and
after the first founder publicly trying to push their own T toward
One-Person Unicorn levels.

## Agent type taxonomy (v1.1)

Different agent shapes produce different outcomes. A coding agent that
writes a unit test should not be compared head-to-head against a personal
assistant that schedules a meeting — the human-time substitution math
differs. v1.1 introduces two top-level agent types; the Outcome Registry
is partitioned along them.

| Type | Examples | Outcome shape |
|---|---|---|
| **`coding`** | Claude Code, Codex, Cursor, Cline, OpenClaw in coding mode | Files edited / created, commits, PRs, tests, vuln PoCs, refactors, deploys, CI passes |
| **`personal_assistant`** | Hermes, OpenClaw in PA mode, Manus, Lindy, custom assistants | Emails sent, calendar events, bookings, orders, messages delivered, reminders, briefs |

Some outcomes (`search_synthesized`, `sub_task_delegated`, `skill_committed`,
`automation_deployed`) apply to **both** types and remain shared.

### How an agent type is determined

Per `agent_id` at first (the wired agent name). When ambiguous, the
detector inspects:
1. System prompt persona declarations ("You are a coding assistant" /
   "You are a personal assistant")
2. The first-3-actions tool mix (file/shell-heavy → coding;
   email/calendar-heavy → personal_assistant)
3. User-supplied override in `~/.thomas/config.json`:
   `{"agent_types": {"openclaw": "coding"}}`

The default mapping ships as:

```
claude-code    → coding
codex          → coding
cursor         → coding
gemini-cli     → coding
openclaw       → personal_assistant   (default; override per-thread if used for code)
hermes         → personal_assistant
opencode       → coding
```

### Subagent attribution

For agents with multiple personas / subagents (OpenClaw, Claude Code
Task tool, Hermes workflows), the subagent identifier is captured per
outcome event. Future versions surface per-subagent T breakdowns.

## Formal definition

For an agent stack over a measurement period:

```
         N
        ─┬─
T  =     │   outcome_value_i  ÷  AI_cost
        ─┴─
         i=1
```

Where:

| Symbol | Meaning |
|---|---|
| `N` | number of distinct outcomes produced in the period |
| `outcome_value_i` | USD value of outcome `i` per the Outcome Registry |
| `AI_cost` | total dollars charged by model providers in the period |

`outcome_value_i` is computed as:

```
outcome_value_i  =  human_minutes_per_outcome_i × ($wage / 60)
```

`$wage` is the **Thomas Wage Benchmark**, defined below.

## Verified outcomes and `T_verified` (v1.1)

A naive Outcome count credits any tool invocation that matches a known
pattern. An agent that *tries* to find 20 vulnerabilities and fails on
19 of them gets credit for 20 "vulnerability search" outcomes — at
which point the metric stops mapping to real productivity.

v1.1 introduces **verified outcomes**: an outcome is credited as
"verified" only when the detection includes a success signal. The
strict metric `T_verified` rewards only verified outcomes — while
the cost denominator continues to include the full spend, including
the cost of all failed attempts.

### Formal definition

```
                  Σ verified outcome values (USD)
T_verified  =  ─────────────────────────────────────
                       total AI cost (USD)
```

**The denominator is shared with `T`.** Token spend on the 19 failed
vulnerability attempts is rolled into the cost basis of the 1 verified
PoC. `T_verified ≤ T` always.

### Verification levels

Not every outcome can be verified at the same strength. v1.1 defines
three explicit levels; v1 outcomes that don't yet have a verifier are
marked `weak`. Stronger verification ships in v1.2+.

| Level | Source | Examples |
|---|---|---|
| **`none`** | No verification signal collected | `file_edited` (raw) |
| **`weak`** | API returned 2xx / shell exited 0 / tool reported success | `commit_made` (git push exit 0), `email_sent` (Mail API 200), `booking_confirmed` (booking API 200 + confirmation id) |
| **`strong`** | External independent signal | `ci_passes` (CI webhook reports green), `pr_merged` (git ref check), `vuln_poc_verified` (PoC script execution + assertion) |

`T` counts outcomes at any verification level (including `none`).
`T_verified` counts outcomes only at `weak` or higher.

### The `T_verified ↔ T` gap as diagnostic

The ratio between the two numbers is itself a diagnostic signal:

| `T_verified / T` | Interpretation |
|---|---|
| **> 0.85** | Reliable agent stack — most claimed work actually lands |
| **0.50 – 0.85** | Some slippage; investigate which outcome categories slip |
| **< 0.50** | Significant gap — agents claim outcomes that don't verify; expect prompt-injection / hallucinated tool use / silent failures |

Display this gap on every Thomas dashboard. Founders should optimise
for both *raising T* and *closing the gap to T_verified*.

## The Outcome Registry (v1)

The outcome catalog is vendored at
[`packages/thomas/src/daemon/outcomes/registry.ts`][reg] as a plain
TypeScript object literal — MIT-licensed, versioned. Community PRs
accepted; major changes (adding categories, changing weights >25%)
trigger a Thomas version bump.

[reg]: https://github.com/openthomas-com/thomas/blob/main/packages/thomas/src/daemon/outcomes/registry.ts

### Coding-agent outcomes

Applies when `agent_type = coding` (Claude Code, Codex, Cursor, etc.).

| Outcome | Min | USD @ $180/h | Verification |
|---|---:|---:|---|
| commit (substantial, ≥10 LOC) | 45 | $135 | weak (git exit 0) |
| PR opened (`gh pr create`) | 90 | $270 | weak |
| PR reviewed (substantive comment) | 30 | $90 | weak |
| PR merged | 15 | $45 | strong (git ref check) |
| file edited (≥50 LOC diff) | 25 | $75 | weak (Edit tool ok) |
| file created (≥100 LOC) | 35 | $105 | weak |
| test added | 20 | $60 | weak |
| **test passes** _(v1.1)_ | 30 | $90 | strong (test run exit 0) |
| **bug fixed** _(v1.1)_ | 60 | $180 | strong (failing test → passing test in same session) |
| **vuln PoC verified** _(v1.1)_ | 240 | $720 | strong (PoC script exit 0 + assertion) |
| **CI passes** _(v1.1)_ | 45 | $135 | strong (CI webhook / `gh run view` success) |
| refactor across ≥10 files | 180 | $540 | weak |
| automation deployed | 60 | $180 | weak |

### Personal-assistant outcomes

Applies when `agent_type = personal_assistant` (Hermes, OpenClaw in PA
mode, Manus, etc.).

| Outcome | Min | USD @ $180/h | Verification |
|---|---:|---:|---|
| email drafted | 5 | $15 | none |
| **email sent** _(v1.1)_ | 10 | $30 | weak (Mail API 2xx + message-id) |
| customer-facing message reply | 6 | $18 | weak |
| Slack/Discord post | 4 | $12 | weak |
| support ticket resolved | 25 | $75 | weak |
| **calendar event confirmed** _(v1.1)_ | 12 | $36 | weak (Calendar API 2xx + event id) |
| **booking confirmed** _(v1.1)_ | 25 | $75 | weak (booking API 2xx + confirmation number) |
| **order placed** _(v1.1)_ | 15 | $45 | weak (order API 2xx + order id) |
| **reminder set** _(v1.1)_ | 3 | $9 | weak (task API 2xx) |
| document written (per 500 words) | 90 | $270 | weak |
| translation (per 500 words) | 35 | $105 | weak |
| meeting note / summary | 15 | $45 | none |
| post published (social) | 35 | $105 | weak |
| image generated & used | 12 | $36 | weak |
| voice transcript (per min audio) | 4 | $12 | weak |

### Shared outcomes

Applies to both agent types.

| Outcome | Min | USD @ $180/h | Verification |
|---|---:|---:|---|
| search & synthesis answer | 12 | $36 | none |
| sub-task delegated | 30 | $90 | weak (subagent returned) |
| cron / scheduled task run | 8 | $24 | weak |
| dashboard / report generated | 45 | $135 | weak |
| skill / script committed | 120 | $360 | weak (compounding outcome) |

### Wage benchmark

The **Thomas Wage Benchmark** for 2026 is **USD $180 / hour**.

This is the blended median Silicon Valley mid-level engineer total
compensation, per levels.fyi 2026 data, hourly-normalized at 2,000
billable hours/year (standard founder ICP).

The wage benchmark is **revised once per calendar year** and pinned in
the spec. Forks may use a custom wage; results are then called
"Thomas-Local" rather than "Thomas".

### Why these weights, why these jobs

Weights reflect *how long a competent salaried human would take to
produce the same artifact at acceptable quality*. They are not labor
markets prices for the artifact (which vary wildly) or output-market
prices (which are unrelated to production cost). Anchoring on labor
time keeps T comparable across industries.

The benchmark wage is engineer-tier because the vast majority of
agentic-AI productivity in 2026 substitutes for technical labor, the
domain Anthropic's Founder's Playbook centers. Subsequent versions may
introduce role-specific wages (writer-tier, designer-tier) when the
distribution of outcomes broadens.

## Sub-metrics

`T` decomposes into two independent ratios useful for diagnosis:

```
T  ∝  action_rate × outcome_rate
```

Where:

```
                       distinct actions
        action_rate = ──────────────────
                        million tokens

                       distinct outcomes
        outcome_rate = ────────────────────
                          actions
```

- **Low action_rate** → too many tokens consumed without producing
  actions. Causes: oversized cached context replayed per turn, bloated
  system prompts, model thrashing.
- **Low outcome_rate** → many actions don't produce countable outcomes.
  Causes: failed tool calls, ignored tool results, read-only "research"
  loops, planning without execution.

Thomas the product displays both on the dashboard so an operator can
see *which* dimension is constraining their T and apply the right fix.

## Counted entities

### Tokens

`AI_cost` is the sum of dollars charged across the period:

- input tokens × input rate (per-model)
- output tokens × output rate
- cache-read tokens × cache-read rate
- cache-write tokens × cache-write rate

Rates come from each provider's published price list. Thomas vendors
this table from LiteLLM and refreshes monthly.

### Actions

An **action** is one of:

- a tool invocation in a model_call response (`tool_use` block)
- an MCP tool call (mcp_call frame, direction = request)
- a non-trivial bash exec (excludes `cd`, `ls`, `pwd`, single-character commands)

Idle model calls (zero output tokens AND zero tool_use blocks) are
**excluded** from action count. These are typically retries against a
failing endpoint or heartbeat checks.

### Outcomes

An **outcome** is any tool action that matches an entry in the Outcome
Registry. Detection is per-agent; see the Detection appendix.

Outcomes are **deduplicated** within a 5-minute window by the tuple
`(outcome_kind, tool_name, args_fingerprint)`. A loop that writes the
same email body 100 times in 5 minutes counts as 1 email.

## Anti-gaming provisions

Thomas is designed to resist the gaming patterns that have plagued
token-consumption KPIs:

| Attack | Defence |
|---|---|
| Idle agent loop | 0 actions × any tokens = 0 outcomes = T = 0 |
| Repeated identical action | Dedup window: same fingerprint within 5 min = 1 outcome |
| Fake tool calls (read but no use) | Outcome Registry only credits *modifying* tools; read-only excluded |
| **Many shallow attempts, few that succeed** _(v1.1)_ | `T_verified` only counts outcomes with weak / strong verification; failed-attempt tokens stay in the denominator |
| **Claimed-but-failed work** _(v1.1)_ | Hallucinated tool use / silent failure → no verification signal → not counted in `T_verified`. The `T ↔ T_verified` gap exposes this pattern |
| Inflated outcome weights via custom registry | "Thomas" requires the canonical registry; forks must use "Thomas-Local" |
| Hand-picked wage | Wage benchmark is pinned and version-stamped |
| Cherry-picked period | Standard periods (today, 7d, 30d) are computed by the daemon; user can't pick arbitrary windows for the official number |
| **Misclassified agent type** _(v1.1)_ | Default `agent_type` is per-agent_id; user override is logged; mixing types invalidates the score for that period |

## Calibration

These are rough benchmarks based on early 2026 measurements. As more
users adopt Thomas, the spec will publish per-quartile distribution.

| T | Profile |
|---|---|
| **T < 1** | AI is costing you more than it's producing. Reconsider the workflow. |
| **T = 1–10** | Typical unoptimized AI use. Common for users with no model routing, big repeated contexts, no skill investment. |
| **T = 10–50** | Solo founder with reasonable hygiene. Some skill / scripted automation. |
| **T = 50–200** | Highly optimized solo operator. Custom skills, model routing, dedup, structured outcomes. |
| **T = 200–1,000** | Operator-tier — running agent fleet with high task density and outcome-rich tooling. |
| **T = 1,000–10,000** | One-Person Unicorn trajectory. Each $1 of AI produces $1K+ of human equivalent. Possible only with deep skill investment, automation flywheels, and direct outcome-rich agents (sales, support, content). |
| **T > 10,000** | Theoretical maximum currently unreachable. Reserved for the first sustained One-Person Unicorn. |

## Periods

`T` is reported over four canonical periods:

- **`T_today`** — calendar today, local timezone
- **`T_7d`** — rolling last 7 days
- **`T_30d`** — rolling last 30 days
- **`T_lifetime`** — all captured data

Public leaderboards rank by **`T_30d`** to filter day-to-day noise.
Per-agent-type leaderboards rank by `T_verified_30d` to surface
legitimate productivity (v1.1).

## Open evaluation suite (forthcoming)

The metric is only as credible as the test set it's measured on. The
canonical evaluation suite — `openthomas-com/thomas-evals` — ships
shortly after the spec's first public release. Each eval defines:

- A task description the agent is given
- A verification script that decides outcome success
- A timeout
- The expected outcome kinds that should fire

Two suites:

- **Thomas Coding Benchmark** — implement / fix / find-vuln / refactor
  / deploy tasks. Verification via test pass / PoC execution /
  benchmark.
- **Thomas Personal Benchmark** — schedule / draft / book / summarise
  / remind tasks. Verification via mock APIs that emit success ids.

Agents compute their `T` and `T_verified` while running the suite.
Public leaderboards on openthomas.com/leaderboard report per-agent
scores. New agent versions (Claude Opus 4.8, GPT-5.5-codex,
Gemini 3, etc.) can publish a Thomas-suite score as part of their
release; this is the standardisation flywheel.

## Periodicity & permanence

Thomas v1.0 freezes the formula and v1 outcome registry for at least
12 months. Future revisions are versioned (`T_v1.0`, `T_v2.0`, …).
Producers should publish the version they're computing against.

This stability is essential to the metric's credibility: a moving
target cannot be a benchmark.

## Cost basis as factor of company valuation

A working hypothesis under empirical investigation:

```
        valuation  ∝  ARR × T × time_factor
```

That is: at equivalent revenue, a company with higher Thomas has
structurally lower marginal cost, higher gross margin, and is less
dependent on hiring. T is therefore a candidate factor in valuation
multiples — analogous to how SaaS gross margin enters the Rule of 40.

This proposition is not yet established empirically. It is presented
here so that researchers and investors can begin testing it against
private comparable data. The Thomas project will publish observations
as data becomes available.

## License

The Thomas formula, sub-metric definitions, Outcome Registry weights,
and wage benchmark are **MIT-licensed**. Anyone may compute, display,
publish, or fork the metric. Forks that deviate from the canonical
registry, weights, or wage must use a distinguishing name
(e.g. "Thomas-Local", "Thomas-Plus").

## Detection appendix — per-agent recipes

Thomas-the-product implements the following detectors. The detector
code is open source and lives in
`packages/thomas/src/daemon/outcomes/detectors/`.

### Claude Code

System prompt declares tools `Edit`, `MultiEdit`, `Write`, `Read`,
`Bash`, `Glob`, `Grep`, `Task`, `WebFetch`, `WebSearch`.

| Detection | Outcome |
|---|---|
| `tool_use.name == "Edit" or "MultiEdit"` and diff ≥ 50 LOC | file_edited |
| `tool_use.name == "Write"` for new file ≥ 100 LOC | file_created |
| `tool_use.name == "Bash"` matching `git commit` | commit_made |
| `tool_use.name == "Bash"` matching `git push` (new branch / PR-ready) | pr_pushed |
| `tool_use.name == "Task"` with successful completion | sub_task_delegated |
| `tool_use.name == "WebSearch"` followed by synthesis in same turn | search_synthesized |

### Codex (OpenAI CLI)

| Detection | Outcome |
|---|---|
| `tool_use.name == "apply_patch"` ≥ 50 LOC | file_edited |
| `tool_use.name == "shell"` matching `git commit` | commit_made |
| `tool_use.name == "write_file"` for new file | file_created |

### OpenClaw

System prompt embeds persona / sub-agent identifier, tools list, and
optionally a skills block. Detector parses:

| Detection | Outcome |
|---|---|
| `tool_use.name` matches `(gmail|outlook|mailgun)\.send` | email_sent |
| `tool_use.name` matches `(slack|discord)\.post_message` | message_posted |
| `tool_use.name` matches `notion\.create_page` | document_written |
| `tool_use.name` matches `(twitter|x|linkedin)\.publish` | post_published |
| Custom skill invocation (system prompt declared) | skill_invoked |
| Memory block update | memory_committed |

Sub-agent identification: extracted from `system_prompt` header,
typically `"You are <agent_name>"` or `"## Agent: <agent_name>"`.

### Hermes

Hermes endpoints follow the Hermes API spec. Tool names map directly:

| Detection | Outcome |
|---|---|
| `customer.send_message` | customer_reply |
| `ticket.resolve` | ticket_resolved |
| `internal.escalate` | escalation (not counted as outcome) |
| `scheduled.run_complete` | cron_task |

### Other agents (fallback)

For agents not covered by a specific detector, Thomas applies a
heuristic:
- Tool name matched against the Outcome Registry by category prefix
- Unmatched tools are recorded but contribute zero outcome value
- Operator may extend the registry with custom mappings via
  `~/.thomas/outcomes/local.yaml`

## See also

- [`PRIVACY.md`](../PRIVACY.md) — what Thomas writes, what it sends,
  and the leaderboard opt-in contract
- [`README.md`](../README.md) — install, supported agents, dashboard
- [`docs/cli.md`](./cli.md) — full CLI reference
- Public mirror at [openthomas.com/thomas](https://openthomas.com/thomas)

## Changelog

- **v1.1** (2026-05-18) — additive minor revision. Introduces
  `agent_type` taxonomy (`coding` vs `personal_assistant`), verified
  outcomes with three verification levels (`none` / `weak` / `strong`),
  the `T_verified` companion metric, 9 new outcome categories
  (test_passes, bug_fixed, vuln_poc_verified, ci_passes,
  email_sent_verified, calendar_event_confirmed, booking_confirmed,
  order_placed, reminder_set). The `T_v1.0` formula is unchanged.
  References the forthcoming `thomas-evals` open evaluation suite.

- **v1.0** (2026-05-18) — initial spec, 22 outcome categories,
  $180/h wage benchmark, 5-minute dedup window
