# ai_judge node

Sends a prompt to a Claude model, parses the JSON envelope it returns, and
routes the run based on the parsed value. `ai_judge` is the AI-driven
counterpart to [condition](/docs/nodes/condition): use it when the model
needs to decide something, not describe it.

## Purpose

Reach for `ai_judge` when a routing decision depends on judgement that's
hard to express as a deterministic expression: "is this refund request
within policy?", "does this reply match our tone guidelines?", "which
category does this support ticket belong to?". For higher-stakes gates
that should not depend on a single model call, use
[consensus](/docs/nodes/consensus) instead.

## YAML schema

```yaml
states:
  review:
    type: ai_judge # discriminator
    label: Is the change safe? # required
    model: claude-haiku-4-5 # required, must be in roscoe.yaml allow-list
    prompt: | # required, auto-augmented with JSON envelope spec
      Review the diff in {{collect.stdout}} and decide if it is safe to ship.
    validator: # required — drives routing
      kind: boolean # boolean | enum | confidence
    maxRetries: 1 # optional, 0-5, default 1
    on: # transition keys depend on validator
      'true': ship
      'false': revise
```

The prompt you write is automatically wrapped with JSON envelope
instructions so the model returns a parseable shape (`{ result, reasoning }`,
plus `score` for confidence). You don't write that scaffolding yourself.
Describe the question and Roscoe adds the rest.

## Configuration

| Field            | Type                                | Required | Default                                           | Meaning                                                                                                                                  |
| ---------------- | ----------------------------------- | -------- | ------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- |
| `type`           | `'ai_judge'`                        | yes      | —                                                 | Discriminator.                                                                                                                           |
| `label`          | string                              | yes      | —                                                 | Human-readable name for the step.                                                                                                        |
| `model`          | string                              | yes      | —                                                 | Model id; checked against `roscoe.yaml` at pre-flight.                                                                                   |
| `prompt`         | string                              | yes      | —                                                 | Question for the model. Supports `{{node.key}}` interpolation.                                                                           |
| `validator`      | `boolean` \| `enum` \| `confidence` | yes      | —                                                 | Determines the JSON shape and how it routes. See below.                                                                                  |
| `maxRetries`     | integer (0-5)                       | no       | `1`                                               | Retried on parse failure or empty response. Timeouts are not retried.                                                                    |
| `timeoutSeconds` | integer (1-2147483)                 | no       | global default (`defaults.timeoutMs`, else 5 min) | Hard wall-clock timeout for this node, overriding the global default.                                                                    |
| `maxIterations`  | integer (1-50)                      | no       | —                                                 | Per-run iteration cap when this node sits in a feedback loop. See the [closed-loop feedback](/docs/recipes/closed-loop-feedback) recipe. |
| `detectStall`    | boolean                             | no       | `false`                                           | Fail the run early if the node produces an identical verdict three times in a row. Most useful on reviewer-shaped judges.                |
| `on`             | transition map                      | yes      | —                                                 | Required keys depend on validator.                                                                                                       |

### Validators

- `kind: boolean` — model returns `result: true | false`. Routes
  `'true'` / `'false'`.
- `kind: enum` — model picks one of the keys of `on`. Routes to that
  key. The list of valid choices is taken from `Object.keys(on)`.
- `kind: confidence` — model returns a `score` between 0 and 100. Routes
  `'true'` if `score >= validator.threshold`, else `'false'`.

```yaml
validator:
  kind: confidence
  threshold: 75
on:
  'true': ship
  'false': revise
```

See [validators](/docs/authoring/validators) for the JSON envelope shape,
retry semantics, and prompt-augmentation details.

## Outputs

| Key         | Type                        | Notes                                             |
| ----------- | --------------------------- | ------------------------------------------------- |
| `result`    | boolean \| string \| number | The parsed value the model returned.              |
| `reasoning` | string                      | The model's stated reasoning, surfaced in the UI. |
| `score`     | number                      | Only present for `kind: confidence` (0-100).      |

Downstream nodes can read `{{review.result}}`, `{{review.reasoning}}`, or
`{{review.score}}`.

## Transitions

| Validator    | Emitted transition keys                           |
| ------------ | ------------------------------------------------- |
| `boolean`    | `'true'` or `'false'`.                            |
| `enum`       | One of the keys of `on` (the model chose).        |
| `confidence` | `'true'` if `score >= threshold`, else `'false'`. |

If the produced transition isn't a key of `on`, the node fails with a clear
mismatch error (useful for catching `enum` typos).

## Worked example

```yaml
name: ai-judge
version: 1
description: Single-agent gate that asks Claude to judge a fact.
initial: ask
states:
  ask:
    type: ai_judge
    label: Is the sky blue?
    model: claude-haiku-4-5
    prompt: |
      Is the daytime sky on a clear day blue? Answer with the JSON shape
      requested below.
    validator:
      kind: boolean
    maxRetries: 1
    on:
      'true': yes-end
      'false': no-end

  yes-end:
    type: end
    label: Yes
    outcome: success
    message: The judge agreed.

  no-end:
    type: end
    label: No
    outcome: failure
    message: The judge disagreed.
```

## Common pitfalls

- **`enum` keys that the model can't reasonably guess.** The model only sees
  the `on` keys as the menu of choices. Use clear, descriptive keys
  (`approve`, `reject`, `needs_info`), not `a`, `b`, `c`.
- **`confidence` threshold mismatched to scale.** Thresholds are 0-100, not
  0-1. Setting `threshold: 0.8` means "any positive score routes true",
  which is almost certainly not what you wanted.
- **Prompts that ask for free-form prose.** The validator needs JSON. The
  executor adds envelope instructions automatically, but a prompt that
  _insists_ on prose ("answer in one paragraph") will fight the envelope
  and fail to parse.
- **Single-agent gates for high-stakes decisions.** A single judge can be
  fooled or unlucky. For decisions where a wrong route is expensive,
  use [consensus](/docs/nodes/consensus) with a quorum.
- **Treating `result` as branching info downstream.** Downstream nodes see
  the parsed value verbatim. For confidence you usually want
  `{{review.score}}`; for booleans, you usually want to branch in the `on`
  map rather than re-checking the result later.

## Where to next

- [condition](/docs/nodes/condition) — a cheaper deterministic check when no
  judgement is needed.
- [script](/docs/nodes/script) — branch on code instead of a model's
  opinion.
- [Output chaining](/docs/authoring/output-chaining) — read
  `{{node.result}}` or `{{node.score}}` downstream.
- [validators](/docs/authoring/validators) — the full validator reference
  shared with `script` and `consensus`.
