Skip to content

Finished job outcomes

GET /api/job-outcomes uses the dashboard’s existing loopback Host check and ab_ui authentication cookie. An unauthenticated request receives HTTP 403. ?job=<name> and ?run=<name> select one finished record before Git derivation. They return the same response envelope and contractVersion, with one populated entry, empty other collection, and next: null. An unknown or unfinished name returns 404; supplying both filters returns 400. Paired-PC reads accept these filters too. Run lookups read only the selected log body and retain next-turn receipt boundaries from current and archived metadata. A reused background worker inspects receipt databases and Git evidence. Dashboard requests use asynchronous file-stat signatures and a bounded display cache; they do not open receipt databases or spawn Git. Cold inspections wait at most 250 ms for the worker before returning pending evidence. Pending delivery is unknown, and the dashboard labels the merge evidence as checking. Unchanged older evidence is marked stale while it refreshes.

Database/WAL files, delivery and read receipts, decisions, run metadata, and loose or packed Git refs invalidate the cache. The worker resolves linked Git directories and shared common directories, and periodically rechecks ancestry. Each cache holds at most 1,024 entries and 8 MB, with at most 32 pending batches. These observations are for display only: merge and cleanup authority continue to inspect fresh evidence.

It returns JSON with contractVersion: 1:

interface Outcome {
observation?: { state: "ready" | "pending" | "stale"; checkedAt: number | null };
delivery: {
status: "unknown" | "delivered" | "read";
messageId: string | null;
recipient: string | null;
deliveredAt: number | null;
readAt: number | null;
};
merge: {
state: "merged" | "held" | "unmerged" | "discarded";
branch: string | null;
baseBranch: string | null;
branchHead: string | null;
reason: string | null;
checkedAt: number;
decisionAt: number | null;
decisionBy: string | null;
};
}
interface Response {
contractVersion: 1;
jobs: Record<string, { startedAt: number; status: "done" | "failed"; outcome: Outcome }>;
runs: Record<string, Outcome>;
groups: { needsReview: string[]; held: string[]; merged: string[]; discarded: string[] };
}

jobs keys are job names, including archived finished jobs. runs keys are the dashboard’s recent finished run log names, without .log. groups contains job names grouped by merge state; needsReview means unmerged. Timestamps are Unix milliseconds. Unknown historical timestamps and identities are null.

Delivery means the result reached the supervisor’s durable bridge queue or local inbox. Read means bridge consumption through a hook, channel, or tool; it does not prove that the supervisor reviewed the work. Broker results reuse durable read_at receipts from active and archived messages. Local results retain delivery records and use the existing read journal for consumption receipts. Historical read journal entries without timestamps still mean read, with readAt: null. A missing receipt means unknown, not a lost result. Progress, approval messages, sibling copies, and results outside a run’s time window do not count as that run’s result.

Merge state is derived by testing whether the job’s branch tip is an ancestor of its recorded local base branch. New runs save the final branch name and tip, so ancestry can still be checked after cleanup removes the branch and worktree. Missing Git evidence conservatively gives unmerged with an explanatory reason. An explicit supervisor decision takes precedence over Git ancestry.

Requester records marked remote keep delivery unknown and merge unmerged with a paired-PC evidence reason, unless the supervisor explicitly held or discarded them. Remote paths and receipt namespaces are never tested against local Git or local message rows. Paired-PC outcome evidence needs integration with the remote-jobs contract.

The supervisor MCP tool is:

{"job":"codex-job-deadbeef","state":"held","reason":"Wait for CPU A/B"}

Call set_job_outcome with a job name or ID, state: "held" | "discarded", and an optional reason. A hold requires a nonblank trimmed reason of at most 2000 characters. Only the job’s owning supervisor may set outcomes, and only for done or failed jobs. Delegated sessions do not receive this tool. Its result is an MCP text block containing JSON { job, outcome }; validation and ownership failures return an MCP error. The tool records a decision and does not delete or merge anything. A continued job has a separate decision keyed by its new start time. Earlier decisions stay in history.

JSON store format 2 is an additive upgrade from format 1. Reads derive legacy values without rewriting old run files. The first upgraded write backs up the old bytes and retains unknown fields through the existing atomic writer. Future store versions are not overwritten. Supervisor decisions live in job-outcomes/<hash-of-job-and-start>.json; local delivery evidence lives in local-result-receipts/<hash-of-job>/<hash-of-message>.json. Archives remain readable. The dashboard shows pending inspections and identifies stale display evidence without treating it as cleanup authority.