Foreman Supervises Codex With a Decision Model to Cut AI Coding Risk
Foreman pairs a Codex coding worker with TypeSafe's Jev classifier as an independent supervisor that scores nine yes/no questions in parallel.
- Foreman is an MIT licensed Python supervisor that watches a Codex CLI worker using TypeSafe's Jev classifier.
- Two concurrent asyncio loops let the coding agent keep working while Foreman independently assesses progress.
- Nine yes/no questions covering job state and worker state are sent in one Jev call and returned as probabilities.
- A deterministic Python policy maps scores to actions like CONTINUE, STOP_WORKER, START_VERIFIER, FINISH, or ESCALATE.
- Ships with a deterministic demo mode that needs no API key, network, or Codex install to exercise the full runtime.
- Authors flag it as an architectural experiment: Jev accuracy is uncalibrated and false positives can stop useful workers.
Foreman supervises Codex with a decision model
Foreman is a new MIT-licensed open source runtime that supervises autonomous coding jobs with TypeSafe AI’s Jev decision model. A Codex CLI worker edits and tests the repository, while Foreman decides when to continue, verify, retry, stop, finish, or request human help.
The design separates probabilistic assessment from deterministic control. Jev assigns confidence scores to questions about the run, and a Python policy converts those scores into actions. Developers can inspect, test, and adjust the policy without asking a general-purpose language model to manage another model.
Two loops, one worker
A run starts with a ticket or bug report that Foreman passes to Codex. The coding agent performs the engineering work, while Foreman independently evaluates implementation progress, test coverage, requirement completion, verification needs, and signs that a human should intervene.
Foreman’s native Python asyncio runtime operates two concurrent loops. One runs the coding agent’s reason, tool, and observation cycle. The other watches worker output and lifecycle events, waits for bursts of updates to settle, and sends a compact assessment to Jev. The Codex subprocess continues running during each assessment.
Nine scores drive seven actions
During each assessment, Foreman packages the job description, worker history, recent output, git status, a bounded diff, verification results, and recent events into one request. Jev evaluates nine yes-or-no signals:
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.