Documentation contents

Introduction

OpenHarnX checks the work of the coding agent you already use. It protects the tests you agreed on, runs them where the change cannot interfere, and writes a review brief that separates what was verified from what was not.

OpenHarnX 0.1.1

A coding agent reports success when the tests pass. Tests can be made to pass without doing the task: rewrite an assertion to expect the new output, mark the inconvenient test as skipped, delete it as obsolete, or change the configuration so it is not collected. Plain pytest agrees, and the agent says "All tests pass".

OpenHarnX sits between trusting that summary and reading every line yourself.

What it does#

  1. Locks the tests that must keep passing before the agent starts: your existing suite (ohx init --lock-tests), or acceptance tests agreed for one task (ohx contract new). A baseline records which tests passed.
  2. Lets the agent work as it normally does. You keep your agent.
  3. Verifies the change with ohx verify: the locked copies run, not the agent's edited files, inside the srt sandbox, from an interpreter the change cannot replace. Per-test results are compared with the baseline and weakened tests or configuration are flagged.
  4. Writes a review brief: what was asked, which files changed, what was verified (each claim linked to the check's output), what remains unverified and the decisions left to you.

The gate and the report make no model calls and need no API key. Verdicts come from checks that ran.

What a verdict means#

VerdictMeaning
READYEvery mandatory check passed, including acceptance tests agreed for the task
NO REGRESSIONSEvery mandatory check passed and nothing that passed before broke. No acceptance tests were agreed, so nothing shows the task is done
BLOCKEDA mandatory check failed: a broken test, a weakened test, changed check configuration
UNKNOWNA mandatory check could not give a result. Never counted as a pass
INVALIDThe evidence cannot be trusted
STALEThe files or the contract changed after the report was made

Where it runs#

  • Locally, from the command line after your agent finishes, on any agent's work.
  • As a Claude Code Stop hook, so Claude Code is verified when it says it is done and a blocked agent is sent back with what failed.
  • In CI, as a gate on pull requests: a GitHub Action, with GitLab CI and Jenkins examples.

Release 0.1.1 supports Python projects tested with pytest, on macOS (arm64) locally and on Linux in CI, with srt as the sandbox. See supported environments.

Where to go next#

OpenHarnX checks the work of one agent at a time. It is not an agent and does not plan, staff or orchestrate agents.

Esc
Try verify, STALE, approve-tests or GitLab. Common pages: