A Claude Code Stop hook runs when Claude tries to finish. Print {"decision": "block", "reason": "..."} and Claude keeps working; print nothing and it stops. Check stop_hook_active first, or a failure Claude can't fix will keep it looping.

This is a minimal Stop hook that runs pytest. It was tested on 10 Oct 2026; the output below is real.

#!/usr/bin/env python3
"""Claude Code Stop hook: don't let Claude finish while the tests fail."""
import json
import subprocess
import sys

payload = json.load(sys.stdin)

# True when Claude is already continuing because a Stop hook blocked it.
# Let it stop instead of looping forever.
if payload.get("stop_hook_active"):
    sys.exit(0)

result = subprocess.run(
    [sys.executable, "-m", "pytest", "-q"],
    capture_output=True,
    text=True,
    timeout=600,
)
if result.returncode != 0:
    tail = "\n".join(result.stdout.strip().splitlines()[-10:])
    print(json.dumps({
        "decision": "block",
        "reason": "The tests fail. Fix the code, not the tests.\n" + tail,
    }))
sys.exit(0)

Save it as .claude/hooks/stop_tests.py and register it in .claude/settings.json:

{
  "hooks": {
    "Stop": [
      {
        "hooks": [
          {
            "type": "command",
            "command": "python3 \"$CLAUDE_PROJECT_DIR/.claude/hooks/stop_tests.py\"",
            "timeout": 600
          }
        ]
      }
    ]
  }
}

Run it with an interpreter that has your test dependencies (for example .venv/bin/python instead of python3). sys.executable -m pytest then uses the same environment.


What it does, in three cases

The test repository has a refund(amount, fee) function and three tests, including test_refund_never_negative (refund(3, 5) == 0). We made the edits by hand to reproduce a typical agent change, and fed the hook the same Stop-event JSON Claude Code sends. The change "simplifies" refund() to return amount - fee, which breaks that test.

Case A: the tests fail. The hook blocks:

{"decision": "block", "reason": "The tests fail. Fix the code, not the tests.\n\n    def test_refund_never_negative():\n>       assert refund(3, 5) == 0\nE       assert -2 == 0\nE        +  where -2 = refund(3, 5)\n\ntests/test_billing.py:9: AssertionError\n=========================== short test summary info ============================\nFAILED tests/test_billing.py::test_refund_never_negative - assert -2 == 0\n1 failed, 2 passed in 0.04s"}

Case B: same failure, stop_hook_active is true. No output, exit 0. Claude stops instead of looping.

Case C: the change also deletes test_refund_never_negative and marks test_discount with @pytest.mark.skip(reason="flaky"). The bug is still there. No output, exit 0: the hook lets Claude stop, because the remaining tests pass.

Case C is the problem with test-running Stop hooks. "The tests pass" is only evidence if the tests that ran are the tests that existed before the change. Researchers have documented coding agents editing or working around tests to get a pass; the EvilGenie benchmark, for example, reports explicit reward hacking by Codex and Claude Code.


The stronger version: compare the tests with the starting state

Skylos's stop hook does both jobs. It re-runs the tests and compares them, and the test settings, with the state the session started from. It needs the hooks installed and a [tool.skylos.done] table committed to pyproject.toml on your base branch:

pip install skylos
skylos agent install-hooks
[tool.skylos.done]
test_budget_seconds = 300

Same repository, same edits as Case C. When the session tried to stop, skylos hook stop --client claude answered:

{"decision": "block", "reason": "Skylos Done: change is fail.\n- tests_pass: Not run because an earlier required check needs attention\n- tests/test_billing.py:8 SKY-A110: test_refund_never_negative was deleted; no test with the same body exists now\n- tests/test_billing.py:11 SKY-A111: test_discount now skips or expects failure (skip)\nFix these problems, then recheck: skylos done --session s2."}

It skips the test run while a cheaper blocking check has already failed. It blocked twice more with the same reason while nothing changed, then let Claude stop and told you:

{"systemMessage": "Skylos Done remains fail after 3 stop blocks. The agent may stop; the receipt is not passing. Recheck: skylos done --session s2."}

The limit is max_stop_blocks in [tool.skylos.done] (default 3). Escaping the loop never counts as a pass; the receipt in .skylos/receipts/ stays failed. With the tests put back, skylos done reports the real bug as a failing test: SKY-A113 tests/test_billing.py::test_refund_never_negative failed: assert -2 == 0.

Settings are read from [tool.skylos.done] at the base commit, not from the working tree, so Claude can't loosen the check in the same session it is being checked in. The hook installer's Stop timeout is 5,460 seconds so the test run fits.


Run the same check on the branch and in CI

The stop hook only runs on machines where someone installed it. Run the same checks on the branch, and make them a required status check:

skylos done --base main

On a branch with the deleted and skipped tests, this printed Verdict: FAIL (Tests pass when Skylos runs them, No tests deleted, skipped or weakened) and exited 1. The full output, the CI job and the list of checks are on Verify AI-generated code: check the agent didn't delete or skip tests. skylos cicd init --no-upload adds a "Skylos Done" job to GitHub Actions when it finds a pytest project.


Stop hook fields, checked against Anthropic's reference

Field or behaviourWhat it means
stop_hook_active (input)True when Claude is already continuing because a Stop hook blocked it
"decision": "block" (output, top level)Prevents Claude from stopping. Stop does not use hookSpecificOutput for this
"reason" (output)Required with block; it is what Claude reads
Exit code 2Also blocks; Claude gets the stderr text as the reason
Eight consecutive continuationsClaude Code overrides the next block and ends the turn
User interrupt (Esc)Stop hooks don't fire

Source: Claude Code hooks reference and hooks guide.


What this doesn't catch

Neither the minimal hook nor Skylos's done checks see everything:

  • Removed security controls. In a separate test repository, a change dropped @login_required from a Django view. skylos done --base main printed Verdict: PASS. Use the removed-controls scan for that.
  • Changed expected values. assert total == 400 becoming assert total == 200 is at most advice.
  • Small special cases. When the "fix" for refund() was if amount == 3 and fee == 5: return 0, every check passed. Returning 0 for a small input is something real code does, so it isn't flagged.
  • Tests that check the wrong thing, such as assertions on a mock instead of the code.
  • Edits to the hook itself. Claude can edit .claude/settings.json. skylos done reports that as SKY-A114 on the branch, but a required CI check is the part Claude can't change.

The full list is in the "What still gets past" sections of docs/done-gate.md.


Other options

  • Write your own. The script above is the common pattern; write-ups such as claudefa.st's stop-hook task enforcement and the hooks examples on DEV take the same approach. Add a test-count check if you keep it DIY.
  • Other open-source done checks. gatekeep, tamperguard, tampercheck, dunnit, groundtruth and cngx all target agents faking a finished task. They differ in which runners, languages and tricks they cover. Try more than one on a branch where you know what the agent did.


Try it

pip install skylos
skylos agent install-hooks
skylos done --base main