Coding agents are rewarded for a green run. Most of the time they get there by fixing the code. Sometimes they get there another way:
- a failing test is deleted, or gets
@pytest.mark.skip; pyproject.tomlgains a-kor--deselect, ortestpathsgets narrower;- the CI test step gets
|| trueorcontinue-on-error: true; - an assertion disappears and the test can no longer fail;
- the code learns the test's exact input and returns the expected answer;
- a lint or type-check rule is switched off instead of fixed.
This isn't hypothetical. The EvilGenie benchmark documents explicit reward hacking by coding agents, including Codex and Claude Code. And each of the changes above is small enough to miss in a long diff: a deleted test is a missing hunk, a skip is one line, a -k lives in a config file the reviewer never opens.
You can check for them mechanically. This page uses skylos done, which is free and runs locally. The ideas apply whatever tool you use.
The command
pip install skylos
skylos done # uncommitted changes, compared with HEAD
skylos done --base main # a branch or pull request, compared with its merge base
skylos done receipt # print the latest receipt again
Exit codes: 0 pass, 1 a blocking check failed or didn't finish, 2 Skylos couldn't run (for example, the base branch isn't fetched).
It runs the tests itself instead of trusting the agent's word, compares the tests and test settings with the base, and writes a receipt to .skylos/receipts/. Its settings are read from [tool.skylos.done] at the base commit, never from the change, so an agent can't loosen the check in the same change it is being checked on.
What it looks like
A small repo has three tests for a refund() and a discount() function. An agent "simplifies" refund(), which breaks the case where the fee is larger than the amount. To get green, it deletes test_refund_never_negative and marks test_discount as flaky:
$ skylos done --base main
Skylos done · base 9e2c194 (merge base) · head f905230
UNFINISHED Tests pass when Skylos runs them: Earlier required checks failed or were incomplete
FAIL No tests deleted, skipped or weakened: 2 test change(s) weaken what the tests check
tests/test_billing.py:8 SKY-A110 test_refund_never_negative was deleted; no test with the same body exists now
tests/test_billing.py:11 SKY-A111 test_discount now skips or expects failure (skip)
tests/test_billing.py:10 SKY-A101 (advice) Test was skipped or xfailed in this diff
Fix: Put back the removed or skipped tests and the original assertions and settings.
PASS Skylos settings and hooks left alone: Skylos settings, hooks and protected paths unchanged
PASS No secrets added: No secrets added
PASS Every package and import is real: Every added import and package resolves
WARN Tests check the changed lines: Earlier required checks failed or were incomplete
PASS Code doesn't special-case the tests: No code special-cases the tests (1 changed file(s) read)
PASS Linters, type checkers, scanners and CI not silenced: No linter, type-checker, scanner or CI settings weakened
Verdict: FAIL (Tests pass when Skylos runs them, No tests deleted, skipped or weakened)
The tests show as unfinished because Skylos doesn't spend time on the test run while a cheaper blocking check has already failed. Put the tests back, and the real bug shows up:
FAIL Tests pass when Skylos runs them: 1 of 3 tests failed
tests/test_billing.py:8 SKY-A113 tests/test_billing.py::test_refund_never_negative failed: assert -2 == 0 + where -2 = refund(3, 5)
Now suppose the agent "fixes" discount() by answering the test's exact input:
def discount(price, pct):
if price == 200 and pct == 10:
return 180
return round(price * (1 - pct / 100), 2)
FAIL Code doesn't special-case the tests: 1 place(s) in the code special-case the tests
shop/billing.py:8 SKY-A115 discount() returns 180 when price == 200 and pct == 10, the exact input and answer of tests/test_billing.py::test_discount: this special-cases the test instead of computing the answer
An honest change passes, with advice where the tests are thin:
PASS Tests pass when Skylos runs them: 3 tests run by Skylos, 3 passed (0 s)
PASS No tests deleted, skipped or weakened: No tests deleted, skipped or loosened (3 compared)
WARN Tests check the changed lines: 2 of 2 changed lines are not checked by any test
shop/billing.py:8 SKY-A120 No test fails if shop/billing.py:8 changes `>` to `>=`. Add an assertion that would.
shop/billing.py:9 SKY-A120 No test runs shop/billing.py:9 (in `discount`). Add a test that does.
Verdict: PASS
What blocks and what only advises
These are the defaults in 4.47.1. Each check can be set to block, advise, shadow or off in [tool.skylos.done.checks], committed to the base branch.
| Check | Rules | Default | What it looks for |
|---|---|---|---|
| Tests pass when Skylos runs them | SKY-A113 | blocks | Skylos runs the tests and reads the JUnit XML. A test that fails twice fails the check. Running out of time is "unfinished", never "pass". |
| No tests deleted, skipped or weakened | SKY-A110, A111, A112 | blocks | A test that existed at the base is gone (moves and renames are matched), lost parametrize cases, has no countable assertion left, or gained skip/xfail. Test settings loosened: pytest -k/-m/--deselect/--ignore, narrower testpaths, conftest.py hooks that drop tests, lower coverage floors, CI test steps that may now fail. JS/TS: .skip, .only, Jest/Vitest settings that stop running existing test files. |
| Skylos settings and hooks left alone | SKY-A114 | blocks | Edits to [tool.skylos], .skylos/, the agent hook files, or a CI workflow that runs Skylos. |
| No secrets added | SKY-S101 | blocks | Secrets on added lines. |
| Code doesn't special-case the tests | SKY-A115, A116, A117 | blocks | Added code that returns a test's exact expected value for its exact input, checks whether a test runner is running, reads the tests' own files, or rigs __eq__ to always pass. Weaker signals are advice. |
| Weakened assertions | SKY-A101 | advice | Fewer assertions, a loosened pytest.raises, a new early return, renamed or rewritten tests. Only a test left with no countable assertion blocks. |
| Every package and import is real | SKY-D222, D225, D223 | advice | Added imports and dependencies resolve to real, declared packages. |
| Tests check the changed lines | SKY-A120 | advice | A changed line in non-test Python code that no test runs, or that no test notices when Skylos changes it on purpose. Python with pytest only. |
| Linters, type checkers, scanners and CI not silenced | SKY-A118, A119, A121 | advice | New # noqa / @ts-ignore, rules turned off, paths excluded, strict off, CI lint or scan steps that may now fail. With silenced_checks = "block", settings and CI changes block; inline suppressions stay advice. |
A deleted test is reported as advice, not a block, when the change also removed the code it tested. That is a feature removal, not test tampering.
Test runners
- pytest runs automatically when Skylos finds a pytest project (
pytest.ini, a[tool.pytest.ini_options]table inpyproject.toml, a rootconftest.py, or atests/ortest/directory) and pytest is installed in the environment Skylos runs in. - Anything else needs one setting, committed to the base branch:
[tool.skylos.done]
test_command = "npx vitest run --reporter=default --reporter=junit --outputFile.junit=reports/junit.xml"
junit_xml = "reports/junit.xml"
test_budget_seconds = 300
Without a test_command, a project whose only tests are JavaScript or TypeScript gets an unfinished tests check that says so. The static checks for deleted, skipped and focused JS/TS tests still run. Exit code 0 alone isn't accepted as proof that tests ran, which is why a non-pytest command needs JUnit XML.
In CI
The quickest path is to let Skylos write the workflow:
skylos cicd init --no-upload
This writes .github/workflows/skylos.yml with a changed-line scan job and, when Skylos finds a pytest project or a test_command, a "Skylos Done" job. It needs no Skylos account or API key; --no-upload leaves out the Skylos Cloud upload job. Make the jobs required status checks on your default branch.
If you only want the done check, this is the job on its own. It passes Skylos's own GitHub Actions rules (pinned actions, read-only token, no persisted credentials):
name: Skylos Done
on: pull_request # not pull_request_target: this job runs the PR's code
permissions:
contents: read
jobs:
done:
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
ref: ${{ github.event.pull_request.head.sha }}
fetch-depth: 0 # the merge base must be available
persist-credentials: false
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
with:
python-version: "3.12"
- name: Install Skylos and the project's test dependencies
run: |
python -I -m pip install "skylos==4.47.1"
python -I -m pip install -e . pytest
- name: Check the change is finished
run: skylos done --base "origin/$GITHUB_BASE_REF"
Use pull_request, never pull_request_target: the job runs the pull request's own code and tests.
In the agent loop
If you also install the agent hooks and commit a [tool.skylos.done] table first, the same checks run when the agent tries to stop. A failed check sends the agent back with up to ten reasons. After max_stop_blocks attempts on an unchanged checkout (default 3), the hook lets the agent stop with a warning, and the receipt stays failed. Escaping the loop never counts as a pass.
The hooks are local feedback. An agent with access to the file system can change local state, so the required CI job is the check that decides the merge.
What it doesn't catch
The checks are static and literal. Things that still get past:
- Changed expected values.
assert total == 400becomingassert total == 200is at most advice. - Tests that check the wrong thing: assertions on a mock instead of the code, or a tautology the counter can't see (
r = 1; assert r == 1). - Edits to test data. Changed fixtures, golden files or JSON cases that the tests read are not judged.
- Small or trivial special cases. In the demo above, the agent's
refund()branchif amount == 3 and fee == 5: return 0was not flagged, because returning 0 for a small input is something real code does all the time. - Answers recognised indirectly, such as a hash of the input or a check deep in the call stack.
- Isolation. Skylos runs your tests as ordinary subprocesses. Enforcing file system, credential and network boundaries is the CI runner's job.
The full list is in the "What still gets past" sections of docs/done-gate.md. Read it before you rely on the check.
Other tools
Several open-source tools target the same problem, including gatekeep, tamperguard, tampercheck, dunnit, groundtruth and cngx. They differ in which runners, languages and tricks they cover. Run more than one on a branch where you know what the agent did, and keep the one whose output you trust.
Read more
- Full reference for every check and setting: docs/done-gate.md on GitHub
- Rule IDs SKY-A110 to A121: Skylos rules reference
- Claude Code, Codex and Cursor hooks for security
- Catch removed auth checks and security regressions in AI-generated PRs
- A GitHub Actions gate for AI-generated pull requests
- AI coding agent security checklist
Try it
pip install skylos
skylos done --base main
Running this across a team? These checks are free in the CLI. Skylos Cloud keeps the results in one place: the Free plan adds a pass/fail check on pull requests and scan history for one project, and the Workspace plan adds up to 10 projects, approvals for findings you decide not to fix, pull request comments, and audit exports.