AI coding assistants produce code that looks right, passes linting, and reaches 90%+ line coverage. Yet that code regularly ships with off-by-one boundary errors, vacuous test assertions, and behavioral regressions that surface only in production. Multi-tier CI/CD verification gates address this gap by layering structural analysis, property-based testing, and mutation testing into a single GitHub Actions pipeline that automatically annotates AI-authored pull requests with actionable SARIF reports.
How to Build Multi-Tier CI/CD Verification Gates for AI Pull Requests
- Configure your repository with Node.js v22+, TypeScript, and required devDependencies (
@babel/parser,fast-check,vitest,@stryker-mutator/core). - Label AI-authored PRs with an
ai-generatedtag to selectively trigger the verification pipeline. - Build a Tier 1 AST diff analyzer using
@babel/parserand@babel/traverseto detect empty catch blocks, identical branches, and console statements. - Write Tier 2 property-based tests with fast-check and Vitest to verify invariants across randomly generated inputs.
- Add a custom Vitest assertion-guard reporter to fail tests that pass with zero assertions.
- Configure Tier 3 Stryker mutation testing scoped to changed PR files, with SARIF output for inline annotations.
- Assemble the three tiers into a sequential GitHub Actions workflow with
needs:dependencies and fail-fast design. - Calibrate mutation score thresholds against your baseline and promote Tier 3 from soft gate to hard gate.
Table of Contents
- Prerequisites
- Why Standard CI Fails AI-Generated Code
- Architecture of a Three-Tier Verification Pipeline
- Tier 1: AST Diff Analysis with Babel
- Tier 2: Property-Based Testing with fast-check and Vitest
- Tier 3: Mutation Testing with Stryker Mutator
- The Complete Workflow: Assembling ai-pr-verify.yml
- Tuning, Pitfalls, and Production Hardening
- Trusting AI Code Through Verification, Not Hope
Prerequisites
Before starting, ensure the following are in place:
- Node.js v22+ and a compatible version of npm
- A GitHub repository with Actions enabled and Code Scanning enabled (required for SARIF upload; on private repos this requires GitHub Advanced Security)
- A
package.jsonwith the following devDependencies (pin versions to ensure reproducibility):
{
"devDependencies": {
"@babel/parser": "^7.23.0",
"@babel/traverse": "^7.23.0",
"fast-check": "^3.15.0",
"vitest": "^1.6.0",
"@stryker-mutator/core": "^8.2.0",
"@stryker-mutator/vitest-runner": "^8.2.0"
}
}
- A
tsconfig.jsonwith"moduleResolution": "NodeNext"and"module": "NodeNext"(required for.jsextension imports from.tssource files):
{
"compilerOptions": {
"target": "ES2022",
"module": "NodeNext",
"moduleResolution": "NodeNext",
"strict": true,
"outDir": "dist",
"rootDir": "."
},
"include": ["src", "tests", "reporters"]
}
- A
vitest.config.ts(shown in the Tier 2 section below) - A
src/sort.tsfile exporting asortNumbersfunction (or equivalent module under test) - PRs must have the
ai-generatedlabel applied (or be authored by a known AI bot account)
Why Standard CI Fails AI-Generated Code
The False Confidence of Green Builds
AI-generated pull requests can hit high line coverage while containing zero actual assertions or off-by-one logic errors. The tests pass and the linter is clean. The formatter has nothing to say. And the code is wrong. This happens because coverage measures execution, not correctness. A test that calls a function and never checks its return value still counts as a covered line. AI code compiles and lints cleanly but contains logic errors that no existing rule flags. It generates plausible control flow, reasonable variable names, and idiomatic patterns, all of which sail through traditional CI pipelines without triggering a single warning.
A test that calls a function and never checks its return value still counts as a covered line.
What Multi-Tier Verification Gates Solve
The three-tier model addresses distinct failure classes in sequence. Dead branches, duplicated conditional arms, and empty catch blocks slip past regex-based linting; Tier 1 catches them through structural AST analysis. Tier 2 applies property-based testing with fast-check to generate random inputs (100 by default; configure via numRuns for broader coverage), exposing edge cases that hand-written unit tests never consider. Tier 3 runs Stryker mutation testing, injecting small code changes to verify that tests actually detect behavioral differences rather than merely executing code paths. Each tier catches what the previous tier cannot. The tech stack throughout: GitHub Actions, Node.js v22+, TypeScript 5.x, @babel/parser and @babel/traverse for AST work, fast-check for property-based testing, Stryker Mutator for mutation analysis, and Vitest as the test runner.
Architecture of a Three-Tier Verification Pipeline
Tier Overview and Fail-Fast Design
The pipeline runs each tier in strict sequence: PR opened → Tier 1 gate (AST diff analysis) → Tier 2 gate (property-based tests) → Tier 3 gate (mutation score evaluation) → SARIF annotation upload → human review. Each tier is a separate GitHub Actions job connected via needs: dependencies. If Tier 1 fails, Tiers 2 and 3 never execute, saving runner minutes and providing immediate feedback. This fail-fast design is critical when Tier 3 (Stryker) is CPU-intensive.
For selective triggering, the pipeline uses a labeling strategy. PRs from known AI bot authors (Copilot, Cursor, Codeium) or PRs manually tagged with an ai-generated label trigger the verification workflow. Human-authored PRs skip the extra gates entirely, avoiding unnecessary latency on standard contributions.
When to Gate vs. When to Advise
Not every tier should block a merge. Hard gates prevent merging until the check passes. Soft gates post a PR comment or annotation but allow the merge to proceed. For most repositories, Tier 1 (AST analysis) works well as a hard gate because structural anti-patterns like empty catch blocks are rarely intentional; suppress the few that are via comment annotations. Tier 2 (property tests) also functions effectively as a hard gate once the property suite is established. Start Tier 3 as a soft gate that posts the mutation score as a PR annotation. Mutation thresholds require calibration: run Stryker once on the existing test suite, record the baseline score, and set break 10 points below it before enabling hard gating.
Tier 1: AST Diff Analysis with Babel
Why AST-Level Checks Beat Regex Linting
AI-generated code passes ESLint and Prettier effortlessly. Those tools enforce style and catch common syntax issues, but they do not inspect program structure deeply enough to detect dead branches, duplicated conditional arms, or vacuous try/catch blocks where the catch body does nothing. These patterns require tree-level inspection. Using @babel/parser with the TypeScript and JSX plugins and @babel/traverse for node walking provides the structural depth needed to flag these AI-specific anti-patterns in Node.js v22+ ESM scripts.
Building the AST Diff Analyzer
// scripts/verify-pr-ast.mjs
import { readFileSync, writeFileSync } from 'node:fs';
import { execSync } from 'node:child_process';
import { parse } from '@babel/parser';
import _traverse from '@babel/traverse';
// @babel/traverse ships CJS with a default export; under Node ESM interop
// the callable may be on .default. Normalise here.
const traverse = typeof _traverse === 'function'
? _traverse
: (_traverse.default ?? (() => { throw new Error('traverse not callable'); }));
const baseSha = process.env.BASE_SHA;
if (!baseSha) {
console.error('BASE_SHA environment variable is required. Set it to the PR base branch SHA.');
process.exit(1);
}
const changedFiles = execSync(`git diff --name-only "${baseSha}" HEAD`, { timeout: 10000 })
.toString()
.split('
')
.filter(f => /\.(ts|js|mjs)$/.test(f) && f.trim().length > 0);
// Initialize SARIF with empty results before processing so upload-sarif
// always finds the file even if all files throw parse errors.
const sarif = {
$schema: 'https://docs.oasis-open.org/sarif/sarif/v2.1.0/errata01/os/schemas/sarif-schema-2.1.0.json',
version: '2.1.0',
runs: [{ tool: { driver: { name: 'ast-diff-analyzer', version: '1.0.0', rules: [] } }, results: [] }],
};
writeFileSync('ast-results.sarif', JSON.stringify(sarif, null, 2));
const findings = [];
for (const file of changedFiles) {
let code;
try {
code = readFileSync(file, 'utf-8');
} catch {
continue;
}
let ast;
try {
ast = parse(code, {
sourceType: 'module',
plugins: ['typescript', 'jsx'],
});
} catch (e) {
findings.push({
file, line: 1, rule: 'parse-error',
message: `File could not be parsed: ${e.message}`,
});
continue;
}
traverse(ast, {
CatchClause(path) {
if (path.node.body.body.length === 0) {
findings.push({
file, line: path.node.loc.start.line,
rule: 'empty-catch', message: 'Empty catch block swallows errors silently',
});
}
},
IfStatement(path) {
const cons = JSON.stringify(path.node.consequent);
const alt = path.node.alternate ? JSON.stringify(path.node.alternate) : null;
if (alt && cons === alt) {
findings.push({
file, line: path.node.loc.start.line,
rule: 'identical-branches', message: 'If/else branches are identical — dead logic',
});
}
},
CallExpression(path) {
const callee = path.node.callee;
if (
callee.type === 'MemberExpression' &&
callee.object.type === 'Identifier' &&
callee.object.name === 'console'
) {
findings.push({
file, line: path.node.loc.start.line,
rule: 'console-in-prod', message: 'console.* call left in production code path',
});
}
},
});
}
// Update SARIF with actual findings before exiting
sarif.runs[0].results = findings.map(f => ({
ruleId: f.rule,
message: { text: f.message },
locations: [{
physicalLocation: {
artifactLocation: { uri: f.file },
region: { startLine: f.line },
},
}],
level: 'error',
}));
writeFileSync('ast-results.sarif', JSON.stringify(sarif, null, 2));
if (findings.length > 0) process.exit(1);
The script parses each changed file from the PR diff, walks the AST, and detects three categories of AI-specific anti-patterns. Empty catch blocks are pervasive in AI output because models optimize for "compiles without errors" and silently swallow exceptions. Identical if/else branches indicate the model duplicated logic without differentiating paths. Console statements in production code paths reveal debugging scaffolding the model forgot to remove. The output is formatted as SARIF JSON, which GitHub's code scanning integration renders as inline PR annotations. The script writes an initial empty SARIF file before processing so that the upload step always finds the file, even if all files fail to parse. Parse errors are captured as findings rather than crashing the script.
Wiring Tier 1 into the Workflow
tier1-ast:
if: contains(github.event.pull_request.labels.*.name, 'ai-generated')
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- uses: actions/setup-node@v4
with:
node-version: '22'
- run: npm ci
- run: node scripts/verify-pr-ast.mjs
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
continue-on-error: false
- uses: github/codeql-action/upload-sarif@v3
if: always()
with:
sarif_file: ast-results.sarif
Setting continue-on-error: false makes this a hard gate. The fetch-depth: 0 fetches the full history so that git diff against the PR base SHA works reliably, including for squash-merge and first-commit PRs (a shallow fetch-depth: 2 can produce empty or incorrect diffs for merge commits). The workflow sets BASE_SHA from the pull request event context, and the script uses it to compute the correct diff. The SARIF upload step runs with if: always() so annotations appear even when the analysis step fails the job.
Tier 2: Property-Based Testing with fast-check and Vitest
Why Property Tests Catch What Unit Tests Miss
Unit tests assert known examples: given input X, expect output Y. Property tests generate random inputs (100 by default; configure via numRuns for broader coverage) and verify invariants that must hold across all of them. This distinction matters enormously for AI-generated code, which tends to handle the "happy path" examples the model was trained on while failing on boundary conditions, empty arrays, negative numbers, or Unicode strings. Properties fall into three categories: invariants (a sort function's output length must equal its input length), round-trip properties (encode then decode returns the original), and oracle comparisons, where the new implementation must match a known-correct reference across all generated inputs.
Writing Effective Property Suites for AI Code
// tests/sort.property.test.ts
import { describe, it, expect } from 'vitest';
import * as fc from 'fast-check';
import { sortNumbers } from '../src/sort.js';
// Note: the .js extension import requires "moduleResolution": "NodeNext"
// and "module": "NodeNext" in tsconfig.json (see Prerequisites).
describe('sortNumbers property tests', () => {
it('preserves array length', () => {
fc.assert(
fc.property(fc.array(fc.integer()), (arr) => {
expect(sortNumbers(arr)).toHaveLength(arr.length);
})
);
});
it('produces sorted output', () => {
fc.assert(
fc.property(fc.array(fc.integer()), (arr) => {
const result = sortNumbers(arr);
for (let i = 1; i < result.length; i++) {
expect(result[i]).toBeGreaterThanOrEqual(result[i - 1]);
}
})
);
});
it('output is a permutation of input', () => {
fc.assert(
fc.property(fc.array(fc.integer()), (arr) => {
const result = sortNumbers(arr);
expect([...result].sort((a, b) => a - b)).toEqual([...arr].sort((a, b) => a - b));
})
);
});
});
This suite tests a sorting utility, one of the most commonly AI-generated modules. Three properties together provide strong guarantees: length preservation, ordering, and permutation correctness. Note the numeric comparator (a, b) => a - b in the permutation test — JavaScript's default Array.prototype.sort() uses lexicographic ordering, which would make the assertion vacuous for integer arrays (e.g., [10, 2, 1].sort() yields [1, 10, 2], not [1, 2, 10]). A vacuous assertion pattern, where a test body simply asserts true or calls the function without any expect(), would pass coverage checks but catch nothing. Property tests make this pattern structurally harder to produce because the framework expects a predicate that exercises the return value.
A module can achieve 95% line coverage while its tests fail to detect any of these behavioral changes, producing a mutation score below 10%.
Detecting Zero-Assertion and Vacuous Assertion Paths
A custom Vitest reporter can count assertion calls per test and flag any test that completes with zero assertions. The exact reporter API depends on your Vitest version; the following example targets Vitest ≥1.4.0 using the onTaskUpdate hook:
// reporters/assertion-guard.ts
import type { Reporter } from 'vitest';
const assertionGuard: Reporter = {
onTaskUpdate(packs) {
// packs: Array<[taskId: string, result: TaskResult | undefined]>
for (const [, result] of packs) {
if (
result?.state === 'pass' &&
result?.type !== 'suite' &&
(result?.assertionCount ?? -1) === 0
) {
console.error(
'[assertion-guard] FAIL: test passed with zero assertions'
);
process.exitCode = 1;
}
}
},
};
export default assertionGuard;
Note: The assertionCount field on TaskResult is available in Vitest ≥1.4.0. For earlier 1.x versions the field may be absent; the ?? -1 sentinel causes the guard to skip those entries safely rather than producing false positives. Pin your Vitest version in package.json to avoid breaking changes.
Register this reporter in vitest.config.ts:
// vitest.config.ts
import { defineConfig } from 'vitest/config';
export default defineConfig({
test: {
reporters: ['default', './reporters/assertion-guard.ts'],
},
});
This ensures that vacuous tests, a hallmark of AI-generated test suites, fail the pipeline even though Vitest itself would mark them as passing.
Wiring Tier 2 into the Workflow
tier2-property:
if: contains(github.event.pull_request.labels.*.name, 'ai-generated')
needs: tier1-ast
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: '22'
- run: npm ci
- run: npx vitest run --reporter=json --outputFile=property-results.json 'tests/**/*.property.test.ts'
- uses: actions/github-script@v7
if: always()
with:
script: |
const fs = require('fs');
let summary = '**Tier 2 Property Tests**: results unavailable (file missing or unreadable)';
try {
const raw = fs.readFileSync('property-results.json', 'utf-8');
const results = JSON.parse(raw);
summary = `**Tier 2 Property Tests**: ${results.numPassedTests ?? 'N/A'} passed, ${results.numFailedTests ?? 'N/A'} failed`;
} catch (e) {
summary = `**Tier 2 Property Tests**: failed to read results — ${e.message}`;
}
github.rest.issues.createComment({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: context.issue.number,
body: summary,
});
The needs: tier1-ast dependency enforces ordering. The if: guard ensures this tier is also skipped when the ai-generated label is absent. Note the quoted glob 'tests/**/*.property.test.ts' — without quotes, bash may not expand ** correctly unless globstar is enabled, causing Vitest to find zero test files and exit successfully with no tests run. Vitest runs only property test files (matched by the *.property.test.ts glob), and the results post as a PR comment via actions/github-script. The comment step safely handles the case where property-results.json is missing or unreadable (e.g., when Vitest exits non-zero before writing the file), posting a descriptive fallback message instead of crashing.
Tier 3: Mutation Testing with Stryker Mutator
The Mutation Score as a Code Correctness Proxy
Mutation testing works by injecting small code changes, called mutants, into the source: flipping > to >=, replacing + with -, swapping true for false. If the test suite still passes after a mutation, that mutant "survived," indicating the tests do not truly verify that behavior. This is especially revealing for AI-generated code. A module can achieve 95% line coverage while its tests fail to detect any of these behavioral changes, producing a mutation score below 10%. The mutation score is a far more honest proxy for test correctness than coverage percentage.
Configuring Stryker for Targeted PR Scope
Important: Using testRunner: 'vitest' requires the separate package @stryker-mutator/vitest-runner to be installed as a devDependency alongside @stryker-mutator/core. Without it, Stryker will fail with "Cannot find test runner plugin for vitest."
// stryker.config.mjs
function parseChangedFiles() {
const raw = process.env.CHANGED_FILES;
if (!raw) return ['src/**/*.ts'];
try {
const parsed = JSON.parse(raw);
if (!Array.isArray(parsed) || parsed.length === 0) return ['src/**/*.ts'];
return parsed;
} catch {
console.error('[stryker.config] CHANGED_FILES is not valid JSON; falling back to src/**/*.ts');
return ['src/**/*.ts'];
}
}
/** @type {import('@stryker-mutator/api/core').PartialStrykerOptions} */
export default {
mutate: parseChangedFiles(),
testRunner: 'vitest',
reporters: ['progress', 'sarif'],
reporterOptions: {
sarif: { fileName: 'mutation-results.sarif' },
},
thresholds: {
high: 80,
low: 60,
break: 0, // Increase to 60 after establishing baseline; start as soft gate
},
concurrency: Number(process.env.STRYKER_CONCURRENCY) || 2,
timeoutMS: 30000,
excludedMutations: ['StringLiteral'],
};
The mutate array restricts Stryker to files changed in the PR, passed via the CHANGED_FILES environment variable from the workflow as a JSON array. The parseChangedFiles function safely handles malformed or missing JSON by falling back to the default glob rather than crashing with an unhandled SyntaxError. This scoping is essential for performance; running Stryker across an entire codebase is prohibitively slow for a CI gate. Set the break threshold to 0 initially (soft gate) to avoid blocking all merges before a baseline mutation score exists; increase it to 60 or higher after calibration. Configure the SARIF reporter output path under reporterOptions, not as a top-level key. The StringLiteral exclusion prevents false positives from mutations in log messages and UI text, which do not indicate test quality gaps. Teams should review excluded mutation types when adding new feature categories to the codebase to avoid masking genuine gaps.
Interpreting Stryker SARIF Output in PR Annotations
When Stryker outputs SARIF and it uploads to GitHub's code scanning, GitHub renders surviving mutants as inline annotations on the PR diff. Each annotation identifies the specific line, the type of mutation applied (conditional boundary, arithmetic operator, block statement removal), and the fact that tests failed to detect the change. In AI-generated code, conditional boundary mutants (changing < to <=) and arithmetic operator mutants (changing + to -) are the most common survivors because AI-authored tests frequently check only the happy path with a single representative input.
Wiring Tier 3 into the Workflow
tier3-mutation:
if: contains(github.event.pull_request.labels.*.name, 'ai-generated')
needs: tier2-property
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- uses: actions/setup-node@v4
with:
node-version: '22'
- run: npm ci
- name: Clean Stryker temp directory
run: rm -rf .stryker-tmp
- name: Set changed files
run: |
CHANGED=$(git diff --name-only "$BASE_SHA" HEAD | grep '\.ts$')
TOTAL=$(echo "$CHANGED" | grep -c '.' || true)
if [ "$TOTAL" -gt 20 ]; then
echo "::warning::Mutation scope truncated: $TOTAL changed .ts files found, only first 20 analyzed."
fi
SCOPED=$(echo "$CHANGED" | head -20 | jq -R . | jq -sc .)
echo "CHANGED_FILES=$SCOPED" >> "$GITHUB_ENV"
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
- run: npx stryker run
- uses: github/codeql-action/upload-sarif@v3
if: always()
with:
sarif_file: mutation-results.sarif
The timeout-minutes: 15 cap prevents Stryker from consuming unlimited runner time. The head -20 limit on changed files provides an additional safeguard against oversized PRs that would generate thousands of mutants. When more than 20 files change, the step emits a GitHub Actions ::warning:: annotation so reviewers know that mutation scope was truncated. The workflow sets CHANGED_FILES as a JSON array so that stryker.config.mjs can parse it safely. The rm -rf .stryker-tmp step prevents stale temporary files from a previous cancelled run from corrupting the current run.
The Complete Workflow: Assembling ai-pr-verify.yml
Full Workflow File
# .github/workflows/ai-pr-verify.yml
name: AI PR Verification Gates
on:
pull_request:
types: [opened, synchronize, labeled]
paths: ['src/**', 'tests/**']
permissions:
security-events: write
pull-requests: write
contents: read
concurrency:
group: ai-verify-${{ github.head_ref }}
cancel-in-progress: true
jobs:
tier1-ast:
if: contains(github.event.pull_request.labels.*.name, 'ai-generated')
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- uses: actions/setup-node@v4
with:
node-version: '22'
- uses: actions/cache@v4
with:
path: node_modules
key: ${{ runner.os }}-node-${{ hashFiles('package-lock.json') }}
restore-keys: |
${{ runner.os }}-node-
- run: npm ci
- run: node scripts/verify-pr-ast.mjs
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
- uses: github/codeql-action/upload-sarif@v3
if: always()
with:
sarif_file: ast-results.sarif
tier2-property:
if: contains(github.event.pull_request.labels.*.name, 'ai-generated')
needs: tier1-ast
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: '22'
- uses: actions/cache@v4
with:
path: node_modules
key: ${{ runner.os }}-node-${{ hashFiles('package-lock.json') }}
restore-keys: |
${{ runner.os }}-node-
- run: npm ci
- run: npx vitest run --reporter=json --outputFile=property-results.json 'tests/**/*.property.test.ts'
- uses: actions/github-script@v7
if: always()
with:
script: |
const fs = require('fs');
let summary = '**Tier 2 Property Tests**: results unavailable (file missing or unreadable)';
try {
const raw = fs.readFileSync('property-results.json', 'utf-8');
const results = JSON.parse(raw);
summary = `**Tier 2 Property Tests**: ${results.numPassedTests ?? 'N/A'} passed, ${results.numFailedTests ?? 'N/A'} failed`;
} catch (e) {
summary = `**Tier 2 Property Tests**: failed to read results — ${e.message}`;
}
github.rest.issues.createComment({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: context.issue.number,
body: summary,
});
tier3-mutation:
if: contains(github.event.pull_request.labels.*.name, 'ai-generated')
needs: tier2-property
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- uses: actions/setup-node@v4
with:
node-version: '22'
- uses: actions/cache@v4
with:
path: node_modules
key: ${{ runner.os }}-node-stryker-${{ hashFiles('package-lock.json', 'stryker.config.mjs') }}
restore-keys: |
${{ runner.os }}-node-stryker-
- run: npm ci
- name: Clean Stryker temp directory
run: rm -rf .stryker-tmp
- name: Set changed files
run: |
CHANGED=$(git diff --name-only "$BASE_SHA" HEAD | grep '\.ts$')
TOTAL=$(echo "$CHANGED" | grep -c '.' || true)
if [ "$TOTAL" -gt 20 ]; then
echo "::warning::Mutation scope truncated: $TOTAL changed .ts files found, only first 20 analyzed."
fi
SCOPED=$(echo "$CHANGED" | head -20 | jq -R . | jq -sc .)
echo "CHANGED_FILES=$SCOPED" >> "$GITHUB_ENV"
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
- run: npx stryker run
- uses: github/codeql-action/upload-sarif@v3
if: always()
with:
sarif_file: mutation-results.sarif
summary:
if: always() && contains(github.event.pull_request.labels.*.name, 'ai-generated')
needs: [tier1-ast, tier2-property, tier3-mutation]
runs-on: ubuntu-latest
steps:
- uses: actions/github-script@v7
env:
TIER1_RESULT: ${{ needs.tier1-ast.result }}
TIER2_RESULT: ${{ needs.tier2-property.result }}
TIER3_RESULT: ${{ needs.tier3-mutation.result }}
with:
script: |
const tiers = [
{ name: 'tier1-ast', result: process.env.TIER1_RESULT },
{ name: 'tier2-property', result: process.env.TIER2_RESULT },
{ name: 'tier3-mutation', result: process.env.TIER3_RESULT },
];
const rows = tiers.map(t => {
const status = t.result === 'success' ? '✅' : '❌';
return `| ${t.name} | ${status} |`;
});
const table = `| Tier | Status |
|---|---|
${rows.join('
')}`;
github.rest.issues.createComment({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: context.issue.number,
body: `### AI PR Verification Summary
${table}`,
});
Selective Triggering for AI-Authored PRs Only
Every job in the pipeline carries the if: contains(github.event.pull_request.labels.*.name, 'ai-generated') condition, ensuring the entire workflow only runs for labeled PRs. If tier1-ast is skipped because the label is absent, all downstream jobs are also skipped, not failed. Teams can automate this label via a separate lightweight workflow that detects PR authors matching known bot accounts or commit trailers indicating AI tool usage. This avoids penalizing human-authored PRs with extra gate latency while maintaining strict verification on AI output.
Tuning, Pitfalls, and Production Hardening
Managing Runner Costs and Timeout Budgets
Stryker is CPU-intensive by design. The timeout-minutes: 15 cap on the Tier 3 job prevents runaway costs. The concurrency group with cancel-in-progress: true ensures that pushing a new commit to the same PR branch cancels any in-flight verification runs rather than queuing them. Be aware that canceling a Stryker run mid-execution may leave temporary files in .stryker-tmp; the workflow includes a cleanup step (rm -rf .stryker-tmp) before each Stryker run to prevent this from causing failures. For large repositories, restricting the mutate glob to only changed files (via the CHANGED_FILES environment variable) and capping the number of files with head -20 keeps mutation testing within a practical time budget.
Common False Positives and How to Suppress Them
The AST analyzer will flag empty catch blocks that are intentional, such as those used in decorator patterns or accompanied by // eslint-disable comments. Adding a check for leading comments in the catch clause body allows the script to skip these cases. For Stryker, string-literal mutants in log messages and user-facing text generate noise without indicating test gaps. The excludedMutations: ['StringLiteral'] configuration in stryker.config.mjs suppresses these. Teams should review excluded mutation types when adding new feature categories to the codebase to avoid masking genuine gaps.
Extending the Pipeline
Note: The following extensions are out of scope for this tutorial and are mentioned only as potential future directions, not as implemented components of the workflow above.
Before any static analysis runs, a Tier 0 step could invoke an LLM-based self-review, asking the model to critique its own diff for logical inconsistencies. After mutation testing completes, a Tier 4 step could deploy runtime contract verification using zod schemas in a staging environment, validating that AI-generated functions conform to their declared input/output types under real traffic patterns. These extensions remain experimental but represent natural progressions of the verification-gate model.
Trusting AI Code Through Verification, Not Hope
Coverage metrics lie. Structural analysis, property invariants, and mutation scores tell the truth.
Coverage metrics lie. Structural analysis, property invariants, and mutation scores tell the truth. Adjust the AST detection rules to match your codebase's patterns, calibrate the mutation score thresholds against your existing test quality, and deploy the pipeline against the next AI-authored pull request.

