swebench

package
v0.44.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 7, 2026 License: Apache-2.0 Imports: 14 Imported by: 0

Documentation

Overview

Package swebench exchanges predictions and reports with the official SWE-bench harness. It does not execute candidates, parse test logs, or grade patches. The Host supplies a frozen dataset, runs the pinned harness, and grants read access to its completed artifacts.

NewSelection freezes dataset content identity and the exact denominator. NewSubmission binds that selection, an explicit attempt, and exact candidate bytes to a content-derived RunID. Export predictions with WritePredictions, run the official harness using that RunID and the selected instance IDs, then import the complete run with Collect. A regrade requires a new AttemptID.

Task digests must cover the complete task and grading environment contract: problem input, base commit, test patch, expected tests, parser, evaluation script, consumed image digest and platform, and execution policy. Keep gold patches, hidden tests, and grading assets out of the candidate's input. These identities bind artifacts; this package cannot attest what a container ran. The Host must make the actual harness consume pinned images and dependencies.

Index

Examples

Constants

View Source
const HarnessRevision = "02e7a74ffd0b707aab73d203fe87bdc7c76afc8e"

HarnessRevision identifies the upstream protocol implemented by this package. Its report and cache behavior are defined in swebench/harness/{grading, reporting,run_evaluation}.py in github.com/SWE-bench/SWE-bench at this commit.

Variables

View Source
var (
	ErrInvalidSelection  = errors.New("eval/swebench: invalid selection")
	ErrInvalidSubmission = errors.New("eval/swebench: invalid submission")
	ErrInvalidArtifacts  = errors.New("eval/swebench: invalid artifacts")
	ErrArtifactMismatch  = errors.New("eval/swebench: artifact identity mismatch")
	ErrNoGrade           = errors.New("eval/swebench: no official grade")
)

Functions

This section is empty.

Types

type CaseResult

type CaseResult struct {
	// contains filtered or unexported fields
}

CaseResult is immutable evidence for one selected task. Completed reflects the upstream summary's report-file-exists count: a malformed file can be completed and still StatusError. Only an imported, valid OfficialReport owns an evaluation result. Candidate and report digests refer to exact bytes.

func (CaseResult) CandidateDigest

func (c CaseResult) CandidateDigest() string

func (CaseResult) Completed

func (c CaseResult) Completed() bool

func (CaseResult) Err

func (c CaseResult) Err() error

func (CaseResult) Failure

func (c CaseResult) Failure() Failure

func (CaseResult) GradeID

func (c CaseResult) GradeID() string

GradeID binds the exact official report to its candidate, task, selection, harness, and attempt. It is empty when no valid official grade was imported.

func (CaseResult) Official

func (c CaseResult) Official() (OfficialReport, bool)

func (CaseResult) Report

func (c CaseResult) Report() (eval.Report, error)

Report projects an imported official decision into eval's general contract. Missing predictions, empty patches, and harness errors return ErrNoGrade; they do not manufacture zero quality scores. A legitimate unresolved report has score zero even when the official harness also advises an infra failure. The mean of available grades is conditional on grade availability; the official resolution rate instead divides Summary.Resolved by Summary.Total.

func (CaseResult) ReportDigest

func (c CaseResult) ReportDigest() string

func (CaseResult) Status

func (c CaseResult) Status() Status

func (CaseResult) Task

func (c CaseResult) Task() Task

type Collection

type Collection struct {
	// contains filtered or unexported fields
}

Collection owns results in the Selection's stable instance order. It can be constructed only by importing and reconciling the official complete run.

func (Collection) Cases

func (c Collection) Cases() []CaseResult

func (Collection) Summary

func (c Collection) Summary() Summary

type Failure

type Failure struct {
	Kind   FailureKind
	Reason string
}

Failure preserves the official run summary's additive diagnostic advice. The harness can classify logs for an error even when no case report exists.

type FailureKind

type FailureKind string
const (
	FailureUnspecified    FailureKind = ""
	FailureInfrastructure FailureKind = "infrastructure"
	FailureAmbiguous      FailureKind = "ambiguous"
)

type OfficialReport

type OfficialReport struct {
	PatchIsNone              bool
	PatchExists              bool
	PatchSuccessfullyApplied bool
	Resolved                 bool
	InfraFailure             bool
	InfraFailureReason       string
	TestsStatus              *TestsStatus
}

OfficialReport is a detached view of one imported harness report. In the pinned harness, PatchSuccessfullyApplied is set after test logs are accepted; it is not merely the patch command's exit status. InfraFailure and its reason are the case report's advice; CaseResult.Failure preserves the independently generated run summary's classification, which can also inspect harness logs.

type Prediction

type Prediction struct {
	InstanceID string
	Patch      string
}

Prediction preserves the exact UTF-8 candidate patch. An empty string is an explicit empty submission; an absent Prediction is a missing submission. Trials are separate Submissions, never duplicate instance IDs in one JSONL.

func (Prediction) Digest

func (p Prediction) Digest() string

type Selection

type Selection struct {
	// contains filtered or unexported fields
}

Selection is an immutable set ordered by instance ID. Caller mutations of the constructor input or Tasks result cannot change its identity.

func NewSelection

func NewSelection(config SelectionConfig) (Selection, error)

func (Selection) Dataset

func (s Selection) Dataset() string

func (Selection) Digest

func (s Selection) Digest() string

func (Selection) Len

func (s Selection) Len() int

func (Selection) Revision

func (s Selection) Revision() string

func (Selection) Split

func (s Selection) Split() string

func (Selection) Tasks

func (s Selection) Tasks() []Task

type SelectionConfig

type SelectionConfig struct {
	Dataset  string
	Revision string
	Split    string
	Tasks    []Task
}

SelectionConfig identifies content independently from an upstream wire schema version. Revision is an immutable 40- or 64-digit hexadecimal commit or content revision; names such as main and latest cannot freeze a dataset.

type Status

type Status string

Status separates an official quality outcome from absent predictions and harness failures. Infrastructure advice never changes a resolved denominator or turns an official unresolved result into an absent grade.

const (
	StatusMissingPrediction Status = "missing_prediction"
	StatusEmptyPatch        Status = "empty_patch"
	StatusError             Status = "error"
	StatusResolved          Status = "resolved"
	StatusUnresolved        Status = "unresolved"
)

type Submission

type Submission struct {
	// contains filtered or unexported fields
}

Submission owns one official run's identity. AttemptID identifies a new execution or regrade even when all candidate bytes are unchanged. Model may contain slashes; after the harness replaces them with __ it must be a safe log-directory component. The full, unmodified model name participates in the run identity so that the upstream directory encoding cannot alias two runs.

func NewSubmission

func NewSubmission(config SubmissionConfig) (Submission, error)

func (Submission) Arguments

func (s Submission) Arguments(datasetPath, predictionsPath string) ([]string, error)

Arguments returns the official harness module and selection flags as separate process arguments, never shell text. datasetPath must name a frozen dataset matching Selection and predictionsPath must hold WritePredictions output.

func (Submission) AttemptID

func (s Submission) AttemptID() string

func (Submission) Collect

func (s Submission) Collect(artifacts fs.FS) (Collection, error)

Collect imports a completed official run from a trusted, immutable snapshot of the harness working directory. Confine disk reads with os.OpenRoot and Root.FS; os.DirFS alone does not prevent symlink escapes.

Collect requires one results.json covering the entire Selection; for cases run separately, invoke the official make_run_report once over all of them. Missing or malformed case reports remain StatusError without a grade, while a candidate mismatch, an unbound report, or an inconsistent summary rejects the collection. Matching artifacts do not prove what the harness executed.

Example
package main

import (
	"bytes"
	"context"
	"crypto/sha256"

	jsonv2 "encoding/json/v2"
	"fmt"
	"path"
	"testing/fstest"

	"github.com/Tangerg/scope/eval"
	"github.com/Tangerg/scope/eval/swebench"
)

func main() {
	check := func(err error) {
		if err != nil {
			panic(err)
		}
	}
	ctx := context.Background()
	// A Host solver has finished collecting a frozen patch, independently of
	// grading. Real task identity covers the full task and environment contract.
	patch := "diff --git a/code.py b/code.py\n--- a/code.py\n+++ b/code.py\n@@ -1 +1 @@\n-old\n+new\n"
	execution := eval.Execution[string]{Status: eval.ExecutionCompleted, Output: &patch}
	check(execution.Validate())
	selection, err := swebench.NewSelection(swebench.SelectionConfig{
		Dataset:  "SWE-bench/SWE-bench_Verified",
		Revision: "3d07b464b7b311a0cbfb5ed5b2d8a3b96f84a33d", Split: "test",
		Tasks: []swebench.Task{{InstanceID: "case-resolved", Digest: fmt.Sprintf("sha256:%x", sha256.Sum256([]byte("fixture task and environment")))}},
	})
	check(err)
	submission, err := swebench.NewSubmission(swebench.SubmissionConfig{
		Selection: selection, Model: "host/solver", AttemptID: "trial-1",
		Predictions: []swebench.Prediction{{InstanceID: "case-resolved", Patch: *execution.Output}},
	})
	check(err)
	var predictions bytes.Buffer
	check(submission.WritePredictions(&predictions))
	arguments, err := submission.Arguments("frozen-dataset.json", "predictions.jsonl")
	check(err)
	fmt.Println("harness:", arguments[1])

	// A production Host writes predictions, runs HarnessRevision using the
	// arguments, and collects an immutable artifact snapshot. This checked
	// example uses official-protocol fixtures; it executes no Docker or model.
	var row struct {
		InstanceID string `json:"instance_id"`
		Model      string `json:"model_name_or_path"`
		Patch      string `json:"model_patch"`
	}
	check(jsonv2.Unmarshal(predictions.Bytes(), &row))
	runDirectory := path.Join("logs", "evaluation", submission.RunID())
	caseDirectory := path.Join(runDirectory, "host__solver", row.InstanceID)
	artifacts := fstest.MapFS{
		path.Join(caseDirectory, "patch.diff"): {Data: []byte(row.Patch)},
		path.Join(caseDirectory, "report.json"): {Data: []byte(`{
		  "case-resolved": {
		    "patch_is_None": false, "patch_exists": true,
		    "patch_successfully_applied": true, "resolved": true, "infra_failure": false,
		    "tests_status": {
		      "FAIL_TO_PASS": {"success": ["test_fixed"], "failure": []},
		      "PASS_TO_PASS": {"success": ["test_existing"], "failure": []},
		      "FAIL_TO_FAIL": {"success": [], "failure": []},
		      "PASS_TO_FAIL": {"success": [], "failure": []}
		    }
		  }
		}`)},
		path.Join(runDirectory, "results.json"): {Data: []byte(`{
		  "total_instances": 1, "submitted_instances": 1, "completed_instances": 1,
		  "resolved_instances": 1, "unresolved_instances": 0,
		  "infra_failure_instances": 0, "ambiguous_failure_instances": 0,
		  "empty_patch_instances": 0, "error_instances": 0,
		  "completed_ids": ["case-resolved"], "incomplete_ids": [], "empty_patch_ids": [],
		  "submitted_ids": ["case-resolved"], "resolved_ids": ["case-resolved"],
		  "unresolved_ids": [], "infra_failure_ids": [], "ambiguous_failure_ids": [],
		  "failure_reasons": {}, "error_ids": [], "schema_version": 2
		}`)},
	}
	collection, err := submission.Collect(artifacts)
	check(err)

	// Assessment identity remains declared even when a case has no grade.
	// Reassessment of collected evidence does not invoke the target again.
	suite, err := eval.NewSuite(eval.SuiteConfig[swebench.CaseResult]{
		Assessments: []eval.Assessment[swebench.CaseResult]{{
			ID: "official-resolution",
			Evaluator: eval.EvaluatorFunc[swebench.CaseResult](func(ctx context.Context, result swebench.CaseResult) (eval.Report, error) {
				if contextErr := ctx.Err(); contextErr != nil {
					return eval.Report{}, contextErr
				}
				return result.Report()
			}),
		}},
	})
	check(err)
	var cases []eval.Case[swebench.CaseResult]
	for _, result := range collection.Cases() {
		cases = append(cases, eval.Case[swebench.CaseResult]{ID: eval.CaseID(result.Task().InstanceID), Subject: result})
	}
	dataset, err := eval.NewDataset(selection.Digest(), cases...)
	check(err)
	experiment, err := eval.NewExperiment(eval.ExperimentConfig[swebench.CaseResult]{Dataset: dataset, Suite: suite})
	check(err)
	report, err := experiment.Run(ctx)
	check(err)
	fmt.Printf("official: %d/%d resolved\n", collection.Summary().Resolved, collection.Summary().Total)
	fmt.Println("assessment:", report.Cases()[0].Result.Results[0].Status())
	fmt.Println("decision:", report.Cases()[0].Result.Results[0].Report.Verdict())
}
Output:
harness: swebench.harness.run_evaluation
official: 1/1 resolved
assessment: completed
decision: pass

func (Submission) Model

func (s Submission) Model() string

func (Submission) Predictions

func (s Submission) Predictions() []Prediction

func (Submission) RunID

func (s Submission) RunID() string

func (Submission) Selection

func (s Submission) Selection() Selection

func (Submission) WritePredictions

func (s Submission) WritePredictions(writer io.Writer) error

WritePredictions encodes every row of the official JSONL before writing any bytes and never normalizes patch whitespace. An I/O failure can still leave partial output; the Host owns atomic publication.

type SubmissionConfig

type SubmissionConfig struct {
	Selection   Selection
	Model       string
	AttemptID   string
	Predictions []Prediction
}

SubmissionConfig binds one attempt to immutable selection and candidate values. NewSubmission snapshots Predictions; omitted cases remain missing.

type Summary

type Summary struct {
	Total                  int
	Submitted              int
	Completed              int
	Resolved               int
	Unresolved             int
	MissingPredictions     int
	EmptyPatches           int
	Errors                 int
	InfrastructureFailures int
	AmbiguousFailures      int
}

Summary preserves official counts for the complete frozen selection. InfrastructureFailures and AmbiguousFailures overlap Unresolved or Errors; neither removes tasks from Total. Completed can include malformed reports. This is a validated import of results.json, not a replacement benchmark score.

type Task

type Task struct {
	InstanceID string `json:"instance_id"`
	Digest     string `json:"digest"`
}

Task binds an official instance ID to the SHA-256 digest of its complete task and grading contract, written as sha256:.

type TestResults

type TestResults struct {
	Success []string `json:"success"`
	Failure []string `json:"failure"`
}

TestResults contains the official grader's test identities without recalculating its parser-specific resolution or maintenance rules.

type TestsStatus

type TestsStatus struct {
	FailToPass TestResults `json:"FAIL_TO_PASS"`
	PassToPass TestResults `json:"PASS_TO_PASS"`
	FailToFail TestResults `json:"FAIL_TO_FAIL"`
	PassToFail TestResults `json:"PASS_TO_FAIL"`
}

TestsStatus retains the four official test-transition groups. Values returned by CaseResult.Official own their slices and may be edited by the caller.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL