Documentation
¶
Overview ¶
Package swebench exchanges predictions and reports with the official SWE-bench harness. It does not execute candidates, parse test logs, or grade patches. The Host supplies a frozen dataset, runs the pinned harness, and grants read access to its completed artifacts.
NewSelection freezes dataset content identity and the exact denominator. NewSubmission binds that selection, an explicit attempt, and exact candidate bytes to a content-derived RunID. Export predictions with WritePredictions, run the official harness using that RunID and the selected instance IDs, then import the complete run with Collect. A regrade requires a new AttemptID.
Task digests must cover the complete task and grading environment contract: problem input, base commit, test patch, expected tests, parser, evaluation script, consumed image digest and platform, and execution policy. Keep gold patches, hidden tests, and grading assets out of the candidate's input. These identities bind artifacts; this package cannot attest what a container ran. The Host must make the actual harness consume pinned images and dependencies.
Index ¶
- Constants
- Variables
- type CaseResult
- func (c CaseResult) CandidateDigest() string
- func (c CaseResult) Completed() bool
- func (c CaseResult) Err() error
- func (c CaseResult) Failure() Failure
- func (c CaseResult) GradeID() string
- func (c CaseResult) Official() (OfficialReport, bool)
- func (c CaseResult) Report() (eval.Report, error)
- func (c CaseResult) ReportDigest() string
- func (c CaseResult) Status() Status
- func (c CaseResult) Task() Task
- type Collection
- type Failure
- type FailureKind
- type OfficialReport
- type Prediction
- type Selection
- type SelectionConfig
- type Status
- type Submission
- func (s Submission) Arguments(datasetPath, predictionsPath string) ([]string, error)
- func (s Submission) AttemptID() string
- func (s Submission) Collect(artifacts fs.FS) (Collection, error)
- func (s Submission) Model() string
- func (s Submission) Predictions() []Prediction
- func (s Submission) RunID() string
- func (s Submission) Selection() Selection
- func (s Submission) WritePredictions(writer io.Writer) error
- type SubmissionConfig
- type Summary
- type Task
- type TestResults
- type TestsStatus
Examples ¶
Constants ¶
const HarnessRevision = "02e7a74ffd0b707aab73d203fe87bdc7c76afc8e"
HarnessRevision identifies the upstream protocol implemented by this package. Its report and cache behavior are defined in swebench/harness/{grading, reporting,run_evaluation}.py in github.com/SWE-bench/SWE-bench at this commit.
Variables ¶
var ( ErrInvalidSelection = errors.New("eval/swebench: invalid selection") ErrInvalidSubmission = errors.New("eval/swebench: invalid submission") ErrInvalidArtifacts = errors.New("eval/swebench: invalid artifacts") ErrArtifactMismatch = errors.New("eval/swebench: artifact identity mismatch") ErrNoGrade = errors.New("eval/swebench: no official grade") )
Functions ¶
This section is empty.
Types ¶
type CaseResult ¶
type CaseResult struct {
// contains filtered or unexported fields
}
CaseResult is immutable evidence for one selected task. Completed reflects the upstream summary's report-file-exists count: a malformed file can be completed and still StatusError. Only an imported, valid OfficialReport owns an evaluation result. Candidate and report digests refer to exact bytes.
func (CaseResult) CandidateDigest ¶
func (c CaseResult) CandidateDigest() string
func (CaseResult) Completed ¶
func (c CaseResult) Completed() bool
func (CaseResult) Err ¶
func (c CaseResult) Err() error
func (CaseResult) Failure ¶
func (c CaseResult) Failure() Failure
func (CaseResult) GradeID ¶
func (c CaseResult) GradeID() string
GradeID binds the exact official report to its candidate, task, selection, harness, and attempt. It is empty when no valid official grade was imported.
func (CaseResult) Official ¶
func (c CaseResult) Official() (OfficialReport, bool)
func (CaseResult) Report ¶
func (c CaseResult) Report() (eval.Report, error)
Report projects an imported official decision into eval's general contract. Missing predictions, empty patches, and harness errors return ErrNoGrade; they do not manufacture zero quality scores. A legitimate unresolved report has score zero even when the official harness also advises an infra failure. The mean of available grades is conditional on grade availability; the official resolution rate instead divides Summary.Resolved by Summary.Total.
func (CaseResult) ReportDigest ¶
func (c CaseResult) ReportDigest() string
func (CaseResult) Status ¶
func (c CaseResult) Status() Status
func (CaseResult) Task ¶
func (c CaseResult) Task() Task
type Collection ¶
type Collection struct {
// contains filtered or unexported fields
}
Collection owns results in the Selection's stable instance order. It can be constructed only by importing and reconciling the official complete run.
func (Collection) Cases ¶
func (c Collection) Cases() []CaseResult
func (Collection) Summary ¶
func (c Collection) Summary() Summary
type Failure ¶
type Failure struct {
Kind FailureKind
Reason string
}
Failure preserves the official run summary's additive diagnostic advice. The harness can classify logs for an error even when no case report exists.
type FailureKind ¶
type FailureKind string
const ( FailureUnspecified FailureKind = "" FailureInfrastructure FailureKind = "infrastructure" FailureAmbiguous FailureKind = "ambiguous" )
type OfficialReport ¶
type OfficialReport struct {
PatchIsNone bool
PatchExists bool
PatchSuccessfullyApplied bool
Resolved bool
InfraFailure bool
InfraFailureReason string
TestsStatus *TestsStatus
}
OfficialReport is a detached view of one imported harness report. In the pinned harness, PatchSuccessfullyApplied is set after test logs are accepted; it is not merely the patch command's exit status. InfraFailure and its reason are the case report's advice; CaseResult.Failure preserves the independently generated run summary's classification, which can also inspect harness logs.
type Prediction ¶
Prediction preserves the exact UTF-8 candidate patch. An empty string is an explicit empty submission; an absent Prediction is a missing submission. Trials are separate Submissions, never duplicate instance IDs in one JSONL.
func (Prediction) Digest ¶
func (p Prediction) Digest() string
type Selection ¶
type Selection struct {
// contains filtered or unexported fields
}
Selection is an immutable set ordered by instance ID. Caller mutations of the constructor input or Tasks result cannot change its identity.
func NewSelection ¶
func NewSelection(config SelectionConfig) (Selection, error)
type SelectionConfig ¶
SelectionConfig identifies content independently from an upstream wire schema version. Revision is an immutable 40- or 64-digit hexadecimal commit or content revision; names such as main and latest cannot freeze a dataset.
type Status ¶
type Status string
Status separates an official quality outcome from absent predictions and harness failures. Infrastructure advice never changes a resolved denominator or turns an official unresolved result into an absent grade.
type Submission ¶
type Submission struct {
// contains filtered or unexported fields
}
Submission owns one official run's identity. AttemptID identifies a new execution or regrade even when all candidate bytes are unchanged. Model may contain slashes; after the harness replaces them with __ it must be a safe log-directory component. The full, unmodified model name participates in the run identity so that the upstream directory encoding cannot alias two runs.
func NewSubmission ¶
func NewSubmission(config SubmissionConfig) (Submission, error)
func (Submission) Arguments ¶
func (s Submission) Arguments(datasetPath, predictionsPath string) ([]string, error)
Arguments returns the official harness module and selection flags as separate process arguments, never shell text. datasetPath must name a frozen dataset matching Selection and predictionsPath must hold WritePredictions output.
func (Submission) AttemptID ¶
func (s Submission) AttemptID() string
func (Submission) Collect ¶
func (s Submission) Collect(artifacts fs.FS) (Collection, error)
Collect imports a completed official run from a trusted, immutable snapshot of the harness working directory. Confine disk reads with os.OpenRoot and Root.FS; os.DirFS alone does not prevent symlink escapes.
Collect requires one results.json covering the entire Selection; for cases run separately, invoke the official make_run_report once over all of them. Missing or malformed case reports remain StatusError without a grade, while a candidate mismatch, an unbound report, or an inconsistent summary rejects the collection. Matching artifacts do not prove what the harness executed.
Example ¶
package main
import (
"bytes"
"context"
"crypto/sha256"
jsonv2 "encoding/json/v2"
"fmt"
"path"
"testing/fstest"
"github.com/Tangerg/scope/eval"
"github.com/Tangerg/scope/eval/swebench"
)
func main() {
check := func(err error) {
if err != nil {
panic(err)
}
}
ctx := context.Background()
// A Host solver has finished collecting a frozen patch, independently of
// grading. Real task identity covers the full task and environment contract.
patch := "diff --git a/code.py b/code.py\n--- a/code.py\n+++ b/code.py\n@@ -1 +1 @@\n-old\n+new\n"
execution := eval.Execution[string]{Status: eval.ExecutionCompleted, Output: &patch}
check(execution.Validate())
selection, err := swebench.NewSelection(swebench.SelectionConfig{
Dataset: "SWE-bench/SWE-bench_Verified",
Revision: "3d07b464b7b311a0cbfb5ed5b2d8a3b96f84a33d", Split: "test",
Tasks: []swebench.Task{{InstanceID: "case-resolved", Digest: fmt.Sprintf("sha256:%x", sha256.Sum256([]byte("fixture task and environment")))}},
})
check(err)
submission, err := swebench.NewSubmission(swebench.SubmissionConfig{
Selection: selection, Model: "host/solver", AttemptID: "trial-1",
Predictions: []swebench.Prediction{{InstanceID: "case-resolved", Patch: *execution.Output}},
})
check(err)
var predictions bytes.Buffer
check(submission.WritePredictions(&predictions))
arguments, err := submission.Arguments("frozen-dataset.json", "predictions.jsonl")
check(err)
fmt.Println("harness:", arguments[1])
// A production Host writes predictions, runs HarnessRevision using the
// arguments, and collects an immutable artifact snapshot. This checked
// example uses official-protocol fixtures; it executes no Docker or model.
var row struct {
InstanceID string `json:"instance_id"`
Model string `json:"model_name_or_path"`
Patch string `json:"model_patch"`
}
check(jsonv2.Unmarshal(predictions.Bytes(), &row))
runDirectory := path.Join("logs", "evaluation", submission.RunID())
caseDirectory := path.Join(runDirectory, "host__solver", row.InstanceID)
artifacts := fstest.MapFS{
path.Join(caseDirectory, "patch.diff"): {Data: []byte(row.Patch)},
path.Join(caseDirectory, "report.json"): {Data: []byte(`{
"case-resolved": {
"patch_is_None": false, "patch_exists": true,
"patch_successfully_applied": true, "resolved": true, "infra_failure": false,
"tests_status": {
"FAIL_TO_PASS": {"success": ["test_fixed"], "failure": []},
"PASS_TO_PASS": {"success": ["test_existing"], "failure": []},
"FAIL_TO_FAIL": {"success": [], "failure": []},
"PASS_TO_FAIL": {"success": [], "failure": []}
}
}
}`)},
path.Join(runDirectory, "results.json"): {Data: []byte(`{
"total_instances": 1, "submitted_instances": 1, "completed_instances": 1,
"resolved_instances": 1, "unresolved_instances": 0,
"infra_failure_instances": 0, "ambiguous_failure_instances": 0,
"empty_patch_instances": 0, "error_instances": 0,
"completed_ids": ["case-resolved"], "incomplete_ids": [], "empty_patch_ids": [],
"submitted_ids": ["case-resolved"], "resolved_ids": ["case-resolved"],
"unresolved_ids": [], "infra_failure_ids": [], "ambiguous_failure_ids": [],
"failure_reasons": {}, "error_ids": [], "schema_version": 2
}`)},
}
collection, err := submission.Collect(artifacts)
check(err)
// Assessment identity remains declared even when a case has no grade.
// Reassessment of collected evidence does not invoke the target again.
suite, err := eval.NewSuite(eval.SuiteConfig[swebench.CaseResult]{
Assessments: []eval.Assessment[swebench.CaseResult]{{
ID: "official-resolution",
Evaluator: eval.EvaluatorFunc[swebench.CaseResult](func(ctx context.Context, result swebench.CaseResult) (eval.Report, error) {
if contextErr := ctx.Err(); contextErr != nil {
return eval.Report{}, contextErr
}
return result.Report()
}),
}},
})
check(err)
var cases []eval.Case[swebench.CaseResult]
for _, result := range collection.Cases() {
cases = append(cases, eval.Case[swebench.CaseResult]{ID: eval.CaseID(result.Task().InstanceID), Subject: result})
}
dataset, err := eval.NewDataset(selection.Digest(), cases...)
check(err)
experiment, err := eval.NewExperiment(eval.ExperimentConfig[swebench.CaseResult]{Dataset: dataset, Suite: suite})
check(err)
report, err := experiment.Run(ctx)
check(err)
fmt.Printf("official: %d/%d resolved\n", collection.Summary().Resolved, collection.Summary().Total)
fmt.Println("assessment:", report.Cases()[0].Result.Results[0].Status())
fmt.Println("decision:", report.Cases()[0].Result.Results[0].Report.Verdict())
}
Output: harness: swebench.harness.run_evaluation official: 1/1 resolved assessment: completed decision: pass
func (Submission) Model ¶
func (s Submission) Model() string
func (Submission) Predictions ¶
func (s Submission) Predictions() []Prediction
func (Submission) RunID ¶
func (s Submission) RunID() string
func (Submission) Selection ¶
func (s Submission) Selection() Selection
func (Submission) WritePredictions ¶
func (s Submission) WritePredictions(writer io.Writer) error
WritePredictions encodes every row of the official JSONL before writing any bytes and never normalizes patch whitespace. An I/O failure can still leave partial output; the Host owns atomic publication.
type SubmissionConfig ¶
type SubmissionConfig struct {
Selection Selection
Model string
AttemptID string
Predictions []Prediction
}
SubmissionConfig binds one attempt to immutable selection and candidate values. NewSubmission snapshots Predictions; omitted cases remain missing.
type Summary ¶
type Summary struct {
Total int
Submitted int
Completed int
Resolved int
Unresolved int
MissingPredictions int
EmptyPatches int
Errors int
InfrastructureFailures int
AmbiguousFailures int
}
Summary preserves official counts for the complete frozen selection. InfrastructureFailures and AmbiguousFailures overlap Unresolved or Errors; neither removes tasks from Total. Completed can include malformed reports. This is a validated import of results.json, not a replacement benchmark score.
type Task ¶
Task binds an official instance ID to the SHA-256 digest of its complete
task and grading contract, written as sha256:
type TestResults ¶
TestResults contains the official grader's test identities without recalculating its parser-specific resolution or maintenance rules.
type TestsStatus ¶
type TestsStatus struct {
FailToPass TestResults `json:"FAIL_TO_PASS"`
PassToPass TestResults `json:"PASS_TO_PASS"`
FailToFail TestResults `json:"FAIL_TO_FAIL"`
PassToFail TestResults `json:"PASS_TO_FAIL"`
}
TestsStatus retains the four official test-transition groups. Values returned by CaseResult.Official own their slices and may be edited by the caller.