draft-detective

Eval Scores Report

Current Inspect AI eval numbers across every eval in evals_inspectai/e2e/.

[!IMPORTANT] This is the current baseline, measured on gpt-5.6-terra. On 28 Aug 2026 every agent moved off the three-tier gpt-5.4-mini / gpt-5.4 / gpt-5.5 stack onto a single model. The numbers those tiers last scored are kept in docs/eval-scores-gpt-5.4-5.5.md for comparison; they no longer describe the running system.

[!NOTE] All evals are end-to-end: they trigger the real workflow through the API, so the backend must be running (uv run dev.py). Scorer numbers are accuracy (mean over samples, 0.001.00), shown as value ±stderr.

Run an eval with: uv run inspect eval evals_inspectai/e2e/<eval>/<eval>_e2e.py --epochs=<n>

The raw Inspect .eval log for each recorded run is copied under docs/evals/ and linked in the Log column. Open one with uv run inspect view --log-dir docs/evals.

Results

Each eval defines its own scorers. The Scorer results column lists every scorer that eval runs, by its Inspect name, with the per-scorer accuracy — see the footnotes for what each scorer checks. Expand a cell to read them; the summary line gives the count and the spread on each side of the deterministic / model-graded split, which is what Overall avg flattens. That average is computed over every metric in the cell, not only the ones a reader expands. Overall avg is the unweighted mean of that eval’s per-scorer accuracies (a rough headline number: scorers measure different things, so it is not a rigorous aggregate). Model-graded scores use openai/gpt-5.4 as the grader, except the two review-assistant suites (17 and 18), which pin openai/gpt-5.6-terra.

Three suites — results_extraction (15) and the two review-assistant ones (17 and 18) — report one metric per check rather than a single blended score, so every rule they enforce is listed separately. A broken rule is a defect and a lower judged score is a trend; averaging them together hides both.

Every run below completed in full: no sample was dropped, errored or retried.

# Eval Samples Epochs Scorer results Overall avg Date Log
1 abbreviation_checker 26 3
<summary>2 deterministic 0.999–1.000 · 1 judged 0.987</summary>structured_output_scorer 0.999 ±0.0011
structured_output_scorer1 1.000 ±0.0002
model_graded_check 0.987 ±0.0093
0.995 2026-08-26 …_9L7LFJgP4rM2Z6oAmDgCg8.eval
2 about_this_ger 13 3
<summary>2 deterministic 0.983–0.987 · 1 judged 0.936</summary>structured_output_scorer 0.983 ±0.0174
structured_output_scorer1 0.987 ±0.0135
model_graded_check 0.936 ±0.0523
0.969 2026-08-26 …_RDABwXSRrsEUEeTGCDyPAX.eval
3 advocacy_tone_v2 14 3
<summary>1 deterministic 1.000 · 1 judged 1.000</summary>structured_output_scorer 1.000 ±0.0006
model_graded_check 1.000 ±0.0003
1.000 2026-08-26 …_TR6MVXLpKk5aABsjrrd3BG.eval
4 claim_reference_validation_v2 7 3
<summary>2 deterministic all 1.000 · 1 judged 1.000</summary>citation_alignment_match 1.000 ±0.0007
citation_count_match 1.000 ±0.0008
model_graded_check 1.000 ±0.0003
1.000 2026-08-26 …_P6c7epQuUUuYnrhLJszKwY.eval
5 document_structure 5 3
<summary>1 deterministic 0.933 · 1 judged 0.800</summary>structured_output_scorer 0.933 ±0.0679
model_graded_check 0.800 ±0.2003
0.867 2026-08-28 …_gVQDPn5E5ru4cJJdYpuGBk.eval
6 figures_tables_check 19 3
<summary>1 deterministic 0.710 · 1 judged 0.851</summary>structured_output_scorer 0.710 ±0.0689
model_graded_check 0.851 ±0.0443
0.780 2026-08-28 …_Z9HXQpWvUiJedwbjCZCXDn.eval
7 inference_validation_v2 6 3
<summary>1 deterministic 1.000 · 1 judged 1.000</summary>structured_output_scorer 1.000 ±0.00010
model_graded_check 1.000 ±0.0003
1.000 2026-08-27 …_ETnaSF9iyWUPQYf6Ac7VHm.eval
8 literature_review_v2 4 3
<summary>1 deterministic 1.000 · 1 judged 0.958</summary>structured_output_scorer 1.000 ±0.00011
model_graded_check 0.958 ±0.0423
0.979 2026-08-26 …_jkRHzj4yivsY66vtVbCo4a.eval
9 live_reports_v2 3 3
<summary>1 deterministic 1.000 · 1 judged 1.000</summary>structured_output_scorer 1.000 ±0.00011
model_graded_check 1.000 ±0.0003
1.000 2026-08-26 …_B8DravaK7UWAGrngodBEco.eval
10 methodological_alignment 2 3
<summary>1 deterministic 1.000 · 1 judged 1.000</summary>structured_output_scorer 1.000 ±0.00012
model_graded_check 1.000 ±0.0003
1.000 2026-08-26 …_gUqTHjY4RMC2DJyPMuLers.eval
11 recommendation_check 6 3
<summary>1 deterministic 0.988 · 1 judged 0.917</summary>structured_output_scorer 0.988 ±0.01213
model_graded_check 0.917 ±0.0833
0.953 2026-08-26 …_4PUbGQZhFYt9PfBHf8TyXe.eval
12 reference_downloader 31 3
<summary>1 deterministic 0.849</summary>structured_output_scorer 0.849 ±0.06014
0.849 2026-08-28 …_YHhh9v2dMmLZVkBdoGLaNi.eval
13 reference_text_extractor 7 3
<summary>1 deterministic 0.878</summary>structured_output_scorer 0.878 ±0.08415
0.878 2026-08-28 …_EcCoBLxfCqnUGhVwdMwCzR.eval
14 reference_validation_v2 70 3
<summary>1 deterministic 0.814 · 1 judged 0.824</summary>structured_output_scorer 0.814 ±0.04416
model_graded_check 0.824 ±0.0333
0.819 2026-08-26 …_LVNQd5h6f5bUWnD2eYogUb.eval
15 results_extraction17 10 3
<summary>11 deterministic 0.942–1.000 · 2 judged 0.867–0.950</summary>report 1.000 ±0.000
inventory_table 1.000 ±0.000
result_count 0.967 ±0.033
labels 1.000 ±0.000
severity_split 1.000 ±0.000
line_ranges 1.000 ±0.000
no_duplicates 1.000 ±0.00018
completeness 0.992 ±0.008
class_accuracy 0.942 ±0.039
no_extras 0.979 ±0.011
severity_ordering 1.000 ±0.00019
classification_grounded 0.950 ±0.025
sample_expectations 0.867 ±0.06920
0.977 2026-08-31 …_ZVwrG4t2tTbjbj6YbvCsWt.eval
16 reviewer_2 2 3
<summary>1 deterministic 1.000 · 1 judged 1.000</summary>structured_output_scorer 1.000 ±0.00021
model_graded_check 1.000 ±0.0003
1.000 2026-08-26 …_i3Bm2FbzhmJAwqas6MUyri.eval
17 reviewer_coverage_report 5 3
<summary>9 deterministic all 1.000 · 4 judged 0.800–0.967</summary>verbatim 1.000 ±0.000
quoted 1.000 ±0.000
id_scheme 1.000 ±0.000
self_contained 1.000 ±0.000
two_part_layout 1.000 ±0.000
voice 1.000 ±0.00022
verdict_table 1.000 ±0.000
verdict_vocabulary 1.000 ±0.000
recommendation 1.000 ±0.00023
verdicts_correct 0.800 ±0.062
part1_is_decision_grade 0.900 ±0.041
evidence_and_location 0.967 ±0.033
scenario_trap 0.800 ±0.13324
0.959 2026-08-26 …_nnzWgW7JK6KsFFqorqCSck.eval
18 revision_planning_summary 5 3
<summary>6 deterministic all 1.000 · 4 judged 0.800–1.000</summary>verbatim 1.000 ±0.000
quoted 1.000 ±0.000
id_scheme 1.000 ±0.000
self_contained 1.000 ±0.000
two_part_layout 1.000 ±0.000
voice 1.000 ±0.00022
locations_by_content 0.967 ±0.033
part1_triage 0.800 ±0.062
planning_notes 1.000 ±0.000
scenario_trap 0.800 ±0.09725
0.957 2026-08-26 …_iUkias5YSUsBuY2EoRoKqm.eval
  Mean across all evals       0.943    

What the model switch changed

Against the last gpt-5.4 / gpt-5.5 figures. Six evals are unchanged, four improved and eight fell, for a suite mean 0.009 lower. Only one movement is large enough to matter on its own.

Eval gpt-5.4 / gpt-5.5 gpt-5.6-terra Δ
abbreviation_checker 0.993 0.995 +0.002
about_this_ger 0.983 0.969 -0.014
advocacy_tone_v2 0.994 1.000 +0.006
claim_reference_validation_v2 1.000 1.000
document_structure 0.867 0.867
figures_tables_check 0.864 0.780 -0.084
inference_validation_v2 1.000 1.000
literature_review_v2 0.944 0.979 +0.035
live_reports_v2 1.000 1.000
methodological_alignment 0.958 1.000 +0.042
recommendation_check 0.955 0.953 -0.002
reference_downloader 0.903 0.849 -0.054
reference_text_extractor 0.922 0.878 -0.044
reference_validation_v2 0.848 0.819 -0.029
results_extraction 1.000 1.000
reviewer_2 1.000 1.000
reviewer_coverage_report 0.972 0.959 -0.013
revision_planning_summary 0.967 0.957 -0.010
Mean across all evals 0.954 0.945 -0.009

Four things worth knowing before reading that table:

Scorer reference

  1. abbreviation_checker · deterministic match of the extracted abbreviations list against the target (inline definition, line span, section definition, ignored flag). 

  2. abbreviation_checker · deterministic check that the “Abbreviations section found” boolean matches the target. 

  3. model_graded_check — an LLM grader compares the workflow’s full output against the target answer, with partial credit. Some evals grade against a target_answer in sample metadata; the mechanism is otherwise identical across evals.  2 3 4 5 6 7 8 9 10 11 12 13

  4. about_this_ger · deterministic match of the flagged preface / “About This” issue titles against the target. 

  5. about_this_ger · deterministic match of the flagged author-biography issue titles against the target. 

  6. advocacy_tone_v2 · deterministic match of the count of flagged issue titles against the target. 

  7. claim_reference_validation_v2 · checks each citation’s support label aligns with the target (supported / partially / unsupported / unverifiable). 

  8. claim_reference_validation_v2 · checks the number of citations found matches the target. 

  9. document_structure and figures_tables_check · deterministic match of the detected issue titles against the target.  2

  10. inference_validation_v2 · deterministic match of the count of reported invalid inferences against the target. An informational (none) issue fails the sample outright rather than counting towards the total: this assessment reports invalid inferences only, so sound reasoning is reported as nothing at all. 

  11. literature_review_v2 and live_reports_v2 · structural checks averaged into a [0,1] score (result present with non-empty report, issue count within the expected band, sane line ranges, citation-like detail when recommendations are expected). Exact sources aren’t asserted because web search is non-deterministic.  2

  12. methodological_alignment · shape check that the analysis ran and populated a reproducibility class plus the field-alignment section (comparison prose is free-form, so exact wording isn’t scored). 

  13. recommendation_check · deterministic match of the counts of recommendations by severity against the target. 

  14. reference_downloader · deterministic match of the final download conclusion against the target. 

  15. reference_text_extractor · deterministic match of the extracted bibliographic references against the target. 

  16. reference_validation_v2 · deterministic match of the final validation result label against the target. 

  17. results_extraction · read this as a dev-split figure, not an unbiased estimate. These ten samples were built and then tuned against this workflow across roughly fifteen runs on 31 Aug: three documents, the skill’s class-selection guidance, the judged criteria and the workflow’s reasoning effort were all edited while these same samples were being scored. Eight of the ten documents are synthetic. Two independent audits found ground-truth defects that had been scoring the workflow down for being right — an engineering estimate that did not follow from its own coefficients, a reservoir yield its own inflow record could not supply, a matching variable the document never stated, and paid-access sources cited in a fixture whose expected class requires openly available ones. Each is fixed, and tests/unit/evals/test_reproducibility_dataset_arithmetic.py plus evals_inspectai/e2e/results_extraction/verify_reservoir_yields.py now re-derive the numbers behind every fully-reproducible expectation. But a dataset repaired wherever the system under test objected to it is a dataset biased towards that system: the gain from 0.899 to 0.977 across this day is mostly eval error being removed, not the workflow improving, and nothing here has been measured on samples the workflow has never influenced. Treat a single reading here as approximate. This configuration has been run four times and class_accuracy came out 0.879, 0.922, 0.942 and 0.992; the recorded run is the 0.942 one, which sits above the median of the four, so read the row as an indication rather than a ceiling. On ten samples a five-point swing between runs is ordinary, and a regression has to clear that before it means anything. Held-out samples are the outstanding work — blind-authored documents, one adversarial case that falsely claims reproducibility, and one full-length real report. Until those exist, treat this eval as a tripwire for large regressions rather than a quality score. 

  18. results_extraction · seven deterministic checks on the shape of the delivery: a substantive markdown report, an inventory table in it (found by the reproducibility column its header names, not by being the widest table in the report) with at least a row per reported result, at least as many results as the dataset declares, a recognised reproducibility label in every issue title, the severity split the skill mandates (anything reproducible is informational none; only not-reproducible carries a real severity), line ranges that land inside the document, and no result reported twice. None of these say whether a classification is right

  19. results_extraction · four deterministic checks against the dataset’s own ground-truth inventory, which names every result each document presents together with the class it should be given and how much the document rests on it: whether every expected result was found, the fraction given the right reproducibility class, whether anything was reported that the document does not present as a result, and whether the severities of the non-reproducible results are ordered by importance. Which of low/medium/high a result earns is the agent’s judgement, so only the ordering is asserted, not the level. 

  20. results_extraction · two judged criteria, each graded three times with the median taken, because single grader calls disagreed with themselves across repeats of an unchanged output. classification_grounded asks only whether each rationale names the ingredients present or missing and how a reader would obtain them; it is given the four class definitions and the list of absences the skill says are not deficiencies, since without them the grader marked down rationales for correctly declining to treat an unrendered figure or a fixed-seed simulation as a gap. sample_expectations is the dataset’s own per-sample rubric. Both use a requirement-shaped prompt rather than Inspect’s model-graded-fact template, which asks whether the submission contains the expert answer and misgrades a criterion that states a property the output must have. 

  21. reviewer_2 · shape check that both the peer review and the rebuttal were produced and are substantive (the model-graded scorer judges whether they cover strengths, weaknesses, and next steps). 

  22. reviewer_coverage_report and revision_planning_summary · the rules the review-assistant skill states outright, checked deterministically against the report’s HTML: every reviewer memo reproduced verbatim, that reproduced text sitting inside a marked quote, a valid per-reviewer point-ID scheme numbered from 1 with no gaps, a self-contained document (no external stylesheets, fonts, scripts or images, and no <script>), a visible two-part split with a short first part, and none of the generic-assistant tells the voice-and-tone skill bans, counted only outside quotes so the reviewer’s own punctuation is not held against it.  2

  23. reviewer_coverage_report · the arithmetic of the summary table, checked deterministically: all four verdict categories present including the ones that scored zero, every point accounted for exactly once and at one granularity with the stated counts matching the IDs listed, the four-point scale actually used in Part 2 rather than only declared in the table header, and Part 1 stating the sign-off decision outright. Whether an individual verdict is right is judged separately. 

  24. reviewer_coverage_report · four judged criteria, each graded in its own call so one weak area cannot colour the rest: each point’s verdict is correct against both drafts, Part 1 is decision-grade for a QAM, each verdict cites evidence and a location, and a per-scenario trap criterion for the specific failure that scenario is built to provoke. 

  25. revision_planning_summary · four judged criteria, graded one call each: reviewer points located by content rather than by numbers the revision will move, Part 1 triaging substantial asks apart from quick fixes, a planning note under each quoted point carrying scope and location and a suggestion, and a per-scenario trap criterion.