Current Inspect AI eval numbers across every eval in evals_inspectai/e2e/.
gpt-5.6-terra on every agent[!NOTE] All evals are end-to-end: they trigger the real workflow through the API, so the backend must be running (
uv run dev.py). Scorer numbers are accuracy (mean over samples,0.00–1.00).Run an eval with:
uv run inspect eval evals_inspectai/e2e/<eval>/<eval>_e2e.py --epochs=<n>The raw Inspect
.evallog for each recorded run is copied underdocs/evals/. The Log column opens it in the hosted log viewer, which is rebuilt fromdocs/evals/on every push tomain. To browse them locally instead, runuv run inspect view --log-dir docs/evals.
Each eval defines its own scorers. The Scorers column gives how many are
deterministic and how many are model-graded, with the spread of their accuracies.
Per-scorer numbers are in the linked log, and what each scorer checks is in its
docstring under evals_inspectai/. Overall avg is the unweighted mean of the
eval’s per-scorer accuracies: a rough headline, since scorers measure different
things. Model-graded scores use openai/gpt-5.4 as the grader, except the two
review-assistant suites (18 and 19), which pin openai/gpt-5.6-terra.
A dev-split figure comes from a dataset written and labelled while the workflow was being built, so it shows the workflow does what its dataset asks, not how it performs on documents it has never seen.
| # | Eval | Samples | Epochs | Scorers | Overall avg | Date | Log |
|---|---|---|---|---|---|---|---|
| 1 | abbreviation_checker1 |
36 | 3 | 22 deterministic all 1.000 | 1.000 | 2026-09-30 | …_S5R5oXb3Fq85dm2bdn4KcL.eval |
| 2 | abbreviation_checker_report2 |
10 | 3 | 20 deterministic 0.922–1.000 | 0.984 | 2026-09-30 | …_Uthi6G3k5Hgc8tA7K3pXLV.eval |
| 3 | about_this_ger3 |
24 | 3 | 14 deterministic 0.982–1.000 · 2 judged 0.636–0.808 | 0.962 | 2026-09-25 | …_DsUmAvj3AQGumAhaCdivTd.eval |
| 4 | advocacy_tone_v24 |
26 | 3 | 21 deterministic 0.995–1.000 · 2 judged 0.530–1.000 | 0.979 | 2026-09-25 | …_TjsNM97tmR26QHfuteQemZ.eval |
| 5 | claim_reference_validation_v25 |
22 | 3 | 16 deterministic 0.889–1.000 · 2 judged 0.984–1.000 | 0.989 | 2026-09-29 | …_LyGi33T94LaBu9MQmppN2C.eval |
| 6 | document_structure6 |
19 | 3 | 7 deterministic 0.900–1.000 · 2 judged 0.884–1.000 | 0.976 | 2026-09-25 | …_ZdxFzLrfbrZ3SgVBvKkRo4.eval |
| 7 | figures_tables_check7 |
33 | 3 | 20 deterministic 0.955–1.000 · 2 judged 0.901–0.989 | 0.989 | 2026-09-25 | …_JBJwF4HJR7XeSwbD6zXCgc.eval |
| 8 | inference_validation_v28 |
32 | 3 | 18 deterministic 0.667–1.000 · 2 judged 0.986–1.000 | 0.980 | 2026-09-28 | …_TpWJeAd34PgKfxrTWu57fM.eval |
| 9 | literature_review_v29 |
12 | 3 | 13 deterministic 0.989–1.000 · 2 judged 0.953–0.964 | 0.994 | 2026-09-29 | …_Uothj6wWS7DWx6SepoJasz.eval |
| 10 | live_reports_v210 |
12 | 3 | 13 deterministic 0.983–1.000 · 2 judged all 1.000 | 0.998 | 2026-09-29 | …_Bcc3cpuMzT2raQb6bJsbxg.eval |
| 11 | methodological_alignment11 |
10 | 3 | 7 deterministic 0.983–1.000 · 3 judged 0.939–0.994 | 0.987 | 2026-09-29 | …_afsNxhi3foWMbkqzopjgpQ.eval |
| 12 | recommendation_check12 |
15 | 3 | 18 deterministic 0.927–1.000 · 3 judged all 1.000 | 0.997 | 2026-09-25 | …_N9crApAnp4ogUdpnNppY3y.eval |
| 13 | reference_downloader13 |
35 | 3 | 7 deterministic 0.940–1.000 | 0.988 | 2026-09-30 | …_TsCQX8LiVwYLLqgvtgcARb.eval |
| 14 | reference_text_extractor14 |
16 | 3 | 7 deterministic 0.987–1.000 | 0.997 | 2026-09-28 | …_82Tfy9oyDuKFjyYuWoTE5L.eval |
| 15 | reference_validation_v215 |
80 | 3 | 14 deterministic 0.923–1.000 · 1 judged 0.879 | 0.963 | 2026-09-29 | …_oCRYhcopfNiN35ypKPmSFD.eval |
| 16 | results_extraction16 |
11 | 3 | 12 deterministic 0.968–1.000 · 2 judged 0.879–0.970 | 0.985 | 2026-09-03 | …_iMzDaGm4SpnSJPpWfTAJM3.eval |
| 17 | reviewer_217 |
12 | 3 | 4 deterministic all 1.000 · 6 judged 0.920–1.000 | 0.986 | 2026-09-28 | …_7i4fQLe3xsBqNzQ4X8AUga.eval |
| 18 | reviewer_coverage_report18 |
7 | 3 | 11 deterministic all 1.000 · 4 judged 0.857–0.952 | 0.975 | 2026-10-01 | …_KwhsEUTvuHYYkiYTtaFTtm.eval |
| 19 | revision_planning_summary19 |
5 | 3 | 6 deterministic all 1.000 · 4 judged 0.800–1.000 | 0.957 | 2026-08-26 | …_iUkias5YSUsBuY2EoRoKqm.eval |
| 20 | active_voice20 |
20 | 3 | 24 deterministic 0.794–1.000 · 3 judged 0.863–0.889 | 0.952 | 2026-09-24 | …_g3nQCy7RwyF7JJv8GRyH3y.eval |
| 21 | concision_precision21 |
15 | 3 | 26 deterministic 0.333–1.000 · 2 judged 0.998–1.000 | 0.953 | 2026-09-24 | …_UyVYbdkwsAjdmfigyyk2nn.eval |
| 22 | writing_consistency22 |
15 | 3 | 25 deterministic 0.963–1.000 · 1 judged all 1.000 | 0.998 | 2026-09-24 | …_J7agYBJdaYwYcCKrQuEfvM.eval |
| 23 | narrative_synthesis23 |
15 | 3 | 22 deterministic 0.989–1.000 · 2 judged 0.992–1.000 | 0.999 | 2026-09-24 | …_biQJ9p29PzEsZy3rPHLqrb.eval |
| 24 | headers_skimmability24 |
16 | 3 | 19 deterministic 0.991–1.000 · 2 judged 0.912–0.985 | 0.995 | 2026-09-24 | …_Tmx6QQmfCh8HGgmah9FXkq.eval |
| 25 | audience_fit25 |
16 | 3 | 21 deterministic all 1.000 · 5 judged all 1.000 | 1.000 | 2026-09-24 | …_GqprsEpnAVJe5gSWCWcHsU.eval |
| Mean across all evals | 0.983 |
Checks that abbreviations are defined at first use, listed in an Abbreviations section, defined the same way there and inline, and not given two meanings. Scored at two layers. The occurrence catalogue the agent extracts is matched occurrence by occurrence (by location, then by occurrence number), with recall, precision and per-field accuracy for the inline definition, lines and section definition. Exempt occurrences (headings, references, footnotes, lists of figures and tables, exempt classes, the always-excluded names) are not catalogued at all, so recording one counts against precision. The issues a user sees, which the workflow builds from that catalogue and stores, are fetched from the project and scored against an issue inventory, with decoys for exemptions and for later or consistent uses. The 36 samples include a document with no abbreviations, one-line samples for each exempt class, each rule failing on its own, an abbreviation used before it is defined, and a 157-line report that spans two extraction chunks; a full-length report is scored separately2. Read the new samples as a dev-split figure. Every metric is 1.000 in every epoch, so the set no longer discriminates; the full-length report carries the harder cases. Two gaps are not scored here: the no-section issue carries no line (the skill says line 1, and it is matched on its title), and definitions are compared by a partial match looser than the skill’s “trivial differences”, which only the full-length report tests. ↩
The same workflow, scorers and rules as abbreviation_checker, on a published 817-line research report about dual-use biology benchmarks that the workflow catalogues in many chunks, plus nine copies of it with first inline definitions removed (one for each of eight abbreviations, and one with all eight). Each copy differs from the report only on its edited lines, and adds a “not defined at first use” issue wherever a first use becomes bare. It is its own task because every sample is long. The catalogue (138 occurrences) was labelled against the extraction skill: benchmark acronyms are abbreviations like any other, so the seven first used in the appendix benchmark table (MMLU, GPQA and others) each fail “not defined at first use” and “missing from Abbreviations section” there, while model names are exempt and the list of figures and tables is not recorded. The judgement calls are listed at the top of dataset_report.yaml. Read this as a dev-split figure. Issue recall (0.945) is capped by one issue the workflow cannot raise yet: the report defines IRT as “Item Response Theory” and later as “Item response time”, two meanings under the skill that the partial-match comparison accepts as one, so it is missed in every epoch; the other misses are single first uses (ML in three runs of thirty, DNA in two, LitQA in one). Issue precision (0.924) is set by the benchmark table’s sub-labels, which the agent sometimes reads as abbreviations to define: “Supp” in “Lab-Bench Supp” in eleven runs, “Chem” in “MMLU Chem” and “Lab-Bench” itself in seven each. The labels treat them as names and shortened words, not acronyms. Occurrence recall (0.922) is repeat uses of AI on busy lines, dropped without affecting any issue. ↩ ↩2
Checks a report’s preface (context, objectives, audience, prior work, contribution, scope) and author biographies (three sentences, position and affiliation, research focus, highest degree) across two validators scored together. Scored against an issue inventory: a failed preface element or a missing section is matched on its title, and an author issue is anchored on the bio; where one author fails two rules the second is optional, since the skill allows one combined issue. The 24 samples cover single and multiple failures, one to four authors, alternative headings, sentence counting around abbreviations, professional degrees and short non-bio lines. Own checks confirm titles come from the skills’ set and name a real author, every issue is medium, and a “section not found” issue stands alone. Read this as a dev-split figure. The detection misses are the objectives element, which the agent occasionally reads into a report’s motivation. The judged scores are the lowest: preface actions rarely draw on the report (“add a sentence naming the intended audience”), and bio actions do not say which sentence to cut. Whether “Dr.” alone satisfies the degree rule is not stated by the skill; one sample expects it not to. ↩
Checks for non-neutral language: certainty without evidence, advocacy, and subjective tone. Scored against an issue inventory: one expected issue per flagged sentence with its check’s title and severity, and decoys for every carve-out (quoted regulations, terms of art, technical requirements, methods language, existing obligations, factual policy mentions, hedged claims, supported recommendations, descriptive uses, and matches inside skipped sections such as acknowledgements, author notes, references and appendices). Where a flagged sentence also carries an evaluative phrase, a second Subjective Tone issue is accepted but not required. Own checks, over every reported issue: titles from the skill’s three, severity by title, a single-sentence range, nothing in a skipped section. Read this as a dev-split figure. Detection is near perfect; the lowest score is action_concrete (0.53): suggested actions name the word but rarely give neutral wording (“replace ‘clearly’ with qualified language”). action_faithful sits at 1.000 largely because there is little proposed wording to be unfaithful to. ↩
Checks whether each in-text citation is supported by the source it cites. The workflow reports every citation with an evidence level, so every citation is an expected issue in an inventory, anchored on the cited claim and titled with the level a correct run assigns; title_correct is level accuracy and label_<level> breaks it down by level. Decoys cover bibliography and footnote entries, a footnote to commentary and uncited claims. The 22 samples cover every level and the boundaries the skill spells out (faithful inference, rounding, an omitted interval; scope overreach and mixed evidence; silence, a contradicted value or direction, the right number on a different measure; a missing source), bracketed, caret and superscript markers, one source cited for claims it does and does not back, a finding attributed to the wrong paper, and a six-citation memo. Own checks confirm evidence quotes are verbatim in a source and cited text verbatim in the document; judged criteria compare the rationale with the labeller’s account of the source, and check the action for an unsupported citation against that account. Citations are paired with the run’s records on the cited claim alone, so a swapped level is scored as wrong. Read this as a dev-split figure. Detection is perfect; the misses are a superscript footnote with no reference list, reported unverifiable in every epoch because the source is never linked to it (so its rationale is graded wrong too), and in one epoch a benchmark figure the source calls “consistent with prior reports” labelled partially supported, where the claim calls it a consensus. ↩
Checks that a document has its required top-level sections, and an appendix when the body refers to one. Scored against an issue inventory, each expected issue matched on its “Missing Section” title since the section has no line to point at, with severity high. The 19 samples cover every section missing on its own, the alternative headings the skill lists, deep headings and bold-labelled blocks, traps where a section’s name appears without the section, and every branch of the conditional appendix. Two own checks confirm titles come from the skill’s set and that an appendix issue points to the sentence citing the appendix. Read the synthetic samples as a dev-split figure. On the two published reports, whose findings sit under a Summary and under finding-titled chapters rather than a Results heading, the agent sometimes still reports a missing Results section (every epoch on one report in this run), which sets clean_document_untouched. The lowest judged score is action_specific: suggested actions list what any such section contains (data source, participants, measures) rather than what this report’s section should say. ↩
Checks that every figure and table is captioned, consistently numbered, and referenced in the text, and that every reference resolves. Scored against an issue inventory: an issue about one element is anchored where it sits and matched on the rule its title names, a numbering issue on “Inconsistent Numbering” alone, with severity where the issues skill fixes it (a missing element high; unreferenced and numbering issues medium). Decoys cover elements a correct run leaves alone (abbreviation tables, appendix-, chapter- and supplementary-prefixed numbering, “Fig.”, parenthetical and plural references, a conversion artefact, a captioned but unnumbered figure’s title). Own checks confirm titles use the skill’s prefixes and name the right element. Four samples embed a chart26. Read this as a dev-split figure. The one detection miss is figures that appear out of numerical order, which the skill’s numbering rule does not list. Suggested actions sometimes renumber without naming the new number or the citing sentence, which sets action_concrete. ↩
Checks reasoning for conclusions that do not follow from their premises. Scored against an issue inventory: each invalid inference is anchored on the sentence that draws it and carries the labeller’s account of the flaw, which a judged criterion compares with the reported analysis; decoys are sound inferences by the adjudicator’s reasons for rejecting them (bounded, adequately supported, valid deduction, a probability sample, hedged, a stated limitation, a cited finding). The 32 samples cover sixteen named fallacies, sound documents of the same shapes, a sentence with two flaws that the skill reports once, section-length reports mixing sound and flawed inferences, and a pair of documents with identical text whose chart contradicts or supports the premise26. Own checks confirm no informational issue, the flaw named in the title and a verbatim key sentence. Read this as a dev-split figure. Recall is perfect; precision misses are a hedged causal attribution flagged in two epochs of three and a recommendation that restates the flaw of the sentence before it as a second issue. ↩
Finds sources a document should cite or discuss but does not, using web search. Scored against an issue inventory: each claim that needs a source (a well-studied claim left uncited, a one-sided claim with a known conflicting literature, a claim a reference already in the bibliography supports but is never cited for) is anchored on the claim with the labeller’s account of the literature it should engage with; claims a run may fairly give context to are optional, and decoys are sentences no source can bear on. One issue is reported per source, so every issue covering a claim counts toward precision and is judged. Sources are not asserted, since search results change: shared source checks require a year and a DOI or URL, a year no later than the publication date on the seven dated records, and the source named in the report. The 12 samples include two documents with no claim. Read this as a dev-split figure. The one detection miss is a button-label claim skipped in one epoch of three. The judged scores are the lowest: a source on the broader topic rather than the claim (a design standard for a claim about one label), and an action that adds a citation to an overstated claim without qualifying it. Decoys are few by design: the skill asks for contextual sources too, and the pilot runs showed a local plan or result is fair ground for one, so those sentences are optional. ↩
Finds literature published after a document that updates or challenges its claims, using web search. Scored against an issue inventory: each record dates the project, and each claim later evidence overturned or qualified (a superseded standard, a reversed guideline, a broken record, a revised estimate, an association later read as non-causal) is anchored on the claim with the labeller’s account of what changed; claims that still hold are optional, and decoys are historical and local facts. Sources are not asserted: shared source checks require a full citation, a year no earlier than the publication date, the source listed in the report and not already in the document’s references. The 12 samples include two documents with nothing to update. Read this as a dev-split figure. Recall is perfect; the misses are an undated web page given as a source once, and an issue recommending the current HTTP/1.1 specification for the reference list rather than for a claim. At 0.998 the set is close to saturated even with its subtler records (the alcohol J-curve, the bacteria-to-human-cell ratio); claims with mixed newer evidence, and documents so recent that little newer work exists, are the outstanding work. ↩
Compares a paper’s methodology with standard practice in its field, using web search for the field baseline. Each of the 10 samples is the methods of a short paper in a different field with two to four planted risks, written the way a paper states its choice (neutrally, or by omission) and anchored on that sentence with the labeller’s account of what the field does instead, and one to three sound choices, some of which look like gaps (no margin of error for an opt-in sample, a sample restricted to the establishments a mandate covers). The skill reports every gap it finds, so there is no precision against the plants: a grader reads every issue on a plant’s line and keeps the best grade (gap_identified), grades that issue’s action, and grades each sound choice against every issue reported. Deterministic checks confirm the report’s five sections, web links, no informational issue, a valid line range on every issue, an issue on each plant’s line and a fitting severity. Read this as a dev-split figure. The misses are real: a random utterance-level split of 3,100 users’ data not flagged as leakage in one epoch, missing baselines in another (the run reported out-of-scope evaluation on that line instead), and a single coder named only as a lack of detail. A first version that named each flaw in the text (“so utterances from the same user can appear in both sets”) scored 1.000 on every metric. ↩
Grades each recommendation as supported, partially supported or unsupported by the document’s own findings, and flags recommendations that are not actionable, name no one to act, or number more than three. Read the actionability, audience and count kinds as a dev-split figure: several records were written and labelled against this workflow. severity_correct misses are one-step disagreements on the same few borderline recommendations (near-miss reporting, permanent briefings, the flat chart), a different one or two per run, so read anything between 0.93 and 1.0 there as variance. Two samples embed a chart26. ↩
Finds and downloads the full text of a cited reference. Each of the 35 references accepts the conclusions a correct run may reach (found, found but not accessible, not found), since which copy counts or whether a site lets an automated download through is often open. A source_found is checked against the file the app kept, read back through the app’s file listing: it is a supporting document, its text carries phrases of the work’s title (distinctive: no other record’s file has them), and it carries phrases from the work’s last part, so a preview of the first pages fails. Nothing is kept on any other conclusion, and an inaccessible source says why. The solver checks the text and keeps only the verdicts, so the log stores no downloaded text of its own. Four fabricated references, one behind a plausible URL, must come back not found without a lookalike kept. Read this as a dev-split figure. The misses are real: a Scribd preview of a journal article (its first five pages) kept as found in one epoch, and in two epochs a 172- or 353-character stub of the DMDC spreadsheet kept and called verified although its cells were never read. The file listing does not show transient candidate files, so a download the cleanup leaves behind is not visible to the eval. ↩
Extracts every bibliographic entry from a document’s reference sections. Extracted and expected entries are matched one to one after normalising whitespace, entities and quotes, with recall, precision, F1, a clean-document check, exact text over matched entries, whether each extracted entry appears in the document (catching invented or merged entries), and whether its line range brackets its text. Every expected entry is verified against its source document when the dataset loads. The 16 samples cover numbered, bulleted, hanging-indent and bracketed lists, entries wrapped onto URL lines, repeated-author placeholders, two reference sections, an appendix after the list, endnotes that are not references, an empty section, and two published reports of 152 and 309 entries. The one miss is the two-section document, read as a single list in one epoch; the skill’s “the reference section” leaves that case open. ↩
Checks a citation’s author, title, publisher, year and identifier against authoritative sources found online. Each of the 80 references is labelled field by field against the reference-validation skill: the final result, each field’s problem type, phrases a correction must carry, and whether the reference is fabricated. Deterministic checks score the result, each field (overall and per field), the fields a correct run flags and the ones it must leave alone, the corrections, and the skill’s contract (the result follows from the fields; an updated reference exactly when an incorrect found work needs one; a URL for a found work); a judged criterion compares the reasoning with the labeller’s account. The skill now states three rules the old labels assumed, each with cases on both sides: an organization that is author and publisher is named once (as RAND’s own reference lists and APA do), a title may omit its subtitle or deck while it still identifies the work, and a cited URL that returns 404 with a title no search finds is a fabricated reference, not a similar work with errors. Read this as a dev-split figure. The remaining misses: an embassy’s September 23 repost of a State Department media note is attributed to the State Department; two dates in the same year (July 10 against July 14, 2010) are reported as a wrong year, against the skill’s year rule; a Perficient URL behind a bot checkpoint and a Medium article behind a redirect are reported wrong; and the fabricated Google post is still matched to a similar 2023 post in two epochs of three. ↩
Extracts a document’s main results and classifies how reproducible each is. Read this as a dev-split figure, not an unbiased estimate: ten of the eleven samples were built and tuned against this workflow, and eight of those ten documents are synthetic. class_accuracy varies between about 0.88 and 0.99 across runs of an unchanged configuration, so a regression has to clear that before it means anything. Held-out samples (blind-authored documents, an adversarial case that falsely claims reproducibility, a full-length real report) are the outstanding work. One sample embeds a chart26. ↩
Produces a four-section peer review and a devil’s-advocate rebuttal. The output is prose, so each sample lists the substantive weaknesses planted in the document and its genuine strengths, and a judge grades each on its own: raised, credited, not attacked. Three more criteria are graded once per sample (what the review says the document says or lacks is accurate, the next steps are concrete, the rebuttal concedes where it is weakest). Deterministic checks confirm both documents, the four sections, five to seven next steps and the header blocks. The 12 samples span the disciplines the persona channels (an incentive design, a historical case study, a claim about AI and security, interviews, a volunteer-panel survey, a theoretical argument its own example undermines), internal inconsistencies, alternative explanations the document reports itself, a truncated chart axis26 and a strong pre-registered trial. Read this as a dev-split figure. The lowest score is strengths_not_attacked: reviews question a strength itself, mostly the teacher-bonus study’s comparison group, which is matched on poverty alone. claims_accurate misses are minor overstatements of what a document says, none that a critique rests on. ↩
Writes the report a quality-assurance manager reads to decide whether a revision answered its reviewers. The deterministic scorers check the rules the review-assistant skill states outright (verbatim memos, point IDs, a self-contained two-part document, the verdict table’s arithmetic) and how the report handles the author’s response memos: two of the seven samples upload them to the revised draft, and there each reply must be reproduced verbatim and marked as the author’s, with an unanswered point’s slot saying so; without them, the header must say none were supplied and no reply may be invented. The response samples plant replies that claim changes the draft does not make (an added threshold, cost estimate and reconciliation paragraph; a recommendation said to be dropped that is still there) and a decline whose only reason is in the reply, which the judged scenario_trap grades; the grader sees the replies alongside both drafts. Every planted claim is caught and every reply quoted in every epoch. The judged scorers are the lowest. scenario_trap (0.857) is set mostly by the commentary scenario, where asks a commentary does not carry are still counted as gaps and the report returns the revision for another pass rather than signing off, in all three epochs; the other trap miss is one small correction (a figure reference) judged partially addressed. part1_is_decision_grade (0.881) is partly the grader counting the verdict table and the header’s note on response memos as extra Part 1 material, though the skill requires both. evidence_and_location misses are locations given by number (“Table 1”, “the second numbered recommendation”) rather than by content. ↩
Writes the summary an author uses to plan a revision from reviewer memos. The same structural rules as the coverage report are checked deterministically; the judged scorers check triage in Part 1 and each scenario’s trap, which are the lowest scores. ↩
Flags passive voice and ambiguous actors, and proposes the active rewrite when the actor is known. Read this as a dev-split figure: the 20 records were labelled against this workflow. Known gaps: a comma stranded after a moved footnote marker in every epoch, a stative “are located in the Midwest” flagged, generic modal passives flagged in two epochs of three, and an ambiguous-actor sentence missed or given a guessed actor. Its two edit criteria are the only graders here calibrated against human labels (14 edit pairs). ↩
Flags wordy constructions, run-ons, throat-clearing, vague references, empty framing and statements of the obvious, with a tighter rewrite or deletion when the fix is fully determined. Read this as a dev-split figure, on the same terms as 20. Two behaviours set the current score. In one epoch of the section-length fixture the agent quoted whole paragraphs as each edit’s original text instead of the sentence it changes, so every decoy in those paragraphs counts as touched; in the Word export that is a whole-paragraph redline for a one-sentence change. And an ambiguous “This” is given a specific referent in two epochs of three instead of an action asking the author, which is the only expectation behind edit_absent_when_not_expected (0.333). ↩
Flags inconsistent terms, spelling, hyphenation, number style, tense and tone across a whole document, and the house-style compound decisionmaking, with word swaps as edits. Read this as a dev-split figure, on the same terms as 20. The misses are one term pair missed once on the section-length fixture and one edit that changes more than three words. ↩
Flags illogical order, data reported without interpretation, unanswered framing questions (why it matters, what is new, what happens next), restated points, and a main body over about 15,000 words. Read this as a dev-split figure, on the same terms as 20. Not in the inventory: length, because the agent estimates word counts rather than counting (on two published reports of 28,680 and 24,685 words it estimated 26,000 and 25,000), and the five-issue volume cap, which no record exceeds. ↩
Flags a missing bottom line, vague headers, headers without a takeaway or that do not match their section, weak bold lead sentences, and an overlong key findings box, with suggested wording as a comment. Read this as a dev-split figure, on the same terms as 20. The lowest score is judged: suggestions that borrow the introduction’s causal wording, and chapter headers that compress a mixed finding into 2 to 8 words. ↩
Settles on the document’s audience, flags a missing, vague or conflicting one, and flags technical language in the main body for a non-technical audience. Read this as a dev-split figure, on the same terms as 20. At 1.000 on every metric in every epoch the set does not discriminate; harder samples (a mixed policy-and-research audience, borderline field terms, a long methods chapter) are the outstanding work. ↩
Figure samples embed a PNG chart, drawn by evals_inspectai/files/figures/generate_fixtures.py and inlined at upload, that carries information the text does not (a logo that is not a figure, a caption inside the image, a placeholder box, values only in a bar chart, a truncated axis). Where the transcript is available, tool_called records whether the agent actually looked at the image. Judged criteria receive the embedded images alongside the text, so a proposed caption or chart reading is graded against the figure itself. ↩ ↩2 ↩3 ↩4 ↩5