Draft Detective is an open-source assistant for reviewing research documents before they are published. It runs a set of independent checks over a draft (citations, reasoning, methods, editorial requirements and language) and reports what it finds as issues pinned to the lines of the text they concern. The author reads each issue, decides what to change, and marks it resolved.
This project is funded by RAND’s CAST Center (RAND Center on AI, Security, and Technology).
This page covers the project’s approach and design. For setup and usage, see the README and DEVELOPMENT files in the GitHub repository.
A three-minute walkthrough of the app: uploading a draft, choosing assessments, reading the findings in the Document Explorer, and exporting them to Word. The demo report is invented; its three references are real publications.
If the player doesn’t load, watch it on YouTube.
Automated scholarly paper review is a growing research area that applies language technology to parts of peer review: checking claims against evidence, analysing citations, and assessing structure and clarity 1. Recent surveys find that large language models make much of this practical, from generating structured comments to verifying checklists and catching technical errors, while raising concerns about bias, inaccuracy, privacy and disclosure 2.
Draft Detective turns that research into a tool an author or reviewer can use on a real draft. It does not score a paper or recommend acceptance. It points at specific passages, says what is wrong with each and why, and leaves the judgement to a person.
Each check, called an assessment, answers one narrow question about the document. They fall into four groups:
| Group | The kind of question it asks | Examples |
|---|---|---|
| Citation check | Are the references real and correctly cited? | Reference Error Checker |
| Substantive review | Do the claims, reasoning, methods and recommendations hold up? | Claim Reference Validation, Internal Inference Validation, Methodological Alignment, Recommendation Check |
| Editorial and style review | Does the document meet structural and house-style requirements? | Figures & Tables Check, Abbreviation Scan, Headers & Skimmability |
| Language | Is the prose neutral, direct and consistent? | Advocacy & Tone, Active Voice & Clear Actors, Concision & Precision |
Beyond these, a peer-review workspace helps authors respond to reviewers: it turns reviewer memos into a revision plan, drafts a response memo per reviewer, and reports which points a revised draft addresses. A simulated reviewer (Reviewer 2) writes a whole-document critique with a devil’s-advocate rebuttal.
The set of assessments grows regularly, so this page does not list them all. The current list, with what each one measures, is in the app’s Run assessments dialog, in the skills/ folder, and in the eval scores report.
One check, one question. An assessment reads the document for a single kind of problem and nothing else. Narrow checks are easier to write, test and trust than one prompt that reviews everything, and a user can run only the ones that matter for a given draft. Presets such as Standard Review and Editorial Review select a useful group in one click.
The rules are written as skills. The instructions for an assessment live in a plain-language SKILL.md file: what to look for, how to judge it, how severe each kind of problem is, and what a good fix looks like. The app’s agents load these files at run time, and the same files install on their own as a plugin for Claude Code or Codex, so a check behaves the same inside the app and in a chat assistant. A new single-pass check needs no Python code: a SKILL.md with a short block of frontmatter registers it, and regenerating the frontend’s API types makes it available in the app (see adding a check).
Every finding is an issue in the text. However different the checks are, they all report in one shape:
{
"title": "Partially supported: the claim is broader than its source",
"description": "Okafor and Brandt (2022) report faster case handling in three of the five departments they studied. The sentence says compressed schedules improve delivery in every department.",
"severity": "medium",
"start_line": 23,
"end_line": 23,
"suggested_action": "Narrow the claim to what the source found.",
"edits": []
}
Severity is high, medium, low, or none for a check that passed and is worth confirming. The line numbers point into a normalised copy of the document that every check shares, which is what lets the web app, the Word export and the MCP server show any check’s findings without knowing anything about the check that produced them.
Critique, don’t write. A check may propose an exact text replacement (an edit) only when the finding fully determines the fix, as with a passive sentence or a mislabelled figure. It never invents a citation, a number or a paragraph. When the right fix needs the author’s knowledge, the issue says what is needed and stops there.
Evidence before verdicts. Claims are checked against the full text of the sources they cite, not against the model’s memory, and the author confirms which file belongs to which reference before those checks run. References are checked against what can be found on the web. Any check that searches the web asks for consent first, because parts of the document are sent as search queries.
Measured end to end. Most assessments have an evaluation suite that runs it through the real API on labelled documents (see Evaluation).
The screens below are synthetic examples: the document, its authors and its findings are invented to show the interface. They are rendered from HTML mockups of the app in docs/mockups/.
The Document Explorer shows the draft with a line gutter. A coloured rule marks every paragraph with an issue, and the issues sit in the margin beside the text they concern. Here Claim Reference Validation has found that a sentence claims more than its cited source supports, while other checks have flagged a correlation presented as a cause, a passive sentence, and advocacy language.
When a fix is purely mechanical, the issue carries a proposed edit shown as a word-level diff. Proposed edits can be exported to Word as tracked changes, which the author accepts or rejects there.
The Reference Error Checker searches for each reference and compares its author, title, publisher, year and identifier with what it finds. Here it confirms one reference, catches a real paper cited under the wrong journal, and fails to find a reference that does not exist.
Users choose which assessments to run, individually or through a preset. Each one is labelled if it searches the web, needs the full text of references, or proposes edits, and web-searching checks need explicit consent before they run.
Most assessments have an end-to-end evaluation suite under evals_inspectai/e2e/, built on Inspect AI. Every sample triggers the real workflow through the API, so a run exercises the same pipeline a user would. Each suite pairs a labelled dataset with two kinds of scorer:
Current scores for every suite are in the eval scores report. The raw Inspect logs, with every sample, transcript and score, can be browsed in the hosted log viewer.
text-embedding-3-large, stored in PostgreSQL with pgvector.Lin, J., Song, J., Zhou, Z., Chen, Y., & Shi, X. (2023). Automated Scholarly Paper Review: Concepts, Technologies, and Challenges. arXiv preprint arXiv:2111.07533. https://arxiv.org/pdf/2111.07533 ↩
Zhuang, Z., Chen, J., Xu, H., Jiang, Y., & Lin, J. (2025). Large language models for automated scholarly paper review: A survey. arXiv preprint arXiv:2501.10326. https://arxiv.org/html/2501.10326v1 ↩