← All research projects

Computational social science · 2026

From Fluency to Accountability

A claim-lifecycle analysis of how AI-like peer-review concerns move from review text into author response, reviewer follow-up, and the final institutional rationale.

Four-stage claim lifecycle from review claim to meta-review uptake
The unit of analysis is an individual review claim and its downstream trace—not a binary label assigned to an entire reviewer or review.
53,463public official reviews
416,862extracted review claims
13,738accept/reject papers
3downstream lifecycle outcomes

Beyond AI-text detection

Most discussions of AI-assisted reviewing ask whether a review sounds generic, cites evidence, or changes a score. This project instead treats peer review as an accountability process: a criticism matters institutionally when participants can identify it, answer it, contest it, and carry it into a decision rationale.

The project does not label reviews as AI-generated. It uses a continuous, unsupervised AI-likeness style proxy and studies observational associations.

Claim-level study design

The public ICLR 2026 process is represented as a sequence. Each official review is decomposed into claims; those claims are aligned with author-response units, reviewer follow-up, and meta-review text. The text layer measures specificity, factual support, within-paper redundancy, uniqueness, and marginal decision contribution.

The process layer estimates within-paper fixed-effects linear probability models. Holding the paper fixed, it asks whether otherwise comparable claims with different AI-likeness scores remain visible later in deliberation.

Where the clearer difference appears

The evidence does not show a broad collapse in surface quality, grounding, originality, or marginal decision relevance for more AI-like reviews. The sharper pattern appears downstream.

Across the full 0–1 AI-likeness range, the within-paper estimates associate higher AI-likeness with a 6.63 percentage-point decrease in author-response coverage, a 0.56 point increase in reviewer follow-up, and a 1.67 point decrease in meta-review uptake, conditional on the included controls.

Case studies suggest that focus may matter more than fluency: long, polished checklists can disperse attention, while a smaller set of consequential and answerable concerns is easier to track through the deliberation chain.

Interpretive boundaries

  • The AI-likeness score is a stylistic proxy, not proof that a reviewer used AI.
  • Fixed effects remove stable paper-level differences, but the estimates remain observational rather than causal.
  • Lexical alignment can miss implicit, merged, or heavily paraphrased responses and uptake.
  • ICLR 2026 was an unusual conference year with a security incident and disrupted discussion; cross-year and cross-venue replication is necessary.
  • Reviewer follow-up and explicit uptake have low base rates, so small absolute changes can be estimated precisely.