Beyond AI-text detection
Most discussions of AI-assisted reviewing ask whether a review sounds generic, cites evidence, or changes a score. This project instead treats peer review as an accountability process: a criticism matters institutionally when participants can identify it, answer it, contest it, and carry it into a decision rationale.
The project does not label reviews as AI-generated. It uses a continuous, unsupervised AI-likeness style proxy and studies observational associations.
Claim-level study design
The public ICLR 2026 process is represented as a sequence. Each official review is decomposed into claims; those claims are aligned with author-response units, reviewer follow-up, and meta-review text. The text layer measures specificity, factual support, within-paper redundancy, uniqueness, and marginal decision contribution.
The process layer estimates within-paper fixed-effects linear probability models. Holding the paper fixed, it asks whether otherwise comparable claims with different AI-likeness scores remain visible later in deliberation.
Where the clearer difference appears
The evidence does not show a broad collapse in surface quality, grounding, originality, or marginal decision relevance for more AI-like reviews. The sharper pattern appears downstream.
Across the full 0–1 AI-likeness range, the within-paper estimates associate higher AI-likeness with a 6.63 percentage-point decrease in author-response coverage, a 0.56 point increase in reviewer follow-up, and a 1.67 point decrease in meta-review uptake, conditional on the included controls.
Case studies suggest that focus may matter more than fluency: long, polished checklists can disperse attention, while a smaller set of consequential and answerable concerns is easier to track through the deliberation chain.
Interpretive boundaries
- The AI-likeness score is a stylistic proxy, not proof that a reviewer used AI.
- Fixed effects remove stable paper-level differences, but the estimates remain observational rather than causal.
- Lexical alignment can miss implicit, merged, or heavily paraphrased responses and uptake.
- ICLR 2026 was an unusual conference year with a security incident and disrupted discussion; cross-year and cross-venue replication is necessary.
- Reviewer follow-up and explicit uptake have low base rates, so small absolute changes can be estimated precisely.