← PaperForge · Blog

Reference integrity and experiment evidence that survive review

dateModified: 2026-09-27

Part 3 · Experiment design and reporting standard

3.1 Baselines: pick them to answer a question, not to look busy

Each baseline exists to rule out an alternative explanation:

|---|---|---|

Baseline typeQuestion it answersExample
Current strong / SOTAAre we better than what people actually use?Best published model with the same data budget
Canonical / classicalIs the improvement real, or an artefact of comparing only against weak models?Logistic regression, BM25, ResNet-50
Same-parameter / same-computeIs it your idea, or your compute budget?Stronger baseline matched on FLOPs or parameters
Ablated version of your own methodDoes each component carry its weight?Remove module A, then B
Upper bound / oracleHow much headroom exists?Gold labels at inference

Rules:

3.2 Ablation design

3.3 Statistical reporting (the "is it noise?" defence)

Minimum standard for a top venue:

3.4 Tables and figures

3.5 The three highest-value "extras"

---

Part 4 · Academic English for non-native writers

Academic English is a register, not a talent: restricted vocabulary, predictable sentence frames, and heavy use of nominalisation. You can learn the register without speaking like a native, and reviewers reward clarity far more than flourish.

4.1 Four tense zones (get these right and 30% of your grammar errors vanish)

|---|---|---|

ZoneTenseExample
Established facts, model architecture, your definitionsPresent simple"The decoder *attends* over the encoder states."
Your experiments and their outcomesPast simple"We *trained* the model for 40 epochs; it *converged* in 12."
Literature that other people didPresent perfect (or past for a specific result)"Several studies *have addressed* this; Chen et al. *reported* a 3% gain."
Statements about this manuscriptPresent simple"Section 4 *describes* ..." (never "will describe")

Tense discipline is a credibility signal: present tense for method and past tense for results is what a native-reading reviewer expects, and drifting between them reads as careless.

4.2 Sentence frames that always work

Introducing a problem

Positioning against prior work

Describing your method

Reporting results

Honest hedging (mandatory)

4.3 Transition words in review-proof writing

|---|---|

RelationUse
AdditionFurthermore / In addition / Moreover (use sparingly — one per paragraph maximum)
ContrastHowever / In contrast / Nevertheless / Yet
Cause–effectTherefore / Consequently / Hence
ConcessionAlthough / While / Even though
IllustrationFor instance / Specifically / To illustrate
Consequence for designThis suggests / This motivates / Accordingly
SummaryOverall / Taken together / In summary

A single connection word per sentence is plenty. Two ("However, moreover, therefore...") is a stylistic tell of unedited prose.

4.4 Six errors that mark a text as unproofed

4.5 Cutting 15% of your word count without losing meaning

Replace these with one word each:

|---|---|

WordyCrisp
due to the fact thatbecause
in order toto
has the ability tocan
a large number ofmany
in the event thatif
the filler construction that announces importance(delete)
we can see that Table 2 showsTable 2 shows
is indicative of the fact thatindicates

Then read each paragraph asking: what is new information here? Delete the sentence that is not new. Most first drafts contain 10–15% sentences that only restate the previous paragraph.

4.6 How to use assistance honestly

---

Part 5 · Venue selection and submission strategy

5.1 Comparative picture (verified 2026 figures)

|---|---|---|---|---|---|

VenueDomainModelReview timelineReviewers valueRecent stats
NeurIPSMLSingle submission deadline, rebuttal phase, then AC discussion~4 monthsTechnical soundness, novelty, beer-view clarity of the claim2025: 21,575 valid → 5,290 accepted, **24.52%** [ref:neurips2025]
ICLRRepresentation learning / DLOpen peer review on OpenReview, public rebuttal~4 monthsNovelty × depth × presentation × reproducibility; rigorous open discussion2026: 19,525 valid, 779 desk rejects, **27.4%** accept [ref:iclr2026]
ICMLMLMultiple submission windows; rebuttal~4 monthsTheory and empirical rigour in balancecheck official statistics page (verify)
ACL family (ACL / EMNLP / NAACL / EACL / AACL, COLING)NLP / CL**ACL Rolling Review (ARR)**: review first, then *commit to a venue*; a paper may be revised across cycles10-week ARR cyclesExperimental completeness, analysis quality, linguistic insightsee [ref:arrcfp] for current policy
CVPR / ICCV / ECCVComputer visionClassic single-track review, rebuttal~3 monthsVisual results, scale of experiments, breadth of evaluation2026: 16,092 → 4,089 accepted, **≈25.2%** [ref:cvpr2026]
AAAI / IJCAIAI (broad)Review with rebuttal~3 monthsBreadth, applicability, clear contribution to the AI audiencevaries by year (verify)

Verify every row before you rely on it. CFP dates, page limits, and policy details move every cycle; acceptance rates move every year. The stable parts are the *culture* columns.

5.2 The ARR model changes your calendar (NLP only, but worth understanding generally)

Under ARR you submit to a review pool, receive reviews and a meta-review after a 10-week cycle, may revise and resubmit with reviewer continuity, and then commit the reviewed paper to a participating venue, whose programme chairs decide acceptance [ref:arrdates]. Consequences:

5.3 Choosing a venue: a decision tree

1. Which community reads this?  (method-centric → ML venues; language-centric → ACL family; visual → CV venues)
2. Which reviewer would consider this a strong result?  (their yardstick decides novelty)
3. Is your strongest evidence breadth or depth?
     depth/theory → ICML / ICLR / NeurIPS;   breadth/application → AAAI / IJCAI / applied tracks
4. Does the venue's page limit destroy your content?  (8-page ACL main vs. 9-page NeurIPS body + appendix)
5. Is your work a negative result, a replication, or a resource?
     → look for dedicated tracks (Datasets & Benchmarks, Reproducibility, Findings-equivalent outlets)
6. Check the CFP for: LLM-use policy, limitations requirement, compute/reporting requirements,
   supplementary-material rules, and whether anonymous preprint posting is permitted before review.

5.4 Avoiding desk rejection (the 10-minute pre-flight)

Before you press submit, verify mechanically:

5.5 Timeline template

|---|---|

Time before the final ARR/deadline monthMilestone
T-8 weeksResults stable: main table reproduced twice, errors known and quantified
T-6 weeksDraft complete: all sections written, even badly; contribution bullets frozen
T-4 weeksInternal review cycle: a colleague not involved in the work reads it and writes a referee report
T-3 weeksReference integrity pass (Appendix B); ablations completed; efficiency table added
T-2 weeksFigures final, notation pass, completeness pass against the claim→evidence matrix
T-1 weekVenue-format pass, policy/statement pass, PDF checks, supplementary packaged
T-3 daysFinal read-aloud of abstract and introduction; submit **before** the last hour
T+?Calendar the rebuttal window immediately after submission; reserve the time before reviews arrive

---

Part 6 · Rebuttal: turning three reviewers into a decision

A rebuttal is not a defence; it is additional evidence delivered under time pressure. Reviewers read it to decide the paper's fate, and at the large venues the decision is made in area-chair discussion where your answer is the last input.

6.1 Principles

6.2 Comment taxonomy → response pattern

|---|---|---|

Reviewer comment typePatternExtra requirement
Missing experiment / baseline"We agree this strengthens the paper" → run it → give the numberPut results in a compact table; compare honestly, even if modest
Clarification / misunderstandingRestate the premise → point to line → say what you changedAlways edit the manuscript; clarity problems are yours even if the reviewer's reading was unlucky
Novelty challengeSituate against the cited works in one sentence per workDifference must be structural, not rhetorical
Writing / presentationAcknowledge → state the concrete fixDo not argue with writing criticism; fix it
Scope / unfair expectationAccept that it is future work → note it explicitly in LimitationsDistinguish "outside scope" from "we cannot do it"
Disagreement on interpretationOffer your evidence, concede the plausible part, add the caveat to the paperKeep it short; stop after two exchanges

6.3 Templates

Agree and fix

> Reviewer X: The comparison omits [baseline], which is the standard method for this task.

We agree. We have now run [baseline] using the authors' released code with the default settings
(Table 4, rows 3–4). Our method remains ahead by 1.8 points, and the gap narrows on [subset],
which we now discuss in Section 5.3. We have removed the sentence claiming X and replaced it with
the measured statement.

Clarify a misreading

> Reviewer Y: The method requires gold labels at inference time, which makes it unusable.

We see how the draft read that way, and we have rewritten Section 3.2. To be explicit: gold labels
are used **only** during training (lines 214–219). At inference only [signal] is available; we
report this setting in Table 2, rows 5–6. The confusion came from our use of the term "supervision"
without qualification, which we have now corrected throughout.

Scope / future work

We agree this is beyond the current scope. We have added it explicitly to Limitations:
"Our analysis assumes a fixed vocabulary; extending to open-vocabulary settings is left to future work."

6.4 Budgeting the rebuttal

6.5 After a rejection

---

Part 7 · Tool stack and safe automation

Prices change constantly; treat all pricing columns as "check current price".

7.1 Core stack

|---|---|---|

ToolWhat it is forFit
Overleaf (https://www.overleaf.com)Collaborative LaTeX editing, Git sync, reviewer-friendly PDFsEssential for most submissions
Official venue LaTeX templates (ACM / IEEE / venue style files)Format complianceEssential — download from the venue's CFP, never from a random repo
Zotero (https://www.zotero.org) or JabRef (https://www.jabref.org)Reference management; BibTeX hygieneEssential
Semantic Scholar (https://www.semanticscholar.org) / DBLP / Crossref (https://www.crossref.org)Forward snowballing ("Cited by"), author disambiguation, DOI truth sourceStrongly recommended
arXiv (https://arxiv.org) daily listings for cs.LG / cs.CL / cs.CV / cs.AIRecency sweepEssential in fast fields
Connected Papers (https://www.connectedpapers.com) / ResearchRabbit (https://www.researchrabbit.ai)Literature maps; finding clusters you missedRecommended
scite (https://scite.ai)How papers are cited — supported, contrasted, mentionedRecommended for Related Work precision (link unreachable from some networks; verify)
Elicit (https://elicit.com) / SciSpaceSemantic search with extracted result rowsRecommended for first-pass screening
draw.io (https://www.drawio.com) / InkscapeArchitecture and pipeline figuresRecommended
matplotlib (https://matplotlib.org) / PGF-TikZExperiment plots as vector graphicsEssential
Mathpix (https://mathpix.com)OCR maths to LaTeXOptional (check that output is correct — it is often subtly wrong)
Writefull (https://www.writefull.com) / Paperpal (https://paperpal.com) / Trinka (https://www.trinka.ai)Academic-English corpora-trained language checkingOptional; verify after every suggestion
Git + a `Makefile`Reproducible figures and tables built from committed scriptsEssential for honest reproducibility

7.2 Automate the boring verification

Three automations pay for themselves on the first submission:

7.3 Where language models belong in this workflow (and where they do not)

Safe uses: paraphrasing *your* sentences; checking consistency of tense and notation; producing a first draft of a check-list you verify; summarising your own paper into abstract bullets that you then rewrite; generating alt-text and captions you check.

Unsafe uses: generating related work from memory (it fabricates titles); writing citations (this is the exact failure that led to desk rejections at ICLR 2026 [ref:iclr2026]); making up numbers; summarising a paper you have not read.

Rule of thumb: a language model may touch your prose, never your evidence. Every output touching evidence needs an external check — DOI for references, files for numbers.

---

Appendix A · Submission-readiness checklist (run 60 minutes before submitting)

Claims and structure

Method

Experiments

Related work

English and format

Policy

Appendix B · Reference integrity: how to avoid a hallucinated-citation desk rejection

ICLR 2026 desk-rejected submissions whose references pointed to papers that do not exist, after checking every flagged reference against multiple bibliographic databases and web search, with human confirmation and an appeal channel [ref:iclr2026]. The cost of an unverified bibliography is now the whole submission.

Ten-minute procedure:

Automated version: keep this as a script (check_refs.py) run in CI; see §7.2 item 3.

Appendix C · Simulating your own reviewers

Before submission, run the following exercise once per co-author and collect answers independently, then compare. Where the three readings disagree, you have found a writing defect, not three opinions.

Reviewer-simulation prompts for this exercise are in PROMPT_PACK.md, with guardrails preventing fabricated citations.

Appendix D · Glossary

Appendix E · References (verified 2026-09-27)

All links checked on 2026-09-27. Facts labelled "verify" move each cycle; consult the venue's official page before relying on them.

**Cited facts must be re-verified every submission cycle.** Acceptance rates, page limits, policy fields and dates change. A stated number that is one cycle out of date is a factual error in your own manuscript.

FAQ

Does PaperForge invent missing citations?

Never. It lists unverified .bib entries missing DOI/arXiv/URL for you to confirm.

How many seeds are enough for the free check?

The rubric looks for reported seeds/runs and variance (± / CI); venues still set the scientific bar.

Where are the venue links?

Appendix E of the guide — NeurIPS, ICLR, CVPR, ARR, and arXiv moderation pages verified 2026-09-27.

References (verified 2026-09-27)