Giving a language model a longer context window does not make its literature reviews ready for publication, according to an evaluation of AI-assisted academic writing. Two researchers scored 20 generated reviews across 15 dimensions after supplying source material from Semantic Scholar and arXiv.
Longer contexts helped the models include a broader range of information and maintain coherence across more material. They also amplified familiar weaknesses: repeated points, omission of important work and a tendency to describe papers one by one instead of synthesizing evidence into an argument. The generated text could provide a useful starting overview, but it did not consistently meet academic publishing standards without expert revision.
The study reinforces a distinction between retrieving more text and understanding a research field. A large context can expose the model to additional sources, yet it does not guarantee that the system weighs methods, conflicting results or gaps appropriately. The experiment covered 20 reviews, two evaluators and a particular set of models and domains, so it cannot establish the same effect for every scholarly workflow. Researchers can use generated drafts to map themes or organize reading, but they still need to verify citations, identify omitted work and rewrite descriptive summaries into a defensible synthesis before treating the output as scholarship.