Researchers have introduced LongNovel, a benchmark for studying hallucinations in summaries of long novels. The paper argues that novels are useful for this problem because their length, characters, and narrative structure create many opportunities for subtle errors.
Long-context models can now process much more text than earlier systems, but a larger window does not guarantee faithful summaries. A model may still invent events, merge characters, or miss how a detail changes across the story.
Benchmarks like LongNovel help separate impressive context length from actual reliability. The limitation is that performance on fiction does not automatically transfer to legal, scientific, or business documents, but it gives researchers a controlled way to measure failures that short articles rarely expose.