Context Breakdown as an Object of Analysis
Hallucinations in LLMs are linked to disrupted context exchange between tokens during causal generation. This connection is investigated in the paper arXiv:2609.21096, published on September 17, 2026. The authors are Amir Jalilifard, Anderson Rocha, Eric Wong, and Marcos Medeiros Raimundo.
The paper analyzes topological patterns of information flows within attention graphs. The goal is to distinguish hallucinated from non-hallucinated responses.
Key elements of the approach:
- Object of analysis — the attention graph, not the finished response text.
- Feature — the topology of information flows between tokens.
- Task — distinguishing hallucinated from non-hallucinated responses.
Conclusion: disrupted context exchange leaves a measurable structural trace in the attention graph.
Forman-Ricci Curvature and Information Bottlenecks
The key measure in the paper is Forman-Ricci curvature. It identifies structural patterns that indicate information bottlenecks in attention graphs.
The method takes into account semi-local and global characteristics of information flow across attention heads. These characteristics are associated with hallucinated responses.

What goes into the calculation:
- Semi-local characteristics of information flow across attention heads.
- Global characteristics of information flow across attention heads.
- The relationship between these characteristics and hallucinated responses.
Practical criterion: value comes from assessing the flow as a whole, including its semi-local and global characteristics.
Three Signatures of a Hallucinating Response
Hallucinated responses exhibit characteristic attention patterns. These are most pronounced in the final transformer layer.
| Pattern | What the information flow looks like | Where it is most pronounced |
|---|---|---|
| Excessive reliance on self-attention | Tokens loop back to themselves instead of exchanging context | Final transformer layer |
| Diffuse context retrieval | Information from earlier tokens is gathered without focus | Final transformer layer |
| Excessive information compression | The context flow loses detail | Final transformer layer |
Conclusion: three distinct attention patterns converge on one issue — disrupted context exchange.
Single-Pass Detection vs. Baselines
The method was evaluated on several LLMs and widely used benchmarks. The single-pass approach consistently outperforms existing baselines on two hallucination detection benchmarks.
Baselines rely on attention features or multiple responses. Across different LLM architectures, the method achieves competitive results.
| Approach | Basis | Evaluation result |
|---|---|---|
| Forman-Ricci curvature method | Topological features of attention graphs | Outperforms baselines on two benchmarks; competitive results across different LLM architectures |
| Attention-based baselines | Attention features | Outperformed on two benchmarks |
| Multiple-response baselines | Multiple responses | Outperformed on two benchmarks |
Practical criterion: single-pass detection does not require multiple responses and is applicable across different LLM architectures.



