An External Bayesian Loop for LLM Planning with Incomplete Data

28 September 202612 views

The article explores an architecture in which a separate module estimates the probabilities of hidden environmental states and passes them to a language model instead of the full sequence of past events. The authors explain its connection to POMDP theory and report more accurate beliefs and better results on two test tasks compared with six alternatives.

An External Bayesian Loop for LLM Planning with Incomplete Data

Why LLMs Make Mistakes Under Partial Observability

The authors link errors made by LLM agents in partially observable environments to the absence of an explicit belief distribution over the hidden state. Typically, an agent acts according to a policy conditioned on its history of actions and observations.

The article describes three characteristic errors:

  • making a decision prematurely in response to ambiguous feedback;
  • shifting toward the wrong hypothesis after an informative observation;
  • policy drift as the history grows.

The problem is not just the quality of the answer to an individual query. The agent lacks an explicit representation of which hidden states it considers possible.

How the External Bayesian Loop Works

The Belief-State Engine (BSE) is an inference module outside the LLM. It maintains a Bayesian posterior distribution over the hidden states of a given POMDP—a partially observable Markov decision process.

At each step, the loop works as follows:

  1. The BSE updates the belief distribution over the hidden state.
  2. The LLM receives only this distribution.
  3. The original log of actions and observations remains inaccessible to the LLM.

This approach separates uncertainty tracking from the language model. The key architectural condition is that the LLM does not receive the original history.

Which Guarantees Depend on the Architecture

The authors define a minimal specification of four axioms for a belief-consistent internal state. The axioms themselves are not disclosed here.

The authors prove that the LLM and BSE together form a valid Markov policy on the belief MDP induced by the original POMDP. Provided the LLM never has access to the original history, the system inherits the Bellman optimality guarantees of classical POMDP theory.

Practical criterion: these guarantees apply to the described system and depend on restricting access to the history. They cannot be separated from the architectural conditions.

What the Evaluation Showed

The architecture was tested on the Tiger POMDP and a red-team attack-graph task. The comparison included six baseline methods.

Baseline methodBSE result in both domains
Reactive LLMHigher task return, better belief calibration, and greater decision consistency
Chain-of-ThoughtHigher task return, better belief calibration, and greater decision consistency
ReActHigher task return, better belief calibration, and greater decision consistency
Natural-language belief trackerHigher task return, better belief calibration, and greater decision consistency
QMDPHigher task return, better belief calibration, and greater decision consistency
POMCPHigher task return, better belief calibration, and greater decision consistency

The abstract also states that the effect is not specific to a single model. Ten targeted ablations are intended to isolate the contribution of individual architectural choices.

When the Approach May Be Worth Considering

The BSE supports a distribution over the hidden states of a given POMDP. Therefore, describing the task within such a model is a central criterion for whether the approach is applicable.

The article's evaluation covers the two specified domains. It does not establish that the results transfer to all tasks involving incomplete data.

What Is Available for Verification

The preprint Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability includes code, environment specifications, prompt templates, and seed logs.

Arnab Chattopadhayay and Debdipta Halder submitted the paper to arXiv on September 9, 2026. The practical takeaway should be assessed in light of the environments, comparisons, and conditions governing LLM access to the history.

Frequently asked questions

An External Bayesian Loop for LLM Planning with Incomplete Data