Task-structure preferences

Authors
Affiliation

Alex Lepauvre

TUD Dresden University of Technology

Florian Ott

TUD Dresden University of Technology

Stefan Kiebel

TUD Dresden University of Technology

Published

July 28, 2026

Abstract

[]

Keywords

Decision making, Reinforcement learning

1 Introduction

A central idea in Bayesian accounts of perception and sensorimotor control is that the brain does not rely on current sensory evidence alone (Knill and Pouget 2004; Lee and Mumford 2003; Rao and Ballard 1999; Clark 2013; Friston 2010) because sensory evidence is often noisy, incomplete, or costly to acquire with high precision. Prior information can reduce this uncertainty by providing an efficient summary of regularities learned from experience and from the statistical structure of the environment. The Bayesian solution is therefore to combine current evidence with prior information. As formulated in the ‘Bayesian brain hypothesis’, the influence of prior information should depend on uncertainty: when current evidence is unreliable, behavior should rely more strongly on the prior. Conversely, when current evidence is reliable, it should dominate behavior. For example, Körding and Wolpert (2004) provided a canonical demonstration of this principle, showing that participants combine learned priors with sensory feedback according to their relative uncertainty.

A similar logic may apply to planning-based decision making. In this case, the role of sensory evidence is taken by an internally generated estimate of action value. In sequential decision tasks, such estimates are assumed to be generated by internal forward planning, which provides state-specific information about the long-value of choosing an action. However, generating these estimates is assumed to be computationally costly because it requires evaluating future outcomes in complex and dynamic environments. In analogy to the Bayesian view on perception, agents may therefore combine their evidence, i.e. their internal estimates, with stored prior information about which actions are usually appropriate in a given task. In this view, action selection reflects a balance between a costly state-specific estimate and a stored prior tendency to act in a particular way. This idea has been applied previously in a Bayesian account of the balance between habitual and goal-directed control, in which habitual tendencies were modelled as priors over policies that are balanced against goal-directed information about current outcome contingencies (Schwöbel et al. 2021, 2024). This approach also resonates with resource-rational accounts of planning, which postulates that the brain adapts its planning computations to the costs and benefits of using its relatively limited cognitive resources (Gershman et al. 2015; Lai and Gershman 2021; Callaway et al. 2022; Tishby and Polani 2010; Parush et al. 2011).

Building on Bayesian accounts of prior–evidence integration and resource-rational accounts of costly planning, we assume that choices in complex sequential decision making reflect the integration of two sources of information about action. First, participants have acquired, for example during training, task-structured preferences that reflect regularities in the task structure and their own experience within the task. Second, as is typically modelled in reinforcement learning, for each trial, participants use forward planning to derive values for each available action in an online fashion. Assuming a Bayesian approach, we hypothesize that these two sources are integrated by weighting the contribution of planning against preferences by their respective uncertainties. This hypothesis implies that on trials with strong preferences, these should guide choices and produce faster responses because less planning is needed. On trials with weak or uncertain preferences, decision making should be more weighted towards forward planning and consequently, response times should be increased due to the increase in computations.

In this paper, we test this hypothesis in a complex sequential decision-making task, using choice model comparisons and response-time analyses. Specifically, we use optimal forward planning as a benchmark and ask whether participants’ deviations from this benchmark are systematic and better explained by task-structured preferences. We compare against several alternative models of how participants may combine planning-derived decision values with potential preferences. Our results show that participants’ choices are best captured by a model that combines planning-derived decision values with state-dependent preferences constructed from multiple task features. We further show in additional analyses that these preferences interact with planning such that choices are guided more strongly by planning when preferences are weak, and more strongly by preferences when they are decisive. Together, these findings suggest that, in complex sequential decision problems, costly forward planning is integrated with structured, task-informed preferences that enable faster and less computationally demanding choices when they provide a clear action tendency.

2 Methods

2.1 Data

To test our hypothesis, we reused data from a previously published study by Ott et al. (2022a) and colleagues, in which 40 participants (22, mean age=24.4, SD=4.6) took part in a sequential decision making task (Ott et al. 2022a, 2022b). The study was approved by the Institutional Review Board of the Technische Universität Dresden and conducted in accordance to ethical standards of the Declaration of Helsinki.

Figure 1 shows a schematic representation of the task. Participants were instructed to gather as many points as possible throughout the task. In each trial, they were presented with an offer associated with a reward of either 1, 2, 3 or 4 points, which they could either accept or reject. To accept the offer in a trial, participants had to pay an energy cost of either 1 or 2 energy points (low cost, LC and high cost, LC respectively). When they rejected the offer, they gained 1 energy point. Participants’ energy level was capped at 6, so their energy could range from 0 (depleted energy) to 6 (max energy). Offers varied in each trial following a uniform distribution (equal probability of each offer in each trial). Costs remained fixed for 4 consecutive trials, which we refer to as a segment (see Figure 1 B), and participants were informed of the price of the current segment and the next segment. The costs of the next segment was pseudorandomized to equate the probability of each transition (from current HC or LC to future HC or LC, see Figure 1 C) and ensure that each transition occurred 15 times. When participants accepted an offer that they could not afford (current energy level below current cost), no reward was awarded and a warning was displayed.

Experimental design
Figure 1: Experimental design: (A) Single trial procedure, each frame represents a step within the trial with duration of each step above. Top row depicts participants accepting, bottom row depicts participants rejecting. In each trial, participants are presented with N golden cups, the number of cups symbolizes the amount of reward. The blue bar represents energy level, they yellow bar represents accumulated reward so far, lightning symbols at the bottom right of each frame represent current cost (left) and future cost (right). (B) Between trial costs dependencies. A segment consists of 4 trials within which cost is fixed, and participants are aware of the cost in the current and future 4 trials but not beyond. (C) Transition structure. Cost can transition from low cost (LC) to high cost (HC), LC to LC, HC to LC and HC to HC. Reproduced from Ott et al. (2022a)

Accordingly, each trial was defined by several task factors: the current offer, the participant’s current energy level, the current cost of accepting, and the cost in the upcoming segment. Participants had to decide whether to accept or reject the offer under the specific combination of these factors. Because accepting depleted energy while rejecting replenished it, each choice also changed the state in which the participant entered subsequent trials. Maximizing overall return was therefore non-trivial and required considering the future consequences of current choices.

In the beginning of each trial, a fixation cross was presented for 0.5s in the middle of the screen. Then, an offer appeared represented by the number of “cups” presented in the middle of the screen and choice options were presented on each side of the screen, surrounded by a black frame. Participants had 5s to provide an answer, by selecting one or the other options (accept or reject), in which case the frame around the chosen option disappeared. Following their choice, the resulting change in energy and points were displayed (green indicating an increase in point and/or energy, red indicating a decrease in energy) for 1s. The trial ended with an inter-trial interval (ITI) ranging from 2 to 5s (uniformly sampled, rounded to the first decimal) in which the offer disappeared and unframed choice options remained on screen. If participants took longer than 5s to provide their answer, a warning message was displayed and the next trial started.

To familiarize themselves with the task, participants conducted a training session of 144 trials, corresponding to 36 segments. Participants then conducted the experiment consisting of 240 trials (60 segments). Participants started with an energy budget of 3, and their reward points were accumulated throughout the entire experiment. Trial sequence was pseudorandomized, such that each costs transitions (LC to LC, LC to HC, HC to LC and HC to HC, see Figure 1 C) occurred equally often within the training and test sessions, resulting in 9 repetition of each transition in the training session and 15 in the experiment. In addition, the offer were sampled to be equally distributed within each transition. The sequence of costs transitions and offers was kept the same across participants. Participants were offered two to three minutes breaks inside the scanner every 20 segments. The test sessions lasted around 40min.

2.2 Forward planning component

To model optimal behaviour in the task, we operationalized our task as a Markov Decision Process (MDP) with the following tuples:

\[ MDP = (S, A, P, R) \]

Where \(S\) represents all possible states in our task, where each state \(s\) consists of a combination of the experimental variables in our task: \(S=E \times O \times CC \times FC \times T\), with energy \(E={0, 1, ..., 6}\), offer \(O={1, 2, 3, 4}\), current segment cost \(CC={1, 2}\), future segment cost \(FC={1, 2}\) and trial \(T={1, 2,..., 13}\). Time steps (i.e. trials) had to be incorporated in the state space to account for the cost being trial dependent. \(A\) is the set of possible actions in our task (\(accept=1, reject=0\)), \(P\) is the transitional probability from a given state to all other states (\(P=P_{\forall s \in S}(s'|s, a)\)), and \(R\) is the reward function characterizing the reward participants can obtain in a given state for each action (\(R=R_{\forall s \in S}(s|a)\)).

While optimal behaviour would require considering outcomes up until the end of the task, we modelled planning over a limited horizon, following the approach used by Ott et al. (2022a). Participants had explicit information about the current and upcoming segment, corresponding to an eight-trial horizon. Solving the task only over these eight trials would implicitly treat the task as ending after the upcoming segment, leading the optimal policy to spend all remaining energy by the end of that horizon. To avoid this terminal artefact, we included one additional segment beyond the instructed horizon. Thus, action values were computed over a twelve-trial planning horizon, with a terminal value set after the twelfth trial, yielding 13 time points in the dynamic-programming solution. In the first 8 trials, the probability of the next state reflects the offer-related stochasticity, while costs are known. For trials beyond the known horizon, the uncertainty associated with the offer is compounded by the uncertainty related to the next cost transition (see Ott et al. 2022a for more details).

We applied backward induction to compute the long-term value associated with each action in any given state (i.e. the \(Q\) function Sutton et al. 1998). Participants maximizing return should choose the action with the largest value. As there were only two possible actions in our task (accept or reject), we computed decision values (DV) as:

\[ DV_s = Q(s, a=1) - Q(s, a=0) \]

In states where the DV are negative, participants should reject, and in states where DV are positive, participants should accept the offer.

2.3 Modelling of participants responses

Participants behaviour were modeled using a logistic regression: \[ P(a=1) = \frac{1}{1+e^{-\eta}} \tag{1}\]

We specified four models that differed in how the linear predictor \(\eta\) was specified. All models included planning-derived DV, but differed in the additional regressors used to explain systematic deviations from optimal planning. We describe these models in the following, starting with the pure planning model as a benchmark and then adding increasingly structured accounts of participants’ deviations from optimal planning.

2.3.1 Planning model

To test whether there is any systematicity in participants deviations from optimality, we first specified a pure planning model, in which participants responses is modeled as a function of DV only:

\[ \begin{aligned} \eta_{planning} &= \beta_{plan} \times DV \end{aligned} \tag{2}\]

This model acted as a benchmark. If the suboptimality (i.e. deviations from forward planning only) in participants responses reflects unstructured noise, alternative models with additional regressors would only increase model complexity but not accuracy.

2.3.2 Frequency prior model

Alternatively, participants deviations from optimality might reflect a recency bias, whereby participants choice on a given trial is biased towards the most frequent action in recent history. Recent accounts of planning consider that priors consist of a state independent recency bias (Schwöbel et al. 2021, 2024; Kuperwajs and Ma 2021). Similarly, under the policy compression framework, the optimal policy integrates planning and marginal action probability across states (Lai and Gershman 2021, 2024). To test whether in our task participants responses are biased towards most frequent action taken in previous trials, and approximate the policy compression framework, we fitted a frequency model, which models participants responses as a function of DV combined with a recency bias:

\[ \begin{aligned} \eta_{frequency} =& \beta{plan} \times DV + \beta{prior} \times P(A) \end{aligned} \tag{3}\]

We computed \(P(A)_t\) in each trial \(t\) using a linearly decaying moving average going back \(k=20\) trials, assuming temporally discounted effects of past actions (see Supplementary 1 for exploration of the fit at different \(k\)):

\[ \begin{aligned} P(A)_t =& \frac{\sum_{i=t-k}^{t-1}x_i \times w_i}{\sum_{i=t-k}^{t-1}w_i} \end{aligned} \tag{4}\]

Where \(x_i\) denotes the action taken on trial i coded as accept or reject, and \(w=[1, 2, ..., k]\) so that the weights increase linearly from 1 for the oldest included trial to \(k\) for the most recent trial.

2.3.3 Context model

In addition, we used the best-fitting model (among 21 models) from Ott et al. (2022a), to test whether participants treat groups of states as separate contexts in which they plan to different extents. We refer to this model as the context model. Here, the planning component (DV) is split between extreme and intermediate offers (offers [1, 4] and [2, 3] respectively) to reflect the fact that participants treat those as separate contexts, with different control requirements. In addition, the model comprise offer specific and energy related flexible biases regressors:

\[ \begin{aligned} \eta_{context} =& \beta{plan23} \times DV_{23} + \beta{plan14} \times DV_{14} + \\ &\beta_{O1}\mathbf{I}_{O1} + \beta_{O2}\mathbf{I}_{O2} + \beta_{O3}\mathbf{I}_{O3} + \beta_{O4}\mathbf{I}_{O4} + \\ &\beta_{basic}\mathbf{I}_{basic} + \beta_{maxE}\mathbf{I}_{maxE} + \\ &\beta_{minE_{LC}}\mathbf{I}_{minE_{LC}} + \beta_{minE_{HC}}\mathbf{I}_{minE_{HC}} \end{aligned} \tag{5}\]

\(\mathbf{I}\) indicate dummy regressors encoding specific experimental conditions: \(\mathbf{I}_{basic}\) encodes trials where energy is sufficient to accept the offer and inferior to 6, \(\mathbf{I}_{maxE}\) encodes trials where energy is equal to 6, \(\mathbf{I}_{minE_{LC}}\) for trials where energy is too low to accept the offer when cost is low, \(\mathbf{I}_{minE_{HC}}\) for trials where energy is too low to accept the offer when cost is high, \(\mathbf{I}_{O1-4}\) for trials with offer 1-4 where energy is sufficient to accept.

2.3.4 Preferences model

Finally, we specified a preference model to test whether participants’ choices reflect the integration of forward planning with task-structured preferences. The model contains three components. First, as in the planning model, choices depend on the planning-derived decision value. Second, choices depend on a state-dependent preference score. This preference score is constructed by assigning weights to the levels of the main task features, namely offer, energy, current cost, and future cost, and linearly combining the corresponding weights for the current state. Third, the model includes a modulation term that tests whether the influence of planning depends on the uncertainty of the preference score. We quantified this uncertainty as the entropy of the preference-derived action tendency. The model therefore tests the prediction that participants rely more strongly on planning when their preferences are weak or uncertain:

\[ \begin{aligned} \eta_{pref} &= \beta_{plan} \times DV + pref + \beta_{modulation} \times(H(Preferprefences) \times DV) \end{aligned} \tag{6}\]

Where: \[ \begin{aligned} pref = \mathbf{X_{pref}}\mathbf{\beta} \end{aligned} \]

And \(H(pref)\) is the entropy of the preference score:

\[ \begin{aligned} H(pref) = -\sigma(pref) * log(\sigma(pref)) - (1-\sigma(pref)) * log(1-\sigma(pref)) \end{aligned} \tag{7}\]

With:

\[ \begin{aligned} \sigma(x) = \frac{1}{1+e^{-x}} \end{aligned} \tag{8}\]

Participants features dependent preferences are estimated as latent variables. \(\mathbf{X_{pref}}\) is a \([M \times N]\) (M=measurements, N=experimental levels) matrix dummy coding each factors of the experimental setup (energy, offer, current and future costs), and \(\mathbf{\beta}\) is a vector of weights estimated for each level of each factor.

2.4 Model fitting and comparison

Models were fitted to participants’ response data as Bayesian hierarchical logistic regressions using PYMC (Abril-Pla et al. 2023) and Bambi (Capretto et al. 2022). We excluded trials where participants did not provide a response or when reaction time (RT) exceeded the response window (5s). Across models, the following parameters were modelled hierarchically, with separate participant-level estimates drawn from group-level distributions: \(\beta_{plan},\ \beta_{plan23},\ \beta_{plan14},\ \beta_{pref},\ \beta_{Interaction},\ \beta_{basic},\ \beta_{O1},\ \beta_{O1},\ \beta_{O2},\ \beta_{O3},\ \beta_{O4}\), while the rest were fixed across participants. For all parameters, we used weakly informative hyperprior distributions \(\mu \sim \mathcal{N}(0, 2)\) and \(\sigma \sim Halfnormal(0, 2)\). All models were fitted using 4 chains of 2,000 samples each (1,000 warmups), resulting in 4,000 samples in total.

We used Pareto-smoothed importance sampling to approximate leave-one-out cross-validation (PSIS-LOO, Vehtari et al. 2017) to estimate the expected log pointwise predictive density (elpd) which we used to compare the fit of these different models accounting for model complexity. Furthermore, we perform model comparison within each single subject to estimate the robustness of the winning model in our population sample, by computing the sum of the pointwise predictive accuracy for each participant and model, yielding a score for each participant and model.

Each of the four main models presented in the main text was compared against alternative parameterizations to ensure the robustness of our findings and rule out alternative explanations. For example, the frequency prior model was tested with varying temporal integration windows (see Figure S3 and Table S4), and the context model was selected from a prior study (Ott et al. 2022a) after comparison with 21 alternatives. Similarly, reduced versions of the preferences model were tested against the full model by removing each component (planning, prefefrences and interactions) as well as individual task factors (see Figure S2 and Table S3). For clarity, we focus on the four best-fitting models (planning, frequency prior, context, and preferences) in the main text, with comparisons to all alternatives detailed in the supplementary material.

2.5 Experimental factors relevance analysis

The preference model estimates feature-specific preference weights for each task factor (energy, offer, current and future costs), revealing that participants’ deviations from optimal planning are more strongly driven by some factors of the task (energy and offers) than others (current and future costs). To quantify the behavioral contribution of each factor, we compared the full preference model to reduced models in which all preference regressors associated with one task factor were removed. A large decrease in predictive accuracy after removing a factor indicates that preferences associated with this factor explain a substantial part of participants’ choices beyond the planning-derived values.

The observation that participants deviation from optimal planning are systematically influenced by certain task features raises the question of whether preferences constitute unhelpful biases (suboptimal tendencies without adaptive value in the task) or whether they reflect a structured yet simplified solution to the task? In other words, could participant’s preferences constitute a low dimensional decomposition of the optimal policy, providing a cheap alternative to guide behavior in the absence of planning? To test this hypothesis, we conducted a series of analyses to test whether factors’ weights in participants preferences relate to the structure of the task.

First asked whether the factors that are strongly weighted in participants’ preferences are more relevant for reward maximization. For this purpose, we constructed reduced MDPs in which one task factor was omitted by averaging transition probabilities and rewards across the levels of that factor. We then computed the optimal policy in the reduced MDP, projected this policy back onto the full MDP, and calculated the expected return. This analysis measures the reward sensitivity of each task factor: if ignoring a factor leads to a large loss in return, optimal behavior depends strongly on information about that factor.

Second, we asked whether participants’ preferences themselves define a useful policy in the absence of planning. We derived policies from the preference component of the full model and from reduced preference models in which one task factor was omitted. We then computed the expected return of each preference-derived policy relative to the optimal policy. If task-structured preferences provide an adaptive approximation to optimal choice, the full preference policy should outperform a stochastic policy and removing behaviorally important factors should reduce expected return.

Finally, we explored a possible source of these preferences. In the policy-compression framework, compressed policies can be described in terms of marginal action probabilities. We therefore computed, under the optimal policy, the probability of accepting each level of each task factor after marginalizing over all other factors. For example, we computed how often an offer of 1 should be accepted across all energy and cost states. We then compared these factor-wise marginal action probabilities with the fitted preference weights. A close correspondence would suggest that participants’ preferences may approximate feature-specific marginal action tendencies of the optimal policy.

2.6 Reaction times analysis

In our task, exact computation of optimal policy is prohibitively expensive, given the size of task’s state space (1456 states), the number of steps to be planned ahead and the highly probabilistic nature of the task. Preferences derived from task knowledge or from task experience are presumably readily available to the participants, or require minimal computation. A part of our hypothesis is that instead of systematically considering all future outcomes before making a choice, participants consider the strength of their preferences to determine how much they should engage in forward planning. If this were the case, we would expect participants’ RT to reflect the strength of their preferences, as stronger preferences should be associated with less planning, hence faster RT.

Ott et al. (2022a) showed that response times were slower when the planning-derived values of accepting and rejecting were similar, and that this conflict effect depended on task context, being stronger for intermediate than for extreme offers. Under our hypothesis, this context dependence should arise from a more general principle: the influence of planning conflict on response time should depend on the strength of participants’ task-structured preferences. When preferences are strong, participants should rely less on planning, reducing the influence of planning-derived value differences on choice and response time. In contrast, when their preferences are weak, they should engage in planning, making their choices and response times more sensitive to planning-derived value differences.

To test whether the effect of planning conflict should depend on preferences strength, we modelled participants’ RTs as a function of planning conflict (i.e. \(conflict = - |DV|\)), preference entropy (fitted in the preference response choice model) and the modulation of planning by preference entropy:

\[ log(RT) = \beta_{conflict} * Conflict + \beta_{pref} * H(preferences) + \beta_{interaction} * (Conflict : H(preferences)) \tag{9}\]

We predicted a positive modulation, which would indicate that the larger the preference entropy (i.e. the weaker the preference), the larger the reliance on planning and the larger the RT.

3 Results

Performances indicate that participants were able to perform the task well. Participants provided their responses with an average reaction time of 0.90s \(\pm\) 0.26 (s.d.) and failed to give their response before time out in only 65.00% \(\pm\) 133.32 (s.d.) out of 240 trials. On average, participants obtained 332.43 \(\pm\) 6.64 (s.d.) points on average throughout the experiment.

3.1 Choice behaviour

We modelled participants response choices using 4 different models, to investigate the systematicity in participants deviations from optimal planning, which would be achieved by pure planning. As already found in (Ott et al. 2022a), participants’ responses systematically deviate from optimality: the pure planning model has the lowest fit compared to the other three models (\(\Delta ELPD_{pref - planning}\)=585.98 \(\pm\) 55.58 s.e., see Figure 2 A). Furthermore, our results indicate that participants biases do not reflect frequency prior (\(\Delta ELPD_{pref - frequency}\)=580.89 \(\pm\) 55.57), nor a splitting of the task into separate contexts (\(\Delta ELPD_{pref - context}\)=164.21 \(\pm\) 52.46). Rather, participants exhibit preferences derived from the structure of the task, as the Preference model was found to fit the data best (see Figure 2 A).

The improved fit of the preference model is reflected in Figure 2 C, showing the proportion of match between observed and predicted responses separately for each model and offer values, where we can see that the proportion of matches is on average higher for the preference compared to all other models. This result is highly consistent across participants: the preference model fits the data best for 35 out of 40 participants (see Figure 2 B).

The estimated preference weights associated with each task factor reveal systematic biases in participants’ behavior (see Figure 2 D and Table S2 for posterior of each fitted parameter). Participants tend to reject low offers (1 and 2) more often than optimal planning would suggest, while they are biased towards accepting high offers (3 and 4). These biases are stronger for extreme offers (1 and 4) compared to intermediate ones (2 and 3).Energy state also plays a critical role in shaping participants’ decisions. When energy is maxed out (e = 6), participants almost always accept offers, even when rejecting would be more reward-maximizing. This behavior is intuitive: rejecting offers at max energy does yield additional energy, so participants prioritize immediate rewards, even at the cost of future ones. In contrast, when energy is at 0, participants almost never accept offers, even though the planner predicts a small probability of acceptance (~0.25). This discrepancy arises because the planner calculates decision values (DVs) based on long-term returns, where the difference between accepting and rejecting at energy = 0 is negative but not particularly large. However, participants appear to adopt a stricter rejection policy, likely reflecting an aversion to accepting offers that yield no immediate benefit. Importantly, energy = 0 occurs in only 17.85% \(\pm\) 4.02 of trials, so this counter-intuitive effect at energy 0 cannot alone explain the success of the preference model. Finally, costs (both current and future) have a smaller impact on participants’ behavior, with a slight bias towards accepting offers when costs are low.

Figure 2: Results of participants choice behaviour models. (A) Result of the model comparison, displaying the expected log pointwise predictive density (ELPD) associated with each model (closer to 0 indicate better fit, see Table S1 for model comparison results). (B) Model fit for each participant, color coded per model of best fit. (C) Observed and fitted probability of accepting the offer separately for each offer by each of the models. The single dots represent the average acceptance probability for each subject. (D) Fitted preference parameters of the preference model. Top left figure depicts the parameters associated with offer parameters, top right associated with each cost, bottom with each energy level. The bar in the middle of each distribution is the median and the dotted lines represent the 2.5 and 97.5% quantiles. The dots represent the mean of the parameter estimated for each participant
Source: Article Notebook

We next tested the central prediction that the influence of planning should depend on the certainty of the preference signal. The preference model revealed a positive modulation of planning by preference entropy (\(\beta_{modulation}\)=1.31, CI=[0.36, 2.25]), indicating that planning-derived DV had a stronger influence on choices when preferences were more uncertain. Conversely, when preferences were strong, choices were less sensitive to planning-derived decision values and were more strongly guided by the preference component.

To visualize this effect, we plotted the fitted probability of accepting as a function of planning-derived DV for different preferences values (see Figure 3 A). This visualization shows the model-implied effect directly: the choice curve is steeper when preferences entropy is high (preference scores close to 0), indicating stronger reliance on planning, and flatter are preferences entropy increases (preference scores away from 0, deep blue and deep red lines in Figure 3 A), indicating that choices are already largely determined by the preference signal. Figure 3 B shows the observed probability of accepting the offer as a function of DV, separately when preferences were positive (upper panel) and negative (lower panel), further showing the effect of preferences on participants response choice: participants accept offer at lower DV when preference scores are positive compared to when preference scores are negative (Figure S1 shows participants responses and fitted values as a function of DV and preferences without binning).

Figure 3: Relationship between preferences and DV. (A) Fitted probability of accepting an offer as a function of planning DV, separately for preference score of [-4, -2, -1, 0, 1, 2, 4], to highlight the interaction between preferences scores and DV. Line represents mean of posterior distribution, shaded area is 95% C.I. (B) Offer acceptance probability (y-axis) binned per decision values (x-axis), separately for trials in which preference score is positive (top row, red bars) (bottom row, blue bars).

We next asked whether these model components were behaviourally meaningful beyond predicting individual choices. Specifically, we tested whether overall task performance was related to participants’ reliance on planning, the strength of their task-structured preferences, and the extent to which preference uncertainty modulated planning. As larger \(\beta_{planning}\) values indicate stronger reliance on planning, we expected participants with larger \(\beta_{planning}\) to obtain higher returns. In contrast to planning, preferences might not necessarily maximize reward, so participants with strong preferences (i.e. low preference entropy) should obtain lower return on average. Finally, larger \(\beta_{modulation}\) indicate that participants plan more when their preferences are weak (i.e. preferences entropy is larger), and should accordingly obtain larger returns compared to participants’ who do not engage more strongly in planning when their preferences are weak. Accordingly, we predicted that \(\beta_{planning}\), preferences entropy and \(\beta_{modulation}\) should be positively correlated with return.

In line with our prediction, we observed a positive correlation between \(\beta_{planning}\) and return (r=0.57, p=0.000, Figure 4 A), showing that participants who rely more on planning get larger return. Furthermore, \(\beta_{modulation}\) is strongly and positively correlated with return (r=0.66, p=0.000, Figure 4 B). However, contrary to our prediction, we observed a negative correlation though only marginally significant correlation between preferences entropy and return (r=-0.28, p=0.082, Figure 4 C). Visual inspection of Figure 4 D revealed that the correlation appears to be driven from the subject with very low return (2.63 standard deviations from population mean). When removing this subject, the correlation between preferences entropy and return is not marginally significant (r=-0.08, p=0.648), while it remains strongly significant for planning (r=0.46, p=0.003) and modulation (r=0.60, p=0.000) when the same subject is removed.

Figure 4: Fitted parameters and performances (A) Correlation between participants return and the fitted \(\beta_{planning}\) parameter. Each dot represent a single participant, the line is a correlation fitted using optimal least square, legend indicates pearson correlation alongside p value. (B) Correlation between participants performances and the average preference entropy. Same convention as (A). (C) Correlation between participants performances and \(\beta_{modulation}\). Same convention as (A)

3.2 Experimental factors relevance

Our results indicate that participants’ choices are shaped by two sources of information: planning-derived values and task-structured preferences. These preferences are not a single global bias. Instead, participants appear to assign preferences to different features of the task, such as the offer, the current cost context, and upcoming transitions, and combine these feature-specific preferences into an overall preference for each action in a given state. Importantly, not all features of the task are weighted equally. Figure 2 C shows that the parameters associated with offers and energy weigh strongly in participants’ preferences compared to costs (both current and future). The fitted preference weights already suggested that participants’ preferences were structured primarily by offer value and energy. We next asked whether the same pattern was evident from a complementary model-comparison perspective. To do so, we compared the full preference model with reduced models in which the preference terms for one task factor were removed at a time. This effectively tests how decision making changes when participants are assumed to ignore one specific sources of task information, such as offer value, action costs, or future transitions.

Figure 5 A displays the difference in fit between the full preference model against models in which the preferences related to each experimental factor were removed one by one. Larger values indicate that the specific factor is more relevant for the full model’s fit. These results indicate that offer (\(\Delta ELPD_{full - offer}\)=137.35 \(\pm\) 54.44 s.e.), followed by energy (\(\Delta ELPD_{full - energy}\)=133.96 \(\pm\) 51.57 s.e.), has a much larger impact on participants final responses, whereas costs-related preferences have a smaller impact on the overall fit (\(\Delta ELPD_{full - CC}\)=47.00 \(\pm\) 49.89 s.e. and \(\Delta ELPD_{full - FC}\)=27.34 \(\pm\) 48.62 s.e.).

Figure 5: Importance of each parameters in participants response and optimal solution. (A) Difference in fit between the full preference model and models in which each of the experimental factors were removed from preference estimations. Values closer to 0 indicate that the factor does not lead to a strong improvement of the model fit (i.e. the factor weight less in participants final decision). (B) Loss in return when planning when ignoring each task factor (0 indicates no loss compared to optimal policy). This figure represents the reward sensitivity of each task factor: the larger the return loss, the more actions are dependent on the levels of the factor for return maximization. For example, the large loss associated with energy reflects the fact that when all energy levels are treated the same leads to accepting offers when energy is low, leading to low energy, preventing accepting future returns. (C) Loss in return of policies derived from the preference component of the full and reduced preference models omitting each task factors. For each model (full and reduced), we derive a policy from the preference component alone and calculated the expected return under such a policy, relative to the optimal policy. The y-axis represents the loss in relative return, 0 indicates that a preference policy performs exactly as well as the optimal policy. The grey line represents the return associated with a fully stochastic policy (i.e. equal probability of each action in each state). (D) Comparison between preference (in probability space) parameters and the marginal action probability of each task level.

We next asked whether the fitted preferences could serve as adaptive biases when planning is absent. As a first step, we quantified how strongly optimal return depends on each task factor. To do so, we derived policies from reduced planning models in which one task factor was ignored, and computed their return loss relative to the optimal policy (Figure 5 B). Our results show that ignoring current and future costs when planning leads only to minimal losses (0.78% and 0.98% respectively), while ignoring the offer level leads to a 13.94% loss. The biggest loss is incurred when ignoring energy levels. This leads to a loss in returns of 72.04%, which is even larger than a purely stochastic policy (i.e. equal probability of accepting and rejecting an offer on each trial). This indicates that reward maximization under planning is highly sensitive to energy. This large loss reflects the structure of the task. A policy that ignores energy cannot adapt choices to the current energy budget. As a result, it can deplete energy in low-energy states as its behaviour is driven by offer levels. The large reward sensitivity of energy is not mirrored in Figure 5 A because that analysis measures the additional predictive contribution of preference terms after accounting for planning-derived decision-values. Energy is already strongly reflected for by the planning component, so removing energy-related preference terms has a smaller effect on model fit than its importance for optimal return would suggest.

Further, we asked whether the fitted preferences could support adaptive choice when planning is absent. If preferences capture useful task structure, policies derived from the preference component alone should be suboptimal, but should still produce lower return loss than a stochastic policy. In addition, removing task factors that strongly shape participants’ preferences should reduce the return of these preference policies. To quantify this, we computed the expected return loss of policies derived from the preference component alone. The full preference policy produced, as expected, a non-zero return loss of 28.53%, indicating that preferences did not fully replace planning, see Figure 5 C. However, this loss was smaller than the loss of a stochastic policy (grey line), showing that the fitted preferences captured task structure that was useful for return maximization (Figure 5 C). Removing current or future cost from the preference policy produced little additional loss relative to the full preference policy (29.72 and 28.38% respectively). In contrast, removing offer value led to a substantial loss in return of 51.52%. The largest loss occurred when energy was removed (69.80% loss), with performance falling below that of the stochastic policy. This suggests that an energy-blind preference policy is not merely less informative, but can become systematically maladaptive by accepting attractive offers without regard to the current energy budget.

Together, these results suggest that participants’ preferences were not arbitrary biases, but adaptive approximations to optimal choice. They supported relatively good performance in the absence of planning, while remaining insufficient to fully replace planning-derived decision values.

How do participants derive such adaptive preferences? A potential explanation is that participants have access to some approximation of the marginal action probability associated with each factor of the task under the optimal policy (of how often should an offer of 1 be accepted, across all energy and costs levels for example). To explore whether this might be contributing to an underlying mechanism for learning participants’ preferences, we computed the factor-dependent marginal action probability under the optimal policy. Figure 5 D plots the preferences parameters side-by-side with the marginal action probability, showing that participants’ preferences relatively closely reflect the feature specific marginal action probability of the optimal policy, with some notable deviations. For offer levels 1 and 2, preferences are lower than marginal action probability, and higher when for offer levels 3 and 4. Across all other tasks levels, participants preferences are larger than marginal action probability (except for energy=4), suggestive of an optimistic bias (i.e. tendency to accept offers more often than they should). Nonetheless, the overall similarities between both variables suggest that this is a plausible explanation of participants preferences.

3.3 Reaction times

Under the proposed hypothesis, participants’ reliance on planning should depend on the strength of their preferences. Strong preferences should reduce the need for planning and therefore lead to faster responses, whereas weak or uncertain preferences should increase reliance on planning and lead to slower responses. Previous studies have shown that RT is positively correlated with planning conflict, indicating that less planning is required when the value associated with each action is more clearly distinct (Ott et al. 2022a). We therefore expected preference strength to attenuate the effect of planning conflict on RT: when participants’ preferences provide a clear action tendency, conflict between planning-derived action values should have a weaker influence on response times.

Indeed, we found that participants’ RT was influenced by (i) conflict (\(\beta_{conflict}\)=0.11, CI=[0.09, 0.13]), (ii) entropy of preferences (\(\beta_{pref}\)=0.44, CI=[0.41, 0.49]) and (iii) the interaction between the two (\(\beta_{modulation}\)=0.11, CI=[0.41, 0.49], see Figure 6 A). This result extends on previous findings showing that participants do not rely on planning to the same extent across all states of the task. Our results suggest that rather than splitting the task in independent contexts and engaging in planning to a different extent across those (Ott et al. 2022a), the differential amount of planning across states instead reflects the interaction between strength of preferences and conflict. The positive interaction (see Figure 6 A) between conflict and entropy of preferences indicates that the reliance on planning is modulated by the strength of participants’ preferences. Specifically, as can be seen in Figure 6 B, planning depends less on conflict when preferences are strong (blue line), but more so when preferences are weaker (orange and green lines).

In other words, when preferences are strong and prescribe a clear action tendency, participants do not engage much effort in planning, even if conflict between decision values is high. Participants rely most on planning when preferences are weak. The highest RTs are observed when preferences are weak (green line) and conflict is high.

Together, these results converge with the choice model and show that participants relied more strongly on planning-derived values when their task-structured preferences were weak, and this increased reliance was reflected in slower responses under planning conflict. Conversely, strong preferences were associated with weaker effects of planning conflict and faster choices.

Figure 6: Results of reaction time models. (A) Comparison of the model of RT as a function of decision values and intermediate offer against the model of RT as a function of decision values and preferences. (B) Posterior distribution of the fitted parameter. (C) Posterior predictive curve depicting the predicted reaction time as a function of the conflict values, separately for preference entropy of 0, 0.5 and 1.0

4 Discussion

We tested the hypothesis that deviations from optimal forward planning in complex sequential decision making reflect structured preference-based guidance rather than unsystematic errors or simple history- and context-dependent biases. Across choice models, participants’ decisions were best explained by a model that combined planning-derived decision values with task-structured preferences over offers, energy, and costs. Critically, the influence of planning depended on the strength of these preferences: planning-derived values shaped choices more strongly when preferences were uncertain. Similarly, response-time analyses showed stronger effects of planning conflict when preferences were uncertain. Additional policy analyses indicated that these preferences were adaptive in the sense that they preserved reward-relevant task structure and supported above-stochastic returns even without planning. Policies relying solely on preferences achieved substantially higher returns than stochastic choice, although they remained below the optimal policy based on planning. Feature-wise analyses showed that these preferences captured reward-relevant structures of the task. Together, these results suggest that participants used task-structured preferences as efficient approximations to optimal choice, while relying more strongly on forward planning when these preferences did not provide a clear action tendency.

A methodological contribution of the present study is that we combined a planning-derived decision-value term with a factorized preference component defined over the task features. Computational models of sequential decision making often include additional parameters capturing response biases, perseveration, or other deviations from value-based choice (Ott et al. 2022a, 2020; Ho et al. 2022; Huys et al. 2012; Schwartenbeck et al. 2015). However, these terms are typically low-dimensional and are not usually decomposed according to the experimental structure of the task. Conversely, factorial analyses of experimental data can quantify how different task factors influence behaviour or neural responses (Barr et al. 2013; Baayen and Milin 2010; Ashburner et al. 2014; Singmann and Kellen 2019), but they do not by themselves specify how these effects relate to forward planning. By combining these approaches, our model treats task-factor effects not as residual noise around planning, but as structured action tendencies that can guide choices and modulate reliance on planning-derived values.

Our approach was motivated by Bayesian accounts of perception and sensorimotor control, in which current evidence is integrated with prior information according to their relative uncertainty (Knill and Pouget 2004; Lee and Mumford 2003; Rao and Ballard 1999; Clark 2013; Friston 2010). An intuitive interpretation of our results is that task-structured preferences play the role of action priors, whereas planning-derived values provide the likelihood-like, state-specific evidence about which action will lead to better future outcomes, and the strength of the priors determines the extent of forward planning. This interpretation is particularly evident in how preferences guide behavior in different states. In states where the relevant features jointly imply a clear action tendency (e.g. high offer, low cost and high energy), preferences alone may be sufficient to support efficient decision making, requiring minimal forward planning. In other states, different features support opposing actions. For instance, a moderate offer may favour rejection, whereas low cost and high energy may favour acceptance, resulting in a weaker and more uncertain preference signal. In such cases, forward planning can help resolve the remaining ambiguity by evaluating the future consequences of the available actions.

The factorized preference component of our model constitutes an extension of existing Bayesian accounts of decision making (Huys et al. 2012; Kuperwajs and Ma 2021; Schwöbel et al. 2021), most explicitely to the policy compression framework which assumes that choice reflects the bayesian integration of planning derived value with state-independent action priors, reflecting the marginal action probability of the current policy (Lai and Gershman 2021). While such priors work well with small state spaces or tasks with limited downstream consequences (Lai and Gershman 2024; Bari and Gershman 2023; Liu and Gershman 2025; Lai et al. 2022; Gershman and Lak 2025), they may be insufficient in more complex task sequential decision making tasks, where states can require different actions and where action–state combinations have downstream implications. Our results suggest that in such tasks, participants rely on state-dependent preferences. Our factorized preferences strike a balance between fully state-specific priors, which would be overly costly in terms of memory and completely state independent action priors, which would be too coarse grained to be useful in complex tasks. Existing Bayesian models of decision making should be extended to enable priors structure reflecting the relevant dimensions of the task at hands.

This framing also extends on the work of Ott et al. (2022a), who found that participants exert more cognitive control for intermediate (2 and 3) compared to extreme (1 and 4) offers. They interpreted this finding as evidence that participants split the state space into coarse contexts and allocate context-dependent control. Our results offer a more nuanced interpretation: intermediate offers correspond to preferences with higher entropy (see Figure 2), prompting greater planning effort to resolve uncertainty. Rather than splitting the state space into coarse context (Ott et al. 2022a), participants allocate planning efforts dynamically based on the strength of their structured priors. Under this view, cognitive control is invertly proportional to priors certainty and reflects how much they engage in forward planning to resolve choice uncertainty.

Our results contribute to the debate about whether deviations from optimality reflects systematic errors or adaptive heuristics. Across psychology and perceptual decision making, deviations from normative models have often been described as systematic errors (Tversky and Kahneman 1974; Gilovich and Griffin 2002; Rahnev and Denison 2018; Berthet and De Gardelle 2023). In contrast, others argue that biases reflect adaptive heuristics that serve as computation-less shortcuts enabling satisficing solutions in complex tasks (Gigerenzer et al. 2000; Tsetsos et al. 2016; Gigerenzer and Brighton 2009). In line with the adaptive heuristic interpretation, we find that preferences are not arbitrary residuals around an optimal planning process. Instead, they represent structured action tendencies derived from task features, supporting above-stochastic performance in the absence of planning. However, participants do not rely on preferences alone. Rather, biases from optimal decision making reflect action priors which are an inexpensive and quickly available first-pass yielding satisficing returns, which are complemented by planning when priors are uncertain and planning-derived action values are in conflict. This implies that there is no strict separation between a fast heuristic system and a slow planning system and that instead, both work in concert in the context of sequential decision making.

Importantly, while our model suggest that task-structured preferences can function like prior action tendencies, the model is purely descriptive and does not establish that the fitted preferences are priors in a formal Bayesian sense. In addition, our model assumes that the preferences have been established prior to the start of the task and does not specify the learning dynamics or updating of these preferences. The similarity between preferences and factor levels-dependent marginal action probability (see Figure 5 D) suggest that preferences might arise by keeping track of action performed in each level of each factor under an approximately optimal policy, in the training session for example. Furthermore general schemas from encounter with similar problem in the environment (budgetting decision in everyday life), or explicit instructions highlighting relevant features might also play a role in learning of preferences. A key question for future research is to establish whether the preferences truly constitute action priors, as well as determining how agents determine which features are relevant for constructing structured priors, especially in real-world settings where task structure is not explicit.

In addition, while our model uses backward induction to derive state-action values, this does not imply that the brain implements forward planning in this way. Biologically plausible planning likely relies on approximate algorithm such as sampling, partial tree search and approximate value computations (Sutton et al. 1998; Silver et al. 2016; Huys et al. 2012; Schwöbel et al. 2021). We are not committed to any one of these alternatives, as our main interest in the paper was to investigate the structure underlying participants’ deviations from planning. The observed interaction between RT, preferences strength and decision conflict suggests that whichever algorithm implements action value estimation in the brain, the extent of planning-related computation can be adaptively adjusted to obtain more or less precise estimates depending on preferences strength. Developing RL or active inference agents equipped with such structured priors, and testing their predictions against human data will be crucial for understanding the emergence and utility of efficient, adaptive decision strategies in complex environments.

The presented account makes several testable predictions about when preference-based guidance should dominate. Reliance on preferences should increase when task features provide a clear action tendency, when planning time is limited, or when the marginal value of additional computation is low. Conversely, reliance on planning should increase when task features provide conflicting action tendencies, when future consequences are difficult to infer from local features alone, or when the cost of an error is high. Future experiments could test these predictions by manipulating time pressure, training, instructions, and task familiarity, planning horizon, task volatility, or the reliability of feature-based cues to disentangle these sources and examine how structured priors are formed, adapted and integrated with forward planning.

5 Acknowledgements

6 Data and code statement

All data used in this paper were previously published as part of Ott et al. (2022a), and are available at https://zenodo.org/records/6328296. Code can be found on github https://github.com/AlexLepauvre/state_abstraction_paper. In addition, we’ve created a custom toolbox to implement reinforcement learning algorithm, which can be found here: https://github.com/AlexLepauvre/state-abstraction

7 References

Abril-Pla, Oriol, Virgile Andreani, Colin Carroll, et al. 2023. “PyMC: A Modern, and Comprehensive Probabilistic Programming Framework in Python.” PeerJ Computer Science 9: e1516.
Ashburner, John, Gareth Barnes, Chun-Chuan Chen, et al. 2014. “SPM12 Manual.” Wellcome Trust Centre for Neuroimaging, London, UK 2464 (4): 53.
Baayen, R Harald, and Petar Milin. 2010. “Analyzing Reaction Times.” International Journal of Psychological Research 3 (2): 12–28.
Bari, Bilal A, and Samuel J Gershman. 2023. “Undermatching Is a Consequence of Policy Compression.” Journal of Neuroscience 43 (3): 447–57.
Barr, Dale J, Roger Levy, Christoph Scheepers, and Harry J Tily. 2013. “Random Effects Structure for Confirmatory Hypothesis Testing: Keep It Maximal.” Journal of Memory and Language 68 (3): 255–78.
Berthet, Vincent, and Vincent De Gardelle. 2023. “The Heuristics-and-Biases Inventory: An Open-Source Tool to Explore Individual Differences in Rationality.” Frontiers in Psychology 14: 1145246.
Callaway, Frederick, Bas Van Opheusden, Sayan Gul, et al. 2022. “Rational Use of Cognitive Resources in Human Planning.” Nature Human Behaviour 6 (8): 1112–25.
Capretto, Tomás, Camen Piho, Ravin Kumar, Jacob Westfall, Tal Yarkoni, and Osvaldo A Martin. 2022. “Bambi: A Simple Interface for Fitting Bayesian Linear Models in Python.” Journal of Statistical Software 103 (15): 1–29. https://doi.org/10.18637/jss.v103.i15.
Clark, Andy. 2013. “Whatever Next? Predictive Brains, Situated Agents, and the Future of Cognitive Science.” Behavioral and Brain Sciences 36 (3): 181–204.
Friston, Karl. 2010. “The Free-Energy Principle: A Unified Brain Theory?” Nature Reviews Neuroscience 11 (2): 127–38.
Gershman, Samuel J, Eric J Horvitz, and Joshua B Tenenbaum. 2015. “Computational Rationality: A Converging Paradigm for Intelligence in Brains, Minds, and Machines.” Science 349 (6245): 273–78.
Gershman, Samuel J, and Armin Lak. 2025. “Policy Complexity Suppresses Dopamine Responses.” Journal of Neuroscience 45 (9).
Gigerenzer, Gerd, and Henry Brighton. 2009. “Homo Heuristicus: Why Biased Minds Make Better Inferences.” Topics in Cognitive Science 1 (1): 107–43.
Gigerenzer, Gerd, Peter M Todd, the ABC Research Group, et al. 2000. Simple Heuristics That Make Us Smart. Oxford University Press.
Gilovich, Thomas, and Dale W Griffin. 2002. Heuristics and Biases: The Psychology of Intuitive Judgment. Cambridge university press.
Ho, Mark K, David Abel, Carlos G Correa, Michael L Littman, Jonathan D Cohen, and Thomas L Griffiths. 2022. “People Construct Simplified Mental Representations to Plan.” Nature 606 (7912): 129–36.
Huys, Quentin JM, Neir Eshel, Elizabeth O’Nions, Luke Sheridan, Peter Dayan, and Jonathan P Roiser. 2012. “Bonsai Trees in Your Head: How the Pavlovian System Sculpts Goal-Directed Choices by Pruning Decision Trees.” PLoS Computational Biology 8 (3): e1002410.
Knill, David C, and Alexandre Pouget. 2004. “The Bayesian Brain: The Role of Uncertainty in Neural Coding and Computation.” TRENDS in Neurosciences 27 (12): 712–19.
Körding, Konrad P, and Daniel M Wolpert. 2004. “Bayesian Integration in Sensorimotor Learning.” Nature 427 (6971): 244–47.
Kuperwajs, Ionatan, and Wei Ji Ma. 2021. “Planning to Plan: A Bayesian Model for Optimizing the Depth of Decision Tree Search.” Proceedings of the Annual Meeting of the Cognitive Science Society 43.
Lai, Lucy, and Samuel J Gershman. 2021. “Policy Compression: An Information Bottleneck in Action Selection.” In Psychology of Learning and Motivation, vol. 74. Elsevier.
Lai, Lucy, and Samuel J Gershman. 2024. “Human Decision Making Balances Reward Maximization and Policy Compression.” PLOS Computational Biology 20 (4): e1012057.
Lai, Lucy, Ann Zixiang Huang, and Samuel J Gershman. 2022. “Action Chunking as Policy Compression.” PsyArXiv.
Lee, Tai Sing, and David Mumford. 2003. “Hierarchical Bayesian Inference in the Visual Cortex.” Journal of the Optical Society of America A 20 (7): 1434–48.
Liu, Shuze, and Samuel Joseph Gershman. 2025. “Action Subsampling Supports Policy Compression in Large Action Spaces.” PLOS Computational Biology 21 (9): e1013444.
Ott, Florian, Eric Legler, and Stefan J Kiebel. 2022a. “Forward Planning Driven by Context-Dependant Conflict Processing in Anterior Cingulate Cortex.” NeuroImage 256: 119222. https://doi.org/10.1016/j.neuroimage.2022.119222.
Ott, Florian, Eric Legler, and Stefan J Kiebel. 2022b. Forward Planning Driven by Context-Dependent Conflict Processing in Anterior Cingulate Cortex - Analysis Code and Datasets. V. 1.2. Zenodo, released. https://doi.org/10.5281/zenodo.6328296.
Ott, Florian, Dimitrije Marković, Alexander Strobel, and Stefan J Kiebel. 2020. “Dynamic Integration of Forward Planning and Heuristic Preferences During Multiple Goal Pursuit.” PLoS Computational Biology 16 (2): e1007685.
Parush, Naama, Naftali Tishby, and Hagai Bergman. 2011. “Dopaminergic Balance Between Reward Maximization and Policy Complexity.” Frontiers in Systems Neuroscience 5: 22.
Rahnev, Dobromir, and Rachel N Denison. 2018. “Suboptimality in Perceptual Decision Making.” Behavioral and Brain Sciences 41: e223.
Rao, Rajesh PN, and Dana H Ballard. 1999. “Predictive Coding in the Visual Cortex: A Functional Interpretation of Some Extra-Classical Receptive-Field Effects.” Nature Neuroscience 2 (1): 79–87.
Schwartenbeck, Philipp, Thomas HB FitzGerald, Christoph Mathys, Ray Dolan, and Karl Friston. 2015. “The Dopaminergic Midbrain Encodes the Expected Certainty about Desired Outcomes.” Cerebral Cortex 25 (10): 3434–45.
Schwöbel, Sarah, Dimitrije Marković, Michael N Smolka, and Stefan Kiebel. 2024. “Joint Modeling of Choices and Reaction Times Based on Bayesian Contextual Behavioral Control.” PLoS Computational Biology 20 (7): e1012228.
Schwöbel, Sarah, Dimitrije Marković, Michael N Smolka, and Stefan J Kiebel. 2021. “Balancing Control: A Bayesian Interpretation of Habitual and Goal-Directed Behavior.” Journal of Mathematical Psychology 100: 102472.
Silver, David, Aja Huang, Chris J Maddison, et al. 2016. “Mastering the Game of Go with Deep Neural Networks and Tree Search.” Nature 529 (7587): 484–89.
Singmann, Henrik, and David Kellen. 2019. “An Introduction to Mixed Models for Experimental Psychology.” In New Methods in Cognitive Psychology. Routledge.
Sutton, Richard S, Andrew G Barto, et al. 1998. Reinforcement Learning: An Introduction. Vol. 1. MIT press Cambridge.
Tishby, Naftali, and Daniel Polani. 2010. “Information Theory of Decisions and Actions.” In Perception-Action Cycle: Models, Architectures, and Hardware. Springer.
Tsetsos, Konstantinos, Rani Moran, James Moreland, Nick Chater, Marius Usher, and Christopher Summerfield. 2016. “Economic Irrationality Is Optimal During Noisy Decision Making.” Proceedings of the National Academy of Sciences 113 (11): 3102–7.
Tversky, Amos, and Daniel Kahneman. 1974. “Judgment Under Uncertainty: Heuristics and Biases: Biases in Judgments Reveal Some Heuristics of Thinking Under Uncertainty.” Science 185 (4157): 1124–31.
Vehtari, Aki, Andrew Gelman, and Jonah Gabry. 2017. “Practical Bayesian Model Evaluation Using Leave-One-Out Cross-Validation and WAIC.” Statistics and Computing 27 (5): 1413–32.

8 Supplementary material

8.1 Supplementary figures:

Figure S1: Interaction between preferences and DV. (A) Participants responses as a function of decision values of preference score. The response in the off-diagonal quadrants (top left and bottom right) are color coded per response given, to highlight that participants responses reflect the strongest signal (tend to accept despite DV being negative if preferences are high enough). (B) Fitted probability of accepting as a function of DV and preference score. Contours with annotations represent a probability of accepting. The shape of the contour lines highlight that participants choices reflect the modulation of DV by preference score. The contour trace indicate that when preference and DV are opposed, participants follow the strongest signal.

8.2 Supplementary tables:

Table S1
Main model comparison, comparing the fit of the preference, context, frequency and pure planning models
elpd_loo p_loo elpd_diff se dse
Preference -1492.49 256.526 0 48.7651 0
Context -1656.69 159.175 164.209 52.4615 19.671
Frequency prior -2073.38 77.0896 580.889 55.5712 35.9363
Planning -2078.46 68.3847 585.976 55.58 36.1456
Table S2
Summary of the parameters of the preference model. The table displays the mean and 95% highest density interval (HDI) for the population level parameters assopciated with the planning values, the preference regressors and their interaction.
mean sd hdi_2.5% hdi_97.5%
\(\beta_{plan}\) 2.072 0.242 1.614 2.573
\(\beta_{O=1}\) -2.374 0.971 -4.232 -0.426
\(\beta_{O=2}\) -1.625 0.938 -3.535 0.092
\(\beta_{O=3}\) 1.627 0.945 -0.22 3.453
\(\beta_{O=4}\) 2.923 0.976 1.054 4.842
\(\beta_{CC=1}\) 0.765 1.161 -1.488 3.105
\(\beta_{CC=2}\) -0.127 1.165 -2.482 2.144
\(\beta_{FC=1}\) 0.457 1.149 -1.86 2.557
\(\beta_{FC=1}\) 0.21 1.149 -2.051 2.359
\(\beta_{E=0}\) -4.438 0.935 -6.247 -2.627
\(\beta_{E=1}\) -0.48 0.781 -1.994 1.076
\(\beta_{E=2}\) -0.02 0.75 -1.473 1.492
\(\beta_{E=3}\) 0.203 0.748 -1.228 1.688
\(\beta_{E=4}\) 0.3 0.758 -1.186 1.799
\(\beta_{E=5}\) 0.793 0.769 -0.602 2.44
\(\beta_{E=6}\) 4.195 0.962 2.363 6.132
\(\beta_{interaction}\) 1.306 0.508 0.258 2.256

8.3 Supplementary model comparison

8.3.1 Preferences model comparison:

The preference model was observed to fit the data best, showing that participants behaviour reflects the integration of a planning component together with priors reflecting the structure of the task, shedding light on the structure of the priors themselves. To further show that participants priors reflect the whole structure of the task, we modelled the data removing one experimental factor from the preferences fitting one at a time and compared to the full model, to show that our results cannot be explained by the fact that participants’s priors do not reflect a single nor a subset of experimental factors. To further show that participants reliance on planning is modulated by the strength of their preferenecs, a model without the interaction was fitted to highlight the increase in fit associated with the interaction. Finally, to show that participants response is best explained by a combination of planning and preference component, we also fitted a model with only the preference component omitting the planning regressor.

Figure S2: Results of preferences model comparison. Model abbreviations: (No_interaction) model comprising planning and preference component, without interaction between the two, (No_FC) model where future cost factor was omitted from preferences, (No_CC) model where future cost factor was omitted from preferences, (No_Planning) model where DV was omitted, (No_Energy) model where energy factor was omitted from preferences, (No_Offer) model where offer factor was omitted from preferences
Table S3
Results of the comparison between different preferences models
elpd_loo p_loo elpd_diff se dse
full_preferences_model -1491.09 253.49 0 48.6156 0
No_Interaction -1501.44 245.861 10.3472 48.9103 5.41809
No_FC -1517.79 227.725 26.7017 48.5354 7.5994
No_CC -1539.49 262.803 48.3996 49.8892 9.67725
No_Planning -1581.23 217.308 90.1381 48.8039 17.2262
No_Energy -1626.44 211.155 135.351 51.5682 17.8535
No_Offer -1629.84 235.203 138.745 54.4382 19.333

8.3.2 Frequency model comparison:

In the main text, we compared the preference model with a model in which priors are computed based on action frequency. In this model, the priors are computed using a linearly decaying moving average, discounting actions that occured further in the past. We presented the results obtained when the priors are derived from the 20 previous trials. This number of trials was selected by exploring priors witk \(k=[2, 5, 10, 15, 20, 30, 50]\) to identify the values that fit the data the best:

Figure S3: Results of frequency model comparison, comparing the temporal integration window, going back k=[2, 5, 10, 15, 20, 30, 50]
Table S4
Results of the frequency model comparisons with different temporal integration windows
elpd_loo p_loo elpd_diff se dse
Linear, k=20 -2069.6 78.3873 0 55.6858 0
Linear, k=15 -2073.74 74.2703 4.13799 55.8878 5.24083
Linear, k=2 -2074.98 75.5677 5.37448 55.6833 6.29665
Linear, k=5 -2076.3 78.84 6.69933 55.6484 5.58139
Linear, k=10 -2077.34 74.0643 7.73213 55.5341 5.65329
Linear, k=50 -2080.71 78.7902 11.1052 55.7161 5.09526
Linear, k=30 -2081.1 75.786 11.4922 55.7602 5.46973

8.4 Parameters recovery:

In the main text, the preference model was observed to fit participants behaviour the best, supporting the hypothesis that participants decision reflects the integration of forward planning with state dependent preferences. The preferences model is rather complex, with preference parameters being estimated as latent variable, and entailing an additional entropy transformation in the interaction term. To establish that despite the complexity of the model, it is still capable of accurately estimate all parameters, we perform a parameter recovery analysis in which we simulate data using the mean parameters recovered from the original data. We then fit the same model to the simulated data. We can then compare the estimated parameters from the simulated data against ground truth to establish the model capacity to identify parameters correctly.

Figure S4: Population level parameters recovery

Across all participants, the ground truth parameters is well within the credible intervals of the retrieved parameters from the data, indicating that our model is indeed capable of recovering the parameters of the underlying generative model.