Robust Bayes Treatment Choice with Partial Identification

07-2026
Abstract

We study a class of binary treatment choice problems with partial identification through the lens of robust (multiple prior) Bayesian analysis. We use a convenient set of prior distributions to derive ex-ante and ex-post robust Bayes decision rules, both for decision makers who can randomize and for decision makers who cannot.
Our main messages are as follows: First, ex-ante and ex-post robust Bayes decision rules do not agree in general, whether or not randomized rules are allowed. Second, randomized treatment assignment for some data realizations can be optimal in both ex-ante and, perhaps more surprisingly, ex-post problems. Therefore, it is usually with loss of generality to exclude randomized rules from consideration, even when regret is evaluated ex post.
We apply our results to a stylized problem where a policy maker uses experimental data to choose whether to implement a new policy in a population of interest, but is concerned about the external validity of the experiment at hand, and to aggregation of data generated by multiple randomized control trials in different sites to make a policy choice in a population for which no experimental data are available.

Keywords: treatment choice, partial identification, robust Bayes, Gamma-minimax, posterior robustness

1 Introduction

A policy maker must decide between implementing a new policy or preserving the status quo. Her data provide information about the potential benefits of these two options. Unfortunately, these data only partially identify payoff-relevant parameters and may therefore not reveal, even in large samples, the correct course of action. Such treatment choice problems with partial identification have recently received growing interest; for example, see D’Adamo (2021), Ishihara and Kitagawa (2021), Yata (2021), Christensen, Moon, and Schorfheide (2022), Kido (2022) or Manski (2024). Several interesting problems that arise in empirical research can be recast using this framework. See Montiel Olea, Qiu, and Stoye (2025) and the end of this section for references.

This paper applies the robust Bayes approach—which interpolates between Bayesian and agnostic minimax analyses by evaluating minimax risk over a set of priors—to a class of treatment choice problems with partial identification. The use of the robust Bayes approach has drawn recent attention in problems that feature partial identification Giacomini and Kitagawa (2021); Giacomini, Kitagawa, and Read (2021); Christensen, Moon, and Schorfheide (2022). Indeed, due to partial identification, integrating Bayesian and minimax elements into decision making can be particularly attractive (see Poirier (1998); Moon and Schorfheide (2012) and references therein).

The robust Bayes approach can be applied ex-ante or ex-post, depending on whether the multiple priors are used to evaluate payoffs before or after seeing the data.1 These concepts represent two different ways of resolving model ambiguity and sampling uncertainty, and both have been proposed to improve Bayesian robustness in decision problems.

By the well-known dynamic consistency of Bayesian decision making, ex-ante and ex-post robust Bayes coincide with each other—and with standard Bayes optimality—if the set of priors is a singleton. Indeed, this equivalence is used to calculate Bayes optimal decisions in practice because ex-post rules are usually computed but ex-ante Bayes optimality is claimed. As pointed out for the present context by Giacomini, Kitagawa, and Read (2021), they do not in general agree otherwise.2 However, to what extent this inequivalence affects treatment choice problems with partial identification is far from clear. The class of priors that we consider has a Cartesian product structure resembling rectangularity, a condition under which maximin welfare loss is known to be dynamically consistent.3 While this result does not apply here—the priors are not rectangular in the strict technical sense, and the theoretical results were not established for regret loss—one might wonder if the conclusion holds anyway or else, what qualitative and quantitative relationships exist between ex-ante and ex-post robust Bayes in the presence of partial identification. Using a convenient class of priors allows us to exhaustively answer these questions in a case that we believe holds some interest.

To this end, we use the same framework as Yata (2021) and Montiel Olea, Qiu, and Stoye (2025) but impose a simple instance of Giacomini and Kitagawa (2021)’s set of priors, namely a symmetric and uniform two-point prior for reduced-form parameters and no restriction at all on unidentified parameters given reduced-form parameters. Working with regret, we formally define ex-ante and ex-post robust Bayes and, following Berger (1985), label them as “\(\Gamma\)-minimax regret” (\(\Gamma\)-MMR) and “\(\Gamma\)-posterior expected regret” (\(\Gamma\)-PER).4 We then precisely characterize when these notions coincide and when they disagree. The main qualitative insights are as follows:

Our point is not to advocate for either notion of robust Bayes criterion. We also do not aim to solve for robust Bayes criteria for more general sets of priors as this would get much more involved but (we suspect) not much more instructive to illustrate the points discussed above. What we hope to illustrate is when, and how, \(\Gamma\)-MMR and \(\Gamma\)-PER criteria differ. We also relate these results to timing assumptions in a fictitious game between the policy maker and an adversarial Nature.

Several auxiliary findings might be of independent interest. First, for \(\Gamma\)-MMR, even if we restrict the set of decision rules to be a class of non-randomized threshold rules based on the “efficient” linear index, the optimal threshold is not always zero. Given the apparent symmetry of the problem, we find this feature rather surprising. Second, whenever the dimension of the signal is larger than \(1\), there always exist (regardless of the parameter space and the variance of the signals) non-randomized linear-index threshold rules (with a threshold equal to zero) that are \(\Gamma\)-MMR optimal (among all decision rules). This is in stark contrast to Montiel Olea, Qiu, and Stoye (2025), in which no linear index rule is globally minimax regret optimal if the degree of partial identification is severe. The intuition is that the prior much reduces the state space; the signal space then becomes so rich relative to the state space that even linear threshold rules can effectively mimic randomization.

The literature on treatment choice with partially identified parameters has been growing since Manski (2004) and Dehejia (2005). For partial identification with known distribution of data, Manski (2000)b, (2005), (2007)a and Stoye (2007) find minimax regret optimal treatment rules. For finite-sample minimax regret results with model ambiguity and sampling uncertainty, see Stoye (2012)a, (2012)b, Yata (2021), Ishihara and Kitagawa (2021) and Montiel Olea, Qiu, and Stoye (2025); Kido (2023)’s analysis is asymptotic. Bayes and robust Bayes approaches are analyzed by Chamberlain (2011), Giacomini and Kitagawa (2021), Giacomini, Kitagawa, and Read (2021), Christensen, Moon, and Schorfheide (2022), among others. Earlier investigations of ex-ante and ex-post \(\Gamma\)-minimax estimators include DasGupta and Studden (1989) and Betrò and Ruggeri (1992). See Vidakovic (2000) for a review. For different settings with point-identified welfare, finite- and large-sample results on optimal treatment choice rules were derived by Canner (1970), Chen and Guggenberger (2025), Hirano and Porter (2009), (2020), Kitagawa, Lee, and Qiu (2022), Schlag (2006), Stoye (2009), and Tetenov (2012)b. There is also a large literature on optimal policy learning with covariates containing results with point identified Bhattacharya and Dupas (2012); Kitagawa and Tetenov (2018), (2021); Mbakop and Tabord-Meehan (2021); Kitagawa and Wang (2023); Athey and Wager (2021); Kitagawa, Sakaguchi, and Tetenov (2021); Ida et al. (2025) as well as partially identified Kallus and Zhou (2018); Ben-Michael et al. (2021); Ben-Michael, Imai, and Jiang (2022); D’Adamo (2021); Christensen, Moon, and Schorfheide (2022); Adjaho and Christensen (2022); Kido (2022); Lei, Sahoo, and Wager (2023) parameters. Guggenberger, Mehta, and Pavlov (2024) and Manski and Tetenov (2023) analyze related problems but focus on quantile, as opposed to expected, loss; Song (2014) considers partial identification but mean squared error regret loss.

The rest of this paper is organized as follows. Section 2 sets up the problem, provides examples, and defines both versions of robust Bayes optimality. Section 3 contains complete solutions for all aforementioned scenarios and relates them to timing assumptions in the “Games against Nature” interpretation of minimax theory. Section 4 concludes. Proofs and auxiliary results are collected in the Appendix.

2 Framework

2.1 Actions, Payoffs, Statistical Model, and Decisions

Our setup follows Montiel Olea, Qiu, and Stoye (2025), who in turn follow Ferguson (1967) and others. Consider a policy maker who needs to choose an action \(a\in[0,1]\) interpreted as probability of assigning treatment in the target population.5

Her payoff when taking action \(a\in[0,1]\) is captured by the welfare function

\[ W(a,\theta):=aW(1,\theta)+(1-a)W(0,\theta), \tag{1}\]

where \(\theta\in\Theta\) is an unknown state of the world or parameter and the functions \(W(1,\cdotp):\Theta\rightarrow\mathbb{R}\) and \(W(0,\cdotp):\Theta\rightarrow\mathbb{R}\) are known. Here, we may interpret \(W(1,\theta)\) and \(W(0,\theta)\) as the welfare of actions \(a=1\) (treating everyone in the population) and action \(a=0\) (treating no one in the population). Therefore, Equation 1 implies that welfare is linear in actions, a standard assumption in the literature. Denote by \(U(\theta):=W(1,\theta)-W(0,\theta)\) the welfare contrast at \(\theta\). If \(U(\theta)\) were known to the policy maker, her optimal action would simply be

\[ \mathbf{1}\left\{ U(\theta)\geq0\right\}. \tag{2}\]

The policy maker does not know \(U(\theta)\) but can learn about \(\theta\). Specifically, we assume that she observes a random vector \(Y\in\mathbb{R}^{n}\) with multivariate normal distribution

\[ Y \sim N(m(\theta),\Sigma), \tag{3}\]

where the function \(m(\cdotp):\Theta\rightarrow\mathbb{R}^{n}\) and the positive definite matrix \(\Sigma\) are known.

Our focus is on the case when the data are not entirely informative about the sign of \(U(\theta)\): Even if the policy maker perfectly learned \(m(\theta)\), she could not (necessarily) pin down the sign of \(U(\theta)\). To formally model such treatment choice problems with (decision-relevant) partial identification, let

\[ M:=\left\{ \mu\in\mathbb{R}^{n}:m(\theta)=\mu,\theta\in\Theta\right\} \tag{4}\]

collect all the means of \(Y\) that can be generated as \(\theta\) ranges over \(\Theta\). We refer to elements \(\mu\in M\) as reduced-form parameters because they are identified in the statistical model Equation 3 without further assumptions. Define the identified set for the welfare contrast given \(\mu\) as

\[ I(\mu):=\left\{ u\in\mathbb{R}:U(\theta)=u,m(\theta)=\mu,\theta\in\Theta\right\} \tag{5}\]

and the corresponding upper and lower bounds as

\[ \overline{I}(\mu):=\sup I(\mu),\quad\underline{I}(\mu):=\inf I(\mu). \tag{6}\]

Henceforth, when we refer to a treatment choice problem with partial identification, we mean there exists some nonempty open set in \(M\) such that for all \(\mu\) in that open set, \(\underline{I}(\mu) <0<\overline{I}(\mu)\). For simplicity, we also assume that the infimum and supremum in Equation 5 are attained.

A decision rule \(d:\mathbb{R}^{n}\rightarrow[0,1]\) is a (measurable) mapping from data \(Y\) to the unit interval \([0,1]\). We call \(d\) non-randomized if it (almost surely, a.s.) maps into \(\{0,1\}\); otherwise, we call \(d\) randomized, including if it randomizes for some but not all data realizations. We use \(\mathcal{D}_n\) to denote the set of all decision rules, and we consider decision rules the same if they a.s. agree. As a result, a rule is unique only up to a.s. agreement. The oracle policy \(\mathbf{1}\{U(\theta) \geq 0\}\) is of special interest and for any given \(\theta\) is contained in \(\mathcal{D}_n\), but is not feasible in the statistical sense because \(U(\theta)\) is not known.

In general, there will not be an unambiguously best feasible decision rule, a problem that gave rise to a large literature on different optimality criteria and their implementation. Before introducing the robust Bayes approach, we give two examples that fit into our general framework.

2.2 Examples

Example 1 (Stoye (2012)a) This is the one-dimensional version of the general setup and has been frequently analyzed before Manski (2000)a; Brock (2006); Stoye (2012)a; Tetenov (2012)a; Kitagawa, Lee, and Qiu (2023). One motivation for it is to think of a policy maker who uses experimental data to choose whether to implement a new policy in a population of interest, but is concerned about the external validity of the experiment at hand. The treatment effect of action \(a=1\) is \(\mu^*\in\mathbb{R}\), while the effect of action \(a=0\) is normalized to \(0\); thus, the policy maker’s expected payoff equals \(W(a,\mu^*):= a \cdot \mu^*\). The policy maker observes a realization of the one-dimensional statistic

\[ \hat{\mu} \sim N(\mu,\sigma^2), \tag{7}\]

where \(\sigma>0\) is known and where \(\mu\in\mathbb{R}\) is an identifiable treatment effect, i.e. it could be perfectly learned from infinite data. In this example, \(\theta=(\mu,\mu^*)^{\top}\), \(\Theta\subseteq \mathbb{R}^2\), \(m(\theta)=\mu\), and \(U(\theta)=\mu^*\).

Since the target population and the population from which data Equation 7 is collected can be different, partial identification naturally arises. We assume that the identifiable and true treatment effects are constrained by \(\left\vert \mu^*-\mu \right\vert \leq k\) for some known \(k \geq 0\), implying

\[ I(\mu)=[\mu-k,\mu+k],\quad \overline{I}(\mu)=\mu+k,\quad \underline{I}(\mu)=\mu-k,\quad\forall \mu \in \mathbb{R}. \]

The planner must choose a statistical decision rule \(d\in\mathcal{D}_1:\mathbb{R}\rightarrow[0,1]\) that maps observed data \(\hat{\mu}\) to an action \(a\in[0,1]\). Taking Equation 7 as an approximation, this stylized example could reflect model uncertainty (e.g., a treatment effect is estimated in a possibly somewhat misspecified model), external validity concerns (e.g., a randomized clinical trial was performed on volunteers), or a shift in the environment (e.g., we transfer estimates from study populations to treatment populations with slightly different covariates or are concerned about distributional drift over time). □

Example 2 (Ishihara and Kitagawa (2021)) This example is taken from Ishihara and Kitagawa (2021)’s (see also Manski (2020)) “evidence aggregation” framework. A policy maker is interested in implementing a new policy in country \(i=0\) and observes estimates of the policy’s effect for countries \(i=1,...,n\). Let \(Y=(Y_1,...,Y_n)^\top \in \mathbb{R}^n\) denote these estimates and let \((x_0,\ldots,x_n)\) be nonrandom, \(d\)-dimensional baseline covariates. The policy maker is willing to extrapolate from her data by assuming that the welfare contrast of interest equals \(U(\theta)=\theta(x_0)\) and that

\[ Y = \begin{pmatrix} Y_1 \\ \vdots \\ Y_n \end{pmatrix} \sim N(m(\theta), \Sigma), \quad m(\theta) = \begin{pmatrix} \theta(x_1) \\ \vdots \\ \theta(x_n) \end{pmatrix}, \quad \Sigma = \operatorname{diag}(\sigma_1^2,\ldots,\sigma_n^2), \]

where \(\theta: \mathbb{R}^d \rightarrow \mathbb{R}\) is an unknown Lipschitz function with known constant \(C\). For notational simplicity, we can further write \(\mu_i\) for \(\theta(x_i)\). Thus, \(Y_i \sim N(\mu_i,\sigma_i^2)\), \(Y\sim N(\mu,\Sigma)\) and \(U(\theta)=\mu_0\). The policy question is: Given data \(Y\), what proportion of the population in country \(i=0\) should be assigned the new policy?

Let \(\left\Vert \beta\right\Vert:=\sqrt{\beta^{\top}\beta}\) be the Euclidean norm of a vector \(\beta\). In this example, the identified set for the welfare contrast \(\mu_0\) is

\[ I(\mu) = \{ u \in \mathbb{R} : \left\vert \mu_i-u \right\vert \leq C\left\Vert x_i-x_0\right\Vert, ~~ i=1,\ldots,n \}. \]

The lower and upper bounds on \(I(\mu)\) are simply intersection bounds:

\[ \underline{I}(\mu) = \max_{i=1,\ldots,n} \{ \mu_i - C\left\Vert x_i-x_0 \right\Vert \}, \quad \overline{I}(\mu) = \min_{i=1,\ldots,n} \left \{ \mu_i + C \left\Vert x_i-x_0 \right\Vert \right \}. \]

2.3 Robust Bayes Optimality

Our setting up to here is as in Montiel Olea, Qiu, and Stoye (2025). We now connect it to the robust Bayes literature by imposing a set of priors \(\Gamma\) on \(\theta\). Following Giacomini and Kitagawa (2021), we choose a particular single proper prior \(\pi_{\mu}\) for \(\mu\in M\) but leave the conditional prior of \(\theta\) given \(\mu\), denoted as \(\pi_{\theta \mid \mu}\), unrestricted except for

\[ \pi_{\theta \mid \mu}\{U(\theta)\in I(\mu)\}=1,\quad \pi_{\mu}\text{-a.s.} \tag{8}\]

Then, the class of priors \(\Gamma\) consists of all priors on \(\theta\) induced by the single prior \(\pi_{\mu}\) and any conditional prior \(\pi_{\theta \mid \mu}\) that meets Equation 8. Intuitively, we choose a single prior on the point-identified parameter and place no new restriction on the partially identified parameter \(U(\theta)\). One can pick any proper prior \(\pi_{\mu}\); for tractability, we let \(\pi_{\mu}\) be supported on two symmetric points \(\{\bar{\mu},-\bar{\mu}\}\) with equal probability, where \(\bar{\mu}\in\mathbb{R}^n\) is chosen by the decision maker. Henceforth, \(\Gamma\) is understood to refer to the implied set of priors:

\[ \Gamma:=\left\{\pi_\theta=\int \pi_{\theta\mid \mu}d\pi_{\mu}: \pi_{\mu} \sim \text{unif}(\{-\bar{\mu},\bar{\mu}\}),\pi_{\theta \mid \mu}\text{ satisfies the condition above}\right\}. \tag{9}\]

While the set of priors \(\Gamma\) broadly puts us into the “robust Bayes” territory, it still does not pin down a uniquely best decision rule because the sign of \(U(\theta)\) can remain ambiguous. Let

\[ L(a,\theta):=\sup_{a^{\prime}\in[0,1]} W(a^{\prime},\theta) - W(a,\theta)=U(\theta)\left\{ \mathbf{1}\{U(\theta)\geq0\}-a\right\} \]

be the regret of action \(a\in[0,1]\). We evaluate decision rules \(d\) by their expected regret, defined as

\[ R(d,\theta) := \mathbb{E}_{m(\theta)}[L(d(Y),\theta)] = U(\theta)\left\{ \mathbf{1}\{U(\theta)\geq0\}-\mathbb{E}_{m(\theta)}[d(Y)]\right\}, \tag{10}\]

where for any \(x\in\mathbb{R}^n\), \(\mathbb{E}_{x}[\cdot]\) denotes expectation taken over \(Y\sim N(x,\Sigma)\).

Even with the set of priors \(\Gamma\) given and commitment to expected regret, the robust Bayes literature contains multiple decision criteria that do not in general agree. The difference lies in when (and how) the expectations regarding the unknown parameter \(\theta\in\Theta\) are taken. For decision rule \(d\in \mathcal{D}_n\), let

\[ r(d,\pi):=\int_{\theta\in\Theta} R(d,\theta)d\pi(\theta) \]

be its Bayes expected regret under a prior \(\pi\in\Gamma\). Following Berger (1985), Definition 12, p. 216, we introduce the first robust Bayes optimality notion.

Definition 1 (Ex-ante \(\Gamma\)-Minimax Regret) A decision rule \(d^* \in \mathcal{D}_n\) is \(\Gamma\)-minimax regret (henceforth \(\Gamma\)-MMR) optimal if

\[ \sup_{\pi \in \Gamma}r(d^*,\pi)= \inf_{d \in \mathcal{D}_n} \sup_{\pi \in \Gamma} r(d,\pi). \]

Definition 2 (Ex-post \(\Gamma\)-Posterior Expected Regret) An alternative “posterior” robustness notion is also common in Bayesian analysis. For each action \(a\in [0,1]\), define posterior expected regret under prior \(\pi\) Berger (1985), Definition 8, p. 159 as

\[ \rho(a,\pi_{\theta\mid Y}):=\int_{\tilde{\theta}\in\Theta}L(a,\tilde{\theta})d\pi_{\theta \mid Y}(\tilde{\theta}), \]

where \(\pi_{\theta\mid Y}\) is the posterior distribution of \(\theta\) given prior \(\pi\) and data \(Y\).6 Then we have the following, alternative optimality criterion Berger (1985), Definition 10, p. 205:

A decision rule \(d^* \in \mathcal{D}_n\) is \(\Gamma\)-posterior expected regret (henceforth \(\Gamma\)-PER) optimal if

\[ \sup_{\pi \in \Gamma} \rho(d^*(Y), \pi_{\theta \mid Y}) = \inf_{a \in [0,1]} \sup_{\pi \in \Gamma} \rho(a, \pi_{\theta \mid Y}), \quad \forall Y \in \mathbb{R}^n. \]

If the decision maker’s action space is restricted to \(\{0,1\}\), i.e., randomization is not allowed, the above definitions are revised by replacing \([0,1]\) with \(\{0,1\}\).

The labeling of \(\Gamma\)-MMR as “ex-ante” versus \(\Gamma\)-PER as “ex-post” can be related to the timing of a fictitious game against an adversarial Nature; see Section Section 3.3 for additional discussion. If \(\Gamma\) were a singleton, the criteria would coincide and would also agree with (single-prior) Bayes optimality. They do not in general agree otherwise. The term “Gamma minimax” usually (and even “robust Bayes” more often than not) refers to \(\Gamma\)-MMR; for example, see Berger (1985), sec. 4.7.6.7

While it is not our agenda to advocate for either criterion, some possible considerations are as follows. The ex-ante approach may be perceived as more suitable if the planner has commitment power and has also been justified axiomatically Hayashi (2008); Stoye (2011). We also find some numerical and theoretical evidence that the \(\Gamma\)-MMR rule may have desirable frequentist properties, e.g. in states of the world off the prior’s support; see discussions at the end of Section 3.2. Regarding computational feasibility, the ex-post approach is often easier due to the applicability of backward induction and is routinely employed to quantify posterior robustness of statistical decisions. On the other hand, recent work on numerical discovery of minimax rules Aradillas Fernández et al. (2025); Guggenberger and Huang (2025) may render the ex-ante approach more scalable. See additional discussions in Section 5.5.

The two-point structure of \(\pi_\mu\) in Equation 9 appears in several related treatment choice problems with minimax regret optimality criteria. For example, the least favorable prior takes such a form in a completely unconstrained minimax regret problem with point identified Stoye (2009) or partially identified Stoye (2012)a; Yata (2021); Montiel Olea, Qiu, and Stoye (2025) welfare contrast. In analogously constrained minimax regret problems in which we restrict \(\mu\in[-\left|\bar{\mu}\right|,\left|\bar{\mu}\right|]\), extending the analyses in the preceding literature would also imply a similar two-point symmetric structure for the least favorable prior. Given these precedents, we think our choice of \(\pi_\mu\) represents a general feature of this class of minimax regret decisions, in addition to offering computational tractability.

3 Main Results

In this section, we solve for the two versions of robust Bayes optimality under the set of priors Equation 9. Following Yata (2021), we assume:

  1. \(\Theta\) is convex, centrosymmetric (i.e., \(\theta\in\Theta\) implies \(-\theta\in\Theta\)) and nonempty.
  2. \(m(\cdot)\) and \(U(\cdot)\) are linear.
Assumption 1

These conditions are restrictive but encompass many examples of empirical relevance; see this paper’s introduction, Montiel Olea, Qiu, and Stoye (2025), and Yata (2021). Exploiting symmetry of the setting, we also set

\[ \overline{I}(\bar{\mu})+\underline{I}(\bar{\mu})>0. \]

This is a normalization because, by Lemma Lemma 9, Assumption 1 implies \(\overline{I}(-\bar{\mu})+\underline{I}(-\bar{\mu})=-(\underline{I}(\bar{\mu})+\overline{I}(\bar{\mu}))\) and we could always replace \(m(\cdot)\mapsto-m(\cdot)\); furthermore, the case of \(\overline{I}(\bar{\mu})+\underline{I}(\bar{\mu})=0\) gives rise to trivial solutions.8

We say a rule is a linear-index threshold rule if it has the form \(\mathbf{1}\left\{ \beta^{\top}Y\geq c\right\}\) for some \(\beta\in\mathbb{R}^{n}\) and \(c\in\mathbb{R}\). Linear-index threshold rules are nonrandomized. They are of particular interest because they form a complete class when \(U(\theta)\) is point-identified Karlin and Rubin (1956) and have received particular attention in the recent literature Ishihara and Kitagawa (2021); Montiel Olea, Qiu, and Stoye (2025). For both \(\Gamma\)-MMR and \(\Gamma\)-PER, we will clarify when linear-index threshold rules are optimal and when they are not. For reasons that will become obvious, the following linear-index threshold rule is of particular interest:

\[ d_{w,0}^{*}:=d_{w,0}^{*}(Y):=\mathbf{1}\left\{ w^{\top}Y\geq0\right\},\text{ }w:=\Sigma^{-1}\bar{\mu}. \]

For each vector \(\beta\in\mathbb{R}^n\), let \(\left\Vert \beta\right\Vert_{\Sigma}:=\sqrt{\beta^{\top}\Sigma\beta}\). Thus, \(\lVert w \rVert_{\Sigma} =\sqrt{w^{\top}\Sigma w}=\sqrt{\bar{\mu}^{\top}\Sigma^{-1}\bar{\mu}}\). Denote by \(\Phi(\cdot)\) the standard normal c.d.f. and by \(\Phi^{-1}(\cdot)\) its inverse, i.e., the corresponding quantile function.

3.1 Ex-ante Robust Bayes Optimality

Theorem 1 (\(\Gamma\)-MMR optimal decisions) Consider a treatment choice problem with welfare function Equation 1, statistical model Equation 3, and set of priors Equation 9, that satisfy Assumption 1.

  1. If

\[ \frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\geq\Phi(\lVert w \rVert_{\Sigma} ), \tag{11}\]

then \(d_{w,0}^{*}\) is uniquely \(\Gamma\)-MMR optimal.

  1. If

\[ \frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}<\Phi(\lVert w \rVert_{\Sigma} ), \tag{12}\]

then a rule \(d\in\mathcal{D}_{n}\) attains \(\Gamma\)-MMR if, and only if, it implies:

\[ \mathbb{E}_{-\bar{\mu}}[d^{*}(Y)] = \frac{-\underline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}, \tag{13}\]

\[ \mathbb{E}_{\bar{\mu}}[d^{*}(Y)] = \frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}. \tag{14}\]

In particular, any rule of the form \(\mathbf{1}\left\{ w^{\top}Y\geq c\right\}\) for some \(c\in\mathbb{R}\) is not \(\Gamma\)-MMR optimal. The following rules, and any convex combination of them, are \(\Gamma\)-MMR optimal:

\[ d^*_\text{RT}:=\Phi\left(\frac{w^{\top}Y}{\tilde{\sigma}}\right), \]

\[ d^*_\text{linear}:= \begin{cases} 0, & w^{\top}Y<-\rho^{*},\\ \frac{w^{\top}Y+\rho^{*}}{2\rho^{*}}, & -\rho^{*}\leq w^{\top}Y\leq\rho^{*},\\ 1, & w^{\top}Y>\rho^{*}, \end{cases} \]

\[ d^*_\text{step}:= \begin{cases} \frac{1}{2} - \beta^*, & w^{\top}Y<0,\\ \frac{1}{2} + \beta^*, & w^{\top}Y\geq0, \end{cases} \]

where

\[ \tilde{\sigma}=\sqrt{\left[\frac{\lVert w \rVert_{\Sigma} ^{2}}{\Phi^{-1}\left(\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\right)}\right]^{2}-\lVert w \rVert_{\Sigma} ^{2}}, \]

\(\rho^{*}>0\) is unique and solves:

\[ \int_{0}^{1}\Phi\left(\frac{2\rho^{*}x-\rho^{*}-\lVert w \rVert_{\Sigma} ^{2}}{\lVert w \rVert_{\Sigma} }\right)dx=\frac{-\underline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}, \]

and

\[ \beta^{*}=\frac{\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}-\frac{1}{2}}{2\Phi(\lVert w \rVert_{\Sigma} )-1}\in\left(0,\frac{1}{2}\right). \]

  1. In case Equation 12, among linear-index threshold rules of the form \(\mathbf{1}\left\{ w^{\top}Y\geq c\right\}\), the optimal thresholds are \(\pm c^{*}\), where

\[ c^{*}:=\lVert w \rVert_{\Sigma} ^2-\lVert w \rVert_{\Sigma} \Phi^{-1}\left(\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\right)>0. \]

  1. In case Equation 12 and if \(n>1\), the following linear-index threshold rule is also \(\Gamma\)-MMR optimal:

\[ d_{w_{t^*},0}^{*}(Y):=\mathbf{1}\left\{ w_{t^*}^{\top}Y\geq0\right\}, \]

where \(w_{t^*}=\Sigma^{-1}\left(t^*\bar{\mu}+(1-t^*)\dot{\mu}\right)\), \(\dot{\mu}\neq0\) is such that \(\dot{\mu}^{\top}\Sigma^{-1}\bar{\mu}=0\),

\[ t^*:=\frac{1}{1\pm\sqrt{\frac{(1-s^{*})}{s^{*}}\frac{\lVert w \rVert_{\Sigma} ^{2}}{\lVert \Sigma^{-1} \dot{\mu}\rVert_{\Sigma} ^{2}}}}, \tag{15}\]

and

\[ s^{*}:=\frac{\left[\Phi^{-1}\left(\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\right)\right]^{2}}{\lVert w \rVert_{\Sigma} ^{2}}\in(0,1). \]

Thus, \(d_{w_{t^*},0}^{*}\) is also optimal among all linear threshold rules.

Theorem 1 reveals that the \(\Gamma\)-MMR rules qualitatively change depending on whether condition Equation 11 is met or not. This condition admits an intuitive interpretation: Up to clamping to the unit interval (i.e., values outside this interval are mapped to its edges), the left-hand side of Equation 11 equals the unique minimax regret optimal rule for known \(\mu\) Manski (2007)b; therefore, it arguably measures the model’s identification strength. The right-hand side of Equation 11 can be interpreted as the informativeness of the data about which of \(\{-\bar{\mu},\bar{\mu}\}\) obtained. Therefore, Theorem 1 says that if the model’s identification power is sufficiently large compared to the informativeness of the data, then, the unique \(\Gamma\)-MMR optimal rule is \(d_{w,0}^{*}\), a non-randomized linear index rule with threshold \(0\) that effectively ignores the partial identification issue. In contrast, if the model’s identification power is small compared to the informativeness of data, there are infinitely many \(\Gamma\)-MMR optimal rules, all of which satisfy Equation 13 and Equation 14. Examples include suitably smoothed versions of \(d_{w,0}^{*}\) like \(d_{\text{RT}}^{*}\) and \(d_{\text{linear}}^{*}\) (the functional forms of which showed up in Montiel Olea, Qiu, and Stoye (2025)) as well as \(d_{\text{step}}^{*}\) (which is a new result).

To understand how these two “regimes” arise, it is helpful to think of a zero-sum game against an adversarial Nature in which a MMR decision rule and a distribution over parameter values (the least favorable prior) form a Nash equilibrium.9 If identification is strong in the sense of Equation 11, this game has an informative equilibrium: the least favorable prior evenly randomizes over two points \((\mu,U(\theta)) \in \{(\bar{\mu},\overline{I}(\bar{\mu})),(-\bar{\mu},\underline{I}(-\bar{\mu}))\}\), to which the unique Bayes response (and, therefore, uniquely optimal MMR rule) is \(d_{w,0}^*\). In all other cases, the equilibrium is uninformative: the least favorable prior is supported on four points \((\mu, U(\theta))\in\{(\bar{\mu},\overline{I}(\bar{\mu})),(\bar{\mu},\underline{I}(\bar{\mu})),(-\bar{\mu},\overline{I}(-\bar{\mu})),(-\bar{\mu},\underline{I}(-\bar{\mu}))\}\) with a probability profile such that the posterior expectation of \(U(\theta)\) always equals \(0\). While any decision rule best responds to this prior, not every decision rule is MMR because the least favorable prior must also best respond to the decision rule. In this very structured setting, the latter is guaranteed by conditions Equation 13 and Equation 14, establishing claim (ii) and nonuniqueness of optimal decision rules. Indeed, even the nonrandomized threshold rules from Theorem 1(iv) do not reflect any updating along the game’s equilibrium path; they rather use uninformative features of the data as a randomization device.

In the uninformative equilibrium, we encounter several additional findings. First, despite the problem’s apparent symmetry, \(d_{w,0}^*\) is not optimal even among linear-index threshold rules that use the index \(w^{\top}Y\). Instead, Theorem 1(iii) characterizes exactly two optimal thresholds, one positive and one negative. Second, while the particular linear-index threshold rule \(d_{w,0}^*\) is not \(\Gamma\)-MMR optimal, in higher dimensional problems (\(n>1\)), there do exist linear-index threshold rules that are. This finding is in stark contrast to Montiel Olea, Qiu, and Stoye (2025), who find that, in a large class of special cases, no linear index threshold rule is MMR optimal. The crucial difference in settings is that the set of priors much constrains the decision theoretic problem’s state space; as a result, the signal space is much richer than the state space, and this can be exploited to mimic randomization without nominally randomizing. Compare Manski (2024)’s abstract observation that, if a policy maker is not allowed to explicitly randomize, sampling uncertainty can be beneficial by providing an implicit randomization device. We note that this phenomenon is reminiscent of classic “purification” results in game theory Dvoretzky, Wald, and Wolfowitz (1951); Khan, Rath, and Sun (2006).10 In contrast, it is not deeply related to the general intuition that “Bayesians don’t randomize.”

We next apply Theorem 1 to Example 1 and immediately get the following results.

Corollary 1 (\(\Gamma\)-MMR decisions in Example 1) In Example 1, the following statements are true:

  1. If

\[ \frac{\bar{\mu}+k}{2k}\geq\Phi\left(\bar{\mu}/\sigma\right), \tag{16}\]

then

\[ d_0^*(\cdot):=\mathbf{1}\{\hat{\mu}\geq0\} \]

is the unique \(\Gamma\)-MMR optimal decision rule.

  1. If Equation 16 fails, then a rule \(d \in \mathcal{D}_1\) attains \(\Gamma\)-MMR if, and only if, it implies

\[ \begin{aligned} \mathbb{E}[d(\hat{\mu}) \mid \mu=-\bar{\mu}] &= \frac{-\bar{\mu}+k}{2k}, \\ \mathbb{E}[d(\hat{\mu}) \mid \mu=\bar{\mu}] &= \frac{\bar{\mu}+k}{2k}. \end{aligned} \]

In particular, no linear threshold rule is \(\Gamma\)-MMR optimal. The following rules, and any convex combination of them, are all \(\Gamma\)-MMR optimal:

\[ \begin{aligned} d^*_\text{RT}&:=\Phi\left(\frac{\hat{\mu}}{\tilde{\sigma}}\right),\\ d^*_\text{linear}&:=\begin{cases} 0, &\hat{\mu}<-\frac{\sigma^{2}\rho^{*}}{\bar{\mu}},\\ \frac{\bar{\mu}\hat{\mu}+\sigma^{2}\rho^{*}}{2\sigma^{2}\rho^{*}}, & -\frac{\sigma^{2}\rho^{*}}{\bar{\mu}}\leq \hat{\mu}\leq\frac{\sigma^{2}\rho^{*}}{\bar{\mu}},\\ 1, & \hat{\mu}>\frac{\sigma^{2}\rho^{*}}{\bar{\mu}}, \end{cases}\\ d^*_\text{step}&:=\begin{cases} \frac{1}{2} - \frac{\bar{\mu}}{2k\left(2\Phi\left(\frac{\bar{\mu}}{\sigma}\right)-1\right)}, & \hat{\mu}<0,\\ \frac{1}{2} + \frac{\bar{\mu}}{2k\left(2\Phi\left(\frac{\bar{\mu}}{\sigma}\right)-1\right)}, & \hat{\mu}\geq0,\\ \end{cases} \end{aligned} \]

where

\[ \tilde{\sigma}=\sigma\sqrt{\left[\frac{\bar{\mu}}{\sigma \Phi^{-1}\left(\frac{\bar{\mu}+k}{2k}\right)}\right]^{2}-1 }, \]

and \(\rho^{*}>0\) is unique and solves \(\int_{0}^{1}\Phi\left(\frac{2\rho^{*}x-\rho^{*}-(\frac{\bar{\mu}}{\sigma}) ^{2}}{\frac{\bar{\mu}}{\sigma} }\right)dx=\frac{-\bar{\mu}+k}{2k}\).

  1. In case (ii), the best linear threshold rules in terms of \(\Gamma\)-minimax regret are \(\mathbf{1}\{\hat{\mu}\geq \pm c^*\}\), where

\[ c^* = \bar{\mu} - \sigma \Phi^{-1}\left(\frac{\bar{\mu}+k}{2k}\right). \]

Code
# Load necessary library
library(ggplot2)
library(latex2exp)

# Run from research/two-point-prior/

# Global parameters
alpha <- 0.7  # Transparency value for plotting
A_col <- "#0A567D"
B_col <- "#963C3C"
C_col <- "#2D6D66"
D_col <- "#6B5957"
E_col <- "#1C84D1"
F_col <- "#A2B1B9"
G_col <- "#FDFBF7"

# Parameters
params <- list(
  list(k = 1, sigma = 1, overline.mu = 1.5),
  list(k = 1, sigma = 1, overline.mu = 0.5),
  list(k = 1, sigma = 0.5, overline.mu = 0.5),
  list(k = 2, sigma = 1, overline.mu = 0.5)
)

# Function to check whether we are in the "small" or large regime
check_condition <- function(overline.mu, sigma, k) {
  return((overline.mu + k) / (2 * k) >= pnorm(overline.mu / sigma))
}

# Decision rules
d_0 <- function(u) {
  as.numeric(u >= 0)
}

d_RT <- function(u, overline.mu, sigma, k) {
  sigma.tilda <-  sigma* sqrt((overline.mu / (sigma * qnorm((overline.mu + k) / (2 * k))))^2 - 1)
  return(pnorm(u / sigma.tilda))
}

rho.finder <- function(rho, overline.mu, sigma, k) {
  w_norm <- overline.mu / sigma
  A <- (rho - w_norm^2) / w_norm
  B <- (-rho - w_norm^2) / w_norm
  integration.output <- (w_norm / (2 * rho)) * (pnorm(A) * A - pnorm(B) * B + (1 / sqrt(2 * pi)) * (exp(-A^2 / 2) - exp(-B^2 / 2)))
  return(integration.output - ((-overline.mu + k) / (2 * k)))
}

d_linear <- function(u, rho.star,overline.mu,sigma) {
  pmax(0, pmin(1, (u*overline.mu  + rho.star*sigma*sigma) / (2 * rho.star*sigma*sigma)))
}

d_step <- function(u, overline.mu, sigma, k) {
  ifelse(u < 0,
         0.5 - overline.mu / (2 * k * (2 * pnorm(overline.mu / sigma) - 1)),
         0.5 + overline.mu / (2 * k * (2 * pnorm(overline.mu / sigma) - 1)))
}

# Generate grid for plotting
u_grid <- seq(-3, 3, 0.01)

# Function to create plots
create_plot <- function(k, sigma, overline.mu) {
  if (check_condition(overline.mu, sigma, k)) {
    d_0_grid <- sapply(u_grid, d_0)
    plot(u_grid, d_0_grid, type = "l", lty = 2, lwd = 5.5, col = A_col,
         xlab = "", ylab = "d(.)", ylim = c(-0.002, 1.002), yaxt = "n", cex.axis = 1.25, yaxs="i")
  } else {
    rho.star <- uniroot(rho.finder, c(0.01, 10), overline.mu = overline.mu, sigma = sigma, k = k)$root
    d_RT_grid <- sapply(u_grid, d_RT, overline.mu = overline.mu, sigma = sigma, k = k)
    d_linear_grid <- sapply(u_grid, d_linear, rho.star = rho.star,overline.mu = overline.mu, sigma = sigma)
    d_step_grid <- sapply(u_grid, d_step, overline.mu = overline.mu, sigma = sigma, k = k)
    
    plot(u_grid, d_RT_grid, type = "l", lty = 1, lwd = 5.5, col = B_col,
         xlab = "", ylab = "d(.)", ylim = c(-0.002, 1.002), yaxt = "n", cex.axis = 1.25, yaxs="i")
    lines(u_grid, d_linear_grid, lty = 2, lwd = 5.5, col = C_col, yaxs="i")
    lines(u_grid, d_step_grid, lty = 4, lwd = 5.5, col = D_col, yaxs="i")
  }
  
  axis(1, cex.axis = 1.25)  # Customize the x-axis ticks
  axis(2, at = seq(0, 1, 0.1), las = 2, cex.axis = 1.25)  # Customize the y-axis ticks
  
  title(main = "")
  mtext(side = 1, line = 3.5, at = mean(par("usr")[1:2]), text = expression(hat(mu)), cex = 1.5)
  mtext(side = 3, line = 0.5, at = mean(par("usr")[1:2]), 
        text = bquote(paste("k = ", .(k), ", ", sigma, " = ", .(sigma), ", ", bar(mu), " = ", .(overline.mu))), 
        cex = 1.5)
  # Add light grid lines
  abline(v = seq(-3, 3, by = 1), col = E_col, lty = 3)
  abline(h = seq(0, 1, by = 0.1), col = E_col, lty = 3)
}

# Set up plotting parameters
svg("figures/fig1.svg", width = 12, height = 13.5)
par(mfrow = c(2, 2), oma = c(8, 2, 2, 2), mar = c(5, 4, 4, 2) + 0.1, bg=G_col)

# Create plots
mapply(function(p) create_plot(p$k, p$sigma, p$overline.mu),
       params)

# Add legend with increased text size and adjusted inset
par(fig = c(0, 1, 0.02, 0.12), oma = c(0, 0, 0, 0), mar = c(0, 0, 0, 0), new = TRUE)
plot(0, 0, type = "n", bty = "n", xaxt = "n", yaxt = "n")
legend("bottom", 
       legend = TeX(c(
         '$d_{0}^{*}$',
         '$d^*_{RT}$',
         '$d^*_{linear}$',
         '$d^*_{step}$'
       )), 
       col = c(A_col, B_col, C_col, D_col),
       lty = c(2, 1, 2, 4), lwd = 6, ncol = 4, bty = "n", xpd = TRUE, 
       title = "", cex = 2.25, title.cex = 0.5)  # Adjust text.width to position lines above text

dev.off()
Figure 1: \(\Gamma\)-MMR optimal rules in Example 1
Notes: This figure reports the \(\Gamma\)-MMR optimal rules in Example 1 for various combinations of parameter values. In the top two panels, the combinations of parameter values satisfy Equation 16. Therefore, the unique \(\Gamma\)-MMR optimal rule is \(d^*_0\). In the bottom two panels, Equation 16 fails. As a result, \(d^*_0\) is no longer \(\Gamma\)-MMR optimal. Instead, \(d^*_{\text{RT}}\), \(d^*_{\text{linear}}\) and \(d^*_{\text{step}}\) are all \(\Gamma\)-MMR optimal.

Note that there is no analog to Theorem 1’s case (iv); indeed, no linear threshold rule is optimal in case (ii). This is because the scalar nature of the signal \(Y\) shuts down the aforementioned purification mechanism. See Figure 1 for an illustration of different \(\Gamma\)-MMR optimal rules in Example 1 for selected parameter values of \(k,\sigma\) and \(\bar{\mu}\).

3.2 Ex-post Robust Bayes Optimality

Theorem 2 (\(\Gamma\)-PER optimal decisions) Suppose all conditions in Theorem 1 hold true. Then:

  1. If \(\underline{I}(\bar{\mu})<0<\overline{I}(\bar{\mu})\), the unique \(\Gamma\)-PER optimal rule is

\[ d_{\operatorname{PER}}^{*}(Y)=\begin{cases} \frac{-\underline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}, & \text{if }w^{\top}Y<0, \\ \frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}, & \text{if }w^{\top}Y\geq0. \end{cases} \]

Otherwise, \(d_{w,0}^{*}\) is \(\Gamma\)-PER optimal.

  1. \(d_{w,0}^{*}\) is always the \(\Gamma\)-PER optimal non-randomized rule.

The results of Theorem 2 offer some important clarifications regarding \(\Gamma\)-PER in treatment choice problems with partial identification. First, even if regret is evaluated according to the posterior distribution, it is not necessarily true that optimal rules are non-randomized. In fact, whenever there is model ambiguity regarding the sign of \(U(\theta)\) (i.e., \(\underline{I}(\bar{\mu})<0<\overline{I}(\bar{\mu})\)), the unique \(\Gamma\)-PER optimal rule is randomized. Therefore, restricting the action space to \(\{0,1\}\) in such a problem is not without loss of generality even under the \(\Gamma\)-PER criterion. Comparing results with Theorem 1 also allows for instructive observations on when \(\Gamma\)-PER and \(\Gamma\)-MMR optimal rules agree or disagree; we elaborate these in Corollary 3. Applying Theorem 2 to Example 1, we finally obtain:

Corollary 2 (\(\Gamma\)-PER rules in Example 1) In Example 1, the following statements are true:

  1. If \(\bar{\mu}<k\), then \(\Gamma\)-PER is uniquely minimized by

\[ d^*_{\operatorname{PER}}(\hat{\mu}) = \begin{cases} \frac{k + \bar{\mu}}{2k} & \text{if } \hat{\mu} \geq 0\\ \frac{k-\bar{\mu}}{2k} & \text{if } \hat{\mu} < 0 \end{cases} \]

  1. If \(\bar{\mu}\geq k\), then \(d_0^*\) is \(\Gamma\)-PER optimal. Furthermore, it is always the \(\Gamma\)-PER optimal threshold rule.
Code
# Load necessary library
library(ggplot2)
library(latex2exp)

# Run from research/two-point-prior/

# Global parameters
alpha <- 0.7  # Transparency value for plotting
A_col <- "#0A567D"
B_col <- "#963C3C"
C_col <- "#2D6D66"
D_col <- "#6B5957"
E_col <- "#1C84D1"
F_col <- "#A2B1B9"
G_col <- "#FDFBF7"

# Parameters
params <- list(
  list(k = 1, sigma = 1, overline.mu = 1.5),
  list(k = 1, sigma = 1, overline.mu = 0.5),
  list(k = 1, sigma = 0.5, overline.mu = 0.5),
  list(k = 2, sigma = 1, overline.mu = 0.5)
)

# Function to check whether we are in the "small" or large regime
check_condition <- function(overline.mu, sigma, k) {
  return(overline.mu >= k)
}

# Decision rules
d_0 <- function(u) {
  as.numeric(u >= 0)
}

d_PER <- function(u, overline.mu, sigma, k){
  ifelse(u>=0,
         (k+overline.mu)/(2*k),
         (k-overline.mu)/(2*k))
}

d_step <- function(u, overline.mu, sigma, k) {
  ifelse(u < 0,
         0.5 - overline.mu / (2 * k * (2 * pnorm(overline.mu / sigma) - 1)),
         0.5 + overline.mu / (2 * k * (2 * pnorm(overline.mu / sigma) - 1)))
}

# Generate grid for plotting
u_grid <- seq(-3, 3, 0.01)

# Function to create plots
create_plot <- function(k, sigma, overline.mu) {
  if (check_condition(overline.mu, sigma, k)) {
    d_0_grid <- sapply(u_grid, d_0)
    plot(u_grid, d_0_grid, type = "l", lty = 2, lwd = 5.5, col = A_col,
         xlab = "", ylab = "d(.)", ylim = c(-0.002, 1.002), yaxt = "n", cex.axis = 1.25, yaxs="i")
  } 
  else {
    d_PER_grid <- sapply(u_grid, d_PER, overline.mu = overline.mu, sigma = sigma, k = k)
    d_step_grid <- sapply(u_grid, d_step, overline.mu = overline.mu, sigma = sigma, k = k)
    d_0_grid <- sapply(u_grid, d_0)
    
    plot(u_grid, d_PER_grid, type = "l", lty = 1, lwd = 5.5, col = B_col,
         xlab = "", ylab = "d(.)", ylim = c(-0.002, 1.002), yaxt = "n", cex.axis = 1.25, yaxs="i")
    if (k>sigma) {
      lines(u_grid, d_step_grid, lty = 4, lwd = 5.5, col = D_col, yaxs="i")
    }
    else {
      lines(u_grid, d_0_grid, lty = 2, lwd = 5.5, col = A_col, yaxs="i")
    }
  }
  
  axis(1, cex.axis = 1.25)  # Customize the x-axis ticks
  axis(2, at = seq(0, 1, 0.1), las = 2, cex.axis = 1.25)  # Customize the y-axis ticks
  
  title(main = "")
  mtext(side = 1, line = 3.5, at = mean(par("usr")[1:2]), text = expression(hat(mu)), cex = 1.5)
  mtext(side = 3, line = 0.5, at = mean(par("usr")[1:2]), 
        text = bquote(paste("k = ", .(k), ", ", sigma, " = ", .(sigma), ", ", bar(mu), " = ", .(overline.mu))), 
        cex = 1.5)
  # Add light grid lines
  abline(v = seq(-3, 3, by = 1), col = E_col, lty = 3)
  abline(h = seq(0, 1, by = 0.1), col = E_col, lty = 3)
}

# Set up plotting parameters
svg("figures/fig2.svg", width = 12, height = 13.5)
par(mfrow = c(2, 2), oma = c(8, 2, 2, 2), mar = c(5, 4, 4, 2) + 0.1, bg=G_col)

# Create plots
mapply(function(p) create_plot(p$k, p$sigma, p$overline.mu),
       params)

# Add legend with increased text size and adjusted inset
par(fig = c(0, 1, 0.02, 0.12), oma = c(0, 0, 0, 0), mar = c(0, 0, 0, 0), new = TRUE)
plot(0, 0, type = "n", bty = "n", xaxt = "n", yaxt = "n")
legend("bottom", 
       legend = TeX(c(
         '$d_{0}^{*}$',
         '$d^*_{PER}$',
         '$d^*_{step}$'
       )), 
       col = c(A_col, B_col, D_col),
       lty = c(2, 1, 4), lwd = 6, ncol = 3, bty = "n", xpd = TRUE, 
       title = "", cex = 2.25, title.cex = 0.5)  # Adjust text.width to position lines above text

dev.off()
Figure 2: \(\Gamma\)-PER optimal rules in Example 1
Notes: This figure reports the \(\Gamma\)-PER optimal rules for the same parameter values considered in Figure 1. In the top left panel, \(\Gamma\)-PER and \(\Gamma\)-MMR optimal rules coincide and are both \(d^*_{0}\). For the rest of the panels, the unique \(\Gamma\)-PER optimal rules are all \(d^*_{\text{PER}}\) and are different from any \(\Gamma\)-MMR optimal rules.

In Figure 2, we depict \(\Gamma\)-PER optimal rules for Example 1 with the same parameter values considered in Figure 1. We see clearly that \(\Gamma\)-MMR and -PER optimal rules coincide (if randomization is allowed) only in the special case when \(k\leq\bar{\mu}\), an observation we generalize in Corollary 3. More specifically, in the top left panel of Figure 2, as \(k\leq\bar{\mu}\), \(\Gamma\)-MMR and -PER coincide and are the non-randomized threshold rule \(d^*_{0}\). In the top right panel, \(k>\bar{\mu}\) and the \(\Gamma\)-PER optimal rule becomes \(d^*_{\text{PER}}\). However, since Equation 16 still holds, \(d^*_0\) is still \(\Gamma\)-MMR optimal. For the bottom two panels, as it still holds \(k>\bar{\mu}\), \(d^*_{\text{PER}}\) is still \(\Gamma\)-PER optimal. However, the associated parameter values imply Equation 16 fails. As a result, \(d^*_{0}\) is no longer \(\Gamma\)-MMR optimal and many \(\Gamma\)-MMR rules exist. But even in these cases, \(\Gamma\)-MMR and -PER rules differ, as among the class of step function rules (which contain \(d^*_{\text{PER}}\)), only \(d^*_{\text{step}}\) is \(\Gamma\)-MMR optimal, still different from \(d^*_{\text{PER}}\).

Code
# Profiled Regret ---------------------------------------------------------

# Packages
rm(list = ls())
library(ggplot2)
library(MASS)
library(latex2exp)
library(foreach)
library(doParallel)

# Run from research/two-point-prior/

alpha <- 0.7  # Transparency value for plotting
A_col <- "#0A567D"
B_col <- "#963C3C"
C_col <- "#2D6D66"
D_col <- "#6B5957"
E_col <- "#1C84D1"
F_col <- "#A2B1B9"
G_col <- "#FDFBF7"

# Values taken from Pepe's code
sigma <- 3.9
LC <- 2.5
x1_x0 <- abs(-7.459)

caux <- (LC * x1_x0) / (sigma * sqrt(pi / 2)) # Controls the gap between "C" and "K"
C <- sigma * sqrt(pi / 2)

k <- round(caux * C, 1)

bar.mu_list <- c(round(0.25 * k, 1), round(0.75 * k, 1))

# Function for profiled regret, special cases
profiled_regret_special_case <- function(k, gamma, sigma, E_d) {
  if (gamma < -k) {
    profiled_regret <- (-gamma + k) * E_d
  } else if (gamma >= -k & gamma <= k) {
    profiled_regret <- max((-gamma + k) * E_d, (gamma + k) * (1 - E_d))
  } else {
    profiled_regret <- (gamma + k) * (1 - E_d)
  }
  return(profiled_regret)
}

# Functions to compute expectations
expectation_d_linear <- function(k, gamma, sigma) {
  rho_star <- rho_star_fun(sigma, k)
  E_d <- pnorm((gamma - rho_star) / sigma) +
    (sigma / (2 * rho_star)) * (1 / sqrt(2 * pi)) * (
      exp((-1 / 2) * (((rho_star + gamma) / sigma) ^ 2)) -
        exp((-1 / 2) * (((rho_star - gamma) / sigma) ^ 2))
    ) +
    (gamma + rho_star) * (1 / (2 * rho_star)) * (
      pnorm((rho_star - gamma) / sigma) -
        pnorm((-rho_star - gamma) / sigma)
    )
  return(E_d)
}

rho_star_fun <- function(sigma, k) {
  rho_star <- uniroot(
    function(x) 1 - (k / x) * (1 - 2 * pnorm(-x / sigma)),
    lower = 0.001,
    upper = k
  )$root
  return(rho_star)
}

expectation_d_PER <- function(k, gamma, sigma, bar.mu) {
  E_d <- 0.5 + (bar.mu / (2 * k)) * (1 - 2 * pnorm(-gamma / sigma))
  return(E_d)
}

expectation_d_Gamma_MMR_linear <- function(k, gamma, sigma, bar.mu) {
  rho_star <- uniroot(rho.finder_Gamma_MMR, c(0.01, 2 * k), bar.mu = bar.mu, sigma = sigma, k = k)$root
  gamma_1 <- bar.mu * gamma / (sigma ^ 2)
  sigma_1 <- bar.mu / sigma

  E_d <- pnorm((gamma_1 - rho_star) / sigma_1) +
    (sigma_1 / (2 * rho_star)) * (1 / sqrt(2 * pi)) * (
      exp((-1 / 2) * (((rho_star + gamma_1) / sigma_1) ^ 2)) -
        exp((-1 / 2) * (((rho_star - gamma_1) / sigma_1) ^ 2))
    ) +
    (gamma_1 + rho_star) * (1 / (2 * rho_star)) * (
      pnorm((rho_star - gamma_1) / sigma_1) -
        pnorm((-rho_star - gamma_1) / sigma_1)
    )

  return(E_d)
}

rho.finder_Gamma_MMR <- function(rho, bar.mu, sigma, k) {
  w_norm <- bar.mu / sigma
  A <- (rho - w_norm ^ 2) / w_norm
  B <- (-rho - w_norm ^ 2) / w_norm
  integration.output <- (w_norm / (2 * rho)) * (pnorm(A) * A - pnorm(B) * B + (1 / sqrt(2 * pi)) * (exp(-A ^ 2 / 2) - exp(-B ^ 2 / 2)))
  return(integration.output - ((-bar.mu + k) / (2 * k)))
}

# Function to create overlaid plots for Profiled Regret
create_profiled_regret_plot <- function(k, sigma, bar.mu) {
  u_grid <- seq(-30, 30, 0.1)
  
  # Compute profile regret values for PER, Global MMR, and Gamma MMR
  profile_regret_PER <- sapply(u_grid, function(gamma) profiled_regret_special_case(k, gamma, sigma, expectation_d_PER(k, gamma, sigma, bar.mu)))
  profile_regret_Global_MMR <- sapply(u_grid, function(gamma) profiled_regret_special_case(k, gamma, sigma, expectation_d_linear(k, gamma, sigma)))
  profile_regret_Gamma_MMR <- sapply(u_grid, function(gamma) profiled_regret_special_case(k, gamma, sigma, expectation_d_Gamma_MMR_linear(k, gamma, sigma, bar.mu)))
  
  # Plot the first line (PER)
  plot(u_grid, profile_regret_PER, type = "l", lty = 1, lwd = 5.5, col = B_col,
       xlab = "", ylab = "Profiled Regret", ylim = c(-0.002, 15), yaxt = "n", cex.axis = 1.25, yaxs="i")
  
  # Add the second line (Global MMR)
  lines(u_grid, profile_regret_Global_MMR, lty = 4, lwd = 5.5, col = D_col)
  
  # Add the third line (Gamma MMR)
  lines(u_grid, profile_regret_Gamma_MMR, lty = 2, lwd = 5.5, col = C_col)
  
  # Add axes and labels
  axis(1, cex.axis = 1.25)
  axis(2, at = seq(0, 15, by = 3), las = 2, cex.axis = 1.25)
  
  title(main = "")
  mtext(side = 1, line = 3.5, at = mean(par("usr")[1:2]), text = expression(mu), cex = 1.5)
  mtext(side = 3, line = 0.5, at = mean(par("usr")[1:2]), 
        text = bquote(paste("k = ", .(k), ", ", sigma, " = ", .(sigma), ", ", bar(mu), " = ", .(bar.mu))), 
        cex = 1.5)
  
  # Add grid lines
  abline(v = seq(-30, 30, by = 5), col = E_col, lty = 3)
  abline(h = seq(0, 15, by = 3), col = E_col, lty = 3)
}

# Set up plotting parameters to create a PDF with two figures and a shared legend
svg("figures/fig3.svg", width = 15, height = 9)
par(mfrow = c(1, 2), oma = c(8, 2, 2, 2), mar = c(5, 4, 4, 2) + 0.1, bg=G_col)

# Create overlaid plots for both bar.mu values
for (t in 1:2) {
  bar.mu <- bar.mu_list[t]
  create_profiled_regret_plot(k, sigma, bar.mu)
}

# Add common legend
par(fig = c(0.0, 1, 0.08, 0.18), oma = c(0, 0, 0, 0), mar = c(0, 0, 0, 0), new = TRUE)
plot(0, 0, type = "n", bty = "n", xaxt = "n", yaxt = "n")
legend("bottom", 
       legend = TeX(c(
         '$d^*_{PER}$',
         '$d^*_{linear}$ (global MMR)',
         '$d^*_{linear}$ (ex-ante $Gamma$ MMR)'
       )), 
       col = c(B_col, D_col, C_col),
       lty = c(1, 4, 2), lwd = 5, ncol = 3, bty = "n", xpd = TRUE, 
       title = "", cex = 1.5, title.cex = 0.5,
       inset = c(0, 0)) # Adjust the inset for better positioning

dev.off()
Figure 3: Profiled regret of \(\Gamma\)-PER and other rules in Example 1
Notes: This figure reports the (frequentist) profiled regrets, as a function of the true mean \(\mu\) of data \(\hat{\mu}\), of \(\Gamma\)-PER rule, \(\Gamma\)-MMR linear rule, and the least randomizing global MMR optimal rule Montiel Olea, Qiu, and Stoye (2025) for the parameter values considered in Montiel Olea, Qiu, and Stoye (2025), fig. 2. In the left plot, we see \(\Gamma\)-PER rule is visually dominated. In both plots, the profiled regret of \(\Gamma\)-MMR linear rule and the least randomizing global MMR optimal rule look very similar and are essentially overlapping with each other.

Applying Montiel Olea, Qiu, and Stoye (2025), Theorem 1 to the current setting, we can conclude that all rules, including both \(\Gamma\)-MMR and \(\Gamma\)-PER optimal rules, are at least admissible. Moreover, by definition, the \(\Gamma\)-MMR rule will have the lower value of the (ex ante) game. Therefore, it might be more useful to compare the (frequentist) profiled regrets Montiel Olea, Qiu, and Stoye (2025) of ex-post \(\Gamma\)-PER and other rules as a function of the true but unknown mean \(\mu\) of data \(\hat{\mu}\). In the context of Example 1, we can report \(\bar{R}(d,\mu):=\sup_{\mu^*\in[\mu-k,\mu+k]}R(d,\mu,\mu^*)\) as \(\mu\) varies for each rule \(d\), where \(R(d,\mu,\mu^*)\) is the expected regret of rule \(d\) as a function of \(\mu\) and \(\mu^{*}\). We report several findings (we emphasize that these are not obvious: we took an ex-ante perspective but did not restrict ourselves to those values of \(\mu\) that the prior allows). First, in both panels of Figure 3, the profiled expected regret of the \(\Gamma\)-PER rule exceeds that of the \(\Gamma\)-MMR rule. Therefore, in this particular example, \(\Gamma\)-MMR arguably outperforms \(\Gamma\)-PER from a broader frequentist point of view. Second, we in fact prove that, whenever \(\bar{\mu}\) is sufficiently small and \(k\) is sufficiently large, the \(\Gamma\)-PER rule is profiled-regret dominated, i.e., there exists a rule \(d\neq d^*_{\text{PER}}\) such that

\[ \overline{R}(d,\mu)\leq\overline{R}(d^*_{\text{PER}},\mu) \tag{17}\]

for all \(\mu\in\mathbb{R}\) with the inequality strict for some \(\mu\); see Lemma Lemma 10 in Section 5.4 for an exact statement. Intuitively, the \(\Gamma\)-PER rule mixes between a coin flip rule and the naive threshold rule (\(d_0^{*}\)). When \(\bar{\mu}\) is small, \(\Gamma\)-PER rule is more analogous to the coin flip rule, which Montiel Olea, Qiu, and Stoye (2025) show to be dominated in terms of profiled regret; also, its profiled regret fails to vanish as \(\mu\to\pm\infty\) even though the optimal treatment is known ex ante in this case. Furthermore, Figure 3 reveals that the \(\Gamma\)-MMR rule \(d^{*}_{\text{linear}}\), though not globally MMR optimal, in some cases has essentially the same profiled risk function as the least randomizing global MMR optimal rule Montiel Olea, Qiu, and Stoye (2025). This may lend a Robust Bayes interpretation to the latter.

3.3 Ex-ante and Ex-post in the Game Against Nature

The different approaches analyzed here can all be expressed as different timing assumptions in the statistical game. The distinction is quite obvious for ex-ante versus ex-post: In the former (and original Waldian) perspective, this game is simultaneous move; in particular, Nature moves before data \(Y\) are realized. In the ex-post perspective, Nature sees \(Y\) before choosing a prior. It is immediately clear that this may be easier to solve because it allows for backward induction. There is also an immediate sense that solutions might not agree, as we indeed found.

But whether the decision maker is allowed to (or at least wants to) randomize or not can equally be thought of as changing the game’s timing, providing another way to think about the action space for the ex-ante and ex-post approaches. The more standard perspective is that the decision maker may randomize over decision rules and Nature must move before learning the outcome of this randomization. By the nature of zero-sum games, this setup will frequently yield randomized solutions. In consequence, it is essential to define the decision maker’s action space as \([0,1]\). In contrast, if Nature is allowed to move after the decision maker’s randomization is realized, then any incentive to randomize is gone and we may as well restrict the action space to \(\{0,1\}\).

Theorem 1 and Theorem 2 clarify that these distinctions actually matter in an interesting example. We spell this out in Corollary 3, which considers both cases when randomization is allowed and not allowed. The bottom line is that the assessment is quite sensitive to how the problem is set up. If underlying parameters lead to sufficiently small identification power, the criteria disagree.

Corollary 3 Consider a treatment choice problem with welfare function Equation 1 and statistical model Equation 3 that satisfies Assumption 1.

  1. Suppose randomization is allowed. Then the \(\Gamma\)-MMR and \(\Gamma\)-PER optimal rules coincide if, and only if, \(\underline{I}(\bar{\mu})\geq0\).

  2. Suppose randomization is not allowed. Then the \(\Gamma\)-MMR and \(\Gamma\)-PER optimal rules coincide if, and only if,

\[ \frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\geq \Phi(\lVert w \rVert_{\Sigma} ). \tag{18}\]

Thus, if randomization is allowed, \(\Gamma\)-MMR and \(\Gamma\)-PER optimal rules coincide only in the somewhat trivial case in which we a priori know that \(U(\theta)\) and \(\mu\) have the same sign, so that optimal treatment choice reduces to Bayesian inference on the point identified \(\mu\). They disagree in all other cases, including in settings where there are infinitely many \(\Gamma\)-MMR optimal rules.

One might expect more agreement once randomization is excluded; after all, this leads to a much simpler action space. Part (ii) shows that there is some truth to this: The condition for agreement changes from \(\underline{I}(\bar{\mu})\geq0\) to the strictly weaker Equation 18. However, the criteria continue to disagree in many cases. It may be instructive to think of these cases in terms of “comparative statics.” For example, consider holding all parameters of the problem fixed but scaling the signal variance \(\Sigma\) by a positive scalar, say to reflect a change in sample size. Then Equation 18 will hold if, and only if, \(\Sigma\) is large enough; hence, as long as \(\underline{I}(\bar{\mu})<0\), increasing sample size will eventually cause disagreement between \(\Gamma\)-MMR and -PER even if randomization is excluded. Similarly, for fixed \(\Sigma\), \(\Gamma\)-MMR and -PER rules will always disagree if the model’s identification power is small enough.

4 Conclusion

We studied treatment choice problems that display partial identification through the lens of the robust Bayes criteria. To do so, we take the general framework of Yata (2021) and others and embed in it a simple example of the set of priors advocated by Giacomini and Kitagawa (2021). We describe and contrast (ex-ante) \(\Gamma\)-minimax regret and (ex-post) \(\Gamma\)-posterior expected regret and analytically derive optimal solutions with and without randomization.

Our results contain two key messages that we think are valuable to the literature. First, with partial identification and multiple priors, ex-ante and ex-post assessments do not agree in general, whether or not randomized rules are allowed. This may at first seem expected due to dynamic inconsistency of multiple prior Bayes criteria, but was not obvious in view of the specific structure of the set of priors. Second, randomization can be optimal in both ex-ante and ex-post problems—it is with loss of generality to exclude them even when regret is evaluated ex-post. The contrast between the results also illustrates a need to better understand the comparative advantages—whether from a theoretical or practical perspective—of using one criterion over the other.

An obvious limitation lies in our use of a convenient but restrictive prior on \(\mu\). As we discovered in Section 3.2, the \(\Gamma\)-MMR rule also performs well across other values of \(\mu\); however, this may be related to the fact that the globally least favorable prior in this setting has a two-point structure on \(\mu\) as well and would obviously not generalize to arbitrary uses of restrictive priors. We provide some insight on more general priors in Section 5.5. In short, the tractability advantage of \(\Gamma\)-PER may become pronounced in such settings, although we hope that recent computational developments Aradillas Fernández et al. (2025); Guggenberger and Huang (2025) will attenuate this concern.

5 Appendix

5.1 Proofs of Main Results

Proofs will frequently claim and verify equilibria of the fictitious game against Nature. Recall that, from basic facts about zero-sum games, if a decision rule uniquely best responds to some least favorable prior, it must be the unique equilibrium rule.

5.1.1 Proof of Theorem 1

Statement (i)

The expected regret of decision rule \(d\) is

\[ R(d,\theta)=U(\theta)\left(\mathbf{1}\{U(\theta)\geq0\}-\mathbb{E}_{m(\theta)}[d(Y)]\right), \quad \theta\in\Theta. \]

Recall that, by the definition of \(\Gamma\), we have \(\pi_{\mu}\sim\text{unif}(\{-\bar{\mu},\bar{\mu}\})\) and, given \(\mu=\pm\bar{\mu}\), \(\pi_{\theta\mid\mu}(U(\theta)\in I(\mu))=1\). Therefore, the Bayes expected regret of \(d\) under prior \(\pi\in\Gamma\) equals

\[ \begin{align*} r(d,\pi) & =\frac{1}{2}\cdot\left[\int_{\tilde{\theta}\in\Theta}U(\tilde{\theta})\left(\mathbf{1}\{U(\tilde{\theta})\geq0\}-\mathbb{E}_{\bar{\mu}}[d]\right)d\pi_{\theta\mid\bar{\mu}}(\tilde{\theta})\right]\\ & +\frac{1}{2}\cdot\left[\int_{\tilde{\theta}\in\Theta}U(\tilde{\theta})\left(\mathbf{1}\{U(\tilde{\theta})\geq0\}-\mathbb{E}_{-\bar{\mu}}[d]\right)d\pi_{\theta\mid-\bar{\mu}}(\tilde{\theta})\right], \end{align*} \]

where \(\mathbb{E}_{\bar{\mu}}[d]:=\mathbb{E}_{\bar{\mu}}[d(Y)]\), \(\mathbb{E}_{-\bar{\mu}}[d]:=\mathbb{E}_{-\bar{\mu}}[d(Y)]\). One can easily solve for

\[ \begin{align*} \sup_{\pi\in\Gamma}r(d,\pi) & =\frac{1}{2}\max\left\{ \overline{I}(\bar{\mu})(1-\mathbb{E}_{\bar{\mu}}[d]),-\underline{I}(\bar{\mu})\mathbb{E}_{\bar{\mu}}[d]\right\} \\ & +\frac{1}{2}\max\left\{ \overline{I}(-\bar{\mu})(1-\mathbb{E}_{-\bar{\mu}}[d]),-\underline{I}(-\bar{\mu})\mathbb{E}_{-\bar{\mu}}[d]\right\}. \end{align*} \]

The least favorable prior \(\pi^*\) equals

\[ \pi^{*}=\left\{ \pi^{*}_{\mu}\sim\text{unif}(\{-\bar{\mu},\bar{\mu}\}),\pi^{*}_{\theta\mid\bar{\mu}}(U(\theta)=\overline{I}(\bar{\mu}))=1,\pi^{*}_{\theta\mid-\bar{\mu}}(U(\theta)=\underline{I}(-\bar{\mu}))=1\right\}. \tag{19}\]

Lemma Lemma 1 shows that the unique Bayes rule against \(\pi^*\) is

\[ d_{w,0}^{*}=\mathbf{1}\{w^{\top}Y\geq0\},\quad\text{where }w=\Sigma^{-1}\bar{\mu}. \]

Lemma Lemma 2 establishes that \(\sup_{\pi\in\Gamma}r(d_{w,0}^{*},\pi)=r(d_{w,0}^{*},\pi^{*})\) as long as Equation 11 holds true. This establishes the claim.

Statement (ii)
Step 1

We show that when Equation 12 holds, any rule \(d\in\mathcal{D}_{n}\) is \(\Gamma\)-MMR optimal if

\[ \mathbb{E}_{\bar{\mu}}[d]=\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})},\quad\mathbb{E}_{-\bar{\mu}}[d]=\frac{-\underline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}. \tag{20}\]

The least favorable prior \(\pi^{*}\) is such that \(\mu\sim\text{unif}(\{-\bar{\mu},\bar{\mu}\})\), and when \(\mu=\bar{\mu}\),

\[ U(\theta)=\begin{cases} \overline{I}(\bar{\mu}), & \text{with probability (w.p.) }p_{1},\\ \underline{I}(\bar{\mu}), & \text{w.p. }1-p_{1}, \end{cases} \]

where \(p_{1}>0\) is such that \(p_{1}\overline{I}(\bar{\mu})+(1-p_{1})\underline{I}(\bar{\mu})=0\), and when \(\mu=-\bar{\mu}\),

\[ U(\theta)=\begin{cases} \overline{I}(-\bar{\mu}), & \text{w.p. }p_{2},\\ \underline{I}(-\bar{\mu}), & \text{w.p. }1-p_{2}, \end{cases} \]

where \(p_{2}\) is such that \(p_{2}\overline{I}(-\bar{\mu})+(1-p_{2})\underline{I}(-\bar{\mu})=0\). Lemma Lemma 3 establishes that any decision rule is Bayes against this prior (intuitively because the data are uninformative), and Lemma Lemma 4 further shows that, for any rule \(d\) that satisfies Equation 20, \(\sup_{\pi\in\Gamma}r(d,\pi)=r(d,\pi^*)\) obtains. This establishes the claim.

Step 2

We next verify the “only if” statement. Recall that the least favorable prior \(\pi^*\) must best respond to any optimal decision rule, i.e., for any MMR optimal rule \(d\), \(\pi^*\) must solve

\[ \begin{align*} \sup_{\pi\in\Gamma}r(d,\pi) & =\frac{1}{2}\max\left\{ \overline{I}(\bar{\mu})(1-\mathbb{E}_{\bar{\mu}}[d]),-\underline{I}(\bar{\mu})\mathbb{E}_{\bar{\mu}}[d]\right\} \\ & +\frac{1}{2}\max\left\{ \overline{I}(-\bar{\mu})(1-\mathbb{E}_{-\bar{\mu}}[d]),-\underline{I}(-\bar{\mu})\mathbb{E}_{-\bar{\mu}}[d]\right\}. \end{align*} \]

This, however, requires that Nature is indifferent between \(\overline{I}(\bar{\mu})\) and \(\underline{I}(\bar{\mu})\) when \(\mu=\bar{\mu}\) and similarly between \(\overline{I}(-\bar{\mu})\) and \(\underline{I}(-\bar{\mu})\) when \(\mu=-\bar{\mu}\). That is, we must have

\[ \overline{I}(\bar{\mu})(1-\mathbb{E}_{\bar{\mu}}[d])=-\underline{I}(\bar{\mu})\mathbb{E}_{\bar{\mu}}[d],\quad\overline{I}(-\bar{\mu})(1-\mathbb{E}_{-\bar{\mu}}[d])=-\underline{I}(-\bar{\mu})\mathbb{E}_{-\bar{\mu}}[d], \]

which is equivalent to Equation 20.

Step 3

We show a rule of form \(\mathbf{1}\left\{ w^{\top}Y\geq c\right\}\) for some \(c\in\mathbb{R}\) cannot be \(\Gamma\)-MMR optimal when Equation 12 holds. Note \(w=\Sigma^{-1}\bar{\mu}\). Thus,

\[ \mathbb{E}_{\mu}[\mathbf{1}\{w^{\top}Y\geq c\}]=1-\Phi\left(\frac{c-w^{\top}\mu}{\sqrt{w^\top \Sigma w}}\right). \]

Suppose by contradiction that a rule \(\mathbf{1}\left\{ w^{\top}Y\geq c\right\}\) is optimal, then by statement (ii) we have

\[ \begin{align} \qquad & 1-\Phi\left(\frac{c-w^\top \bar{\mu}}{\sqrt{w^\top \Sigma w}}\right)=\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\\ \qquad & 1-\Phi\left(\frac{c+w^\top \bar{\mu}}{\sqrt{w^\top \Sigma w}}\right)=\frac{-\underline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}. \end{align} \]

By symmetry, the two equations can both hold only if \(c=0\). But Equation 12 then implies that

\[ \Phi\left(\frac{w^\top \bar{\mu}}{\sqrt{w^\top \Sigma w}}\right)=\Phi(\lVert w \rVert_{\Sigma} )>\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}, \]

so that the first equation cannot in fact hold when \(c=0\), a contradiction.

Step 4

We verify that \(d_{\text{RT}}^*\), \(d_{\text{linear}}^*\) and \(d_{\text{step}}^*\) are all \(\Gamma\)-MMR optimal. Due to symmetry, it suffices to show that

\[ \mathbb{E}_{\bar{\mu}}[d_{\text{RT}}^{*}]=\mathbb{E}_{\bar{\mu}}[d_{\text{linear}}^{*}]=\mathbb{E}_{\bar{\mu}}[d_{\text{step}}^{*}]=\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}. \]

To see \(\mathbb{E}_{\bar{\mu}}[d_{\text{RT}}^*]=\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\), consider the random threshold rule \(\mathbf{1}\left\{ w^\top Y\geq\xi\right\}\), where \(\xi\sim N(0,\tilde{\sigma}^{2})\) is independent of \(Y\). As \(w^\top Y-\xi\sim N\left(\lVert w \rVert_{\Sigma} ^2,\lVert w \rVert_{\Sigma} ^2+\tilde{\sigma}^2\right)\), algebra shows

\[ \mathbb{E}_{\bar{\mu}}\left[\mathbf{1}\left\{ w^\top Y\geq\xi\right\} \right]=\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})} \]

as required. To see \(\mathbb{E}_{\bar{\mu}}[d_{\text{linear}}^*]=\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\), let \(\rho>0\) and

\[ d_{\text{linear},\rho}:=\begin{cases} 0, & w^{\top}Y<-\rho,\\ \frac{w^{\top}Y+\rho^{*}}{2\rho^{*}}, & -\rho\leq w^{\top}Y\leq\rho,\\ 1, & w^{\top}Y>\rho. \end{cases} \]

Applying Lemma C.7 in Montiel Olea, Qiu, and Stoye (2025) and integration by parts yield

\[ \begin{align*} f(\rho):=\mathbb{E}_{\bar{\mu}}[d_{\text{linear},\rho}] &=1-\int_{0}^{1}\Phi\left(\frac{2\rho x-\rho-\lVert w \rVert_{\Sigma} ^{2}}{\lVert w \rVert_{\Sigma} }\right)dx\\ &=1-\frac{\lVert w \rVert_{\Sigma} }{2\rho}\int_{\frac{-\rho-\lVert w \rVert_{\Sigma} ^{2}}{\lVert w \rVert_{\Sigma} }}^{\frac{\rho-\lVert w \rVert_{\Sigma} ^{2}}{\lVert w \rVert_{\Sigma} }}\Phi(t)dt. \end{align*} \]

Note that \(\lim_{\rho\downarrow0}f(\rho)=1-\Phi(-\lVert w \rVert_{\Sigma} )=\Phi(\lVert w \rVert_{\Sigma} )>\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\), while L’Hopital’s rule implies

\[ \begin{align*} \lim_{\rho\rightarrow\infty}f(\rho) &=1-\frac{1}{2}\lim_{\rho\rightarrow\infty}\left\{\Phi\left(\frac{\rho-\lVert w \rVert_{\Sigma} ^{2}}{\lVert w \rVert_{\Sigma} }\right)-\Phi\left(\frac{-\rho-\lVert w \rVert_{\Sigma} ^{2}}{\lVert w \rVert_{\Sigma} }\right)\right\}\\ &=\frac{1}{2}<\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}, \end{align*} \]

where the last inequality follows from \(\overline{I}(\bar{\mu})+\underline{I}(\bar{\mu})>0\) and \(1>\Phi(\lVert w \rVert_{\Sigma} )>\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\). Furthermore, by applying the chain rule, \(\frac{\partial f(\rho)}{\partial\rho}<0\). Therefore, \(f(\cdotp)\) is strictly decreasing in \((0,\infty)\). We conclude that there must exist some unique \(\rho^{*}>0\) such that \(f(\rho^{*})=\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\), or equivalently,

\[ \int_{0}^{1}\Phi\left(\frac{2\rho^{*}x-\rho^{*}-\lVert w \rVert_{\Sigma} ^{2}}{\lVert w \rVert_{\Sigma} }\right)dx=\frac{-\underline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}, \]

which implies \(\mathbb{E}_{\bar{\mu}}[d_{\text{linear}}^{*}]=\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\).

Finally, we verify that \(\mathbb{E}_{\bar{\mu}}[d_{\text{step}}^{*}]=\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\). For any \(\beta\in\left(0,\frac{1}{2}\right)\), consider the following step function rule:

\[ d_{\text{step},\beta}:=\left(\frac{1}{2}-\beta\right)\mathbf{1}\left\{w^{\top}Y<0\right\}+\left(\frac{1}{2}+\beta\right)\mathbf{1}\left\{w^{\top}Y\geq0\right\}. \]

One then has

\[ \begin{align*} \mathbb{E}_{\bar{\mu}}[d_{\text{step},\beta}] &=\left(\frac{1}{2}-\beta\right)\Phi(-\lVert w \rVert_{\Sigma} )+\left(\frac{1}{2}+\beta\right)\left(1-\Phi(-\lVert w \rVert_{\Sigma} )\right)\\ &=\frac{1}{2}+\beta\left(2\Phi(\lVert w \rVert_{\Sigma} )-1\right). \end{align*} \]

Setting \(\mathbb{E}_{\bar{\mu}}[d_{\text{step},\beta^{*}}]=\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\) yields \(\beta^{*}=\frac{\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}-\frac{1}{2}}{2\Phi(\lVert w \rVert_{\Sigma} )-1}\). As \(\overline{I}(\bar{\mu})+\underline{I}(\bar{\mu})>0\) and Equation 12 holds, \(\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}>\frac{1}{2}\) and therefore \(\beta^{*}>0\). Furthermore, \(\beta^{*}<\frac{1}{2}\) holds due to Equation 12 as well. Since \(d_{\text{step}}^{*}=d_{\text{step},\beta^{*}}\), we conclude that \(\mathbb{E}_{\bar{\mu}}[d_{\text{step}}^{*}]=\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\).

Statement (iii)

For any rule of form \(d_{w,c}(Y):=\mathbf{1}\{w^{\top}Y\geq c\}\) where \(w=\Sigma^{-1}\bar{\mu}\) and \(c\in\mathbb{R}\), we may calculate

\[ \mathbb{E}_{\bar{\mu}}[d_{w,c}(Y)]=1-\Phi\left(\frac{c-w^{\top}\bar{\mu}}{\sqrt{w^{\top}\Sigma w}}\right) \]

and, due to Lemma Lemma 9 and recalling \(\lVert w \rVert_{\Sigma} ^{2}=w^{\top}\Sigma w=\bar{\mu}^{\top}\Sigma^{-1}\bar{\mu}\),

\[ \begin{align*} g(c):=\sup_{\pi\in\Gamma} r(d_{w,c},\pi) &=\frac{1}{2}\max\left\{ \overline{I}(\bar{\mu})\Phi\left(\frac{-\lVert w \rVert_{\Sigma} ^{2}+c}{\lVert w \rVert_{\Sigma} }\right),-\underline{I}(\bar{\mu})\Phi\left(\frac{\lVert w \rVert_{\Sigma} ^{2}-c}{\lVert w \rVert_{\Sigma} }\right)\right\}\\ &\quad+\frac{1}{2}\max\left\{ -\underline{I}(\bar{\mu})\Phi\left(\frac{\lVert w \rVert_{\Sigma} ^{2}+c}{\lVert w \rVert_{\Sigma} }\right),\overline{I}(\bar{\mu})\Phi\left(\frac{-\lVert w \rVert_{\Sigma} ^{2}-c}{\lVert w \rVert_{\Sigma} }\right)\right\}. \end{align*} \]

Lemma Lemma 5 shows that \(g\) is decreasing on \([0,c^*]\) and increasing on \([c^{*},\infty)\), implying that the optimal threshold rule is \(d_{w,c^{*}}\) when \(c\in[0,\infty)\). By symmetry, \(d_{w,-c^*}\) is optimal when \(c\in(-\infty,0]\), and \(d_{w,-c^*}\) and \(d_{w,c^*}\) share the same worst-case expected regret.

Statement (iv)

In case Equation 12 and when \(n>1\), there exists \(\dot{\mu}\neq\mathbf{0}\) such that \(\dot{\mu}^{\top}\Sigma^{-1}\bar{\mu}=0\), i.e., \(\dot{\mu}\) is orthogonal to \(\Sigma^{-1}\bar{\mu}\). For any \(t\in\mathbb{R}\), let

\[ d_{w_{t},0}(Y)=\mathbf{1}\{w_{t}^{\top}Y\geq0\},\text{ where }w_{t}=\Sigma^{-1}(t\bar{\mu}+(1-t)\dot{\mu}). \tag{21}\]

Lemma Lemma 6 shows that when \(t=t^{*}\), \(\mathbb{E}_{\bar{\mu}}[d_{w_{t^{*}},0}(Y)]=\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\). By symmetry, one then also has \(\mathbb{E}_{-\bar{\mu}}[d_{w_{t^{*}},0}(Y)]=\frac{-\underline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\). Applying statement (ii) yields that \(d_{w_{t^{*}},0}\) is MMR optimal.

5.1.2 Proof of Corollary 1

In Example 1, \(\overline{I}(\bar{\mu})=\bar{\mu}+k,\underline{I}(\bar{\mu})=\bar{\mu}-k\), \(\Sigma=\sigma^{2}\). The results of the corollary follow directly from Theorem 1(i)-(iii).

5.1.3 Proof of Theorem 2

Recall that the \(\Gamma\)-PER optimal rule solves

\[ \inf_{a\in[0,1]}\sup_{\pi\in\Gamma}\int_{\tilde{\theta}\in\Theta}L(a,\tilde{\theta})d\pi_{\theta \mid Y}(\tilde{\theta}),\hspace{1em}\forall Y\in\mathbb{R}^{n}, \]

where \(L(a,\theta)=U(\theta)(\mathbf{1}\{U(\theta)\geq0\}-a)\), and \(\pi_{\theta\mid Y}\) is the posterior distribution of \(\theta\) given \(Y\). If randomization is not allowed, the \(\Gamma\)-PER optimal rule solves

\[ \inf_{a\in\{0,1\}}\sup_{\pi\in\Gamma}\int_{\tilde{\theta}\in\Theta}L(a,\tilde{\theta})d\pi_{\theta \mid Y}(\tilde{\theta}),\hspace{1em}\forall Y\in\mathbb{R}^{n}. \]

Statement (i) then follows from Lemma Lemma 7; statement (ii) follows from Lemma Lemma 8.

5.1.4 Proof of Corollary 2

Directly follows from Theorem 2.

5.1.5 Proof of Corollary 3

(i)

“If”: If \(\underline{I}(\bar{\mu})\geq0\), then Theorem 1(i) and Theorem 2(i) apply and establish that \(d_{w,0}^*\) is both the unique \(\Gamma\)-MMR and the unique \(\Gamma\)-PER optimal rule.

“Only if”: If \(\underline{I}(\bar{\mu})<0\), then \(d_{\text{PER}}^{*}\) is uniquely \(\Gamma\)-PER optimal by Theorem 2(i). If condition Equation 11 holds as well, then Theorem 1(i) implies that \(d_{w,0}^*\) is uniquely \(\Gamma\)-MMR optimal; hence, \(\Gamma\)-MMR and \(\Gamma\)-PER optimal rules disagree. If condition Equation 12 applies, then, by Theorem 1(ii), a \(\Gamma\)-MMR optimal rule must satisfy Equation 13 and Equation 14. But \(d_{\text{PER}}^*\) can be written as

\[ d_{\operatorname{PER}}^{*}=d_{\operatorname{step},\beta_{\text{PER}}}=\begin{cases} \frac{1}{2}-\beta_{\text{PER}}, & \text{if }w^{\top}Y<0,\\ \frac{1}{2}+\beta_{\text{PER}}, & \text{if }w^{\top}Y\geq0, \end{cases} \]

where \(\beta_{\text{PER}}=\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}-\frac{1}{2}\). Step 4 for the proof of Theorem 1(ii) shows that the only \(\Gamma\)-MMR optimal rule of form \(d_{\text{step},\beta}\) is \(d_{\text{step}}^*\neq d_{\text{PER}}^*\).

(ii)

“If”: When randomization is not allowed, Theorem 2(ii) shows that \(d_{w,0}^*\) is always \(\Gamma\)-PER optimal. Under condition Equation 18, Theorem 1(i) establishes that \(d_{w,0}^*\) is \(\Gamma\)-MMR optimal as well.

“Only if”: Again, \(d_{w,0}^*\) is always \(\Gamma\)-PER optimal when randomization is not allowed. If Equation 18 fails (i.e., Equation 12 holds) and \(n>1\), it follows by Theorem 1(iv) that the non-randomized threshold rule \(d_{w_{t^*},0}^*\) is \(\Gamma\)-MMR optimal. If \(n=1\), Theorem 1(iii) implies that \(d_{w,0}^*\) is not optimal even among linear threshold rules. Therefore, when Equation 12 holds, \(\Gamma\)-MMR and \(\Gamma\)-PER optimal rules disagree with and without randomization.

5.2 Technical Lemmas Supporting Ex-ante Analysis

Lemma 1 (Lemma) The Bayes rule supported by prior Equation 19 is \(d_{w,0}^*\).

Proof. Note \(\overline{I}(\bar{\mu})>0\) due to \(\overline{I}(\bar{\mu})+\underline{I}(\bar{\mu})>0\) and \(\underline{I}(-\bar{\mu})=-\overline{I}(\bar{\mu})\) by Lemma Lemma 9. Given \(\pi^{*}\), the Bayes optimal rule must solve the posterior problem

\[ \begin{align*} &\min_{a\in[0,1]}&\int_{\tilde{\theta}\in\Theta}L(a,\tilde{\theta})d\pi^*_{\theta \mid Y}(\tilde{\theta}), \\ &&\int_{\tilde{\theta}\in\Theta}L(a,\tilde{\theta})d\pi^*_{\theta \mid Y}(\tilde{\theta})\propto\overline{I}(\bar{\mu})(1-a)\cdot\frac{1}{2}\cdot f(Y|\bar{\mu})+\underline{I}(-\bar{\mu})(-a)\cdot\frac{1}{2}\cdot f(Y|-\bar{\mu}), \end{align*} \]

where \(f(Y|\bar{\mu})\) and \(f(Y|-\bar{\mu})\) are the likelihood of \(Y\) at \(\bar{\mu}\) and \(-\bar{\mu}\). This problem is equivalent to

\[ \min_{a\in[0,1]}\overline{I}(\bar{\mu})f(Y|\bar{\mu})+a\underbrace{\overline{I}(\bar{\mu})}_{>0}(f(Y|-\bar{\mu})-f(Y|\bar{\mu})). \]

Since \(\overline{I}(\bar{\mu})>0\), the unique Bayes optimal rule is \(\mathbf{1}\{f(Y|-\bar{\mu})-f(Y|\bar{\mu})\leq0\}\), which is equivalent to \(d^*_{w,0}\) after further algebra.

Lemma 2 (Lemma) Consider decision rule \(d^*_{w,0}\) and prior \(\pi^{*}\) defined in Equation 19. If Equation 11 holds, then

\[ \sup_{\pi\in\Gamma}r(d_{w,0}^{*},\pi)=r(d_{w,0}^{*},\pi^{*}). \tag{22}\]

Proof. As \(w^{\top}Y\sim N(w^{\top}\mu,w^{\top}\Sigma w)\), algebra shows

\[ \mathbb{E}_{\mu}[d_{w,0}^{*}] =\Phi\left(\frac{w^{\top}\mu}{\sqrt{w^{\top}\Sigma w}}\right) \]

for all \(\mu\in M\). In particular,

\[ \mathbb{E}_{\bar{\mu}}[d_{w,0}^{*}]=\Phi\left(\frac{w^{\top}\bar{\mu}}{\sqrt{w^{\top}\Sigma w}}\right),\quad \mathbb{E}_{-\bar{\mu}}[d_{w,0}^{*}]=\Phi\left(-\frac{w^{\top}\bar{\mu}}{\sqrt{w^{\top}\Sigma w}}\right). \]

It follows that \(\sup_{\pi\in\Gamma}r(d_{w,0}^{*},\pi)=r(d_{w,0}^{*},\pi^{*})\) as long as

\[ \overline{I}(\bar{\mu})\Phi\left(-\frac{w^{\top}\bar{\mu}}{\sqrt{w^{\top}\Sigma w}}\right)\geq-\underline{I}(\bar{\mu})\Phi\left(\frac{w^{\top}\bar{\mu}}{\sqrt{w^{\top}\Sigma w}}\right) \]

and

\[ \overline{I}(-\bar{\mu})\Phi\left(\frac{w^{\top}\bar{\mu}}{\sqrt{w^{\top}\Sigma w}}\right)\leq-\underline{I}(-\bar{\mu})\Phi\left(-\frac{w^{\top}\bar{\mu}}{\sqrt{w^{\top}\Sigma w}}\right). \]

Both inequalities are the same as Equation 11 after further algebra, recalling that \(\overline{I}(-\bar{\mu})=-\underline{I}(\bar{\mu})\) and \(-\underline{I}(-\bar{\mu})=\overline{I}(\bar{\mu})\) by Lemma Lemma 9, \(w=\Sigma^{-1}\bar{\mu}\), and \(\lVert w \rVert_{\Sigma} =\sqrt{\bar{\mu}^{\top}\Sigma^{-1}\bar{\mu}}\).

Lemma 3 (Lemma) Consider the prior \(\pi^{*}\) such that \(\mu\sim\operatorname{unif}(\{-\bar{\mu},\bar{\mu}\})\), and when \(\mu=\bar{\mu}\),

\[ U(\theta)=\begin{cases} \overline{I}(\bar{\mu}), & \text{w.p. }p_{1},\\ \underline{I}(\bar{\mu}), & \text{w.p. }1-p_{1}, \end{cases} \]

where \(p_{1}\) is such that \(p_{1}\overline{I}(\bar{\mu})+(1-p_{1})\underline{I}(\bar{\mu})=0\), and when \(\mu=-\bar{\mu}\),

\[ U(\theta)=\begin{cases} \overline{I}(-\bar{\mu}), & \text{w.p. }p_{2},\\ \underline{I}(-\bar{\mu}), & \text{w.p. }1-p_{2}, \end{cases} \]

where \(p_{2}\) is such that \(p_{2}\overline{I}(-\bar{\mu})+(1-p_{2})\underline{I}(-\bar{\mu})=0\). Given this prior, any decision rule is Bayes optimal under Equation 12.

Proof. As Equation 12 holds, we have \(\underline{I}(\bar{\mu})<0<\overline{I}(\bar{\mu})\). Analogous to Lemma Lemma 1, a Bayes rule must solve

\[ \begin{align*} \min_{a\in[0,1]}\quad & \frac{1}{2}\cdot f(Y\mid \bar{\mu})\left[p_{1}\overline{I}(\bar{\mu})(1-a)+(1-p_{1})(-\underline{I}(\bar{\mu}))a\right]\\ &+\frac{1}{2} \cdot f(Y\mid-\bar{\mu})\left[p_{2}\overline{I}(-\bar{\mu})(1-a)+(1-p_{2})(-\underline{I}(-\bar{\mu}))a\right]. \end{align*} \]

Since \(p_{1}\overline{I}(\bar{\mu})+(1-p_{1})\underline{I}(\bar{\mu})=0\) and \(p_{2}\overline{I}(-\bar{\mu})+(1-p_{2})\underline{I}(-\bar{\mu})=0\), the objective is constant in \(a\), hence the claim.

Lemma 4 (Lemma) Consider the prior \(\pi^{*}\) in Lemma Lemma 3. Then, when Equation 12 holds, we have

\[ \sup_{\pi\in\Gamma}r(d,\pi)=r(d,\pi^*) \]

for any decision rule \(d\in\mathcal{D}_{n}\) such that Equation 20 is true.

Proof. For any decision rule \(d\in\mathcal{D}_{n}\) such that Equation 20 is true, algebra shows

\[ \overline{I}(\bar{\mu})(1-\mathbb{E}_{\bar{\mu}}[d])=-\underline{I}(\bar{\mu})\mathbb{E}_{\bar{\mu}}[d]=-\frac{\overline{I}(\bar{\mu})\underline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}, \]

and

\[ \overline{I}(-\bar{\mu})(1-\mathbb{E}_{-\bar{\mu}}[d])=-\underline{I}(-\bar{\mu})\mathbb{E}_{-\bar{\mu}}[d]=-\frac{\overline{I}(\bar{\mu})\underline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}. \]

Thus, \(\sup_{\pi\in\Gamma}r(d,\pi)=-\frac{\overline{I}(\bar{\mu})\underline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\). Meanwhile, for any \(d\) satisfying Equation 20,

\[ r(d,\pi^*)=\frac{1}{2}p_{1}\overline{I}(\bar{\mu})+\frac{1}{2}p_{2}\overline{I}(-\bar{\mu})=-\frac{\overline{I}(\bar{\mu})\underline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}. \]

The desired conclusion follows.

Lemma 5 (Lemma) In case Equation 12, the function

\[ \begin{align*} g(c) & =\frac{1}{2}\max\left\{ \overline{I}(\bar{\mu})\Phi\left(\frac{-\lVert w \rVert_{\Sigma} ^{2}+c}{\lVert w \rVert_{\Sigma} }\right),-\underline{I}(\bar{\mu})\Phi\left(\frac{\lVert w \rVert_{\Sigma} ^{2}-c}{\lVert w \rVert_{\Sigma} }\right)\right\} \\ & +\frac{1}{2}\max\left\{ -\underline{I}(\bar{\mu})\Phi\left(\frac{\lVert w \rVert_{\Sigma} ^{2}+c}{\lVert w \rVert_{\Sigma} }\right),\overline{I}(\bar{\mu})\Phi\left(\frac{-\lVert w \rVert_{\Sigma} ^{2}-c}{\lVert w \rVert_{\Sigma} }\right)\right\} \end{align*} \]

is decreasing in \([0,c^{*}]\) and increasing in \([c^{*},\infty)\).

Proof. When \(c\in[0,c^{*})\) and Equation 12 holds, the two maxima select the terms involving \(-\underline{I}(\bar{\mu})\), so

\[ g(c)=-\frac{\underline{I}(\bar{\mu})}{2}\left[\Phi\left(\frac{\lVert w \rVert_{\Sigma} ^{2}-c}{\lVert w \rVert_{\Sigma} }\right)+\Phi\left(\frac{\lVert w \rVert_{\Sigma} ^{2}+c}{\lVert w \rVert_{\Sigma} }\right)\right], \]

and

\[ \frac{\partial g(c)}{\partial c}=-\frac{\underline{I}(\bar{\mu})}{2\lVert w \rVert_{\Sigma} }\left[\phi\left(\frac{\lVert w \rVert_{\Sigma} ^{2}+c}{\lVert w \rVert_{\Sigma} }\right)-\phi\left(\frac{\lVert w \rVert_{\Sigma} ^{2}-c}{\lVert w \rVert_{\Sigma} }\right)\right]<0. \]

When \(c\in(c^{*},\infty)\), the first maximum switches and

\[ g(c)=\frac{1}{2}\left\{ \overline{I}(\bar{\mu})\Phi\left(\frac{-\lVert w \rVert_{\Sigma} ^{2}+c}{\lVert w \rVert_{\Sigma} }\right)-\underline{I}(\bar{\mu})\Phi\left(\frac{\lVert w \rVert_{\Sigma} ^{2}+c}{\lVert w \rVert_{\Sigma} }\right)\right\}, \]

so

\[ \frac{\partial g(c)}{\partial c}=\frac{1}{2\lVert w \rVert_{\Sigma} }\left\{ \overline{I}(\bar{\mu})\phi\left(\frac{-\lVert w \rVert_{\Sigma} ^{2}+c}{\lVert w \rVert_{\Sigma} }\right)-\underline{I}(\bar{\mu})\phi\left(\frac{\lVert w \rVert_{\Sigma} ^{2}+c}{\lVert w \rVert_{\Sigma} }\right)\right\}>0. \]

Since \(g\) is continuous at \(c=c^{*}\), the claim follows.

Lemma 6 (Lemma) In case Equation 12 and when \(n>1\), let

\[ d_{w_{t},0}(Y)=\mathbf{1}\{w_{t}^{\top}Y\geq0\},\text{ where }w_{t}=\Sigma^{-1}(t\bar{\mu}+(1-t)\dot{\mu}),\dot{\mu}\neq\mathbf{0}, \]

and \(\dot{\mu}^{\top}\Sigma^{-1}\bar{\mu}=0\). Then, \(\mathbb{E}_{\bar{\mu}}[d_{w_{t^{*}},0}(Y)]=\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\), where \(t^{*}\) is defined in Equation 15.

Proof. Write \(f(t):=\mathbb{E}_{\bar{\mu}}[d_{w_{t},0}(Y)]=\Phi\left(\frac{w_{t}^{\top}\bar{\mu}}{\sqrt{w_{t}^{\top}\Sigma w_{t}}}\right)\) and \(k^{*}:=\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\). It suffices to show \(f(t^{*})=k^{*}\), equivalently,

\[ \frac{\left(w_{t^{*}}^{\top}\bar{\mu}\right)^{2}}{w_{t^{*}}^{\top}\Sigma w_{t^{*}}}=\left(\Phi^{-1}\left(k^{*}\right)\right)^{2}. \]

Since \(\dot{\mu}^{\top}\Sigma^{-1}\bar{\mu}=0\),

\[ w_{t^{*}}^{\top}\bar{\mu}=t^{*}\lVert w \rVert_{\Sigma} ^{2}, \]

and

\[ w_{t^{*}}^{\top}\Sigma w_{t^{*}}=(t^{*})^{2}\lVert w \rVert_{\Sigma} ^{2}+(1-t^{*})^{2}\lVert \Sigma^{-1} \dot{\mu}\rVert_{\Sigma}^{2}. \]

Therefore \(t^{*}\) should satisfy

\[ \frac{\left(\Phi^{-1}(k^{*})\right)^{2}}{\lVert w \rVert_{\Sigma} ^{2}}=\frac{(t^{*})^{2}\lVert w \rVert_{\Sigma} ^{2}}{(t^{*})^{2}\lVert w \rVert_{\Sigma} ^{2}+(1-t^{*})^{2}\lVert \Sigma^{-1} \dot{\mu}\rVert_{\Sigma}^{2}}. \]

As \(s^{*}:=\frac{\left(\Phi^{-1}(k^{*})\right)^{2}}{\lVert w \rVert_{\Sigma} ^{2}}\in(0,1)\) due to Equation 12, solving the equation yields Equation 15.

5.3 Technical Lemmas Supporting Ex-post Analysis

Lemma 7 (Lemma) Suppose all the conditions of Theorem 1 hold. If \(\underline{I}(\bar{\mu})<0<\overline{I}(\bar{\mu})\), then the unique \(\Gamma\)-PER optimal rule is

\[ d_{\operatorname{PER}}^{*}(Y)=\begin{cases} \frac{-\underline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}, & \text{if }w^{\top}Y<0,\\ \frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}, & \text{if }w^{\top}Y\geq0. \end{cases} \]

Otherwise, \(d_{w,0}^{*}\) is \(\Gamma\)-PER optimal.

Proof. Let \(f(Y\mid\mu)\) be the likelihood of \(Y\). For each action \(a\in[0,1]\),

\[ \begin{align*} V_{\Gamma}(a) &:=f(Y\mid\bar{\mu})\max\left\{ \overline{I}(\bar{\mu})(1-a),-\underline{I}(\bar{\mu})a\right\}\\ &\quad+f(Y\mid-\bar{\mu})\max\left\{ \overline{I}(\bar{\mu})a,-\underline{I}(\bar{\mu})(1-a)\right\}, \end{align*} \]

using Lemma Lemma 9. If \(\underline{I}(\bar{\mu})\geq0\), the problem reduces to

\[ \inf_{a\in[0,1]}\overline{I}(\bar{\mu})\left\{ f(Y\mid\bar{\mu})+\left[f(Y\mid-\bar{\mu})-f(Y\mid\bar{\mu})\right]a\right\}, \]

whose solution is \(\mathbf{1}\left\{ f(Y\mid-\bar{\mu})-f(Y\mid\bar{\mu})\leq0\right\}=d_{w,0}^{*}\).

If \(\underline{I}(\bar{\mu})<0<\overline{I}(\bar{\mu})\), then

\[ 0<\frac{-\underline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}<\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}<1. \]

The piecewise-linear objective \(V_{\Gamma}(a)\) is minimized at \(\frac{-\underline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\) when \(f(Y\mid-\bar{\mu})>f(Y\mid\bar{\mu})\) and at \(\frac{\overline{I}(\bar{\mu})}{\overline{I}(\bar{\mu})-\underline{I}(\bar{\mu})}\) when \(f(Y\mid-\bar{\mu})<f(Y\mid\bar{\mu})\). Since \(f(Y\mid-\bar{\mu})\leq f(Y\mid\bar{\mu})\) is equivalent to \(\bar{\mu}^{\top}\Sigma^{-1}Y\geq0\), the stated rule follows, up to a null set.

Lemma 8 (Lemma) Suppose all conditions of Theorem 1 hold true. Then, \(d_{w,0}^{*}\) is always the \(\Gamma\)-PER optimal non-randomized rule.

Proof. When randomization is not allowed, we solve \(\inf_{a\in\{0,1\}}V_{\Gamma}(a)\), where \(V_{\Gamma}(a)\) is defined in the proof of Lemma Lemma 7. As

\[ V_{\Gamma}(1)=-f(Y\mid\bar{\mu})\underline{I}(\bar{\mu})+f(Y\mid-\bar{\mu})\overline{I}(\bar{\mu}), \]

and

\[ V_{\Gamma}(0)=f(Y\mid\bar{\mu})\overline{I}(\bar{\mu})-f(Y\mid-\bar{\mu})\underline{I}(\bar{\mu}), \]

the optimal action is \(a=1\) if and only if \(V_{\Gamma}(1)\leq V_{\Gamma}(0)\), which is equivalent to \(d_{w,0}^{*}\) after further algebra.

5.4 Additional Results

Lemma 9 (Lemma) Consider a treatment choice problem with welfare function Equation 1, statistical model Equation 3 and a set of priors Equation 9, that satisfies Assumption 1. Then, the following statements hold:

  1. \(\overline{I}(-\bar{\mu})=-\underline{I}(\bar{\mu})\), \(\underline{I}(-\bar{\mu})=-\overline{I}(\bar{\mu})\).

  2. \(\overline{I}(-\bar{\mu})+\underline{I}(-\bar{\mu})=-(\underline{I}(\bar{\mu})+\overline{I}(\bar{\mu}))\).

Proof. For statement (i), we show only that \(\overline{I}(-\bar{\mu})=-\underline{I}(\bar{\mu})\); the other equality follows analogously. By definition,

\[ \overline{I}(-\bar{\mu})=\sup_{\{\theta\in\Theta:m(\theta)=-\bar{\mu}\}}U(\theta). \]

Since \(\Theta\) is centrosymmetric and both \(U(\cdotp)\) and \(m(\cdot)\) are linear,

\[ \begin{align*} \sup_{\{\theta\in\Theta:m(\theta)=-\bar{\mu}\}}U(\theta) &=\sup_{\{-\theta\in\Theta:-m(-\theta)=-\bar{\mu}\}}-U(-\theta)\\ &=\sup_{\{\tilde{\theta}\in\Theta:-m(\tilde{\theta})=-\bar{\mu}\}}-U(\tilde{\theta})\\ &=-\inf_{\{\tilde{\theta}\in\Theta:m(\tilde{\theta})=\bar{\mu}\}}U(\tilde{\theta})\\ &=-\underline{I}(\bar{\mu}). \end{align*} \]

Statement (ii) follows by summing both equalities in statement (i).

Lemma 10 (Lemma) In Example 1, suppose \(k>\sqrt{\frac{\pi}{2}}\sigma\). Then, if \(\bar{\mu}>0\) is sufficiently small, the corresponding \(\Gamma\)-PER optimal rule is dominated in terms of profiled regret.

Proof. Denote by \(d_{MMR,\text{linear}}^{*}\) the least randomizing global MMR optimal rule derived in Montiel Olea, Qiu, and Stoye (2025). We aim to show that when \(k>\sqrt{\frac{\pi}{2}}\sigma\) and \(\bar{\mu}>0\) is sufficiently small, the associated \(\Gamma\)-PER optimal rule \(d^*_{\text{PER}}\) is such that

\[ \bar{R}(d_{\operatorname{PER}}^{*},\mu)\geq\bar{R}(d_{MMR,\text{linear}}^{*},\mu),\text{ for all }\mu\geq0, \tag{23}\]

with the inequality strict for all \(\mu>0\). A symmetry argument then immediately implies that \(d_{\operatorname{PER}}^{*}\) is dominated in terms of profiled regret.

Pick any \(0<\bar{\mu}<\sqrt{\frac{\pi}{2}}\sigma\). We show that

\[ \bar{R}(d_{\operatorname{PER}}^{*},\mu)=(\mu+k)\left(1-\mathbb{E}_{\mu}\left[d_{\operatorname{PER}}^{*}(\hat{\mu})\right]\right) \]

for all \(\mu\geq0\). By results in Montiel Olea, Qiu, and Stoye (2025), Appendix B.3.1, the profiled regret of a rule \(d\) is

\[ \bar{R}(d,\mu)= \begin{cases} (-\mu+k)\mathbb{E}_{\mu}\left[d(\hat{\mu})\right], & \text{if }\mu<-k,\\ \max\left\{(\mu+k)\left(1-\mathbb{E}_{\mu}\left[d(\hat{\mu})\right]\right),(-\mu+k)\mathbb{E}_{\mu}\left[d(\hat{\mu})\right]\right\}, & \text{if }-k\leq\mu\leq k,\\ (\mu+k)(1-\mathbb{E}_{\mu}\left[d(\hat{\mu})\right]), & \text{if }\mu>k. \end{cases} \]

For the \(\Gamma\)-PER rule,

\[ d_{\operatorname{PER}}^{*}(\hat{\mu}) =\frac{k+\overline{\mu}}{2k}\mathbf{1}\left\{ \hat{\mu}\geq0\right\} +\frac{k-\overline{\mu}}{2k}\mathbf{1}\left\{ \hat{\mu}<0\right\}, \]

and

\[ \mathbb{E}_{\mu}\left[d_{\operatorname{PER}}^{*}(\hat{\mu})\right] =\frac{1}{2}+\frac{\overline{\mu}}{2k}\left(1-2\Phi\left(-\frac{\mu}{\sigma}\right)\right). \]

For any \(0\leq\mu\leq k\), the relevant maximum is attained by \((\mu+k)(1-\mathbb{E}_{\mu}[d_{\operatorname{PER}}^{*}])\) if and only if

\[ \frac{\mu}{1-2\Phi\left(-\frac{\mu}{\sigma}\right)}\geq\overline{\mu}. \]

The left-hand side is increasing in \(\mu\) and has limit \(\sigma\sqrt{\frac{\pi}{2}}\) as \(\mu\downarrow0\), so the condition holds for \(0<\bar{\mu}<\sqrt{\frac{\pi}{2}}\sigma\). Hence, for all \(\mu\geq0\),

\[ \bar{R}(d_{\operatorname{PER}}^{*},\mu) =(\mu+k)\left(\frac{1}{2}-\frac{\overline{\mu}}{2k}\left(1-2\Phi\left(-\frac{\mu}{\sigma}\right)\right)\right). \]

For sufficiently small \(\bar{\mu}>0\), this function is strictly increasing in \(\mu\geq0\). Montiel Olea, Qiu, and Stoye (2025) show that

\[ \sup_{\mu\geq0}\bar{R}(d_{MMR,\text{linear}}^{*},\mu)=\bar{R}(d_{MMR,\text{linear}}^{*},0)=\frac{k}{2}. \]

Since \(\bar{R}(d_{\operatorname{PER}}^{*},0)=\frac{k}{2}\) and \(\bar{R}(d_{\operatorname{PER}}^{*},\mu)\) is strictly increasing for \(\mu>0\), Equation 23 follows, completing the proof.

5.5 General Nature of Our Main Results

In the main text, we derived finite-sample \(\Gamma\)-MMR and -PER optimal rules for a class of priors such that \(\pi_{\mu}\sim\text{unif}(\{-\bar{\mu},\bar{\mu}\})\). In this section, we discuss the implications of our results when \(\pi_{\mu}\) has a more general structure (e.g., with a continuous support).

It is relatively straightforward to extend our \(\Gamma\)-PER results (Theorem 2) to other forms of \(\pi_{\mu}\). Let

\[ V_{\Gamma}(a):=V_{\Gamma}(a,Y):=\sup_{\pi\in\Gamma}\int_{\tilde{\theta}\in\Theta}L(a,\tilde{\theta})d_{\theta\mid Y}(\tilde{\theta}). \]

The ex-post \(\Gamma\)-PER criterion aims to solve \(\min_{a\in[0,1]}V_{\Gamma}(a)\). Given a general \(\pi_{\mu}\) and the normal likelihood of \(Y\), we can derive the posterior distribution of the reduced-form parameter \(\mu\) given \(Y\), written as \(\pi_{\mu\mid Y}\), using the standard Bayes rule. In light of the structure of the class of priors \(\Gamma\),

\[ \begin{align*} V_{\Gamma}(a) &= \int\sup_{\mu^{*}\in I(x)}\left\{ \mu^{*}\left[\mathbf{1}\left\{ \mu^{*}\geq0\right\} -a\right]\right\} d\pi_{\mu\mid Y}(x)\\ &= \int\max\left\{ \overline{I}(x)\left(1-a\right),-\underline{I}(x)a\right\} d\pi_{\mu\mid Y}(x). \end{align*} \]

Note \(\overline{I}(x)(1-a)\geq-\underline{I}(x)a\) if and only if \(\overline{I}(x)\geq a(\overline{I}(x)-\underline{I}(x))\), equivalent to \(a\leq\frac{\overline{I}(x)}{\overline{I}(x)-\underline{I}(x)}\) when \(\overline{I}(x)-\underline{I}(x)>0\) on the support of \(\pi_{\mu\mid Y}\). Thus, for each \(a\in[0,1]\),

\[ \begin{align*} V_{\Gamma}(a) &=\left(1-a\right)\int_{\left\{ x\in\mathbb{R}^{n}:\frac{\overline{I}(x)}{\overline{I}(x)-\underline{I}(x)}\geq a\right\}}\overline{I}(x)d\pi_{\mu\mid Y}(x)\\ &\quad-a\int_{\left\{ x\in\mathbb{R}^{n}:\frac{\overline{I}(x)}{\overline{I}(x)-\underline{I}(x)}<a\right\}}\underline{I}(x)d\pi_{\mu\mid Y}(x). \end{align*} \]

The key observation is that the value of \(V_{\Gamma}(a)\) for each \(a\in[0,1]\) can be numerically solved, although it may not have a closed-form solution. Therefore, in general, we can still find \(d_{\text{PER}}^{*}(Y)\) by numerically solving \(\min_{a\in[0,1]}V_{\Gamma}(a)\). Since the solution may or may not be at the corners \(\{0,1\}\), it is in general not the case that the associated \(\Gamma\)-PER rule always randomizes.

In contrast, the \(\Gamma\)-MMR optimal rule is more difficult to find once we allow general forms of \(\pi_{\mu}\). Similar to the unconstrained MMR optimality problem, one often needs to resort to the “guess-and-verify” strategy by forming a least favorable prior (LFP)—unlike the unconstrained scenario, for the \(\Gamma\)-MMR problem, the LFP has to come from \(\Gamma\), which restricts Nature’s strategy significantly. The particular form of \(\pi_{\mu}\sim\text{unif}(\{-\bar{\mu},\bar{\mu}\})\) makes our guess of the LFP much easier, while for other cases, it may be more difficult. That said, although the exact analytic form of \(\Gamma\)-MMR optimal rule is difficult to acquire for general \(\pi_{\mu}\), one may utilize recent computational techniques Aradillas Fernández et al. (2025); Guggenberger and Huang (2025) to numerically find an \(\varepsilon\)-\(\Gamma\)-MMR optimal rule with near optimal convergence property. Moreover, we think that the main qualitative insights regarding \(\Gamma\)-MMR optimal rules may extend to other forms of \(\pi_{\mu}\). For example, we conjecture that whenever the identification power of the model is sufficiently large compared to the informativeness of the data (in a sense to be suitably defined), \(\Gamma\)-MMR optimal rule will take a form of a non-randomized threshold rule; otherwise, the optimal rule may be randomized. We leave the verification of these conjectures for future research.

6 References

Adjaho, Christopher, and Timothy Christensen. 2022. “Externally Valid Treatment Choice.” arXiv Preprint arXiv:2205.05561.
Amarante, Massimiliano, and Marciano Siniscalchi. 2019. “Recursive Maxmin Preferences and Rectangular Priors: A Simple Proof.” Economic Theory Bulletin 7 (1): 125–29.
Aradillas Fernández, Andrés, José Blanchet, José Luis Montiel Olea, Chen Qiu, Jörg Stoye, and Lezhi Tan. 2025. \(\epsilon\)-Minimax Solutions of Statistical Decision Problems.” arXiv Preprint arXiv:2509.08107. https://arxiv.org/abs/2509.08107.
Athey, Susan, and Stefan Wager. 2021. “Efficient Policy Learning with Observational Data.” Econometrica 89 (1): 133–61.
Ben-Michael, Eli, D James Greiner, Kosuke Imai, and Zhichao Jiang. 2021. “Safe Policy Learning Through Extrapolation: Application to Pre-Trial Risk Assessment.” arXiv Preprint arXiv:2109.11679.
Ben-Michael, Eli, Kosuke Imai, and Zhichao Jiang. 2022. “Policy Learning with Asymmetric Utilities.” arXiv Preprint arXiv:2206.10479.
Berger, J. O. 1985. Statistical Decision Theory and Bayesian Analysis. Springer.
Betrò, Bruno, and Fabrizio Ruggeri. 1992. “Conditional!‘-Minimax Actions Under Convex Losses.” Communications in Statistics-Theory and Methods 21 (4): 1051–66.
Bhattacharya, Debopam, and Pascaline Dupas. 2012. “Inferring Welfare Maximizing Treatment Assignment Under Budget Constraints.” Journal of Econometrics 167 (1): 168–96. https://doi.org/http://dx.doi.org/10.1016/j.jeconom.2011.11.007.
Brock, William A. 2006. “Profiling Problems with Partially Identified Structure.” The Economic Journal 116 (515): F427–40.
Canner, Paul L. 1970. “Selecting One of Two Treatments When the Responses Are Dichotomous.” Journal of the American Statistical Association 65 (329): 293–306. http://www.jstor.org/stable/2283593.
Chamberlain, Gary. 2011. Bayesian Aspects of Treatment Choice.” In The Oxford Handbook of Bayesian Econometrics. Oxford University Press. https://doi.org/10.1093/oxfordhb/9780199559084.013.0002.
Chen, Haoning, and Patrik Guggenberger. 2025. “A Note on Minimax Regret Rules with Multiple Treatments in Finite Samples.” Econometric Theory, forthcoming.
Christensen, Timothy, Hyungsik Roger Moon, and Frank Schorfheide. 2022. “Optimal Discrete Decisions When Payoffs Are Partially Identified.” arXiv Preprint arXiv:2204.11748.
D’Adamo, Riccardo. 2021. “Policy Learning Under Ambiguity.” arXiv Preprint arXiv:2111.10904.
DasGupta, Anirban, and William J Studden. 1989. “Frequentist Behavior of Robust Bayes Estimates of Normal Means.” Statistics & Risk Modeling 7 (4): 333–62.
Dehejia. 2005. “Program Evaluation as a Decision Problem.” Journal of Econometrics 125:141–73.
Dvoretzky, A., A. Wald, and J. Wolfowitz. 1951. Elimination of Randomization in Certain Statistical Decision Procedures and Zero-Sum Two-Person Games.” The Annals of Mathematical Statistics 22 (1): 1–21. https://doi.org/10.1214/aoms/1177729689.
Epstein, Larry G, and Martin Schneider. 2003. “Recursive Multiple-Priors.” Journal of Economic Theory 113 (1): 1–31.
Ferguson, T. S. 1967. Mathematical Statistics: A Decision Theoretic Approach. Vol. 7. Academic Press New York.
Giacomini, Raffaella, and Toru Kitagawa. 2021. “Robust Bayesian Inference for Set-Identified Models.” Econometrica 89 (4): 1519–56. https://doi.org/https://doi.org/10.3982/ECTA16773.
Giacomini, R., T. Kitagawa, and M. Read. 2021. “Robust Bayesian Analysis for Econometrics.” CEPR Discussion Paper No. 16488.
Guggenberger, Patrik, and Jiaqi Huang. 2025. “On the Numerical Approximation of Minimax Regret Rules via Fictitious Play.” arXiv Preprint arXiv:2503.10932.
Guggenberger, Patrik, Nihal Mehta, and Nikita Pavlov. 2024. “Minimax Regret Treatment Rules with Finite Samples When a Quantile Is the Object of Interest.” The Pennsylvania State University.
Hayashi, Takashi. 2008. “Regret Aversion and Opportunity Dependence.” Journal of Economic Theory 139 (1): 242–68.
Hirano, Keisuke, and Jack R. Porter. 2009. “Asymptotics for Statistical Treatment Rules.” Econometrica 77 (5): 1683–1701. https://doi.org/10.3982/ECTA6630.
———. 2020. “Asymptotic Analysis of Statistical Decision Rules in Econometrics.” In Handbook of Econometrics, Volume 7A, edited by Steven N. Durlauf, Lars Peter Hansen, James J. Heckman, and Rosa L. Matzkin, 7:283–354. Handbook of Econometrics. Elsevier. https://doi.org/https://doi.org/10.1016/bs.hoe.2020.09.001.
Ida, Takanori, Takunori Ishihara, Koichiro Ito, Daido Kido, Toru Kitagawa, Shosei Sakaguchi, and Shusaku Sasaki. 2025. “Choosing Who Chooses: Selection-Driven Targeting in Energy Rebate Programs.” Econometrica, forthcoming.
Ishihara, Takuya, and Toru Kitagawa. 2021. “Evidence Aggregation for Treatment Choice.” arXiv Preprint arXiv:2108.06473.
Kallus, Nathan, and Angela Zhou. 2018. “Confounding-Robust Policy Improvement.” Advances in Neural Information Processing Systems 31.
Karlin, Samuel, and Herman Rubin. 1956. “The Theory of Decision Procedures for Distributions with Monotone Likelihood Ratio.” The Annals of Mathematical Statistics, 272–99.
Khan, M. Ali, Kali P. Rath, and Yeneng Sun. 2006. “The Dvoretzky-Wald-Wolfowitz Theorem and Purification in Atomless Finite-Action Games.” International Journal of Game Theory 34:91–104.
Kido, Daido. 2022. “Distributionally Robust Policy Learning with Wasserstein Distance.” arXiv Preprint arXiv:2205.04637.
———. 2023. “Locally Asymptotically Minimax Statistical Treatment Rules Under Partial Identification.” arXiv Preprint arXiv:2311.08958.
Kitagawa, Toru. 2012. “Estimation and Inference for Set-Identified Parameters Using Posterior Lower Probability.” Manuscript, UCL.
Kitagawa, Toru, Sokbae Lee, and Chen Qiu. 2022. “Treatment Choice with Nonlinear Regret.” arXiv Preprint arXiv:2205.08586.
———. 2023. “Treatment Choice, Mean Square Regret and Partial Identification.” The Japanese Economic Review 74 (4): 573–602.
Kitagawa, Toru, Shosei Sakaguchi, and Aleksey Tetenov. 2021. “Constrained Classification and Policy Learning.” arXiv Preprint.
Kitagawa, Toru, and Aleksey Tetenov. 2018. “Who Should Be Treated? Empirical Welfare Maximization Methods for Treatment Choice.” Econometrica 86 (2): 591–616.
———. 2021. “Equality-Minded Treatment Choice.” Journal of Business Economics and Statistics 39 (2): 561–74.
Kitagawa, Toru, and Guanyi Wang. 2023. “Who Should Get Vaccinated? Individualized Allocation of Vaccines over SIR Network.” Journal of Econometrics 232 (1): 109–31.
Lehmann, Erich L, and George Casella. 1998. Theory of Point Estimation. Second. Springer Science & Business Media.
Lei, Lihua, Roshni Sahoo, and Stefan Wager. 2023. “Policy Learning Under Biased Sample Selection.” arXiv Preprint arXiv:2304.11735.
Manski, Charles F. 2000a. “Identification Problems and Decisions Under Ambiguity: Empirical Analysis of Treatment Response and Normative Analysis of Treatment Choice.” Journal of Econometrics 95 (2): 415–42.
Manski, Charles F. 2000b. “Identification Problems and Decisions Under Ambiguity: Empirical Analysis of Treatment Response and Normative Analysis of Treatment Choice.” Journal of Econometrics 95:415–42.
———. 2004. “Statistical Treatment Rules for Heterogeneous Populations.” Econometrica 72 (4): 1221–46.
———. 2005. Social Choice with Partial Knowledge of Treatment Response. Princeton University Press.
———. 2007a. Identification for Prediction and Decision. Harvard University Press.
———. 2007b. “Minimax-Regret Treatment Choice with Missing Outcome Data.” Journal of Econometrics 139:105–15.
———. 2020. “Towards Credible Patient-Centered Meta-Analysis.” Epidemiology 31:345–52.
———. 2024. “Identification and Statistical Decision Theory.” Econometric Theory, in press.
Manski, Charles F., and Aleksey Tetenov. 2007. “Admissible Treatment Rules for a Risk-Averse Planner with Experimental Data on an Innovation.” Journal of Statistical Planning and Inference 137 (6): 1998–2010.
Manski, Charles F, and Aleksey Tetenov. 2023. “Statistical Decision Theory Respecting Stochastic Dominance.” The Japanese Economic Review 74:447–69.
Mbakop, Eric, and Max Tabord-Meehan. 2021. “Model Selection for Treatment Choice: Penalized Welfare Maximization.” Econometrica 89 (2): 825–48.
Montiel Olea, José Luis, Chen Qiu, and Jörg Stoye. 2025. “Decision Theory for Treatment Choice Problems with Partial Identification.” Review of Economic Studies, forthcoming.
Moon, Hyungsik Roger, and Frank Schorfheide. 2012. “Bayesian and Frequentist Inference in Partially Identified Models.” Econometrica 80 (2): 755–82. https://doi.org/https://doi.org/10.3982/ECTA8360.
Poirier, Dale J. 1998. “Revising Beliefs in Nonidentified Models.” Econometric Theory 14 (4): 483–509. http://www.jstor.org/stable/3533214.
Savage, L. 1951. “The Theory of Statistical Decision.” Journal of the American Statistical Association 46:55–67.
Schlag, Karl H. 2006. ELEVEN - Tests Needed for a Recommendation.” European University Institute Working Paper, ECO No. 2006/2.
Song, Kyungchul. 2014. “Point Decisions for Interval-Identified Parameters.” Econometric Theory 30 (2): 334–56. https://doi.org/10.1017/S0266466613000327.
Stoye, Jörg. 2007. “Minimax Regret Treatment Choice with Incomplete Data and Many Treatments.” Econometric Theory 23 (1): 190–99. https://doi.org/10.1017/S0266466607070089.
———. 2009. “Minimax Regret Treatment Choice with Finite Samples.” Journal of Econometrics 151 (1): 70–81.
———. 2011. “Axioms for Minimax Regret Choice Correspondences.” Journal of Economic Theory 146 (6): 2226–51.
———. 2012a. “Minimax Regret Treatment Choice with Covariates or with Limited Validity of Experiments.” Journal of Econometrics 166 (1): 138–56.
———. 2012b. “New Perspectives on Statistical Decisions Under Ambiguity.” Annu. Rev. Econ. 4 (1): 257–82.
Tetenov, Aleksey. 2012a. “Measuring Precision of Statistical Inference on Partially Identified Parameters.” Discuss. Pap., Coll. Carlo Alberto, Torino.
———. 2012b. “Statistical Treatment Choice Based on Asymmetric Minimax Regret Criteria.” Journal of Econometrics 166 (1): 157–65.
Vidakovic, Brani. 2000. \(\Gamma\)-Minimax: A Paradigm for Conservative Robust Bayesians.” Robust Bayesian Analysis, 241–59.
Wakai, Katsutoshi. 2007. “A Note on Recursive Multiple-Priors.” Journal of Economic Theory 135 (1): 567–71.
Wald, Abraham. 1945. “Statistical Decision Functions Which Minimize the Maximum Risk.” Annals of Mathematics 46 (2): 265–80. http://www.jstor.org/stable/1969022.
Yata, Kohei. 2021. “Optimal Decision Rules Under Partial Identification.” arXiv Preprint arXiv:2111.04926.

  1. Giacomini, Kitagawa, and Read (2021) discuss both notions; they refer to the ex-ante and ex-post problems as “Gamma-minimax” and “Conditional Gamma-minimax”, respectively. Christensen, Moon, and Schorfheide (2022) focuses on the ex-post problem for treatment choice problems with partial identification in a restricted class of decision rules.↩︎

  2. For an estimation problem with a quadratic loss, Kitagawa (2012), Appendix B derives the ex-post \(\Gamma\)-minimax estimator and shows that it is not ex-ante \(\Gamma\)-minimax optimal.↩︎

  3. See Epstein and Schneider (2003) and also Wakai (2007), Amarante and Siniscalchi (2019), and references therein.↩︎

  4. Among others, see Savage (1951), Manski (2004), Stoye (2012)b, and Montiel Olea, Qiu, and Stoye (2025) for justifications of focusing on regret in treatment choice problems. In particular, while minimax loss can be an attractive alternative to minimax regret, it leads to trivial recommendations in treatment choice settings including our examples.↩︎

  5. Randomization could be i.i.d. across future potential treatment recipients, fractional in the sense of randomly assigning a certain fraction of the treatment population (in this sense, \(a\in[0,1]\) can also be interpreted as the fraction of the population receiving the treatment), or an “all or nothing” randomization for the entire treatment population. While these might not be practically equivalent in all applications, they are in the current decision theoretic framework. See Manski and Tetenov (2007) for an exception in the related literature.↩︎

  6. In our setting, information from data \(Y\) does not revise the conditional prior \(\pi_{\theta\mid\mu}\) Giacomini and Kitagawa (2021). For any event \(A\) in the \(\sigma\)-algebra of \(\Theta\), we therefore have \(\pi_{\theta\mid Y}(A)=\int\pi_{\theta\mid\mu}(A)d\pi_{\mu\mid Y}\), where \(\pi_{\mu\mid Y}\) is the posterior distribution of \(\mu\) given \(Y\).↩︎

  7. Giacomini, Kitagawa, and Read (2021) discuss both criteria for general loss functions and refer to Definitions Definition 1 and Definition 2 as the “unconditional \(\Gamma\)-minimax” and “conditional \(\Gamma\)-minimax” problems, respectively. For treatment choice problems with partial identification, Christensen, Moon, and Schorfheide (2022) optimize the \(\Gamma\)-PER criterion, restricting the action space to be \(\{0,1\}\).↩︎

  8. In this case, the identified set for \(U(\theta)\) does not change with the true value of \(\mu\), and under all decision criteria considered here, the solution will be a no-data rule that simply flips a coin.↩︎

  9. Analyzing this game is also how results are formally proved. See Wald (1945) for what may be the first clear statement of this and Lehmann and Casella (1998), Theorem 1.4, Chap. 5 for a formalization. Within the literature on decisions under partial identification, the proof technique was first explicitly used in Stoye (2007) and the equilibrium structure with a noninformative and an informative regime was first encountered in Stoye (2012)a.↩︎

  10. We thank Elliot Lipnowski for reminding us of this literature.↩︎