HTML from LaTeXML, with custom CSS/JS. The PDF is more accurate.

Beyond Unbounded Beliefs:
How Preferences and Information Interplay in Social Learning111We thank Nageeb Ali, Marina Halac, Ben Golub, Andreas Kleiner, Elliot Lipnowski, José Montiel Olea, Xiaosheng Mu, Harry Pei, Jacopo Perego, Evan Sadler, Lones Smith, and Peter Sørensen for helpful comments. We also received useful feedback from various seminar and conference audiences. César Barilla, John Cremin, and Zikai Xu provided excellent research assistance. Kartik gratefully acknowledges support from NSF Grant SES-2018948.

Navin Kartik Department of Economics, Columbia University. E-mail: nkartik@columbia.edu.    SangMok Lee Department of Economics, Washington University in St. Louis. E-mail: sangmoklee@wustl.edu.    Tianhao Liu Department of Economics, Columbia University. E-mail: tl3014@columbia.edu.    Daniel Rappoport Booth School of Business, University of Chicago. E-mail: Daniel.Rappoport@chicagobooth.edu.
April 2024
Abstract

When does society eventually learn the truth, or take the correct action, via observational learning? In a general model of sequential learning over social networks, we identify a simple condition for learning dubbed excludability. Excludability is a joint property of agents’ preferences and their information. We develop two classes of preferences and information that jointly satisfy excludability: (i) for a one-dimensional state, preferences with single-crossing differences and a new informational condition, directionally unbounded beliefs; and (ii) for a multi-dimensional state, intermediate preferences and subexponential location-shift information. These applications exemplify that with multiple states “unbounded beliefs” is not only unnecessary for learning, but incompatible with familiar informational structures like normal information. Unbounded beliefs demands that a single agent can identify the correct action. Excludability, on the other hand, only requires that a single agent must be able to displace any wrong action, even if she cannot take the correct action.

Keywords: social learning; herds; information cascades; single crossing; Euclidean preferences; location-shift information; unbounded beliefs.

1 Introduction

This paper concerns the classic sequential observational or social learning model initiated by Banerjee (1992) and Bikhchandani et al. (1992). There is an unknown payoff-relevant state (e.g., product quality). Each of many agents has homogeneous preferences over her own action and the state (e.g., all prefer products of higher quality). Agents act in sequence, each receiving her own private information about the state and observing some subset of her predecessors’ actions. The central economic question is about asymptotic learning: do Bayesian agents eventually learn to take the correct action (e.g., will the highest quality product eventually prevail)?

One would anticipate that whether there is social learning depends on the combination of agents’ preferences and their information structure. But, at least for finite action sets, economists have largely emphasized the latter dimension alone.222Unless noted otherwise, our introduction should be understood as referring to the canonical sequential social learning model with a finite action set, homogeneous preferences, and no direct payoff externalities. It is well recognized that variations in those aspects can also matter for social learning; see for example, Lee (1993) on infinite action spaces, Avery and Zemsky (1998) and Eyster et al. (2014) on endogenous prices or congestion costs, and Goeree et al. (2006) on heterogeneous preferences. The reason is inextricably tied to focusing on models with two states. With only two states, there is social learning given any (nontrivial) preferences if and only if there is learning for all preferences. For, with two states, even the former requires private signals/beliefs to be unbounded (Smith and Sørensen, 2000; Acemoglu et al., 2011). Unbounded beliefs says that given any full-support prior it should be possible for a single private signal, however unlikely it is, to make an agent arbitrarily close to certain about the true state.

With multiple—i.e., more than two—states, unbounded beliefs still characterizes learning for all preferences (Arieli and Mueller-Frank, 2021).333Arieli and Mueller-Frank (2021, Theorem 1) refer to the condition as “totally unbounded beliefs”. They establish their result for a complete network, i.e., when each agent observes the actions of all predecessors. A by-product of our analysis is to establish it for general networks (Corollary 1 in Section 3). However, it is now a very demanding condition. Consider, for instance, the canonical example of normal information: the state is ωΩ and agents’ signals are drawn independently from a normal distribution with mean ω and fixed variance. With only two states, there is unbounded beliefs because a very high signal makes one arbitrarily convinced of the high state, while a very low signal makes one arbitrarily convinced of the low state. But with multiple states, normal information fails unbounded beliefs: given any full-support prior, there is an upper bound on how certain one can become about any non-extremal state based on observing one signal.444So binary states is special because all states are extreme states. There is nothing exceptional about normal information violating unbounded beliefs; see Remark 2 in Section 3. Is social learning doomed with multiple states for familiar information structures like normal information?

Our paper shows that the answer is no. With multiple states, whether society eventually learns to take the correct action depends on the interplay of preferences and information. Crucially, learning can obtain under standard preferences with familiar information structures that fail unbounded beliefs. Figure 1 illustrates an example of normal information with state space Ω={1,2,3}, action set A={a1,a2}, and a uniform prior μ0. The failure of unbounded beliefs is reflected in the set of posteriors, represented by the black curve, being bounded away from state 2’s vertex. For concreteness, suppose that each agent observes all predecessors’ actions. In 1(a), preferences violate single crossing—defined formally in Section 4—because action a1 is optimal in both states 1 and 3, whereas a2 is optimal in state 2. Here, learning fails: since action a1 is optimal after any signal the first agent receives, society is stuck with all agents taking a1. By contrast, in 1(b), agents have single-crossing preferences; specifically an agent who takes action ai gets the quadratic-loss utility (iω)2. Now, at any belief at which learning the state would be useful (i.e., a belief that puts positive probability on both state 1, where a1 is optimal, and either state 2 or 3, where a2 is optimal), no single action is optimal after all signals. This property yields social learning; see Theorem 1 in Section 3.

(a) u(a1,ω)=12, u(a2,ω)=(2ω)2
(b) u(ai,ω)=(iω)2
Figure 1: Belief simplex for state space Ω={1,2,3}. The curve depicts the set of posteriors for a single agent under normal information with prior μ0. The action set is A={a1,a2} and each agent’s utility is u(a,ω). The shaded regions depict optimal actions under uncertainty.

Excludability.

Our paper develops a simple joint condition on information and preferences, which we call excludability, that is not only sufficient for social learning on general observational networks (satisfying a mild condition known as expanding observations), but in a sense also necessary; see Theorem 2 in Section 3.

Roughly speaking, excludability requires that for each pair of actions, a and a, a single agent must be able to receive a signal that makes her arbitrarily convinced that a is better than a, no matter which (full-support) belief she starts with. Put differently, information must be able to distinguish the set of states in which a is better than a from the set in which a is better than a. Excludability implies that society can never get stuck on a wrong action: if an action is suboptimal at the true state, then some agent will receive a private signal convincing her not to take that action. We establish that this property of displacing wrong actions leads to social learning. Notably, an agent can displace wrong actions even if she cannot take the correct action, i.e., the optimal action at the true state. (See Figure 2 in Section 3 for a concrete example.) We view the distinction of social learning arising from the individual capacity to displace wrong actions rather than to discover the correct action as a key insight; this distinction cannot be seen with only two states, where the two notions are equivalent.

Excludability provides a useful perspective on existing ideas in the literature. For instance, as detailed in Section 3, an information structure yields excludability for all preferences if and only if that information structure has unbounded beliefs. But more importantly, we can use excludability to deduce weaker informational conditions that yield social learning for canonical classes of preferences.555Although this approach of obtaining more tenable conditions by restricting preferences to some broad class is novel to social learning, it is classical in other areas of economics. For instance, first-order stochastic dominance is weakened to second-order by restricting to concave (and increasing) utility functions.

Single-crossing preferences.

Our leading application of excludability is to preferences with single-crossing differences (SCD). Here we show that learning obtains when the information structure satisfies directionally unbounded beliefs (DUB). SCD is a familiar property (Milgrom and Shannon, 1994) that is widely assumed in economics: it captures settings in which there are no preference reversals as the state increases. By contrast, DUB appears to be a new condition on information structures, although Milgrom (1979) utilizes a related property in the context of auction theory. Like SCD, DUB is formulated for a (totally) ordered state space. It requires that for any state ω and any prior that puts positive probability on ω, there exist both: (i) signals that make one arbitrarily certain that the state is at least ω; and (ii) signals that make one arbitrarily certain that the state is at most ω. Crucially, no signal need make one arbitrarily certain about ω (unlike unbounded beliefs). For the normal information structure discussed earlier, requirements (i) and (ii) are met for any state by arbitrarily high and arbitrarily low signals, respectively.

Proposition 1 in Section 4 shows that SCD preferences and DUB information are jointly sufficient for excludability, and hence learning. For a direct intuition on the SCD-DUB interplay, consider normal information again. There are preferences (like those in 1(a)) under which society can get stuck at some belief at which agents are taking an incorrect action, but only a strong signal about an intermediate state would change the action—alas no such signal is available. However, under SCD preferences (like those in 1(b)), if knowing that the state is some intermediate ω would change the action, then so would knowing that the state is at least ω or at most ω. Normal information, or more generally DUB, guarantees that there are strong signals approximating such knowledge.

Intermediate preferences.

Our second application in Section 4 is to intermediate preferences in multidimensional spaces (Grandmont, 1978), where the state is ωd and the action is ad. These subsume both constant-elasticity-of-substitution preferences common in many areas of economics and Euclidean preferences invoked in political economy and communication/delegation models.

Using excludability, we show that social learning obtains under intermediate preferences so long as information is given by a subexponential location-shift family. Location-shift families are widely-used information structures: for some density g:dd, the signal distribution in any state ω is given by g(sω). Loosely, the subexponential condition requires that the tail of g must be thin enough, eventually decreasing faster than an exponential rate. We establish that this thin-tails property combined with intermediate preferences yields excludability. Notably, multidimensional normal information (i.e., normally distributed signals with mean equal to the state and some fixed covariance matrix) satisfies the subexponential requirement.

Methodology.

A significant contribution of our paper is also methodological. We develop an approach to tackle learning, and more generally, asymptotic social welfare with multiple states in general observational networks. Theorem 1 in Section 3 is the backbone by which we tie learning to excludability. Theorem 1 reduces the complex dynamic problem of social learning in networks to a much simpler “static” problem. The theorem says that there is learning if and only if every stationary belief has adequate knowledge. A stationary belief is one at which there is an action that is optimal no matter an agent’s signal, and an adequate-knowledge belief is one at which there is an action that is optimal no matter the state in the belief’s support. Excludability is a simple sufficient—and necessary, in a sense explained later—condition for all stationary beliefs to have adequate knowledge.

Theorem 1 itself is a consequence of Theorem 3 in Section 5, which provides a welfare lower bound even when learning fails. The theorem roughly says that for any preferences and information (and given expanding observations), agents eventually obtain at least their cascade utility. Cascade utility is the minimum expected utility an agent can get from any Bayes-plausible distribution of stationary beliefs. Theorem 3 implies that learning obtains when the cascade utility equals the utility obtained from taking the correct action in each state, which leads to Theorem 1.

Related literature.

A number of papers on sequential Bayesian social learning only consider the complete observational network: each agent observes all her predecessors’ actions. For that case and with binary states, Smith and Sørensen (2000) show that, given any nontrivial preferences, there is learning if and only if beliefs are unbounded. For the complete network but with multiple states, Arieli and Mueller-Frank (2021) show that unbounded beliefs—which they call “totally unbounded beliefs”—is sufficient for learning, and also necessary if learning must obtain no matter society’s preferences.666The early work of Bikhchandani, Hirshleifer, and Welch (1992) allowed for multiple states, but they only identified failures of learning because they implicitly restricted attention to “bounded beliefs”; more precisely, they assumed finite signals with full-support distributions. The approach of both Smith and Sørensen (2000) and Arieli and Mueller-Frank (2021) rests on the social belief—an agent’s belief based on observing her predecessors’ actions, before observing her own signal—being a martingale in the complete network.

Gale and Kariv (2003) and Çelen and Kariv (2004) depart from the complete network, noting that martingale methods now fail. Both these papers also depart from the canonical setting in other ways, however: in Gale and Kariv (2003) agents choose actions repeatedly, while in Çelen and Kariv (2004) private signals are not independent conditional on the true state. Acemoglu, Dahleh, Lobel, and Ozdaglar (2011) provide a general treatment of observational networks in an otherwise classical setting. But they only allow for binary states and binary actions. They introduce the condition of expanding observations, explaining that this property of the network is necessary for learning. They establish that it is also sufficient for learning with unbounded beliefs. Building on Banerjee and Fudenberg (2004), a key contribution of Acemoglu, Dahleh, Lobel, and Ozdaglar (2011) is to use a welfare improvement principle to deduce learning; this approach works even though martingale arguments fail. Lobel and Sadler (2015) introduce a notion of “information diffusion” and use the improvement principle to establish information diffusion even when learning fails.

The analysis in both Acemoglu, Dahleh, Lobel, and Ozdaglar (2011) and Lobel and Sadler (2015) relies on their binary-state binary-action structure.777Banerjee and Fudenberg (2004) and Smith and Sørensen (2020) consider “unordered” random sampling models that also only allow for binary states and actions. We believe ours is the first paper to consider the canonical sequential social learning problem with general observational networks and general state and action spaces. At a methodological level, we develop a novel analysis based on continuity and compactness—rather than monotonicity or other properties that are specific to binary states or actions—that uncovers the fundamental logic underlying a general improvement principle.

Substantively, our focus on multiple states and actions allows us to shed light on how preferences and information jointly shape social learning. As already noted, their interplay in determining learning has not received attention in the prior literature because of its focus on binary states. The only exception we are aware of is Arieli and Mueller-Frank (2021, Theorem 3), discussed in Section 3; their result assumes a special utility function and is only for the complete network.

2 Model

There is a countable state space Ω, endowed with the discrete topology, and standard Borel spaces of actions A and signals S. We allow each of these three sets to be finite or infinite. An information or signal structure is given by a collection of probability measures over S, one for each state, denoted by F(|ω). Assume that for any ω and ω, F(|ω) and F(|ω) are mutually absolutely continuous. It follows that each F(|ω) has a density f(|ω); more precisely, this is the Radon-Nikodym derivative of F(|ω) with respect to some reference measure that is mutually absolutely continuous with every F(|ω). Without further loss of generality we assume f(|)>0, so that no signal rules out any state.

The game.

At the outset, a state ω is drawn from a common prior probability mass function μ0ΔΩ.888For any topological space X, ΔX denotes the set of Borel probability measures over X. Then, an infinite sequence of agents, indexed by n=1,2,, sequentially select actions. An agent n observes both a private signal sn drawn from f(|ω) and the actions of some subset of her predecessors Bn{1,2,,n1}, and then chooses her action anA. Agents’ private signals are drawn independently conditional on the state, and no agent observes either the state or any of her predecessors’ signals. Each observational neighborhood Bn is stochastically generated according to a probability distribution Qn over all subsets of {1,2,,n1}, assumed to be independent across n, independent of the state ω, and independent of any private signals. The distributions (Qn)n constitute the observational network structure and are common knowledge, but the realized neighborhood Bn is the private information of agent n.

Agent n’s information set thus consists of her signal sn, neighborhood Bn, and the actions chosen by the neighbors (ak)kBn.999While we assume that each agent observes the identities of her neighbors as well as their chosen actions, the Conclusion explains how our analysis extends to various cases of “random sampling” in which neighbors’ identities are not observed. Our analysis also applies if agents receive arbitrary information about their predecessors’ realized neighborhoods. Let n denote the set of all possible information sets for agent n. A strategy for agent n is a (measurable) function σn:nΔA.

All agents are expected utility maximizers and have common preferences that depend only on their own action and the state, represented by the utility function u:A×Ω. We assume that utility is bounded: there is u¯0 such that |u(,)|u¯.

We study the Bayes Nash equilibria—hereafter simply equilibria—of this game. We assume that for every belief there is an optimal action, so that an equilibrium exists.101010Existence of optimal actions is assured under standard assumptions, e.g., if A is compact and u(,) is suitably continuous. We also note that as there are no direct payoff externalities, strategic interaction is minimal: any σn affects other agents only insofar as affecting how n’s successors update about signal sn from the observation of action an. Hence, we could just as well adopt (weak) Perfect Bayesian equilibrium or refinements.

Remark 1.

Appendix A describes a more general setting in which our main results are proved. For example, Ω can be a closed subset of and each u(a,) piecewise continuous with A finite. We also do not require the signal distributions to be mutually absolutely continuous.

Adequate learning.

The full-information expected utility given a belief μ is the expected utility under that belief if the state will be revealed before an action is chosen:

u(μ):=ωΩmaxaAu(a,ω)μ(ω).

Given a prior μ0 and a strategy profile σ, agent n’s utility un is a random variable. Let 𝔼σ,μ0[un] be agent n’s ex-ante expected utility. We say there is adequate learning if for every prior μ0 and every equilibrium σ, 𝔼σ,μ0[un]u(μ0). In words, adequate learning requires that given any prior and equilibrium, no matter which state is realized, eventually agents take actions that are arbitrarily close to optimal in that state.111111Our notion of adequate learning is different from Arieli and Mueller-Frank’s (2021), who require learning for all utility functions. Following Aghion et al. (1991), we use “adequate” to signify that learning the state precisely is not necessary when some action is optimal in multiple states. We say there is inadequate learning if adequate learning fails.121212That we deem learning to be inadequate if there is some equilibrium in which learning fails, rather than in every equilibrium, is innocuous given that there is no strategic interaction (cf. fn. 10). On the other hand, the issue of whether learning fails at every prior rather than only at some priors is substantive. We return to this issue in our Conclusion.

We will also be interested in situations in which agents choose from a subset of actions, referred to as a choice set.131313We restrict attention to choice sets such that for every belief there is an optimal action. We say that there is (in)adequate learning for a choice set A~A if there is (in)adequate learning when agents are restricted to choose from actions in A~.

Expanding observations.

As observed by Acemoglu, Dahleh, Lobel, and Ozdaglar (2011), a necessary condition for adequate learning is that the network structure has expanding observations:

K:limnQn(Bn{1,,K})=0. (1)

The reason is that a failure of expanding observations means that for some K, there is an infinite number of agents each of whom, with probability uniformly bounded away from 0, observes at most actions a1,,aK. In that event, the agent cannot do better than choosing her action based on only K+1 signals.

Accordingly, we assume expanding observations. Leading examples of network structures with expanding observations include: (i) the classic complete network in which each agent’s neighborhood is all her predecessors (formally, Qn(Bn={1,,n1})=1); (ii) each agent only observes her immediate predecessor (Qn(Bn={n1})=1); and (iii) each agent observes a random predecessor (Qn(Bn={k})=1/(n1) for all k{1,,n1}).

3 Characterizations of Learning

3.1 Stationary Beliefs and Adequate Knowledge

The key to all our results on learning is Theorem 1 below, which simplifies the question of adequate learning to a “one-shot updating” property of beliefs. To state that result, we require two concepts concerning the value of information.

For any belief μΔΩ, let c(μ):=argmaxaA𝔼μ[u(a,ω)] denote the set of optimal actions under that belief. Abusing notation, for a degenerate belief on state ω we write c(ω). Denoting the posterior after signal s when starting from belief μ by μs, we say that belief μ is stationary if there is ac(μ) such that ac(μs) for μ-a.e. signal s. We say that belief μ has adequate knowledge if there is ac(μ) such that ac(ω) for all ωSuppμ. So a belief is stationary if an agent holding that belief does not benefit from observing a signal from the given information structure.141414Some readers may find it helpful to note that in their setting, Smith and Sørensen (2000) refer to stationary beliefs as “cascade beliefs”. On the other hand, a belief has adequate knowledge if the agent would not benefit from observing a signal from any information structure, in particular learning the state.

Any adequate-knowledge belief, such as a belief that puts probability one on a single state, is stationary. In general, there can be stationary beliefs without adequate knowledge, as seen in 1(a).

Theorem 1.

There is adequate learning if and only if all stationary beliefs have adequate knowledge.

Theorem 1 provides a characterization of adequate learning that holds regardless of the observational network structure, given our maintained assumption of expanding observations. Its “only if” direction is straightforward because our notion of learning considers all priors: if the prior is stationary and has inadequate knowledge, then society is stuck with all agents taking the prior-optimal action even though it is suboptimal in some states. More important and subtle is the theorem’s “if” direction. It is inspired by earlier results, particularly Arieli and Mueller-Frank (2021, Lemma 1) and Lobel and Sadler (2015, Theorem 1), but the logic in the current general setting of arbitrary networks and multiple states and actions is novel. We defer this logic to Section 5, instead turning now to how we can build on Theorem 1 for a more practicable characterization of learning. In particular, we seek a more transparent condition on the combinations of preferences and information that yield adequate learning.

3.2 Excludability

A key notion is whether information allows an agent to become arbitrarily sure about a subset of states Ω relative to another subset Ω′′. To make that precise, let μs(Ω) denote the posterior on states Ω induced by belief μ and signal s, and Prμ(S) be the probability of signal set S induced by belief μ.

Definition 1.

A set Ω is distinguishable from another set Ω′′ if for any ε>0 and μΔ(ΩΩ′′) with μ(Ω)>0, it holds that Prμ(s:μs(Ω)>1ε)>0.

Note that Ω is distinguishable from Ω′′ if and only if every ωΩ is distinguishable from Ω′′. Moreover, if Ω is distinguishable from Ω′′, then every subset of Ω is distinguishable from every subset of Ω′′. The following observation essentially reinterprets distinguishability directly in terms of the signal structure rather than posteriors.

Lemma 1.

Ω is distinguishable from Ω′′ if for every ωΩ and ε>0, there is a positive-probability set of signals S such that

ω′′Ω′′,sS:f(s|ω′′)f(s|ω)<ε.

Conversely, this condition is also necessary if Ω′′ is finite.

We emphasize that the set S in the lemma cannot depend on ω′′Ω′′; for Ω to be distinguished from Ω′′, each ωΩ must be distinguished from all ω′′Ω′′ simultaneously. Consider the example of normal information: Ω and signals are normally distributed on with mean ω and fixed variance. When Ω={1,2,3}, state 2 is distinguishable from 1 because f(s|1)/f(s|2)0 as s, and state 2 is distinguishable from 3 because f(s|3)/f(s|2)0 as s. But state 2 cannot be distinguished from both 1 and 3 simultaneously, because min{f(s|1)/f(s|2),f(s|3)/f(s|2)} is bounded away from 0.

Distinguishability of each state from its complement is the condition of unbounded beliefs; this is termed “totally unbounded beliefs” by Arieli and Mueller-Frank (2021) and is the multi-state extension of the two-state notion introduced by Smith and Sørensen (2000). But with multiple states, unbounded beliefs is incompatible with familiar information structures.

Remark 2.

Under any monotone likelihood ratio property (MLRP) information structure, no state ω is distinguishable from {ω,ω′′} with ω<ω<ω′′.151515For ordered state and signals spaces, the MLRP holds if s>s and ω>ω, f(s|ω)/f(s|ω)f(s|ω)/f(s|ω). Consequently, if |Ω|>2, unbounded beliefs fails under the MLRP.

Fortunately, learning only requires certain subsets of states to be distinguished from each other. For any two actions a1 and a2, let the preferred set Ωa1,a2:={ω:u(a1,ω)>u(a2,ω)} be the set of states in which a1 is strictly preferred to a2.

Definition 2.

A utility function and an information structure jointly satisfy excludability if for every a1 and a2, Ωa1,a2 is distinguishable from Ωa2,a1.

Excludability is a joint condition on preferences and information. It requires that for any pair of actions, a single agent can become arbitrarily certain that one action is strictly better than the other, starting from any belief that does not exclude that event. Since excludability is defined using preferred sets, it is straightforward to deduce which sets must be distinguishable for any given preferences; Lemma 1 then provides a set of likelihood-ratio conditions on the information structure, without reference to beliefs.

Unbounded beliefs implies excludability for any preferences. Conversely, if unbounded beliefs fails, then there is some state ω that is not distinguishable from its complement, and excludability fails when preferences are such that for some a1 and a2, Ωa1,a2={ω} while Ωa2,a1=Ω{ω}. Hence, excludability for all preferences is equivalent to unbounded beliefs. But with multiple states, excludability can be substantially weaker for any given (class of) preferences, as developed in Section 4.161616With only two states, Ω={ω1,ω2}, excludability under any given nontrivial preferences is equivalent to unbounded beliefs. (Nontrivial means that no action is optimal at all states.) For, there must be actions a1 and a2 such that Ωa1,a2={ω1} and Ωa2,a1={ω2}; excludability requires these sets to be mutually distinguishable, which is unbounded beliefs. This matters because:

Theorem 2.

Excludability implies adequate learning for every choice set. If excludability fails and the number of states is finite, then there is inadequate learning for some choice set.

(See Theorem 2 in the appendix for a more general version of Theorem 2 that does not require finiteness in the second statement. Hereafter, for brevity, we leave it as implicit that it is Theorem 2 rather than Theorem 2 we are invoking when discussing the necessity of excludability for learning in an infinite state space.)

Excludability is sufficient for adequate learning because it ensures that wrong actions can always be “displaced”, which by Theorem 1 is the key to social learning. More precisely, excludability guarantees that, no matter the choice set, all stationary beliefs have adequate knowledge. Suppose a belief μ has inadequate knowledge, so that c(μ)c(ω) for some state ωSuppμ. (For simplicity, assume c(μ) and c(ω) are singletons.) Excludability implies that preferred set Ωc(ω),c(μ) is distinguishable from Ωc(μ),c(ω). Hence, with positive probability, an agent who starts with belief μ will obtain a posterior that puts arbitrarily large probability on Ωc(ω),c(μ) relative to Ωc(μ),c(ω), in which event she strictly prefers c(ω) to c(μ). Consequently, μ is not stationary.

We highlight that excludability does not guarantee that a wrong action can always be displaced by the correct action. In other words, even though excludability guarantees that given any wrong action—say, c(μ) when the true state is ω—a single agent can receive a signal convincing her that c(μ) is worse than the correct action c(ω), there may be no signal that leads the agent to take c(ω). When there are two states and finite actions, always being able to displace a wrong action and always being able to take the correct action are equivalent, as they both reduce to unbounded beliefs. But more generally, it is displacing wrong actions that is fundamental for learning.

To illustrate the point concretely, consider the example depicted in Figure 2. There are three states and three actions, Ω=A={1,2,3}. The signal structure and preferences are detailed in the figure’s caption. The correct action in each state ω is a=ω. Importantly, unbounded beliefs fails yet there is excludability.171717Unbounded beliefs fails because under normal information the middle state is not distinguishable from its complement. Excludability can be verified by checking distinguishability of the preferred sets for each pair of actions; alternatively, we note that the preferences satisfy single-crossing differences (SCD), and as explained in Subsection 4.1, SCD and normal information imply excludability. Let agent n’s social belief be her belief about the state given only the history of her neighbors’ actions, prior to observing her own private signal. When each agent observes all predecessors’ actions, Figure 2 shows two representative numerically-simulated paths of social beliefs given the true state ω=2. The social belief starts at the prior, marked by a star in the figure, and then evolves as agents take actions, as indicated by either of the arrowed paths. There is a range of beliefs, shaded in grey, such that for any social belief in that range no signal can lead an agent to take the correct action 2. As the prior is in this range, the first agent necessarily takes a wrong action: either 1 (which occurs in the red path) or 3 (the blue path). Nevertheless, even though no agent can take the correct action 2 for a while, society never gets stuck at a wrong action: given that an agent’s predecessor chose a{1,3}, there are signals (very high if a=1 and very low if a=3) that convince the agent that a is worse than the correct action 2, and hence the agent will not take action a. At some point, after enough switching between actions 1 and 3, the social belief is driven outside the grey region and it becomes possible for an agent to take the correct action 2. Eventually, society settles on that action.

Refer to caption
Figure 2: Two simulated social belief paths—one in red and one in blue—in a complete network. There are three states labeled 1,2,3, and there is normal information (with standard deviation 1.2). There are three actions with respective state-contingent utilities (1,0,0.3), (0,0.2,0), and (0.3,0,1). The optimal action under uncertainty is delineated by the dashed lines. The true state is 2, and society starts with the prior (0.35,0.1,0.55), marked by the black star. The grey shaded region indicates beliefs at which no single signal can lead to state 2’s correct action. On each path, a dot represents the social belief after an agent has acted, and arrows indicate the sequencing.

Turning to necessity in Theorem 2: for a fixed choice set, all stationary beliefs can have adequate knowledge (and hence there is adequate learning, by Theorem 1) even absent excludability. But when excludability fails, there is some preferred set Ωa1,a2 that cannot be distinguished from Ωa2,a1. If Ω is finite, this means that when the choice set is {a1,a2}, a belief that puts small probability on Ωa1,a2 relative to Ωa2,a1 is stationary and has inadequate knowledge. Hence, Theorem 1 implies that excludability is necessary for learning when we seek learning for all choice sets. The following example illustrates these points using an infinite action set for convenience.

Example 1.

Consider Ω={0,1}, A=[0,1], and u(a,ω)=(aω)2. This is an example of “responsive preferences” (Lee, 1993; Ali, 2018). Fix any nontrivial signal structure and any observational network structure satisfying expanding observations.

Adequate learning obtains by Theorem 1, because the only stationary beliefs have certainty on one of the two states. For, given any nondegenerate belief, with positive probability the posterior-optimal action will be different from the prior-optimal action, as the uniquely optimal action equals the posterior expected state. However, excludability is equivalent to the signal structure having unbounded beliefs, as for any a1<a2, Ωa1,a2={0} and Ωa2,a1={1}. So excludability is not necessary for adequate learning at choice set A. But absent excludability there is inadequate learning at any non-singleton finite choice set. For, there is then some state such that any prior that puts probability close to 1 on that state will be stationary, but this prior has inadequate knowledge.

The choice-set variation required by Theorem 2 comes “for free” when we seek an informational condition that ensures learning for a broad-enough class of preferences. Specifically, it is sufficient that for any utility function in the class and any choice set, there is another utility function that is identical on that set but makes all other actions dominated. Since the class of all preferences has this property, and excludability for all preferences is equivalent to the information structure having unbounded beliefs, Theorem 2 immediately implies:

Corollary 1.

An information structure yields adequate learning for all preferences if and only if it has unbounded beliefs.

This corollary extends results from the prior literature, which are either for the complete network (Arieli and Mueller-Frank, 2021, Theorem 1) or general networks but with only two states (Acemoglu, Dahleh, Lobel, and Ozdaglar, 2011, Theorem 2).

To our knowledge, the only prior exception to unbounded beliefs driving learning with a discrete action space is the interesting example of Arieli and Mueller-Frank (2021, Theorem 3). They consider the complete network and a special utility function, which they call “simple utility”, in which the payoff is 1 if the action matches the state and 0 otherwise. For this case, they show that pairwise distinguishability—for any pair of states, each is distinguishable from the other—is sufficient for learning. This result also follows from Theorem 2; indeed, the theorem implies that learning obtains for general observational networks. For, under simple utility, the preferred sets for actions a1a2 are just {a1} and {a2}, which means excludability is equivalent to pairwise distinguishability.

4 Applications

Excludability permits a study of informational conditions that assure adequate learning for broad and widely-used classes of preferences. This section presents two such applications: one with a one-dimensional state, and one with a multi-dimensional state.

4.1 Learning in a One-Dimensional World

In this subsection we assume a totally ordered state space: for simplicity, Ω. A function h:Ω is single crossing if either: (i) for all ω<ω, h(ω)>0h(ω)0; or (ii) for all ω<ω, h(ω)<0h(ω)0. That is, a single-crossing function switches sign between strictly positive and strictly negative at most once.

Definition 3.

Preferences represented by u:A×Ω have single-crossing differences (SCD) if for all a and a, the difference u(a,)u(a,) is single crossing.

SCD is an ordinal property closely related to notions in Milgrom and Shannon (1994) and Athey (2001), but, following Kartik et al. (2023), the formulation is without an order on A.181818SCD is equivalent to there existing some order on A with respect to which Athey’s (2001) “weak single-crossing property of incremental returns” holds. Ignoring indifferences, SCD requires that the preference over any pair of actions can only flip once as the state changes monotonically. SCD is widely satisfied in economic models; in particular, it is assured by supermodularity of u.

The key informational condition is that of distinguishing upper and lower sets from each other. More precisely, we require that for any ω, {ω:ωω} and {ω:ω<ω} are distinguishable from each other, and {ω:ω>ω} and {ω:ωω} are distinguishable from each other. But since a set Ω is distinguishable from Ω′′ if and only if each ωΩ is distinguishable from Ω′′, we can simplify as follows.

Definition 4.

An information structure has directionally unbounded beliefs (DUB) if every ω is distinguishable from {ω:ω<ω} and also from {ω:ω>ω}.

Crucially, DUB does not require any state ω to be distinguishable from any subset of states containing both a higher and a lower state than ω. Rather, using Lemma 1, we can view DUB as only requiring that for any state ω, there are signals that are arbitrarily more likely in ω relative to all ω<ω, and also other signals that are arbitrarily more likely in ω relative to all ω>ω.

A leading example of DUB information is normal information. More generally, for any MLRP information structure, DUB can be easily checked because it reduces to pairwise distinguishability.191919Regardless of the MLRP, DUB implies pairwise distinguishability. To see why the converse is true given the MLRP, consider the case of finite states. Note that for any ω>ω, f(s|ω)/f(s|ω) as ssupS (the ratio is increasing by MLRP, and it diverges by pairwise distinguishability); similarly, the ratio goes to 0 as sinfS. Hence, for any ω and ε>0, the condition in Lemma 1 is met for Ω={ω} and Ω′′={ω′′:ω′′<ω} when S is any sufficiently small upper set of signals, while for Ω′′={ω′′:ω′′>ω} the condition is met when S is any sufficiently small lower set. For an infinite state space, the intuition is the same but we appeal to the monotone convergence theorem. We note that when A is finite, pairwise distinguishability is inescapable (even without the MLRP) for adequate learning in any rich-enough class of preferences.202020“Rich-enough” here means that for any two states, there is a preference in the class such that the optimal actions in those two states are disjoint.

Our main result in this subsection is:

Proposition 1.

If preferences have SCD and the information structure has DUB, then there is adequate learning. Conversely, if the information structure violates DUB and there are at least two actions, then there are SCD preferences for which there is inadequate learning.

The result says that not only is DUB a sufficient informational condition for adequate learning under any SCD preferences, but it also necessary to assure learning for all SCD preferences.

Here is the logic for sufficiency. Recall that Ωa,a denotes the states in which action a is strictly preferred to a. SCD implies non-reversal of strict preferences: for any a and a, either infΩa,asupΩa,a or infΩa,asupΩa,a. DUB says that every upper (resp., lower) set of states and its strict lower (resp., strict upper) set are distinguishable from each other. Therefore, SCD and DUB together guarantee excludability, and so Proposition 1’s first statement follows from Theorem 2.

We would like to caution against the following intuition. Under SCD preferences, any inadequate-knowledge belief μ has distinct optimal actions at the extreme states of μ’s support. DUB information then guarantees learning because the extreme states can be distinguished from their complements, and so μ is not stationary. While valid for finite states, this is not a generally applicable intuition. Indeed, the following example shows that pairing DUB with distinct optimal actions at all states is not a robust principle for learning.

Example 2.

Let Ω= and A={a}. In any state ω, the utility from any integer action a is given by quadratic loss, u(a,ω)=(aω)2, whereas the action a is a “safe action”, u(a,ω)=ε for a small constant ε>0.212121Strictly speaking, quadratic-loss utility with Ω= violates our maintained assumption of bounded utility, but we ignore that to keep the example succinct. So any action ω is uniquely optimal in state ω but worse than the safe action a in every other state. Plainly, SCD is violated.

Consider normal information. There are full-support priors μ such that the posterior probability μs(ω) is uniformly bounded away from 1 across signals s and states ω (see Supplementary Appendix SA.3 for details). For any such prior, for small enough ε>0, the safe action a is optimal after every signal. In other words, any such prior is stationary but has inadequate knowledge. So Theorem 1 implies inadequate learning.

The argument for the necessity of DUB in Proposition 1 is as follows. Take any state ω, any two actions a1a2, and consider the following SCD utility: for all ω<ω, u(a1,ω)=1 and u(a2,ω)=0; for all ωω, u(a1,ω)=0 and u(a2,ω)=1; and otherwise u(a,ω)=1. Since all actions except a1 and a2 are dominated and can be ignored, Theorem 2 implies that for there to be adequate learning, Ωa1,a2={ω:ω<ω} and Ωa2,a1={ω:ωω} must be distinguishable. In particular, ω is distinguishable from its lower set. An analogous argument shows that ω is distinguishable from its upper set. Since ω is arbitrary, DUB holds.

While our main point in this subsection is that DUB is the correct informational condition for adequate learning under SCD preferences, it is also worth noting that for any preferences violating SCD, one can show that there are DUB information structures—e.g., normal information—with inadequate learning at some choice set. In this sense SCD and DUB are a minimal pair of sufficient conditions.

4.2 Learning in a Multi-Dimensional World

We now turn to a multi-dimensional environment: A,Ωd for some integer d1.222222We view any xd as a column vector and denote its transposition by x and its standard Euclidean norm by x. For instance, A={1,2,3}2 can represent a set of feasible policies, Ω={1,2,3}2 society’s ideal policy, and individuals have quadratic-loss preferences u(a,ω)=aω2. Is there a natural class of information structures for which learning obtains?

More generally, consider the following class of preferences:

Definition 5.

Preferences are intermediate if for all a1a2, either Ωa1,a2= or Ωa1,a2=Ω or there are hd and c such that Ωa1,a2={ω:hω>c}.

Introduced by Grandmont (1978), intermediate preferences have preferred sets that are either trivial or half spaces; so if ω,ωΩa1,a2, then for any λ(0,1) and ω′′=λω+(1λ)ωΩ, it holds that ω′′Ωa1,a2. A leading family, subsuming quadratic-loss preferences, is weighted Euclidean preferences: u(a,ω)=l((aω)W(aω)), for some d×d symmetric positive definite matrix W and strictly increasing loss function l:++.232323To confirm that these are intermediate preferences, note that by simple algebraic manipulation, (a1ω)W(a1ω)(a2ω)W(a2ω)=(a1a2)W(a1+a22ω). Hence, ωΩa1,a2 if and only if (a1a2)W(a1+a22ω)>0, or equivalently, hω>c where h=2(a2a1)W and c=(a2a1)W(a1+a2). Another salient example, discussed by Caplin and Nalebuff (1988, Section 5), is the constant-elasticity-of-substitution utility u(a,ω)=(i=1dω(i)(a(i))r)1/r where a(i) and ω(i) denote the respective i-th coordinates, and r0 is a parameter.

Turning to information, we focus on the familiar class of location-shift information structures: S=d and there is a density g:d++, called the standard density, such that f(s|ω)=g(sω). We restrict attention to standard densities that are uniformly continuous. The following property will be crucial.

Definition 6.

A location-shift information structure is subexponential if there are p>1 and M>0 such that g(s)<exp(sp) for all s>M.

A subexponential density has a thin tail in the sense that it eventually decays strictly faster than the exponential density. Our leading example of a subexponential location-shift information structure is multivariate normal information: there is some covariance matrix Σ such that the distribution of signals in state ω is 𝒩(ω,Σ). Here the standard density is that of 𝒩(0,Σ), and Definition 6 is verified by taking any exponent p(1,2) and any large M>0. Subexponential information can fail unbounded beliefs; for example, this is the case for normal information when Ω contains non-extreme states, i.e., there is some state in the interior of the convex hull of Ω.

The main result of this subsection is:

Proposition 2.

If preferences are intermediate and the information structure is subexponential location-shift, then there is adequate learning.

The result follows from Theorem 2 and the next lemma, which says that all half spaces are distinguishable from their complements under subexponential location-shift information. Since the nontrivial preferred sets for intermediate preferences are half spaces, the lemma implies that this combination of preferences and information yields excludability.

Lemma 2.

For a subexponential location-shift information structure, the sets {ω:hω>c} and {ω:hω<c} are distinguishable from each other for any hd and c.

The exponent p being strictly larger than 1 in the definition of subexponential is essential for the lemma. To see that, consider the double-exponential standard density g(s)=cexp(s) with c>0 a constant of integration. This density is not subexponential, and indeed the conclusion of Lemma 2 fails: no two states ωω are distinguishable from each other because f(s|ω)/f(s|ω)=g(sω)/g(sω)exp(ωω) for any signal s. The failure of pairwise distinguishability implies inadequate learning even with a binary state when the action set is discrete and preferences are nontrivial.

We can provide an intuition for Lemma 2 by considering a bivariate normal standard density, g(s)=exp(sΣs/2)/2π with Σ a 2×2 covariance matrix. Take an arbitrary hyperplane h, as illustrated in Figure 3. We seek to distinguish the half space to the right of h from its complementary half space to the left. It is sufficient to distinguish an arbitrary single state ω1 to the right of h from all the states to the left. Figure 3 shows how to construct a sequence of signals verifying that distinguishability. For a sequence of cn0, select sn on the iso-density ellipse of level cn given state ω1 so that the direction of h is tangent with the ellipse at sn. For all n, the “ellipsoid distance” between sn and ω1, (snω1)Σ(snω1), is then smaller than the ellipsoid distance between sn and any state to the left of h (such as ω2 and ω3) by some fixed amount. Due to the normal distribution being subexponential, as cn0 the likelihood ratio g(snω)g(snω1)0 uniformly across ω to the left of h.


Figure 3: The logic underlying Lemma 2 for a bivariate normal standard density. We seek to distinguish ω1 from the solid black line. The ellipses are iso-density signals of a given level at the states ω1, ω2, and ω3. As sn grows along the dotted line, corresponding to lower iso-density levels, min{f(sn|ω2)/f(sn|ω1),f(sn|ω3)/f(sn|ω1)}0.

We make two further comments regarding Proposition 2. First, in the one-dimensional environment of Subsection 4.1, SCD is more or less equivalent to preferred sets being half spaces, and DUB is equivalent to the distinguishability of half spaces from their complements in the sense of Lemma 2. Proposition 2 can thus be viewed as an extension of Proposition 1 to a multi-dimensional world; the restriction to location-shift information allows us to unpack the kind of information that yields the requisite half-space distinguishability.

Second, a location-shift information structure does not have to be subexponential to guarantee learning for all intermediate preferences. But it can be shown that if the standard density g:d is superexponential in the sense that there are p(0,1) and M>0 such that g(s)exp(sp) for all s>M, then learning fails for all nontrivial intermediate preferences when A is finite.

5 Theorem 1 and a General Welfare Bound

We now return to the general characterization of adequate learning, Theorem 1, to explain how it is derived. The theorem is best understood as a corollary of a welfare bound regardless of whether there is learning. Stating that result requires some notation. Abusing notation, let

u(μ):=maxaAωu(a,ω)μ(ω)

be an agent’s expected utility when she takes an optimal action under belief μ. Recalling that μs denotes the posterior given a belief μ and signal s, let

I(μ):=(ωΩSu(μs)dF(s|ω)μ(ω))u(μ)

be the expected utility improvement from observing a private signal at belief μ. Observe that I(μ)=0 for any stationary belief μ. We write ΦBPΔΔΩ to denote the set of Bayes-plausible distributions of beliefs: φΦBP 𝔼φ[μ]=μ0. Again abusing notation, we write u(φ):=𝔼φ[u(μ)] for the expected utility of an agent under the distribution of beliefs φ, and analogously write I(φ):=𝔼φ[I(μ)]. It follows that

ΦS:={φΦBP:I(φ)=0}

is the set of Bayes-plausible distributions of beliefs that are supported on the set of stationary beliefs. (We have suppressed the dependence of ΦBP and ΦS on the prior μ0.)

Building on a notion mentioned by Lobel and Sadler (2015), we can now define the cascade utility level as

u(μ0):=infφΦSu(φ).

In words, u(μ0) is the lowest utility level that an agent can get if her Bayes-plausible distribution of beliefs is supported on stationary beliefs. Our welfare bound is that eventually all agents are assured a utility level of at least u(μ0). More precisely:

Theorem 3.

In any equilibrium σ, liminfn𝔼σ,μ0[un]u(μ0).

The “if” direction of Theorem 1 readily follows from Theorem 3: when all stationary beliefs have adequate knowledge, a correct action is taken almost surely for any distribution of stationary beliefs, hence u(μ0)=u(μ0), and we have adequate learning.

The conclusion of Theorem 3 would be straightforward if we were assured that agents eventually hold stationary beliefs. However, there are networks (with expanding observations) in which with positive probability the beliefs of an infinite number of agents are bounded away from the set of stationary beliefs; see Example SA.1 in Supplementary Appendix SA.1.

Instead, we prove Theorem 3 via an improvement principle, as suggested by Banerjee and Fudenberg (2004) and developed by Acemoglu, Dahleh, Lobel, and Ozdaglar (2011) and others. The foundation in our general setting is a novel compactness-continuity argument. First, Lemma 4 in the appendix establishes that ΦBP is compact when both ΔΩ and ΔΔΩ are endowed with the Prohorov metric generated from the metric on Ω. The idea when Ω is countable is that although the prior μ0 can be supported on an infinite set, it must concentrate an arbitrarily large mass on only finitely many states. Consequently, for any δ>0, there is a finite subset of states Ω such that any Bayes-plausible distribution of beliefs puts at least 1δ probability on beliefs that put at least 1δ probability on Ω. Using Prohorov’s Theorem, we then deduce that ΦBP is compact. Second, we show that the utility function u(φ) and the improvement function I(φ) are continuous (Lemma 5 in the appendix), and thus uniformly continuous on ΦBP.

Now consider any ε-neighborhood of the set of Bayes-plausible distributions supported on stationary beliefs, call it (ΦS)ε. If an agent’s distribution of beliefs is in (ΦS)ε, then her ex-ante expected utility is at least close to u, as u(φ)u on ΦS and u(φ) is uniformly continuous. On the other hand, if the distribution is not in (ΦS)ε, then there is some strictly positive minimum utility improvement that the agent obtains (as the complement of (ΦS)ε is closed, hence compact, and I(φ) is continuous).

We can then apply an improvement principle. The idea is as follows, where we consider deterministic networks for simplicity. Expanding observations guarantees that we can partition society into “generations” such that an agent in one generation observes a predecessor who is in either the previous generation or the current generation. We inductively argue that the lowest ex-ante utility in each generation is either close to u or increases by a fixed amount compared to the previous generation. Consider an agent’s distribution of social beliefs, φ. Her utility u(φ) must be at least the lowest ex-ante expected utility of the previous generation, because the current agent can just mimic the agent with the largest index she observes.242424With stochastic networks, the fact that an agent can obtain any observed predecessor’s ex-ante expected utility through mimicking relies on our assumption that players’ observation neighborhoods are drawn independently. Otherwise, whether a player has observed some predecessor may correlate with that predecessor realizing a lower utility. Then, as explained in the previous paragraph, either u(φ) is at least close to u (when φ is in (ΦS)ε), or the agent can improve upon u(φ) by at least some fixed amount. Thus, the lowest ex-ante expected utility in each generation increases by a fixed amount until it becomes at least close to u. Since ε was arbitrary, it follows that eventually all agents’ utility must be higher than a level arbitrarily close to u, which is the conclusion of Theorem 3.

Although previous authors have deduced versions of Theorem 1 and Theorem 3 in special environments, what allows us to establish these two general results is our novel proof methodology. We highlight two distinctions with Lobel and Sadler (2015, Theorem 1), which is the most related existing result to Theorem 3. They consider a binary-state binary-action model. In that setting, they establish a welfare bound of “diffusion utility”, which is the utility obtained by a hypothetical agent who observes an information structure that contains only the strongest signals (an “expert agent”, in their terminology). Our cascade utility is more fundamentally tied to when learning stops, as it is defined using stationary beliefs. It is not hard to see that in general, no matter the number of states or actions, cascade utility is always at least as high as (the natural extension of) diffusion utility; Remark 5 in Appendix C elaborates. Typically the ranking will be strict, although Lobel and Sadler (2015) note that cascade and diffusion utilities coincide in their binary-state binary-action setting. Methodologically, Lobel and Sadler’s argument for a minimum improvement, like that of Acemoglu, Dahleh, Lobel, and Ozdaglar (2011), owes to certain monotonicity that does not extend beyond their binary-binary setting.

Remark 3.

Our approach to proving Theorem 3 can be adapted to address belief convergence. Since expanding observations is compatible with the observational network having multiple components, one cannot expect the social belief to converge even in probability.252525Consider an observational network consisting of two disjoint complete subnetworks: every odd agent observes only all odd predecessors, and symmetrically for even agents. Given any specification in which learning would fail on a complete network—such as the canonical binary state/binary action herding example—there is positive probability of the limit belief among odd agents being different from that among even agents. Furthermore, there can be a positive probability that the social belief is not eventually even in a neighborhood of the set of stationary beliefs, as already noted. Nevertheless, there are reasonable conditions under which convergence to the stationary set does obtain. Consider deterministic networks and assume that society can be covered by finitely many subsequences such that in each subsequence agent nk observes nk1. Then, denoting agent n’s (random) social belief by μn, it holds that for all ε>0, limnPr(μnSε)=1, where Sε denotes the ε-neighborhood of the set of stationary beliefs. See Proposition SA.1 in Supplementary Appendix SA.1. We note that this result applies, in particular, to the immediate-predecessor network and the complete network. The latter is special because the social belief is then a martingale, which is assured to converge almost surely by the martingale convergence theorem. For this case, Arieli and Mueller-Frank (2021, Lemma 1) have established that the limit is stationary.

Remark 4.

Theorem 3 can be used to quantify how a failure of excludability impacts welfare. Proposition SA.2 in Supplementary Appendix SA.2 provides a formal result in this vein. In particular, that result implies a sense in which an environment with “approximate excludability” ensures that, eventually, agents’ ex-ante expected utilities are close to the full-information utility.

6 Concluding Remarks

This paper has studied a general model of sequential social learning on observational networks. Our main theme has been how learning turns jointly on preferences and information when there are multiple states. We close by commenting on certain aspects of our approach.

First, our model assumes “non-anonymous sampling”, i.e., whenever an agent sees the action of some predecessor, she knows the identity of that predecessor. However, our methodology extends to anonymous sampling, i.e., when each agent observes only the frequencies of actions in their realized neighborhood, as in Smith and Sørensen (2020). Our results apply in that case when expanding observations (condition (1)) holds for the “induced network structure” (Q~n)n where each Q~n is defined by first drawing a neighborhood Bn according to Qn and then uniform-randomly drawing a single agent from Bn. Appendix A (fn. 29) explains why. Interestingly, the condition coincides with Smith and Sørensen’s (2020) “non-over-sampling” requirement. Note that expanding observations for the induced network (Q~n) is more demanding than expanding observations for (Qn); this is not surprising since agents have less information when they cannot observe identities. Nevertheless, the requirement is satisfied, for example, when each agent observes the action of a uniform-randomly drawn predecessor or the actions of all predecessors—in either case, not observing their identities. But the requirement is violated when each agent n observes either agent 1 or agent n1, but doesn’t observe the identity (whereas expanding observations holds here when the identity is observed).

Second, the notion of learning we have adopted considers all possible priors. While this strengthens our sufficiency results, it correspondingly weakens our necessity results. With only two states, learning at any single (nondegenerate) prior is equivalent to learning at all priors. Our earlier working paper (Kartik, Lee, Liu, and Rappoport, 2022, Supplementary Appendix SA.1) provides some analysis concerning the extent to which this is true with multiple states.

Third, our analysis has not touched on the speed of learning/welfare convergence. For binary states and the complete network, Rosenberg and Vieille (2019) deduce the condition on the likelihood of extreme posteriors that determines whether learning is, in certain senses, efficient; they point out that their condition is violated by normal information. See Hann-Caruthers et al. (2018) as well.

Lastly, our work only addresses Bayesian learning with correctly specified agents. There is a large literature on non-Bayesian social learning, surveyed by Golub and Sadler (2016). There has also been recent interest in (mis)learning among misspecified Bayesian agents; see, for example, Frick et al. (2020) and Bohren and Hauser (2021).

References

  • D. Acemoglu, M. A. Dahleh, I. Lobel, and A. Ozdaglar (2011) Bayesian Learning in Social Networks. Review of Economic Studies 78 (4), pp. 1201–1236. External Links: Document, https://academic.oup.com/restud/article-pdf/78/4/1201/18376066/rdr004.pdf, ISSN 0034-6527, Link Cited by: §1, §1, §1, §2, §3.2, §5, §5.
  • P. Aghion, P. Bolton, C. Harris, and B. Jullien (1991) Optimal learning by experimentation. Review of Economic Studies 58 (4), pp. 621–654. External Links: Link Cited by: footnote 11.
  • S. N. Ali (2018) On the role of responsiveness in rational herds. Economics Letters 163, pp. 79–82. Cited by: Example 1.
  • I. Arieli and M. Mueller-Frank (2021) A general analysis of sequential social learning. Mathematics of Operations Research 46 (4), pp. 1235–1249. Cited by: §1, §1, §1, §3.1, §3.2, §3.2, §3.2, Remark 3, footnote 11, footnote 3.
  • S. Athey (2001) Single crossing properties and the existence of pure strategy equilibria in games of incomplete information. Econometrica 69 (4), pp. 861–889. External Links: Document Cited by: §4.1, footnote 18.
  • S. Athey (2002) Monotone comparative statics under uncertainty. Quarterly Journal of Economics 117 (1), pp. 187–223. External Links: Document Cited by: Appendix SA.4.
  • C. Avery and P. Zemsky (1998) Multi-dimensional uncertainty and herd behavior in financial markets. American Economic Review 88 (4), pp. 724–748. Cited by: footnote 2.
  • A. Banerjee and D. Fudenberg (2004) Word-of-mouth learning. Games and Economic Behavior 46 (1), pp. 1–22. Cited by: §1, §5, footnote 7.
  • A. V. Banerjee (1992) A simple model of herd behavior. Quarterly Journal of Economics 107 (3), pp. 797–817. Cited by: §1.
  • S. Bikhchandani, D. Hirshleifer, and I. Welch (1992) A theory of fads, fashion, custom, and cultural change as informational cascades. Journal of Political Economy 100 (5), pp. 992–1026. Cited by: §1, footnote 6.
  • V. I. Bogachev (2018) Weak convergence of measures. American Mathematical Society, Providence, Rhode Island. Cited by: §A.3.
  • J. A. Bohren and D. N. Hauser (2021) Learning with heterogeneous misspecified models: characterization and robustness. Econometrica 89 (6), pp. 3025–3077. External Links: Document, https://onlinelibrary.wiley.com/doi/pdf/10.3982/ECTA15318, Link Cited by: §6.
  • A. Caplin and B. Nalebuff (1988) On 64%-majority rule. Econometrica 56 (4), pp. 787–814. Cited by: §4.2.
  • B. Çelen and S. Kariv (2004) Observational learning under imperfect information. Games and Economic Behavior 47 (1), pp. 72–86. External Links: Document, ISSN 0899-8256, Link Cited by: §1.
  • H. Crauel (2002) Random probability measures on polish spaces. Vol. 11, CRC press. Cited by: §A.1.
  • R. Durrett (2019) Probability: theory and examples. Vol. 49, Cambridge University Press. Cited by: §A.1.
  • E. Eyster, A. Galeotti, N. Kartik, and M. Rabin (2014) Congested observational learning. Games and Economic Behavior 87, pp. 519–538. Cited by: footnote 2.
  • M. Frick, R. Iijima, and Y. Ishii (2020) Misinterpreting others and the fragility of social learning. Econometrica 88 (6), pp. 2281–2328. External Links: Document, Link Cited by: §6.
  • D. Gale and S. Kariv (2003) Bayesian learning in social networks. Games and Economic Behavior 45 (2), pp. 329–346. Cited by: §1.
  • J. K. Goeree, T. R. Palfrey, and B. W. Rogers (2006) Social learning with private and common values. Economic Theory 26 (2), pp. 245–264. Cited by: footnote 2.
  • B. Golub and E. Sadler (2016) Learning in social networks. In The Oxford Handbook of the Economics of Network, Y. Bramoullé, A. Galeotti, and B. W. Rogers (Eds.), pp. 504–542. Cited by: §6.
  • J. Grandmont (1978) Intermediate preferences and the majority rule. Econometrica 46 (2), pp. 317–330. Cited by: §1, §4.2.
  • W. Hann-Caruthers, V. V. Martynov, and O. Tamuz (2018) The speed of sequential asymptotic learning. Journal of Economic Theory 173, pp. 383–409. External Links: Document, ISSN 0022-0531, Link Cited by: §6.
  • N. Kartik, S. Lee, T. Liu, and D. Rappoport (2022) Beyond unbounded beliefs: how preferences and information interplay in social learning. Note: unpublished Cited by: §6.
  • N. Kartik, S. Lee, and D. Rappoport (2023) Single-Crossing Differences in Convex Environments. Review of Economic Studies, pp. forthcoming. External Links: ISSN 0034-6527, Document, Link, https://academic.oup.com/restud/advance-article-pdf/doi/10.1093/restud/rdad103/54092237/rdad103.pdf Cited by: §4.1.
  • I. H. Lee (1993) On the convergence of informational cascades. Journal of Economic Theory 61 (2), pp. 395–411. External Links: Document, ISSN 0022-0531, Link Cited by: Example 1, footnote 2.
  • I. Lobel and E. Sadler (2015) Information diffusion in networks through social learning. Theoretical Economics 10 (3), pp. 807–851. External Links: Document, https://onlinelibrary.wiley.com/doi/pdf/10.3982/TE1549, Link Cited by: §1, §1, §3.1, §5, §5, Remark 5, Remark 5.
  • P. R. Milgrom (1979) A convergence theorem for competitive bidding with differential information. Econometrica 47 (3), pp. 679–688. External Links: ISSN 00129682, 14680262, Link Cited by: §1.
  • P. Milgrom and C. Shannon (1994) Monotone comparative statics. Econometrica 62 (1), pp. 157–180. Cited by: §1, §4.1.
  • D. Rosenberg and N. Vieille (2019) On the efficiency of social learning. Econometrica 87 (6), pp. 2141–2168. External Links: Document, Link Cited by: §6.
  • L. Smith and P. Sørensen (2000) Pathological outcomes of observational learning. Econometrica 68 (2), pp. 371–398. Cited by: §1, §1, §3.2, footnote 14.
  • L. Smith and P. Sørensen (2020) Rational social learning with random sampling. Note: unpublished External Links: Link Cited by: §6, footnote 7.

Appendices

Appendix A contains the proofs for Theorems 13. Appendix B contains the proofs of Proposition 1 and Lemma 2 (which proves Proposition 2). Appendix C contains the proof of Lemma 1 and Remark 5 on cascade vs. diffusion utility .

Appendix A Backbone Results

In this section, we prove our three theorems in the following setting, which is more general than that described in the main text:

  • The action space and signal space (A,𝒜),(S,𝒮) are standard Borel spaces;

  • The state space Ω is equipped with a metric d and its Borel sigma-algebra, (Ω), such that (Ω,d) is a sigma-compact Polish space;262626That is, (Ω,d) is a complete and separable metric space that can be represented as a countable union of compact sets.

  • The utility function u(a,ω) has absolute value uniformly bounded by u¯ and it is pointwise equicontinuous when regarded as a collection of functions of ω indexed by a; moreover, for every belief (Borel probability distribution over Ω), there exists an optimal action;

  • The information/signal structure F(|ω) is a Markov kernel from (Ω,(Ω)) to (S,𝒮) that is continuous in ω in the total variation (TV) sense;

  • The network structure is given by Q(Qn)n, where each Qn is a probability measure over all neighborhoods, i.e., all subsets of {1,2,,n1}, independent across n, independent of the state ω, and independent of any private signals.

When Ω is countable as in the main text, we endow it with the discrete metric so that the sigma-compactness and continuity requirements are trivially satisfied.

Discontinuous utilities.

While we make a continuity assumption on preferences, our main results hold for utilities satisfying the following condition that permits discontinuities (cf.  Remark 1):

Condition 1.

There is a countable partition of Ω into Borel sets Bi and pointwise equicontinuous functions vi:Ω uniformly bounded by u¯ such that vi|Bi=u.

To obtain our results for such utilities, we can define a new state space Ω~ as a disjoint union: Ω~:=Ωi, where each Ωi is a copy of Ω. Choose any metric on Ω~ that induces the disjoint union topology. Define a utility u~ on Ω~ by u~|Ωi:=vi for each i. It follows that Ω~ is sigma-compact Polish and u~ is pointwise equicontinuous and uniformly bounded. The information structure is defined such that on each Ωi it is the same as before. Using our results for the new setting, one can deduce Theorems 13 for the original setting.272727More specifically, the results in the original setting are equivalent to the corresponding results in the new setting restricted to priors/beliefs that put zero probability on the added states Ωi\Bi. We can use our methodology to derive Theorems 13 in the new setting for such restricted beliefs.

A.1 Overarching Probability Space and Beliefs

We now formalize the overarching probability space over all realizations of the state, signals, observation neighborhoods, and actions. We also define formal objects corresponding to agents’ social and posterior beliefs and distributions of beliefs.

Overarching probability space.

Our probability space is constructed from three components: the Markov kernel F and probability space (Ω,(Ω),μ0); the network structure Q(Qn)n; each agent n’s strategy σn(|aBn,Bn) as a Markov kernel from (A|Bn|,𝒜|Bn|) to (A,𝒜) for each realization of neighborhood Bn.

Taken together, for the first n agents, we can define a probability space that describes the joint distribution of their neighborhoods, signals, actions, and the states. Since all these elements lie in standard Borel spaces, the Kolmogorov Extension Theorem guarantees existence of an overarching probability space (H,,) that is consistent with each finite probability space (i.e., up to each agent n). We suppress the dependence of on σ and μ0.

Beliefs.

Given this overarching probability space, agent n’s social belief (i.e., her belief after observing her neighbors and their actions, but before observing her private signal) is (|aBn,Bn) and her posterior belief is (|aBn,Bn,sn). These beliefs are well-defined because, as a countable product of standard Borel spaces, the overarching probability space is a standard Borel space, and hence there exist regular conditional probabilities (Durrett, 2019, Theorem 4.1.17).

Distribution of beliefs.

We denote by ΔΩ the space of beliefs (Borel probability measures on Ω) equipped with the Prohorov metric, and by ΔΔΩ the space of belief distributions (Borel probability measures on ΔΩ) also equipped with the Prohorov metric.

The social belief of agent n, μn, as a regular conditional probability, can be regarded as a measurable function from (H,,) to (ΔΩ,(ΔΩ)); see Crauel (2002, Remark 3.20). As Ω is a Polish space, so is ΔΩ. We define agent n’s distribution of social beliefs, φn, as the push-forward measure of μn. Hence, φnΔΔΩ since it is by definition a Borel probability measure on ΔΩ.

A.2 Space of Bayes-Plausible Belief Distributions is Compact

Given a prior μ0ΔΩ and a strategy profile σ, any agent’s belief distribution φΔΔΩ must be Bayes plausible:

Aμ(A)dφ(μ)=μ0(A),A(Ω). (2)

Let ΦBPΔΔΩ be the set of Bayes-plausible belief distributions; note that we suppress the dependence of ΦBP on μ0.

Our goal is to establish (Lemma 4 below) that even though the set of belief distributions ΔΔΩ need not be compact, the subset of Bayes-plausible distributions ΦBP is. A key step is the following lemma, which shows that any belief distribution φΦBP has to put a large probability on a compact subset of ΔΩ.

Lemma 3.

Let δ>0 and {Ωi}i be a sequence of compact sets with μ0(Ωi)1(δ2i)2, i. Defining Vδ:={μΔΩ:μ(Ωi)1δ2i,i}, it holds that:

  1. 1.

    Vδ is compact;

  2. 2.

    φ(μVδ)<δ, φΦBP.

Intuitively, in the lemma’s statement, the set Vδ contains all beliefs that put high probability on a set of states that the prior μ0 ascribes high probability to. The lemma concludes that the set Vδ is compact and that any Bayes-plausible belief distribution must put high probability on Vδ.

  • Proof.

    (Part 1) First, Vδ is closed. To see this, take any μkμ and μkVδ. Since each Ωi is compact (and thus closed), weak convergence implies

    lim supkμk(Ωi)μ(Ωi),i,

    which implies μ(Ωi)1δ2i. Thus, μVδ, and hence Vδ is closed.

    Next, the beliefs in Vδ are tight by definition. Hence, by Prohorov’s theorem, the closure of Vδ, which is Vδ itself, is compact.

    (Part 2) Note that φ(μVδ)=φ(i{μ(Ωic)>δ2i})iφ(μ(Ωic)>δ2i). For each i, we view μ(Ωic) as a non-negative random variable with distribution induced by φ. Since φ is Bayes plausible, 𝔼φ[μ(Ωic)]=μ0(Ωic)(δ2i)2, which implies (using Markov’s inequality) that φ(μ(Ωic)>δ2i)<δ2i. This implies that φ(μVδ)<iδ2i=δ. ∎

Given Lemma 3, we can use Prohorov’s theorem again to show:

Lemma 4.

ΦBP is compact.

  • Proof.

    First, we prove that ΦBP is closed. Take any φkφ and φkΦBP, and want to show that φΦBP, i.e., 𝔼φ[μ(W)]=μ0(W),W(Ω).

    Take any open set W(Ω). For any μkμ, it holds that μ(W)lim infμk(W). In other words, μ(W) (as a function of μ) is lower semi-continuous. By properties of weak convergence, it follows that 𝔼φ[μ(W)]lim inf𝔼φk[μ(W)]=μ0(W). That is, the mean measure of φ ascribes a smaller probability than μ0 to any open set.

    Now observe that WcxWcB1/n(x) for any n. Hence,

    𝔼φ[μ(Wc)]limn𝔼φ[μ(xWcB1/n(x))]limnμ0(xWcB1/n(x))=μ0(Wc),

    where the second inequality is from the previous result applied to open sets xWcB1/n(x), and the last equality follows from Wc=nxWcB1/n(x) (and this equality holds because Wc is closed). Therefore, 𝔼φ[μ(W)]=μ0(W).

    Since 𝔼φ[μ] and μ0 agree on all open sets, and open sets generate (Ω), 𝔼φ[μ] and μ0 agree on all sets in (Ω). This establishes that φΦBP.

    Finally, Ω being sigma-compact implies that for any δ, there is an increasing sequence of compact sets {Ωi}i such that Ω=iΩi, and this sequence {Ωi} satisfies the hypotheses in Lemma 3. The lemma guarantees that there is a compact set Vδ such that φ(Vδ)<δ for all φΦBP, hence ΦBP is tight. Prohorov’s theorem now implies that the closure of ΦBP, which is ΦBP itself, is compact. ∎

A.3 Continuity of Various Functions

We next define some functions of interest, some of which were already defined in the main text but are now defined for the more general setting considered in the appendix.

Let u(μ) be the expected utility that an agent can get at belief μ:

u(μ):=supaAΩu(a,ω)dμ(ω).

Let uF(μ) be the expected utility that an agent can get at belief μ, if she can choose an action after observing her private signal:

uF(μ):=supβ:SAΩSu(β(s),ω)dF(s|ω)dμ(ω).

Finally, let u(μ) be the full information utility at μ:

u(μ):=ΩsupaAu(a,ω)dμ(ω).

Our continuity assumptions on the utility function and the information structure allow us to prove:

Lemma 5.

u,uF,u are continuous in μ.

To prove Lemma 5, we use Theorem 2.2.8 in Bogachev (2018), which we restate without proof for our context as the following claim:

Claim 1.

Let μkμ. If Γ is a uniformly bounded and pointwise equicontinuous family of functions on Ω, then

limksupfΓ|ΩfdμkΩfdμ|=0.
  • Proof of Lemma 5.

    By assumption, Γ:={u(a,ω)}aA, viewed as a family of functions of ω indexed by a, is uniformly bounded and pointwise equicontinuous.

    Consider the function u. Since the supremum of the pointwise equicontinuous functions u(ω):=supau(a,ω) is continuous in ω, the definition of weak convergence implies that u(μ) is continuous in μ.

Now consider the function u. Its continuity follows from

|u(μk)u(μ)|=|supfΓΩfdμksupfΓΩfdμ|supfΓ|ΩfdμkΩfdμ|,

which converges to 0 as μkμ, by Claim 1.

Lastly, suppose we establish that ΓF:={Su(β(s),ω)dF(s|ω)}β:SA, as a family of functions of ω indexed by β, is pointwise equicontinuous.282828Here we assume β are (measurable) pure strategies for notation clarity. The same argument works for mixed strategies, in which case β would be Markov kernels. Then, as ΓF is uniformly bounded, Claim 1 implies that uF(μ) is continuous, proving the lemma.

To establish the pointwise equicontinuity of ΓF, observe that ω,ω and β,

|Su(β(s),ω)dF(s|ω)Su(β(s),ω)dF(s|ω)| (3)
|S(u(β(s),ω)u(β(s),ω))dF(s|ω)|+|Su(β(s),ω)dF(s|ω)Su(β(s),ω)dF(s|ω)|.

Fix any ω and any ε>0. Since {u(a,ω)}aA is pointwise equicontinuous, there exists δ1 such that d(ω,ω)<δ1 implies the first term on the right-hand side of inequality (3) to be smaller than ε/2 (regardless of β(s)). The second term is smaller than 2u¯dTV(F(|ω),F(|ω)) (where TV represents total variation), and by the continuity assumption of the information structure, there exists δ2>0 such that d(ω,ω)<δ2 implies dTV(F(|ω),F(|ω))<ε/4u¯. Therefore, if d(ω,ω)<min{δ1,δ2}, then the right-hand side of inequality (3) is less than ε (regardless of β(s)). It follows that ΓF is pointwise equicontinuous. ∎

Now define the utility improvement I(μ) and the utility gap G(μ) at μ as:

I(μ):=uF(μ)u(μ),G(μ):=u(μ)u(μ).

By Lemma 5, I(μ) and G(μ) are continuous. Lastly, with an abuse of notation, define u(φ):=𝔼φ[u(μ)], I(φ):=𝔼φ[I(μ)], and G(φ):=𝔼φ[G(μ)] as the corresponding functions over distributions of beliefs. Since u(μ), I(μ), and G(μ) are continuous, so are u(φ), I(φ), and G(φ).

We note that a belief μ is stationary if and only if I(μ)=0, and a belief μ has adequate knowledge if and only if G(μ)=0. To confirm these points, consider stationary beliefs. If there is an action that is a.s. optimal regardless of the signal, then I(μ)=0. Conversely, if no action is a.s. optimal regardless of the signal, then for any action there is a positive-probability set of signals for which that action is strictly suboptimal; hence uF(μ)>u(μ), and I(μ)>0. The argument for adequate knowledge beliefs is similar.

A.4 Proofs for Backbone Results

Logically, Theorem 3 Theorem 1 Theorem 2. So we prove the results in that order.

We prove the result in two steps. In Step 1 below, we prove that if agent n’s social belief distribution φn, which is her belief distribution incorporating the observation of her neighborhood’s actions but not her private signal, is not close to being supported on only stationary beliefs, then her utility 𝔼σ,μ0[un], which is the ex-ante expected utility under equilibrium σ after observing the private signal, improves from u(φn) by some positive amount bounded away from zero. In Step 2 below, we use the expanding observations assumption to establish that this minimum improvement propagates through the network until eventually agents obtain at least arbitrarily close to their cascade utility level.

Step 1: Recall the set of Bayes-plausible belief distributions that are supported by stationary beliefs, ΦS:={φΦBP:I(φ)=0}, and the cascade utility, u:=infφΦSu(φ).

Take any ε>0, and let (ΦS)ε denote the ε-neighborhood of ΦS. An agent n’s belief distribution φn must be Bayes plausible, so φnΦBP. Since u(φ) is uniformly continuous on ΦBP (as u(φ) is continuous, and ΦBP is compact), if φn(ΦS)ε, then u(φn)uγ(ε) for some γ() such that γ(ε)0 when ε0. If, on the other hand, φnΦ+:=ΦBP\(ΦS)ε, then I(φn)>0 because φn puts positive probability on {μ:I(μ)>0}. Since (ΦS)ε is open, Φ+ is a closed subset of a compact set ΦBP; hence Φ+ is compact, and since I(φ) is continuous, it attains a minimum over Φ+ at some φ¯Φ+. Thus, if φnΦ+ the agent obtains an improvement I(φn)I(φ¯)>0.

Step 2: We will argue that for any ε>0, 𝔼σ,μ0[un]uγ(ε) once n is large enough. Since ε is arbitrary, taking ε0 implies lim infn𝔼σ,μ0[un]u, which completes the proof.

For a given ε>0, let δ=I(φ¯)4u¯>0, let N0=1, and define Nk for k=1,2, sequentially such that for all nNk, Qn(maxbBnb<Nk1)<δ. Expanding observations ensures that such Nk exist.

We claim that, for any agent nNk, 𝔼σ,μ0[un]αk:=min{uγ(ε),kI(φ¯)2u¯}. Since α0=u¯, clearly 𝔼σ,μ0[un]α0 for any nN0. Suppose the claim holds for all agents nNk1. Take any agent nNk. Agent n’s neighborhood is drawn independently of everything that has happened before, so conditional on agent n observing an agent nNk1, even without her private signal agent n can achieve a utility of at least αk1 by imitating agent n. Hence, u(φn)(1δ)αk1+δ(u¯).292929 If agents do not observe the identities associated with the observed actions of their predecessors, an agent can uniform-randomly select one of the actions they observe to imitate. So long as the “induced network structure” (i.e., a network structure (Q~n) wherein each Q~n is defined by first drawing a neighborhood Bn from Qn and then uniform-randomly drawing a single agent from Bn) satisfies expanding observations, the current proof goes through without change using the induced network structure. If φn(ΦS)ε, then by definition u(φn)uγ(ε), and thus 𝔼σ,μ0[un]u(φn)uγ(ε)αk. If φn(ΦS)ε, then agent n can improve her utility by at least I(φ¯), and so

𝔼σ,μ0[un] (1δ)αk1+δ(u¯)+I(φ¯)
αk1+I(φ¯)2(because αk1u¯ and δ=I(φ¯)4u¯)
αk.

Since the definition of αk implies that there is a finite K such that for all kK, αk=uγ(ε), it follows that for all nNK, 𝔼σ,μ0[un]uγ(ε). ∎

  • Proof of Theorem 1.

    The “only if” direction is straightforward. If there is a stationary belief without adequate knowledge, then when the prior is that belief there is an equilibrium where each agent ignores her signal and action history and obtains a utility that is strictly below the full-information utility level.

    For the ”if” direction, fix any prior μ0 and equilibrium σ. Since all stationary beliefs have adequate knowledge, I(μ)=0 implies G(μ)=0. Thus, for any φΦS, φ({μ:I(μ)=0})=φ({μ:G(μ)=0})=1, which implies G(φ)=u(φ)u(φ)=0. Moreover, because μ0 is the mean measure of φ,

    u(φ)=𝔼φ[Ωsupau(a,ω)dμ]=Ωsupau(a,ω)dμ0=u(μ0),

    which implies u(φ)=u(μ0). As a result, u(μ0)=infφΦSu(φ)=u(μ0). It follows from Theorem 3 that lim infn𝔼σ,μ0[un]u(μ0). Since 𝔼σ,μ0[un]u(μ0) for all n, it further follows that 𝔼σ,μ0[un]u(μ0). As μ0 and σ are arbitrarily, we have adequate learning. ∎

Next we state and prove a more general version of Theorem 2. For any n, define Ωa1,a2n:={ω:u(a1,ω)u(a2,ω)>1n}.

Theorem 2.

Excludability implies adequate learning at every choice set. There is inadequate learning for choice set {a1,a2} if Ωa1,a2 is not distinguishable from Ωa2,a1n for some n.

Note that when Ω is finite, or the utility difference between any pair of actions is bounded away from zero, a failure of excludability is equivalent to the condition for necessity in the theorem holding for some a1,a2. Hence Theorem 2 is implied by Theorem 2.

  • Proof of Theorem 2.

    (First statement) First note that excludability (under the full choice set A) implies excludability under any choice subset AA. So we fix an arbitrary AA and show that excludability under that subset implies adequate learning at that choice subset. In what follows, the domain of actions should be understood as A, and we denote a typical element by a.

Theorem 1 implies that we need only show that any μΔΩ with inadequate knowledge is not stationary. So take any μΔΩ with inadequate knowledge and any ac(μ). Since there is inadequate knowledge, μ(aΩa,a)>0, i.e., there is a positive measure of states where a is not optimal. The continuity of u(a,ω)u(a,ω) implies that Ωa,an are open sets for any a and n. Since Ω is Polish, it is second-countable and hence has a countable basis. Therefore, each open set Ωa,an, and hence the open set aΩa,a(=anΩa,an), is a union of countably many basic open sets. Since μ(aΩa,a)>0, at least one basic open set contained in Ωa,an for some a and n has strictly positive measure, i.e., μ(Ωa,an)>0.

Now denote μ():=μ(|Ωa,aΩa,an) as the corresponding conditional probability. Since Ωa,a is distinguishable from Ωa,a by excludability, so is Ωa,an.303030In fact, excludability is equivalent to: Ωa1,a2n is distinguishable from Ωa2,a1 for all a1,a2 and n. Therefore, for any ε>0 there exists a set of signals S such that Prμ(S)>0 and μs(Ωa,an)>1ε for all sS. The utility improvement upon observing any sS by switching from a to a is therefore bounded below by (1n(1ε)2u¯ε)μs(Ωa,aΩa,an), as the expected improvement on Ω\(Ωa,aΩa,an) is nonnegative. For small ε>0, 1n(1ε)2u¯ε>0. Furthermore, integrating μs(Ωa,aΩa,an) over sS yields Prμ(S)μ(Ωa,aΩa,an)>0. Hence, the ex-ante improvement is bounded below by (1n(1ε)2u¯ε)Prμ(S)μ(Ωa,aΩa,an)>0. It follows that I(μ)>0, and thus μ is not stationary.

(Second statement) Suppose there are two actions a1,a2 and an n such that Ωa1,a2 is not distinguishable from Ωa2,a1n. This means there exists μΔ(Ωa1,a2Ωa2,a1n) with μ(Ωa1,a2)>0 such that μs(Ωa1,a2)1ε for some ε>0 and μ-a.e. s. Consider μΔ(Ωa1,a2Ωa2,a1n) with a small μ(Ωa1,a2)>0 such that μ(|Ωa1,a2)=μ(|Ωa1,a2) and μ(|Ωa2,a1n)=μ(|Ωa2,a1n). Under μ, upon observing signal s, the posterior on Ωa1,a2 satisfies

μs(Ωa1,a2)μs(Ωa2,a1n)=μs(Ωa1,a2)/μ(Ωa1,a2)μs(Ωa2,a1n)/μ(Ωa2,a1n)μ(Ωa1,a2)μ(Ωa2,a1n)1εεμ(Ωa2,a1n)μ(Ωa1,a2)μ(Ωa1,a2)μ(Ωa2,a1n)

for μ-a.e. s. Hence, by choosing μ so that μ(Ωa1,a2)μ(Ωa2,a1n) is arbitrarily small, the ratio μs(Ωa1,a2)μs(Ωa2,a1n) can be made arbitrarily small uniformly over s.

Under μ, after observing s, the expected improvement by switching from a2 to a1 is bounded above by 2u¯μs(Ωa1,a2)1nμs(Ωa2,a1n), which is strictly negative when μs(Ωa1,a2)μs(Ωa2,a1n) is small. Therefore, for μ-a.e. s, a2 is strictly better than a1, and thus μ is stationary for choice set {a1,a2}. However, since μ(Ωa1,a2)>0, the belief μ has inadequate knowledge. Theorem 1 implies there is inadequate learning for choice set {a1,a2}. ∎

Appendix B Applications

We now specialize to the main text’s setting: Ω is countable, endowed with the discrete metric, and F(|ω) are absolutely continuous with respect to each to other, and so there are densities f(|ω)>0.

B.1 SCD Preferences & DUB Information

Sufficiency follows directly from Theorem 2. For necessity, first observe that if the information structure fails DUB, then there exists some state ω such that ω is not distinguishable from its lower set (or from its upper set, which has a symmetric argument). Fix any pair of distinct actions a1 and a2, and define the following SCD preferences: for ω<ω, u(a1,ω)=1 and u(a2,ω)=0; for ωω, u(a1,ω)=0 and u(a2,ω)=1; and any other actions are strictly dominated. It follows that Ωa2,a1 is not distinguishable from {ω:u(a1,ω)u(a2,ω)>12}. By Theorem 2, there is inadequate learning when the choice is {a1,a2}, and since all other actions are strictly dominated, also for the full choice set A. ∎

B.2 Intermediate Preferences & Location-Shift Information

The proof of Lemma 2 is more involved than the intuition given in the main text using Figure 3, because in general one cannot explicitly identify the sequence of signals that establishes distinguishability of the relevant two sets.

We will use the following claim in proving Lemma 2. For any h,xd, let xh:=hx be the “signed distance” of x in direction h, i.e., between x and the hyperplane {z:hz=0}. Note that h is linear, so xxh=xhxh.

Claim 2.

If a standard density g is subexponential, then for any s¯ with s¯h>0, and ε(0,1), there is s with ss¯h1 such that:

  1. 1.

    sup{s:ssh1/s¯h}g(s)g(s)<ε; and

  2. 2.

    sup{s:0<ssh<1/s¯h}g(s)g(s)<2.

  • Proof.

    Suppose not, to contradiction. Then there exists s¯ with s¯h>0 and ε(0,1) with the following property: for every s with ss¯h1, we can find s with ssh>0 such that either (i) ssh1/s¯h and g(s)g(s)ε, or (ii) 0<ssh<1/s¯h and g(s)g(s)2. For an arbitrary choice of s given s, we define ks:=ssh. That means, for each s with ss¯h1, we have ks>0 and a signal s with ssh=ks such that either (i) g(s)g(s)εεkss¯h (because kss¯h1), or (ii) g(s)g(s)2>εkss¯h (because ε<1).

    We construct a sequence of signals (si)i=1. First, take any s1 such that s1s¯h=1. Then, for all i>1, take any si given si1 as explained in the previous paragraph. Note that for all i, sisi1h=ksi1, so sih=(s¯h+1)+j=1i1ksj.

    First, suppose that i=1ksi=, so that limisih=. It holds that for all si, g(si)g(s¯)g(s1)g(s¯)ε(ksi1++ks1)s¯h=g(s1)g(s¯)ε(sihs¯h1)s¯h, which in turn implies that

    (sihs¯h1)s¯hlog(ε)+log(g(s1))log(g(si)). (4)

    However, since g is subexponential, and sihsih, there is p>1 such that for all large enough i,

    log(g(si))<(sihh)p. (5)

    The left-hand side of inequality (4) is linear in sih while the right-hand side of inequality (5) has exponent p>1, so for large enough i these inequalities are in contradiction.

    Next, suppose instead limisih<. Then there is N such that for all iN, we have ksi<1/s¯h and thus g(si+1)g(si)2. It follows that limig(si)g(sN)limi2iN=. This contradicts the boundedness of g (being a density, g is bounded because it is uniformly continuous). ∎

Without loss, we only prove that {ω:hω>c} is distinguishable from {ω:hω<c}.

We use Claim 2 iteratively to construct a signal sequence (si)i=1. Choose any s1 with s1h>0, and for i>1, choose any si such that sisi1h1 that satisfies (i) sup{s:ssih1/si1h}g(s)g(si)<1i1 and (ii) sup{s:0<ssih<1/si1h}g(s)g(si)<2. This construction is well-defined by Claim 2, with limisih=.

As noted after Definition 1, it is sufficient to prove that any ω¯{ω:hω>c} is distinguishable from {ω:hω<c}.313131We note that this uses the assumption of countable states. So take any such ω¯ and μ with μ(ω¯)>0. Define s¯i:=si+ω¯. It follows that for all i,

ωω¯h<0f(s¯i|ω)f(s¯i|ω¯)=g(s¯iω)g(s¯iω¯)=g(si+(ω¯ω))g(si)<2, (6)

and

ωω¯h1si1hf(s¯i|ω)f(s¯i|ω¯)=g(si+(ω¯ω))g(si)<1i1, (7)

and thus,

μ({ω:hω<c}|s¯i)μ(ω¯|s¯i)ωω¯h<0μ(ω)f(s¯i|ω)μ(ω¯)f(s¯i|ω¯)<1i1ωω¯h1/si1hμ(ω)μ(ω¯)+21/si1h<ωω¯h<0μ(ω)μ(ω¯). (8)

The last expression can be taken arbitrarily small because si1h as i.

It remains only to show that the above argument holds for a positive measure of signals rather than just a single s¯i. Since g is uniformly continuous, there is a neighborhood of s¯i, say S¯i, over which (6) and (7) hold with slightly relaxed bounds; for instance, the bounds can be relaxed to 4 and 2/(i1), respectively. This establishes the analog of inequality (8) for all signals in S¯i with the relaxed bounds. Since we have assumed g()>0, each S¯i has positive measure, so we conclude ω¯ is distinguishable from {ω:hω<c}. ∎

Appendix C Other Material

  • Proof of Lemma 1.

    As noted before the lemma, Ω is distinguishable from Ω′′ if and only if each ωΩ is distinguishable from Ω′′. So fix any ωΩ.

    We first prove that if the lemma’s condition holds, then ω is distinguishable from Ω′′. Take any probability measure μΔ({ω}Ω′′) such that μ(ω)>0. By assumption, for any ε>0 there exists a positive-probability set of signals S such that f(s|ω′′)f(s|ω)<ε,ω′′Ω′′,sS. It follows that for all sS,

    μ(ω|s)=f(s|ω)μ(ω)ω~{ω}Ω′′f(s|ω~)μ(ω~)=μ(ω)μ(ω)+ω~Ω′′f(s|ω~)f(s|ω)μ(ω~)>μ(ω)μ(ω)+ε.

    Since for any ε>0 we can find a positive-probability set of signals S satisfying the above inequality, we conclude that for any ε>0, Prμ(s:μs(Ω)>1ε)>0.

    We next prove that if ω is distinguishable from Ω′′, and Ω′′ is finite, then the lemma’s condition holds. Consider any μ uniformly distributed over {ω}Ω′′. The distinguishability of ω from Ω′′ implies that for every ε>0 there is a positive-probability set of signals S such that sS we have ω~Ω′′f(s|ω~)f(s|ω)<ε, and so f(s|ω~)f(s|ω)<ε for every ω~Ω′′. ∎

Remark 5.

Lobel and Sadler’s (2015) definition of diffusion utility is tailored to their binary-binary model. In general, we can define it as the highest utility an agent can obtain from any Bayes-plausible distribution of beliefs that is supported on the set of feasible posteriors (i.e., those available under the given information structure and the prior); call the corresponding signal structure the expert signal structure.

Let us now argue that diffusion utility is lower than cascade utility. Notice that diffusion utility must be lower than first drawing a posterior from an arbitrary Bayes-plausible distribution of stationary beliefs and then drawing a signal from the expert signal structure, because this “combined” signal structure is Blackwell more informative than just the expert signal structure. But in the combined structure, the expert signal has no value by definition of stationary beliefs, and so the combined signal structure provides a utility equal to that from the (arbitrary Bayes-plausible) distribution of stationary belief distributions.

Diffusion utility and cascade utility coincide with two states and two actions, as noted by Lobel and Sadler’s (2015). But adding even one action can break this coincidence, e.g., if the third action is a “safe” action—one that is optimal only for some interval of interior beliefs—that shrinks the set of stationary beliefs. For a starker example, recall Example 1 with Ω={0,1}, A=[0,1], and u(a,ω)=(aω)2. Any nontrivial information structure leads to learning, with the stationary beliefs being just 0 and 1. So the cascade utility is the full-information utility of 0, whereas diffusion utility will be strictly lower absent unbounded beliefs.

Supplementary Appendices

Appendix SA.1 Belief Convergence

This section elaborates on Remark 3. Our discussion in this section focuses on deterministic networks.

One may expect the social belief to be eventually close to the stationary set with high probability: after all, when an agent’s social belief is not close to the stationary set, her private information gives her a welfare improvement bounded away from zero; expanding observations should propagate these improvements, which implies (since utility is bounded) that they must eventually vanish. However, the following is a counterexample.323232Absent expanding observations, there are trivial counterexamples using the empty network.

Example SA.1.

Consider binary states with a uniform prior, binary signals with symmetric precision (less than 1), and binary actions with simple utility. The network is as follows: agents 1 and 2 observe no one; for odd n3, agent n observes agent n2; for even n>3, agent n observes agent n1 and agent 2. So there is expanding observations. In this network, the odd agents form an immediate-predecessor network and there is an equilibrium where a cascade along this subsequence starts from agent 3.

Now consider even agents. Consider the positive-probability event in which agents 1 and 2 take different actions. An even agent n>3 observes agents n1 and 2, which, given the equilibrium behavior of odd agents, is equivalent to observing agents 1 and 2. So the social belief of every even agent n>3 equals the prior, which is bounded away from the stationary set.333333The example illustrates that with positive probability social beliefs may not eventually converge to the set of stationary beliefs. But the point also holds for posterior beliefs, not just social beliefs. For simplicity, consider the same example but with an additional signal that is uninformative. Call the two actions a and b. Consider an equilibrium in which the first agent plays a upon receiving the uninformative signal, while the second agent plays b upon receiving the uninformative signal. Then, in the event that the first agent plays b and the second agent plays a, the path of social beliefs for agents n3 is identical to the example above: odd agents are in a cascade, while even agents’ social belief is just the prior. With positive probability, an even agent will now receive an uninformative signal, whereafter her posterior belief lies outside the stationary set.

The “problem” in Example SA.1 is that even though each of the even agents (n>3) is getting a welfare improvement bounded away from zero, these improvements are not passed on to any future agents, and all future even agents continue to have social beliefs bounded away from the stationary set. In other words, expanding observations is not enough to validate the intuition described before the example. The following proposition identifies a reasonable condition on the network that is sufficient.

Proposition SA.1.

Assume there exist finitely many subsequences of agents {nk,j}k=1Nj (j=1,,J<, 1Nj) such that agent nk,j observes nk1,j, and every agent in society is in at least one of the subsequences. Then, for all ε>0, limnφn(μnSε)=1.

The proposition’s assumption encompasses canonical examples like the complete network and the k-immediate-predecessor networks (i.e., every agent observes the last k agents) for any k1. But it rules out any network in which infinite number of agents are not observed by any of their successors, which explains why it does not apply to Example SA.1.

  • Proof of Proposition SA.1.

    Along any subsequence j, u(φnk,j)u(φnk1,j)+I(φnk1,j) by the improvement principle, given that nk1,j is observable to nk,j. It follows that k=1NjI(φnk,j)2u¯. Hence, society’s total improvement is bounded: nI(φn)2u¯J.

    Now fix any ε,δ>0. Consider Vδ/2 defined in Lemma 3. The lemma established that Vδ/2 is compact and φ(μVδ/2)<δ/2,φΦBP. Since Sε is open, K:=(Sε)cVδ/2 is compact. Next we argue (μnKi.o.)=0. Suppose, to contradiction, (μnKi.o.)>0. Then n(μnK)= by the Borel-Cantelli lemma. Since K is compact and I()>0 on K, I() achieves its minimum in K at some μ¯K with I(μ¯)>0. So the total improvement is nI(φn)I(μ¯)n(μnK)=, which contradicts nI(φn)2u¯J.

    Observe that (μnKi.o.)=0 implies φn(μnK)<δ/2 for all large n. Therefore, for all large n, φn(μn(Sε)c)φn(μnK)+φn(μnVδ/2)<δ. We conclude that for all ε>0, limnφn(μnSε)=1. ∎

Remark 6.

If ΔΩ is compact (e.g., Ω itself is compact), we can replace Vδ/2 in the proof with ΔΩ, so that K=(Sε)c. Then the argument in the proof’s second paragraph shows that (μn(Sε)ci.o.)=0, i.e., the social belief converges to the stationary set almost surely rather than only in probability.

Appendix SA.2 ε-Excludability

This section elaborates on Remark 4. Say that for any ε(0,1/2) a set of states Ω is ε-distinguishable from Ω′′ if for any μΔ(ΩΩ′′) with μ(Ω)>ε, there is a positive-measure set of signals S such that μ(Ω|s)>1ε for all sS. A utility function and an information structure jointly satisfy ε-excludability if Ωa1,a2 and Ωa2,a1 are ε-distinguishable from each other, for any pair of actions a1,a2. Note that ε-excludability implies ε-excludability for all ε>ε, and excludability is equivalent to ε-excludability for all ε>0.

Proposition SA.2.

Let Ω be finite. For all ε(0,1/2), ε-excludability implies that in any equilibrium σ, liminfn𝔼σ,μ0[un]u(μ0)2u¯ε1ε|Ω|.

Before proving Proposition SA.2, we give an example illustrating the result’s use.

Example SA.2.

There are three states, ω{1,2,3}, SCD preferences, and Laplace information:

f(s|ω)=12bexp(|sω|b),

where b>0 is a scale parameter; a smaller b corresponds to more precise information.

It is straightforward to verify that no two states can be distinguished from each other.343434For any pair of states ωω, and any signal s, the likelihood ratio f(s|ω)/f(s|ω)exp(2/b). Therefore, not every stationary belief has adequate knowledge (so long as preferences are nontrivial), and by Theorem 1 there is inadequate learning.

Nonetheless, we claim that ε-excludability holds for any ε such that ε>11+exp(12b). To see this, observe that since the information structure has MLRP and preferences satisfy SCD, we can focus on ε-distinguishing state 3 from 2 (or, equally, 2 from 1).353535By MLRP, only arbitrarily large signals can distinguish a state from a lower state, and for large s the likelihood ratio f(s|3)/f(s|2)<f(s|3)/f(s|1), so considering adjacent states is sufficient for ε-excludability. When ε>11+exp(12b), we have ε1εexp(1/b)>1εε, so there exist signals that move the prior (0,1ε,ε) to a posterior of at least 1ε on state 3, which implies ε-distinguishability of state 3 from 2.

Proposition SA.2 implies that in any equilibrium, liminf𝔼σ,μ0[un]u(μ0)6u¯exp(12b). This quantitative welfare bound yields, in particular, convergence to the full-information utility u(μ0) as b0.

  • Proof of Proposition SA.2.

    Take any stationary belief μ, and let a be an optimal action at belief μ. For each state ω, take any aωc(ω), and consider μω():=μ(|{ω}Ωa,aω). If μω(ω)ε, then μ(ω)ε, so (u(aω,ω)u(a,ω))μ(ω)2u¯ε.

    Consider the other case of μω(ω)>ε. For any sS, because u(a,ω)u(aω,ω)0 for each ωΩa,aω, and μ is stationary,

    ω{ω}Ωa,aω(u(a,ω)u(aω,ω))μ(ω|s)ωΩ(u(a,ω)u(aω,ω))μ(ω|s)0.

    Then,

    (u(aω,ω)u(a,ω))μω(ω|s) ωΩa,aω(u(a,ω)u(aω,ω))μω(ω|s)
    2u¯(ωΩa,aωμω(ω|s))=2u¯(1μω(ω|s)).

    By ε-excludability, there exists a positive-measure set of signals S such that, for any sS, μω(ω|s)>1ε, which implies that u(aω,ω)u(a,ω)2u¯ε1ε.

    In either case (μω(ω)ε or μω(ω)>ε), we have (u(aω,ω)u(a,ω))μ(ω)2u¯ε1ε. Since Ω is finite,

    ωΩ(u(aω,ω)u(a,ω))μ(ω)2u¯ε1ε|Ω|.

    Namely, the utility gap u(μ)u(μ)2u¯ε1ε|Ω|, for any stationary belief μ.

    Finally, for any φΦS,

    u(μ0)u(φ)=𝔼φ[u(μ)u(μ)]2u¯ε1ε|Ω|.

    By taking infimum of u(φ) across φΦS, we obtain u(μ0)u(μ0)2u¯ε1ε|Ω|, and subsequently by invoking Theorem 3, we conclude that in any equilibrium σ, liminfn𝔼σ,μ0[un]u(μ0)2u¯ε1ε|Ω|. ∎

Appendix SA.3 Details on Example 2

For Example 2, we show here how to construct a full-support prior such that the posterior probability is uniformly bounded away from 1 across signals and states. Take any prior μ such that for some c>0, min{μ(n1)μ(n),μ(n+1)μ(n)}>c for all n (e.g., a double-sided geometric distribution). Denoting the posterior after signal s by μs, the posterior likelihood ratio satisfies

μs({n1,n+1})μs(n)=f(s|n1)f(s|n)μ(n1)μ(n)+f(s|n+1)f(s|n)μ(n+1)μ(n)>c(f(s|n1)f(s|n)+f(s|n+1)f(s|n)).

As the last expression is the sum of a strictly positive decreasing function of s and a strictly positive increasing function of s, it is bounded away from 0 in s. The bound is independent of n because normal information is a location-shift family of distributions. Therefore, the posterior likelihood ratio is uniformly bounded away from 0, and hence, the posterior μs(n) is uniformly bounded away from 1.

Appendix SA.4 Learning at a Fixed Prior

For tractability, our discussion in this section assumes a complete network.

The issue.

The definition of adequate learning we adopted in Section 2 requires that there is learning for all priors. In general, one may be interested in whether there is adequate learning at some given (full-support) prior.363636To be complete: there is adequate learning at prior μ0 if for every equilibrium strategy profile σ, 𝔼σ,μ0[un]u(μ0). Adequate (or inadequate) learning at a prior for a choice set is defined analogously. Of course, our sufficient conditions—e.g., Theorem 2’s excludability—remain sufficient, but fixing a prior raises the question of whether the conditions are necessary. With only two states, the distinction between some prior and all priors is immaterial: if adequate learning fails at any prior, then the only adequate-knowledge beliefs are those with certainty on some state, and there is an open ball of stationary beliefs around certainty on one of the states; hence, given any full-support prior, there cannot be a belief path that converges to certainty on that state, implying a failure of adequate learning at all full-support priors.

However, with multiple states, a failure of adequate learning at some prior does not imply an open ball of stationary beliefs around any adequate-knowledge belief. To illustrate, consider Figure SA.1. Action a is optimal at states 2 and 3 while a¯ is optimal at state 1. Adequate learning fails when the prior is μ because μ, which has support {1,2}, is stationary but has inadequate knowledge.373737In this example, preferences have SCD. But, consistent with Proposition 1, DUB is violated because state 2 is not distinguishable from state 1 or state 3. Yet there is no open ball of stationary beliefs: no full-support belief is stationary because the optimal actions are distinct at the extreme states 1 and 3, and the extreme states are distinguishable from their complements. This raises the possibility that there is learning at some—or even all—full-support priors, with on-path sequences of beliefs (which necessarily have full support at every finite time) converging almost surely to adequate-knowledge beliefs without ever hitting any stationary belief (all of which have non-full-support).

Figure SA.1: Preference regions among the actions a¯ and a shaded in blues. Under belief μ, posteriors are given by the black curve, while under belief μ, posteriors are given by the red line.

Some partial analysis.

While we are unable to characterize learning at a fixed prior in general, we provide some partial analysis below that we hope will be useful for future research. We focus on obtaining an analog of Theorem 2 for a fixed prior.

First, we provide a lemma (Lemma SA.1) that connects the existence of certain on-path histories to distinguishability. Second, we conjecture a result (SA.1) and show in Proposition SA.3 that, if the result is true, it combines with the lemma to deliver a fixed-prior analog of Theorem 2. Third, we show that the conjecture is true in a class of problems (Claim SA.1).

Lemma SA.1.

Take an aribtrary ωΩ and set of states ΩΩ\{ω}. State ω is distinguishable from Ω if there exist an equilibrium under a full-support prior and a history of actions h such that (h|ω)=0, and (h|ω) is bounded away from 0 across ωΩ.

The lemma ties the asymptotic probabilities of on-path histories to the information structure of an individual agent. The formal proof of the lemma is provided at the end of this appendix, but to see the intuition, suppose that the relevant h in the hypothesis of the lemma is some eventual herd on some action aA, i.e., h={hm,a,a,a,,} with hm some finite subhistory. If h has 0 probability in state ω, it must be because in an infinite number of periods, agents have positive probability of obtaining signals which overturn the herd on a, i.e., result in them taking some other action than a. However, the probability of this history is positive for states in Ω. This means that the probability of signals that overturn the herd must vanish over time at a fast enough rate in ω, but either not vanish or vanish at a slow enough rate in each ωΩ. In particular, there must exist overturning signals whose probability gets arbitrarily large in state ω relative to those in every ωΩ, which means ω is distinguishable from Ω.

Conjecture SA.1.

Take any a1,a2A, any full-support prior, and any equilibrium. If there is adequate learning at that prior, choice set {a1,a2}, and equilibrium,383838That is, under the given prior μ0 and equilibrium σ, 𝔼σ,μ0[un]u(μ0). then

h and ε>0:(h|ω)>ε,ωΩa1,a2. (SA.1)

The conjecture says that given any full-support prior, any binary choice set {a1,a2}, and any equilibrium in which there is adequate learning, we can find a single history that occurs with probability bounded away from 0 in all states in which a1 is strictly preferred (and analogously, a different history for the states in which a2 is strictly preferred). To appreciate the conjecture, let us focus for discussion on the case of finite states, nontrivial information, and nontrivial preferences. First note that if Ωa1,a2 is a singleton—as is the case with binary states—then it is straightforward that there is such a history, as there is a herd almost surely and every herd begins at some finite time. When Ωa1,a2 is not a singleton, given adequate learning, the same logic shows that for each state in Ωa1,a2, there is a history that has positive probability in that state, namely one with a herd on a1. But SA.1 demands more: a single history that has positive probability in all states in Ωa1,a2. Nonetheless, the conjecture seems intuitive: (up to tie-breaking issues) it would be surprising for every infinite history that has positive probability in some ωΩa1,a2 to have zero probability in some other ωΩa1,a2, given that agents’ have the same ordinal preferences over the binary actions in both ω and ω. For instance, consider a fully-informative information structure and any nontrivial preferences. Clearly, given any choice set {a1,a2}, there are only two possible histories: either a1 in every period or a2 in every period. The former has probability 1 in each ωΩa1,a2, and the latter has probability 1 in each ωΩa2,a1 and so SA.1 holds. Even though individuals’ private information distinguishes states perfectly, the public history does not.

Proposition SA.3.

If SA.1 is true, then not only does excludability imply adequate learning at every prior for every choice set, but moreover, if excludability fails, then there exists a choice set with inadequate learning at every full-support prior.

  • Proof.

    That excludability implies adequate learning at every prior for every choice set is implied by Theorem 2, with no need to invoke SA.1. So we only prove the second portion of the proposition, doing so by contraposition.

    To that end, assume that for every choice set, there is some full-support prior at which there is adequate learning in some equilibrium. For every binary choice set {a1,a2}, SA.1 implies the existence of a history h satisfying (SA.1) at the full-support prior and equilibrium at which there is adequate learning. Since there is adequate learning, eventually all agents must be taking a1 in h, which implies that (h|ω)=0 for each ωΩa2,a1. Then, taking Ω=Ωa1,a2 in Lemma SA.1 yields that ω is distinguishable from Ωa1,a2. Since a1,a2 and ωΩa1,a2 are arbitrary, there is excludability. ∎

We have not been able to establish SA.1 in general. However, we are able to establish it when preferences satisfy SCD and the information structure satisfies the strict MLRP (assuming, only for convenience, that the state space is discrete):

Claim SA.1.

Assume Ω. If preferences satisfy SCD and the information structure satisfies the strict MLRP, then SA.1 is true.

The proof is at the end of this appendix. Combining Claim SA.1 and Proposition SA.3, we see that under a complete network, the signal structure and preferences in Figure SA.1 entail inadequate learning at every full-support prior, such as μ in the figure. Note that the figure’s signal structure satisfies strict MLRP because the black curve in Figure SA.1 is concave vis-à-vis the 13 edge and approaches the 1 and 3 vertices.

Omitted Proofs

  • Proof of Lemma SA.1.

    Suppose not. Then there exist a belief μΔ(Ω{ω}) with μ(ω)>0 and a small ε>0 such that for almost every signal s, the posterior μ(ω|s)1ε. By taking the conditional distribution of μ on Ω, call it μ~, and z:=εμ(ω)1μ(ω)(0,1), we obtain for almost every s,

    Ωf(s|ω)dμ~(ω)zf(s|ω). (SA.2)

    Suppose there exist an equilibrium σ under a full support prior and history h such that (h|ω)=0, and (h|ω) is bounded away from 0 across ωΩ. Let an be the action taken by agent n along h and An:=A\{an}. Let (an|hn,ω):=Sσ(an|s,hn)f(s|ω)ds be the probability that agent n plays action an when the state is ω and the sub-history is hn. It holds that:

    n=1log(1z(An|hn,ω)) n=1log(1Ω(An|hn,ω)dμ~(ω))(using (SA.2))
    =n=1log(Ω(an|hn,ω)dμ~(ω))
    n=1Ωlog((an|hn,ω))dμ~(ω)(by Jensen’s inequality)
    =Ωn=1log((an|hn,ω))dμ~(ω)(by Tonelli’s theorem)
    =Ωlog(n=1(an|hn,ω))dμ~(ω)
    >(as log(h|ω) is bounded across ωΩ). (SA.3)

    Below we will invoke the fact that for arbitrary sequences (Sn) and (Tn) and constant c>0, if limnSnTn=c>0 and nSn<, then nTn<.393939For any c<c there exists N such that for all n>N, Sn/Tnc, or TnSn/c. So nTnnNTn+n>NSn/c<. Let Sn=log(1z(An|hn,ω)) and Tn=log(1(An|hn,ω)). Note that limnSnTn=z(0,1) because limx0log(1zx)log(1x)=z and (SA.3) implies limn(An|hn,ω)=0. The aforementioned mathematical fact implies that

    n=1log(1(An|hn,ω))>.

    As (an|hn,ω)=1(An|hn,ω), it further follows that n=1(an|hn,ω)>0, which contradicts (h|ω)=0. ∎

  • Proof of Claim SA.1.

    Take any information structure f with strict MLRP, and a utility function u that has SCD. Then, take an equilibrium σ under a full support prior μ and a binary choice set {a1,a2}. Since u has SCD, Ωa1,a2 and Ωa2,a1 are either an upper and lower set or the other way around. We consider the case that a2 is preferred in higher states and a1 is preferred in lower states. We omit an analogous proof for the other case.

We observe that (a1|hn,ω):=Sσ(a1|hn,s)f(s|ω)ds, the probability that agent n plays action a1 given any finite history hn, decreases in ω. First, the probability σ(a1|hn,s) decreases in s. The social belief μ(|hn) has full support, so strict MLRP of the information structure implies that s<s, the posterior μ(|hn,s) strictly monotone likelihood-ratio dominates μ(|hn,s). By Theorem 2 of Athey (2002),

D(s):=ω(u(a2,ω)u(a1,ω))μ(ω|hn,s)

is strictly single crossing in s, i.e., D(s)0D(s)>0, s>s. Hence, σ(a1|hn,s) is decreasing in s. Since f satisfies strict MLRP, (a1|hn,ω) is decreasing in ω.

Next, suppose there is adequate learning. So for each state ωΩa1,a2, there is an infinite history with a herd on a1, h=(,a1,a1,) that occurs with positive probability in ω. In particular, if we let ω~=maxΩa1,a2, then for any finite sub-history hn of h,

ωω~:(a1|hn,ω)(a1|hn,ω~)>0,

and since h has positive probability at ω~Ωa1,a2, it follows that

ωω~:n=1(a1|hn,ω)n=1(a1|hn,ω~)>0. (SA.4)

This means (h|ω) is uniformly bounded away from 0 for {ω:ωω~}. Since Ωa1,a2{ω:ωω~}, this establishes SA.1. ∎

HTML from LaTeXML, with custom CSS/JS. The PDF is more accurate.