HTML from LaTeXML, with custom CSS/JS. The PDF is more accurate.

Optimal Contracts for Experimentationthanks: We thank Andrea Attar, Patrick Bolton, Pierre-André Chiappori, Bob Gibbons, Alex Frankel, Zhiguo He, Supreet Kaur, Alessandro Lizzeri, Suresh Naidu, Derek Neal, Alessandro Pavan, Andrea Prat, Canice Prendergast, Jonah Rockoff, Andy Skrzypacz, Lars Stole, Pierre Yared, various seminar and conference audiences, and anonymous referees and the Co-editor for helpful comments. We also thank Johannes Hörner and Gustavo Manso for valuable discussions of the paper. Sébastien Turban provided excellent research assistance. Kartik gratefully acknowledges the hospitality of and funding from the University of Chicago Booth School of Business during a portion of this research; he also thanks the Sloan Foundation for financial support through an Alfred P. Sloan Fellowship.

Marina Halac Graduate School of Business, Columbia University and Department of Economics, University of Warwick. Email: mhalac@columbia.edu.    Navin Kartik Department of Economics, Columbia University. Email: nkartik@columbia.edu.    Qingmin Liu Department of Economics, Columbia University. Email: qingmin.liu@columbia.edu.
January 2016
Abstract

This paper studies a model of long-term contracting for experimentation. We consider a principal-agent relationship with adverse selection on the agent’s ability, dynamic moral hazard, and private learning about project quality. We find that each of these elements plays an essential role in structuring dynamic incentives, and it is only their interaction that generally precludes efficiency. Our model permits an explicit characterization of optimal contracts.

1 Introduction

Agents need to be incentivized to work on, or experiment with, projects of uncertain feasibility. Particularly with uncertain projects, agents are likely to have some private information about their project-specific skills.11 1 Other forms of private information, such as beliefs about the project feasibility or personal effort costs, are also relevant; see Subsection 7.4. Incentive design must deal with not only dynamic moral hazard, but also adverse selection (pre-contractual hidden information) and the inherent process of learning. To date, there is virtually no theoretical work on contracting in such settings. How well can a principal incentivize an agent? How do the environment’s features affect the shape of optimal incentive contracts? What distortions, if any, arise? An understanding is relevant not only for motivating research and development, but also for diverse applications like contract farming, technology adoption, and book publishing, as discussed subsequently.

This paper provides an analysis using a simple model of experimentation. We show that the interaction of learning, adverse selection, and moral hazard introduces new conceptual and analytical issues, with each element playing a role in structuring dynamic incentives. Their interaction affects social efficiency: the principal typically maximizes profits by inducing an agent of low ability to end experimentation inefficiently early, even though there would be no distortion without either adverse selection or moral hazard. Furthermore, despite the intricacy of the problem, intuitive contracts are optimal. The principal can implement the second best by selling the project to the agent and committing to buy back output at time-dated future prices; these prices must increase over time in a manner calibrated to deal with moral hazard and learning.

Our model builds on the now-canonical two-armed “exponential bandit” version of experimentation (Keller, Rady, and Cripps, 2005).22 2 As surveyed by Bergemann and Välimäki (2008), learning is often modeled in economics as an experimentation or bandit problem since Rothschild (1974). The project at hand may either be good or bad. In each period, the agent privately chooses whether to exert effort (work) or not (shirk). If the agent works in a period and the project is good, the project is successful in that period with some probability; if either the agent shirks or the project is bad, success cannot obtain in that period. In the terminology of the experimentation literature, working on the project in any period corresponds to “pulling the risky arm”, while shirking is “pulling the safe arm”; the opportunity cost of pulling the risky arm is the effort cost that the agent incurs. Project success yields a fixed social surplus, accrued by the principal, and obviates the need for any further effort. We introduce adverse selection by assuming that the probability of success in a period (conditional on the agent working and the project being good) depends on the agent’s ability—either high or low—which is the agent’s ex-ante private information or type. Our baseline model assumes no other contracting frictions, in particular we set aside limited liability and endow the principal with full ex-ante commitment power: she maximizes profits by designing a menu of contracts to screen the agent’s ability.33 3 Subsection 7.2 studies the implications of limited liability. The importance of limited liability varies across applications; we also view it as more insightful to separate its effects from those of adverse selection.

Since beliefs about the project’s quality decline so long as effort has been exerted but success not obtained, the first-best or socially efficient solution is characterized by a stopping rule: the agent keeps working (so long as he has not succeeded) up until some point at which the project is permanently abandoned. An important feature for our analysis is that the efficient stopping time is a non-monotonic function of the agent’s ability. The intuition stems from two countervailing forces: on the one hand, for any given belief about the project’s quality, a higher-ability agent provides a higher marginal benefit of effort because he succeeds with a higher probability; on the other hand, a higher-ability agent also learns more from the lack of success over time, so at any point he is more pessimistic about the project than the low-ability agent. Hence, depending on parameter values, the first-best stopping time for a high-ability agent may be larger or smaller than that of a low-ability agent (cf. Bobtcheff and Levy, 2015).

Turning to the second best, the key distinguishing feature of our setting from a canonical (static) adverse selection problem is the dynamic moral hazard and its interaction with the agent’s private learning. Recall that in a standard buyer-seller adverse selection problem, there is no issue about what quantity the agent of one type would consume if he were to deviate and take the other type’s contract: it is simply the quantity specified by the chosen contract. By contrast, in our setting, it is not a priori clear what “consumption bundle”, i.e. effort profile, each agent type will choose after such an off-the-equilibrium path deviation. Dealing with this problem would not pose any conceptual difficulty if there were a systematic relationship between the two types’ effort profiles, for instance if there were a “single-crossing condition” ensuring that the high type always wants to experiment at least as long as the low type. However, given the nature of learning, there is no such systematic relationship in an arbitrary contract. As effort off the equilibrium path is crucial when optimizing over the menu of contracts—because it affects how much “information rent” the agent gets—and the contracts in turn influence the agent’s off-path behavior, we are faced with a non-trivial fixed point problem.

Theorem 2 establishes that the principal optimally screens the agent types by offering two distinct contracts, each inducing the agent to work for some amount of time (so long as success has not been obtained) after which the project is abandoned. Compared to the social optimum, an inefficiency typically obtains: while the high-ability type’s stopping time is efficient, the low-ability type experiments too little. This result is reminiscent of the familiar “no distortion at the top but distortion below” in static adverse selection models, but the distortion arises here only from the conjunction of adverse selection and moral hazard; we show that absent either one, the principal would implement the first best (Theorem 1). Moreover, because of the aforementioned lack of a single-crossing property, it is not immediate in our setting that the principal shouldn’t have the low type over-experiment to reduce the high type’s information rent, particularly when the first best entails the high type stopping earlier than the low type.

Theorem 2 is indirect in the sense that it establishes the (in)efficiency result without elucidating the form of second-best contracts. Our methodology to characterize such contracts distinguishes between the two orderings of the first-best stopping times. We first study the case in which the efficient stopping time for a high-ability agent is larger than that of a low-ability agent. Here we show that although there is no analog of the single-crossing condition mentioned above in an arbitrary contract, such a condition must hold in an optimal contract for the low type. This allows us to simplify the problem and fully characterize the principal’s solution (Theorem 3 and Theorem 4). The case in which the first-best stopping time for the high-ability agent is lower than that of the low-ability agent proves to be more challenging: now, as suggested by the first best, an optimal contract for the low type is often such that the high type would experiment less than the low type should he take this contract. We are able to fully characterize the solution in this case under no discounting (Theorem 5 and Theorem 6).

The second-best contracts we characterize take simple and intuitive forms, partly owing to the simple underlying primitives. In any contract that stipulates experimentation for TT periods it suffices to consider at most T+1T+1 transfers. The reason is that the parties share a common discount factor and there are T+1T+1 possible project outcomes: a success can occur in each of the TT periods or never. One class of contracts are bonus contracts: the agent pays the principal an up-front fee and is then rewarded with a bonus that depends on when the project succeeds (if ever). We characterize the unique sequence of time-dependent bonuses that must be used in an optimal bonus contract for the low-ability type.44 4 For the high type, there are multiple optimal contracts even within a given class such as bonus contracts. The reason for the asymmetry is that the low type’s contract is pinned down by information rent minimization considerations, unlike the high type’s contract. Of course, the high type’s contract cannot be arbitrary either. This sequence is increasing over time up until the termination date. The shape, and its exact calibration, arises from a combination of the agent becoming more pessimistic over time (absent earlier success) and the principal’s desire to avoid any slack in the provision of incentives, while crucially taking into account that the agent can substitute his effort across time.

The optimal bonus contract can be viewed as a simple “sale-with-buyback contract”: the principal sells the project to the agent at the outset for some price, but commits to buy back the project’s output (that obtains with a success) at time-dated future prices. It is noteworthy that contract farming arrangements, widely used in developing countries between agricultural companies and farm producers (Barrett et al., 2012), are often sale-with-buyback contracts: the company sells seeds or other technology (e.g., fertilizers or pesticides) to the farmer and agrees to buy back the crop at pre-determined prices, conditional on this output meeting certain quality standards and delivery requirements (Minot, 2007). The contract farming setting involves a profit-maximizing firm (principal) and a farmer (agent). Miyata, Minot, and Hu (2009) describe the main elements of these environments, focusing on the case of China. It is initially unknown whether the new seeds or technology will produce the desired outcomes in a particular farm, which maps into our project uncertainty.55 5 Besley and Case (1993) study how farmers learn about a new technology over time given the realization of yields from past planting decisions, and how they in turn make dynamic choices. Besides the evident moral hazard problem, there is also adverse selection: farmers differ in unobservable characteristics, such as industriousness, intelligence, and skills.66 6 Beaman et al. (2015) provide evidence of such unobservable characteristics using a field experiment in Mali. Our analysis not only shows that sale-with-buyback contracts are optimal in the presence of uncertainty, moral hazard, and unobservable heterogeneity, but elucidates why. Moreover, as discussed further in Subsection 5.3, our paper offers implications for the design of such contracts and for field experiments on technology adoption more broadly. In particular, field experiments might test our predictions regarding the rich structure of optimal bonus contracts and how the calibration depends on underlying parameters.77 7 We should highlight that our paper is not aimed at studying all the institutional details of contract farming or technology adoption. For example, we do not address multi-agent experimentation and social learning, which has been emphasized by the empirical literature (e.g., Conley and Udry, 2010).

Another class of optimal contracts that we characterize are penalty contracts: the agent receives an up-front payment and is then required to pay the principal some time-dependent penalty in each period in which a success does not obtain, up until either the project succeeds or the contract terminates.88 8 There is a flavor here of “clawbacks” that are sometimes used in practice when an agent is found to be negligent. In our setting, it is the lack of project success that is treated like evidence of negligence (i.e. shirking); note, however, that in equilibrium the principal knows that the agent is not actually negligent. Analogous to the optimal bonus contract, we identify the unique sequence of penalties that must be used in an optimal penalty contract for the low-ability type: the penalty increases over time with a jump at the termination date. These types of contracts correspond to those used, for example, in arrangements between publishers and authors: authors typically receive advances and are then required to pay the publisher back if they do not succeed in completing the book by a given deadline (Owen, 2013). This application fits into our framework when neither publisher nor author may initially be sure whether a commercially-viable book can be written in the relevant timeframe (uncertain project feasibility); the author will have superior information about his suitability or comparative advantage in writing the book (adverse selection about ability); and how much time he actually devotes to the task is unobservable (moral hazard).99 9 Not infrequently, authors fail to deliver in a timely fashion (Suddath, 2012). That private information can be a substantive issue is starkly illustrated by the case of Herman Rosenblat, whose contract with Penguin Books to write a Holocaust survivor memoir was terminated when it was discovered that he fabricated his story.

Our results have implications for the extent of experimentation and innovation across different economic environments. An immediate prediction concerns the effects of asymmetric information: we find that environments with more asymmetric information (either moral hazard or adverse selection) should feature less experimentation, lower success rates, and more dispersion of success rates. We also find that the relationship between success rates and the underlying environment can be subtle. Absent any distortions, “better environments” lead to more innovation. Specifically, an increase in the proportion of high-ability agents or an increase in the ability of both types of the agent yields a higher probability of success in the first best. In the presence of moral hazard and adverse selection, however, the opposite can be true: these changes can induce the principal to distort the low-ability type’s experimentation by more, to the extent that the average success probability goes down in the second best. Consequently, observing higher innovation rates in contractual settings like those we study is neither necessary nor sufficient to deduce a better underlying environment. As discussed in Subsection 5.3, these results may contribute an agency-theoretic component to the puzzle of low technology adoption rates in developing countries.

Related literature.

Broadly, this paper fits into literatures on long-term contracting with either dynamic moral hazard and/or adverse selection. Few papers combine both elements, but two recent exceptions are Sannikov (2007) and Gershkov and Perry (2012).1010 10 Some earlier papers with adverse selection and dynamic moral hazard, such as Laffont and Tirole (1988), focus on the effects of short-term contracting. There is also a literature on dynamic contracting with adverse selection and evolving types but without moral hazard or with only one-shot moral hazard, such as Baron and Besanko (1984) or, more recently, Battaglini (2005), Boleslavsky and Said (2013), and Eső and Szentes (2015). Pavan, Segal, and Toikka (2014) provide a rather general treatment of dynamic mechanism design without moral hazard. These papers are not concerned with learning/experimentation and their settings and focus differ from ours in many ways.1111 11 Demarzo and Sannikov (2011), He et al. (2014), and Prat and Jovanovic (2014) study private learning in moral-hazard models following Holmström and Milgrom (1987), but do not have adverse selection. Sannikov (2013) also proposes a Brownian-motion model and a first-order approach to deal with moral hazard when actions have long-run effects, which raises issues related to private learning. Chassang (2013) considers a general environment and develops an approach to find detail-free contracts that are not optimal but instead guarantee some efficiency bounds so long as there is a long horizon and players are patient. More narrowly, starting with Bergemann and Hege (1998, 2005), there is a fast-growing literature on contracting for experimentation. Virtually all existing research in this area addresses quite different issues than we do, primarily because adverse selection is not accounted for.1212 12 See Bonatti and Hörner (2011, 2015), Manso (2011), Klein (2012), Ederer (2013), Hörner and Samuelson (2013), Kwon (2013), Guo (2014), Halac, Kartik, and Liu (2015), and Moroni (2015). The only exception we are aware of is the concurrent work of Gomes, Gottlieb, and Maestri (2015). They do not consider moral hazard; instead, they introduce two-dimensional adverse selection. Under some conditions they obtain an “irrelevance result” on the dimension of adverse selection that acts similar to our agent’s ability, a conclusion that is similar to our benchmark that the first best obtains in our model when there is no moral hazard.

Outside a pure experimentation framework, Gerardi and Maestri (2012) analyze how an agent can be incentivized to acquire and truthfully report information over time using payments that compare the agent’s reports with the ex-post observed state; by contrast, we assume the state is never observed when experimentation is terminated without a success. Finally, our model can also be interpreted as a problem of delegated sequential search, as in Lewis and Ottaviani (2008) and Lewis (2011). The main difference is that, in our context, these papers assume that the project’s quality is known and hence there is no learning about the likelihood of success (cf. Subsection 7.3); moreover, they do not have adverse selection.

2 The Model

Environment.

A principal needs to hire an agent to work on a project. The project’s quality—synonymous with the state—may either be good or bad, a binary variable. Both parties are initially uncertain about the project’s quality; the common prior on the project being good is β0(0,1)\beta_{0}\in(0,1). The agent is privately informed about whether his ability is low or high, θ{L,H}\theta\in\{L,H\}, where θ=H\theta=H represents “high”. The principal’s prior on the agent’s ability being high is μ0(0,1)\mu_{0}\in(0,1). In each period, t{1,2,}t\in\{1,2,\ldots\}, the agent can either exert effort (work) or not (shirk); this choice is never observed by the principal. Exerting effort in any period costs the agent c>0c>0. If effort is exerted and the project is good, the project is successful in that period with probability λθ\lambda^{\theta}; if either the agent shirks or the project is bad, success cannot obtain in that period. Success is observable and once a project is successful, no further effort is needed.1313 13 Subsection 7.1 establishes that our results apply without change if success is privately observed by the agent but can be verifiably disclosed. We assume 1>λH>λL>01>\lambda^{H}>\lambda^{L}>0. A success yields the principal a payoff normalized to 11; the agent does not intrinsically care about project success. Both parties are risk neutral, have quasi-linear preferences, share a common discount factor δ(0,1]\delta\in(0,1], and are expected-utility maximizers.

Contracts.

We consider contracting at period zero with full commitment power from the principal. To deal with the agent’s hidden information at the time of contracting, the principal’s problem is, without loss of generality, to offer the agent a menu of dynamic contracts from which the agent chooses one. A dynamic contract specifies a sequence of transfers as a function of the publicly observable history, which is simply whether or not the project has been successful to date. To isolate the effects of adverse selection, we do not impose any limited liability constraints until Subsection 7.2. We assume that once the agent has accepted a contract, he is free to work or shirk in any period up until some termination date that is specified by the contract.1414 14 There is no loss of generality here. If the principal has the ability to block the agent from choosing whether to work in some period—“lock him out of the laboratory”, so to speak—this can just as well be achieved by instead stipulating that project success in that period would trigger a large payment to the principal. Throughout, we follow the convention that transfers are from the principal to the agent; negative values represent payments in the other direction.

Formally, a contract is given by 𝐂=(T,W0,𝒃,𝒍)\mathbf{C}=\left(T,W_{0},\bm{b},\bm{l}\right), where T{0,1,}T\in\mathbb{N}\equiv\{0,1,\ldots\} is the termination date of the contract, W0W_{0}\in\mathbb{R} is an up-front transfer (or wage) at period zero, 𝒃=(b1,,bT)\bm{b}=\left(b_{1},\ldots,b_{T}\right) specifies a transfer btb_{t}\in\mathbb{R} made at period tt conditional on the project being successful in period tt, and analogously 𝒍=(l1,,lT)\bm{l}=\left(l_{1},\ldots,l_{T}\right) specifies a transfer ltl_{t}\in\mathbb{R} made at period tt conditional on the project not being successful in period tt (nor in any prior period).1515 15 We thus restrict attention to deterministic contracts. Throughout, symbols in bold typeface denote vectors. W0W_{0} and TT are redundant because W0W_{0} can be effectively induced by suitable modifications to b1b_{1} and l1l_{1}, while TT can be effectively induced by setting bt=lt=0b_{t}=l_{t}=0 for all t>Tt>T. However, it is expositionally convenient to include these components explicitly in defining a contract. Furthermore, there is no loss in assuming that TT\in\mathbb{N}; as we show, it is always optimal for the principal to stop experimentation at a finite time, so she cannot benefit from setting T=T=\infty.,1616 16 As the principal and agent share a common discount factor, what matters is only the mapping from outcomes to transfers, not the dates at which transfers are made. Our convention facilitates our exposition. We refer to any btb_{t} as a bonus and any ltl_{t} as a penalty. Note that btb_{t} is not constrained to be positive nor must ltl_{t} be negative; however, these cases will be focal and hence our choice of terminology. Without loss of generality, we assume that if T>0T>0 then T=max{t:either bt0 or lt0}T=\max\{t:\text{either $b_{t}\neq 0$ or $l_{t}\neq 0$}\}. The agent’s actions are denoted by 𝐚=(a1,,aT)\mathbf{a}=\left(a_{1},\ldots,a_{T}\right), where at=1a_{t}=1 if the agent works in period tt and at=0a_{t}=0 if the agent shirks.

Payoffs.

The principal’s expected discounted payoff at time zero from a contract 𝐂=(T,W0,𝒃,𝒍)\mathbf{C}=\left(T,W_{0},\bm{b},\bm{l}\right), an agent of type θ\theta, and a sequence of the agent’s actions 𝐚\mathbf{a} is denoted Π0θ(𝐂,𝐚)\Pi_{0}^{\theta}(\mathbf{C},\mathbf{a}), which can be computed as:

Π0θ(𝐂,𝐚):=W0(1β0)t=1Tδtlt+β0t=1Tδt[s<t(1asλθ)][atλθ(1bt)(1atλθ)lt].\Pi_{0}^{\theta}\left(\mathbf{C},\mathbf{a}\right):=-W_{0}-(1-\beta_{0})\sum% \limits_{t=1}^{T}\delta^{t}l_{t}+\beta_{0}\sum\limits_{t=1}^{T}\delta^{t}\left% [\prod\limits_{s<t}\left(1-a_{s}\lambda^{\theta}\right)\right]\left[a_{t}% \lambda^{\theta}\left(1-b_{t}\right)-\left(1-a_{t}\lambda^{\theta}\right)l_{t}% \right]. (1)

Formula (1) is understood as follows. W0W_{0} is the up-front transfer made from the principal to the agent. With probability 1β01-\beta_{0} the state is bad, in which case the project never succeeds and hence the entire sequence of penalties 𝒍\bm{l} is transferred. Conditional on the state being good (which occurs with probability β0\beta_{0}), the probability of project success depends on both the agent’s effort choices and his ability; s<t(1asλθ)\prod\limits_{s<t}\left(1-a_{s}\lambda^{\theta}\right) is the probability that a success does not obtain between period 11 and t1t-1 conditional on the good state. If the project were to succeed at time tt, then the principal would earn a payoff of 11 in that period, and the transfers would be the sequence of penalties (l1,,lt1)(l_{1},\ldots,l_{t-1}) followed by the bonus btb_{t}.

Through analogous reasoning, bearing in mind that the agent does not directly value project success but incurs the cost of effort, the agent’s expected discounted payoff at time zero given his type θ\theta, contract 𝐂\mathbf{C}, and action profile 𝐚\mathbf{a} is

U0θ(𝐂,𝐚):=W0+(1β0)t=1Tδt(ltatc)+β0t=1Tδt[s<t(1asλθ)][at(λθbtc)+(1atλθ)lt].U_{0}^{\theta}\left(\mathbf{C},\mathbf{a}\right):=W_{0}+(1-\beta_{0})\sum% \limits_{t=1}^{T}\delta^{t}\left(l_{t}-a_{t}c\right)+\beta_{0}\sum\limits_{t=1% }^{T}\delta^{t}\left[\prod\limits_{s<t}\left(1-a_{s}\lambda^{\theta}\right)% \right]\left[a_{t}\left(\lambda^{\theta}b_{t}-c\right)+\left(1-a_{t}\lambda^{% \theta}\right)l_{t}\right]. (2)

If a contract is not accepted, both parties’ payoffs are normalized to zero.

Bonus and penalty contracts.

Our analysis will make use of two simple classes of contracts. A bonus contract is one where aside from any initial transfer there is at most only one other transfer, which occurs when the agent obtains a success. Formally, a bonus contract is 𝐂=(T,W0,𝒃,𝒍)\mathbf{C}=\left(T,W_{0},\bm{b},\bm{l}\right) such that lt=0l_{t}=0 for all t{1,,T}t\in\{1,\ldots,T\}. A bonus contract is a constant-bonus contract if, in addition, there is some constant bb such that bt=bb_{t}=b for all t{1,,T}t\in\{1,\ldots,T\}. When the context is clear, we denote a bonus contract as just 𝐂=(T,W0,𝒃)\mathbf{C}=(T,W_{0},\bm{b}) and a constant-bonus contract as 𝐂=(T,W0,b)\mathbf{C}=(T,W_{0},b). By contrast, a penalty contract is one where the agent receives no payments for success and instead is penalized for failure. Formally, a penalty contract is 𝐂=(T,W0,𝒃,𝒍)\mathbf{C}=\left(T,W_{0},\bm{b},\bm{l}\right) such that bt=0b_{t}=0 for all t{1,,T}t\in\{1,\ldots,T\}. A penalty contract is a onetime-penalty contract if, in addition, lt=0l_{t}=0 for all t{1,,T1}t\in\{1,\ldots,T-1\}. That is, while in a general penalty contract the agent may be penalized for each period in which he fails to obtain a success, in a onetime-penalty contract the agent is penalized only if a success does not obtain by the termination date TT. We denote a penalty contract as just 𝐂=(T,W0,𝒍)\mathbf{C}=(T,W_{0},\bm{l}) and a onetime-penalty contract as 𝐂=(T,W0,lT)\mathbf{C}=(T,W_{0},l_{T}).

Although each of these two classes of contracts will be useful for different reasons, there is an isomorphism between them; furthermore, either class is “large enough” in a suitable sense. More precisely, say that two contracts, 𝐂=(T,W0,𝒃,𝒍)\mathbf{C}=(T,W_{0},\bm{b},\bm{l}) and ^𝐂=(T,W^0,^𝒃,𝒍^)\widehat{}\mathbf{C}=(T,\widehat{W}_{0},\widehat{}\bm{b},\widehat{\bm{l}}), are equivalent if for all θ{L,H}\theta\in\{L,H\} and 𝐚=(a1,,aT)\mathbf{a}=\left(a_{1},\ldots,a_{T}\right): U0θ(𝐂,𝐚)=U0θ(𝐂^,𝐚) and Π0θ(𝐂,𝐚)=Π0θ(^𝐂,𝐚).U_{0}^{\theta}(\mathbf{C},\mathbf{a})=U_{0}^{\theta}(\widehat{\mathbf{C}},% \mathbf{a})\text{ \ and \ }\Pi_{0}^{\theta}\left(\mathbf{C},\mathbf{a}\right)=% \Pi_{0}^{\theta}(\widehat{}\mathbf{C},\mathbf{a}).

Proposition 1.

For any contract 𝐂=(T,W0,𝒃,𝒍)\mathbf{C}=(T,W_{0},\bm{b},\bm{l}) there exist both an equivalent penalty contract 𝐂^=(T,W^0,𝐥^)\widehat{\mathbf{C}}=(T,\widehat{W}_{0},\widehat{\bm{l}}) and an equivalent bonus contract 𝐂~=(T,W~0,𝐛~)\widetilde{\mathbf{C}}=(T,\widetilde{W}_{0},\widetilde{\bm{b}}).

Proof.

See the Supplementary Appendix. ∎

Proposition 1 implies that it is without loss to focus either on bonus contracts or on penalty contracts. The proof is constructive: given an arbitrary contract, it explicitly derives equivalent penalty and bonus contracts. The intuition is that all that matters in any contract is the induced vector of discounted transfers for success occurring in each possible period (and never), and these transfers can be induced with bonuses or penalties.1717 17 For example, in a two-period contract 𝐂=(2,W0,𝒃,𝒍)\mathbf{C}=(2,W_{0},\bm{b},\bm{l}), the agent’s discounted transfer is W0+δb1W_{0}+\delta b_{1} if he succeeds in period one, W0+δl1+δ2b2W_{0}+\delta l_{1}+\delta^{2}b_{2} if he succeeds in period two, and W0+δl1+δ2l2W_{0}+\delta l_{1}+\delta^{2}l_{2} if he does not succeed in either period. The same transfers are induced by a penalty contract 𝐂^=(2,W^0,𝒍^)\widehat{\mathbf{C}}=(2,\widehat{W}_{0},\widehat{\bm{l}}) with W^0=W0+δb1\widehat{W}_{0}=W_{0}+\delta b_{1}, l^1=l1b1+δb2\widehat{l}_{1}=l_{1}-b_{1}+\delta b_{2}, and l^2=l2b2\widehat{l}_{2}=l_{2}-b_{2}, and by a bonus contract ~𝐂=(2,W~0,~𝒃)\widetilde{}\mathbf{C}=(2,\widetilde{W}_{0},\widetilde{}\bm{b}) with W~0=W0+δl1+δ2l2\widetilde{W}_{0}=W_{0}+\delta l_{1}+\delta^{2}l_{2}, b~1=b1l1δl2\widetilde{b}_{1}=b_{1}-l_{1}-\delta l_{2}, and b~2=b2l2\widetilde{b}_{2}=b_{2}-l_{2}. The proof also shows that when δ=1\delta=1, onetime-penalty contracts are equivalent to constant-bonus contracts.

3 Benchmarks

3.1 The first best

Consider the first-best solution, i.e. when the agent’s type θ\theta is commonly known and his effort in each period is publicly observable and contractible. Since beliefs about the state being good decline so long as effort has been exerted but success not obtained, the first-best solution is characterized by a stopping rule such that an agent of ability θ\theta keeps exerting effort so long as success has not obtained up until some period tθt^{\theta}, whereafter effort is no longer exerted.1818 18 More precisely, the first best can always be achieved using a stopping rule for each type; when and only when δ=1\delta=1, there are other rules that also achieve the first best. Without loss, we focus on stopping rules. Let βtθ\beta_{t}^{\theta} be a generic belief on the state being good at the beginning of period tt (which will depend on the history of effort), and β¯tθ\overline{\beta}_{t}^{\theta} be this belief when the agent has exerted effort in all periods 1,,t11,\ldots,t-1. The first-best stopping time tθt^{\theta} is given by

tθ=maxt0{t:β¯tθλθc},t^{\theta}=\max_{t\geq 0}\left\{t:\overline{\beta}_{t}^{\theta}\lambda^{\theta% }\geq c\right\}, (3)

where, for each θ\theta, β¯0θ:=β0\overline{\beta}^{\theta}_{0}:=\beta_{0}, and for t1t\geq 1, Bayes’ rule yields

β¯tθ=β0(1λθ)t1β0(1λθ)t1+(1β0).\overline{\beta}_{t}^{\theta}=\frac{\beta_{0}\left(1-\lambda^{\theta}\right)^{% t-1}}{\beta_{0}\left(1-\lambda^{\theta}\right)^{t-1}+\left(1-\beta_{0}\right)}. (4)

Note that (3) is only well-defined when cβ0λθc\leq\beta_{0}\lambda^{\theta}; if c>β0λθc>\beta_{0}\lambda^{\theta}, it would be efficient to not experiment at all, i.e. stop at tθ=0t^{\theta}=0. To focus on the most interesting cases, we assume:

Assumption 1.

Experimentation is efficient for both types: for θ{L,H}\theta\in\{L,H\}, β0λθ>c.\beta_{0}\lambda^{\theta}>c.

If parameter values are such that β¯tθθλθ=c\overline{\beta}_{t^{\theta}}^{\theta}\lambda^{\theta}=c,1919 19 We do not assume this condition in our analysis, but it is convenient for the current discussion. equations (3) and (4) can be combined to derive the following closed-form solution for the first-best stopping time for type θ\theta:

tθ=1+log(cλθc1β0β0)log(1λθ).t^{\theta}=1+\frac{\log\left(\frac{c}{\lambda^{\theta}-c}\frac{1-\beta_{0}}{% \beta_{0}}\right)}{\log\left(1-\lambda^{\theta}\right)}. (5)

Equation (5) yields intuitive monotonicity of the first-best stopping time as a function of the prior that the project is good, β0\beta_{0}, and the cost of effort, cc.2020 20 One may also notice that the discount factor, δ\delta, does not enter (5). In other words, unlike the traditional focus of experimentation models, there is no tradeoff here between “exploration” and “exploitation”, as the first-best strategy is invariant to patience. Our model and subsequent analysis can be generalized to incorporate this tradeoff, but the additional burden does not yield commensurate insight. But it also implies a fundamental non-monotonicity as a function of the agent’s ability, λθ\lambda^{\theta}, as shown in Figure 1. (For simplicity, the figure ignores integer constraints on tθt^{\theta}.) This stems from the interaction of two countervailing forces. On the one hand, for any given belief about the state, the expected marginal benefit of effort is higher when the agent’s ability is higher; on the other hand, the higher is the agent’s ability, the more informative is a lack of success in a period in which he works. Hence, at any time t>1t>1, a higher-ability agent is more pessimistic about the state (given that effort has been exerted in all prior periods), which has the effect of decreasing the expected marginal benefit of effort. Altogether, this makes the first-best stopping time non-monotonic in ability; both tH>tLt^{H}>t^{L} and tH<tLt^{H}<t^{L} are robust possibilities that arise for different parameters. As we will see, this has substantial implications.

Refer to caption
Figure 1: The first-best stopping time.

The first-best expected discounted surplus at time zero from type θ\theta is

t=1tθδt[β0(1λθ)t1(λθc)(1β0)c].\sum\limits_{t=1}^{t^{\theta}}\delta^{t}\left[\beta_{0}\left(1-\lambda^{\theta% }\right)^{t-1}\left(\lambda^{\theta}-c\right)-(1-\beta_{0})c\right].

3.2 No adverse selection or no moral hazard

Our model has two sources of asymmetric information: adverse selection and moral hazard. To see that their interaction is essential, it is useful to understand what would happen in the absence of either one.

Consider first the case without adverse selection, i.e. assume the agent’s ability is observable but there is moral hazard. The principal can then use a constant-bonus contract to effectively sell the project to the agent at a price that extracts all the (ex-ante) surplus. Specifically, suppose the principal offers the agent of type θ\theta a constant-bonus contract 𝐂θ=(tθ,W0θ,1)\mathbf{C}^{\theta}=(t^{\theta},W^{\theta}_{0},1), where W0θW_{0}^{\theta} is chosen so that conditional on the agent exerting effort in each period up to the first-best termination date (as long as success has not obtained), the agent’s participation constraint at time zero binds:

U0θ(𝐂θ,𝟏)=t=1tθδt[β0(1λθ)t1(λθc)(1β0)c]+W0θ=0,U_{0}^{\theta}\left(\mathbf{C}^{\theta},\mathbf{1}\right)=\sum\limits_{t=1}^{t% ^{\theta}}\delta^{t}\left[\beta_{0}\left(1-\lambda^{\theta}\right)^{t-1}\left(% \lambda^{\theta}-c\right)-(1-\beta_{0})c\right]+W_{0}^{\theta}=0,

where the notation 𝟏\mathbf{1} denotes the action profile of working in every period of the contract. Plainly, this contract makes the agent fully internalize the social value of success and hence achieves the first-best level of experimentation, while the principal keeps all the surplus.

Consider next the case with adverse selection but no moral hazard: the agent’s effort in any period still costs him c>0c>0 but is observable and contractible. The principal can then implement the first best and extract all the surplus by using simple contracts that pay the agent for effort rather than outcomes. Specifically, the principal can offer the agent a choice between two contracts that involve no bonuses or penalties, with each paying the agent cc for every period that he works. The termination date is tLt^{L} in the contract intended for the low type and tHt^{H} in the contract intended for the high type. Plainly, the agent’s payoff is zero regardless of his type and which contract and effort profile he chooses. Hence, the agent is willing to choose the contract intended for his type and work until either a success is obtained or the termination date is reached.2121 21 The same idea underlies Gomes et al.’s (2015) Lemma 2. While this mechanism makes the agent indifferent over the contracts, there are more sophisticated optimal mechanisms, detailed in earlier versions of our paper, that satisfy the agent’s self-selection constraint strictly.

To summarize:

Theorem 1.

If there is either no moral hazard or no adverse selection, the principal optimally implements the first best and extracts all the surplus.

A proof is omitted in light of the simple arguments preceding the theorem. Theorem 1 also holds when there are many types; that both kinds of information asymmetries are essential to generate distortions is general in our experimentation environment.2222 22 We note that learning is also important in generating distortions: in the absence of learning (i.e. if the project were known to be good, β0=1\beta_{0}=1), the principal may again implement the first best. For expositional purposes, we defer this discussion to Subsection 7.3.

4 Second-Best (In)Efficiency

We now turn to the setting with both moral hazard and adverse selection. In this section, we formalize the principal’s problem and deduce the nature of second-best inefficiency. We provide explicit characterizations of optimal contracts in Section 5 and Section 6.

Without loss, we assume that the principal specifies a desired effort profile along with a contract. An optimal menu of contracts maximizes the principal’s ex-ante expected payoff subject to incentive compatibility constraints for effort (ICaθ{}^{\theta}_{a} below), participation constraints (IRθ below), and self-selection constraints for the agent’s choice of contract (ICθθ{}^{\theta\theta^{\prime}} below). Denote

𝜶θ(𝐂):=argmax𝐚U0θ(𝐂,𝐚)\bm{\alpha}^{\theta}\left(\mathbf{C}\right):=\operatorname*{arg\,max}\limits_{% \mathbf{a}}\ U_{0}^{\theta}\left(\mathbf{C},\mathbf{a}\right)

as the set of optimal action plans for the agent of type θ\theta under contract 𝐂\mathbf{C}. With a slight abuse of notation, we will write U0θ(𝐂,𝜶θ(𝐂))U^{\theta}_{0}(\mathbf{C},\bm{\alpha}^{\theta}(\mathbf{C})) for the type-θ\theta agent’s utility at time zero from any contract 𝐂\mathbf{C}. The principal’s program is:

max(𝐂H,𝐂L,𝐚H,𝐚L)μ0Π0H(𝐂H,𝐚H)+(1μ0)Π0L(𝐂L,𝐚L)\max_{\left(\mathbf{C}^{H},\mathbf{C}^{L},\mathbf{a}^{H},\mathbf{a}^{L}\right)% }\mu_{0}\Pi_{0}^{H}\left(\mathbf{C}^{H},\mathbf{a}^{H}\right)+\left(1-\mu_{0}% \right)\Pi_{0}^{L}\left(\mathbf{C}^{L},\mathbf{a}^{L}\right)

subject to, for all θ,θ{L,H}\theta,\theta^{\prime}\in\left\{L,H\right\},

𝐚θ\displaystyle\mathbf{a}^{\theta} 𝜶θ(𝐂θ),\displaystyle\in\bm{\alpha}^{\theta}(\mathbf{C}^{\theta}), (ICaθ{}^{\theta}_{a})
U0θ(𝐂θ,𝐚θ)\displaystyle U_{0}^{\theta}(\mathbf{C}^{\theta},\mathbf{a}^{\theta}) 0,\displaystyle\geq 0, (IRθ)
U0θ(𝐂θ,𝐚θ)\displaystyle U_{0}^{\theta}(\mathbf{C}^{\theta},\mathbf{a}^{\theta}) U0θ(𝐂θ,𝜶θ(𝐂θ)).\displaystyle\geq U_{0}^{\theta}(\mathbf{C}^{\theta^{\prime}},\bm{\alpha}^{% \theta}(\mathbf{C}^{\theta^{\prime}})). (ICθθ{}^{\theta\theta^{\prime}})

Adverse selection is reflected in the self-selection constraints (ICθθ{}^{\theta\theta^{\prime}}), as is familiar. Moral hazard is reflected directly in the constraints (ICaθ{}^{\theta}_{a}) and also indirectly in the constraints (ICθθ{}^{\theta\theta^{\prime}}) via the term 𝜶θ(𝐂θ)\bm{\alpha}^{\theta}(\mathbf{C}^{\theta^{\prime}}). To get a sense of how these matter, consider the agent’s incentive to work in some period tt. This is shaped not only by the transfers that are directly tied to success/failure in period tt (btb_{t} and ltl_{t}) but also by the transfers tied to subsequent outcomes, through their effect on continuation values. In particular, ceteris paribus, raising the continuation value (say, by increasing either bt+1b_{t+1} or lt+1l_{t+1}) makes reaching period t+1t+1 more attractive and hence reduces the incentive to work in period tt: this is a dynamic agency effect.2323 23 Mason and Välimäki (2011), Bhaskar (2012, 2014), Hörner and Samuelson (2013), and Kwon (2013) also highlight dynamic agency effects, but in settings without adverse selection. Note moreover that the continuation value at any point in a contract depends on the agent’s type and his effort profile; hence it is not sufficient to consider a single continuation value at each period. Furthermore, besides having an effect on continuation values, the agent’s type also affects current incentives for effort because the expected marginal benefit of effort in any period differs for the two types. Altogether, the optimal plan of action will generally be different for the two types of the agent, i.e. for an arbitrary contract 𝐂\mathbf{C}, we may have 𝜶H(𝐂)𝜶L(𝐂)=\bm{\alpha}^{H}(\mathbf{C})\cap\bm{\alpha}^{L}(\mathbf{C})=\emptyset.2424 24 Related issues arise in static models that allow for both adverse selection and moral hazard; see for example the discussion in Laffont and Martimort (2001, Chapter 7).

Our result on second-best (in)efficiency is as follows:

Theorem 2.

In any optimal menu of contracts, each type θ{L,H}\theta\in\{L,H\} is induced to work for some number of periods, t¯θ\overline{t}^{\theta}. Relative to the first-best stopping times, tHt^{H} and tLt^{L}, the second best has t¯H=tH\overline{t}^{H}=t^{H} and t¯LtL\overline{t}^{L}\leq t^{L}.

Proof.

See Appendix A. ∎

Theorem 2 says that relative to the first best, there is no distortion in the amount of experimentation by the high-ability agent whereas the low-ability agent may be induced to under-experiment. It is interesting that this is a familiar “no distortion (only) at the top” result from static models of adverse selection, even though the inefficiency arises here from the conjunction of adverse selection and dynamic moral hazard (cf. Theorem 1). Moral hazard generates an “information rent” for the high type but not for the low type. As will be elaborated subsequently, reducing the low type’s amount of experimentation allows the principal to reduce the high type’s information rent. The optimal t¯L\overline{t}^{L} trades off this information rent with the low type’s efficiency. For typical parameters, it will be the case that t¯L{1,,tL1}\overline{t}^{L}\in\{1,\ldots,t^{L}-1\}, so that the low type engages in some experimentation but not as much as socially efficient; however, it is possible that the low type is induced to not experiment at all (t¯L=0\overline{t}^{L}=0) or to experiment for the first-best amount of time (t¯L=tL\overline{t}^{L}=t^{L}). The former possibility arises for reasons akin to exclusion in the standard model (e.g. the prior, μ0\mu_{0}, on the high type is sufficiently high); the latter possibility is because time is discrete. Indeed, if the length of each time interval shrinks and one takes a suitable continuous-time limit, then there will be some distortion, i.e. t¯L<tL\overline{t}^{L}<t^{L}.

The proof of Theorem 2 does not rely on characterizing second-best contracts.2525 25 Note that when δ<1\delta<1, efficiency requires each type to use a “stopping strategy” (i.e., work for a consecutive sequence of periods beginning with period one). The proof technique for Theorem 2 does not allow us to establish that the low type uses a stopping strategy in the second-best solution; however, it shows that one can take the high type to be doing so. That the low type can also be taken to use a stopping strategy (with the second-best stopping time) will be deduced subsequently in those cases in which we are able to characterize second-best contracts. We establish t¯H=tH\overline{t}^{H}=t^{H} by proving that the low type’s self-selection constraint can always be satisfied without creating any distortions. The idea is that the principal can exploit the two types’ differing probabilities of success by making the high type’s contract “risky enough” to deter the low type from taking it, while still satisfying all other constraints.2626 26 Specifically, given an optimal contract for the high type, the principal can increase the magnitude of the penalties while adjusting the time-zero transfer so that the high type’s expected payoff and effort profile do not change. Making the penalties severe enough (i.e., negative enough) then ensures that the low type’s payoff from taking the high type’s contract is negative and hence (ICLH) is satisfied at no cost. Crucially, an analogous construction would not work for the high type’s self-selection constraint: the high type’s payoff under the low type’s contract cannot be lower than the low type’s, as the high type can always generate the same distribution of project success as the low type by suitably mixing over effort. From the point of view of correlated-information mechanism design (Cremer and McLean, 1985, 1988; Riordan and Sappington, 1988), the issue is that because of moral hazard, the signal correlated with the agent’s type is not independent of the agent’s report. In a different setting, Obara (2008) has also noted this effect of hidden actions. While Obara (2008) shows that in his setting approximate full surplus extraction may be achieved by having agents randomize over their actions, this is not generally possible here because the feasible set of distributions of project success for the high type is a superset of that of the low type. We establish t¯LtL\overline{t}^{L}\leq t^{L} by showing that any contract for the low type inducing t¯L>tL\overline{t}^{L}>t^{L} can be modified by “removing” the last period of experimentation in this contract and concurrently reducing the information rent for the high type. Due to the lack of structure governing the high type’s behavior upon deviating to the low type’s contract, we prove the information-rent reduction no matter what action plan the high type would choose upon taking the low type’s contract. It follows that inducing over-experimentation by the low type cannot be optimal: not only would that reduce social surplus but it would also increase the high type’s information rent.

While Theorem 2 has implications for the extent of experimentation and innovation in different economic environments, we postpone such discussion to Subsection 5.3, after describing optimal contracts and their comparative statics.

5 Optimal Contracts when tH>tLt^{H}>t^{L}

We characterize optimal contracts by first studying the case in which the first-best stopping times are ordered tH>tLt^{H}>t^{L}, i.e. when the speed-of-learning effect that pushes the first-best stopping time down for a higher-ability agent does not dominate the productivity effect that pushes in the other direction. Any of the following conditions on the primitives is sufficient for tH>tLt^{H}>t^{L}, given a set of other parameters: (i) β0\beta_{0} is small enough, (ii) λL\lambda^{L} and λH\lambda^{H} are small enough, or (iii) cc is large enough. We maintain the assumption that tH>tLt^{H}>t^{L} implicitly throughout this section.

5.1 The solution

A class of solutions to the principal’s program described in Section 4 when tH>tLt^{H}>t^{L} is as follows:

Theorem 3.

Assume tH>tLt^{H}>t^{L}. There is an optimal menu in which the principal separates the two types using penalty contracts. In particular, the optimum can be implemented using a onetime-penalty contract for type HH, 𝐂H=(tH,W0H,ltHH)\mathbf{C}^{H}=(t^{H},W^{H}_{0},l^{H}_{t^{H}}) with ltHH<0<W0Hl^{H}_{t^{H}}<0<W^{H}_{0}, and a penalty contract for type LL, 𝐂L=(t¯L,W0L,𝐥L)\mathbf{C}^{L}=(\overline{t}^{L},W^{L}_{0},\bm{l}^{L}), such that:

  1. 1.

    For all t{1,,t¯L}t\in\{1,\ldots,\overline{t}^{L}\},

    ltL={(1δ)cβ¯tLλL if t<t¯L,cβ¯t¯LLλL if t=t¯L;{l}_{t}^{L}=\begin{cases}-\left(1-\delta\right)\frac{c}{\overline{\beta}_{t}^{% L}\lambda^{L}}&\text{ if }t<\overline{t}^{L}\text{,}\\ -\frac{c}{\overline{\beta}_{\overline{t}^{L}}^{L}\lambda^{L}}&\text{ if }t=% \overline{t}^{L}\text{;}\end{cases} (6)
  2. 2.

    W0L>0W^{L}_{0}>0 is such that the participation constraint, (IRL), binds;

  3. 3.

    Type HH gets an information rent: U0H(𝐂H,𝜶H(𝐂H))>0U^{H}_{0}(\mathbf{C}^{H},\bm{\alpha}^{H}(\mathbf{C}^{H}))>0;

  4. 4.

    𝟏𝜶H(𝐂H)\mathbf{1}\in\bm{\alpha}^{H}(\mathbf{C}^{H}); 𝟏𝜶L(𝐂L)\mathbf{1}\in\bm{\alpha}^{L}(\mathbf{C}^{L}); and 𝟏=𝜶H(𝐂L)\mathbf{1}=\bm{\alpha}^{H}(\mathbf{C}^{L}).

Generically, the above contract is the unique optimal contract for type LL within the class of penalty contracts.

Proof.

See Appendix B. ∎

The optimal contract for the low type characterized by (6) is a penalty contract in which the magnitude of the penalty is increasing over time, with a “jump” in the contract’s final period. The jump highlights dynamic agency effects: by obtaining a success in a period tt, the agent not only avoids the penalty ltLl^{L}_{t} but also the penalty lt+1Ll^{L}_{t+1} and those after. The last period’s penalty needs to compensate for the absence of future penalties. Figure 2 depicts the low type’s contract; the comparative statics seen in the figure will be discussed subsequently. Only when there is no discounting does the low type’s contract reduce to a onetime-penalty contract where a penalty is paid only if the project has not succeeded by t¯L\overline{t}^{L}. For any discount factor, the high type’s contract characterized in Theorem 3 is a onetime-penalty contract in which he only pays a penalty to the principal if there is no success by the first-best stopping time tHt^{H}. On the equilibrium path, both types of the agent exert effort in every period until their respective stopping times; moreover, were the high type to take the low type’s contract (off the equilibrium path), he would also exert effort in every period of the contract. This implies that the high type gets an information rent because he would be less likely than the low type to incur any of the penalties in 𝐂L\mathbf{C}^{L}.

Although the optimal contract for the low type is (generically) unique among penalty contracts, there are a variety of optimal penalty contracts for the high type. The reason is that the low type’s optimal contract is pinned down by the need to simultaneously incentivize the low type’s effort and yet minimize the information rent obtained by the high type. This leads to a sequence of penalties for the low type, given by (6), that make him indifferent between working and shirking in each period of the contract, as we explain further in Subsection 5.2. On the other hand, the high type’s contract only needs to be made unattractive to the low type subject to incentivizing effort from the high type and providing the high type a utility level given by his information rent. There is latitude in how this can be done: the onetime penalty in the high type’s contract of Theorem 3 is chosen to be severe enough so that this contract is “too risky” for the low type to accept.

Remark 1.

The proof of Theorem 3 provides a simple algorithm to solve for an optimal menu of contracts. For any t^{0,,tL}\hat{t}\in\{0,\ldots,t^{L}\}, we characterize an optimal menu that solves the principal’s program subject to an additional constraint that the low type must experiment until period t^\hat{t}. The low type’s contract in this menu is given by (6) with the termination date t^\hat{t} rather than t¯L\overline{t}^{L}. An optimal (unconstrained) menu is then obtained by maximizing the principal’s objective function over t^{0,,tL}\hat{t}\in\{0,\ldots,t^{L}\}.

The characterization in Theorem 3 yields the following comparative statics:

Proposition 2.

Assume tH>tLt^{H}>t^{L} and consider changes in parameters that preserve this ordering. The second-best stopping time for type LL, t¯L\overline{t}^{L}, is weakly increasing in β0\beta_{0} and λL\lambda^{L}, weakly decreasing in cc and μ0\mu_{0}, and can increase or decrease in λH\lambda^{H}. The distortion in this stopping time, measured by tLt¯Lt^{L}-\overline{t}^{L}, is weakly increasing in μ0\mu_{0} and can increase or decrease in β0\beta_{0}, λL\lambda^{L}, λH\lambda^{H}, and cc.

Proof.

See the Supplementary Appendix. ∎

Figure 2 illustrates some of the conclusions of Proposition 2. The comparative static of t¯L\overline{t}^{L} in μ0\mu_{0} is intuitive: the higher the ex-ante probability of the high type, the more the principal benefits from reducing the high type’s information rent and hence the more she shortens the low type’s experimentation. Matters are more subtle for other parameters. Consider, for example, an increase in β0\beta_{0}. On the one hand, this increases the social surplus from experimentation, which suggests that t¯L\overline{t}^{L} should increase. But there are two other effects: holding fixed t¯L\overline{t}^{L}, penalties of lower magnitude can be used to incentivize effort from the low type because the project is more likely to succeed (cf. equation (6)), which has an effect of decreasing the information rent for the high type; yet, a higher β0\beta_{0} also has a direct effect of increasing the information rent because the differing probability of success for the two types is only relevant when the project is good. Nevertheless, Proposition 2 establishes that it is optimal to (weakly) increase t¯L\overline{t}^{L} when β0\beta_{0} increases.

Since the high type’s information rent is increasing in λH\lambda^{H}, one may expect the principal to reduce the low type’s experimentation when λH\lambda^{H} increases. However, a higher λH\lambda^{H} means that the high type is likely to succeed earlier when deviating to the low type’s contract. For this reason, an increase in λH\lambda^{H} can reduce the incremental information-rent cost of extending the low type’s contract, to the extent that the gain in efficiency from the low type makes it optimal to increase t¯L\overline{t}^{L}.

Refer to caption
Figure 2: The optimal penalty contract for type LL under different values of μ0\mu_{0} and β0\beta_{0}. Both graphs have δ=0.5\delta=0.5, λL=0.1\lambda^{L}=0.1, λH=0.12\lambda^{H}=0.12, and c=0.06c=0.06. The left graph has β0=0.89\beta_{0}=0.89, μ0=0.3\mu_{0}=0.3, and μ0=0.6\mu_{0}^{\prime}=0.6; the right graph has β0=0.85\beta_{0}=0.85, β0=0.89\beta_{0}^{\prime}=0.89, and μ0=0.3\mu_{0}=0.3. The first-best entails tL=15t^{L}=15 on the left graph, and tL=12t^{L}=12 (for β0\beta_{0}) and tL=15t^{L}=15 (for β0\beta^{\prime}_{0}) on the right graph.

Turning to the magnitude of distortion, tLt¯Lt^{L}-\overline{t}^{L}: since the first-best stopping time tLt^{L} does not depend on the probability of a high type, μ0\mu_{0}, while t¯L\overline{t}^{L} is decreasing in this parameter, it is immediate that the distortion is increasing in μ0\mu_{0}. The time tLt^{L} is also independent of the high type’s ability, λH\lambda^{H}; thus, since t¯L\overline{t}^{L} may increase or decrease in λH\lambda^{H}, the same is true for tLt¯Lt^{L}-\overline{t}^{L}. Finally, with respect to β0\beta_{0}, λL\lambda^{L}, and cc, the distortion’s ambiguous comparative statics stem from the fact that tLt^{L} and t¯L\overline{t}^{L} move in the same direction when these parameters change. For example, increasing β0\beta_{0} can reduce tLt¯Lt^{L}-\overline{t}^{L} when μ0\mu_{0} is low but increase tLt¯Lt^{L}-\overline{t}^{L} when μ0\mu_{0} is high; the reason is that a larger ex-ante probability of the high type makes increasing t¯L\overline{t}^{L} more costly in terms of information rent.

Theorem 3 utilizes penalty contracts in which the agent is required to pay the principal when he fails to obtain a success. While these contracts prove analytically convenient (as explained in Subsection 5.2), a weakness is that they do not satisfy interim participation constraints: in the implementation of Theorem 3, the agent of either type θ\theta would “walk away” from his contract in any period t{1,,t¯θ}t\in\{1,\ldots,\overline{t}^{\theta}\} if he could. The following result provides a remedy:

Theorem 4.

Assume tH>tLt^{H}>t^{L}. The second best can also be implemented using a menu of bonus contracts. Specifically, the principal offers type LL the bonus contract 𝐂L=(t¯L,W0L,𝐛L)\mathbf{C}^{L}=(\overline{t}^{L},W^{L}_{0},\bm{b}^{L}) wherein for any t{1,,t¯L}t\in\{1,\ldots,\overline{t}^{L}\},

btL=s=tt¯Lδst(lsL),b^{L}_{t}=\sum_{s=t}^{\overline{t}^{L}}\delta^{s-t}(-l_{s}^{L}), (7)

where 𝐥L\bm{l}^{L} is the penalty sequence in the optimal penalty contract given in Theorem 3, and W0LW^{L}_{0} is chosen to make the participation constraint, (IRL), bind. For type HH, the principal can use a constant-bonus contract 𝐂H=(tH,W0H,bH)\mathbf{C}^{H}=(t^{H},W^{H}_{0},b^{H}) with a suitably chosen W0HW^{H}_{0} and bH>0b^{H}>0.

Generically, the above contract is the unique optimal contract for type LL within the class of bonus contracts. This implementation satisfies interim participation constraints in each period for each type, i.e. each type θ\theta’s continuation utility at the beginning of any period t{1,,t¯θ}t\in\{1,\ldots,\overline{t}^{\theta}\} in 𝐂θ\mathbf{C}^{\theta} is non-negative.

A proof is omitted because the proof of Proposition 1 can be used to verify that each bonus contract in Theorem 4 is equivalent to the corresponding penalty contract in Theorem 3, and hence the optimality of those penalty contracts implies the optimality of these bonus contracts. Using (6), it is readily verified that in the bonus sequence (7),

bt¯LL=cβ¯t¯LLλL and btL=(1δ)cβ¯tLλL+δbt+1Lfor any t{1,,t¯L1},b^{L}_{\overline{t}^{L}}=\frac{c}{\overline{\beta}^{L}_{\overline{t}^{L}}% \lambda^{L}}\text{ and }b^{L}_{t}=\frac{(1-\delta)c}{\overline{\beta}^{L}_{t}% \lambda^{L}}+\delta b^{L}_{t+1}\ \text{for any $t\in\{1,\ldots,\overline{t}^{L% }-1\}$}, (8)

and hence the reward for success increases over time. When δ=1\delta=1, the low type’s bonus contract is a constant-bonus contract, analogous to the penalty contract in Theorem 3 being a onetime-penalty contract.

An interpretation of the bonus contracts in Theorem 4 is that the principal initially sells the project to the agent at some price (the up-front transfer W0W_{0}) with a commitment to buy back the output generated by a success at time-dated future prices (the bonuses 𝒃\bm{b}).

5.2 Sketch of the proof

We now sketch in some detail how we prove Theorem 3. The arguments reveal how the interaction of adverse selection, dynamic moral hazard, and private learning jointly shape optimal contracts. This subsection also serves as a guide to follow the formal proof in Appendix B.

While we have defined a contract as 𝐂=(T,W0,𝒃,𝒍)\mathbf{C}=(T,W_{0},\bm{b},\bm{l}), it will be useful in this subsection alone (so as to parallel the formal proof) to consider a larger space of contracts, where a contract is given by 𝐂=(Γ,W0,𝒃,𝒍)\mathbf{C}=\left(\Gamma,W_{0},\bm{b},\bm{l}\right). The first element here is a set of periods, Γ{0}\Gamma\subseteq\mathbb{N}\setminus\{0\}, at which the agent is not “locked out,” i.e. at which he is allowed to choose whether to work or shirk. As discussed in fn. 14, this additional instrument does not yield the principal any benefit, but it will be notationally convenient in the proof. The termination date of the contract is now 0 if Γ=\Gamma=\emptyset and otherwise maxΓ\max\Gamma. We say that a contract is connected if Γ={1,,T}\Gamma=\{1,\ldots,T\} for some TT; in this case we refer to TT as the length of the contract, and TT is also the termination date. The agent’s actions are denoted by 𝐚=(at)tΓ\mathbf{a}=\left(a_{t}\right)_{t\in\Gamma}.

As justified by Proposition 1, we solve the principal’s problem (stated at the outset of Section 4) by restricting attention to menus of penalty contracts: for each θ{L,H}\theta\in\{L,H\}, 𝐂θ=(Γθ,W0θ,𝒍θ)\mathbf{C}^{\theta}=(\Gamma^{\theta},W^{\theta}_{0},\bm{l}^{\theta}). Penalty contracts are analytically convenient to deal with the combination of adverse selection and dynamic moral hazard for reasons explained in Step 4 below.

Step 1: We simplify the principal’s program by (i) focussing on contracts for type LL that induce him to work in every non-lockout period, i.e. on contracts in the set {𝐂L:𝟏𝜶L(𝐂L)}\{\mathbf{C}^{L}:\mathbf{1}\in\bm{\alpha}^{L}(\mathbf{C}^{L})\}; and (ii) ignoring the constraints (IRH) and (ICLH). It is established in the proof of Theorem 2 that a solution to this simplified program also solves the original program.2727 27 The idea for (i) is as follows: fix any contract, 𝐂L\mathbf{C}^{L}, in which there is some period, tΓLt\in\Gamma^{L}, such that it would be suboptimal for type LL to work in period tt. Since type LL will not succeed in period tt, one can modify 𝐂L\mathbf{C}^{L} to create a new contract, 𝐂^L\widehat{\mathbf{C}}^{L}, in which tΓ^Lt\notin\widehat{\Gamma}^{L}, and ltLl^{L}_{t} is “shifted up” by one period with an adjustment for discounting. This ensures that the incentives for type LL in all other periods remain unchanged, and critically, that no matter what behavior would have been optimal for type HH under contract 𝐂L\mathbf{C}^{L}, the new contract is less attractive to type HH. As for (ii), we show that type HH always has an optimal action plan under contract 𝐂L\mathbf{C}^{L} that yields him a higher payoff than that of type LL under 𝐂L\mathbf{C}^{L}, and hence (IRH) is implied by (ICHL) and (IRL). Finally, we show that (ICLH) can always be satisfied while still satisfying the other constraints in the principal’s program by making the high type’s contract “risky enough” to deter the low type from taking it. Call this program [P1].

It is not obvious a priori what action plan the high type may use when taking the low type’s contract. Accordingly, we tackle a relaxed program, [RP1], that replaces (ICHL) in program [P1] by a relaxed version, called (Weak-ICHL), that only requires type HH to prefer taking his contract and following an optimal action plan over taking type LL’s contract and working in every period. Formally, (ICHL) requires U0H(𝐂H,𝜶H(𝐂H))U0H(𝐂L,𝜶H(𝐂L))U_{0}^{H}(\mathbf{C}^{H},\bm{\alpha}^{H}(\mathbf{C}^{H}))\geq U_{0}^{H}(% \mathbf{C}^{L},\bm{\alpha}^{H}(\mathbf{C}^{L})) whereas (Weak-ICHL) requires only U0H(𝐂H,𝜶H(𝐂H))U0H(𝐂L,𝟏)U_{0}^{H}(\mathbf{C}^{H},\bm{\alpha}^{H}(\mathbf{C}^{H}))\geq U_{0}^{H}(% \mathbf{C}^{L},\mathbf{1}). We emphasize that this restriction on type HH’s action plan under type LL’s contract is not without loss for an arbitrary contract 𝐂L\mathbf{C}^{L}; i.e., given an arbitrary 𝐂L\mathbf{C}^{L} with 𝟏𝜶L(𝐂L)\mathbf{1}\in\bm{\alpha}^{L}(\mathbf{C}^{L}), it need not be the case that 𝟏𝜶H(𝐂L)\mathbf{1}\in\bm{\alpha}^{H}(\mathbf{C}^{L})—it is in this sense that there is no “single-crossing property” in general. The reason is that because of their differing probabilities of success from working in future periods (conditional on the good state), the two types trade off current and future penalties differently when considering exerting effort in the current period. In particular, the desire to avoid future penalties provides more of an incentive for the low type to work in the current period than the high type.2828 28 To substantiate this point, consider any two-period penalty contract under which it is optimal for both types to work in each period. It can be verified that changing the first-period penalty by ε1>0\varepsilon_{1}>0 while simultaneously changing the second period penalty by ε2<0-\varepsilon_{2}<0 would preserve type θ\theta’s incentive to work in period one if and only if ε1(1λθ)δε2\varepsilon_{1}\leq(1-\lambda^{\theta})\delta\varepsilon_{2}. Note that because ε2<0-\varepsilon_{2}<0, both types will continue to work in period two independent of their action in period one. Consequently, the initial contract can always be modified in a way that preserves optimality of working in both periods for the low type, but makes it optimal for the high type to shirk in period one and work in period two.

Relaxing (ICHL) to (Weak-ICHL) is motivated by a conjecture that even though the high type may choose to work less than the low type in an arbitrary contract, this will not be the case in an optimal contract for the low type. This relaxation is a critical step in making the program tractable because it severs the knot in the fixed point problem of optimizing over the low type’s contract while not knowing what action plan the high type would follow should he take this contract. The relaxation works because of the efficiency ordering tH>tLt^{H}>t^{L}, as elaborated subsequently.

In the relaxed program [RP1], it is straightforward to show that (Weak-ICHL) and (IRL) must bind at an optimum: otherwise, time-zero transfers in one of the two contracts can be profitably lowered without violating any of the constraints. Consequently, one can substitute from the binding version of these constraints to rewrite the objective function as the sum of total surplus less an information rent for the high type, as in the standard approach. We are left with a relaxed program, [RP2], which maximizes this objective function and whose only constraints are the direct moral hazard constraints (ICHa{}_{a}^{H}) and (ICLa{}_{a}^{L}), where type LL must work in all periods. This program is tractable because it can be solved by separately optimizing over each type’s penalty contract. The following steps 2–5 derive an optimal contract for type LL in program [RP2] that has useful properties.

Step 2: We show that there is an optimal penalty contract for type LL that is connected. A rough intuition is as follows.2929 29 For the intuition that follows, assume that all penalties being discussed are negative transfers, i.e. transfers from the agent to the principal. Because type LL is required to work in all non-lockout periods, the value of the objective function in program [RP2] can be improved by removing any lockout periods in one of two ways: either by “shifting up” the sequence of effort and penalties or by terminating the contract early (suitably adjusting for discounting in either case). Shifting up the sequence of effort and penalties eliminates inefficient delays in type LL’s experimentation, but it also increases the rent given to type HH, because the penalties—which are more likely to be borne by type LL than type HH—are now paid earlier. Conversely, terminating the contract early reduces the rent given to type HH by lowering the total penalties in the contract, but it also shortens experimentation by type LL. It turns out that either of these modifications may be beneficial to the principal, but at least one of them will be if the initial contract is not connected.

Step 3: Given any termination date TLT^{L}, there are many penalty sequences that can be used by a connected penalty contract of length TLT^{L} to induce the low-ability agent to work in each period 1,,TL1,\ldots,T^{L}. We construct the unique sequence, call it 𝒍¯(TL)\overline{\bm{l}}(T^{L}), that ensures the low type’s incentive constraint for effort binds in each period of the contract, i.e. in any period t{1,,TL}t\in\{1,\ldots,T^{L}\}, the low type is indifferent between working (and then choosing any optimal effort profile in subsequent periods) and shirking (and then choosing any optimal effort profile in subsequent periods), given the past history of effort. The intuition is straightforward: in the final period, TLT^{L}, there is obviously a unique such penalty as it must solve l¯TLL(TL)=c+(1β¯TLLλL)l¯TLL(TL)\overline{l}^{L}_{T^{L}}(T^{L})=-c+(1-\overline{\beta}^{L}_{T^{L}}\lambda^{L})% \overline{l}^{L}_{T^{L}}(T^{L}). Iteratively working backward using a one-step deviation principle, this pins down penalties in each earlier period through the (forward-looking) incentive constraint for effort in each period. Naturally, for any TLT^{L} and t{1,,TL}t\in\{1,\ldots,T^{L}\}, l¯tL(TL)<0\overline{l}^{L}_{t}(T^{L})<0, i.e. as suggested by the term “penalty”, the agent pays the principal each time there is a failure.

Step 4: We show that any connected penalty contract for type LL that solves program [RP2] must use the penalty structure 𝒍¯L()\overline{\bm{l}}^{L}(\cdot) of Step 3. The idea is that any slack in the low type’s incentive constraint for effort in any period can be used to modify the contract to strictly reduce the high type’s expected payoff from taking the low type’s contract (without affecting the low type’s behavior or expected payoff), based on the high type succeeding with higher probability in every period when taking the low type’s contract.3030 30 This is because the constraint (Weak-ICHL) in program [RP2] effectively constrains the high type in this way, even though, as previously noted, it may not be optimal for the high type to work in each period when taking an arbitrary contract for the low type.

Although this logic is intuitive, a formal argument must deal with the challenge that modifying a transfer in any period to reduce slack in the low type’s incentive constraint for effort in that period has feedback on incentives in every prior period—the dynamic agency problem. Our focus on penalty contracts facilitates the analysis here because penalty contracts have the property that reducing the incentive to exert effort in any period tt by decreasing the severity of the penalty in period tt has a positive feedback of also reducing the incentive for effort in earlier periods, since the continuation value of reaching period tt increases. Due to this positive feedback, we are able to show that the low type’s incentive for effort in a given period of a connected penalty contract can be modified without affecting his incentives in any other period by solely adjusting the penalties in that period and the previous one. In particular, in an arbitrary connected penalty contract 𝐂L\mathbf{C}^{L}, if type LL’s incentive constraint is slack in some period tt, we can increase ltLl^{L}_{t} and reduce lt1Ll^{L}_{t-1} in a way that leaves type LL’s incentives for effort unchanged in every period sts\neq t while still being satisfied in period tt. We then verify that this “local modification” strictly reduces the high type’s information rent.3131 31 By contrast, bonuses have a negative feedback: reducing the bonus in a period tt increases the incentive to work in prior periods because the continuation value of reaching period tt decreases. Consequently, keeping incentives for effort in earlier periods unchanged after reducing the bonus in period tt would require a “global modification” of reducing the bonus in all prior periods, not just the previous period. This makes the analysis with bonus contracts less convenient.

Step 5: In light of Steps 2–4, all optimal connected penalty contracts for type LL in program [RP2] can be found by just optimizing over the length of connected penalty contracts with the penalty structure 𝒍¯L()\overline{\bm{l}}^{L}(\cdot). By Theorem 2, the optimal length, t¯L\overline{t}^{L}, cannot be larger than the first-best stopping time: t¯LtL\overline{t}^{L}\leq t^{L}. In this step, we further establish that t¯L\overline{t}^{L} is generically unique, and that generically there is no optimal penalty contract for type LL that is not connected.

Step 6: Let 𝐂¯L\overline{\mathbf{C}}^{L} be the contract for type LL identified in Steps 2–5.3232 32 The initial transfer in 𝐂¯L\overline{\mathbf{C}}^{L} is set to make the participation constraint for type LL bind. In the non-generic cases where there are multiple optimal lengths of contract, 𝐂¯L\overline{\mathbf{C}}^{L} uses the largest one. Recall that [RP1] differs from the principal’s original program [P1] in that it imposes (Weak-ICHL) rather than (ICHL). In this step, we show that any solution to [RP1] using 𝐂¯L\overline{\mathbf{C}}^{L} satisfies (ICHL) and hence is also a solution to program [P1]. Specifically, we show that 𝜶H(𝐂¯L)=𝟏\bm{\alpha}^{H}(\overline{\mathbf{C}}^{L})=\mathbf{1}, i.e. if type HH were to take contract 𝐂¯L\overline{\mathbf{C}}^{L}, it would be uniquely optimal for him to work in all periods 1,,t¯L1,\ldots,\overline{t}^{L}. The intuition is as follows: under contract 𝐂¯L\overline{\mathbf{C}}^{L}, type HH has a higher expected probability of success from working in any period tt¯Lt\leq\overline{t}^{L}, no matter his prior choices of effort, than does type LL in period tt given that type LL has exerted effort in all prior periods (recall 𝟏𝜶L(𝐂¯L)\mathbf{1}\in\bm{\alpha}^{L}(\overline{\mathbf{C}}^{L})). The argument relies on Theorem 2 having established that t¯LtL\overline{t}^{L}\leq t^{L}, because tH>tLt^{H}>t^{L} then implies that for any t{1,,t¯L}t\in\{1,\ldots,\overline{t}^{L}\}, βtHλH>β¯tLλL\beta^{H}_{t}\lambda^{H}>\overline{\beta}^{L}_{t}\lambda^{L} for any history of effort by type HH in periods 1,,t11,\ldots,t-1. Using this property, we verify that because 𝐂¯L\overline{\mathbf{C}}^{L} makes type LL indifferent between working and shirking in each period up to t¯L\overline{t}^{L} (given that he has worked in all prior periods), type HH would find it strictly optimal to work in each period up to t¯L\overline{t}^{L} no matter his prior history of effort.

5.3 Implications and applications

Asymmetric information and success.

Our results offer predictions on the extent of experimentation and innovation. An immediate implication concerns the effects of asymmetric information. Compare a setting with either no moral hazard or no adverse selection, as in Theorem 1, with a setting where both features are present, as in Theorem 2. The theorems reveal that, other things equal, the amount of experimentation will be lower in the latter, and, consequently, the average probability of success will also be lower. Furthermore, because low-ability agents’ experimentation is typically distorted down whereas that of high-ability agents is not, we predict a larger dispersion in success rates across agents and projects when both forms of asymmetric information are present.3333 33 While this is readily evident when tH>tLt^{H}>t^{L}, it is also true when tHtLt^{H}\leq t^{L}. In the latter case, even though the second best may narrow the gap in the types’ duration of experimentation, the gap in their success rates widens.

Our analysis also bears on the relationship between innovation rates and the quality of the underlying environment. Absent any distortions, “better environments” lead to more success. In particular, an increase in the agent’s average ability, μ0λH+(1μ0)λL\mu_{0}\lambda^{H}+(1-\mu_{0})\lambda^{L}, yields a higher probability of success in the first best.3434 34 Although tLt^{L} and tHt^{H} may increase or decrease in λL\lambda^{L} and λH\lambda^{H} respectively, one can show that the first-best probability of success is always increasing in μ0\mu_{0}, λL\lambda^{L}, and λH\lambda^{H}. However, contracts designed in the presence of moral hazard and adverse selection need not produce this property. The reason is that an improvement in the agent’s average ability can make it optimal for the principal to distort experimentation by more: as shown in Proposition 2, t¯L\overline{t}^{L} decreases in μ0\mu_{0} and, for some parameter values, in λH\lambda^{H}. Such a reduction in t¯L\overline{t}^{L} can decrease the second-best average success probability when the agent’s average ability increases. Consequently, observing higher innovation rates in contractual settings is neither necessary nor sufficient to deduce a better underlying environment.

Contract farming and technology adoption.

Though our model is not developed to explain a particular application, our framework speaks to contract farming and, more broadly, technology adoption in developing countries. Technology adoption is inherently a dynamic process of experimentation and learning. Understanding the adoption of agricultural innovations in low-income countries, and the obstacles to it, has been a central topic in development economics (Feder, Just, and Zilberman, 1985; Foster and Rosenzweig, 2010). Practitioners, policymakers, and researchers have long recognized the importance of contractual arrangements to provide proper incentives, because farmers typically don’t internalize the broader benefits of their experimentation.

As described in the Introduction, contract farming is a common practice in developing countries; it involves a profit-maximizing firm, which is typically a large-scale buyer such as an exporter or a food processor, and a farmer, who may be a small or a large grower. The contractual environment features not only learning about the quality of new seeds or a new technology, but also moral hazard and unobservable heterogeneity (Miyata et al., 2009).3535 35 Using a field experiment, Kelsey (2013) shows that landholders have private information relevant to their performance under a contract that offers incentives for afforestation, and that efficiency can be increased by using an allocation mechanism that induces self-selection. The arrangements used between agricultural firms and farmers resemble the contracts characterized in Theorem 4, with firms committing to time-dated future prices for an output of a certain quality delivered by a given deadline. Our analysis shows why such contracts are optimal in the presence of uncertainty, moral hazard, and unobservable heterogeneity, and how the shape of the contract hinges on the interaction of these three key features. Theorem 4 and formula (8) reveal how an optimal pattern of outcome-contingent buyback prices should be determined. In principle, these predicted contracts could be subject to empirical testing.

Much of the recent research on technology adoption uses controlled field experiments to study the incentives of potential adopters. Our results may inform the design of experimental work, particularly with regards to dynamic considerations, which are receiving increasing attention. For example, Jack et al. (2014) use a field experiment to study both the initial take-up decision and the subsequent investment (follow-through) decisions in the context of agricultural technology (tree species) adoption in Zambia. The authors consider simple contracts to investigate the interplay between the uncertainty of a technology’s profitability, the self-selection of farmers, and learning of new information. In their experimental design, contracts specify the initial price of the technology and an outcome-contingent payment tied to the survival of trees by the end of one year. The study uses variation of the contracts in the two dimensions (initial price and contingent payment) to evaluate their performance. The authors find that 35% of farmers who pay a positive price for take-up have no trees one year later; in addition, among farmers who follow-through, the tree survival rate responds to learning over time.

The contract form used in Jack et al. (2014) shares features with what emerges as an optimal contract in our model, and their basic findings are also consistent with our results. Their controlled experiment is simple in that performance is assessed and a reward is paid only at the end of one year. Our model shows that to optimally incentivize experimentation, agents must be compensated with continual rewards contingent on the time of success, up until an optimally chosen termination date which may differ from the efficient stopping time. Moreover, perhaps counterintuitively, Theorem 4 shows that higher rewards must be offered for later success, with the rate of increase depending on the rate of learning (and other factors).3636 36 In particular, formula (8) reveals that rewards will optimally increase more sharply over time, up until the contract termination, if the rate of learning is higher. Our results thus point to a new dimension that can improve follow-through rates; this could be tested in future field experiments.

Finally, many scholars study the puzzle of low technology adoption rates and its potential solutions (e.g., Suri, 2011, and the references therein). Our paper adds to the discussion by relating adoption rates to the underlying contractual environment. As mentioned earlier, we predict less experimentation, lower success rates, and more dispersion of success rates in settings with more asymmetric information; the lower (and more dispersed) success rates translate into lower (and more dispersed) adoption rates. We also find that the relationship between adoption rates and the underlying environment can be subtle, with “better environments” possibly leading to less experimentation and lower adoption in the second best. Our results thus provide a novel explanation for the low adoption rate puzzle. Empirical researchers have recently been interested in how agency contributes to the puzzle (e.g., Atkin et al., 2015); our work contributes to the theoretical background for such lines of inquiry.

Naturally, there are dimensions of contract farming and technology adoption that our analysis does not cover. For example, social learning among farmers affects adoption (Conley and Udry, 2010), and agricultural companies will want to take this into account when designing contracts.3737 37 Another aspect is the choice of farmer size: as discussed in Miyata et al. (2009), there are different advantages to contracting with small versus large growers, and the optimal farmer size for a firm may change as parties experiment and learn over time. A deeper understanding of optimal contracts for multiple experimenting agents who can learn from each other would be useful for this application.3838 38 Recent work on this agenda, albeit without adverse selection, includes Frick and Ishii (2015) and Moroni (2015). While this and similar extensions may yield new insights, we expect our main results to be robust: to reduce the information rent of high-ability types, the principal will benefit from distorting the length of experimentation of low-ability types, and from setting payments so that their incentive constraint for effort binds at each time. This suggests that, under appropriate conditions, an agent will still receive a higher reward for succeeding later rather than earlier.

Book contracts.

As mentioned in the Introduction, some contractual relationships between a publisher and author have the features we study: it is initially uncertain whether a satisfactory book can be written in the relevant timeframe; the author may be privately informed about his suitability for the task; and how much time the author spends on this is not observable to the publisher. It is common for real-world publishing contracts to resemble the penalty contracts characterized in Theorem 3: book contracts pay an advance to the author that the publisher can recoup if the author fails to deliver on time (according to a delivery-of-manuscript clause) or if the book is unacceptable (according to a satisfactory-manuscript clause); see Bunnin (1983) and Fowler (1985). There is substantial dispersion in both the deadlines and the advances that authors are given; Kuzyk,Raya (2006) notes that publishing houses try to assess an author’s chances of succeeding when determining these terms.

6 Optimal Contracts when tHtLt^{H}\leq t^{L}

We now turn to characterizing optimal contracts when the first-best stopping times are ordered tHtLt^{H}\leq t^{L}. Any of the following conditions on the primitives is sufficient for this case given a set of other parameters: (i) β0\beta_{0} is large enough, (ii) λH\lambda^{H} is large enough, or (iii) cc is small enough.

The principal’s program remains as described in Section 4, but solving the program is now substantially more difficult than when tH>tLt^{H}>t^{L}. To understand why, consider Figure 3, which depicts the two types’ “no-shirk expected marginal product” curves, β¯tθλθ\overline{\beta}^{\theta}_{t}\lambda^{\theta}, as a function of time. (For simplicity, the figure is drawn ignoring integer constraints.) For any parameters, these curves cross exactly once as shown in the figure; the crossing point tt^{*} is the unique solution to

β¯tHλHβ¯tLλL0>β¯t+1HλHβ¯t+1LλL.\mbox{$\overline{\beta}^{H}_{t^{*}}\lambda^{H}-\overline{\beta}^{L}_{t^{*}}% \lambda^{L}\geq 0>\overline{\beta}^{H}_{t^{*}+1}\lambda^{H}-\overline{\beta}^{% L}_{t^{*}+1}\lambda^{L}$}.

Parameters under which tH>tLt^{H}>t^{L} entail tL<tt^{L}<t^{*}, as seen with the high effort cost in Figure 3. When tLtt^{L}\leq t^{*}, it holds at any ttLt\leq t^{L} that the high type has a higher expected marginal product than the low type conditional on the agent working in all prior periods. It is this fact that allowed us to prove Theorem 3 by conjecturing that the high type would work in every period when taking the low type’s contract.

By contrast, tHtLt^{H}\leq t^{L} implies tLtt^{L}\geq t^{*}, as seen with the low effort cost in Figure 3. Since the second-best stopping time for the low type can be arbitrarily close to his first-best stopping time (e.g. if the prior on the low type, 1μ01-\mu_{0}, is sufficiently large), it is no longer valid to conjecture that the high type will work in every period when taking the low type’s optimal contract—in this sense, “single crossing” need not hold even at the optimum. The reason is that at some period after tt^{*}, given that both types have worked in each prior period, the high type can be sufficiently more pessimistic than the low type that the high type finds it optimal to shirk in some or all of the remaining periods, even though λH>λL\lambda^{H}>\lambda^{L} and the low type would be willing to work for the contract’s duration.3939 39 More precisely, the relaxed program, [RP1], described in Step 1 of the proof sketch of Theorem 3 can yield a solution that is not feasible in the original program, because the constraint (ICHL) is violated; the high type would deviate from accepting his contract to accepting the low type’s contract and then shirk in some periods. Indeed, this will necessarily be true in the last period of the low type’s contract if this period is later than tt^{*} and the contract makes the low type just indifferent between working and shirking in this period as in the characterization of Theorem 3.

Refer to caption
Figure 3: No-shirk expected marginal product curves with β0=0.99,λL=0.28,λH=0.35\beta_{0}=0.99,\lambda^{L}=0.28,\lambda^{H}=0.35.

Solving the principal’s program without being able to restrict attention to some suitable subset of action plans for the high type when he takes the low type’s contract appears intractable. For an arbitrary δ\delta, we have been unable to find a valid restriction. The following example elucidates the difficulties.

Example 1.

For an open and dense set of parameters {β0,c,λL,λH}\{\beta_{0},c,\lambda^{L},\lambda^{H}\} with t¯L=tL=3\overline{t}^{L}=t^{L}=3,4040 40 It suffices for the parameters to satisfy the following four conditions: 1. The first-best stopping time for type LL is tL=3t^{L}=3 (i.e., β¯3LλL>c>β¯4LλL\overline{\beta}^{L}_{3}\lambda^{L}>c>\overline{\beta}^{L}_{4}\lambda^{L}) and the probability of type LL is large enough (i.e., μ0\mu_{0} is sufficiently small) that it is not optimal to distort the stopping time of type LL: t¯L=tL=3\overline{t}^{L}=t^{L}=3. 2. The expected marginal product for type HH after one period of work is less than that of type LL after one period of work, but larger than that of type LL after two periods of work: β¯3LλL<β¯2HλH<β¯2LλL.\overline{\beta}^{L}_{3}\lambda^{L}<\overline{\beta}^{H}_{2}\lambda^{H}<% \overline{\beta}^{L}_{2}\lambda^{L}. 3. Ex-ante, type HH is more likely to succeed by working in one period than type LL is by working in two periods: 1λH<(1λL)2.1-\lambda^{H}<(1-\lambda^{L})^{2}. 4. There is some δ(0,1)\delta^{*}\in(0,1) such that 1β0λL1β0λH=δ(1λL)(1β¯2HλH1β¯2LλL).\frac{1}{\beta_{0}\lambda^{L}}-\frac{1}{\beta_{0}\lambda^{H}}=\delta^{*}(1-% \lambda^{L})\left(\frac{1}{\overline{\beta}^{H}_{2}\lambda^{H}}-\frac{1}{% \overline{\beta}^{L}_{2}\lambda^{L}}\right). there is a δ(0,1)\delta^{*}\in(0,1) such that the optimal penalty contract for type LL as a function of the discount factor, 𝐂L(δ)=(3,W0L(δ),𝐥L(δ))\mathbf{C}^{L}(\delta)=(3,W^{L}_{0}(\delta),\bm{l}^{L}(\delta)), has the property that the optimal action plans for type HH under this contract are given by

𝜶H(𝐂L(δ))={{(1,1,0),(1,0,1)}if δ(0,δ){(1,1,0),(1,0,1),(0,1,1)}if δ=δ{(1,0,1),(0,1,1)}if δ(δ,1){(1,1,0),(1,0,1),(0,1,1)}if δ=1.\bm{\alpha}^{H}(\mathbf{C}^{L}(\delta))=\begin{cases}\{(1,1,0),(1,0,1)\}&\mbox% {if }\delta\in(0,\delta^{*})\\ \{(1,1,0),(1,0,1),(0,1,1)\}&\mbox{if }\delta=\delta^{*}\\ \{(1,0,1),(0,1,1)\}&\mbox{if }\delta\in(\delta^{*},1)\\ \{(1,1,0),(1,0,1),(0,1,1)\}&\mbox{if }\delta=1.\\ \end{cases}

Figure 4 depicts the contract and type HH’s optimal action plans as a function of δ\delta for a particular set of other parameters.4141 41 The initial transfer W0LW^{L}_{0} in each case is determined by making the participation constraint of type LL bind. Notice that the only action plan that is optimal for type HH for all δ\delta is the non-consecutive-work plan (1,0,1)(1,0,1), but for each value of δ\delta at least one other plan is also optimal. Interestingly, the stopping strategy (1,1,0)(1,1,0) is not optimal for type HH when δ(δ,1)\delta\in(\delta^{*},1) although it is when δ=1\delta=1. The lack of lower hemi-continuity of 𝛂H(𝐂L(δ))\bm{\alpha}^{H}(\mathbf{C}^{L}(\delta)) at δ=1\delta=1 is not an accident, as we will discuss subsequently.

Refer to caption
Figure 4: The optimal penalty contract for type LL in Example 1 with β0=0.86\beta_{0}=0.86, c=0.1c=0.1, λL=0.75\lambda^{L}=0.75, λH=0.95\lambda^{H}=0.95 (left graph) and the optimal action profiles for type HH under this contract (right graph).

Nevertheless, we are able to solve the problem when δ=1\delta=1.

Theorem 5.

Assume δ=1\delta=1 and tHtLt^{H}\leq t^{L}. There is an optimal menu in which the principal separates the two types using onetime-penalty contracts, 𝐂H=(tH,W0H,ltHH)\mathbf{C}^{H}=(t^{H},W^{H}_{0},l^{H}_{t^{H}}) with ltHH<0<W0Hl^{H}_{t^{H}}<0<W^{H}_{0} for type HH and 𝐂L=(t¯L,W0L,lt¯LL)\mathbf{C}^{L}=(\overline{t}^{L},W^{L}_{0},l^{L}_{\overline{t}^{L}}) for type LL, such that:

  1. 1.

    lt¯LL=min{cβ¯t¯LLλL,cβ¯tHLHλH}l^{L}_{\overline{t}^{L}}=\min\left\{-\frac{c}{\overline{\beta}^{L}_{\overline{% t}^{L}}\lambda^{L}},-\frac{c}{\overline{\beta}^{H}_{t^{HL}}\lambda^{H}}\right\}, where tHL:=max𝐚𝜶H(𝐂L)#{n:an=1}t^{HL}:=\max\limits_{\mathbf{a}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right)}% \#\left\{n:a_{n}=1\right\};

  2. 2.

    W0L>0W^{L}_{0}>0 is such that the participation constraint, (IRL), binds;

  3. 3.

    Type HH gets an information rent: U0H(𝐂H,𝜶H(𝐂H))>0U^{H}_{0}(\mathbf{C}^{H},\bm{\alpha}^{H}(\mathbf{C}^{H}))>0;

  4. 4.

    𝟏𝜶H(𝐂H)\mathbf{1}\in\bm{\alpha}^{H}(\mathbf{C}^{H}); 𝟏𝜶L(𝐂L)\mathbf{1}\in\bm{\alpha}^{L}(\mathbf{C}^{L}).

Proof.

See Appendix C. ∎

For δ=1\delta=1, the optimal menus of penalty contracts characterized in Theorem 5 for tHtLt^{H}\leq t^{L} share some common properties with those characterized in Theorem 3 for tH>tLt^{H}>t^{L}: in both cases, a onetime-penalty contract is used for the low type and the high type earns an information rent. On the other hand, part 1 of Theorem 5 points to two differences: (i) it will generally be the case in the optimal 𝐂L\mathbf{C}^{L} that when tHtLt^{H}\leq t^{L}, 𝟏𝜶H(𝐂L)\mathbf{1}\notin\bm{\alpha}^{H}(\mathbf{C}^{L}), whereas for tH>tLt^{H}>t^{L}, 𝜶H(𝐂L)=𝟏\bm{\alpha}^{H}(\mathbf{C}^{L})=\mathbf{1}; and (ii) when tHtLt^{H}\leq t^{L}, it can be optimal for the principal to induce the low type to work in each period by satisfying the low type’s incentive constraint for effort with slack (i.e. with strict inequality), whereas when tH>tLt^{H}>t^{L}, the penalty sequence makes this effort constraint bind in each period.

The intuition for these differences derives from information rent minimization considerations. The high type earns an information rent because by following the same effort profile as the low type he is less likely to incur any penalty for failure, and hence has a higher utility from any penalty contract than the low type.4242 42 Strictly speaking, this intuition applies so long as lt0l_{t}\leq 0 for all tt in the penalty contract. Minimizing the rent through this channel suggests minimizing the magnitude of the penalties that are used to incentivize the low type’s effort; it is this logic that drives Theorem 3 and for δ=1\delta=1 leads to a onetime-penalty contract with

lt¯LL=cβ¯t¯LLλL.l^{L}_{\overline{t}^{L}}=-\frac{c}{\overline{\beta}^{L}_{\overline{t}^{L}}% \lambda^{L}}. (9)

However, when t¯L>t\overline{t}^{L}>t^{*} (which is only possible when tHtLt^{H}\leq t^{L}), the high type would find it optimal under this contract to work only for some T<t¯LT<\overline{t}^{L} number of periods. It is then possible—and is true for an open and dense set of parameters—that TT is such that the high type is more likely to incur the onetime penalty than the low type. But in such a case, the penalty given in (9) would not be optimal because the principal can lower lt¯LLl^{L}_{\overline{t}^{L}} (i.e. increase the magnitude of the penalty) to reduce the information rent, which she can keep doing until the high type finds it optimal to work for more periods and becomes less likely to incur the onetime penalty than the low type. This explains part 1 of Theorem 5.

We should note that this possibility arises because time is discrete. It can be shown that when the length of time intervals vanishes, in real-time the tHLt^{HL} and t¯L\overline{t}^{L} in the statement of Theorem 5 are such that β¯t¯LLλLβ¯tHLHλH\overline{\beta}^{L}_{\overline{t}^{L}}\lambda^{L}\leq\overline{\beta}^{H}_{t^% {HL}}\lambda^{H} (in particular, β¯t¯LLλL=β¯tHLHλH\overline{\beta}^{L}_{\overline{t}^{L}}\lambda^{L}=\overline{\beta}^{H}_{t^{HL% }}\lambda^{H} when t¯L>tHL\overline{t}^{L}>t^{HL}, or equivalently when t¯L>t\overline{t}^{L}>t^{*}), and hence lt¯LL=cβ¯t¯LLλLl^{L}_{\overline{t}^{L}}=-\frac{c}{\overline{\beta}^{L}_{\overline{t}^{L}}% \lambda^{L}} is optimal, just as in Theorem 3 when δ=1\delta=1. Intuitively, because learning is smooth in continuous time, the high type would always work long enough upon deviating to the low type’s contract that he is less likely to incur the onetime penalty lt¯LLl^{L}_{\overline{t}^{L}} than the low type. Thus, by the logic above, lowering the onetime penalty below that in (9) would only increase the information rent of the high type in the continuous-time limit.

Remark 2.

The proof of Theorem 5 provides an algorithm to solve for an optimal menu of contracts when tHtLt^{H}\leq t^{L} and δ=1\delta=1. For each pair of integers (s,t)(s,t) such that 0sttL0\leq s\leq t\leq t^{L}, one can compute the principal’s payoff from using the onetime-penalty contract for type LL given by Theorem 5 when t¯L\overline{t}^{L} is replaced by tt and tHLt^{HL} is replaced by ss. Optimizing over (s,t)(s,t) then yields an optimal (unconstrained) menu.

How do we prove Theorem 5 in light of the difficulties described earlier of finding a suitable restriction on the high type’s behavior when taking the low-type’s contract? The answer is that when δ=1\delta=1, one can conjecture that the optimal contract for the low type must be a onetime-penalty contract (as was also true when tH>tLt^{H}>t^{L}). Notice that because of no discounting, any onetime-penalty contract would make the agent of either type indifferent among all action plans that involve the same number of periods of work. In particular, a stopping strategy—an action plan that involves consecutive work for some number of periods followed by shirking thereafter—is always optimal for either type in a onetime-penalty contract. The heart of the proof of Theorem 5 establishes that it is without loss of generality to restrict attention to penalty contracts for the low type under which the high type would find it optimal to use a stopping strategy (see Subsection C.4 in Appendix C). With this in hand, we are then able to show that a onetime-penalty contract for the low type is indeed optimal (see Subsection C.5). Finally, the rent-minimization considerations described above are used to complete the argument. Observe that optimality of a onetime-penalty contract for the low type and that of a stopping strategy for the high type under such a contract is consistent with the solution in Example 1 for δ=1\delta=1, as seen in Figure 4. Moreover, the example plainly shows that such a strategy space restriction will not generally be valid when δ<1\delta<1.4343 43 Due to the agent’s indifference over all action plans that involve the same number of periods of work in a onetime-penalty contract when δ=1\delta=1, the correspondence 𝜶H(𝐂L(δ))\bm{\alpha}^{H}(\mathbf{C}^{L}(\delta)) will generally fail lower hemi-continuity at δ=1\delta=1. In particular, the low type’s optimal contract for δ\delta close to 1 may be such that a stopping strategy is not optimal for the high type under this contract. However, the correspondence 𝜶H(𝐂L(δ))\bm{\alpha}^{H}(\mathbf{C}^{L}(\delta)) is upper hemi-continuous and the optimal contract is continuous at δ=1\delta=1. All these points can be seen in Figure 4.

We provide a bonus-contracts implementation of Theorem 5:

Theorem 6.

Assume δ=1\delta=1 and tHtLt^{H}\leq t^{L}. The second-best can also be implemented using a menu of constant-bonus contracts: 𝐂L=(t¯L,W0L,bL)\mathbf{C}^{L}=(\overline{t}^{L},W^{L}_{0},b^{L}) with bL=lt¯LL>0>W0Lb^{L}=-l^{L}_{\overline{t}^{L}}>0>W^{L}_{0} where lt¯LLl^{L}_{\overline{t}^{L}} is given in Theorem 5, and 𝐂H=(tH,W0H,bH)\mathbf{C}^{H}=(t^{H},W^{H}_{0},b^{H}) with a suitably chosen W0HW^{H}_{0} and bH>0b^{H}>0.

A proof is omitted since this result follows directly from Theorem 5 and the proof of Proposition 1 (using δ=1\delta=1). For similar reasons to those discussed around Theorem 4, the implementation in Theorem 6 satisfies interim participation constraints whereas that of Theorem 5 does not.

We end this section by emphasizing that although we are unable to characterize second-best optimal contracts when δ<1\delta<1 and tHtLt^{H}\leq t^{L}, the (in)efficiency conclusions from Theorem 2 apply for all parameters.

7 Discussion

7.1 Private observability and disclosure

Suppose that project success is privately observed by the agent but can be verifiably disclosed. The principal’s payoff from project success obtains here only when the agent discloses it, and contracts are conditioned not on project success but rather the disclosure of project success. Private observability introduces additional constraints for the principal because the agent must also now be incentivized to not withhold project success. For example, in a bonus contract where δbt+1>bt\delta b_{t+1}>b_{t}, an agent who obtains success in period tt would strictly prefer to withhold it and continue to period t+1t+1, shirk in that period, and then reveal the success at the end of period t+1t+1. Nevertheless, we show in the Supplementary Appendix that private observability does not reduce the principal’s payoff compared to our baseline setting: in each of the menus identified in Theorems 3–6, each of the contracts would induce the agent (of either type) to reveal project success immediately when it is obtained, so these menus remain optimal and implement the same outcome as when project success is publicly observable.4444 44 However, unlike the menus of Theorems 3–6, not every optimal menu under public observability is optimal under private observability. In this sense, these menus have a desirable robustness property that other optimal menus need not.

7.2 Limited liability

To focus on the interaction of adverse selection and moral hazard in experimentation, we have abstracted away from limited-liability considerations. Consider introducing the requirement that all transfers must be above some minimum threshold, say zero. The Supplementary Appendix shows how such a limited-liability constraint alters the second-best solution for the case of tH>tLt^{H}>t^{L} and δ=1\delta=1. This constraint results in both types of the agent acquiring a rent, so long as they are both induced to experiment. Three points are worth emphasizing. First, each type’s second-best stopping time is no larger than his first-best stopping time. The logic precluding over-experimentation, however, is somewhat different—and simpler—than without limited liability: inducing over-experimentation requires paying a bonus of more than one (the principal’s value of success) in the last period of the contract in which the agent works, implying a loss for the principal which under limited liability cannot be offset through an up-front payment. Second, while both types’ second-best stopping times are now (typically) distorted, their ordering is the same as without limited liability (i.e., t¯Lt¯H\overline{t}^{L}\leq\overline{t}^{H}). The reason is that the principal could otherwise improve upon the menu by just offering both types the low type’s contract, which would induce the high type to experiment longer without increasing the high type’s payoff. Third, the principal can implement the second-best stopping time for the low type by using a constant-bonus contract of the form described in Theorem 4 (with δ=1\delta=1). This contract ensures that the low type’s incentive constraint for effort binds in each period, and thus it minimizes both the rent that the low type obtains from his contract and the high type’s payoff from taking the low type’s contract.

We should note that in our dynamic setting, there are less severe forms of limited liability that may be relevant in applications. For example, one may only require that the sum of penalties at any point do not exceed the initial transfer given to the agent.4545 45 Biais et al. (2010) study such a limited-liability requirement in a setting without adverse selection or learning, where large losses arrive according to a Poisson process whose intensity is determined by the agent’s effort. We conjecture that similar conclusions to those discussed above would also emerge under such a requirement, as both types of the agent will again acquire a rent.

7.3 The role of learning

We have assumed that β0(0,1)\beta_{0}\in(0,1). If instead β0=1\beta_{0}=1 then there would be no learning about the project quality and the first best would entail both types working until project success has been obtained. How is the second best affected by β0=1\beta_{0}=1?

Suppose, for simplicity, that there is some (possibly large) exogenous date T¯\overline{T} at which the game ends. The first-best stopping times are then tL=tH=T¯t^{L}=t^{H}=\overline{T}. The principal’s program can be solved here just as in Section 5, because β¯tHλH=λH>β¯tLλL=λL\overline{\beta}^{H}_{t}\lambda^{H}=\lambda^{H}>\overline{\beta}^{L}_{t}% \lambda^{L}=\lambda^{L} for all tT¯t\leq\overline{T}.4646 46 It should be clear that nothing would have changed in the analysis in Section 5 if we had assumed existence of a suitably large end date, in particular so long as T¯max{tH,tL}\overline{T}\geq\max\{t^{H},t^{L}\}. In the absence of learning, the social surplus from the low type working is constant over time. So long as parameters are such that it is not optimal for the principal to exclude the low type (i.e. t¯L>0\overline{t}^{L}>0), it turns out that there is no distortion: t¯L=t¯H=T¯\overline{t}^{L}=\overline{t}^{H}=\overline{T}. We provide a more complete argument in the Supplementary Appendix, but to see the intuition consider a large T¯\overline{T}. Then, even though both types are likely to succeed prior to T¯\overline{T}, the probability of reaching T¯\overline{T} without a success is an order of magnitude higher for the low type because (1λL1λH)t\left(\frac{1-\lambda^{L}}{1-\lambda^{H}}\right)^{t}\rightarrow\infty as tt\rightarrow\infty. Hence, it would not be optimal to locally distort the length of experimentation from T¯\overline{T} because such a distortion would generate a larger efficiency loss from the low type than a gain from reducing the high type’s information rent. By contrast, when β0<1\beta_{0}<1 and there is learning, this logic fails because the incremental social surplus from the low type working vanishes over time. Therefore, learning from experimentation plays an important role in our results: for any parameters with β0<1\beta_{0}<1 under which there is distortion of the low type’s length of experimentation without entirely excluding him, there would instead be no distortion were β0=1\beta_{0}=1.

7.4 Adverse selection on other dimensions

Another important modeling assumption in this paper is that pre-contractual hidden information is about the agent’s ability. An alternative is to suppose that the agent has hidden information about his cost of effort but his ability is commonly known; specifically, the low type’s cost of working in any period is cL>0c^{L}>0 whereas the high type’s cost is cH(0,cL)c^{H}\in(0,c^{L}). It is immediate that the first-best stopping time for the high type would always be larger than that of the low type because there is no speed-of-learning effect. Hence, the problem can be solved following our approach in Section 5 for tH>tLt^{H}>t^{L}.4747 47 This applies to binary effort choices. Another alternative would be for the agent to choose effort from a richer set, e.g. +\mathbb{R}_{+}, and effort costs be convex with one type having a lower marginal cost than the other. The speed-of-learning effect would emerge in this setting because the two types would generally choose different effort levels in any period. Analyzing such a problem is beyond the scope of this paper. However, not only would this alternative model miss the considerations involved with tHtLt^{H}\leq t^{L}, but furthermore, it also obviates interesting features of the problem even when tH>tLt^{H}>t^{L}. For example, in this setting it would be optimal for the high type to work in all periods in any contract in which it is optimal for the low type to work in all periods; recall that this is not true in our model even when tH>tLt^{H}>t^{L} (cf. fn. 28).

Another source of adverse selection would be private information about project quality. Specifically, suppose that the agent’s ability is commonly known but, prior to contracting, he receives a private signal about the true project quality: there is a high type whose belief that the state is good is β0H(0,1)\beta^{H}_{0}\in(0,1) and a low type whose belief is β0L(0,β0H)\beta^{L}_{0}\in(0,\beta^{H}_{0}).4848 48 Private information about project quality is studied by Gomes et al. (2015) in experimentation without moral hazard, and in a different setting by Gerardi and Maestri (2012). Another possibility would be non-common priors between the principal and the agent, which would involve quite distinct considerations. Again, the first-best stopping times here would always have tH>tLt^{H}>t^{L} and the problem can be studied following our approach to this case.

Appendices: Notation and Terminology

It is convenient in proving our results to work with an apparently larger set of contracts than that defined in the main text. Specifically, in the Appendices, we assume that the principal can stipulate binding “lockout” periods in which the agent is prohibited from working. As discussed in fn. 14 of the main text, this instrument does not yield any benefit to the principal because suitable transfers can be used to ensure that the agent shirks in any desired period regardless of his type and action history. Nevertheless, stipulating lockout periods simplifies the phrasing of our arguments; we use it, in particular, to prove that an optimal contract for the low type never induces him to shirk before termination.

Accordingly, we denote a general contract by 𝐂=(Γ,W0,𝒃,𝒍)\mathbf{C}=\left(\Gamma,W_{0},\bm{b},\bm{l}\right), where all the elements are as introduced in the main text, except that instead of having the termination date of the contract in the first component, we now have a set of periods, Γ{0}\Gamma\subseteq\mathbb{N}\setminus\{0\}, at which the agent is not locked out, i.e. at which he is allowed to choose whether to work or shirk. Note that, without loss, 𝒃=(bt)tΓ\bm{b}=(b_{t})_{t\in\Gamma} and 𝒍=(lt)tΓ\bm{l}=(l_{t})_{t\in\Gamma},4949 49 There is no loss in not allowing for transfers in lockout periods. and the agent’s actions are denoted by 𝐚=(at)tΓ\mathbf{a}=\left(a_{t}\right)_{t\in\Gamma}, where at=1a_{t}=1 if the agent works in period tΓt\in\Gamma and at=0a_{t}=0 if the agent shirks. The termination date of the contract is 0 if Γ=\Gamma=\emptyset and is otherwise maxΓ\max\Gamma, which we require to be finite.5050 50 One can show that this restriction does not hurt the principal. We say that a contract is connected if Γ={1,,T}\Gamma=\{1,\ldots,T\} for some TT; in this case we refer to TT as the length of the contract, TT is also the termination date, and we write 𝐂=(T,W0,𝒃,𝒍)\mathbf{C}=\left(T,W_{0},\bm{b},\bm{l}\right).

Given some program for the principal, we say that a simplified program entails no loss of optimality if the value of the two programs is the same.

Appendix A Proof of Theorem 2

Without loss by Proposition 1, we focus on penalty contracts throughout the proof.

A.1 Step 1: Low type always works

We show that it is without loss of optimality to focus on contracts for the low type, 𝐂L=(ΓL,W0L,𝒍L)\mathbf{C}^{L}=\left(\Gamma^{L},W_{0}^{L},\bm{l}^{L}\right), in which the low type works in all periods tΓLt\in\Gamma^{L}. Denote the set of penalty contracts by 𝒞\mathcal{C}, and recall that the principal’s program, with the restriction to penalty contracts, is:

max(𝐂H𝒞,𝐂L𝒞,𝐚H,𝐚L)μ0Π0H(𝐂H,𝐚H)+(1μ0)Π0L(𝐂L,𝐚L)\max_{\left(\mathbf{C}^{H}\in\mathcal{C},\mathbf{C}^{L}\in\mathcal{C},\mathbf{% a}^{H},\mathbf{a}^{L}\right)}\mu_{0}\Pi_{0}^{H}\left(\mathbf{C}^{H},\mathbf{a}% ^{H}\right)+\left(1-\mu_{0}\right)\Pi_{0}^{L}\left(\mathbf{C}^{L},\mathbf{a}^{% L}\right)

subject to, for all θ,θ{L,H}\theta,\theta^{\prime}\in\left\{L,H\right\},

𝐚θ\displaystyle\mathbf{a}^{\theta} 𝜶θ(𝐂θ),\displaystyle\in\bm{\alpha}^{\theta}(\mathbf{C}^{\theta}), (ICθa{}_{a}^{\theta})
U0θ(𝐂θ,𝐚θ)\displaystyle U_{0}^{\theta}(\mathbf{C}^{\theta},\mathbf{a}^{\theta}) 0,\displaystyle\geq 0, (IRθ)
U0θ(𝐂θ,𝐚θ)\displaystyle U_{0}^{\theta}(\mathbf{C}^{\theta},\mathbf{a}^{\theta}) U0θ(𝐂θ,𝜶θ(𝐂θ)).\displaystyle\geq U_{0}^{\theta}(\mathbf{C}^{\theta^{\prime}},\bm{\alpha}^{% \theta}(\mathbf{C}^{\theta^{\prime}})). (ICθθ{}^{\theta\theta^{\prime}})

Suppose there is a solution to this program, (𝐂H,𝐂L,𝐚H,𝐚L)(\mathbf{C}^{H},\mathbf{C}^{L},\mathbf{a}^{H},\mathbf{a}^{L}), with 𝐚L𝟏\mathbf{a}^{L}\neq\mathbf{1} and 𝐂L=(ΓL,W0L,𝒍L)\mathbf{C}^{L}=\left(\Gamma^{L},W_{0}^{L},\bm{l}^{L}\right). It suffices to show that there is another solution to the program, (𝐂H,𝐂^L,𝐚H,𝟏)(\mathbf{C}^{H},\widehat{\mathbf{C}}^{L},\mathbf{a}^{H},\mathbf{1}), where 𝐂^L=(Γ^L,W^0L,𝒍^L)\widehat{\mathbf{C}}^{L}=\left(\widehat{\Gamma}^{L},\widehat{W}_{0}^{L},% \widehat{\bm{l}}^{L}\right) is such that:

  1. (i)

    𝟏𝜶L(𝐂^L)\mathbf{1}\in\bm{\alpha}^{L}(\widehat{\mathbf{C}}^{L});

  2. (ii)

    U0L(𝐂L,𝐚L)=U0L(𝐂^L,𝟏)U^{L}_{0}(\mathbf{C}^{L},\mathbf{a}^{L})=U^{L}_{0}(\widehat{\mathbf{C}}^{L},% \mathbf{1});

  3. (iii)

    Π0L(𝐂L,𝐚L)=Π0L(𝐂^L,𝟏)\Pi^{L}_{0}(\mathbf{C}^{L},\mathbf{a}^{L})=\Pi^{L}_{0}(\widehat{\mathbf{C}}^{L% },\mathbf{1}); and

  4. (iv)

    U0H(𝐂L,𝜶H(𝐂L))U0H(𝐂^L,𝜶H(𝐂^L))U^{H}_{0}(\mathbf{C}^{L},\bm{\alpha}^{H}(\mathbf{C}^{L}))\geq U^{H}_{0}(% \widehat{\mathbf{C}}^{L},\bm{\alpha}^{H}(\widehat{\mathbf{C}}^{L})).

To this end, let t=min{s:as=0}t=\min\{s:a_{s}=0\} and denote the largest preceding period in ΓL\Gamma^{L} as

p(t)={maxΓL{t,t+1,} if sΓL s.t. s<t,0 otherwise.p(t)=\begin{cases}\max\Gamma^{L}\setminus\{t,t+1,\ldots\}&\text{ if }\ \exists s% \in\Gamma^{L}\text{ s.t. }s<t,\\ 0&\text{ otherwise.}\end{cases}

Construct 𝐂^L=(Γ^L,W^0L,𝒍^L)\widehat{\mathbf{C}}^{L}=\left(\widehat{\Gamma}^{L},\widehat{W}_{0}^{L},% \widehat{\bm{l}}^{L}\right) as follows:

Γ^L\displaystyle\widehat{\Gamma}^{L} =\displaystyle= ΓL\{t};\displaystyle\Gamma^{L}\backslash\left\{t\right\};
l^sL\displaystyle\widehat{l}_{s}^{L} =\displaystyle= {lsLif sp(t) and sΓ^L,lsL+δtp(t)ltLif s=p(t)>0;\displaystyle\begin{cases}l_{s}^{L}&\text{if $s\neq p(t)$ and $s\in\widehat{% \Gamma}^{L}$},\\ l_{s}^{L}+\delta^{t-p(t)}l_{t}^{L}&\text{if $s=p(t)>0$;}\end{cases}
W^0L\displaystyle\widehat{W}_{0}^{L} =\displaystyle= {W0Lif p(t)>0,W0L+δtltLif p(t)=0.\displaystyle\begin{cases}W^{L}_{0}&\text{if }p(t)>0,\\ W^{L}_{0}+\delta^{t}l^{L}_{t}&\text{if }p(t)=0.\end{cases}

Notice that under contract 𝐂L\mathbf{C}^{L}, the profile 𝐚L\mathbf{a}^{L} has type LL shirking in period tt and thus receiving ltLl^{L}_{t} with probability one conditional on not succeeding before this period; the new contract 𝐂^L\widehat{\mathbf{C}}^{L} just locks the agent out in period tt and shifts the payment ltLl^{L}_{t} up to the preceding non-lockout period, suitably discounted. It follows that the incentives for effort for type LL remain unchanged in any other period; moreover, since atL=0a^{L}_{t}=0, both the principal’s payoff from type LL under this contract and type LL’s payoff do not change. Finally, observe that for type HH, no matter which action he would take at tt in any optimal action plan under 𝐂L\mathbf{C}^{L} (whether it is work or shirk), his payoff from 𝐂^L\widehat{\mathbf{C}}^{L} must be weakly lower because the lockout in period tt is effectively as though he has been forced to shirk in period tt and receive ltLl^{L}_{t}.

Performing this procedure repeatedly for each period in which the original profile 𝐚L\mathbf{a}^{L} prescribes shirking yields a final contract 𝐂^L\widehat{\mathbf{C}}^{L} which satisfies all the desired properties.

A.2 Step 2: Simplifying the principal’s problem

By Step 1, we can focus on the following program [P]:

max(𝐂H𝒞,𝐂L𝒞,𝐚H)μ0Π0H(𝐂H,𝐚H)+(1μ0)Π0L(𝐂L,𝟏)\max_{(\mathbf{C}^{H}\in\mathcal{C},\mathbf{C}^{L}\in\mathcal{C},\mathbf{a}^{H% })}\mu_{0}\Pi_{0}^{H}\left(\mathbf{C}^{H},\mathbf{a}^{H}\right)+\left(1-\mu_{0% }\right)\Pi_{0}^{L}\left(\mathbf{C}^{L},\mathbf{1}\right) (P)

subject to

𝟏\displaystyle\mathbf{1} 𝜶L(𝐂L)\displaystyle\in\bm{\alpha}^{L}(\mathbf{C}^{L}) (ICLa{}_{a}^{L})
𝐚H\displaystyle\mathbf{a}^{H} 𝜶H(𝐂H)\displaystyle\in\bm{\alpha}^{H}(\mathbf{C}^{H}) (ICHa{}_{a}^{H})
U0L(𝐂L,𝟏)\displaystyle U_{0}^{L}\left(\mathbf{C}^{L},\mathbf{1}\right) 0\displaystyle\geq 0 (IRL)
U0H(𝐂H,𝐚H)\displaystyle U_{0}^{H}\left(\mathbf{C}^{H},\mathbf{a}^{H}\right) 0\displaystyle\geq 0 (IRH)
U0L(𝐂L,𝟏)\displaystyle U_{0}^{L}\left(\mathbf{C}^{L},\mathbf{1}\right) U0L(𝐂H,𝜶L(𝐂H))\displaystyle\geq U_{0}^{L}\left(\mathbf{C}^{H},\bm{\alpha}^{L}\left(\mathbf{C% }^{H}\right)\right) (ICLH)
U0H(𝐂H,𝐚H)\displaystyle U_{0}^{H}\left(\mathbf{C}^{H},\mathbf{a}^{H}\right) U0H(𝐂L,𝜶H(𝐂L)).\displaystyle\geq U_{0}^{H}\left(\mathbf{C}^{L},\bm{\alpha}^{H}\left(\mathbf{C% }^{L}\right)\right). (ICHL)

We first show that it is without loss of optimality to ignore constraints (IRH) and (ICLH).

Step 2a: Consider (IRH). Define a stochastic action plan 𝝈=(σt)tΓL\bm{\sigma}=\left(\sigma_{t}\right)_{t\in\Gamma^{L}} for type HH under contract 𝐂L\mathbf{C}^{L} as follows: σtΔ({0,1})\sigma_{t}\in\Delta\left(\left\{0,1\right\}\right) with σt(1)λLλH\sigma_{t}\left(1\right)\equiv\frac{\lambda^{L}}{\lambda^{H}} and σt(0)1λLλH\sigma_{t}\left(0\right)\equiv 1-\frac{\lambda^{L}}{\lambda^{H}} for all tΓLt\in\Gamma^{L}. In other words, under 𝝈\bm{\sigma}, the agent works in any period of ΓL\Gamma^{L} (so long as he not succeeded before) with probability λL/λH\lambda^{L}/\lambda^{H}. Note that these probabilities are independent across periods. By construction, it holds for all tΓLt\in\Gamma^{L} that 𝔼𝝈[at]=λL\mathbb{E}_{\bm{\sigma}}\left[a_{t}\right]=\lambda^{L}, where 𝔼𝝈\mathbb{E}_{\bm{\sigma}} is the ex-ante expectation with respect to the probability measure induced by 𝝈\bm{\sigma}.

Type HH’s expected payoff under contract 𝐂L\mathbf{C}^{L} given stochastic action plan 𝝈\bm{\sigma} is

U0H(𝐂L,𝝈)\displaystyle U_{0}^{H}\left(\mathbf{C}^{L},\bm{\sigma}\right) =β0tΓLδt𝔼𝝈{[sΓLst1(1asλH)][(1atλH)ltLatc]}+(1β0)tΓLδt𝔼𝝈[(ltLatc)]+W0L\displaystyle=\beta_{0}\sum\limits_{t\in\Gamma^{L}}\delta^{t}\mathbb{E}_{\bm{% \sigma}}\left\{\left[\prod\limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{L}}{s\leq t% -1}}\left(1-a_{s}\lambda^{H}\right)\right]\left[\left(1-a_{t}\lambda^{H}\right% )l_{t}^{L}-a_{t}c\right]\right\}+(1-\beta_{0})\sum\limits_{t\in\Gamma^{L}}% \delta^{t}\mathbb{E}_{\bm{\sigma}}\left[\left(l_{t}^{L}-a_{t}c\right)\right]+W% _{0}^{L}
=β0tΓLδt[sΓLst1𝔼𝝈(1asλH)]𝔼𝝈[(1atλH)ltLatc]+(1β0)tΓLδt𝔼𝝈[(ltLatc)]+W0L\displaystyle=\beta_{0}\sum\limits_{t\in\Gamma^{L}}\delta^{t}\left[\prod% \limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{L}}{s\leq t-1}}\mathbb{E}_{\bm{% \sigma}}\left(1-a_{s}\lambda^{H}\right)\right]\mathbb{E}_{\bm{\sigma}}\left[% \left(1-a_{t}\lambda^{H}\right)l_{t}^{L}-a_{t}c\right]+(1-\beta_{0})\sum% \limits_{t\in\Gamma^{L}}\delta^{t}\mathbb{E}_{\bm{\sigma}}\left[\left(l_{t}^{L% }-a_{t}c\right)\right]+W_{0}^{L}
=β0tΓLδt[sΓLst1(1λL)][(1λL)ltLλLc]+(1β0)tΓLδt(ltLλLc)+W0L\displaystyle=\beta_{0}\sum\limits_{t\in\Gamma^{L}}\delta^{t}\left[\prod% \limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{L}}{s\leq t-1}}\left(1-\lambda^{L}% \right)\right]\left[\left(1-\lambda^{L}\right)l_{t}^{L}-\lambda^{L}c\right]+(1% -\beta_{0})\sum\limits_{t\in\Gamma^{L}}\delta^{t}\left(l_{t}^{L}-\lambda^{L}c% \right)+W_{0}^{L}
β0tΓLδt[sΓLst1(1λL)][(1λL)ltLc]+(1β0)tΓLδt(ltLc)+W0L\displaystyle\geq\beta_{0}\sum\limits_{t\in\Gamma^{L}}\delta^{t}\left[\prod% \limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{L}}{s\leq t-1}}\left(1-\lambda^{L}% \right)\right]\left[\left(1-\lambda^{L}\right)l_{t}^{L}-c\right]+(1-\beta_{0})% \sum\limits_{t\in\Gamma^{L}}\delta^{t}\left(l_{t}^{L}-c\right)+W_{0}^{L}
=U0L(𝐂L,𝟏),\displaystyle=U_{0}^{L}\left(\mathbf{C}^{L},\mathbf{1}\right),

where the second equality follows from the independence of σt\sigma_{t} and σs\sigma_{s} for all t,sΓL,t,s\in\Gamma^{L}, the third equality follows from the fact that 𝔼𝝈[at]=λL\mathbb{E}_{\bm{\sigma}}\left[a_{t}\right]=\lambda^{L} for all tΓLt\in\Gamma^{L}, and the inequality follows from λL<1\lambda^{L}<1.5151 51 As a notational convention, the expression sΓL(1λL)\prod\limits_{{s\in\Gamma^{L}}}\left(1-\lambda^{L}\right) means (1λL)|ΓL|(1-\lambda^{L})^{\left|\Gamma^{L}\right|}, and analogously for similar expressions.

It follows immediately from the above string of (in)equalities that there exists a pure action plan 𝐚=(at)tΓL\mathbf{a=}\left(a_{t}\right)_{t\in\Gamma^{L}} such that

U0H(𝐂L,𝐚)U0H(𝐂L,𝝈)U0L(𝐂L,𝟏)0,U_{0}^{H}\left(\mathbf{C}^{L},\mathbf{a}\right)\geq U_{0}^{H}\left(\mathbf{C}^% {L},\bm{\sigma}\right)\geq U_{0}^{L}\left(\mathbf{C}^{L},\mathbf{1}\right)\geq 0,

where the last inequality follows from (IRL). Therefore, (ICHL) implies that

U0H(𝐂H,𝐚H)U0H(𝐂L,αH(𝐂L))U0H(𝐂L,𝐚)0,U_{0}^{H}\left(\mathbf{C}^{H},\mathbf{a}^{H}\right)\geq U_{0}^{H}\left(\mathbf% {C}^{L},\mathbf{\alpha}^{H}\left(\mathbf{C}^{L}\right)\right)\geq U_{0}^{H}% \left(\mathbf{C}^{L},\mathbf{a}\right)\geq 0,

which establishes (IRH).

Step 2b: Consider next (ICLH). By the same arguments as in Step 1, without loss of optimality we can restrict attention to contracts for the high type 𝐂H\mathbf{C}^{H} in which the high type works in all periods tΓHt\in\Gamma^{H}. If an optimal contract 𝐂H\mathbf{C}^{H} has ΓH=\Gamma^{H}=\emptyset, (ICLH) is trivially satisfied.5252 52 If an optimal contract 𝐂H\mathbf{C}^{H} excludes type HH, then without loss it can be taken to involve no transfers at all, which ensures that it would yield type LL a zero payoff, and hence (ICLH) follows from (IRL). Thus, assume an optimal contract 𝐂H\mathbf{C}^{H} has ΓH\Gamma^{H}\neq\emptyset. Let TH=maxΓHT^{H}=\max\Gamma^{H} and denote type HH’s expected payoff under 𝐂H\mathbf{C}^{H} by U¯0H.\overline{U}_{0}^{H}. We show that there exists a onetime-penalty contract that yields the principal the same expected payoff as 𝐂H\mathbf{C}^{H} and satisfies (ICLH). Consider a family of contracts 𝐂^H=(ΓH,W^0H,l^THH)\widehat{\mathbf{C}}^{H}=(\Gamma^{H},\widehat{W}_{0}^{H},\widehat{l}_{T^{H}}^{% H}), where l^THH\widehat{l}_{T^{H}}^{H} and W^0H\widehat{W}_{0}^{H} jointly ensure that type HH works in all periods tΓHt\in\Gamma^{H} and his expected payoff under ^𝐂H\widehat{}\mathbf{C}^{H} is equal to U¯0H\overline{U}_{0}^{H}:

[tΓH(1λH)β0+(1β0)]δTHl^THHc[β0tΓHδtsΓH,s<t(1λH)(1β0)tΓHδt]+W^0H=U¯0H.\left[\prod\limits_{t\in\Gamma^{H}}\left(1-\lambda^{H}\right)\beta_{0}+(1-% \beta_{0})\right]\delta^{T^{H}}\widehat{l}_{T^{H}}^{H}-c\left[\beta_{0}\sum_{t% \in\Gamma^{H}}\delta^{t}\prod\limits_{s\in\Gamma^{H},s<t}\left(1-\lambda^{H}% \right)-(1-\beta_{0})\sum_{t\in\Gamma^{H}}\delta^{t}\right]+\widehat{W}_{0}^{H% }=\overline{U}_{0}^{H}. (A.1)

It is immediate that any such contract 𝐂^H\widehat{\mathbf{C}}^{H} yields the principal the same expected payoff from type HH as the original contract 𝐂H\mathbf{C}^{H}, as it leaves both type HH’s action plan and type HH’s expected payoff under the new contract unchanged from the original contract. Furthermore, note that the penalty l^THH\widehat{l}_{T^{H}}^{H} can be chosen to be severe enough (i.e. sufficiently negative) to ensure that it is also optimal for type LL to work in all periods after accepting contract 𝐂^H\widehat{\mathbf{C}}^{H}; i.e., we can choose l^THH\widehat{l}_{T^{H}}^{H} so that for all θ{L,H}\theta\in\{L,H\}, 𝜶θ(𝐂^H)=𝟏\bm{\alpha}^{\theta}(\widehat{\mathbf{C}}^{H})=\mathbf{1}. All that remains is to show that a sufficiently severe l^THH\widehat{l}_{T^{H}}^{H} and its corresponding W^0H\widehat{W}_{0}^{H} (determined by (A.1)) also satisfy (ICLH) given that 𝜶L(𝐂^H)=𝟏\bm{\alpha}^{L}(\widehat{\mathbf{C}}^{H})=\mathbf{1}. To show this, note that type LL’s expected payoff from taking contract 𝐂^H\widehat{\mathbf{C}}^{H} and working in all periods tΓHt\in\Gamma^{H} is

U0L(𝐂H,𝟏)=[tΓH(1λL)β0+(1β0)]δTHl^THHc[β0tΓHδtsΓH,s<t(1λL)+(1β0)tΓHδt]+W^0H.U_{0}^{L}\left(\mathbf{C}^{H},\mathbf{1}\right)=\left[\prod\limits_{t\in\Gamma% ^{H}}\left(1-\lambda^{L}\right)\beta_{0}+(1-\beta_{0})\right]\delta^{T^{H}}% \widehat{l}_{T^{H}}^{H}-c\left[\beta_{0}\sum_{t\in\Gamma^{H}}\delta^{t}\prod% \limits_{s\in\Gamma^{H},s<t}\left(1-\lambda^{L}\right)+(1-\beta_{0})\sum_{t\in% \Gamma^{H}}\delta^{t}\right]+\widehat{W}_{0}^{H}. (A.2)

It follows from (A.2) and (A.1) that

U0L(𝐂H,𝟏)=[tΓH(1λL)tΓH(1λH)]β0δTHl^THH+cβ0tΓHδt[sΓH,s<t(1λH)sΓH,s<t(1λL)]+U¯0H.U_{0}^{L}\left(\mathbf{C}^{H},\mathbf{1}\right)=\left[\prod\limits_{t\in\Gamma% ^{H}}\left(1-\lambda^{L}\right)-\prod\limits_{t\in\Gamma^{H}}\left(1-\lambda^{% H}\right)\right]\beta_{0}\delta^{T^{H}}\widehat{l}_{T^{H}}^{H}+c\beta_{0}\sum_% {t\in\Gamma^{H}}\delta^{t}\left[\prod\limits_{s\in\Gamma^{H},s<t}\left(1-% \lambda^{H}\right)-\prod\limits_{s\in\Gamma^{H},s<t}\left(1-\lambda^{L}\right)% \right]+\overline{U}_{0}^{H}.

Since ΓH\Gamma^{H}\neq\emptyset and tΓH(1λL)tΓH(1λH)>0,\prod\limits_{t\in\Gamma^{H}}\left(1-\lambda^{L}\right)-\prod\limits_{t\in% \Gamma^{H}}\left(1-\lambda^{H}\right)>0, l^THH\widehat{l}_{T^{H}}^{H} can be chosen sufficiently negative such that U0L(𝐂H,𝟏)<0,U_{0}^{L}\left(\mathbf{C}^{H},\mathbf{1}\right)<0, establishing (ICLH).


Step 2c: By Step 2a and Step 2b, it is without loss of optimality to ignore (IRH) and (ICLH) in program [P]. Ignoring these two constraints yields the following program [P1]:

max(𝐂H𝒞,𝐂L𝒞,𝐚H)μ0Π0H(𝐂H,𝐚H)+(1μ0)Π0L(𝐂L,𝟏)\max_{(\mathbf{C}^{H}\in\mathcal{C},\mathbf{C}^{L}\in\mathcal{C},\mathbf{a}^{H% })}\mu_{0}\Pi_{0}^{H}\left(\mathbf{C}^{H},\mathbf{a}^{H}\right)+\left(1-\mu_{0% }\right)\Pi_{0}^{L}\left(\mathbf{C}^{L},\mathbf{1}\right) (P1)

subject to

𝟏\displaystyle\mathbf{1} 𝜶L(𝐂L)\displaystyle\in\bm{\alpha}^{L}(\mathbf{C}^{L}) (ICLa{}_{a}^{L})
𝐚H\displaystyle\mathbf{a}^{H} 𝜶H(𝐂H)\displaystyle\in\bm{\alpha}^{H}(\mathbf{C}^{H}) (ICHa{}_{a}^{H})
U0L(𝐂L,𝟏)\displaystyle U_{0}^{L}\left(\mathbf{C}^{L},\mathbf{1}\right) 0\displaystyle\geq 0 (IRL)
U0H(𝐂H,𝐚H)\displaystyle U_{0}^{H}\left(\mathbf{C}^{H},\mathbf{a}^{H}\right) U0H(𝐂L,𝜶H(𝐂L)).\displaystyle\geq U_{0}^{H}\left(\mathbf{C}^{L},\bm{\alpha}^{H}\left(\mathbf{C% }^{L}\right)\right). (ICHL)

It is clear that in any solution to program [P1], (IRL) must be binding: otherwise, the initial time-zero transfer from the principal to the agent in the contract 𝐂L\mathbf{C}^{L} can be reduced slightly to strictly improve the second term of the objective function while not violating any of the constraints. Similarly, (ICHL) must also bind because otherwise the time-zero transfer in the contract 𝐂H\mathbf{C}^{H} can be reduced to improve the first term of the objective function without violating any of the constraints.

Using these two binding constraints, substituting in the formulae from equations (1) and (2), and letting the principal select the optimal action plan the high type should use when taking the low type’s contract (𝐚HL𝜶H(𝐂L\mathbf{a}^{HL}\in\bm{\alpha}^{H}(\mathbf{C}^{L})), we can rewrite the objective function (P1) as the expected total surplus less type HH’s “information rent”, obtaining the following program that we call [P2]:

max(𝐂H𝒞,𝐂L𝒞,𝐚HL𝜶H(𝐂L),𝐚H){μ0{β0tΓHδt[sΓHst1(1asHλH)]atH(λHc)(1β0)tΓHδtatHc}+(1μ0){β0tΓLδt[sΓLst1(1λL)](λLc)(1β0)tΓLδtc}μ0{β0tΓLδtltL[sΓLst(1asHLλH)sΓLst(1λL)]β0ctΓLδtatHL[sΓLst1(1asHLλH)sΓLst1(1λL)]+ctΓLδt(1atHL)[1β0+β0sΓLst1(1λL)]}Information rent of type H}\displaystyle\max\limits_{\begin{subarray}{c}(\mathbf{C}^{H}\in\mathcal{C},\\ \mathbf{C}^{L}\in\mathcal{C},\\ \mathbf{a}^{HL}\in\bm{\alpha}^{H}(\mathbf{C}^{L}),\mathbf{a}^{H})\end{subarray% }}\left\{\begin{array}[]{l}\mu_{0}\left\{\beta_{0}\sum\limits_{t\in\Gamma^{H}}% \delta^{t}\left[\prod\limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{H}}{s\leq t-1}% }\left(1-a_{s}^{H}\lambda^{H}\right)\right]a_{t}^{H}\left(\lambda^{H}-c\right)% -(1-\beta_{0})\sum\limits_{t\in\Gamma^{H}}\delta^{t}a_{t}^{H}c\right\}\\ +\left(1-\mu_{0}\right)\left\{\beta_{0}\sum\limits_{t\in\Gamma^{L}}\delta^{t}% \left[\prod\limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{L}}{s\leq t-1}}\left(1-% \lambda^{L}\right)\right]\left(\lambda^{L}-c\right)-(1-\beta_{0})\sum\limits_{% t\in\Gamma^{L}}\delta^{t}c\right\}\\ -\mu_{0}\underbrace{\left\{\begin{array}[]{l}\beta_{0}\sum\limits_{t\in\Gamma^% {L}}\delta^{t}l_{t}^{L}\left[\prod\limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{L% }}{s\leq t}}\left(1-a_{s}^{HL}\lambda^{H}\right)-\prod\limits_{\genfrac{}{}{0.% 0pt}{}{s\in\Gamma^{L}}{s\leq t}}\left(1-\lambda^{L}\right)\right]\\ -\beta_{0}c\sum\limits_{t\in\Gamma^{L}}\delta^{t}a_{t}^{HL}\left[\prod\limits_% {\genfrac{}{}{0.0pt}{}{s\in\Gamma^{L}}{s\leq t-1}}\left(1-a_{s}^{HL}\lambda^{H% }\right)-\prod\limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{L}}{s\leq t-1}}\left(% 1-\lambda^{L}\right)\right]\\ +c\sum\limits_{t\in\Gamma^{L}}\delta^{t}(1-a_{t}^{HL})\left[1-\beta_{0}+\beta_% {0}\prod\limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{L}}{s\leq t-1}}\left(1-% \lambda^{L}\right)\right]\end{array}\right\}}_{\text{Information rent of type % $H$}}\end{array}\right\} (P2)

subject to

𝟏argmax(at)tΓL{β0tΓLδt[sΓLst1(1asλL)][(1atλL)ltLatc]+(1β0)tΓLδt(ltLatc)+W0L},\displaystyle\mathbf{1}\in\operatorname*{arg\,max}_{\left(a_{t}\right)_{t\in% \Gamma^{L}}}\left\{\begin{array}[]{l}\beta_{0}\sum\limits_{t\in\Gamma^{L}}% \delta^{t}\left[\prod\limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{L}}{s\leq t-1}% }\left(1-a_{s}\lambda^{L}\right)\right]\left[\left(1-a_{t}\lambda^{L}\right)l_% {t}^{L}-a_{t}c\right]+(1-\beta_{0})\sum\limits_{t\in\Gamma^{L}}\delta^{t}\left% (l_{t}^{L}-a_{t}c\right)+W_{0}^{L}\end{array}\right\}, (ICLa{}_{a}^{L})
𝐚Hargmax(at)tΓH{β0tΓHδt[sΓHst1(1asλH)][(1atλH)ltHatc]+(1β0)tΓHδt(ltHatc)+W0H}.\displaystyle\mathbf{a}^{H}\in\operatorname*{arg\,max}_{\left(a_{t}\right)_{t% \in\Gamma^{H}}}\left\{\begin{array}[]{l}\beta_{0}\sum\limits_{t\in\Gamma^{H}}% \delta^{t}\left[\prod\limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{H}}{s\leq t-1}% }\left(1-a_{s}\lambda^{H}\right)\right]\left[\left(1-a_{t}\lambda^{H}\right)l_% {t}^{H}-a_{t}c\right]+(1-\beta_{0})\sum\limits_{t\in\Gamma^{H}}\delta^{t}\left% (l_{t}^{H}-a_{t}c\right)+W_{0}^{H}\end{array}\right\}. (ICHa{}_{a}^{H})

Program [P2] is separable, i.e. it can be solved by maximizing (P2) with respect to (𝐂L,𝐚HL)(\mathbf{C}^{L},\mathbf{a}^{HL}) subject to (ICLa{}_{a}^{L}) and separately maximizing (P2) with respect to (𝐂H,𝐚H)(\mathbf{C}^{H},\mathbf{a}^{H}) subject to (ICHa{}_{a}^{H}).

We denote the information rent of type HH by R(𝐂L,𝐚HL)R\left(\mathbf{C}^{L},\mathbf{a}^{HL}\right). Note that given any action plan 𝐚\mathbf{a} that type HH uses when taking type LL’s contract,

R(𝐂L,𝐚)=U0H(𝐂L,𝐚)U0L(𝐂L,𝟏).R\left(\mathbf{C}^{L},\mathbf{a}\right)=U_{0}^{H}\left(\mathbf{C}^{L},\mathbf{% a}\right)-U_{0}^{L}\left(\mathbf{C}^{L},\mathbf{1}\right).

Hence, R(𝐂L,𝐚)=R(𝐂L,𝐚^)R\left(\mathbf{C}^{L},\mathbf{a}\right)=R\left(\mathbf{C}^{L},\widehat{\mathbf% {a}}\right) whenever both 𝐚,𝐚^𝜶H(𝐂L).\mathbf{a},\widehat{\mathbf{a}}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right).

It will be convenient at various places to consider the difference in information rents under contracts 𝐂^L\widehat{\mathbf{C}}^{L} and 𝐂L\mathbf{C}^{L} and corresponding action plans 𝐚^\widehat{\mathbf{a}} and 𝐚\mathbf{a}:

R(𝐂^L,𝐚^)R(𝐂L,𝐚)\displaystyle R\big{(}\widehat{\mathbf{C}}^{L},\widehat{\mathbf{a}}\big{)}-R% \left(\mathbf{C}^{L},\mathbf{a}\right) =U0H(𝐂^L,𝐚^)U0L(𝐂^L,𝟏)(U0H(𝐂L,𝐚)U0L(𝐂L,𝟏))\displaystyle=U_{0}^{H}\big{(}\widehat{\mathbf{C}}^{L},\widehat{\mathbf{a}}% \big{)}-U_{0}^{L}\big{(}\widehat{\mathbf{C}}^{L},\mathbf{1}\big{)}-\Big{(}U_{0% }^{H}\left(\mathbf{C}^{L},\mathbf{a}\right)-U_{0}^{L}\left(\mathbf{C}^{L},% \mathbf{1}\right)\Big{)}
=(U0H(𝐂^L,𝐚^)U0H(𝐂L,𝐚))(U0L(𝐂^L,𝟏)U0L(𝐂L,𝟏)).\displaystyle=\left(U_{0}^{H}\big{(}\widehat{\mathbf{C}}^{L},\widehat{\mathbf{% a}}\big{)}-U_{0}^{H}\left(\mathbf{C}^{L},\mathbf{a}\right)\right)-\left(U_{0}^% {L}\big{(}\widehat{\mathbf{C}}^{L},\mathbf{1}\big{)}-U_{0}^{L}\left(\mathbf{C}% ^{L},\mathbf{1}\right)\right). (A.11)

Moreover, when the action plan does not change across contracts (i.e. 𝐚=𝐚^\mathbf{a}=\widehat{\mathbf{a}} above), (A.11) specializes to

R(𝐂^L,𝐚)R(𝐂L,𝐚)=β0tΓLδt(l^tLltL)[sΓL,st(1asλH)sΓL,st(1λL)].R\big{(}\widehat{\mathbf{C}}^{L},\mathbf{a}\big{)}-R\left(\mathbf{C}^{L},% \mathbf{a}\right)=\beta_{0}\sum\limits_{t\in\Gamma^{L}}\delta^{t}\left(% \widehat{l}_{t}^{L}-l_{t}^{L}\right)\left[\prod\limits_{s\in\Gamma^{L},s\leq t% }\left(1-a_{s}\lambda^{H}\right)-\prod\limits_{s\in\Gamma^{L},s\leq t}\left(1-% \lambda^{L}\right)\right]. (A.12)

A.3 Step 3: Under-experimentation by the low type

Suppose per contra that 𝐂L=(ΓL,W0L,𝒍L)\mathbf{C}^{L}=(\Gamma^{L},W_{0}^{L},\bm{l}^{L}) is an optimal contract for the low type inducing him to work for |ΓL|>tL\left|\Gamma^{L}\right|>t^{L} periods. This implies ΓL\Gamma^{L}\neq\emptyset. We show that there exists 𝐂^L=(Γ^L,W^0L,𝒍^L)\mathbf{\widehat{C}}^{L}=(\widehat{\Gamma}^{L},\widehat{W}_{0}^{L},\widehat{% \bm{l}}^{L}) that induces the low type to work for |Γ^L|=|ΓL|1|\widehat{\Gamma}^{L}|=\left|\Gamma^{L}\right|-1 periods and strictly increases the principal’s payoff.

Let T=maxΓLT=\max\Gamma^{L} and T^=maxΓL\{T}\widehat{T}=\max\Gamma^{L}\backslash\left\{T\right\} be respectively the last and the second to the last non-lockout periods in contract 𝐂L.\mathbf{C}^{L}. Consider contract 𝐂^L\widehat{\mathbf{C}}^{L} defined as follows:

Γ^L=ΓL\{T},\widehat{\Gamma}^{L}=\Gamma^{L}\backslash\left\{T\right\},
l^tL={ltLif tΓ^L and t<T^lT^L+δTT^(1λL)lTLδTT^cif t=T^,\widehat{l}_{t}^{L}=\begin{cases}l_{t}^{L}&\text{if }t\in\widehat{\Gamma}^{L}% \text{ and }t<\widehat{T}\\ l_{\widehat{T}}^{L}+\delta^{T-\widehat{T}}(1-\lambda^{L})l_{T}^{L}-\delta^{T-% \widehat{T}}c&\text{if }t=\widehat{T},\end{cases}

and W^0\widehat{W}_{0} is such that (IRL) binds in contract 𝐂^L\widehat{\mathbf{C}}^{L}. Note that by construction, 𝐂^L\widehat{\mathbf{C}}^{L} gives the agent a continuation payoff in T^\widehat{T} which is the same the low-type agent would obtain if, given no success in T^\widehat{T}, the agent were to work in period TT. We proceed in two sub-steps.

Step 3a: Type LL works in all periods of 𝐂^L\widehat{\mathbf{C}}^{L}

We first show that type LL works in all periods tΓ^Lt\in\widehat{\Gamma}^{L} in 𝐂^L\widehat{\mathbf{C}}^{L}. Specifically, we show that type LL’s “incentive to work” in any period tΓ^Lt\in\widehat{\Gamma}^{L} under 𝐂^L\widehat{\mathbf{C}}^{L} is the same as his incentive to work in that period tt under the original contract 𝐂L\mathbf{C}^{L}; hence, the fact that type LL is willing to work in all periods tΓLt\in\Gamma^{L} under 𝐂L\mathbf{C}^{L} given that he works in all future periods (by Step 1) implies that is willing to work in all periods tΓ^Lt\in\widehat{\Gamma}^{L} under 𝐂^L\widehat{\mathbf{C}}^{L} given that he works in all future periods.

Type LL’s incentive to work in period T^\widehat{T} under the original contract 𝐂L\mathbf{C}^{L}, given that he works in period TT under such contract, is given by the difference between his continuation payoff from working and his continuation payoff from shirking in T^\widehat{T}:

{c+βT^L(1λL)[lT^L+δTT^(1λL)lTLδTT^c]+(1βT^L)(lT^L+δTT^lTLδTT^c)}\displaystyle\left\{-c+\beta_{\widehat{T}}^{L}(1-\lambda^{L})\left[l_{\widehat% {T}}^{L}+\delta^{T-\widehat{T}}(1-\lambda^{L})l_{T}^{L}-\delta^{T-\widehat{T}}% c\right]+(1-\beta_{\widehat{T}}^{L})\left(l_{\widehat{T}}^{L}+\delta^{T-% \widehat{T}}l_{T}^{L}-\delta^{T-\widehat{T}}c\right)\right\}
{lT^L+δTT^[c+βT^L(1λL)lTL+(1βT^L)lTL]}.\displaystyle-\left\{l_{\widehat{T}}^{L}+\delta^{T-\widehat{T}}\left[-c+\beta_% {\widehat{T}}^{L}(1-\lambda^{L})l_{T}^{L}+(1-\beta_{\widehat{T}}^{L})l_{T}^{L}% \right]\right\}. (A.13)

Note that βT^L=β¯|ΓL|1L\beta_{\widehat{T}}^{L}=\overline{\beta}^{L}_{|\Gamma^{L}|-1} if the agent works in all periods prior to T^\widehat{T}. Type LL works in period T^\widehat{T} only if expression (A.13) is non-negative. With some algebra, expression (A.13) can be simplified to

βT^LλL(lT1L+δ(1λL)lTL)+βT^LλLδcc.-\beta_{\widehat{T}}^{L}\lambda^{L}\left(l_{T-1}^{L}+\delta(1-\lambda^{L})l_{T% }^{L}\right)+\beta_{\widehat{T}}^{L}\lambda^{L}\delta c-c.

More generally, type LL’s incentive to work in any period tΓLt\in\Gamma^{L}, t<Tt<T, under contract 𝐂L\mathbf{C}^{L}, given work in all future periods in ΓL\Gamma^{L} and a belief βtL\beta_{t}^{L} in period tt, is

βtLλLτΓL,τtδτt[sΓL,tsτ(1λL)]lτL+βtLλLτΓL,τ>tδτt[sΓL,t<sτ(1λL)]cc.-\beta_{t}^{L}\lambda^{L}\sum_{\tau\in\Gamma^{L},\tau\geq t}\delta^{\tau-t}% \left[\prod\nolimits_{s\in\Gamma^{L},t\leq s\leq\tau}(1-\lambda^{L})\right]l_{% \tau}^{L}+\beta_{t}^{L}\lambda^{L}\sum_{\tau\in\Gamma^{L},\tau>t}\delta^{\tau-% t}\left[\prod\nolimits_{s\in\Gamma^{L},t<s\leq\tau}(1-\lambda^{L})\right]c-c. (A.14)

Note that βtL=β¯|{st:sΓL}|L\beta_{t}^{L}=\overline{\beta}_{\left|\left\{s\leq t:s\in\Gamma^{L}\right\}% \right|}^{L} if the low type works in all periods prior to t.t. Under contract 𝐂^L\widehat{\mathbf{C}}^{L}, type LL’s incentive to work in any period tΓ^Lt\in\widehat{\Gamma}^{L}, given work in all future periods and a belief βtL\beta_{t}^{L} in period tt, is

βtLλLτΓ^L,τtδτt[sΓ^L,tsτ(1λL)]l^τL+βtLλLτΓ^L,τ>tδτt[sΓ^L,t<sτ(1λL)]cc.-\beta_{t}^{L}\lambda^{L}\sum_{\tau\in\widehat{\Gamma}^{L},\tau\geq t}\delta^{% \tau-t}\left[\prod\nolimits_{s\in\widehat{\Gamma}^{L},t\leq s\leq\tau}(1-% \lambda^{L})\right]\widehat{l}_{\tau}^{L}+\beta_{t}^{L}\lambda^{L}\sum_{\tau% \in\widehat{\Gamma}^{L},\tau>t}\delta^{\tau-t}\left[\prod\nolimits_{s\in% \widehat{\Gamma}^{L},t<s\leq\tau}(1-\lambda^{L})\right]c-c.

By the definition of Γ^L\widehat{\Gamma}^{L} and l^tL\widehat{l}_{t}^{L} above, this expression can be rewritten as

βtLλLτΓL,tτ<T^δτt[sΓ^L,tsτ(1λL)]lτL\displaystyle\qquad-\beta_{t}^{L}\lambda^{L}\sum_{\tau\in\Gamma^{L},t\leq\tau<% \widehat{T}}\delta^{\tau-t}\left[\prod\nolimits_{s\in\widehat{\Gamma}^{L},t% \leq s\leq\tau}(1-\lambda^{L})\right]l_{\tau}^{L}
βtLλLδT^t[sΓL,tsT^(1λL)](lT^L+δTT^(1λL)lTLδTT^c)\displaystyle\qquad-\beta_{t}^{L}\lambda^{L}\delta^{\widehat{T}-t}\left[\prod% \nolimits_{s\in\Gamma^{L},t\leq s\leq\widehat{T}}(1-\lambda^{L})\right]\left(l% _{\widehat{T}}^{L}+\delta^{T-\widehat{T}}(1-\lambda^{L})l_{T}^{L}-\delta^{T-% \widehat{T}}c\right)
+βtLλLτΓL,t<τT^δτt[sΓL,t<sτ(1λL)]cc\displaystyle\qquad+\beta_{t}^{L}\lambda^{L}\sum_{\tau\in\Gamma^{L},t<\tau\leq% \widehat{T}}\delta^{\tau-t}\left[\prod\nolimits_{s\in\Gamma^{L},t<s\leq\tau}(1% -\lambda^{L})\right]c-c
=βtLλLτΓL,tτTδτt[sΓ^L,tsτ(1λL)]lτL+βtLλLτΓL,t<τTδτt[sΓL,t<sτ(1λL)]cc,\displaystyle=-\beta_{t}^{L}\lambda^{L}\sum_{\tau\in\Gamma^{L},t\leq\tau\leq T% }\delta^{\tau-t}\left[\prod\nolimits_{s\in\widehat{\Gamma}^{L},t\leq s\leq\tau% }(1-\lambda^{L})\right]l_{\tau}^{L}+\beta_{t}^{L}\lambda^{L}\sum_{\tau\in% \Gamma^{L},t<\tau\leq T}\delta^{\tau-t}\left[\prod\nolimits_{s\in\Gamma^{L},t<% s\leq\tau}(1-\lambda^{L})\right]c-c,

which is equal to expression (A.14) above. Hence, type LL is willing to work in all periods tΓ^Lt\in\widehat{\Gamma}^{L} under contract 𝐂^L\widehat{\mathbf{C}}^{L}.

Step 3b: Contract 𝐂^L\widehat{\mathbf{C}}^{L} weakly reduces type HH’s information rent

Since |ΓL|>tL\left|\Gamma^{L}\right|>t^{L} and contract 𝐂^L\widehat{\mathbf{C}}^{L} induces type LL to work for |ΓL|1\left|\Gamma^{L}\right|-1 periods, it is immediate that 𝐂^L\widehat{\mathbf{C}}^{L} strictly increases surplus from type LL relative to 𝐂L\mathbf{C}^{L}. To show that 𝐂^L\widehat{\mathbf{C}}^{L} increases the principal’s objective, it is thus sufficient to show that 𝐂^L\widehat{\mathbf{C}}^{L} weakly reduces type HH’s information rent relative to 𝐂L\mathbf{C}^{L}.

Let 𝐚^HL𝜶H(𝐂^L)\widehat{\mathbf{a}}^{HL}\in\bm{\alpha}^{H}\big{(}\widehat{\mathbf{C}}^{L}\big% {)} be an optimal action plan for type HH under contract 𝐂^L\widehat{\mathbf{C}}^{L}, 𝐚^HL=(a^tHL)tΓ^L\widehat{\mathbf{a}}^{HL}=(\widehat{a}_{t}^{HL})_{t\in\widehat{{\Gamma}}^{L}}. Define an action plan 𝐚^HL(1)\widehat{\mathbf{a}}^{HL(1)} for type HH under contract 𝐂L\mathbf{C}^{L} as follows: a^tHL(1)=a^tHL\widehat{a}_{t}^{HL(1)}=\widehat{a}_{t}^{HL} for tΓ^Lt\in\widehat{\Gamma}^{L} and a^THL(1)=1\widehat{a}_{T}^{HL(1)}=1. Note that since 𝐚HL𝜶H(𝐂L)\mathbf{a}^{HL}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right), U0H(𝐂L,𝐚HL)U0H(𝐂L,𝐚^HL(1))U_{0}^{H}\left(\mathbf{C}^{L},\mathbf{a}^{HL}\right)\geq U_{0}^{H}\left(% \mathbf{C}^{L},\widehat{\mathbf{a}}^{HL(1)}\right), and hence

R(𝐂L,𝐚HL)R(𝐂L,𝐚^HL(1)).R\left(\mathbf{C}^{L},\mathbf{a}^{HL}\right)\geq R(\mathbf{C}^{L},\widehat{% \mathbf{a}}^{HL(1)}). (A.15)

Now consider type HH’s information rent under 𝐂^L\widehat{\mathbf{C}}^{L} given optimal action plan 𝐚^HL𝜶H(𝐂^L)\widehat{\mathbf{a}}^{HL}\in\bm{\alpha}^{H}(\widehat{\mathbf{C}}^{L}):

R(𝐂^L,𝐚^HL)=\displaystyle R(\widehat{\mathbf{C}}^{L},\widehat{\mathbf{a}}^{HL})= β0tΓ^Lδtl^tL[sΓ^L,st(1a^sHLλH)sΓ^L,st(1λL)]\displaystyle\beta_{0}\sum\limits_{t\in\widehat{{\Gamma}}^{L}}\delta^{t}% \widehat{l}_{t}^{L}\left[\prod\limits_{s\in\widehat{{\Gamma}}^{L},s\leq t}% \left(1-\widehat{a}_{s}^{HL}\lambda^{H}\right)-\prod\limits_{s\in\widehat{{% \Gamma}}^{L},s\leq t}\left(1-\lambda^{L}\right)\right]
β0ctΓ^Lδta^tHL[sΓ^L,s<t(1a^sHLλH)sΓ^L,s<t(1λL)]\displaystyle-\beta_{0}c\sum\limits_{t\in\widehat{{\Gamma}}^{L}}\delta^{t}% \widehat{a}_{t}^{HL}\left[\prod\limits_{s\in\widehat{{\Gamma}}^{L},s<t}\left(1% -\widehat{a}_{s}^{HL}\lambda^{H}\right)-\prod\limits_{s\in\widehat{{\Gamma}}^{% L},s<t}\left(1-\lambda^{L}\right)\right]
+ctΓ^Lδt(1a^tHL)(1β0+β0sΓ^L,s<t(1λL)).\displaystyle+c\sum\limits_{t\in\widehat{{\Gamma}}^{L}}\delta^{t}(1-\widehat{a% }_{t}^{HL})\big{(}1-\beta_{0}+\beta_{0}\prod\limits_{s\in\widehat{{\Gamma}}^{L% },s<t}\left(1-\lambda^{L}\right)\big{)}.

Using the definition of 𝐂^L\widehat{\mathbf{C}}^{L}, this can be rewritten as

R(𝐂^L,𝐚^HL)=\displaystyle R(\widehat{\mathbf{C}}^{L},\widehat{\mathbf{a}}^{HL})= β0tΓ^L,t<T^δtltL[sΓ^L,st(1a^sHLλH)sΓ^L,st(1λL)]\displaystyle\beta_{0}\sum\limits_{t\in\widehat{{\Gamma}}^{L},t<\widehat{T}}% \delta^{t}l_{t}^{L}\left[\prod\limits_{s\in\widehat{{\Gamma}}^{L},s\leq t}% \left(1-\widehat{a}_{s}^{HL}\lambda^{H}\right)-\prod\limits_{s\in\widehat{{% \Gamma}}^{L},s\leq t}\left(1-\lambda^{L}\right)\right]
+β0δT^(lT^L+δTT^(1λL)lTLδTT^c)l^T^L[sΓ^L(1a^sHLλH)sΓ^L(1λL)]\displaystyle+\beta_{0}\delta^{\widehat{T}}\underbrace{\left(l_{\widehat{T}}^{% L}+\delta^{T-\widehat{T}}(1-\lambda^{L})l_{T}^{L}-\delta^{T-\widehat{T}}c% \right)}_{\widehat{l}_{\widehat{T}}^{L}}\left[\prod\limits_{s\in\widehat{{% \Gamma}}^{L}}\left(1-\widehat{a}_{s}^{HL}\lambda^{H}\right)-\prod\limits_{s\in% \widehat{{\Gamma}}^{L}}\left(1-\lambda^{L}\right)\right]
β0ctΓ^Lδta^tHL[sΓ^L,s<t(1a^sHLλH)sΓ^L,s<t(1λL)]\displaystyle-\beta_{0}c\prod\limits_{t\in\widehat{{\Gamma}}^{L}}\delta^{t}% \widehat{a}_{t}^{HL}\left[\prod\limits_{s\in\widehat{{\Gamma}}^{L},s<t}\left(1% -\widehat{a}_{s}^{HL}\lambda^{H}\right)-\prod\limits_{s\in\widehat{{\Gamma}}^{% L},s<t}\left(1-\lambda^{L}\right)\right]
+ctΓ^Lδt(1a^tHL)(1β0+β0sΓ^L,s<t(1λL)).\displaystyle+c\sum\limits_{t\in\widehat{{\Gamma}}^{L}}\delta^{t}(1-\widehat{a% }_{t}^{HL})\big{(}1-\beta_{0}+\beta_{0}\prod\limits_{s\in\widehat{{\Gamma}}^{L% },s<t}\left(1-\lambda^{L}\right)\big{)}.

Simple algebraic manipulations yield

R(𝐂^L,𝐚^HL)=R(𝐂L,𝐚^HL(1))+β0δTsΓ^L,sT^(1a^sHLλH)(λHλL)lTL.R\big{(}\widehat{\mathbf{C}}^{L},\widehat{\mathbf{a}}^{HL}\big{)}=R\big{(}% \mathbf{C}^{L},\widehat{\mathbf{a}}^{HL(1)}\big{)}+\beta_{0}\delta^{T}\prod% \limits_{s\in\widehat{{\Gamma}}^{L},s\leq\widehat{T}}\left(1-\widehat{a}_{s}^{% HL}\lambda^{H}\right)(\lambda^{H}-\lambda^{L})l_{T}^{L}. (A.16)

Note that since type LL is willing to work in period TT under contract 𝐂L\mathbf{C}^{L} (by Step 1), it holds that lTL<0l_{T}^{L}<0, and thus (A.16) yields

R(𝐂^L,𝐚^HL)R(𝐂L,𝐚^HL(1))<0.R\big{(}\widehat{\mathbf{C}}^{L},\widehat{\mathbf{a}}^{HL}\big{)}-R\big{(}% \mathbf{C}^{L},\widehat{\mathbf{a}}^{HL(1)}\big{)}<0.

Using (A.15), this implies R(𝐂^L,𝐚^HL)<R(𝐂L,𝐚HL)R\big{(}\widehat{\mathbf{C}}^{L},\widehat{\mathbf{a}}^{HL}\big{)}<R\left(% \mathbf{C}^{L},{\mathbf{a}}^{HL}\right).

A.4 Step 4: Efficient experimentation by the high type

The objective in [P2] involving the high type’s contract is social surplus from the high type. Furthermore, when 𝐚H=(1,,1)\mathbf{a}^{H}=(1,\ldots,1), with the sequence having arbitrary finite length, there is obviously a sequence of (sufficiently severe) penalties 𝒍H\bm{l}^{H} to ensure that (ICHa{}_{a}^{H}) is satisfied. It follows that we can take 𝐚H=(1,,1)\mathbf{a}^{H}=(1,\ldots,1) in an optimal contract 𝐂H\mathbf{C}^{H}, where the number of periods of work is tHt^{H}.

Appendix B Proof of Theorem 3

We remind the reader that Subsection 5.2 provides an outline and intuition for this proof. Without loss by Proposition 1, we focus on penalty contracts throughout the proof. In this appendix, we will introduce programs and constraints that have analogies with those used in Appendix A. Accordingly, we often use the same labels for equations as before, but the reader should bear in mind that all references in this appendix to such equations are to those defined in this appendix.

B.1 Step 1: The principal’s program

By Step 1 and Step 2 in the proof of Theorem 2, we work with the principal’s program [P1]. Recall that in this program, without loss, type LL works in all periods tΓLt\in\Gamma^{L} and constraints (ICLH) and (IRH) of program [P] are ignored. In this step, we relax the principal’s program by considering a weak version of (ICHL) in which type HH is assumed to exert effort in all periods tΓLt\in\Gamma^{L} if he chooses 𝐂L\mathbf{C}^{L}. The relaxed program, [RP1], is therefore:

max(𝐂H𝒞,𝐂L𝒞,𝐚H)μ0Π0H(𝐂H,𝐚H)+(1μ0)Π0L(𝐂L,𝟏)\max_{(\mathbf{C}^{H}\in\mathcal{C},\mathbf{C}^{L}\in\mathcal{C},\mathbf{a}^{H% })}\mu_{0}\Pi_{0}^{H}\left(\mathbf{C}^{H},\mathbf{a}^{H}\right)+\left(1-\mu_{0% }\right)\Pi_{0}^{L}\left(\mathbf{C}^{L},\mathbf{1}\right) (RP1)

subject to

𝟏\displaystyle\mathbf{1} 𝜶L(𝐂L)\displaystyle\in\bm{\alpha}^{L}(\mathbf{C}^{L}) (ICLa{}_{a}^{L})
𝐚H\displaystyle\mathbf{a}^{H} 𝜶H(𝐂H)\displaystyle\in\bm{\alpha}^{H}(\mathbf{C}^{H}) (ICHa{}_{a}^{H})
U0L(𝐂L,𝟏)\displaystyle U_{0}^{L}\left(\mathbf{C}^{L},\mathbf{1}\right) 0\displaystyle\geq 0 (IRL)
U0H(𝐂H,𝐚H)\displaystyle U_{0}^{H}\left(\mathbf{C}^{H},\mathbf{a}^{H}\right) U0H(𝐂L,𝟏).\displaystyle\geq U_{0}^{H}\left(\mathbf{C}^{L},\mathbf{1}\right). (Weak-ICHL)

By the same arguments as in Step 2 in the proof of Theorem 2, it is clear that in any solution to program [RP1], (IRL) and (Weak-ICHL) must be binding. Using these two binding constraints and substituting in the formulae from equations (1) and (2), we can rewrite the objective function (RP1) as the sum of expected total surplus less type HH’s “information rent”, obtaining the following explicit version of the relaxed program which we call [RP2]:

max𝐂H𝒞𝐂L𝒞𝐚H{μ0{β0tΓHδt[sΓHst1(1asHλH)]atH(λHc)(1β0)tΓHδtatHc}+(1μ0){β0tΓLδt[sΓLst1(1λL)](λLc)(1β0)tΓLδtc}μ0{β0tΓLδtltL[sΓLst(1λH)sΓLst(1λL)]β0tΓLδtc[sΓLst1(1λH)sΓLst1(1λL)]}Information rent of type H}\displaystyle\max\limits_{\begin{subarray}{c}\mathbf{C}^{H}\in\mathcal{C}\\ \mathbf{C}^{L}\in\mathcal{C}\\ \mathbf{a}^{H}\end{subarray}}\left\{\begin{array}[]{l}\mu_{0}\left\{\beta_{0}% \sum\limits_{t\in\Gamma^{H}}\delta^{t}\left[\prod\limits_{s\in\Gamma^{H}\atop s% \leq t-1}\left(1-a_{s}^{H}\lambda^{H}\right)\right]a_{t}^{H}\left(\lambda^{H}-% c\right)-(1-\beta_{0})\sum\limits_{t\in\Gamma^{H}}\delta^{t}a_{t}^{H}c\right\}% \\ +\left(1-\mu_{0}\right)\left\{\beta_{0}\sum\limits_{t\in\Gamma^{L}}\delta^{t}% \left[\prod\limits_{s\in\Gamma^{L}\atop s\leq t-1}\left(1-\lambda^{L}\right)% \right]\left(\lambda^{L}-c\right)-(1-\beta_{0})\sum\limits_{t\in\Gamma^{L}}% \delta^{t}c\right\}\\ -\mu_{0}\underbrace{\left\{\begin{array}[]{l}\beta_{0}\sum\limits_{t\in\Gamma^% {L}}\delta^{t}l_{t}^{L}\left[\prod\limits_{s\in\Gamma^{L}\atop s\leq t}\left(1% -\lambda^{H}\right)-\prod\limits_{s\in\Gamma^{L}\atop s\leq t}\left(1-\lambda^% {L}\right)\right]\\ -\beta_{0}\sum\limits_{t\in\Gamma^{L}}\delta^{t}c\left[\prod\limits_{s\in% \Gamma^{L}\atop s\leq t-1}\left(1-\lambda^{H}\right)-\prod\limits_{s\in\Gamma^% {L}\atop s\leq t-1}\left(1-\lambda^{L}\right)\right]\end{array}\right\}}_{% \text{Information rent of type $H$}}\end{array}\right\} (RP2)

subject to

𝟏argmax(at)tΓL{β0tΓLδt[sΓLst1(1asλL)][(1atλL)ltLatc]+(1β0)tΓLδt(ltLatc)+W0L},\displaystyle\mathbf{1}\in\operatorname*{arg\,max}_{\left(a_{t}\right)_{t\in% \Gamma^{L}}}\left\{\begin{array}[]{l}\beta_{0}\sum\limits_{t\in\Gamma^{L}}% \delta^{t}\left[\prod\limits_{s\in\Gamma^{L}\atop s\leq t-1}\left(1-a_{s}% \lambda^{L}\right)\right]\left[\left(1-a_{t}\lambda^{L}\right)l_{t}^{L}-a_{t}c% \right]+(1-\beta_{0})\sum\limits_{t\in\Gamma^{L}}\delta^{t}\left(l_{t}^{L}-a_{% t}c\right)+W_{0}^{L}\end{array}\right\}, (ICLa{}_{a}^{L})
𝐚Hargmax(at)tΓH{β0tΓHδt[sΓHst1(1asλH)][(1atλH)ltHatc]+(1β0)tΓHδt(ltHatc)+W0H}.\displaystyle\mathbf{a}^{H}\in\operatorname*{arg\,max}_{\left(a_{t}\right)_{t% \in\Gamma^{H}}}\left\{\begin{array}[]{l}\beta_{0}\sum\limits_{t\in\Gamma^{H}}% \delta^{t}\left[\prod\limits_{s\in\Gamma^{H}\atop s\leq t-1}\left(1-a_{s}% \lambda^{H}\right)\right]\left[\left(1-a_{t}\lambda^{H}\right)l_{t}^{H}-a_{t}c% \right]+(1-\beta_{0})\sum\limits_{t\in\Gamma^{H}}\delta^{t}\left(l_{t}^{H}-a_{% t}c\right)+W_{0}^{H}\end{array}\right\}. (ICHa{}_{a}^{H})

Program [RP2] is separable, i.e. it can be solved by maximizing (RP2) with respect to 𝐂L\mathbf{C}^{L} subject to (ICLa{}_{a}^{L}) and separately maximizing (RP2) with respect to (𝐂H,𝐚H)(\mathbf{C}^{H},\mathbf{a}^{H}) subject to (ICHa{}_{a}^{H}).

B.2 Step 2: Connected contracts for the low type

We claim that in program [RP2], it is without loss to consider solutions in which the low type’s contract is a connected penalty contract, i.e. solutions 𝐂L\mathbf{C}^{L} in which ΓL={1,,TL}\Gamma^{L}=\left\{1,...,T^{L}\right\} for some TL.T^{L}.

To prove this, observe that the optimal 𝐂L\mathbf{C}^{L} is a solution of

max𝐂L{(1μ0){β0tΓLδt[sΓLst1(1λL)](λLc)(1β0)tΓLδtc}μ0β0{tΓLδtltL[sΓLst(1λH)sΓLst(1λL)]tΓLδtc[sΓLst1(1λH)sΓLst1(1λL)]}}\max_{\mathbf{C}^{L}}\left\{\begin{array}[]{l}\left(1-\mu_{0}\right)\left\{% \beta_{0}\sum\limits_{t\in\Gamma^{L}}\delta^{t}\left[\prod\limits_{s\in\Gamma^% {L}\atop s\leq t-1}\left(1-\lambda^{L}\right)\right]\left(\lambda^{L}-c\right)% -(1-\beta_{0})\sum\limits_{t\in\Gamma^{L}}\delta^{t}c\right\}\\ -{\mu_{0}\beta_{0}\left\{\begin{array}[]{l}\sum\limits_{t\in\Gamma^{L}}\delta^% {t}l_{t}^{L}\left[\prod\limits_{s\in\Gamma^{L}\atop s\leq t}\left(1-\lambda^{H% }\right)-\prod\limits_{s\in\Gamma^{L}\atop s\leq t}\left(1-\lambda^{L}\right)% \right]-\sum\limits_{t\in\Gamma^{L}}\delta^{t}c\left[\prod\limits_{s\in\Gamma^% {L}\atop s\leq t-1}\left(1-\lambda^{H}\right)-\prod\limits_{s\in\Gamma^{L}% \atop s\leq t-1}\left(1-\lambda^{L}\right)\right]\end{array}\right\}}\end{% array}\right\} (B.8)

subject to (ICLa{}_{a}^{L}),

𝟏argmax(at)tΓL{β0tΓLδt[sΓL,st1(1asλL)][(1atλL)ltLatc]+(1β0)tΓLδt(ltLatc)+W0L}.\mathbf{1}\in\operatorname*{arg\,max}_{\left(a_{t}\right)_{t\in\Gamma^{L}}}% \left\{\begin{array}[]{l}\beta_{0}\sum\limits_{t\in\Gamma^{L}}\delta^{t}\left[% \prod\limits_{s\in\Gamma^{L},s\leq t-1}\left(1-a_{s}\lambda^{L}\right)\right]% \left[\left(1-a_{t}\lambda^{L}\right)l_{t}^{L}-a_{t}c\right]+(1-\beta_{0})\sum% \limits_{t\in\Gamma^{L}}\delta^{t}\left(l_{t}^{L}-a_{t}c\right)+W_{0}^{L}\end{% array}\right\}. (B.9)

To avoid trivialities, consider any optimal 𝐂L\mathbf{C}^{L} with ΓL\Gamma^{L}\neq\emptyset. First consider the possibility that 1ΓL1\notin\Gamma^{L}. In this case, construct a new penalty contract 𝐂^L\widehat{\mathbf{C}}^{L} that is “shifted up by one period”:

Γ^L\displaystyle\widehat{\Gamma}^{L} =\displaystyle= {s:s+1ΓL},\displaystyle\{s:s+1\in\Gamma^{L}\},
l^sL\displaystyle\widehat{l}_{s}^{L} =\displaystyle= ls+1L for all sΓ^L,\displaystyle l^{L}_{s+1}\text{ for all }s\in\widehat{\Gamma}^{L},
W^0L\displaystyle\widehat{W}_{0}^{L} =\displaystyle= W0L.\displaystyle W^{L}_{0}.

Clearly it remains optimal for the agent to work in every period in Γ^L\widehat{\Gamma}^{L}, and since the value of (B.8) must have been weakly positive under 𝐂L\mathbf{C}^{L}, it is now weakly higher since the modification has just multiplied it by δ1>1\delta^{-1}>1. This procedure can be repeated for all lockout periods at the beginning of the contract, so that without loss, we hereafter assume that 1ΓL1\in\Gamma^{L}. We are of course done if ΓL\Gamma^{L} is now connected, so also assume that ΓL\Gamma^{L} is not connected.

Let tt^{\circ} be the earliest lockout period in 𝐂L\mathbf{C}^{L}, i.e. t=min{t:tΓL and t1ΓL}t^{\circ}=\min\{t:t\notin\Gamma^{L}\text{ and }t-1\in\Gamma^{L}\}. (Such a t>1t^{\circ}>1 exists given the preceding discussion.) We will argue that one of two possible modifications preserves the agent’s incentive to work in all periods in the modified contract and weakly improves the principal’s payoff. This suffices because the procedure can then be applied iteratively to produce a connected contract.

Modification 1: Consider first a modified penalty contract 𝐂^L\widehat{\mathbf{C}}^{L} that removes the lockout period tt^{\circ} and shortens the contract by one period as follows:

Γ^L\displaystyle\widehat{\Gamma}^{L} =\displaystyle= {1,,t1}{s:st and s+1ΓL},\displaystyle\{1,\ldots,t^{\circ}-1\}\cup\{s:s\geq t^{\circ}\text{ and }s+1\in% \Gamma^{L}\},
l^sL\displaystyle\widehat{l}_{s}^{L} =\displaystyle= {lsLif s<t1,lsL+Δ1if s=t1,ls+1Lif st and sΓ^L,\displaystyle\begin{cases}l^{L}_{s}&\text{if }s<t^{\circ}-1,\\ l^{L}_{s}+\Delta_{1}&\text{if }s=t^{\circ}-1,\\ l^{L}_{s+1}&\text{if }s\geq t^{\circ}\text{ and }s\in\widehat{\Gamma}^{L},\end% {cases}
W^0L\displaystyle\widehat{W}_{0}^{L} =\displaystyle= W0L.\displaystyle W^{L}_{0}.

Note that in the above construction, Δ1\Delta_{1} is a free parameter. We will find conditions on Δ1\Delta_{1} such that type LL’s incentives for effort are unchanged and the principal is weakly better off.

For an arbitrary tt, define

S(t)\displaystyle S(t) =\displaystyle= (λLc)sΓL,st1(1λL),\displaystyle\left(\lambda^{L}-c\right)\prod\limits_{s\in\Gamma^{L},s\leq t-1}% \left(1-\lambda^{L}\right),
R(t)\displaystyle R(t) =\displaystyle= sΓL,st(1λH)sΓL,st(1λL).\displaystyle\prod\limits_{s\in\Gamma^{L},s\leq t}\left(1-\lambda^{H}\right)-% \prod\limits_{s\in\Gamma^{L},s\leq t}\left(1-\lambda^{L}\right).

The value of (B.8) under 𝐂L\mathbf{C}^{L} is

V(𝐂L)=(1μ0)[β0tΓLδtS(t)(1β0)tΓLδtc]μ0β0[tΓLδtltLR(t)tΓLδtcR(t1)].\displaystyle V(\mathbf{C}^{L})=\left(1-\mu_{0}\right)\left[\beta_{0}\sum% \limits_{t\in\Gamma^{L}}\delta^{t}S(t)-(1-\beta_{0})\sum\limits_{t\in\Gamma^{L% }}\delta^{t}c\right]-\mu_{0}\beta_{0}\left[\sum\limits_{t\in\Gamma^{L}}\delta^% {t}l_{t}^{L}R(t)-\sum\limits_{t\in\Gamma^{L}}\delta^{t}cR(t-1)\right]. (B.10)

The value of (B.8) after the modification to 𝐂^L\widehat{\mathbf{C}}^{L} is

V(𝐂^L)=\displaystyle V(\widehat{\mathbf{C}}^{L})= (1μ0)[β0(tΓLt<tδtS(t)+δ1tΓLt>tδtS(t))(1β0)(tΓLt<tδtc+δ1tΓLt>tδtc)]\displaystyle\left(1-\mu_{0}\right)\left[\begin{array}[]{l}\beta_{0}\left(% \displaystyle\sum\limits_{t\in\Gamma^{L}\atop t<t^{\circ}}\delta^{t}S(t)+% \delta^{-1}\displaystyle\sum\limits_{t\in\Gamma^{L}\atop t>t^{\circ}}\delta^{t% }S(t)\right)-(1-\beta_{0})\left(\displaystyle\sum\limits_{t\in\Gamma^{L}\atop t% <t^{\circ}}\delta^{t}c+\delta^{-1}\displaystyle\sum\limits_{t\in\Gamma^{L}% \atop t>t^{\circ}}\delta^{t}c\right)\end{array}\right]
μ0β0[tΓLt<t1δtltLR(t)+δ1tΓLt>tδtltLR(t)+δt1(lt1L+Δ1)R(t1)tΓLt<tδtcR(t1)δ1tΓLt>tδtcR(t1)].\displaystyle-\mu_{0}\beta_{0}\left[\begin{array}[]{l}\displaystyle\sum\limits% _{t\in\Gamma^{L}\atop t<t^{\circ}-1}\delta^{t}l_{t}^{L}R(t)+\delta^{-1}% \displaystyle\sum\limits_{t\in\Gamma^{L}\atop t>t^{\circ}}\delta^{t}l_{t}^{L}R% (t)+\delta^{t^{\circ}-1}\left(l_{t^{\circ}-1}^{L}+\Delta_{1}\right)R(t^{\circ}% -1)\\[25.0pt] -\displaystyle\sum\limits_{t\in\Gamma^{L}\atop t<t^{\circ}}\delta^{t}cR(t-1)-% \delta^{-1}\displaystyle\sum\limits_{t\in\Gamma^{L}\atop t>t^{\circ}}\delta^{t% }cR(t-1)\end{array}\right].

Therefore, the modification benefits the principal if and only if

0V(𝐂^L)V(𝐂L)\displaystyle 0\leq V(\widehat{\mathbf{C}}^{L})-V(\mathbf{C}^{L}) =\displaystyle= (1μ0)[β0(δ11)tΓLt>tδtS(t)(1β0)(δ11)tΓLt>tδtc]\displaystyle\left(1-\mu_{0}\right)\left[\begin{array}[]{l}\beta_{0}\left(% \delta^{-1}-1\right)\displaystyle\sum\limits_{t\in\Gamma^{L}\atop t>t^{\circ}}% \delta^{t}S(t)-(1-\beta_{0})\left(\delta^{-1}-1\right)\displaystyle\sum\limits% _{t\in\Gamma^{L}\atop t>t^{\circ}}\delta^{t}c\end{array}\right]
μ0β0[(δ11)tΓLt>tδtltLR(t)+δt1Δ1R(t1)(δ11)tΓLt>tδtcR(t1)].\displaystyle-\mu_{0}\beta_{0}\left[\begin{array}[]{l}\left(\delta^{-1}-1% \right)\displaystyle\sum\limits_{t\in\Gamma^{L}\atop t>t^{\circ}}\delta^{t}l_{% t}^{L}R(t)+\delta^{t^{\circ}-1}\Delta_{1}R(t^{\circ}-1)-\left(\delta^{-1}-1% \right)\displaystyle\sum\limits_{t\in\Gamma^{L}\atop t>t^{\circ}}\delta^{t}cR(% t-1)\end{array}\right].

The above inequality is satisfied for any Δ1\Delta_{1} if δ=1\delta=1, and if δ<1\delta<1, then after rearranging terms, the above inequality is equivalent to

(1μ0)[β0tΓLt>tδtS(t)(1β0)tΓLt>tδtc]\displaystyle\left(1-\mu_{0}\right)\left[\beta_{0}\sum\limits_{t\in\Gamma^{L}% \atop t>t^{\circ}}\delta^{t}S(t)-(1-\beta_{0})\sum\limits_{t\in\Gamma^{L}\atop t% >t^{\circ}}\delta^{t}c\right]
μ0β0[tΓLt>tδtltLR(t)δt1Δ1R(t1)1δ1tΓLt>tδtcR(t1)].\displaystyle\qquad\geq\mu_{0}\beta_{0}\left[\sum\limits_{t\in\Gamma^{L}\atop t% >t^{\circ}}\delta^{t}l_{t}^{L}R(t)-\frac{\delta^{t^{\circ}-1}\Delta_{1}R(t^{% \circ}-1)}{1-\delta^{-1}}-\sum\limits_{t\in\Gamma^{L}\atop t>t^{\circ}}\delta^% {t}cR(t-1)\right]. (B.15)

Now turn to the incentives for effort for the agent of type LL. Clearly, since 𝐂L\mathbf{C}^{L} induces the agent to work in all periods, it remains optimal for the agent to work under 𝐂^L\widehat{\mathbf{C}}^{L} in all periods beginning with tt^{\circ}. Consider the incentive constraint for effort in period t1t^{\circ}-1 under 𝐂^L\widehat{\mathbf{C}}^{L}. Using (B.9), this is given by:

β¯t1LλL{lt1L+Δ1+δ1tΓLt>tδt(t1)[sΓLt1<st1(1λL)][(1λL)ltLc]}c.-\overline{\beta}_{t^{\circ}-1}^{L}\lambda^{L}\left\{l_{t^{\circ}-1}^{L}+% \Delta_{1}+\delta^{-1}\sum\limits_{t\in\Gamma^{L}\atop t>t^{\circ}}\delta^{t-% \left(t^{\circ}-1\right)}\left[\prod\limits_{s\in\Gamma^{L}\atop t^{\circ}-1<s% \leq t-1}\left(1-\lambda^{L}\right)\right]\left[\left(1-\lambda^{L}\right)l_{t% }^{L}-c\right]\right\}\geq c. (B.16)

Analogously, the incentive constraint in period t1t^{\circ}-1 under the original contract 𝐂L\mathbf{C}^{L} is:

β¯t1LλL{lt1L+tΓLt>t1δt(t1)[sΓLt1<st1(1λL)][(1λL)ltLc]}c.-\overline{\beta}_{t^{\circ}-1}^{L}\lambda^{L}\left\{l_{t^{\circ}-1}^{L}+\sum% \limits_{t\in\Gamma^{L}\atop t>t^{\circ}-1}\delta^{t-(t^{\circ}-1)}\left[\prod% \limits_{s\in\Gamma^{L}\atop t^{\circ}-1<s\leq t-1}\left(1-\lambda^{L}\right)% \right]\left[\left(1-\lambda^{L}\right)l_{t}^{L}-c\right]\right\}\geq c. (B.17)

If we choose Δ1\Delta_{1} such that the left-hand side of (B.16) is equal to the left-hand side of (B.17), then since it is optimal to work under the original contract in period t1t^{\circ}-1, it will also be optimal to work under the new contract in period t1t^{\circ}-1. Accordingly, we choose Δ1\Delta_{1} such that:

Δ1\displaystyle\Delta_{1} =\displaystyle= tΓL,t>t1δt(t1)[sΓL,t1<st1(1λL)][(1λL)ltLc]\displaystyle\sum\limits_{t\in\Gamma^{L},t>t^{\circ}-1}\delta^{t-\left(t^{% \circ}-1\right)}\left[\prod\limits_{s\in\Gamma^{L},t^{\circ}-1<s\leq t-1}\left% (1-\lambda^{L}\right)\right]\left[\left(1-\lambda^{L}\right)l_{t}^{L}-c\right] (B.18)
δ1tΓL,t>tδt(t1)[sΓL,t1<st1(1λL)][(1λL)ltLc]\displaystyle-\delta^{-1}\sum\limits_{t\in\Gamma^{L},t>t^{\circ}}\delta^{t-% \left(t^{\circ}-1\right)}\left[\prod\limits_{s\in\Gamma^{L},t^{\circ}-1<s\leq t% -1}\left(1-\lambda^{L}\right)\right]\left[\left(1-\lambda^{L}\right)l_{t}^{L}-% c\right]
=\displaystyle= (1δ1)tΓL,t>t1δt(t1)[sΓL,t1<st1(1λL)][(1λL)ltLc],\displaystyle(1-\delta^{-1})\sum\limits_{t\in\Gamma^{L},t>t^{\circ}-1}\delta^{% t-\left(t^{\circ}-1\right)}\left[\prod\limits_{s\in\Gamma^{L},t^{\circ}-1<s% \leq t-1}\left(1-\lambda^{L}\right)\right]\left[\left(1-\lambda^{L}\right)l_{t% }^{L}-c\right],

where the second equality is because {t:tΓL,t>t1}={t:tΓL,t>t}\{t:t\in\Gamma^{L},t>t^{\circ}-1\}=\{t:t\in\Gamma^{L},t>t^{\circ}\}, since tΓLt^{\circ}\notin\Gamma^{L}. Note that (B.18) implies Δ1=0\Delta_{1}=0 if δ=1\delta=1.

Now consider the incentive constraint for effort in any period τ<t1\tau<t^{\circ}-1. We will show that because Δ1\Delta_{1} is such that the left-hand side of (B.16) is equal to the left-hand side of (B.17), the fact that it was optimal to work in period τ\tau under contract 𝐂L\mathbf{C}^{L} implies that it is optimal to work in period τ\tau under contract 𝐂^L\widehat{\mathbf{C}}^{L}. Formally, the incentive constraint for effort in period τ\tau under 𝐂L\mathbf{C}^{L} is

β¯τLλL{lτL+tΓL,t>τδtτ[sΓL,τ<st1(1λL)][(1λL)ltLc]}c,-\overline{\beta}_{\tau}^{L}\lambda^{L}\left\{l_{\tau}^{L}+\sum\limits_{t\in% \Gamma^{L},t>\tau}\delta^{t-\tau}\left[\prod\limits_{s\in\Gamma^{L},\tau<s\leq t% -1}\left(1-\lambda^{L}\right)\right]\left[\left(1-\lambda^{L}\right)l_{t}^{L}-% c\right]\right\}\geq c, (B.19)

which is satisfied since 𝐂L\mathbf{C}^{L} induces the agent to work in all periods. Analogously, the incentive constraint for effort in period τ\tau under 𝐂^L\widehat{\mathbf{C}}^{L} can be written as

β¯τLλL{l^τL+tΓ^L,t>τδtτ[sΓ^L,τ<st1(1λL)][(1λL)l^tLc]}c.\displaystyle-\overline{\beta}_{\tau}^{L}\lambda^{L}\left\{\widehat{l}_{\tau}^% {L}+\sum\limits_{t\in\widehat{\Gamma}^{L},t>\tau}\delta^{t-\tau}\left[\prod% \limits_{s\in\widehat{\Gamma}^{L},\tau<s\leq t-1}\left(1-\lambda^{L}\right)% \right]\left[\left(1-\lambda^{L}\right)\widehat{l}_{t}^{L}-c\right]\right\}% \geq c.

Algebraic simplification using the definition of ^𝐂L\widehat{}\mathbf{C}^{L} and equation (B.18) shows that this constraint is identical to (B.19), and hence is satisfied.

Thus, if δ=1\delta=1, this modification with Δ1=0\Delta_{1}=0 weakly benefits the principal while preserving the agent’s incentives, and we are done. So hereafter assume δ<1\delta<1, which requires us to also consider another modification.


Modification 2: Now we consider a modified contract ~𝐂L\widetilde{}\mathbf{C}^{L} that eliminates all periods after tt^{\circ}, defined as follows:

W~0L=W0L,Γ~L={1,,t1},l~sL={lsLif s<t1,lsL+Δ2if s=t1.\displaystyle\widetilde{W}_{0}^{L}=W^{L}_{0},\ \ \ \widetilde{\Gamma}^{L}=\{1,% \ldots,t^{\circ}-1\},\ \ \ \widetilde{l}_{s}^{L}=\begin{cases}l^{L}_{s}&\text{% if }s<t^{\circ}-1,\\ l^{L}_{s}+\Delta_{2}&\text{if }s=t^{\circ}-1.\end{cases}

Again, Δ2\Delta_{2} is a free parameter above. We now find conditions on Δ2\Delta_{2} such that type LL’s incentives are unchanged and the principal is weakly better off.

The value of (B.8) under the modification ~𝐂L\widetilde{}\mathbf{C}^{L} is

V(~𝐂L)\displaystyle V(\widetilde{}\mathbf{C}^{L}) =\displaystyle= (1μ0)[β0tΓL,t<tδtS(t)(1β0)tΓL,t<tδtc]\displaystyle\left(1-\mu_{0}\right)\left[\beta_{0}\sum\limits_{t\in\Gamma^{L},% t<t^{\circ}}\delta^{t}S(t)-(1-\beta_{0})\sum\limits_{t\in\Gamma^{L},t<t^{\circ% }}\delta^{t}c\right]
μ0β0[tΓL,t<t1δtltLR(t)+δt1(lt1L+Δ2)R(t1)tΓL,t<tδtcR(t1)].\displaystyle-\mu_{0}\beta_{0}\left[\begin{array}[]{l}\displaystyle\sum\limits% _{t\in\Gamma^{L},t<t^{\circ}-1}\delta^{t}l_{t}^{L}R(t)+\delta^{t^{\circ}-1}% \left(l_{t^{\circ}-1}^{L}+\Delta_{2}\right)R(t^{\circ}-1)-\displaystyle\sum% \limits_{t\in\Gamma^{L},t<t^{\circ}}\delta^{t}cR(t-1)\end{array}\right].

Therefore, recalling (B.10), this modification benefits the principal if and only if

0V(~𝐂L)V(𝐂L)\displaystyle 0\leq V(\widetilde{}\mathbf{C}^{L})-V(\mathbf{C}^{L}) =\displaystyle= (1μ0)[β0tΓL,t>tδtS(t)(1β0)tΓL,t>tδtc]\displaystyle-\left(1-\mu_{0}\right)\left[\beta_{0}\sum\limits_{t\in\Gamma^{L}% ,t>t^{\circ}}\delta^{t}S(t)-(1-\beta_{0})\sum\limits_{t\in\Gamma^{L},t>t^{% \circ}}\delta^{t}c\right]
μ0β0[tΓL,t>tδtltLR(t)+δt1Δ2R(t1)+tΓL,t>tδtcR(t1)],\displaystyle-\mu_{0}\beta_{0}\left[\begin{array}[]{l}-\displaystyle\sum% \limits_{t\in\Gamma^{L},t>t^{\circ}}\delta^{t}l_{t}^{L}R(t)+\delta^{t^{\circ}-% 1}\Delta_{2}R(t^{\circ}-1)+\displaystyle\sum\limits_{t\in\Gamma^{L},t>t^{\circ% }}\delta^{t}cR(t-1)\end{array}\right],

or equivalently after rearranging terms, if and only if

(1μ0)[β0tΓLt>tδtS(t)(1β0)tΓLt>tδtc]\displaystyle\left(1-\mu_{0}\right)\left[\beta_{0}\sum\limits_{t\in\Gamma^{L}% \atop t>t^{\circ}}\delta^{t}S(t)-(1-\beta_{0})\sum\limits_{t\in\Gamma^{L}\atop t% >t^{\circ}}\delta^{t}c\right]
μ0β0[tΓLt>tδtltLR(t)δt1Δ2R(t1)tΓLt>tδtcR(t1)].\displaystyle\qquad\leq\mu_{0}\beta_{0}\left[\sum\limits_{t\in\Gamma^{L}\atop t% >t^{\circ}}\delta^{t}l_{t}^{L}R(t)-\delta^{t^{\circ}-1}\Delta_{2}R(t^{\circ}-1% )-\sum\limits_{t\in\Gamma^{L}\atop t>t^{\circ}}\delta^{t}cR(t-1)\right]. (B.22)

As with the previous modification, the only incentive constraint for effort that needs to be verified in ~𝐂L\widetilde{}\mathbf{C}^{L} is that of period t1t^{\circ}-1, which since it is the last period of the contract is simply:

β¯t1LλL(lt1L+Δ2)c.-\overline{\beta}_{t^{\circ}-1}^{L}\lambda^{L}\left(l_{t^{\circ}-1}^{L}+\Delta% _{2}\right)\geq c. (B.23)

We choose Δ2\Delta_{2} so that the left-hand side of (B.23) is equal to the left-hand side of (B.17):

Δ2\displaystyle\Delta_{2} =\displaystyle= tΓL,t>t1δt(t1)[sΓL,t1<st1(1λL)][(1λL)ltLc]=Δ11δ1,\displaystyle\sum\limits_{t\in\Gamma^{L},t>t^{\circ}-1}\delta^{t-\left(t^{% \circ}-1\right)}\left[\prod\limits_{s\in\Gamma^{L},t^{\circ}-1<s\leq t-1}\left% (1-\lambda^{L}\right)\right]\left[\left(1-\lambda^{L}\right)l_{t}^{L}-c\right]% =\frac{\Delta_{1}}{1-\delta^{-1}}, (B.24)

where the second equality follows from (B.18). But now, observe that (B.24) implies that either (B.15) or (B.22) is guaranteed to hold, and hence either the modification to 𝐂^L\widehat{\mathbf{C}}^{L} or to ~𝐂L\widetilde{}\mathbf{C}^{L} weakly benefits the principal while preserving the agent’s effort incentives.

Remark 3.

Given δ<1\delta<1, the choice of Δ2\Delta_{2} in (B.24) implies that if inequality (B.15) holds with equality then so does inequality (B.22), and vice-versa. In other words, if neither of the modifications strictly benefits the principal (while preserving the agent’s effort incentives), then it must be that both modifications leave the principal’s payoff unchanged (while preserving the agent’s effort incentives).

B.3 Step 3: Defining the critical contract for the low type

Take any connected penalty contract 𝐂L=(TL,W0L,𝒍L)\mathbf{C}^{L}=(T^{L},W^{L}_{0},\bm{l}^{L}) that induces effort from the low type in each period t{1,,TL}t\in\{1,\ldots,T^{L}\}. We claim that the low type’s incentive constraint for effort binds at all periods if and only if 𝒍L=𝒍¯L(TL)\bm{l}^{L}=\overline{\bm{l}}^{L}(T^{L}), where 𝒍¯L(TL)\overline{\bm{l}}^{L}(T^{L}) is defined as follows:

l¯tL={(1δ)cβ¯tLλLif t<TL,cβ¯TLLλLif t=TL.\overline{l}_{t}^{L}=\left\{\begin{tabular}[]{ll}$-\left(1-\delta\right)\frac{% c}{\overline{\beta}_{t}^{L}\lambda^{L}}$&if $t<T^{L}$,\\ $-\frac{c}{\overline{\beta}_{T^{L}}^{L}\lambda^{L}}$&if $t=T^{L}$.\end{tabular% }\right. (B.25)

The proof of this claim is via three sub-steps; for the remainder of this step, since TLT^{L} is given and held fixed, we ease notation by just writing 𝒍¯L\overline{\bm{l}}^{L} instead of 𝒍¯L(TL)\overline{\bm{l}}^{L}(T^{L}).

Step 3a: First, we argue that with the above penalty sequence, the low type is indifferent between working and shirking in each period t{1,,TL}t\in\{1,\ldots,T^{L}\} given that he has worked in all prior periods and will do in all subsequent periods no matter his action at period tt. In other words, we need to show that for all t{1,,TL}t\in\{1,\ldots,T^{L}\}:

β¯tLλL{l¯tL+s=t+1TLδst(1λL)s(t+1)[(1λL)l¯sLc]}=c.-\overline{\beta}_{t}^{L}\lambda^{L}\left\{\overline{l}_{t}^{L}+\sum\limits_{s% =t+1}^{T^{L}}\delta^{s-t}\left(1-\lambda^{L}\right)^{s-\left(t+1\right)}\left[% \left(1-\lambda^{L}\right)\overline{l}_{s}^{L}-c\right]\right\}=c. (B.26)

We prove that (B.26) is indeed satisfied for all tt by induction. First, it is immediate from (B.25) that (B.26) holds for t=TLt=T^{L}. Next, for any t<TLt<T^{L}, assume (B.26) holds for t+1t+1. This is equivalent to

s=t+2TLδs(t+1)(1λL)s(t+2)[(1λL)l¯sLc]=cβ¯t+1LλLl¯t+1L.\sum\limits_{s=t+2}^{T^{L}}\delta^{s-\left(t+1\right)}\left(1-\lambda^{L}% \right)^{s-\left(t+2\right)}\left[\left(1-\lambda^{L}\right)\overline{l}_{s}^{% L}-c\right]=-\frac{c}{\overline{\beta}_{t+1}^{L}\lambda^{L}}-\overline{l}_{t+1% }^{L}. (B.27)

To show that (B.26) holds for tt, it suffices to show that

β¯tLλL{l¯tL+δ[(1λL)l¯t+1Lc]+δ(1λL)s=t+2TLδs(t+1)(1λL)s(t+2)[(1λL)l¯sLc]}=c.-\overline{\beta}_{t}^{L}\lambda^{L}\left\{\overline{l}_{t}^{L}+\delta\left[% \left(1-\lambda^{L}\right)\overline{l}_{t+1}^{L}-c\right]+\delta\left(1-% \lambda^{L}\right)\sum\limits_{s=t+2}^{T^{L}}\delta^{s-\left(t+1\right)}\left(% 1-\lambda^{L}\right)^{s-\left(t+2\right)}\left[\left(1-\lambda^{L}\right)% \overline{l}_{s}^{L}-c\right]\right\}=c.

Using (B.27), the above equality is equivalent to

β¯tLλL{l¯tL+δ[(1λL)l¯t+1Lc]+δ(1λL)[cβ¯t+1LλLl¯t+1L]}=c,-\overline{\beta}_{t}^{L}\lambda^{L}\left\{\overline{l}_{t}^{L}+\delta\left[% \left(1-\lambda^{L}\right)\overline{l}_{t+1}^{L}-c\right]+\delta\left(1-% \lambda^{L}\right)\left[-\frac{c}{\overline{\beta}_{t+1}^{L}\lambda^{L}}-% \overline{l}_{t+1}^{L}\right]\right\}=c,

which simplifies to

l¯tL=cβ¯tLλL+δc+δ(1λL)cβ¯t+1LλL.\overline{l}_{t}^{L}=-\frac{c}{\overline{\beta}_{t}^{L}\lambda^{L}}+\delta c+% \delta\left(1-\lambda^{L}\right)\frac{c}{\overline{\beta}_{t+1}^{L}\lambda^{L}}. (B.28)

Since β¯t+1L=β¯tL(1λL)1β¯tLλL\overline{\beta}_{t+1}^{L}=\frac{\overline{\beta}_{t}^{L}\left(1-\lambda^{L}% \right)}{1-\overline{\beta}_{t}^{L}\lambda^{L}}, (B.28) is in turn equivalent to l¯tL=(1δ)cβ¯tLλL\overline{l}_{t}^{L}=-\left(1-\delta\right)\frac{c}{\overline{\beta}_{t}^{L}% \lambda^{L}}, which is true by the definition of 𝒍¯L\overline{\bm{l}}^{L} in (B.25).


Step 3b: Next, we show that given the sequence 𝒍¯L\overline{\bm{l}}^{L}, it would be optimal for the low type to work in any period no matter the prior history of effort. Consider first the last period, TLT^{L}. No matter the history of prior effort, the current belief is some βTLLβ¯TLL,\beta_{T^{L}}^{L}\geq\overline{\beta}_{T^{L}}^{L}, hence βTLLλLl¯tLβ¯TLLλLl¯tL=c-\beta^{L}_{T^{L}}\lambda^{L}\overline{l}^{L}_{t}\geq\overline{\beta}^{L}_{T^{% L}}\lambda^{L}\overline{l}^{L}_{t}=c (where the equality is by definition), so that it is optimal to work in TLT^{L} .

Now assume inductively that the assertion is true for period t+1TLt+1\leq T^{L}, and consider period t<TLt<T^{L} after any history of prior effort, with current belief βtL\beta_{t}^{L}. Since we already showed that equation (B.26) holds, it follows from βtLβ¯tL\beta_{t}^{L}\geq\overline{\beta}^{L}_{t} that

βtLλL{l¯tL+s=t+1TLδst(1λL)s(t+1)[(1λL)l¯sLc]}c,-\beta_{t}^{L}\lambda^{L}\left\{\overline{l}_{t}^{L}+\sum\limits_{s=t+1}^{T^{L% }}\delta^{s-t}\left(1-\lambda^{L}\right)^{s-\left(t+1\right)}\left[\left(1-% \lambda^{L}\right)\overline{l}_{s}^{L}-c\right]\right\}\geq c,

and hence it is optimal for the agent to work in period tt.


Step 3c: Finally, we argue that any profile of penalties, 𝒍L\bm{l}^{L}, that makes the low type’s incentive constraint for effort bind at every period t{1,,TL}t\in\{1,\ldots,T^{L}\} must coincide with 𝒍¯L\overline{\bm{l}}^{L}, given that the penalty contract must induce work from the low type in each period up to TLT^{L}. Again, we use induction. Since l¯TLL\overline{l}^{L}_{T^{L}} is the unique penalty that makes the agent indifferent between working and shirking at period TLT^{L} given that he has worked in all prior periods, it follows that lTLL=l¯TLLl^{L}_{T^{L}}=\overline{l}^{L}_{T^{L}}. Note from Step 3b that it would remain optimal for the agent to work in period TLT^{L} given any profile of effort in prior periods.

For the inductive step, pick some period t<TLt<T^{L} and assume that in every period x{t,,TL}x\in\{t,\ldots,T^{L}\}, the agent is indifferent between working and shirking given that he has worked in all prior periods, and would also find it optimal to work at xx following any other profile of effort prior to xx. Under these hypotheses, the indifference at period t+1t+1 implies that

β¯t+1LλL{lt+1L+s=t+2TLδs(t+1)(1λL)s(t+2)[(1λL)lsLc]}=c.-\overline{\beta}_{t+1}^{L}\lambda^{L}\left\{l_{t+1}^{L}+\sum\limits_{s=t+2}^{% T^{L}}\delta^{s-\left(t+1\right)}\left(1-\lambda^{L}\right)^{s-\left(t+2\right% )}\left[\left(1-\lambda^{L}\right)l_{s}^{L}-c\right]\right\}=c. (B.29)

Given the inductive hypothesis, the incentive constraint for effort at period tt is

β¯tLλL{ltL+s=t+1TLδst(1λL)s(t+1)[(1λL)lsLc]}c,-\overline{\beta}_{t}^{L}\lambda^{L}\left\{l_{t}^{L}+\sum\limits_{s=t+1}^{T^{L% }}\delta^{s-t}\left(1-\lambda^{L}\right)^{s-\left(t+1\right)}\left[\left(1-% \lambda^{L}\right)l_{s}^{L}-c\right]\right\}\geq c,

which, when set to bind, can be written as

β¯tLλL{ltL+δ[(1λL)lt+1Lc]+δ(1λL)s=t+2TLδs(t+1)(1λL)s(t+2)[(1λL)lsLc]}=c.-\overline{\beta}_{t}^{L}\lambda^{L}\left\{l_{t}^{L}+\delta\left[\left(1-% \lambda^{L}\right)l_{t+1}^{L}-c\right]+\delta\left(1-\lambda^{L}\right)\sum% \limits_{s=t+2}^{T^{L}}\delta^{s-\left(t+1\right)}\left(1-\lambda^{L}\right)^{% s-\left(t+2\right)}\left[\left(1-\lambda^{L}\right)l_{s}^{L}-c\right]\right\}=c. (B.30)

Substituting (B.29) into (B.30) , using the fact that β¯t+1L=β¯tL(1λ)β¯tL(1λ)+1β¯tL\overline{\beta}^{L}_{t+1}=\frac{\overline{\beta}_{t}^{L}(1-\lambda)}{% \overline{\beta}_{t}^{L}(1-\lambda)+1-\overline{\beta}_{t}^{L}}, and performing some algebra shows that ltL=l¯tLl^{L}_{t}=\overline{l}^{L}_{t}. Moreover, by the reasoning in Step 3b, this also ensures that the agent would find it optimal to work in period tt for any other history of actions prior to period tt.

B.4 Step 4: The critical contract is optimal

By Step 2, we can restrict attention in solving program [RP2] to connected penalty contracts for the low type. For any TLT^{L}, Step 3 identified a particular sequence of penalties, 𝒍¯L(TL)\overline{\bm{l}}^{L}(T^{L}). We now show that any connected penalty contract for the low type that solves [RP2] must have precisely this penalty structure.

The proof involves two sub-steps; throughout, we hold an arbitrary TLT^{L} fixed and, to ease notation, drop the dependence of 𝒍¯L()\overline{\bm{l}}^{L}(\cdot) on TLT^{L}.


Step 4a: We first show that any connected penalty contract for the low type of length TLT^{L} that satisfies (ICLa{}_{a}^{L}) and has ltL>l¯tLl^{L}_{t}>\overline{l}^{L}_{t} in some period tTLt\leq T^{L} is not optimal. To prove this, consider any such connected penalty contract. Define

t^=max{t:tTL and ltL>l¯tL}.\hat{t}=\max\left\{t:t\leq T^{L}\text{ and }l_{t}^{L}>\overline{l}_{t}^{L}% \right\}.

Observe that we must have t^<TL\hat{t}<T^{L} because otherwise (ICLa{}_{a}^{L}) would be violated in period TLT^{L}. Furthermore, by definition of t^,\hat{t}, ltLl¯tLl_{t}^{L}\leq\overline{l}_{t}^{L} for all TLt>t^T^{L}\geq t>\hat{t}. We will prove that we can change the penalty structure by lowering lt^Ll^{L}_{\hat{t}} and raising some subsequent lsLl^{L}_{s} for s{t^+1,,TL}s\in\{\hat{t}+1,\ldots,T^{L}\} in a way that keeps type LL’s incentives for effort unchanged, and yet increase the value of the objective function (RP2).

Claim: There exists t~{t^+1,,TL}\widetilde{t}\in\{\hat{t}+1,\ldots,T^{L}\} such that (ICLa{}_{a}^{L}) at t~\widetilde{t} is slack and lt~L<l¯t~L.l_{\widetilde{t}}^{L}<\overline{l}_{\widetilde{t}}^{L}.

Proof: Suppose not, then for each TLt>t^T^{L}\geq t>\hat{t}, either ltL=l¯tL,l_{t}^{L}=\overline{l}_{t}^{L}, or ltL<l¯tLl_{t}^{L}<\overline{l}_{t}^{L} and (ICLa{}_{a}^{L}) binds. Then since whenever ltL<l¯tL,l_{t}^{L}<\overline{l}_{t}^{L}, (ICLa{}_{a}^{L}) binds by supposition, it must be that in all t>t^,t>\hat{t}, (ICLa{}_{a}^{L}) binds (this follows from Step 3). But then (ICLa{}_{a}^{L}) at t^\hat{t} is violated since lt^L>l¯t^Ll_{\hat{t}}^{L}>\overline{l}_{\hat{t}}^{L}. \parallel

Claim: There exists t¯{t^+1,,TL}\overline{t}\in\{\hat{t}+1,\ldots,T^{L}\} such that lt¯L<l¯t¯Ll_{\overline{t}}^{L}<\overline{l}_{\overline{t}}^{L} and for any t{t^+1,,t¯}t\in\left\{\hat{t}+1,...,\overline{t}\right\}, (ICLa{}_{a}^{L}) at tt is slack. In particular, we can take t¯\overline{t} to be the first such period after t^\hat{t}.

Proof: Fix t~\widetilde{t} in the previous claim. Note that (ICLa{}_{a}^{L}) at t^+1\hat{t}+1 must be slack because otherwise (ICLa{}_{a}^{L}) at t^\hat{t} is violated by lt^L>l¯t^Ll_{\hat{t}}^{L}>\overline{l}_{\hat{t}}^{L} and Step 3. There are two cases. (1) lt^+1L<l¯t^+1L;l_{\hat{t}+1}^{L}<\overline{l}_{\hat{t}+1}^{L}; then t^+1\hat{t}+1 is the t¯\overline{t} we want. (2) lt^+1L=l¯t^+1Ll_{\hat{t}+1}^{L}=\overline{l}_{\hat{t}+1}^{L}; in this case, since (ICLa{}_{a}^{L}) is slack at t^+1,\hat{t}+1, it must be that (ICLa{}_{a}^{L}) at t^+2\hat{t}+2 is slack (otherwise, the claim in Step 3 is violated); now if lt^+2L<l¯t^+2L,l_{\hat{t}+2}^{L}<\overline{l}_{\hat{t}+2}^{L}, we are done because t^+2\hat{t}+2 is the t¯\overline{t} we are looking for; if lt^+2L=l¯t^+2L,l_{\hat{t}+2}^{L}=\overline{l}_{\hat{t}+2}^{L}, then we continue to t^+3\hat{t}+3 and so on until we reach t~\widetilde{t} which we know gives us a slack (ICLa{}_{a}^{L}), lt~L<lt~L,l_{\widetilde{t}}^{L}<l_{\widetilde{t}}^{L}, and we are sure that (ICLa{}_{a}^{L}) is slack in all periods of this process before reaching t~.\widetilde{t}. \parallel


Now we shall show that we can slightly reduce lt^L>l¯t^Ll_{\hat{t}}^{L}>\overline{l}_{\hat{t}}^{L} and slightly increase lt¯L<l¯t¯Ll_{\overline{t}}^{L}<\overline{l}_{\overline{t}}^{L} and meanwhile keep the incentives for effort of type LL satisfied for all periods. By the same reasoning as used in Step 2, the incentive constraint for effort in period t^\hat{t} (given that the agent will work in all subsequent periods no matter his behavior at period tt) can be written as

βt^LλL{lt^L+t>t^δtt^(1λL)t(t^+1)[(1λL)ltLc]}c.-\beta_{\hat{t}}^{L}\lambda^{L}\left\{l_{\hat{t}}^{L}+\sum\limits_{t>\hat{t}}% \delta^{t-\hat{t}}\left(1-\lambda^{L}\right)^{t-(\hat{t}+1)}\left[\left(1-% \lambda^{L}\right)l_{t}^{L}-c\right]\right\}\geq c. (B.31)

Observe that if we reduce lt^Ll_{\hat{t}}^{L} by Δ>0\Delta>0 and increase lt¯Ll_{\overline{t}}^{L} by Δδt¯t^(1λL)t¯(t^+1)(1λL)=Δδt¯t^(1λL)t¯t^\frac{\Delta}{\delta^{\overline{t}-\hat{t}}\left(1-\lambda^{L}\right)^{% \overline{t}-(\hat{t}+1)}\left(1-\lambda^{L}\right)}=\frac{\Delta}{\delta^{% \overline{t}-\hat{t}}\left(1-\lambda^{L}\right)^{\overline{t}-\hat{t}}}, then the left-hand side of (B.31) does not change. Moreover, it follows that incentives for effort at t<t^t<\hat{t} are also unchanged (see Step 2), and the incentive condition at t¯\overline{t} will be satisfied if Δ\Delta is small enough because the original (ICLa{}_{a}^{L}) at t¯\overline{t} is slack.

Finally, we show that the modification above leads to a reduction of the rent of type HH in (RP2), i.e. raises the value of the objective. The rent is given by

μ0β0{t=1TLδtltL[(1λH)t(1λL)t]t=1TLδtc[(1λH)t1(1λL)t1]}.\mu_{0}\beta_{0}\left\{\sum\limits_{t=1}^{T^{L}}\delta^{t}l_{t}^{L}\left[\left% (1-\lambda^{H}\right)^{t}-\left(1-\lambda^{L}\right)^{t}\right]-\sum\limits_{t% =1}^{T^{L}}\delta^{t}c\left[\left(1-\lambda^{H}\right)^{t-1}-\left(1-\lambda^{% L}\right)^{t-1}\right]\right\}. (B.32)

Hence, the change in the rent from reducing lt^Ll_{\hat{t}}^{L} by Δ\Delta and increasing lt¯Ll_{\overline{t}}^{L} by Δδt¯t^(1λL)t¯t^\frac{\Delta}{\delta^{\overline{t}-\hat{t}}\left(1-\lambda^{L}\right)^{% \overline{t}-\hat{t}}} is

μ0β0δt^Δ{[(1λH)t^(1λL)t^]+1(1λL)t¯t^[(1λH)t¯(1λL)t¯]}\displaystyle\mu_{0}\beta_{0}\delta^{\hat{t}}\Delta\left\{-\left[\left(1-% \lambda^{H}\right)^{\hat{t}}-\left(1-\lambda^{L}\right)^{\hat{t}}\right]+\frac% {1}{\left(1-\lambda^{L}\right)^{\overline{t}-\hat{t}}}\left[\left(1-\lambda^{H% }\right)^{\overline{t}}-\left(1-\lambda^{L}\right)^{\overline{t}}\right]\right\}
=\displaystyle= μ0β0δt^Δ(1λH)t^(1λL)t¯t^[(1λH)t¯t^(1λL)t¯t^]<0,\displaystyle\mu_{0}\beta_{0}\delta^{\hat{t}}\Delta\frac{\left(1-\lambda^{H}% \right)^{\hat{t}}}{\left(1-\lambda^{L}\right)^{\overline{t}-\hat{t}}}\left[% \left(1-\lambda^{H}\right)^{\overline{t}-\hat{t}}-\left(1-\lambda^{L}\right)^{% \overline{t}-\hat{t}}\right]<0,

where the inequality is because t¯>t^\overline{t}>\hat{t} and 1λH<1λL1-\lambda^{H}<1-\lambda^{L}.


Step 4b: By Step 4a, we can restrict attention to penalty sequences 𝒍L\bm{l}^{L} such that ltLl¯tLl^{L}_{t}\leq\overline{l}^{L}_{t} for all tTLt\leq T^{L}. Now we show that unless 𝒍L()=¯𝒍L()\bm{l}^{L}(\cdot)=\overline{}\bm{l}^{L}(\cdot), the value of the objective (RP2) can be improved while satisfying the incentive constraint for effort, (ICLa{}_{a}^{L}). Recall that by Step 4a, (ICLa{}_{a}^{L}) is satisfied in all periods t=1,,TLt=1,\ldots,T^{L} whenever ltL=l¯tLl_{t}^{L}=\overline{l}_{t}^{L}. Thus, if ltL<l¯tLl_{t}^{L}<\overline{l}_{t}^{L} for any period, we can replace ltLl_{t}^{L} by l¯tL\overline{l}_{t}^{L} without affecting the effort incentives for type LL. Moreover, by doing this we reduce the rent of type HH, given by (B.32) above, and thus raise the value of (RP2).

B.5 Step 5: Generic uniqueness of the optimal contract for the low type

By Step 4, an optimal contract for the low type that solves program [RP2] can be found by optimizing over TLT^{L}, i.e. the length of connected penalty contracts with the penalty structure 𝒍¯L(TL)\overline{\bm{l}}^{L}(T^{L}). By Theorem 2, TLtLT^{L}\leq t^{L}. In this step, we establish generic uniqueness of the optimal contract for the low type. We proceed in two sub-steps.

Step 5a: First, we show that the optimal length TLT^{L} of connected penalty contracts with the penalty structure 𝒍¯L(TL)\overline{\bm{l}}^{L}(T^{L}) is generically unique. The portion of the objective (RP2) that involves TLT^{L} is

V(TL)\displaystyle V\left(T^{L}\right) :=\displaystyle:= (1μ0)[β0t=1TLδt(1λL)t1(λLc)(1β0)t=1TLδtc]\displaystyle\left(1-\mu_{0}\right)\left[\beta_{0}\sum\limits_{t=1}^{T^{L}}% \delta^{t}\left(1-\lambda^{L}\right)^{t-1}\left(\lambda^{L}-c\right)-(1-\beta_% {0})\sum\limits_{t=1}^{T^{L}}\delta^{t}c\right] (B.33)
μ0β0{t=1TLδtl¯tL(TL)[(1λH)t(1λL)t]t=1TLδtc[(1λH)t1(1λL)t1]},\displaystyle-\mu_{0}\beta_{0}\left\{\sum\limits_{t=1}^{T^{L}}\delta^{t}% \overline{l}_{t}^{L}(T^{L})\left[\left(1-\lambda^{H}\right)^{t}-\left(1-% \lambda^{L}\right)^{t}\right]-\sum\limits_{t=1}^{T^{L}}\delta^{t}c\left[\left(% 1-\lambda^{H}\right)^{t-1}-\left(1-\lambda^{L}\right)^{t-1}\right]\right\},\ % \ \ \ \ \ \ \ \ \

where we have used the desired penalty sequence. Note that by Theorem 2 TLtLT^{L}\leq t^{L} in any optimal contract for the low type and hence there is a finite number of maximizers of V(TL)V\left(T^{L}\right). It follows that if we perturb μ0\mu_{0} locally, the set of maximizers will not change. Now suppose that the maximizer of V(TL)V(T^{L}) is not unique. Without loss, pick any two maximizers T~L\widetilde{T}^{L} and T^L\widehat{T}^{L}. We must have V(T~L)=V(T^L)V(\widetilde{T}^{L})=V(\widehat{T}^{L}) and (again by Theorem 2) tLmax{T~L,T^L}t^{L}\geq\max\{\widetilde{T}^{L},\widehat{T}^{L}\}. Without loss, assume T^L>T~L\widehat{T}^{L}>\widetilde{T}^{L}. Note that the first term in square brackets in (B.33) is social surplus from the low type and hence it is strictly increasing in TLT^{L} for TL<tLT^{L}<t^{L}. Therefore, both the first and second terms in V(T^L)V(\widehat{T}^{L}) must be larger than the first and second terms in V(T~L)V(\widetilde{T}^{L}) respectively. But then it is immediate that perturbing μ0\mu_{0} within an arbitrarily small neighborhood will change the ranking of V(T~L)V(\widetilde{T}^{L}) and V(T^L)V(\widehat{T}^{L}), which implies that the assumed multiplicity is non-generic.

It follows that there is generically a unique TLT^{L} that maximizes V(TL)V(T^{L}); hereafter we denote this solution t¯L\overline{t}^{L}. In the non-generic cases in which multiple maximizers exist, we select the largest one.


Step 5b: In Step 5a we showed that among connected penalty contracts, there is generically a unique contract for type LL that solves [RP2]. We now claim that there generically cannot be any other penalty contract for type LL that solves [RP2]. Suppose, to contradiction, that this is false: there is an optimal non-connected penalty contract 𝐂L=(ΓL,W0L,𝒍L)\mathbf{C}^{L}=(\Gamma^{L},W^{L}_{0},\bm{l}^{L}) in which 𝟏𝜶L(𝐂L)\mathbf{1}\in\bm{\alpha}^{L}(\mathbf{C}^{L}). Let t<maxΓLt^{\circ}<\max\Gamma^{L} be the earliest lockout period in 𝐂L\mathbf{C}^{L}. Without loss, owing to genericity, we take δ<1\delta<1. Following the arguments of Step 2, in particular Remark 3, the optimality of 𝐂L\mathbf{C}^{L} implies that there are two connected penalty contracts that are also optimal: ^𝐂L=(T^L,W^0L,^𝒍L)\widehat{}\mathbf{C}^{L}=(\widehat{T}^{L},\widehat{W}^{L}_{0},\widehat{}\bm{l}% ^{L}) obtained from 𝐂L\mathbf{C}^{L} by applying Modification 1 of Step 2 as many times as needed to eliminate all lockout periods, and ~𝐂L=(T~L,W~0L,~𝒍L)\widetilde{}\mathbf{C}^{L}=(\widetilde{T}^{L},\widetilde{W}^{L}_{0},\widetilde% {}\bm{l}^{L}) obtained from 𝐂L\mathbf{C}^{L} by applying Modification 2 of Step 2 to shorten the contract by just eliminating all periods from tt^{\circ} on. Note that the modifications ensure that 𝟏𝜶L(^𝐂L)\mathbf{1}\in\bm{\alpha}^{L}(\widehat{}\mathbf{C}^{L}) and 𝟏𝜶L(~𝐂L)\mathbf{1}\in\bm{\alpha}^{L}(\widetilde{}\mathbf{C}^{L}). But now, the fact that T^L>T~L\widehat{T}^{L}>\widetilde{T}^{L} contradicts the generic uniqueness of connected penalty contracts for the low type that solve [RP2].

B.6 Step 6: Back to the original program

We have shown so far that there is a solution to program [RP2] in which the low type’s contract is a connected penalty contract of length t¯LtL\overline{t}^{L}\leq t^{L} and in which the penalty sequence is given by 𝒍¯L(t¯L)\overline{\bm{l}}^{L}(\overline{t}^{L}). In terms of optimizing over the high type’s contract, note that, as shown in Theorem 2, we can take the solution as inducing the high type to work in each period up to tHt^{H} and no longer: this follows from the fact that the portion of the objective in (RP2) involving the high type’s contract is social surplus from the high type.

Recall that solutions to [RP2] produce solutions to [RP1] by choosing W0LW^{L}_{0} to make (IRL) bind and W0HW^{H}_{0} to make (Weak-ICHL) bind, which can always be done. Accordingly, let 𝐂¯L=(t¯L,W¯0L,𝒍¯L(t¯L))\overline{\mathbf{C}}^{L}=(\overline{t}^{L},\overline{W}^{L}_{0},\overline{\bm% {l}}^{L}(\overline{t}^{L})) be the connected penalty contract where W¯0L\overline{W}^{L}_{0} is set to make (IRL) bind. Recall that [RP1] differs from program [P1] in that it imposes (Weak-ICHL) rather than (ICHL). We will argue that any solution to [RP1] using 𝐂¯L\overline{\mathbf{C}}^{L} satisfies (ICHL) and hence is also a solution to program [P1]. As shown in Step 2 of the proof of Theorem 2, contract 𝐂¯L\overline{\mathbf{C}}^{L} can then be combined with a suitable onetime-penalty contract for type HH to produce a solution to the principal’s original program [P].

We show that given any connected penalty contract of length TLtLT^{L}\leq t^{L} with penalty sequence 𝒍¯L(TL)\overline{\bm{l}}^{L}(T^{L}), it would be optimal for type HH to work in every period 1,,TL1,\ldots,T^{L}, no matter the history of prior effort. Fix any TLtLT^{L}\leq t^{L} and write 𝒍¯L𝒍¯L(TL)\overline{\bm{l}}^{L}\equiv\overline{\bm{l}}^{L}(T^{L}). The argument is by induction. Consider the last period, TLT^{L}. Since β¯TLLλLl¯TLL=c-\overline{\beta}_{T^{L}}^{L}\lambda^{L}\overline{l}_{T^{L}}^{L}=c, it follows from the fact that tH>tLt^{H}>t^{L} (hence β¯tHλH>β¯tLλL\overline{\beta}^{H}_{t}\lambda^{H}>\overline{\beta}^{L}_{t}\lambda^{L} for all t<tHt<t^{H}) that no matter the history of effort, βTLHλHl¯TLLc-\beta_{T^{L}}^{H}\lambda^{H}\overline{l}_{T^{L}}^{L}\geq c, i.e., regardless of the history, type HH will work in period TL.T^{L}. Now assume that it is optimal for type HH to work in period t+1TLt+1\leq T^{L} no matter the history of effort, and consider period tt with belief βtH\beta_{t}^{H}. This inductive hypothesis implies that

βt+1HλH{l¯t+1L+s=t+2TLδs(t+1)(1λH)s(t+2)[(1λH)l¯sLc]}c,-\beta_{t+1}^{H}\lambda^{H}\left\{\overline{l}_{t+1}^{L}+\sum\limits_{s=t+2}^{% T^{L}}\delta^{s-\left(t+1\right)}\left(1-\lambda^{H}\right)^{s-\left(t+2\right% )}\left[\left(1-\lambda^{H}\right)\overline{l}_{s}^{L}-c\right]\right\}\geq c,

or equivalently,

s=t+2TLδs(t+1)(1λH)s(t+2)[(1λH)l¯sLc]cβt+1HλHl¯t+1L.\sum\limits_{s=t+2}^{T^{L}}\delta^{s-\left(t+1\right)}\left(1-\lambda^{H}% \right)^{s-\left(t+2\right)}\left[\left(1-\lambda^{H}\right)\overline{l}_{s}^{% L}-c\right]\leq-\frac{c}{\beta_{t+1}^{H}\lambda^{H}}-\overline{l}_{t+1}^{L}. (B.34)

Therefore, at period t<TLt<T^{L}:

βtHλH{l¯tL+δ[(1λH)l¯t+1Lc]+δ(1λH)s=t+2TLδs(t+1)(1λH)s(t+2)[(1λH)l¯sLc]}\displaystyle-\beta_{t}^{H}\lambda^{H}\left\{\overline{l}_{t}^{L}+\delta\left[% \left(1-\lambda^{H}\right)\overline{l}_{t+1}^{L}-c\right]+\delta\left(1-% \lambda^{H}\right)\sum\limits_{s=t+2}^{T^{L}}\delta^{s-\left(t+1\right)}\left(% 1-\lambda^{H}\right)^{s-\left(t+2\right)}\left[\left(1-\lambda^{H}\right)% \overline{l}_{s}^{L}-c\right]\right\}
\displaystyle\geq βtHλH{l¯tL+δ[(1λH)l¯t+1Lc]+δ(1λH)(cβt+1HλHl¯t+1L)}\displaystyle-\beta_{t}^{H}\lambda^{H}\left\{\overline{l}_{t}^{L}+\delta\left[% \left(1-\lambda^{H}\right)\overline{l}_{t+1}^{L}-c\right]+\delta\left(1-% \lambda^{H}\right)\left(-\frac{c}{\beta_{t+1}^{H}\lambda^{H}}-\overline{l}_{t+% 1}^{L}\right)\right\}
=\displaystyle= βtHλH(l¯tLδc)+δ(1λH)βtHcβt+1H=βtHλHl¯tL+δcβ¯tLλLl¯tL+δc=c,\displaystyle-\beta_{t}^{H}\lambda^{H}\left(\overline{l}_{t}^{L}-\delta c% \right)+\delta\left(1-\lambda^{H}\right)\frac{\beta_{t}^{H}c}{\beta_{t+1}^{H}}% =-\beta_{t}^{H}\lambda^{H}\overline{l}_{t}^{L}+\delta c\geq-\overline{\beta}_{% t}^{L}\lambda^{L}\overline{l}_{t}^{L}+\delta c=c,

where the first inequality uses (B.34), the second equality uses βt+1H=βtH(1λH)1βtH+βtH(1λH)\beta^{H}_{t+1}=\frac{\beta^{H}_{t}(1-\lambda^{H})}{1-\beta^{H}_{t}+\beta^{H}_% {t}(1-\lambda^{H})}, and the final equality uses the fact that l¯tL=(1δ)cβ¯tLλL\overline{l}^{L}_{t}=-\frac{(1-\delta)c}{\overline{\beta}^{L}_{t}\lambda^{L}}.

Appendix C Proof of Theorem 5

We assume throughout this appendix that δ=1\delta=1. Without loss of optimality by Proposition 1, we focus on menus of penalty contracts. In this appendix, we will introduce programs and constraints that have analogies with those used in Appendix B for the case of tH>tLt^{H}>t^{L}. Accordingly, we often use the same labels for equations as before, but the reader should bear in mind that all references in this appendix to such equations are to those defined in this appendix.

Outline

. Since this is a long proof, let us outline the components. We begin in Step 1 by taking program [P2] from the proof of Theorem 2 for the case of δ=1\delta=1; we continue to call this program [P2]. Note that a critical difference here relative to the relaxed program [RP2] in the proof of Theorem 3 is that the current program [P2] does not constrain what the high type must do when taking the low type’s contract.

In Step 2, we show that there is an optimal penalty contract for type LL that is connected. In Step 3, we develop three lemmas pertaining to properties of the set 𝜶H(𝐂L)\bm{\alpha}^{H}(\mathbf{C}^{L}) in any 𝐂L\mathbf{C}^{L} that is an optimal contract for type LL. We then use these lemmas in Step 4 to show that in solving [P2], we can restrict attention to connected penalty contracts 𝐂L\mathbf{C}^{L} for type LL such that 𝜶H(𝐂L)\bm{\alpha}^{H}(\mathbf{C}^{L}) includes a stopping strategy with the most work property, i.e., an action plan that involves consecutive work for some number of periods followed by shirking thereafter, and where the number of work periods is larger than in any action plan in 𝜶H(𝐂L)\bm{\alpha}^{H}(\mathbf{C}^{L}). Building on the restriction to stopping strategies, we then show in Step 5 that there is always an optimal contract for type LL that is a onetime-penalty contract.

The last step, Step 6, is relegated to the Supplementary Appendix. For an arbitrary time TLT^{L}, this step first defines a particular last-period penalty lTLL(TL)l^{L}_{T^{L}}(T^{L}) and an associated time THL(TL)TLT^{HL}(T^{L})\leq T^{L}, and then establishes that if TLT^{L} is the optimal length of experimentation for type LL, there is an optimal onetime-penalty contract for type LL with penalty lTLL(TL)l^{L}_{T^{L}}(T^{L}) and in which type HH’s most-work optimal stopping strategy involves THL(TL)T^{HL}(T^{L}) periods of work. Hence, using lTLL(TL)l^{L}_{T^{L}}(T^{L}) and THL(TL)T^{HL}(T^{L}), an optimal contract for type LL that solves [P2] can be found by optimizing over the length TLT^{L}. By Theorem 2, the optimal length, t¯L\overline{t}^{L}, is no larger than the first-best stopping time, tLt^{L}.

C.1 Step 1: The principal’s program

By Step 1 and Step 2 in the proof of Theorem 2, we work with the principal’s program [P2]. Here we restate this program given δ=1\delta=1:

max𝐂H𝒞𝐂L𝒞𝐚H𝐚HL𝜶H(𝐂L){μ0{β0tΓH[sΓHst1(1asHλH)]atH(λHc)(1β0)tΓHatHc}+(1μ0){β0tΓL[sΓLst1(1λL)](λLc)(1β0)tΓLc}μ0{β0tΓLltL[sΓLst(1asHLλH)sΓLst(1λL)]β0ctΓLatHL[sΓLst1(1asHLλH)sΓLst1(1λL)]+ctΓL(1atHL)[(1β0)+β0sΓLst1(1λL)]}Information rent of type H}\displaystyle\max\limits_{\begin{subarray}{c}\mathbf{C}^{H}\in\mathcal{C}\\ \mathbf{C}^{L}\in\mathcal{C}\\ \mathbf{a}^{H}\\ \mathbf{a}^{HL}\in\bm{\alpha}^{H}(\mathbf{C}^{L})\end{subarray}}\left\{\begin{% array}[]{l}\mu_{0}\left\{\beta_{0}\sum\limits_{t\in\Gamma^{H}}\left[\prod% \limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{H}}{s\leq t-1}}\left(1-a_{s}^{H}% \lambda^{H}\right)\right]a_{t}^{H}\left(\lambda^{H}-c\right)-(1-\beta_{0})\sum% \limits_{t\in\Gamma^{H}}a_{t}^{H}c\right\}\\ +\left(1-\mu_{0}\right)\left\{\beta_{0}\sum\limits_{t\in\Gamma^{L}}\left[\prod% \limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{L}}{s\leq t-1}}\left(1-\lambda^{L}% \right)\right]\left(\lambda^{L}-c\right)-(1-\beta_{0})\sum\limits_{t\in\Gamma^% {L}}c\right\}\\ -\mu_{0}\underbrace{\left\{\begin{array}[]{l}\beta_{0}\sum\limits_{t\in\Gamma^% {L}}l_{t}^{L}\left[\prod\limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{L}}{s\leq t% }}\left(1-a_{s}^{HL}\lambda^{H}\right)-\prod\limits_{\genfrac{}{}{0.0pt}{}{s% \in\Gamma^{L}}{s\leq t}}\left(1-\lambda^{L}\right)\right]\\ -\beta_{0}c\sum\limits_{t\in\Gamma^{L}}a_{t}^{HL}\left[\prod\limits_{\genfrac{% }{}{0.0pt}{}{s\in\Gamma^{L}}{s\leq t-1}}\left(1-a_{s}^{HL}\lambda^{H}\right)-% \prod\limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{L}}{s\leq t-1}}\left(1-\lambda% ^{L}\right)\right]\\ +c\sum\limits_{t\in\Gamma^{L}}\left(1-a_{t}^{HL}\right)\left[(1-\beta_{0})+% \beta_{0}\prod\limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{L}}{s\leq t-1}}\left(% 1-\lambda^{L}\right)\right]\end{array}\right\}}_{\text{Information rent of % type $H$}}\end{array}\right\} (P2)

subject to

𝟏argmax(at)tΓL{β0tΓL[sΓLst1(1asλL)][(1atλL)ltLatc]+(1β0)tΓL(ltLatc)+W0L},\displaystyle\mathbf{1}\in\operatorname*{arg\,max}_{\left(a_{t}\right)_{t\in% \Gamma^{L}}}\left\{\begin{array}[]{l}\beta_{0}\sum\limits_{t\in\Gamma^{L}}% \left[\prod\limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{L}}{s\leq t-1}}\left(1-a% _{s}\lambda^{L}\right)\right]\left[\left(1-a_{t}\lambda^{L}\right)l_{t}^{L}-a_% {t}c\right]+(1-\beta_{0})\sum\limits_{t\in\Gamma^{L}}\left(l_{t}^{L}-a_{t}c% \right)+W_{0}^{L}\end{array}\right\}, (ICLa{}_{a}^{L})
𝐚Hargmax(at)tΓH{β0tΓH[sΓHst1(1asλH)][(1atλH)ltHatc]+(1β0)tΓH(ltHatc)+W0H}.\displaystyle\mathbf{a}^{H}\in\operatorname*{arg\,max}_{\left(a_{t}\right)_{t% \in\Gamma^{H}}}\left\{\begin{array}[]{l}\beta_{0}\sum\limits_{t\in\Gamma^{H}}% \left[\prod\limits_{\genfrac{}{}{0.0pt}{}{s\in\Gamma^{H}}{s\leq t-1}}\left(1-a% _{s}\lambda^{H}\right)\right]\left[\left(1-a_{t}\lambda^{H}\right)l_{t}^{H}-a_% {t}c\right]+(1-\beta_{0})\sum\limits_{t\in\Gamma^{H}}\left(l_{t}^{H}-a_{t}c% \right)+W_{0}^{H}\end{array}\right\}. (ICHa{}_{a}^{H})

As in the second step in the proof of Theorem 2, the information rent of type HH when he takes action plan 𝐚\mathbf{a} under type LL’s contract 𝐂L\mathbf{C}^{L} is given by R(𝐂L,𝐚)=U0H(𝐂L,𝐚)U0L(𝐂L,𝟏)R\left(\mathbf{C}^{L},\mathbf{a}\right)=U_{0}^{H}\left(\mathbf{C}^{L},\mathbf{% a}\right)-U_{0}^{L}\left(\mathbf{C}^{L},\mathbf{1}\right), and R(𝐂L,𝐚)=R(𝐂L,𝐚)R\left(\mathbf{C}^{L},\mathbf{a}\right)=R\left(\mathbf{C}^{L},\mathbf{a}^{% \prime}\right) whenever 𝐚,𝐚𝜶H(𝐂L).\mathbf{a},\mathbf{a}^{\prime}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right). The difference in information rents under contracts 𝐂^L\widehat{\mathbf{C}}^{L} and 𝐂L\mathbf{C}^{L} and corresponding action plans 𝐚^\widehat{\mathbf{a}} and 𝐚\mathbf{a} is:

R(𝐂^L,𝐚^)R(𝐂L,𝐚)=(U0H(𝐂^L,𝐚^)U0H(𝐂L,𝐚))(U0L(𝐂^L,𝟏)U0L(𝐂L,𝟏)).R\big{(}\widehat{\mathbf{C}}^{L},\widehat{\mathbf{a}}\big{)}-R\left(\mathbf{C}% ^{L},\mathbf{a}\right)=\left(U_{0}^{H}\big{(}\widehat{\mathbf{C}}^{L},\widehat% {\mathbf{a}}\big{)}-U_{0}^{H}\left(\mathbf{C}^{L},\mathbf{a}\right)\right)-% \left(U_{0}^{L}\big{(}\widehat{\mathbf{C}}^{L},\mathbf{1}\big{)}-U_{0}^{L}% \left(\mathbf{C}^{L},\mathbf{1}\right)\right). (C.9)

When the action plan does not change across contracts (i.e. 𝐚=𝐚^\mathbf{a}=\widehat{\mathbf{a}} above), (C.9) specializes to

R(𝐂^L,𝐚)R(𝐂L,𝐚)=β0tΓL(l^tLltL)[sΓL,st(1asλH)sΓL,st(1λL)].R\big{(}\widehat{\mathbf{C}}^{L},\mathbf{a}\big{)}-R\left(\mathbf{C}^{L},% \mathbf{a}\right)=\beta_{0}\sum\limits_{t\in\Gamma^{L}}\left(\widehat{l}_{t}^{% L}-l_{t}^{L}\right)\left[\prod\limits_{s\in\Gamma^{L},s\leq t}\left(1-a_{s}% \lambda^{H}\right)-\prod\limits_{s\in\Gamma^{L},s\leq t}\left(1-\lambda^{L}% \right)\right]. (C.10)

C.2 Step 2: Connected contracts for the low type

We now claim that in program [P2], it is without loss to consider solutions in which the low type’s contract is a connected penalty contract, i.e. solutions 𝐂L\mathbf{C}^{L} in which ΓL={1,,TL}\Gamma^{L}=\left\{1,\ldots,T^{L}\right\} for some TL.T^{L}.

To avoid trivialities, consider any optimal non-connected 𝐂L\mathbf{C}^{L} with ΓL\Gamma^{L}\neq\emptyset. Let tt^{\circ} be the earliest lockout period in 𝐂L\mathbf{C}^{L}, i.e. t=min{t:t>0,tΓL}t^{\circ}=\min\{t:t>0,t\notin\Gamma^{L}\}. Consider a modified penalty contract 𝐂^L\widehat{\mathbf{C}}^{L} that removes the lockout period tt^{\circ} and shortens the contract by one period as follows:

W^0L=W0L,Γ^L={1,,t1}{s:st and s+1ΓL},l^sL={lsLif st1,ls+1Lif st and sΓ^L.\displaystyle\widehat{W}_{0}^{L}=W^{L}_{0},\ \ \widehat{\Gamma}^{L}=\{1,\ldots% ,t^{\circ}-1\}\cup\{s:s\geq t^{\circ}\text{ and }s+1\in\Gamma^{L}\},\ \ % \widehat{l}_{s}^{L}=\begin{cases}l^{L}_{s}&\text{if }s\leq t^{\circ}-1,\\ l^{L}_{s+1}&\text{if }s\geq t^{\circ}\text{ and }s\in\widehat{\Gamma}^{L}.\end% {cases}

Given δ=1\delta=1, it is straightforward that it remains optimal for type LL to work in every period in Γ^L\widehat{\Gamma}^{L}, and given any optimal action plan for type HH under the original contract, 𝐚HL𝜶H(𝐂L)\mathbf{a}^{HL}\in\bm{\alpha}^{H}(\mathbf{C}^{L}), the action plan

𝐚^HL=(a^sHL)sΓ^L\displaystyle\widehat{\mathbf{a}}^{HL}=(\widehat{a}^{HL}_{s})_{s\in\widehat{% \Gamma}^{L}} =\displaystyle= {asHLif st1,as+1HLif st and sΓ^L,\displaystyle\begin{cases}a^{HL}_{s}&\text{if }s\leq t^{\circ}-1,\\ a^{HL}_{s+1}&\text{if }s\geq t^{\circ}\text{ and }s\in\widehat{\Gamma}^{L},% \end{cases}

is optimal for type HH under the modified contract, i.e. 𝐚^HL𝜶H(𝐂^L)\widehat{\mathbf{a}}^{HL}\in\bm{\alpha}^{H}(\widehat{\mathbf{C}}^{L}). Given no discounting, it is also immediate that the surplus generated by type LL is unchanged by the modification. It thus follows that the value of (P2) is unchanged by the modification. This procedure can be applied iteratively to all lockout periods to produce a connected contract.

C.3 Step 3: Optimal deviation action plans for the high type

By the previous steps, we can restrict our attention to connected penalty contracts 𝐂L=(TL,W0L,𝒍L)\mathbf{C}^{L}=(T^{L},W^{L}_{0},\bm{l}^{L}) that induce effort from the low type in each period t{1,,TL}t\in\{1,\ldots,T^{L}\}. We now describe properties of an optimal connected penalty contract for the low type (Step 3a) and an optimal action plan for the high type when taking the low type’s contract (Step 3b).

Step 3a: Consider an optimal connected penalty contract for type LL, 𝐂L=(TL,W0L,𝒍L)\mathbf{C}^{L}=(T^{L},W_{0}^{L},\bm{l}^{L}). The next two lemmas describe properties of such a contract.

Lemma 1.

Suppose that 𝐂L=(TL,W0L,𝐥L)\mathbf{C}^{L}=(T^{L},W_{0}^{L},\bm{l}^{L}) is an optimal contract for type LL. Then for any t=1,,TL,t=1,\ldots,T^{L}, there exists an optimal action plan 𝐚𝛂H(𝐂L)\mathbf{a}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right) such that at=1.a_{t}=1.

Proof.

Suppose to the contrary that for some τ{1,,TL},\tau\in\left\{1,\ldots,T^{L}\right\}, aτ=0a_{\tau}=0 for all 𝐚𝜶H(𝐂L)\mathbf{a}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right). For any ε>0\varepsilon>0, define a contract 𝐂L(ε)=(TL,W0L,𝒍L(ε))\mathbf{C}^{L}\left(\varepsilon\right)=(T^{L},W_{0}^{L},\bm{l}^{L}(\varepsilon)) modified from 𝐂L=(TL,W0L,𝒍L)\mathbf{C}^{L}=(T^{L},W_{0}^{L},\bm{l}^{L}) as follows: (i) lτL(ε)=lτLεl_{\tau}^{L}\left(\varepsilon\right)=l_{\tau}^{L}-\varepsilon; (ii) lτ1L(ε)=lτ1L+ε(1λL)l_{\tau-1}^{L}\left(\varepsilon\right)=l_{\tau-1}^{L}+\varepsilon\left(1-% \lambda^{L}\right); and (iii) ltL(ε)=ltLl_{t}^{L}\left(\varepsilon\right)=l_{t}^{L} if t{τ1,τ}t\notin\{\tau-1,\tau\}. We derive a contradiction by showing that for small enough ε>0\varepsilon>0, 𝐂L(ε)\mathbf{C}^{L}\left(\varepsilon\right) together with an original optimal contract for type HH, 𝐂H,\mathbf{C}^{H}, is feasible in [P2] and strictly improves the objective. Note that by construction, (𝐂L(ε),𝐂H)\left(\mathbf{C}^{L}\left(\varepsilon\right),\mathbf{C}^{H}\right) satisfy (ICLa{}_{a}^{L}) and (ICHa{}_{a}^{H}). To evaluate how the objective changes when 𝐂L(ε)\mathbf{C}^{L}\left(\varepsilon\right) is used instead of 𝐂L\mathbf{C}^{L}, we thus only need to consider the difference in the information rents associated with these contracts, R(𝐂L(ε))R(𝐂L)R\left(\mathbf{C}^{L}\left(\varepsilon\right)\right)-R\left(\mathbf{C}^{L}\right).

We first claim that 𝜶H(𝐂L(ε))𝜶H(𝐂L)\bm{\alpha}^{H}\left(\mathbf{C}^{L}\left(\varepsilon\right)\right)\subseteq\bm% {\alpha}^{H}\left(\mathbf{C}^{L}\right) when ε\varepsilon is small enough. To see this, fix any 𝐚𝜶H(𝐂L)\mathbf{a}\in\bm{\alpha}^{H}(\mathbf{C}^{L}). Since the set of action plans is discrete, the optimality of 𝐚\mathbf{a} implies that there is some η>0\eta>0 such that U0H(𝐂L,𝐚)>U0H(𝐂L,𝐚)+ηU_{0}^{H}\left(\mathbf{C}^{L},\mathbf{a}\right)>U_{0}^{H}\left(\mathbf{C}^{L},% \mathbf{a^{\prime}}\right)+\eta for any 𝐚𝜶H(𝐂L).\mathbf{a^{\prime}}\notin\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right). Since U0H(𝐂L(ε),𝐚)U_{0}^{H}\left(\mathbf{C}^{L}\left(\varepsilon\right),\mathbf{a^{\prime}}\right) is continuous in ε,\varepsilon, it follows immediately that for all ε\varepsilon small enough and all 𝐚𝜶H(𝐂L)\mathbf{a^{\prime}}\notin{\bm{\alpha}}^{H}(\mathbf{C}^{L}): U0H(𝐂L(ε),𝐚)>U0H(𝐂L(ε),𝐚)+ηU_{0}^{H}\left(\mathbf{C}^{L}\left(\varepsilon\right),\mathbf{a}\right)>U_{0}^% {H}\left(\mathbf{C}^{L}\left(\varepsilon\right),\mathbf{a^{\prime}}\right)+\eta. Thus, 𝐚𝜶H(𝐂L(ε))\mathbf{a^{\prime}}\notin\bm{\alpha}^{H}\left(\mathbf{C}^{L}(\varepsilon)\right). It follows that 𝜶H(𝐂L(ε))𝜶H(𝐂L)\bm{\alpha}^{H}\left(\mathbf{C}^{L}\left(\varepsilon\right)\right)\subseteq\bm% {\alpha}^{H}\left(\mathbf{C}^{L}\right).

Next, for small enough ε\varepsilon, take 𝐚𝜶H(𝐂L(ε))𝜶H(𝐂L)\mathbf{a}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\left(\varepsilon\right)\right% )\subseteq\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right). Since aτ=0a_{\tau}=0 by assumption, (C.10) implies

R(𝐂L(ε),𝐚)R(𝐂L,𝐚)=\displaystyle R\left(\mathbf{C}^{L}\left(\varepsilon\right),\mathbf{a}\right)-% R\left(\mathbf{C}^{L},\mathbf{a}\right)= β0ε(1λL)[s=1τ1(1asλH)(1λL)τ1]β0ε[s=1τ(1asλH)(1λL)τ]\displaystyle\beta_{0}\varepsilon\left(1-\lambda^{L}\right)\left[\prod\limits_% {s=1}^{\tau-1}\left(1-a_{s}\lambda^{H}\right)-\left(1-\lambda^{L}\right)^{\tau% -1}\right]-\beta_{0}\varepsilon\left[\prod\limits_{s=1}^{\tau}\left(1-a_{s}% \lambda^{H}\right)-\left(1-\lambda^{L}\right)^{\tau}\right]
=\displaystyle= λLβ0εs=1τ1(1asλH)<0.\displaystyle-\lambda^{L}\beta_{0}\varepsilon\prod\limits_{s=1}^{\tau-1}\left(% 1-a_{s}\lambda^{H}\right)<0.

Hence, 𝐂L(ε)\mathbf{C}^{L}\left(\varepsilon\right) strictly improves the objective relative to 𝐂L\mathbf{C}^{L}. ∎

Lemma 2.

Suppose that 𝐂L=(TL,W0L,𝐥L)\mathbf{C}^{L}=(T^{L},W_{0}^{L},\bm{l}^{L}) is an optimal contract for type LL and there is some τ{1,,TL}\tau\in\left\{1,\ldots,T^{L}\right\} such that aτ=1a_{\tau}=1 for all 𝐚𝛂H(𝐂L)\mathbf{a}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right). Then (ICLa{}_{a}^{L}) binds at τ.\tau.

Proof.

Recall from (ICLa{}_{a}^{L}) that 𝐚L=𝟏.\mathbf{a}^{L}=\mathbf{1}. Suppose to the contrary that (ICLa{}_{a}^{L}) is not binding at some τ\tau but aτ=1a_{\tau}=1 for all 𝐚𝜶H(𝐂L).\mathbf{a}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right). For any ε>0\varepsilon>0, define a contract 𝐂L(ε)=(TL,W0L,𝒍L(ε))\mathbf{C}^{L}\left(\varepsilon\right)=(T^{L},W_{0}^{L},\bm{l}^{L}(\varepsilon)) modified from 𝐂L=(TL,W0L,𝒍L)\mathbf{C}^{L}=(T^{L},W_{0}^{L},\bm{l}^{L}) as follows: (i) lτL(ε)=lτL+εl_{\tau}^{L}\left(\varepsilon\right)=l_{\tau}^{L}+\varepsilon; (ii) lτ1L(ε)=lτ1Lε(1λL)l_{\tau-1}^{L}\left(\varepsilon\right)=l_{\tau-1}^{L}-\varepsilon\left(1-% \lambda^{L}\right); and (iii) ltL(ε)=ltLl_{t}^{L}\left(\varepsilon\right)=l_{t}^{L} if t{τ1,τ}t\notin\{\tau-1,\tau\}. We derive a contradiction by showing that for small enough ε>0\varepsilon>0, 𝐂L(ε)\mathbf{C}^{L}\left(\varepsilon\right) together with an original optimal contract for type HH, 𝐂H,\mathbf{C}^{H}, is feasible in [P2] and strictly improves the objective. Note that by construction (ICLa{}_{a}^{L}) is still satisfied under 𝐂L(ε)\mathbf{C}^{L}\left(\varepsilon\right) at t=1,,τ1,τ+1,,TL.t=1,\ldots,\tau-1,\tau+1,\ldots,T^{L}. Moreover, since (ICLa{}_{a}^{L}) is slack at τ\tau under contract 𝐂L\mathbf{C}^{L}, it continues to be slack at τ\tau under 𝐂L(ε)\mathbf{C}^{L}(\varepsilon) for ε\varepsilon small enough.

Now for small enough ε\varepsilon, take any 𝐚𝜶H(𝐂L(ε))𝜶H(𝐂L)\mathbf{a}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}(\varepsilon)\right)\subseteq% \bm{\alpha}^{H}\left(\mathbf{C}^{L}\right), where the subset inequality follows from the arguments in the proof of Lemma 1. Recall that by assumption aτ=1a_{\tau}=1. Using (C.10)\eqref{Rdiff1},

R(𝐂L(ε),𝐚)R(𝐂L,𝐚)=\displaystyle R\left(\mathbf{C}^{L}\left(\varepsilon\right),\mathbf{a}\right)-% R\left(\mathbf{C}^{L},\mathbf{a}\right)= β0ε(1λL)[s=1τ1(1asλH)(1λL)τ1]+β0ε[s=1τ(1asλH)(1λL)τ]\displaystyle-\beta_{0}\varepsilon\left(1-\lambda^{L}\right)\left[\prod\limits% _{s=1}^{\tau-1}\left(1-a_{s}\lambda^{H}\right)-\left(1-\lambda^{L}\right)^{% \tau-1}\right]+\beta_{0}\varepsilon\left[\prod\limits_{s=1}^{\tau}\left(1-a_{s% }\lambda^{H}\right)-\left(1-\lambda^{L}\right)^{\tau}\right]
=\displaystyle= β0ε{s=1τ1(1asλH)[(1λL)+(1λH)]}\displaystyle\beta_{0}\varepsilon\left\{\prod\limits_{s=1}^{\tau-1}\left(1-a_{% s}\lambda^{H}\right)\left[-\left(1-\lambda^{L}\right)+\left(1-\lambda^{H}% \right)\right]\right\}
=\displaystyle= (λHλL)β0εs=1τ1(1asλH)\displaystyle-\left(\lambda^{H}-\lambda^{L}\right)\beta_{0}\varepsilon\prod% \limits_{s=1}^{\tau-1}\left(1-a_{s}\lambda^{H}\right)
<\displaystyle< 0.\displaystyle 0.

Hence, 𝐂L(ε)\mathbf{C}^{L}(\varepsilon) strictly improves the objective relative to 𝐂L\mathbf{C}^{L}. ∎

Step 3b: For any 𝐚𝜶H(𝐂L)\mathbf{a}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right) and s<t,s<t, define

D(s,t,𝐚):=τ=st1lτL(1λH)n=s+1τan.D\left(s,t,\mathbf{a}\right):=\sum\limits_{\tau=s}^{t-1}l^{L}_{\tau}\left(1-% \lambda^{H}\right)^{\sum\nolimits_{n=s+1}^{\tau}a_{n}}.

The next lemma describes properties of any action plan 𝐚𝜶H(𝐂L)\mathbf{a}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right).

Lemma 3.

Suppose 𝐚𝛂H(𝐂L)\mathbf{a}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right) and s<t.s<t.

(1) If D(s,t,𝐚)>0D\left(s,t,\mathbf{a}\right)>0 and as=1,a_{s}=1, then at=1.a_{t}=1.

(2) If D(s,t,𝐚)<0D\left(s,t,\mathbf{a}\right)<0 and as=0,a_{s}=0, then at=0.a_{t}=0.

(3) If D(s,t,𝐚)=0,D\left(s,t,\mathbf{a}\right)=0, then 𝐚𝛂H(𝐂L)\mathbf{a}^{\prime}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right) where as=at,a_{s}^{\prime}=a_{t}, at=as,a_{t}^{\prime}=a_{s}, and aτ=aτa_{\tau}^{\prime}=a_{\tau} if τs,t\tau\neq s,t.

Proof.

Consider the first case of D(s,t,𝐚)>0.D\left(s,t,\mathbf{a}\right)>0. Suppose to the contrary that for some optimal action plan 𝐚\mathbf{a} and two periods s<t,s<t, we have D(s,t,𝐚)>0D\left(s,t,\mathbf{a}\right)>0 and as=1a_{s}=1 but at=0.a_{t}=0. Consider an action plan 𝐚\mathbf{a}^{\prime} such that 𝐚\mathbf{a}^{\prime} and 𝐚\mathbf{a} agree except that as=0a_{s}^{\prime}=0 and at=1.a_{t}^{\prime}=1. That is,

𝐚\displaystyle\mathbf{a} =\displaystyle= (,1period s,,0period t,),\displaystyle(\cdots,\underbrace{1}_{\text{period $s$}},\cdots,\underbrace{0}_% {\text{period $t$}},\cdots),
𝐚\displaystyle\mathbf{a}^{\prime} =\displaystyle= (,0period s,,1period t,).\displaystyle(\cdots,\underbrace{0}_{\text{period $s$}},\cdots,\underbrace{1}_% {\text{period $t$}},\cdots).

Let UsH(𝐂L,𝐚)U_{s}^{H}\left(\mathbf{C}^{L},\mathbf{a}\right) be type HH’s payoff evaluated at the beginning of period ss. Then,

UsH(𝐂L,𝐚)UsH(𝐂L,𝐚)=βsHλHτ=st1lτL(1λH)n=s+1τan=βsHλHD(s,t,𝐚).U_{s}^{H}\left(\mathbf{C}^{L},\mathbf{a}\right)-U_{s}^{H}\left(\mathbf{C}^{L},% \mathbf{a}^{\prime}\right)=-\beta_{s}^{H}\lambda^{H}\sum\nolimits_{\tau=s}^{t-% 1}l_{\tau}^{L}\left(1-\lambda^{H}\right)^{\sum\nolimits_{n=s+1}^{\tau}a_{n}}=-% \beta_{s}^{H}\lambda^{H}D\left(s,t,\mathbf{a}\right).

The intuition for this expression is as follows. Since action plans 𝐚\mathbf{a} and 𝐚\mathbf{a}^{\prime} have the same number of working periods, the assumption of no discounting implies that neither the effort costs nor the penalty sequence matters for the difference in utilities conditional on the bad state. Conditional on the good state, the effort costs again do not affect the difference in utilities; however, the probability with which the agent receives ltτl_{t}^{\tau} for any τ{s,,t1}\tau\in\{s,\ldots,t-1\} is “shifted up” in 𝐚\mathbf{a}^{\prime} as compared to 𝐚\mathbf{a}.

Therefore, UsH(𝐂L,𝐚)UsH(𝐂L,𝐚)<0U_{s}^{H}\left(\mathbf{C}^{L},\mathbf{a}\right)-U_{s}^{H}\left(\mathbf{C}^{L},% \mathbf{a}^{\prime}\right)<0 if D(s,t,𝐚)>0.D\left(s,t,\mathbf{a}\right)>0. But this contradicts the assumption that 𝐚\mathbf{a} is optimal; hence, the claim in part (1) follows. The proof of part (2) is analogous.

Finally, consider part (3). The claim is trivial if as=at.a_{s}=a_{t}. If as=1a_{s}=1 and at=0a_{t}=0, then UsH(𝐂L,𝐚)UsH(𝐂L,𝐚)=0U_{s}^{H}\left(\mathbf{C}^{L},\mathbf{a}\right)-U_{s}^{H}\left(\mathbf{C}^{L},% \mathbf{a}^{\prime}\right)=0 from the argument above; hence, both 𝐚\mathbf{a} and 𝐚\mathbf{a}^{\prime} are optimal. The case of as=0a_{s}=0 and at=1a_{t}=1 is analogous. ∎

C.4 Step 4: Stopping strategies for the high type

We use the following concepts to characterize the solution to [P2]:

Definition 1.

An action plan 𝐚\mathbf{a} is a stopping strategy (that stops at tt) if there exists t1t\geq 1 such that as=1a_{s}=1 for sts\leq t and as=0a_{s}=0 for s>ts>t.

Definition 2.

An optimal action plan for type θ\theta under contract 𝐂\mathbf{C}, 𝐚𝜶θ(𝐂)\mathbf{a}\in\bm{\alpha}^{\theta}(\mathbf{C}), has the most-work property (or is a most-work optimal strategy) if no other optimal action plan under the contract has more work periods; that is, for all 𝐚𝜶θ(𝐂)\mathbf{a^{\prime}}\in\bm{\alpha}^{\theta}(\mathbf{C}), #{n:an=1}#{n:an=1}\#\left\{n:a_{n}=1\right\}\geq\#\left\{n:a^{\prime}_{n}=1\right\}.

Step 3 described properties of optimal contracts for the low type and optimal action plans for the high type under the low type’s contract. We now use these properties to show that in solving program [P2], we can restrict attention to connected penalty contracts for the low type 𝐂L=(TL,W0L,𝒍L)\mathbf{C}^{L}=(T^{L},W_{0}^{L},\bm{l}^{L}) such that there is an optimal action plan for the high type under the contract 𝐚𝜶H(𝐂L)\mathbf{a}\in\bm{\alpha}^{H}(\mathbf{C}^{L}) that is a stopping strategy with the most work property.

Let N=min𝐚𝜶H(𝐂L)#{n:an=0}.N=\min_{\mathbf{a}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right)}\#\left\{n:a_{% n}=0\right\}. That is, among all action plans that are optimal for type HH under contract 𝐂L,\mathbf{C}^{L}, the action plan in which type HH works the largest number of periods involves type HH shirking in NN periods. Let ANA_{N} be the set of optimal action plans that involve type HH shirking in NN periods. Let

AN,k={𝐚AN:at=0 for all t>TLk},A_{N,k}=\left\{\mathbf{a}\in A_{N}:a_{t}=0\text{ for all }t>T^{L}-k\right\},

i.e. any 𝐚AN,k\mathbf{a}\in A_{N,k} contains a total of NN shirking periods, (at least) kk of which are in the tail.

Our goal is to establish the following:

for any k<NAN,kn=k+1NAN,n.\text{for any $k<N$: $A_{N,k}\neq\emptyset\implies\bigcup_{n=k+1}^{N}A_{N,n}% \neq\emptyset$}. (C.11)

In other words, whenever ANA_{N} contains an action plan that has k<Nk<N shirks in the tail, ANA_{N} must contain an action plan that has at least k+1k+1 shirks in the tail. By induction, this implies AN,NA_{N,N}\neq\emptyset, which is equivalent to the existence of an optimal action plan that is a stopping strategy with the most work property.

Suppose to contradiction that (C.11) is not true; i.e. there is some k<Nk<N such that AN,kA_{N,k}\neq\emptyset and yet n=k+1NAN,n=\bigcup_{n=k+1}^{N}A_{N,n}=\emptyset. Then there exists

t^=min{t:𝐚AN,k,at=0,t<TLk,as=1 for each s=t+1,,TLk}.\hat{t}=\min\left\{t:\mathbf{a}\in A_{N,k},a_{t}=0,t<T^{L}-k,a_{s}=1\text{ for% each }s=t+1,\ldots,T^{L}-k\right\}. (C.12)

In words, t^\hat{t} is the smallest shirking period preceding a working period such that there is an optimal action plan 𝐚AN,k\mathbf{a}\in A_{N,k} with k+1k+1 shirking periods from (including) t^\hat{t}. Now take t^0=t^.\hat{t}_{0}=\hat{t}. For n=0,1,,n=0,1,\ldots, whenever {t:at=0,𝐚AN,k,t<t^n}\left\{t:a_{t}=0,\mathbf{a}\in A_{N,k},t<\hat{t}_{n}\right\}\neq\emptyset, define

t^n+1=min{t:at=0,𝐚AN,k,t<t^n,as=1 for each s=t+1,,t^n1}.\hat{t}_{n+1}=\min\left\{t:a_{t}=0,\mathbf{a}\in A_{N,k},t<\hat{t}_{n},a_{s}=1% \text{ for each }s=t+1,\ldots,\hat{t}_{n}-1\right\}. (C.13)

The sequence {t^n}\left\{\hat{t}_{n}\right\} uniquely pins down an action profile 𝐚^AN,k.\widehat{\mathbf{a}}\in A_{N,k}. In words, among all effort profiles in AN,k,A_{N,k}, 𝐚^\widehat{\mathbf{a}} has the earliest nn-th shirk for each n=1,,N.n=1,...,N. Note that 𝐚^\widehat{\mathbf{a}} takes the following form:

period: t^\hat{t} t^+1\hat{t}+1 \cdots TLkT^{L}-k TLk+1T^{L}-k+1 \cdots
𝐚^:\widehat{\mathbf{a}}\mathbf{:} 0 11 \cdots 11 0 \cdots

We will prove that n=k+1NAN,n\bigcup_{n=k+1}^{N}A_{N,n}\neq\emptyset (contradicting the hypothesis above) by showing that we can “move” the shirking in period t^\hat{t} of 𝐚^\widehat{\mathbf{a}} to the end. This is done via three lemmas.

Lemma 4.

Suppose AN,kA_{N,k}\neq\emptyset and n=k+1NAN,n=\bigcup_{n=k+1}^{N}A_{N,n}=\emptyset. Then ltL=0l^{L}_{t}=0 for any t=t^+1,,TLk1.t=\hat{t}+1,\ldots,T^{L}-k-1.

Proof.

We proceed by induction. Take any t{t^+1,,TLk1}t\in\left\{\hat{t}+1,\ldots,T^{L}-k-1\right\} and assume that lsL=0l^{L}_{s}=0 for s=t+1,,TLk1s=t+1,\ldots,T^{L}-k-1. We show that ltL=0l^{L}_{t}=0.

Step 1: ltL0.l^{L}_{t}\geq 0.

Proof of Step 1: Suppose not, i.e., ltL<0.l^{L}_{t}<0. Then the fact that (ICLa{}_{a}^{L}) is satisfied at period t+1t+1 and the hypothesis that ltL<0l^{L}_{t}<0 imply that (ICLa{}_{a}^{L}) is slack at period tt.5454 54 This can be proved along very similar lines to part (2) of Lemma 3. Hence, by Lemma 2, there exists an action plan 𝐚𝜶H(𝐂L)\mathbf{a}^{\prime}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right) such that at=0.a^{\prime}_{t}=0. Now, by the assumption that ltL<0l^{L}_{t}<0 together with the induction hypothesis, we obtain s=tmlsL<0\sum\nolimits_{s=t}^{m}l^{L}_{s}<0 for m{t,,TLk1}m\in\{t,\ldots,T^{L}-k-1\}. By Lemma 3, part (2), as=0a^{\prime}_{s}=0 for any s=t,,TLk.s=t,\ldots,T^{L}-k. Thus, 𝐚𝜶H(𝐂L)\mathbf{a}^{\prime}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right) is as follows:

period: t^\hat{t} t^+1\hat{t}+1 \cdots tt t+1t+1 \cdots TLk1T^{L}-k-1 TLkT^{L}-k TLk+1T^{L}-k+1 \cdots
𝐚^:\widehat{\mathbf{a}}\mathbf{:} 0 11 \cdots 1 1 \cdots 11 11 0 \cdots
𝐚:\mathbf{a}^{\prime}\mathbf{:} 0 0 \cdots 0 0

Claim 1: There exists s>TLks^{\ast}>T^{L}-k such that as=1.a^{\prime}_{s^{\ast}}=1.

Proof: Suppose not. Then as=0a^{\prime}_{s}=0 for all sTLks\geq T^{L}-k (recall aTLk=0a^{\prime}_{T^{L}-k}=0). We claim this implies #{n:an=0}>N\#\left\{n:a^{\prime}_{n}=0\right\}>N. To see this, note that #{n:an=0}N\#\left\{n:a^{\prime}_{n}=0\right\}\geq N by assumption. If #{n:an=0}=N,\#\left\{n:a^{\prime}_{n}=0\right\}=N, then 𝐚AN\mathbf{a}^{\prime}\in A_{N}, and since 𝐚\mathbf{a}^{\prime} contains k+1k+1 shirking periods in its tail, it follows that 𝐚AN,k+1,\mathbf{a}^{\prime}\in A_{N,k+1}, contradicting the assumption that n=k+1NAN,n=\bigcup_{n=k+1}^{N}A_{N,n}=\emptyset. Given that #{n:an=0}>N\#\left\{n:a^{\prime}_{n}=0\right\}>N and as=0a^{\prime}_{s}=0 for sTLk,s\geq T^{L}-k, it follows that βTLkH(𝐚)βTLkH(𝐚^)\beta_{T^{L}-k}^{H}(\mathbf{a}^{\prime})\geq\beta_{T^{L}-k}^{H}(\widehat{% \mathbf{a}}) and taking aTLk=1a^{\prime}_{T^{L}-k}=1 is optimal, a contradiction. \parallel

Now let ss^{\ast} be the first such working period after TLkT^{L}-k. Then,

period: t^\hat{t} t^+1\hat{t}+1 \cdots tt t+1t+1 \cdots TLk1T^{L}-k-1 TLkT^{L}-k TLk+1T^{L}-k+1 \cdots ss^{\ast} \cdots
𝐚^:\widehat{\mathbf{a}}\mathbf{:} 0 11 \cdots 1 1 \cdots 11 11 0 \cdots 0 \cdots
𝐚:\mathbf{a}^{\prime}\mathbf{:} 0 0 \cdots 0 0 0 \cdots 1

Applying parts (1) and (2) of Lemma 3 to 𝐚^\widehat{\mathbf{a}} and 𝐚\mathbf{a}^{\prime}, we obtain s=TLks1lsL=0.\sum\nolimits_{s=T^{L}-k}^{s^{\ast}-1}l^{L}_{s}=0. Now applying part (3) of Lemma 3, we obtain that the agent is indifferent between 𝐚\mathbf{a}^{\prime} and 𝐚′′\mathbf{a}^{\prime\prime} where 𝐚′′\mathbf{a}^{\prime\prime} differs from 𝐚\mathbf{a}^{\prime} only by switching the actions in period TLkT^{L}-k and period s.s^{\ast}. But since s=tTLk1lsL<0\sum\nolimits_{s=t}^{T^{L}-k-1}l^{L}_{s}<0, the optimality of at′′=0,aTLk′′=1a^{\prime\prime}_{t}=0,a^{\prime\prime}_{T^{L}-k}=1 contradicts part (2) of Lemma 3.

Step 2: ltL0.l^{L}_{t}\leq 0.

Proof of Step 2: Assume to the contrary that ltL>0.l^{L}_{t}>0. We have two cases to consider.

Case 1: lTLkL0.l^{L}_{T^{L}-k}\geq 0.

By the induction hypothesis and the assumption that ltL>0l^{L}_{t}>0, we have s=tTLklsL(1λH)n=t+1sa^n>0\sum\nolimits_{s=t}^{T^{L}-k}l^{L}_{s}\left(1-\lambda^{H}\right)^{\sum% \nolimits_{n=t+1}^{s}\widehat{a}_{n}}>0. Therefore, by part (1) of Lemma 3, a^TLk+1=1.\widehat{a}_{T^{L}-k+1}=1. But this contradicts the definition of 𝐚^\widehat{\mathbf{a}}.

Case 2: lTLkL<0.l^{L}_{T^{L}-k}<0.

In this case, (ICLa{}_{a}^{L}) must be slack in period TLkT^{L}-k (since it is satisfied in the next period and lTLkL<0l^{L}_{T^{L}-k}<0). Hence by Lemma 2, there exists 𝐚~\widetilde{\mathbf{a}} such that a~TLk=0.\widetilde{a}_{T^{L}-k}=0.

period: t^\hat{t} t^+1\hat{t}+1 \cdots tt t+1t+1 \cdots TLk1T^{L}-k-1 TLkT^{L}-k \cdots
𝐚^:\widehat{\mathbf{a}}\mathbf{:} 0 11 \cdots 1 1 \cdots 11 11 \cdots
𝐚~:\widetilde{\mathbf{a}}\mathbf{:} 0

Claim 2: a~s=0\widetilde{a}_{s}=0 for any s>TLk.s>T^{L}-k.

Proof: Suppose the claim is not true. Then define

τ:=min{s:s>TLk and a~s=1}.\tau:=\min\left\{s:s>T^{L}-k\text{ and }\widetilde{a}_{s}=1\right\}.

This is shown in the following table:

period: t^\hat{t} t^+1\hat{t}+1 \cdots tt t+1t+1 \cdots TLk1T^{L}-k-1 TLkT^{L}-k \cdots τ1\tau-1 τ\tau \cdots
𝐚^:\widehat{\mathbf{a}}\mathbf{:} 0 11 \cdots 1 1 \cdots 11 11 \cdots 0 0 \cdots
𝐚~:\widetilde{\mathbf{a}}\mathbf{:} 0 \cdots 0 1

Applying parts (1) and (2) of Lemma 3 to 𝐚^\widehat{\mathbf{a}}\mathbf{\ }and 𝐚~\widetilde{\mathbf{a}} respectively, we obtain

s=TLkτ1lsL=0.\sum\nolimits_{s=T^{L}-k}^{\tau-1}l^{L}_{s}=0. (C.14)

But then, by the induction hypothesis and the assumption that ltL>0,l^{L}_{t}>0, we obtain

s=tTLk1lsL(1λH)n=t+1sa^n>0.\sum\nolimits_{s=t}^{T^{L}-k-1}l^{L}_{s}\left(1-\lambda^{H}\right)^{\sum% \nolimits_{n=t+1}^{s}\widehat{a}_{n}}>0. (C.15)

Notice that a^s=0\widehat{a}_{s}=0 for s>TLks>T^{L}-k by definition. Hence, (C.14) and (C.15) imply s=tτ1lsL(1λH)n=t+1sa^n>0.\sum\nolimits_{s=t}^{\tau-1}l^{L}_{s}\left(1-\lambda^{H}\right)^{\sum\nolimits% _{n=t+1}^{s}\widehat{a}_{n}}>0. Now applying part (1) of Lemma 3 to 𝐚^,\widehat{\mathbf{a}}, we reach the conclusion that a^τ=1,\widehat{a}_{\tau}=1, a contradiction. \parallel

Hence, we have established the claim that a~s=0\widetilde{a}_{s}=0 for all s>TLk,s>T^{L}-k, as depicted below:

period: t^\hat{t} t^+1\hat{t}+1 \cdots tt t+1t+1 \cdots TLk1T^{L}-k-1 TLkT^{L}-k \cdots τ1\tau-1 τ\tau \cdots
𝐚^:\widehat{\mathbf{a}}\mathbf{:} 0 11 \cdots 1 1 \cdots 11 11 \cdots 0 0 \cdots
𝐚~:\widetilde{\mathbf{a}}\mathbf{:} 0 \cdots 0 0 \cdots

Claim 3: #{n:a~n=0}=N+1\#\left\{n:\widetilde{a}_{n}=0\right\}=N+1 and βTLkH(𝐚~)=βTLkH(𝐚^).\beta_{T^{L}-k}^{H}\left(\widetilde{\mathbf{a}}\right)=\beta_{T^{L}-k}^{H}% \left(\widehat{\mathbf{a}}\right).

Proof: By definition of N,N, #{n:a~n=0}N.\#\left\{n:\widetilde{a}_{n}=0\right\}\geq N. If #{n:a~n=0}=N,\#\left\{n:\widetilde{a}_{n}=0\right\}=N, then 𝐚~\widetilde{\mathbf{a}} contains k+1k+1 shirking periods in its tail, contradicting the assumption that AN,k+1=A_{N,k+1}=\emptyset. Moreover, if #{n:a~n=0}>N+1\#\left\{n:\widetilde{a}_{n}=0\right\}>N+1, then βTLkH(𝐚~)>βTLkH(𝐚^).\beta_{T^{L}-k}^{H}\left(\widetilde{\mathbf{a}}\right)>\beta_{T^{L}-k}^{H}% \left(\widehat{\mathbf{a}}\right). But then since a^TLk=1,\widehat{a}_{T^{L}-k}=1, we should have a~TLk=1,\widetilde{a}_{T^{L}-k}=1, a contradiction. Therefore, it must be #{n:a~n=0}=N+1.\#\left\{n:\widetilde{a}_{n}=0\right\}=N+1. \parallel

By Claim 3, we can choose 𝐚~\widetilde{\mathbf{a}} such that 𝐚~\widetilde{\mathbf{a}} differs from 𝐚^\widehat{\mathbf{a}} only in period TLk.T^{L}-k. This is shown in the following table:

period: t^\hat{t} t^+1\hat{t}+1 \cdots tt t+1t+1 \cdots TLk1T^{L}-k-1 TLkT^{L}-k \cdots τ1\tau-1 τ\tau \cdots
𝐚^:\widehat{\mathbf{a}}\mathbf{:} 0 11 \cdots 1 1 \cdots 11 11 \cdots 0 0 \cdots
𝐚~:\widetilde{\mathbf{a}}\mathbf{:} 0 1 \cdots 1 1 \cdots 11 0 \cdots 0 0 \cdots

But by assumption ltL>0,l^{L}_{t}>0, and by the induction hypothesis lsL=0l^{L}_{s}=0 for s=t+1,,TLk1.s=t+1,\ldots,T^{L}-k-1. Therefore,

s=tTLk1lsL(1λH)n=t+1sa~n>0.\sum\limits_{s=t}^{T^{L}-k-1}l^{L}_{s}\left(1-\lambda^{H}\right)^{\sum% \nolimits_{n=t+1}^{s}\widetilde{a}_{n}}>0.

Applying part (1) of Lemma 3, we must conclude that a~TLk=1,\widetilde{a}_{T^{L}-k}=1, a contradiction. ∎

Lemma 5.

Suppose AN,kA_{N,k}\neq\emptyset and m=k+1NAN,m=\bigcup_{m=k+1}^{N}A_{N,m}=\emptyset. Then lt^L=0.l^{L}_{\hat{t}}=0.

Proof.

Step 1: lt^L0.l^{L}_{\hat{t}}\geq 0.

Proof of Step 1: To the contrary, suppose lt^L<0.l^{L}_{\hat{t}}<0. Lemma 4 implies s=t^TLk1lsL(1λH)n=t^+1sa^n<0\sum\nolimits_{s=\hat{t}}^{T^{L}-k-1}l^{L}_{s}\left(1-\lambda^{H}\right)^{\sum% \nolimits_{n=\hat{t}+1}^{s}\widehat{a}_{n}}<0. Then part (2) of Lemma 3 implies that a^TLk=0,\widehat{a}_{T^{L}-k}=0, a contradiction.

Step 2: lt^L0.l^{L}_{\hat{t}}\leq 0.

Proof of Step 2: Suppose to the contrary that lt^L>0.l^{L}_{\hat{t}}>0. Note that by Lemma 1, there exists an action plan 𝐚𝜶H(𝐂L)\mathbf{a}^{\prime}\in\bm{\alpha}^{H}\left(\mathbf{C}^{L}\right) such that at^=1a_{\hat{t}}^{\prime}=1. Then since, by Lemma 4, ltL=0l^{L}_{t}=0 for t=t^+1,,TLk1t=\hat{t}+1,\ldots,T^{L}-k-1, it follows from part (1) of Lemma 3 that as=1a_{s}^{\prime}=1 for s=t^+1,,TLks=\hat{t}+1,\ldots,T^{L}-k. Hence, we obtain the following table:

period: t^\hat{t} t^+1\hat{t}+1 \cdots TLk1T^{L}-k-1 TLkT^{L}-k TLk+1T^{L}-k+1 \cdots
𝐚^:\widehat{\mathbf{a}}\mathbf{:} 0 11 \cdots 11 11 0 \cdots
𝐚:\mathbf{a}^{\prime}\mathbf{:} 1 1 \cdots 1 1

Claim: there exists t~<t^,\widetilde{t}<\hat{t}, such that at~=0a_{\widetilde{t}}^{\prime}=0 and a^t~=1.\widehat{a}_{\widetilde{t}}=1.

Proof: since #{t:at=0}N=#{t:a^t=0}\#\left\{t:a_{t}^{\prime}=0\right\}\geq N=\#\left\{t:\widehat{a}_{t}=0\right\}, at^=1a_{\hat{t}}^{\prime}=1, a^t^=0\widehat{a}_{\hat{t}}=0, and a^t=0\widehat{a}_{t}=0 for all t>TLk,t>T^{L}-k, we have #{t:at=0,t<t^}>#{t:a^t=0,t<t^}.\#\left\{t:a_{t}^{\prime}=0,t<\hat{t}\right\}>\#\left\{t:\widehat{a}_{t}=0,t<% \hat{t}\right\}. The claim follows immediately.

We can take t~\widetilde{t} to be the largest period that satisfies the above claim. Hence 𝐚^\widehat{\mathbf{a}} and 𝐚\mathbf{a}^{\prime} are as follows:

period: t~\widetilde{t} \cdots t^\hat{t} t^+1\hat{t}+1 \cdots TLk1T^{L}-k-1 TLkT^{L}-k TLk+1T^{L}-k+1 \cdots
𝐚^:\widehat{\mathbf{a}}\mathbf{:} 1 0 11 \cdots 11 11 0 \cdots
𝐚:\mathbf{a}^{\prime}\mathbf{:} 0 1 1 \cdots 1 1

There are two cases to consider.

Case 1: a^t=at\widehat{a}_{t}=a_{t}^{\prime} for each t=t~+1,,t^1.t=\widetilde{t}+1,\ldots,\hat{t}-1.

Lemma 3 implies that s=t~t^1lsL(1λH)n=t~+1sa^n=0\sum\nolimits_{s=\widetilde{t}}^{\hat{t}-1}l^{L}_{s}\left(1-\lambda^{H}\right)% ^{\sum\nolimits_{n=\widetilde{t}+1}^{s}\widehat{a}_{n}}=0 and the agent is indifferent between 𝐚^\widehat{\mathbf{a}} and 𝐚^\widehat{\mathbf{a}}^{\prime} where 𝐚^\widehat{\mathbf{a}}^{\prime} differs from 𝐚^\widehat{\mathbf{a}} only in that the actions at periods t~\widetilde{t} and t^\hat{t} are switched. But this contradicts the definition of t^\hat{t} (see (C.12)).

Case 2: a^m=0\widehat{a}_{m}=0 and am=1a_{m}^{\prime}=1 for some m{t~+1,,t^1}.m\in\left\{\widetilde{t}+1,\ldots,\hat{t}-1\right\}.

First note that Case 1 and Case 2 are exhaustive because t~\widetilde{t} is taken to be the largest period t<t^t<\hat{t} such that a^t=1\widehat{a}_{t}=1 and at=0.a_{t}^{\prime}=0. Without loss, we take mm to be the smallest possible. Hence a^t=at\widehat{a}_{t}=a_{t}^{\prime} for each t=t~+1,,m1.t=\widetilde{t}+1,\ldots,m-1. Then 𝐚^\widehat{\mathbf{a}} and 𝐚\mathbf{a}^{\prime} are as follows:

period: t~\widetilde{t} \cdots mm \cdots t^\hat{t} t^+1\hat{t}+1 \cdots TLk1T^{L}-k-1 TLkT^{L}-k TLk+1T^{L}-k+1 \cdots
𝐚^:\widehat{\mathbf{a}}\mathbf{:} 1 0 0 11 \cdots 11 11 0 \cdots
𝐚:\mathbf{a}^{\prime}\mathbf{:} 0 11 1 1 \cdots 1 1

But again, by Lemma 3, we can switch the actions at periods t~\widetilde{t} and mm in 𝐚^,\widehat{\mathbf{a}}, contradicting the definition of 𝐚^\widehat{\mathbf{a}} (see (C.13)). ∎

Lemma 6.

If AN,kA_{N,k}\neq\emptyset then n=k+1NAN,n.\bigcup_{n=k+1}^{N}A_{N,n}\neq\emptyset.

Proof.

Suppose to the contrary that n=k+1NAN,n=\bigcup_{n=k+1}^{N}A_{N,n}=\emptyset. Then ltL=0l^{L}_{t}=0 for t=t^,,TLk1,t=\hat{t},\ldots,T^{L}-k-1, by Lemma 4 and Lemma 5. Therefore, by part (3) of Lemma 3, we can switch a^t^\widehat{a}_{\hat{t}} with a^TLk\widehat{a}_{T^{L}-k} to obtain 𝐚^\widehat{\mathbf{a}}^{\prime}. However, since #{t:a^t=0}=#{t:a^t=0}=N,\#\left\{t:\widehat{a}_{t}^{\prime}=0\right\}=\#\left\{t:\widehat{a}_{t}=0% \right\}=N, it follows immediately that 𝐚^AN.\widehat{\mathbf{a}}^{\prime}\in A_{N}. Since a^t=0\widehat{a}_{t}^{\prime}=0 for all t>TLk1,t>T^{L}-k-1, 𝐚^n=k+1NAN,n.\widehat{\mathbf{a}}^{\prime}\in\bigcup_{n=k+1}^{N}A_{N,n}.

C.5 Step 5: Onetime-penalty contracts for the low type

In Step 4, we showed that we can restrict attention in solving program [P2] to connected penalty contracts for the low type 𝐂L=(TL,W0L,𝒍L)\mathbf{C}^{L}=(T^{L},W_{0}^{L},\bm{l}^{L}) such that there is an optimal action plan for the high type 𝐚𝜶H(𝐂L)\mathbf{a}\in\bm{\alpha}^{H}(\mathbf{C}^{L}) that is a stopping strategy with the most work property. We now use this result to show that we can further restrict attention to onetime-penalty contracts for the low type, 𝐂L=(TL,W0L,lTLL)\mathbf{C}^{L}=(T^{L},W_{0}^{L},l_{T^{L}}^{L}).

This result is proved via two lemmas.

Lemma 7.

Let 𝐂L=(TL,W0L,𝐥L)\mathbf{C}^{L}=(T^{L},W_{0}^{L},\bm{l}^{L}) be an optimal contract for the low type with a most-work optimal stopping strategy for the high type 𝐚^\widehat{\mathbf{a}} that stops at t^\hat{t}, i.e. t^=max{t{1,,TL}:a^t=1}\hat{t}=\max\{t\in\{1,\ldots,T^{L}\}:\widehat{a}_{t}=1\}. For each t>t^t>\hat{t}, there is an optimal action plan, 𝐚~𝛂H(𝐂L)\widetilde{\mathbf{a}}\in\bm{\alpha}^{H}(\mathbf{C}^{L}), such that for any ss, a^s=a~ss{t^,t}\widehat{a}_{s}=\widetilde{a}_{s}\iff s\notin\{\hat{t},t\}.

Proof.

Step 1: First, we show that the Lemma’s claim is true for some t>t^t>\hat{t} (rather than for all t>t^t>\hat{t}). Suppose not, to contradiction. Then Lemma 3 implies that

for any n{t^,t^+1,,TL1}n\in\{\hat{t},\hat{t}+1,\ldots,T^{L}-1\}, s=t^nlsL<0\sum_{s=\hat{t}}^{n}l^{L}_{s}<0. (C.16)

Hence, (ICLa{}_{a}^{L}) is slack at t^\hat{t} (since it is satisfied in the next period and lt^L<0l^{L}_{\hat{t}}<0) and, by Lemma 2, there exists an optimal action plan, 𝐚′′\mathbf{a{{}^{\prime\prime}}}, with at^′′=0a^{\prime\prime}_{\hat{t}}=0.

Claim 1: as′′=0a^{\prime\prime}_{s}=0 for all s>t^s>\hat{t}.

Proof: Suppose to contradiction that there exists τ>t^\tau>\hat{t} such that aτ′′=1a^{\prime\prime}_{\tau}=1. Take the smallest such τ\tau. Then it follows from Lemma 3 applied to 𝐚^\widehat{\mathbf{a}} and 𝐚′′\mathbf{a{{}^{\prime\prime}}} that s=t^τ1lsL=0\sum\limits_{s=\hat{t}}^{\tau-1}l^{L}_{s}=0, contradicting (C.16). \parallel

Hence, we obtain that as′′=0a^{\prime\prime}_{s}=0 for all st^s\geq\hat{t}, and it follows from the optimality of a^t^=1\widehat{a}_{\hat{t}}=1 and at^′′=0a^{\prime\prime}_{\hat{t}}=0 that 𝐚′′\mathbf{a{{}^{\prime\prime}}} is a stopping strategy that stops at t^1{\hat{t}}-1:

period: \cdots t^2\hat{t}-2 t^1\hat{t}-1 t^\hat{t} t^+1\hat{t}+1 t^+2\hat{t}+2 \cdots
𝐚^:\widehat{\mathbf{a}}\mathbf{:} \cdots 1 1 1 0 0 \cdots
𝐚′′:\mathbf{a^{\prime\prime}}\mathbf{:} \cdots 1 1 0 0 0 \cdots

Next, note that by Lemma 1, there is an optimal action plan, 𝐚\mathbf{a^{\prime}}, with aTL=1a^{\prime}_{T^{L}}=1.

Claim 2: at^=1a^{\prime}_{\hat{t}}=1.

Proof: Suppose at^=0a^{\prime}_{\hat{t}}=0. Then by (C.16) and Lemma 3, at^+1=0a^{\prime}_{\hat{t}+1}=0. But then again by (C.16) and Lemma 3, at^+2=0a^{\prime}_{\hat{t}+2}=0, and using induction we arrive at the conclusion that aTL=0a^{\prime}_{T^{L}}=0. Contradiction. \parallel

Since aTL=1a^{\prime}_{T^{L}}=1 and at^=1a^{\prime}_{\hat{t}}=1, by the most work property of 𝐚^\widehat{\mathbf{a}}, there must exist a period m<t^m<\hat{t} such that am=0a^{\prime}_{m}=0. Take the largest such period:

period: \cdots mm m+1m+1 \cdots t^1\hat{t}-1 t^\hat{t} t^+1\hat{t}+1 t^+2\hat{t}+2 \cdots TLT^{L}
𝐚^:\widehat{\mathbf{a}}\mathbf{:} \cdots 1 1 \cdots 1 1 0 0 \cdots 0
𝐚′′:\mathbf{a^{\prime\prime}}\mathbf{:} \cdots 1 1 \cdots 1 0 0 0 \cdots 0
𝐚:\mathbf{a^{\prime}}\mathbf{:} 0 1 \cdots 1 1 1

Applying Lemma 3 to 𝐚′′\mathbf{a^{\prime\prime}} and 𝐚\mathbf{a^{\prime}} yields s=mt^1lsL(1λH)n=m+1san=0\sum\limits_{s=m}^{\hat{t}-1}l^{L}_{s}\left(1-\lambda^{H}\right)^{\sum% \nolimits_{n=m+1}^{s}a^{\prime}_{n}}=0. Hence, there exists an optimal action plan 𝐚′′′\mathbf{a^{\prime\prime\prime}} obtained from 𝐚\mathbf{a^{\prime}} by switching ama^{\prime}_{m} and at^a^{\prime}_{\hat{t}}. But then the optimality of 𝐚′′′\mathbf{a^{\prime\prime\prime}} contradicts at^′′′=0a^{\prime\prime\prime}_{\hat{t}}=0, aTL′′′=1a^{\prime\prime\prime}_{T^{L}}=1, (C.16), and Lemma 3.

Step 2: We now prove the Lemma’s claim for t^+1\hat{t}+1. That is, we show that there exists an optimal action plan, call it 𝐚^t^+1\widehat{\mathbf{a}}^{\hat{t}+1}, such that for any ss, a^st^+1=a^ss{t^,t^+1}\widehat{a}^{\hat{t}+1}_{s}=\widehat{a}_{s}\iff s\notin\{\hat{t},\hat{t}+1\}. Suppose, to contradiction, that the claim is false. Then, by Lemma 3, lt^L<0l^{L}_{\hat{t}}<0. Using Step 1, there is some τ>t^\tau>\hat{t} that satisfies the Lemma’s claim; let 𝐚τ\mathbf{a}^{\tau} be the corresponding optimal action plan (which is identical to 𝐚^\widehat{\mathbf{a}} in exactly all periods except from t^\hat{t} and τ\tau). Since by Lemma 1 there exists an optimal action plan, call it 𝐚\mathbf{a^{\prime}}, with at^+1=1a^{\prime}_{\hat{t}+1}=1, Lemma 3 and lt^L<0l^{L}_{\hat{t}}<0 imply at^=1a^{\prime}_{\hat{t}}=1. By the most work property of 𝐚^\widehat{\mathbf{a}}, there must exist a period m<t^m<\hat{t} such that am=0a^{\prime}_{m}=0. Take the largest such period:

period: \cdots mm m+1m+1 \cdots t^1\hat{t}-1 t^\hat{t} t^+1\hat{t}+1 \cdots τ\tau τ+1\tau+1 \cdots
𝐚^:\widehat{\mathbf{a}}{:} \cdots 1 1 \cdots 1 1 0 \cdots 0 0 \cdots
𝐚τ:\mathbf{a}^{\tau}{:} \cdots 1 1 \cdots 1 0 0 \cdots 1 0 \cdots
𝐚:\mathbf{a^{\prime}}{:} 0 1 \cdots 1 1 1

Applying Lemma 3 to 𝐚τ\mathbf{a}^{\tau} and 𝐚\mathbf{a^{\prime}} yields s=mt^1lsL(1λH)n=m+1san=0\sum\nolimits_{s=m}^{\hat{t}-1}l^{L}_{s}\left(1-\lambda^{H}\right)^{\sum% \nolimits_{n=m+1}^{s}a^{\prime}_{n}}=0. Hence, there exists an optimal action plan 𝐚′′\mathbf{a^{\prime\prime}} obtained from 𝐚\mathbf{a^{\prime}} by switching ama^{\prime}_{m} and at^a^{\prime}_{\hat{t}}. But then the optimality of 𝐚′′\mathbf{a^{\prime\prime}} contradicts at^′′=0a^{\prime\prime}_{\hat{t}}=0, at^+1′′=1a^{\prime\prime}_{\hat{t}+1}=1, lt^L<0l^{L}_{\hat{t}}<0, and Lemma 3.

Step 3: Finally, we use induction to prove that the Lemma’s claim is true for any s>t^+1s>\hat{t}+1. (Note the claim is true for t^+1\hat{t}+1 by Step 2.) Take any t+1{t^+2,,TL}t+1\in\{\hat{t}+2,\ldots,T^{L}\}. Assume the claim is true for s=t^+2,,ts=\hat{t}+2,\ldots,t. We show that the claim is true for t+1t+1.

By Step 2 and the induction hypothesis, there exists an optimal action plan, 𝐚^t\widehat{\mathbf{a}}^{t}, such that for any ss, a^st=a^ss{t^,t}\widehat{a}^{t}_{s}=\widehat{a}_{s}\iff s\notin\{\hat{t},t\}. We shall show that there exists an optimal action plan, 𝐚^t+1\widehat{\mathbf{a}}^{t+1}, such that for any ss, a^st+1=a^ss{t^,t+1}\widehat{a}^{t+1}_{s}=\widehat{a}_{s}\iff s\notin\{\hat{t},t+1\}. Suppose, to contradiction, that the claim is false. Note that Step 2, the induction hypothesis, and Lemma 3 imply lsL=0l^{L}_{s}=0 for all s=t^,,t1s=\hat{t},\ldots,t-1. It thus follows from Lemma 3 and the claim being false that ltL<0l^{L}_{t}<0. By Lemma 1 there exists an optimal action plan, call it 𝐚\mathbf{a^{\prime}}, with at+1=1a^{\prime}_{t+1}=1. Then Lemma 3 and ltL<0l^{L}_{t}<0 imply that at=1a^{\prime}_{t}=1.

Claim 3: as=1a^{\prime}_{s}=1 for all s=t^,,t1s=\hat{t},\ldots,t-1.

Proof: Suppose to contradiction that as=0a^{\prime}_{s^{*}}=0 for some s{t^,,t1}s^{*}\in\{\hat{t},\ldots,t-1\}. Then since lsL=0l^{L}_{s}=0 for all s=t^,,t1s=\hat{t},\ldots,t-1, by Lemma 3, there exists an optimal action plan, 𝐚′′\mathbf{a}^{\prime\prime}, obtained from 𝐚\mathbf{a}^{\prime} by switching asa^{\prime}_{s^{*}} and ata^{\prime}_{t}. But then the optimality of 𝐚′′\mathbf{a^{\prime\prime}} contradicts at′′=0a^{\prime\prime}_{t}=0, at+1′′=1a^{\prime\prime}_{t+1}=1, ltL<0l^{L}_{t}<0, and Lemma 3. \parallel

Hence, we obtain as=1a^{\prime}_{s}=1 for all s=t^,,t+1s=\hat{t},\ldots,t+1, and by the most work property of 𝐚^\widehat{\mathbf{a}}, there must exist a period m<t^m<\hat{t} such that am=0a^{\prime}_{m}=0. Take the largest such period:

period: \cdots mm m+1m+1 \cdots t^1\hat{t}-1 t^\hat{t} t^+1\hat{t}+1 \cdots tt t+1t+1 \cdots
𝐚^:\widehat{\mathbf{a}}: \cdots 1 1 \cdots 1 1 0 \cdots 0 0 \cdots
𝐚^t:\widehat{\mathbf{a}}^{t}: \cdots 1 1 \cdots 1 0 0 \cdots 1 0 \cdots
𝐚\mathbf{a}^{\prime}: \cdots 0 1 \cdots 1 1 1 \cdots 1 1

Applying Lemma 3 to 𝐚^t\widehat{\mathbf{a}}^{t} and 𝐚\mathbf{a^{\prime}} yields s=mt^1lsL(1λH)n=m+1san=0\sum\limits_{s=m}^{\hat{t}-1}l^{L}_{s}\left(1-\lambda^{H}\right)^{\sum% \nolimits_{n=m+1}^{s}a^{\prime}_{n}}=0. Since lsL=0l^{L}_{s}=0 for all s=t^,,t1s=\hat{t},\ldots,t-1, we obtain s=mt1lsL(1λH)n=m+1san=0\sum\limits_{s=m}^{t-1}l^{L}_{s}\left(1-\lambda^{H}\right)^{\sum\nolimits_{n=m% +1}^{s}a^{\prime}_{n}}=0. Hence, by Lemma 3, there exists an optimal action plan 𝐚′′\mathbf{a^{\prime\prime}} obtained from 𝐚\mathbf{a^{\prime}} by switching ama^{\prime}_{m} and ata^{\prime}_{t}. But then the optimality of 𝐚′′\mathbf{a^{\prime\prime}} contradicts at′′=0a^{\prime\prime}_{t}=0, at+1′′=1a^{\prime\prime}_{t+1}=1, ltL<0l^{L}_{t}<0, and Lemma 3. ∎

Lemma 8.

If 𝐂L\mathbf{C}^{L} is an optimal contract for the low type with a most-work optimal stopping strategy for the high type, then 𝐂L\mathbf{C}^{L} is a onetime-penalty contract.

Proof.

Fix 𝐂L\mathbf{C}^{L} per the Lemma’s assumptions. Let 𝐚^\widehat{\mathbf{a}} and t^\hat{t} be as defined in the statement of Lemma 7. Then, it immediately follows from Lemma 7 and Lemma 3 that ltL=0l^{L}_{t}=0 for all t{t^,t^+1,,TL1}t\in\{\hat{t},\hat{t}+1,\ldots,T^{L}-1\}. We use induction to prove that ltL=0l^{L}_{t}=0 for all t<t^t<\hat{t}.

Assume ltL=0l^{L}_{t}=0 for all t{m+1,m+2,,TL1}t\in\{m+1,m+2,\ldots,T^{L}-1\} for m<t^m<\hat{t}. We will show that lmL=0l^{L}_{m}=0. First, lmL>0l^{L}_{m}>0 is not possible because then s=mt^lsL(1λH)n=m+1sa^n>0\sum\limits_{s=m}^{\hat{t}}l^{L}_{s}\left(1-\lambda^{H}\right)^{\sum\nolimits_% {n=m+1}^{s}\widehat{a}_{n}}>0 (by Lemma 7 and the inductive assumption), contradicting the optimality of 𝐚^\widehat{\mathbf{a}} and Lemma 3. Second, we claim lmL<0l^{L}_{m}<0 is not possible. Suppose, to contradiction, that lmL<0l^{L}_{m}<0. Then (ICLa{}_{a}^{L}) is slack at mm and, by Lemma 2, there exists an optimal plan 𝐚\mathbf{a^{\prime}} with am=0a^{\prime}_{m}=0. Now by Lemma 3, Lemma 7, and the inductive assumption, as=0a^{\prime}_{s}=0 for all sms\geq m. Hence, βt^H(𝐚)>βt^H(𝐚^)\beta^{H}_{\hat{t}}(\mathbf{a^{\prime}})>\beta^{H}_{\hat{t}}(\widehat{\mathbf{% a}}), and thus the optimality of 𝐚^\widehat{\mathbf{a}} implies that 𝐚\mathbf{a^{\prime}} is suboptimal at t^\hat{t}, a contradiction. ∎

References

  • D. Atkin, A. Chaudhry, S. Chaudry, A. K. Khandelwal, and E. Verhoogen (2015) Organizational barriers to technology adoption: evidence from soccer-ball producers in pakistan. Note: unpublished Cited by: §5.3.
  • D. P. Baron and D. Besanko (1984) Regulation and information in a continuing relationship. Information Economics and Policy 1 (3), pp. 267–302. Cited by: footnote 10.
  • C. Barrett, M. Bachke, M. Bellemare, H. Michelson, S. Narayanan, and T. Walker (2012) Smallholder participation in contract farming: comparative evidence from five countries. World Development 40 (4), pp. 715–730. Cited by: §1.
  • M. Battaglini (2005) Long-term contracting with markovian consumers. American Economic Review 95 (3), pp. 637–658. Cited by: footnote 10.
  • L. Beaman, D. Karlan, B. Thuysbaert, and C. Udry (2015) Selection into credit markets: evidence from agriculture in mali. Note: unpublished Cited by: footnote 6.
  • D. Bergemann and U. Hege (1998) Venture capital financing, moral hazard, and learning. Journal of Banking & Finance 22 (6-8), pp. 703–735. Cited by: §1.
  • D. Bergemann and U. Hege (2005) The financing of innovation: learning and stopping. RAND Journal of Economics 36 (4), pp. 719–752. Cited by: §1.
  • D. Bergemann and J. Välimäki (2008) Bandit problems. In The New Palgrave Dictionary of Economics, S. N. Durlauf and L. E. Blume (Eds.), Cited by: footnote 2.
  • T. Besley and A. Case (1993) Modeling technology adoption in developing countries. American Economic Review, Papers and Proceedings 83 (2), pp. 396–402. Cited by: footnote 5.
  • V. Bhaskar (2012) Dynamic moral hazard, learning and belief manipulation. Note: unpublished Cited by: footnote 23.
  • V. Bhaskar (2014) The ratchet effect re-examined: a learning perspective. Note: unpublished Cited by: footnote 23.
  • B. Biais, T. Mariotti, J. Rochet, and S. Villeneuve (2010) Large risks, limited liability, and dynamic moral hazard. Econometrica 78, pp. 73–118. Cited by: footnote 45.
  • C. Bobtcheff and R. Levy (2015) More haste, less speed: signaling through investment timing. Note: unpublished Cited by: §1.
  • R. Boleslavsky and M. Said (2013) Progressive Screening: Long-Term Contracting with a Privately Known Stochastic Process. Review of Economic Studies 80 (1), pp. 1–34. Cited by: footnote 10.
  • A. Bonatti and J. Hörner (2011) Collaborating. American Economic Review 101 (2), pp. 632–663. Cited by: footnote 12.
  • A. Bonatti and J. Hörner (2015) Career concerns with exponential learning. Note: forthcoming in Theoretical Economics Cited by: footnote 12.
  • B. Bunnin (1983) Author law & strategies. Berkeley, CA: Nolo, 1983. Cited by: §5.3.
  • S. Chassang (2013) Calibrated incentive contracts. Econometrica 81 (5), pp. 1935–1971. External Links: Document, ISSN 1468-0262, Link Cited by: footnote 11.
  • T. G. Conley and C. R. Udry (2010) Learning about a new technology: pineapple in ghana. American Economic Review 100, pp. 35–69. Cited by: §5.3, footnote 7.
  • J. Cremer and R. P. McLean (1985) Optimal selling strategies under uncertainty for a discriminating monopolist when demands are interdependent. Econometrica 53 (2), pp. 345–361. Cited by: footnote 26.
  • J. Cremer and R. P. McLean (1988) Full extraction of the surplus in bayesian and dominant strategy auctions. Econometrica 56 (6), pp. 1247–1257. Cited by: footnote 26.
  • P. Demarzo and Y. Sannikov (2011) Learning in dynamic incentive contracts. Note: unpublished Cited by: footnote 11.
  • F. Ederer (2013) Incentives for parallel innovation. Note: unpublished Cited by: footnote 12.
  • P. Eső and B. Szentes (2015) Dynamic contracting: an irrelevance theorem. Note: forthcoming in Theoretical Economics Cited by: footnote 10.
  • G. Feder, R. Just, and D. Zilberman (1985) Adoption of agricultural innovations in developing countries: a survey. Economic Development and Cultural Change 33 (2), pp. 255–298. Cited by: §5.3.
  • A. D. Foster and M. R. Rosenzweig (2010) Microeconomics of technology adoption. Annual Review of Economics 2. Cited by: §5.3.
  • M. Fowler (1985) The ‘satisfactory manuscript’ clause in book publishing contracts. Columbia VLA Journal of Law & the Arts 10, pp. 119–152. Cited by: §5.3.
  • M. Frick and Y. Ishii (2015) Innovation adoption by forward-looking social learners. Note: unpublished Cited by: footnote 38.
  • D. Gerardi and L. Maestri (2012) A principal-agent model of sequential testing. Theoretical Economics 7 (3), pp. 425–463. Cited by: §1, footnote 48.
  • A. Gershkov and M. Perry (2012) Dynamic contracts with moral hazard and adverse selection. Review of Economic Studies 79 (1), pp. 268–306. Cited by: §1.
  • R. Gomes, D. Gottlieb, and L. Maestri (2015) Experimentation and project selection: screening and learning. Note: unpublished Cited by: §1, footnote 21, footnote 48.
  • Y. Guo (2014) Dynamic delegation of experimentation. Note: unpublished Cited by: footnote 12.
  • M. Halac, N. Kartik, and Q. Liu (2015) Contests for experimentation. Note: unpublished Cited by: footnote 12.
  • Z. He, B. Wei, J. Yu, and F. Gao (2014) Optimal long-term contracting with learning. Note: unpublished Cited by: footnote 11.
  • B. Holmström and P. R. Milgrom (1987) Aggregation and linearity in the provision of intertemporal incentives. Econometrica, pp. 303–328. Cited by: footnote 11.
  • J. Hörner and L. Samuelson (2013) Incentives for experimenting agents. RAND Journal of Economics 44 (4), pp. 632–663. External Links: Document, ISSN 1756-2171, Link Cited by: footnote 12, footnote 23.
  • B. K. Jack, P. Oliva, C. Severen, E. Walker, and S. Bell (2014) Technology adoption under uncertainty: take up and subsequent investment in zambia. Note: unpublished Cited by: §5.3, §5.3.
  • G. Keller, S. Rady, and M. Cripps (2005) Strategic experimentation with exponential bandits. Econometrica 73 (1), pp. 39–68. Cited by: §1.
  • J. Kelsey (2013) Private information and the allocation of land use subsidies in malawi. American Economic Journal: Applied Economics 5 (3), pp. 113–135. Cited by: footnote 35.
  • N. Klein (2012) The importance of being honest. Note: unpublished Cited by: footnote 12.
  • Kuzyk,Raya (2006) Inside publishing. Poets & Writers 34.1 (Jan/Feb). Cited by: §5.3.
  • S. Kwon (2013) Dynamic moral hazard with persistent states. Note: unpublished Cited by: footnote 12, footnote 23.
  • J. Laffont and D. Martimort (2001) The theory of incentives: the principal-agent model. Princeton University Press. Cited by: footnote 24.
  • J. Laffont and J. Tirole (1988) The dynamics of incentive contracts. Econometrica 56 (5), pp. 1153–1175. Cited by: footnote 10.
  • T. R. Lewis and M. Ottaviani (2008) Search agency. Note: unpublished Cited by: §1.
  • T. R. Lewis (2011) A theory of delegated search for the best alternative. Note: unpublished Cited by: §1.
  • G. Manso (2011) Motivating innovation. Journal of Finance 66 (5), pp. 1823–1860. Cited by: footnote 12.
  • R. Mason and J. Välimäki (2011) Dynamic moral hazard and stopping. Note: unpublished Cited by: footnote 23.
  • N. Minot (2007) Contract farming in developing countries: patterns, impact, and policy implications. Case Study 6-3 of the Program, Food Policy for Developing Countries: the Role of Government in the Global Food System. Cited by: §1.
  • S. Miyata, N. Minot, and D. Hu (2009) Impact of contract farming on income: linking small farmers, packers, and supermarkets in china. World Development 37 (11), pp. 1781–1790. Cited by: §1, §5.3, footnote 37.
  • S. Moroni (2015) Experimentation in organizations. Note: unpublished Cited by: footnote 12, footnote 38.
  • I. Obara (2008) The full surplus extraction theorem with hidden actions. The B.E. Journal of Theoretical Economics 8 (1). Cited by: footnote 26.
  • L. Owen (2013) Clark’s publishing agreements: a book of precedents: ninth edition. Bloomsbury Professional. Cited by: §1.
  • A. Pavan, I. Segal, and J. Toikka (2014) Dynamic mechanism design: a myersonian approach. Econometrica 82 (2), pp. 601–653. External Links: Document, ISSN 1468-0262, Link Cited by: footnote 10.
  • J. Prat and B. Jovanovic (2014) Dynamic contracts when agent’s quality is unknown. Theoretical Economics 9 (3), pp. 865–914. Cited by: footnote 11.
  • M. H. Riordan and D. E. M. Sappington (1988) Optimal contracts with public ex post information. Journal of Economic Theory 45 (1), pp. 189–199. Cited by: footnote 26.
  • M. Rothschild (1974) A two-armed bandit theory of market pricing. Journal of Economic Theory 9 (2), pp. 185–202. Cited by: footnote 2.
  • Y. Sannikov (2007) Agency problems, screening and increasing credit lines. Note: unpublished Cited by: §1.
  • Y. Sannikov (2013) Moral hazard and long-run incentives. Note: unpublished Cited by: footnote 11.
  • C. Suddath (2012) Penguin group sues writers over book advances. Bloomberg Businessweek, pp. September 27. Cited by: footnote 9.
  • T. Suri (2011) Selection and comparative advantage in technology adoption. Econometrica 79 (1), pp. 159–209. Cited by: §5.3.

Appendix D Supplementary Appendix for Online Publication Only

D.1 Proof of Proposition 1

We prove the result more generally for contracts with lockouts. Fix a contract 𝐂=(Γ,W0,𝒃,𝒍)\mathbf{C}=(\Gamma,W_{0},\bm{b},\bm{l}). The result is trivial if Γ=\Gamma=\emptyset, so assume Γ\Gamma\neq\emptyset. Let T=maxΓT=\max\Gamma. For any period tΓt\in\Gamma with t<Tt<T, define the smallest successor period in Γ\Gamma as σ(t)=min{t:t>t,tΓ}\sigma(t)=\min\{t^{\prime}:t^{\prime}>t,t^{\prime}\in\Gamma\}; moreover, let σ(0)=minΓ\sigma(0)=\min\Gamma.

Given any action profile for the agent, the agent’s time-zero expected discounted payoff when his type is θ{L,H}\theta\in\{L,H\} and the principal’s time-zero expected discounted payoff only depend upon a contract’s induced vector of discounted transfers, say (τt)tΓ(\tau_{t})_{t\in\Gamma} when success is obtained in period tt and on the discounted transfer when there is no success. Hence, it suffices to construct a penalty contract, ^𝐂\widehat{}\mathbf{C}, and bonus contract, ~𝐂\widetilde{}\mathbf{C}, that induce the same such vector of transfers as 𝐂\mathbf{C}.

To this end, define the penalty contract 𝐂^=(Γ,W^0,𝒍^)\widehat{\mathbf{C}}=(\Gamma,\widehat{W}_{0},\widehat{\bm{l}}) as follows:

  1. (a)

    For any tt such that t<Tt<T and tΓt\in\Gamma, l^t=ltbt+δσ(t)tbσ(t)\widehat{l}_{t}=l_{t}-b_{t}+\delta^{\sigma(t)-t}b_{\sigma(t)}.

  2. (b)

    l^T=lTbT\widehat{l}_{T}=l_{T}-b_{T}.

  3. (c)

    W^0=W0+δσ(0)bσ(0)\widehat{W}_{0}=W_{0}+\delta^{\sigma(0)}b_{\sigma(0)}.

Define the bonus contract ~𝐂=(Γ,W~0,~𝒃)\widetilde{}\mathbf{C}=(\Gamma,\widetilde{W}_{0},\widetilde{}\bm{b}) as follows:

  1. (a)

    For any tΓt\in\Gamma, b~t=btst,sΓδstls\widetilde{b}_{t}=b_{t}-\sum\limits_{s\geq t,s\in\Gamma}\delta^{s-t}l_{s}.

  2. (b)

    W~0=W0+tΓδtlt\widetilde{W}_{0}=W_{0}+\sum\limits_{t\in\Gamma}\delta^{t}l_{t}.

Consider first the discounted transfer induced by each of these three contracts if success is not obtained. For 𝐂\mathbf{C}, it is W0+tΓδtltW_{0}+\sum_{t\in\Gamma}\delta^{t}l_{t}. For ^𝐂\widehat{}\mathbf{C}, it is

W^0+tΓδtl^t=W0+δσ(0)bσ(0)+tΓ,t<Tδt(ltbt+δσ(t)tbσ(t))+δT(lTbT)=W0+tΓδtlt,\displaystyle\widehat{W}_{0}+\sum_{t\in\Gamma}\delta^{t}\widehat{l}_{t}=W_{0}+% \delta^{\sigma(0)}b_{\sigma(0)}+\sum_{t\in\Gamma,t<T}\delta^{t}\left(l_{t}-b_{% t}+\delta^{\sigma(t)-t}b_{\sigma(t)}\right)+\delta^{T}\left(l_{T}-b_{T}\right)% =W_{0}+\sum_{t\in\Gamma}\delta^{t}l_{t},

where the first equality follows from the definition of ^𝐂\widehat{}\mathbf{C} and the second from algebraic simplification. For ~𝐂\widetilde{}\mathbf{C}, since there are no penalties, the corresponding discounted transfer is just W~0=W0+tΓδtlt\widetilde{W}_{0}=W_{0}+\sum_{t\in\Gamma}\delta^{t}l_{t}. Hence, all three contracts induce the same transfer in the event of no success.

Next, for any sΓs\in\Gamma, consider a success obtained in period ss. The discounted transfer in this event in 𝐂\mathbf{C} is W0+tΓ,t<sδtlt+δsbsW_{0}+\sum_{t\in\Gamma,t<s}\delta^{t}l_{t}+\delta^{s}b_{s}. For ^𝐂\widehat{}\mathbf{C}, since there are no bonuses, it is

W^0+tΓ,t<sδtl^t=W0+δσ(0)bσ(0)+tΓ,t<sδt(ltbt+δσ(t)tbσ(t))=W0+tΓ,t<sδtlt+δsbs,\displaystyle\widehat{W}_{0}+\sum_{t\in\Gamma,t<s}\delta^{t}\widehat{l}_{t}=W_% {0}+\delta^{\sigma(0)}b_{\sigma(0)}+\sum_{t\in\Gamma,t<s}\delta^{t}\left(l_{t}% -b_{t}+\delta^{\sigma(t)-t}b_{\sigma(t)}\right)=W_{0}+\sum_{t\in\Gamma,t<s}% \delta^{t}l_{t}+\delta^{s}b_{s},

where again the first equality uses the definition of ^𝐂\widehat{}\mathbf{C} and the second follows from simplification. For ~𝐂\widetilde{}\mathbf{C}, since there are no penalties, the corresponding discounted transfer is

W~0+δsb~s=W0+tΓδtlt+δs(bsts,tΓδtsls)=W0+tΓ,t<sδtlt+δsbs,\displaystyle\widetilde{W}_{0}+\delta^{s}\widetilde{b}_{s}=W_{0}+\sum\limits_{% t\in\Gamma}\delta^{t}l_{t}+\delta^{s}\Bigg{(}b_{s}-\sum\limits_{t\geq s,t\in% \Gamma}\delta^{t-s}l_{s}\Bigg{)}=W_{0}+\sum\limits_{t\in\Gamma,t<s}\delta^{t}l% _{t}+\delta^{s}b_{s},

where again the first equality is by definition of ~𝐂\widetilde{}\mathbf{C} and the second from simplification. Hence, all three contracts induce the same transfer in the event of success in any period sΓs\in\Gamma.

D.2 Proof of Proposition 2

We use a monotone comparative statics argument. Recall expression (B.33), which was the portion of the principal’s objective that involves a stopping time for the low type, TT:

V(T,β0,μ0,c,δ,λL,λH)\displaystyle V(T,\beta_{0},\mu_{0},c,\delta,\lambda^{L},\lambda^{H}) :=\displaystyle:= (1μ0)[β0t=1Tδt(1λL)t1(λLc)(1β0)t=1Tδtc]\displaystyle\left(1-\mu_{0}\right)\left[\beta_{0}\sum\limits_{t=1}^{T}\delta^% {t}\left(1-\lambda^{L}\right)^{t-1}\left(\lambda^{L}-c\right)-(1-\beta_{0})% \sum\limits_{t=1}^{T}\delta^{t}c\right]
μ0β0{t=1Tδtl¯tL(T)[(1λH)t(1λL)t]t=1Tδtc[(1λH)t1(1λL)t1]},\displaystyle-\mu_{0}\beta_{0}\left\{\begin{array}[]{l}\sum\limits_{t=1}^{T}% \delta^{t}\overline{l}_{t}^{L}(T)\left[\left(1-\lambda^{H}\right)^{t}-\left(1-% \lambda^{L}\right)^{t}\right]\\ -\sum\limits_{t=1}^{T}\delta^{t}c\left[\left(1-\lambda^{H}\right)^{t-1}-\left(% 1-\lambda^{L}\right)^{t-1}\right]\end{array}\right\},

where l¯tL(T)\overline{l}_{t}^{L}(T) is given by (6) in Theorem 3. The second-best stopping time, t¯L\overline{t}^{L}, is the TT that maximizes V(T,)V(T,\cdot).5555 55 While the maximizer is generically unique, recall that if multiple maximizers exist we select the largest one. To establish the comparative statics of t¯L\overline{t}^{L} with respect to the parameters, we show that V(T,)V(T,\cdot) has increasing or decreasing differences in TT and the relevant parameter.

Substituting l¯tL(T)\overline{l}_{t}^{L}(T) from (6) into V()V(\cdot) above yields

V(T,β0,μ0,c,δ,λL,λH)\displaystyle V(T,\beta_{0},\mu_{0},c,\delta,\lambda^{L},\lambda^{H}) =(1μ0)[β0t=1Tδt(1λL)t1(λLc)(1β0)t=1Tδtc]\displaystyle=\left(1-\mu_{0}\right)\left[\beta_{0}\sum\limits_{t=1}^{T}\delta% ^{t}\left(1-\lambda^{L}\right)^{t-1}\left(\lambda^{L}-c\right)-(1-\beta_{0})% \sum\limits_{t=1}^{T}\delta^{t}c\right]
μ0{ct=1T1δt(1δ)β0(1λL)t1+1β0λL(1λL)t1[(1λH)t(1λL)t]cδTβ0(1λL)T1+1β0λL(1λL)TL1[(1λH)T(1λL)T]β0t=1Tδtc[(1λH)t1(1λL)t1]}.\displaystyle-\mu_{0}\left\{\begin{array}[]{l}-c\sum\limits_{t=1}^{T-1}\delta^% {t}\left(1-\delta\right)\frac{\beta_{0}\left(1-\lambda^{L}\right)^{t-1}+1-% \beta_{0}}{\lambda^{L}\left(1-\lambda^{L}\right)^{t-1}}\left[\left(1-\lambda^{% H}\right)^{t}-\left(1-\lambda^{L}\right)^{t}\right]\\ -c\delta^{T}\frac{\beta_{0}\left(1-\lambda^{L}\right)^{T-1}+1-\beta_{0}}{% \lambda^{L}\left(1-\lambda^{L}\right)^{T^{L}-1}}\left[\left(1-\lambda^{H}% \right)^{T}-\left(1-\lambda^{L}\right)^{T}\right]\\ -\beta_{0}\sum\limits_{t=1}^{T}\delta^{t}c\left[\left(1-\lambda^{H}\right)^{t-% 1}-\left(1-\lambda^{L}\right)^{t-1}\right]\end{array}\right\}. (D.5)

After some algebraic manipulation, we obtain

V(T+1,β0,μ0,c,δ,λL,λH)V(T,β0,μ0,c,δ,λL,λH)\displaystyle V(T+1,\beta_{0},\mu_{0},c,\delta,\lambda^{L},\lambda^{H})-V(T,% \beta_{0},\mu_{0},c,\delta,\lambda^{L},\lambda^{H})
=δT+1{(1μ0)[β0(1λL)T(λLc)(1β0)c]μ0cβ0(1λL)T+1β0(1λL)TλL(1λH)T(λHλL)}.\displaystyle=\delta^{T+1}\left\{\begin{array}[]{l}\left(1-\mu_{0}\right)\left% [\beta_{0}\left(1-\lambda^{L}\right)^{T}\left(\lambda^{L}-c\right)-(1-\beta_{0% })c\right]\\ -\mu_{0}c\frac{\beta_{0}\left(1-\lambda^{L}\right)^{T}+1-\beta_{0}}{\left(1-% \lambda^{L}\right)^{T}\lambda^{L}}\left(1-\lambda^{H}\right)^{T}\left(\lambda^% {H}-\lambda^{L}\right)\end{array}\right\}. (D.8)

(D.2) implies that V(T,β0,μ0,c,δ,λL,λH)V(T,\beta_{0},\mu_{0},c,\delta,\lambda^{L},\lambda^{H}) has increasing differences in (T,β0)\left(T,\beta_{0}\right), because

β0[V(T+1,β0,)V(T,β0,)]=δT+1{(1μ0)[(1λL)T(λLc)+c]+μ0c1(1λL)T(1λL)TλL(1λH)T(λHλL)}>0.\frac{\partial}{\partial\beta_{0}}\left[V(T+1,\beta_{0},\cdot)-V(T,\beta_{0},% \cdot)\right]=\delta^{T+1}\left\{\begin{array}[]{l}\left(1-\mu_{0}\right)\left% [\left(1-\lambda^{L}\right)^{T}\left(\lambda^{L}-c\right)+c\right]\\ +\mu_{0}c\frac{1-\left(1-\lambda^{L}\right)^{T}}{\left(1-\lambda^{L}\right)^{T% }\lambda^{L}}\left(1-\lambda^{H}\right)^{T}\left(\lambda^{H}-\lambda^{L}\right% )\end{array}\right\}>0.

It thus follows that t¯L\overline{t}^{L} is increasing in β0\beta_{0}. Similarly, (D.2) also implies

c[V(T+1,c,)V(T,c,)]=δT+1{(1μ0)[β0(1λL)T+(1β0)]μ0β0(1λL)T+1β0(1λL)TλL(1λH)T(λHλL)}<0,\frac{\partial}{\partial c}\left[V(T+1,c,\cdot)-V(T,c,\cdot)\right]=\delta^{T+% 1}\left\{\begin{array}[]{l}-\left(1-\mu_{0}\right)\left[\beta_{0}\left(1-% \lambda^{L}\right)^{T}+(1-\beta_{0})\right]\\ -\mu_{0}\frac{\beta_{0}\left(1-\lambda^{L}\right)^{T}+1-\beta_{0}}{\left(1-% \lambda^{L}\right)^{T}\lambda^{L}}\left(1-\lambda^{H}\right)^{T}\left(\lambda^% {H}-\lambda^{L}\right)\end{array}\right\}<0,

and hence t¯L\overline{t}^{L} is decreasing in cc.

To obtain the comparative static of t¯L\overline{t}^{L} in μ0\mu_{0}, we compute

μ0[V(T+1,μ0,)V(T,μ0,)]=δT+1{[β0(1λL)T(λLc)(1β0)c]cβ0(1λL)T+1β0(1λL)TλL(1λH)T(λHλL)}.\frac{\partial}{\partial\mu_{0}}\left[V(T+1,\mu_{0},\cdot)-V(T,\mu_{0},\cdot)% \right]=\delta^{T+1}\left\{\begin{array}[]{l}-\left[\beta_{0}\left(1-\lambda^{% L}\right)^{T}\left(\lambda^{L}-c\right)-(1-\beta_{0})c\right]\\ -c\frac{\beta_{0}\left(1-\lambda^{L}\right)^{T}+1-\beta_{0}}{\left(1-\lambda^{% L}\right)^{T}\lambda^{L}}\left(1-\lambda^{H}\right)^{T}\left(\lambda^{H}-% \lambda^{L}\right)\end{array}\right\}. (D.9)

Recall that the first-best stopping time tLt^{L} is such that β0(1λL)tL1β0(1λL)tL1+1β0λLc\frac{\beta_{0}\left(1-\lambda^{L}\right)^{t^{L}-1}}{\beta_{0}\left(1-\lambda^% {L}\right)^{t^{L}-1}+1-\beta_{0}}\lambda^{L}\geq c, which is equivalent to β0(1λL)tL1(λLc)(1β0)c0.\beta_{0}\left(1-\lambda^{L}\right)^{t^{L}-1}\left(\lambda^{L}-c\right)-\left(% 1-\beta_{0}\right)c\geq 0. Thus, for T+1tLT+1\leq t^{L},

β0(1λL)T(λLc)(1β0)c0.\beta_{0}\left(1-\lambda^{L}\right)^{T}\left(\lambda^{L}-c\right)-\left(1-% \beta_{0}\right)c\geq 0. (D.10)

Combining (D.9) and (D.10) implies

μ0[V(T+1,μ0,)V(T,μ0,)]δT+1cβ0(1λL)T+1β0(1λL)TλL(1λH)T(λHλL)<0.\frac{\partial}{\partial\mu_{0}}\left[V(T+1,\mu_{0},\cdot)-V(T,\mu_{0},\cdot)% \right]\leq-\delta^{T+1}c\frac{\beta_{0}\left(1-\lambda^{L}\right)^{T}+1-\beta% _{0}}{\left(1-\lambda^{L}\right)^{T}\lambda^{L}}\left(1-\lambda^{H}\right)^{T}% \left(\lambda^{H}-\lambda^{L}\right)<0.

It follows that t¯L\overline{t}^{L} is decreasing in μ0\mu_{0}.

We next consider the comparative statics of t¯L\overline{t}^{L} with respect to λL\lambda^{L} and λH\lambda^{H}. For λL\lambda^{L}, note that since T+1tLT+1\leq t^{L} and the first-best stopping time is increasing in ability starting at λL\lambda^{L}, the social surplus from the low type (given by the expression in the first square brackets in (D.2)) has increasing differences in (T,λL)(T,\lambda^{L}), and the low type’s expected marginal product given work up to T+1T+1, β¯T+1LλL\overline{\beta}^{L}_{T+1}\lambda^{L}, is increasing in λL\lambda^{L}. Therefore, substituting β0(1λL)T+1β0(1λL)TλL=β0β¯T+1LλL\frac{\beta_{0}\left(1-\lambda^{L}\right)^{T}+1-\beta_{0}}{\left(1-\lambda^{L}% \right)^{T}\lambda^{L}}=\frac{\beta_{0}}{\overline{\beta}^{L}_{T+1}\lambda^{L}} in (D.2), we obtain

λL[V(T+1,λL,)V(T,λL,)]=δT+1{(1μ0)β0[(1λL)TT(1λL)T1(λLc)]+μ0c(1λH)Tβ0β¯T+1LλL(1+λHλLβ¯T+1LλL(β¯T+1LλL)λL)}>0,\frac{\partial}{\partial\lambda^{L}}\left[V(T+1,\lambda^{L},\cdot)-V(T,\lambda% ^{L},\cdot)\right]=\delta^{T+1}\left\{\begin{array}[]{l}\left(1-\mu_{0}\right)% \beta_{0}\left[\left(1-\lambda^{L}\right)^{T}-T\left(1-\lambda^{L}\right)^{T-1% }\left(\lambda^{L}-c\right)\right]\\ +\mu_{0}c\left(1-\lambda^{H}\right)^{T}\frac{\beta_{0}}{\overline{\beta}^{L}_{% T+1}\lambda^{L}}\left(1+\frac{\lambda^{H}-\lambda^{L}}{\overline{\beta}^{L}_{T% +1}\lambda^{L}}\frac{\partial\left(\overline{\beta}^{L}_{T+1}\lambda^{L}\right% )}{\partial\lambda^{L}}\right)\end{array}\right\}>0,

which implies that t¯L\overline{t}^{L} is increasing in λL\lambda^{L}.

That t¯L\overline{t}^{L} can increase or decrease in λH\lambda^{H} follows from the fact that (D.2) yields

λH[V(T+1,λH,)V(T,λH,)]=δT+1μ0cβ0(1λL)T+1β0(1λL)TλL[(1λH)TT(1λH)T1(λHλL)],\frac{\partial}{\partial\lambda^{H}}\left[V(T+1,\lambda^{H},\cdot)-V(T,\lambda% ^{H},\cdot)\right]=-\delta^{T+1}\mu_{0}c\frac{\beta_{0}\left(1-\lambda^{L}% \right)^{T}+1-\beta_{0}}{\left(1-\lambda^{L}\right)^{T}\lambda^{L}}\left[\left% (1-\lambda^{H}\right)^{T}-T\left(1-\lambda^{H}\right)^{T-1}\left(\lambda^{H}-% \lambda^{L}\right)\right],

whose sign can vary with parameters. Specifically, let (β0,μ0,c,δ,λL)=(0.95,0.1,0.215,0.8,0.25)(\beta_{0},\mu_{0},c,\delta,\lambda^{L})=(0.95,0.1,0.215,0.8,0.25), which results in a first-best stopping time tL=5t^{L}=5. Consider three values of λH\lambda^{H}: λ1H=0.45\lambda^{H}_{1}=0.45, λ2H=0.5\lambda^{H}_{2}=0.5, and λ3H=0.55\lambda^{H}_{3}=0.55. The corresponding first-best stopping times are t1H=6t^{H}_{1}=6, t2H=5t^{H}_{2}=5, and t3H=5t^{H}_{3}=5. One can verify that the low type’s second-best stopping time, t¯L\overline{t}^{L}, increases (from 33 to 44) when λH\lambda^{H} increases from λ1H\lambda^{H}_{1} to λ2H\lambda^{H}_{2} while it decreases (from 44 to 0) when λH\lambda^{H} increases from λ2H\lambda^{H}_{2} to λ3H\lambda^{H}_{3}.

Finally, consider the comparative statics of the distortion, tLt¯Lt^{L}-\overline{t}^{L}. By (3), tLt^{L} is independent of μ0\mu_{0} and λH\lambda^{H}, while we have just shown that t¯L\overline{t}^{L} is decreasing in μ0\mu_{0} and can increase or decrease in λH\lambda^{H}. Therefore, tLt¯Lt^{L}-\overline{t}^{L} is increasing in μ0\mu_{0} and can increase or decrease in λH\lambda^{H} depending on parameters. To see that tLt¯Lt^{L}-\overline{t}^{L} can increase or decrease in β0\beta_{0} as well, take the set of parameters considered in Figure 2, (μ0,c,δ,λL,λH)=(0.3,0.06,0.5,0.1,0.12)(\mu_{0},c,\delta,\lambda^{L},\lambda^{H})=(0.3,0.06,0.5,0.1,0.12). The figure shows that given these parameters, tLt¯Lt^{L}-\overline{t}^{L} decreases (from 1210=212-10=2 to 1514=115-14=1) when β0\beta_{0} increases from 0.850.85 to 0.890.89. If instead we take these parameter values but change only μ0\mu_{0} to μ0=0.7\mu_{0}=0.7, we find that the same increase in β0\beta_{0} leads to an increase in tLt¯Lt^{L}-\overline{t}^{L} (from 121=1112-1=11 to 151=1415-1=14). The comparative static of tLt¯Lt^{L}-\overline{t}^{L} with respect to cc and λL\lambda^{L} can be shown by similar computations.

D.3 Step 6 of Proof of Theorem 5

We remind the reader that Steps 1–5 of the proof of Theorem 5 are in Appendix C of the paper.

By the previous steps in the proof, we restrict attention to onetime-penalty contracts for the low type such that the low type works in all periods t{1,,TL}t\in\{1,\ldots,T^{L}\} and the high type has a most-work optimal stopping strategy. For an arbitrary such contract 𝐂L\mathbf{C}^{L}, let t^(𝐂L)\hat{t}(\mathbf{C}^{L}) denote the high type’s most-work optimal stopping time, i.e. t^(𝐂L):=max{t{1,,TL}:a^s=1 for all s=1,,t,𝐚^𝜶H(𝐂L)}\hat{t}(\mathbf{C}^{L}):=\max\{t\in\{1,\ldots,T^{L}\}:\widehat{a}_{s}=1\text{ % for all }s=1,\ldots,t,\ \widehat{\mathbf{a}}\in\bm{\alpha}^{H}(\mathbf{C}^{L})\}. We now show that given TLT^{L}, there exists an optimal onetime-penalty contract for the low type 𝐂L=(TL,W0L,lTLL)\mathbf{C}^{L}=(T^{L},W_{0}^{L},l_{T^{L}}^{L}) where t^(𝐂L)\hat{t}(\mathbf{C}^{L}) is given by

THL(TL):=min{t{1,,TL}:β¯t+1HλH<β¯TLLλL and (1λH)t(1λL)TL},T^{HL}(T^{L}):=\min\left\{t\in\{1,\ldots,T^{L}\}:\overline{\beta}^{H}_{t+1}% \lambda^{H}<\overline{\beta}^{L}_{T^{L}}\lambda^{L}\text{ and }(1-\lambda^{H})% ^{t}\leq(1-\lambda^{L})^{T^{L}}\right\},

and lTLLl_{T^{L}}^{L} is given by

l¯TLL(TL):=min{cβ¯TLLλL,cβ¯THL(TL)HλH}.\overline{l}^{L}_{T^{L}}(T^{L}):=\min\left\{-\frac{c}{\overline{\beta}^{L}_{T^% {L}}\lambda^{L}},-\frac{c}{\overline{\beta}^{H}_{T^{HL}(T^{L})}\lambda^{H}}% \right\}.

When not essential, we suppress the dependence of t^(𝐂L)\hat{t}(\mathbf{C}^{L}) on 𝐂L\mathbf{C}^{L}. We proceed by proving five claims.

Claim 1: Given any onetime-penalty contract 𝐂L=(TL,W0L,lTLL)\mathbf{C}^{L}=(T^{L},W_{0}^{L},l_{T^{L}}^{L}), β¯t^+1HλHlTLL<c-\overline{\beta}^{H}_{\hat{t}+1}\lambda^{H}l^{L}_{T^{L}}<c.

Proof: Suppose to contradiction that β¯t^+1HλHlTLLc-\overline{\beta}^{H}_{\hat{t}+1}\lambda^{H}l^{L}_{T^{L}}\geq c. Then type HH is willing to work one more period after having worked for t^\hat{t} periods, contradicting the definition of t^\hat{t}. \parallel

Claim 2: Given an optimal onetime-penalty contract 𝐂L=(TL,W0L,lTLL)\mathbf{C}^{L}=(T^{L},W_{0}^{L},l_{T^{L}}^{L}), (1λH)t^(1λL)TL(1-\lambda^{H})^{\hat{t}}\leq(1-\lambda^{L})^{T^{L}}.

Proof: Suppose to contradiction that given an optimal contract 𝐂L=(TL,W0L,lTLL)\mathbf{C}^{L}=(T^{L},W_{0}^{L},l_{T^{L}}^{L}), type HH’s most-work optimal stopping time t^\hat{t} is such that (1λH)t^>(1λL)TL(1-\lambda^{H})^{\hat{t}}>(1-\lambda^{L})^{T^{L}}. Then for any strategy 𝐚~𝜶H(𝐂L)\widetilde{\mathbf{a}}\in\bm{\alpha}^{H}(\mathbf{C}^{L}) where type HH works for a total of t~\widetilde{t} periods, (1λH)t~>(1λL)TL(1-\lambda^{H})^{\widetilde{t}}>(1-\lambda^{L})^{T^{L}}. Now note that given 𝐂L\mathbf{C}^{L} and 𝐚~\widetilde{\mathbf{a}}, type HH’s information rent is

β0lTLL[(1λH)t~(1λL)TL]β0ct=1TLa~t[s=1t1(1a~sλH)(1λL)t1]+ct=1TL(1a~t)[(1β0)+β0(1λL)t1].\begin{array}[]{l}\beta_{0}l_{T^{L}}^{L}\left[(1-\lambda^{H})^{\widetilde{t}}-% (1-\lambda^{L})^{T^{L}}\right]-\beta_{0}c\sum\limits_{t=1}^{T^{L}}\widetilde{a% }_{t}\left[\prod\limits_{s=1}^{t-1}\left(1-\widetilde{a}_{s}\lambda^{H}\right)% -\left(1-\lambda^{L}\right)^{t-1}\right]\\ +c\sum\limits_{t=1}^{T^{L}}\left(1-\widetilde{a}_{t}\right)\left[(1-\beta_{0})% +\beta_{0}\left(1-\lambda^{L}\right)^{t-1}\right].\end{array}

Consider a modification that reduces lTLLl_{T^{L}}^{L} by ε>0\varepsilon>0. By Claim 1, for ε\varepsilon small enough, this modification does not affect incentives, and by (1λH)t~>(1λL)TL(1-\lambda^{H})^{\widetilde{t}}>(1-\lambda^{L})^{T^{L}}, the modification strictly reduces type HH’s information rent. But then 𝐂L\mathbf{C}^{L} cannot be optimal. \parallel

Claim 3: In any onetime-penalty contract 𝐂L=(TL,W0L,lTLL)\mathbf{C}^{L}=(T^{L},W_{0}^{L},l_{T^{L}}^{L}), if 𝟏𝜶L(𝐂L)\mathbf{1}\in\bm{\alpha}^{L}(\mathbf{C}^{L}) then

lTLLmin{cβ¯TLLλL,cβ¯t^(𝐂L)HλH}.{l}^{L}_{T^{L}}\leq\min\left\{-\frac{c}{\overline{\beta}^{L}_{T^{L}}\lambda^{L% }},-\frac{c}{\overline{\beta}^{H}_{\hat{t}(\mathbf{C}^{L})}\lambda^{H}}\right\}. (D.11)

Conversely, given any onetime penalty contract 𝐂L=(TL,W0L,lTLL)\mathbf{C}^{L}=(T^{L},W_{0}^{L},l_{T^{L}}^{L}), if lTLLmin{cβ¯TLLλL,cβ¯tHλH}l_{T^{L}}^{L}\leq\min\left\{-\frac{c}{\overline{\beta}^{L}_{T^{L}}\lambda^{L}}% ,-\frac{c}{\overline{\beta}^{H}_{t}\lambda^{H}}\right\} for some tTLt\leq T^{L}, then t^(𝐂L)t\hat{t}(\mathbf{C}^{L})\geq t and 𝟏𝜶L(𝐂L)\mathbf{1}\in\bm{\alpha}^{L}(\mathbf{C}^{L}).

Proof: For the first part of the claim, assume to contradiction that there is 𝐂L=(TL,W0L,lTLL)\mathbf{C}^{L}=(T^{L},W_{0}^{L},l_{T^{L}}^{L}) such that 𝟏𝜶L(𝐂L)\mathbf{1}\in\bm{\alpha}^{L}(\mathbf{C}^{L}) but (D.11) does not hold. Suppose first that cβ¯TLLλLcβ¯t^HλH-\frac{c}{\overline{\beta}^{L}_{T^{L}}\lambda^{L}}\leq-\frac{c}{\overline{% \beta}^{H}_{\hat{t}}\lambda^{H}}. Then type LL is not willing to work for TLT^{L} periods; having worked for TL1T^{L}-1 periods, type LL’s incentive compatibility constraint for effort in period TLT^{L} is β¯TLLλLlTLLc,-\overline{\beta}^{L}_{T^{L}}\lambda^{L}l^{L}_{T^{L}}\geq c, which is not satisfied with lTLL>cβ¯TLLλLl_{T^{L}}^{L}>-\frac{c}{\overline{\beta}^{L}_{T^{L}}\lambda^{L}}. Suppose next that cβ¯TLLλL>cβ¯t^HλH-\frac{c}{\overline{\beta}^{L}_{T^{L}}\lambda^{L}}>-\frac{c}{\overline{\beta}^% {H}_{\hat{t}}\lambda^{H}}. Then type HH is not willing to work for t^\hat{t} periods; having worked for t^1\hat{t}-1 periods, type HH is willing to work one more period only if β¯t^LλHlTLLc,-\overline{\beta}^{L}_{\hat{t}}\lambda^{H}l^{L}_{T^{L}}\geq c, which is not satisfied with lTLL>cβ¯t^HλLl_{T^{L}}^{L}>-\frac{c}{\overline{\beta}^{H}_{\hat{t}}\lambda^{L}}.

For the second part of the claim, assume lTLLmin{cβ¯TLLλL,cβ¯tHλH}l_{T^{L}}^{L}\leq\min\left\{-\frac{c}{\overline{\beta}^{L}_{T^{L}}\lambda^{L}}% ,-\frac{c}{\overline{\beta}^{H}_{t}\lambda^{H}}\right\}. Consider first type LL. The proof is by induction. Consider the last period, TLT^{L}. Since no matter the history of effort the current belief is some βTLLβ¯TLL\beta_{T^{L}}^{L}\geq\overline{\beta}_{T^{L}}^{L}, it is immediate that βTLLλLlTLLc-{\beta}_{T^{L}}^{L}\lambda^{L}l_{T^{L}}^{L}\geq c, and thus it is optimal for type LL to work in the last period. Now assume inductively that it is optimal for type LL to work in period t+1TLt+1\leq T^{L} no matter the history of effort, and consider period tt with belief βtL\beta_{t}^{L}. The inductive hypothesis implies that

βt+1LλL{lTLL(1λL)TL(t+1)cs=t+2TL(1λL)s(t+2)}c.-\beta_{t+1}^{L}\lambda^{L}\left\{l_{T^{L}}^{L}(1-\lambda^{L})^{T^{L}-(t+1)}-c% \sum\limits_{s=t+2}^{T^{L}}\left(1-\lambda^{L}\right)^{s-\left(t+2\right)}% \right\}\geq c. (D.12)

Therefore, at period tt:

βtLλL{c+(1λL)[lTLL(1λL)TL(t+1)cs=t+2TL(1λL)s(t+2)]}βtLλL[c+(1λL)(cβt+1LλL)]=c,\displaystyle-\beta_{t}^{L}\lambda^{L}\left\{-c+(1-\lambda^{L})\left[l_{T^{L}}% ^{L}(1-\lambda^{L})^{T^{L}-(t+1)}-c\sum\limits_{s=t+2}^{T^{L}}\left(1-\lambda^% {L}\right)^{s-\left(t+2\right)}\right]\right\}\geq-\beta_{t}^{L}\lambda^{L}% \left[-c+(1-\lambda^{L})\left(-\frac{c}{\beta_{t+1}^{L}\lambda^{L}}\right)% \right]=c,

where the inequality uses (D.12) and the equality uses βt+1L=βtL(1λL)1βtL+βtL(1λL)\beta^{L}_{t+1}=\frac{\beta^{L}_{t}(1-\lambda^{L})}{1-\beta^{L}_{t}+\beta^{L}_% {t}(1-\lambda^{L})}.

Finally, consider type HH. By Lemma 3 and the fact that ltL=0l^{L}_{t}=0 for all t=1,,TL1t=1,\ldots,T^{L}-1, type HH is indifferent between any two action plans 𝐚\mathbf{a} and 𝐚\mathbf{a^{\prime}} such that #{t:at=0}=#{t:at=0}\#\left\{t:{a}_{t}=0\right\}=\#\left\{t:a^{\prime}_{t}=0\right\}. Thus, without loss, we restrict attention to stopping strategies, and we only need to show that it is optimal for type HH to stop at sts\geq t. Note that for any s<ts<t, given that type HH has worked consecutively until and including period ss, β¯s+1HλHlTLLc-\overline{\beta}_{s+1}^{H}\lambda^{H}l_{T^{L}}^{L}\geq c, and thus type HH does not want to stop at ss. \parallel

Claim 4: There exists an optimal onetime-penalty contract 𝐂L=(TL,W0L,lTLL)\mathbf{C}^{L}=(T^{L},W_{0}^{L},l_{T^{L}}^{L}) satisfying lTLLmin{cβ¯TLLλL,cβ¯t^HλH}{l}^{L}_{T^{L}}\geq\min\left\{-\frac{c}{\overline{\beta}^{L}_{T^{L}}\lambda^{L% }},-\frac{c}{\overline{\beta}^{H}_{\hat{t}}\lambda^{H}}\right\}.

Proof: Suppose, to contradiction, the claim is false. Given an optimal onetime-penalty contract for type LL, 𝐂L=(TL,W0L,lTLL)\mathbf{C}^{L}=(T^{L},W_{0}^{L},l_{T^{L}}^{L}), and type HH’s most-work optimal stopping strategy 𝐚^\widehat{\mathbf{a}}, type HH’s information rent is

β0lTLL[(1λH)t^(1λL)TL]β0ct=1t^[(1λH)t1(1λL)t1]+ct=t^+1TL[(1β0)+β0(1λL)t1].\begin{array}[]{l}\beta_{0}l_{T^{L}}^{L}[(1-\lambda^{H})^{\hat{t}}-(1-\lambda^% {L})^{T^{L}}]-\beta_{0}c\sum\limits_{t=1}^{\hat{t}}\left[\left(1-\lambda^{H}% \right)^{t-1}-\left(1-\lambda^{L}\right)^{t-1}\right]\\ +c\sum\limits_{t=\hat{t}+1}^{T^{L}}\left[(1-\beta_{0})+\beta_{0}\left(1-% \lambda^{L}\right)^{t-1}\right].\end{array}

Consider a modification that increases lTLLl_{T^{L}}^{L} by ε>0\varepsilon>0. By Claim 4 being false and Claim 3, for ε\varepsilon small enough, working in all periods t=1,,TLt=1,\ldots,T^{L} remains optimal for type LL, and 𝐚^\widehat{\mathbf{a}} remains optimal for type HH. But then Claim 2 implies that type HH’s information rent either goes down or remains unchanged with the modification, and thus there exists an optimal contract 𝐂L=(TL,W0L,lTLL)\mathbf{C}^{L}=(T^{L},W_{0}^{L},l_{T^{L}}^{L}) where the claim is true. \parallel

Claim 5: There is an optimal onetime-penalty contract 𝐂L=(TL,W0L,lTLL)\mathbf{C}^{L}=(T^{L},W_{0}^{L},l_{T^{L}}^{L}) with t^(𝐂L)=THL(TL)\hat{t}(\mathbf{C}^{L})=T^{HL}(T^{L}).

Proof: Take an arbitrary optimal contract 𝐂L=(TL,W0L,lTLL)\mathbf{C}^{L}=(T^{L},W_{0}^{L},l_{T^{L}}^{L}). By Claims 1 and 3, t^(𝐂L)\hat{t}(\mathbf{C}^{L}) satisfies β¯t^(𝐂L)+1HλH<β¯TLLλL\overline{\beta}^{H}_{\hat{t}(\mathbf{C}^{L})+1}\lambda^{H}<\overline{\beta}^{% L}_{T^{L}}\lambda^{L}. By Claim 2, t^(𝐂L)\hat{t}(\mathbf{C}^{L}) satisfies (1λH)t^(𝐂L)(1λL)TL(1-\lambda^{H})^{\hat{t}(\mathbf{C}^{L})}\leq(1-\lambda^{L})^{T^{L}}. Thus, all that remains to be shown is that there exists 𝐂L\mathbf{C}^{L} where t^(𝐂L)\hat{t}(\mathbf{C}^{L}) is the smallest period t{1,,TL}t\in\{1,\ldots,T^{L}\} that satisfies these two conditions. Suppose to contradiction that this claim is false. Then t^(𝐂L)1\hat{t}(\mathbf{C}^{L})-1 also satisfies the conditions; that is, β¯t^(𝐂L)HλH<β¯TLLλL\overline{\beta}^{H}_{\hat{t}(\mathbf{C}^{L})}\lambda^{H}<\overline{\beta}^{L}% _{T^{L}}\lambda^{L} and (1λH)t^(𝐂L)1(1λL)TL(1-\lambda^{H})^{\hat{t}(\mathbf{C}^{L})-1}\leq(1-\lambda^{L})^{T^{L}}. By Claims 3 and 4, lTLL=min{cβ¯TLLλL,cβ¯t^(𝐂L)HλH},{l}^{L}_{T^{L}}=\min\left\{-\frac{c}{\overline{\beta}^{L}_{T^{L}}\lambda^{L}},% -\frac{c}{\overline{\beta}^{H}_{\hat{t}(\mathbf{C}^{L})}\lambda^{H}}\right\}, and thus since β¯t^(𝐂L)HλH<β¯TLLλL\overline{\beta}^{H}_{\hat{t}(\mathbf{C}^{L})}\lambda^{H}<\overline{\beta}^{L}% _{T^{L}}\lambda^{L}, lTLL=cβ¯t^(𝐂L)HλH<cβ¯TLLλL.{l}^{L}_{T^{L}}=-\frac{c}{\overline{\beta}^{H}_{\hat{t}(\mathbf{C}^{L})}% \lambda^{H}}<-\frac{c}{\overline{\beta}^{L}_{T^{L}}\lambda^{L}}. It follows that type HH’s incentive constraint in period t^(𝐂L)\hat{t}(\mathbf{C}^{L}) binds; i.e., type HH is indifferent between working and shirking at t^(𝐂L)\hat{t}(\mathbf{C}^{L}) given that he has worked in all periods t=1,,t^(𝐂L)1t=1,\ldots,\hat{t}(\mathbf{C}^{L})-1 and will shirk in all periods t=t^(𝐂L)+1,,TLt=\hat{t}(\mathbf{C}^{L})+1,\ldots,T^{L}. Hence, both a stopping strategy that stops at t^(𝐂L)\hat{t}(\mathbf{C}^{L}) and a stopping strategy that stops at t^(𝐂L)1\hat{t}(\mathbf{C}^{L})-1 are optimal for type HH given 𝐂L\mathbf{C}^{L}, and type HH’s information rent is the same for either of these two action plans. Type HH’s information rent can thus be written as

β0lTLL[(1λH)t^(𝐂L)1(1λL)TL]β0ct=1t^(𝐂L)1[(1λH)t1(1λL)t1]+ct=t^(𝐂L)TL[(1β0)+β0(1λL)t1].\begin{array}[]{l}\beta_{0}l_{T^{L}}^{L}[(1-\lambda^{H})^{\hat{t}(\mathbf{C}^{% L})-1}-(1-\lambda^{L})^{T^{L}}]-\beta_{0}c\sum\limits_{t=1}^{\hat{t}(\mathbf{C% }^{L})-1}\left[\left(1-\lambda^{H}\right)^{t-1}-\left(1-\lambda^{L}\right)^{t-% 1}\right]\\ +c\sum\limits_{t=\hat{t}(\mathbf{C}^{L})}^{T^{L}}\left[(1-\beta_{0})+\beta_{0}% \left(1-\lambda^{L}\right)^{t-1}\right].\end{array}

Now consider a modified contract, 𝐂^L\widehat{\mathbf{C}}^{L}, obtained from 𝐂L\mathbf{C}^{L} by increasing lTLLl_{T^{L}}^{L} by ε>0\varepsilon>0. Since lTLL=cβ¯t^(𝐂L)HλH{l}^{L}_{T^{L}}=-\frac{c}{\overline{\beta}^{H}_{\hat{t}(\mathbf{C}^{L})}% \lambda^{H}}, a stopping strategy that stops at t^(𝐂L)\hat{t}(\mathbf{C}^{L}) is no longer optimal for type HH under 𝐂^L\widehat{\mathbf{C}}^{L}. Since lTLL<cβ¯TLLλL{l}^{L}_{T^{L}}<-\frac{c}{\overline{\beta}^{L}_{T^{L}}\lambda^{L}} and lTLL<cβ¯t^(𝐂L)1HλH{l}^{L}_{T^{L}}<-\frac{c}{\overline{\beta}^{H}_{\hat{t}(\mathbf{C}^{L})-1}% \lambda^{H}}, for ε\varepsilon small enough, 𝟏𝜶L(𝐂^L)\mathbf{1}\in\bm{\alpha}^{L}(\widehat{\mathbf{C}}^{L}) and a stopping strategy that stops at t^(𝐂L)1\hat{t}(\mathbf{C}^{L})-1 remains optimal for type HH under 𝐂^L\widehat{\mathbf{C}}^{L}. Then t^(𝐂^L)=t^(𝐂L)1\hat{t}(\widehat{\mathbf{C}}^{L})=\hat{t}(\mathbf{C}^{L})-1, and since (1λH)t^(𝐂L)1(1λL)TL(1-\lambda^{H})^{\hat{t}(\mathbf{C}^{L})-1}\leq(1-\lambda^{L})^{T^{L}}, type HH’s information rent either goes down or remains unchanged with the modification, so 𝐂^L\widehat{\mathbf{C}}^{L} is optimal. If t^(𝐂^L)=THL(TL)\hat{t}(\widehat{\mathbf{C}}^{L})=T^{HL}(T^{L}), we are done. Otherwise, we can apply the argument to t^(𝐂^L)\hat{t}(\widehat{\mathbf{C}}^{L}) and repeat until we eventually arrive at the desired contract 𝐂L\mathbf{C}^{L} with t^(𝐂L)=THL\hat{t}(\mathbf{C}^{L})=T^{HL}. \parallel

D.4 Details for Subsection 7.1

Here we provide a formal result for the discussion in Subsection 7.1 of the paper.

Theorem 7.

Even if project success is privately observed by the agent, the menus of contracts identified in Theorems 3–6 remain optimal and implement the same outcome as when project success is publicly observable.

Proof.

It suffices to show that in each of the menus, each of the contracts would induce the agent (of either type) to reveal project success immediately when it is obtained. Consider first the menus of Theorem 3 and Theorem 5: for each θ{L,H}\theta\in\{L,H\}, the contract for type θ\theta, 𝐂θ\mathbf{C}^{\theta}, is a penalty contract in which ltθ0l^{\theta}_{t}\leq 0 for all tt. Hence, no matter which contract the agent takes and no matter his type, it is optimal to reveal a success when obtained. For the implementation in Theorem 4, observe from (8) that type LL’s bonus contract has the property that δbt+1LbtL\delta b^{L}_{t+1}\leq b^{L}_{t} for all t{1,,t¯L1}t\in\{1,\ldots,\overline{t}^{L}-1\}; moreover, this property also holds in type LL’s bonus contract in Theorem 6 and in type HH’s bonus contracts in both Theorem 4 and Theorem 6, as these contracts are constant-bonus contracts. Hence, under all these contracts, it is optimal for the agent of either type to disclose success immediately when obtained. ∎

D.5 Details for Subsection 7.2

Here we provide a formal result for the discussion in Subsection 7.2 of the paper.

Theorem 8.

Assume tH>tLt^{H}>t^{L}, δ=1\delta=1, and that all transfers must be non-negative. In any optimal menu of contracts, each type θ{L,H}\theta\in\{L,H\} is induced to work for some number of periods, t¯θ\overline{t}^{\theta}_{\ell\ell}, where t¯Lt¯H\overline{t}^{L}_{\ell\ell}\leq\overline{t}^{H}_{\ell\ell}. Relative to the first-best stopping times, tHt^{H} and tLt^{L}, the second best has t¯HtH\overline{t}^{H}_{\ell\ell}\leq t^{H} and t¯LtL\overline{t}^{L}_{\ell\ell}\leq t^{L}. The principal can implement the second best using a bonus contract for type HH, 𝐂H=(t¯H,W0H,𝐛H)\mathbf{C}^{H}=(\overline{t}^{H}_{\ell\ell},W^{H}_{0},\bm{b}^{H}), and a constant-bonus contract for type LL, 𝐂L=(t¯L,W0L,bL)\mathbf{C}^{L}=(\overline{t}^{L}_{\ell\ell},W^{L}_{0},b^{L}), such that

  1. 1.

    bL=cβ¯t¯LLλLb^{L}=\frac{c}{\overline{\beta}^{L}_{\overline{t}^{L}_{\ell\ell}}\lambda^{L}};

  2. 2.

    Type HH gets a rent: U0H(𝐂H,𝜶H(𝐂H))>0U^{H}_{0}(\mathbf{C}^{H},\bm{\alpha}^{H}(\mathbf{C}^{H}))>0;

  3. 3.

    If t¯L>0\overline{t}^{L}_{\ell\ell}>0, type LL gets a rent: U0L(𝐂L,𝜶L(𝐂L))>0U^{L}_{0}(\mathbf{C}^{L},\bm{\alpha}^{L}(\mathbf{C}^{L}))>0;

  4. 4.

    𝟏𝜶H(𝐂H)\mathbf{1}\in\bm{\alpha}^{H}(\mathbf{C}^{H}); 𝟏𝜶L(𝐂L)\mathbf{1}\in\bm{\alpha}^{L}(\mathbf{C}^{L}); and 𝟏=𝜶H(𝐂L)\mathbf{1}=\bm{\alpha}^{H}(\mathbf{C}^{L}).

Proof.

The principal’s program is the following, called [Pℓℓ]:

max(𝐂H,𝐂L,𝐚H,𝐚L)μ0Π0H(𝐂H,𝐚H)+(1μ0)Π0L(𝐂L,𝐚L)\max_{\left(\mathbf{C}^{H},\mathbf{C}^{L},\mathbf{a}^{H},\mathbf{a}^{L}\right)% }\mu_{0}\Pi_{0}^{H}\left(\mathbf{C}^{H},\mathbf{a}^{H}\right)+\left(1-\mu_{0}% \right)\Pi_{0}^{L}\left(\mathbf{C}^{L},\mathbf{a}^{L}\right) (Pℓℓ)

subject to, for all θ,θ{L,H}\theta,\theta^{\prime}\in\left\{L,H\right\},

𝐚θ\displaystyle\mathbf{a}^{\theta} 𝜶θ(𝐂θ),\displaystyle\in\bm{\alpha}^{\theta}(\mathbf{C}^{\theta}), (ICaθ{}^{\theta}_{a})
U0θ(𝐂θ,𝐚θ)\displaystyle U_{0}^{\theta}(\mathbf{C}^{\theta},\mathbf{a}^{\theta}) 0,\displaystyle\geq 0, (IRθ)
U0θ(𝐂θ,𝐚θ)\displaystyle U_{0}^{\theta}(\mathbf{C}^{\theta},\mathbf{a}^{\theta}) U0θ(𝐂θ,𝜶θ(𝐂θ)),\displaystyle\geq U_{0}^{\theta}(\mathbf{C}^{\theta^{\prime}},\bm{\alpha}^{% \theta}(\mathbf{C}^{\theta^{\prime}})), (ICθθ{}^{\theta\theta^{\prime}})
W0θ,btθ,ltθ\displaystyle W_{0}^{\theta},b^{\theta}_{t},l^{\theta}_{t} 0 for all tΓθ.\displaystyle\geq 0\text{ for all }t\in\Gamma^{\theta}. (\ell\ellθ)

Note that the limited liability constraint for type θ\theta, (\ell\ellθ), implies that this type’s participation constraint, (IRθ), is satisfied. From now on, we thus ignore the constraints (IRθ).

Step 1: Bonus contracts

We show that it is without loss to focus on bonus contracts. Suppose by contradiction that in the solution to [Pℓℓ], for some θ{L,H}\theta\in\{L,H\}, 𝐂θ=(Γθ,W0θ,𝒃θ,𝒍θ)\mathbf{C}^{\theta}=(\Gamma^{\theta},W^{\theta}_{0},\bm{b}^{\theta},\bm{l}^{% \theta}) is not a bonus contract, i.e. ltθ0l^{\theta}_{t}\neq 0 for some tΓθt\in\Gamma^{\theta}. We can construct an equivalent bonus contract ~𝐂θ=(Γθ,W~0θ,~𝒃θ)\widetilde{}\mathbf{C}^{\theta}=(\Gamma^{\theta},\widetilde{W}^{\theta}_{0},% \widetilde{}\bm{b}^{\theta}) as in the proof of Proposition 1:

  • (a)

    For any tΓθt\in\Gamma^{\theta},  b~tθ=btθst,sΓθlsθ\widetilde{b}^{\theta}_{t}=b^{\theta}_{t}-\sum\limits_{s\geq t,s\in\Gamma^{% \theta}}l^{\theta}_{s},

  • (b)

    W~0θ=W0θ+tΓθltθ\widetilde{W}^{\theta}_{0}=W^{\theta}_{0}+\sum\limits_{t\in\Gamma^{\theta}}l^{% \theta}_{t}.

Note that by the limited liability constraint, 𝐂θ\mathbf{C}^{\theta} has W0θ0W^{\theta}_{0}\geq 0 and ltθ0l_{t}^{\theta}\geq 0 for all tΓθt\in\Gamma^{\theta}. Hence, ~𝐂θ\widetilde{}\mathbf{C}^{\theta} has W~0θ0\widetilde{W}^{\theta}_{0}\geq 0. Moreover, if b~tθ<0\widetilde{b}^{\theta}_{t}<0 for some tΓθt\in\Gamma^{\theta}, then regardless of his type, the agent shirks in period tt under contract ~𝐂θ\widetilde{}\mathbf{C}^{\theta}. Therefore, we can define another bonus contract, ^𝐂θ=(Γ^θ,W~0θ,~𝒃θ){\widehat{}\mathbf{C}^{\theta}}=(\widehat{\Gamma}^{\theta},\widetilde{W}^{% \theta}_{0},\widetilde{}\bm{b}^{\theta}), where tΓ^θt\in\widehat{\Gamma}^{\theta} if and only if tΓθt\in\Gamma^{\theta} and b~tθ0\widetilde{b}_{t}^{\theta}\geq 0. Since under contract ~𝐂θ{\widetilde{}\mathbf{C}}^{\theta} the agent of either type receives zero with probability one in all periods tt in which b~tθ<0\widetilde{b}_{t}^{\theta}<0, the incentives for effort for both agent types and the payoffs for the principal and both agent types are unchanged in the new contract ^𝐂θ{\widehat{}\mathbf{C}}^{\theta} in which the agent is locked out in these periods. It follows that the bonus contract ^𝐂θ\widehat{}\mathbf{C}^{\theta} is equivalent to contract ~𝐂θ{\widetilde{}\mathbf{C}}^{\theta} and thus to the original contract 𝐂θ\mathbf{C}^{\theta}, and it satisfies limited liability.

Step 2: Both types always work

We show that it is without loss to focus on bonus contracts in which each type is prescribed to work in every period under his own contract. Suppose that there is a solution to [Pℓℓ] in which, for some θ{L,H}\theta\in\{L,H\}, 𝐂θ=(Γθ,W0θ,𝒃θ)\mathbf{C}^{\theta}=\left(\Gamma^{\theta},W_{0}^{\theta},\bm{b}^{\theta}\right) induces 𝐚θ𝟏\mathbf{a}^{\theta}\neq\mathbf{1}. Consider contract 𝐂^θ=(Γ^θ,W0θ,𝒃θ)\widehat{\mathbf{C}}^{\theta}=\left(\widehat{\Gamma}^{\theta},{W}_{0}^{\theta}% ,{\bm{b}}^{\theta}\right) where tΓ^θt\in\widehat{\Gamma}^{\theta} if and only if tΓθt\in\Gamma^{\theta} and atθ=1a_{t}^{\theta}=1. Notice that in any period tt in which type θ\theta shirks under contract 𝐂θ\mathbf{C}^{\theta}, he receives zero with probability one; this is the same type θ\theta receives under contract 𝐂^θ\widehat{\mathbf{C}}^{\theta} where he is locked out in period tt. It follows that the incentives for effort for type θ\theta and both the principal’s payoff from type θ\theta and type θ\theta’s payoff do not change with the new contract. Moreover, observe that for type θθ\theta^{\prime}\neq\theta, no matter which action he would take at tt in any optimal action plan under 𝐂θ\mathbf{C}^{\theta}, his payoff from 𝐂^θ\widehat{\mathbf{C}}^{\theta} must be weakly lower because the lockout in period tt effectively forces him to shirk in period tt and receive zero.

Step 3: Connected contracts

It is immediate that given δ=1\delta=1, it is without loss to focus on connected bonus contracts: under no discounting, nothing changes when a period tΓθt\notin\Gamma^{\theta} is removed from type θ\theta’s bonus contract, 𝐂θ=(Γθ,W0θ,𝒃θ)\mathbf{C}^{\theta}=(\Gamma^{\theta},W^{\theta}_{0},\bm{b}^{\theta}). When a lockout period is removed, the future sequence of transfers and effort is shifted up by one period, but this has no effect on the payoffs of the principal and the agent of either type when there is no discounting.

Step 4: Relaxing the principal’s program

By Steps 1-3, we restrict attention to connected bonus contracts that induce each agent type to work in each period under his own contract. We now relax the principal’s problem [Pℓℓ] by considering a weak version of (ICHL) in which type HH is assumed to exert effort in all periods t{1,,TL}t\in\{1,\ldots,T^{L}\} if he takes contract 𝐂L\mathbf{C}^{L}. Ignoring the participation constraints as explained above and denoting the set of connected bonus contracts by 𝒞b\mathcal{C}^{b}, the relaxed program, [RPℓℓ], is

max(𝐂H𝒞b,𝐂L𝒞b)μ0Π0H(𝐂H,𝟏)+(1μ0)Π0L(𝐂L,𝟏)\max_{(\mathbf{C}^{H}\in\mathcal{C}^{b},\mathbf{C}^{L}\in\mathcal{C}^{b})}\mu_% {0}\Pi_{0}^{H}\left(\mathbf{C}^{H},\mathbf{1}\right)+\left(1-\mu_{0}\right)\Pi% _{0}^{L}\left(\mathbf{C}^{L},\mathbf{1}\right) (RPℓℓ)

subject to

𝟏\displaystyle\mathbf{1} 𝜶L(𝐂L),\displaystyle\in\bm{\alpha}^{L}(\mathbf{C}^{L}), (ICLa{}_{a}^{L})
𝟏\displaystyle\mathbf{1} 𝜶H(𝐂H),\displaystyle\in\bm{\alpha}^{H}(\mathbf{C}^{H}), (ICHa{}_{a}^{H})
U0L(𝐂L,𝟏)\displaystyle U_{0}^{L}\left(\mathbf{C}^{L},\mathbf{1}\right) U0L(𝐂H,𝜶L(𝐂H)),\displaystyle\geq U_{0}^{L}\left(\mathbf{C}^{H},\bm{\alpha}^{L}(\mathbf{C}^{H}% )\right), (ICLH)
U0H(𝐂H,𝟏)\displaystyle U_{0}^{H}\left(\mathbf{C}^{H},\mathbf{1}\right) U0H(𝐂L,𝟏),\displaystyle\geq U_{0}^{H}\left(\mathbf{C}^{L},\mathbf{1}\right), (Weak-ICHL)
W0L,btL\displaystyle W_{0}^{L},b^{L}_{t} 0 for all t{1,,TL},\displaystyle\geq 0\text{ for all }t\in\{1,\ldots,T^{L}\}, (\ell\ellL)
W0H,btH\displaystyle W_{0}^{H},b^{H}_{t} 0 for all t{1,,TH}.\displaystyle\geq 0\text{ for all }t\in\{1,\ldots,T^{H}\}. (\ell\ellH)

We will solve this relaxed program and later verify that the solution is feasible in (and hence is a solution to) [Pℓℓ].

Step 5: An optimal contract for the low type

Take any arbitrary connected bonus contract 𝐂=(T,W0,𝒃)\mathbf{C}=(T,W_{0},\bm{b}). It follows from Step 3 of the proof of Theorem 3 and the proof of Proposition 1 that type θ\theta’s incentive constraint for effort binds in each period t{1,,T}t\in\{1,\ldots,T\} under contract 𝐂\mathbf{C} if and only if 𝒃=¯𝒃θ(T)\bm{b}=\overline{}\bm{b}^{\theta}(T), where ¯𝒃θ(T)\overline{}\bm{b}^{\theta}(T) is defined as follows:

b¯tθ(T)=b¯θ(T):=cβ¯Tθλθ for all t{1,,T}.\overline{b}^{\theta}_{t}(T)=\overline{b}^{\theta}(T):=\frac{c}{\overline{% \beta}^{\theta}_{T}\lambda^{\theta}}\text{ for all $t\in\{1,\ldots,T\}$}. (D.13)

We can show that in solving program [RPℓℓ], it is without loss to restrict attention to constant-bonus contracts for type LL with bonus as defined in (D.13). The proof follows from Step 4 in the proof of Theorem 3. Take any arbitrary connected bonus contract 𝐂L=(TL,W0L,𝒃L)\mathbf{C}^{L}=(T^{L},W^{L}_{0},\bm{b}^{L}) that induces type LL to work in each period t{1,,TL}t\in\{1,\ldots,T^{L}\}. We modify this contract into a constant-bonus contract ^𝐂L=(TL,W^0L,b^L)\widehat{}\mathbf{C}^{L}=(T^{L},\widehat{W}^{L}_{0},\widehat{b}^{L}) where b^L=b¯L(TL)\widehat{b}^{L}=\overline{b}^{L}(T^{L}) and the modified initial transfer W^0L\widehat{W}_{0}^{L} is such that U0L(𝐂L,𝟏)=U0L(^𝐂L,𝟏)U^{L}_{0}(\mathbf{C}^{L},\mathbf{1})=U^{L}_{0}(\widehat{}\mathbf{C}^{L},% \mathbf{1}). We can show that this modification relaxes (Weak-ICHL) while keeping all other constraints in [RPℓℓ] unchanged, and thus it allows to weakly increase the objective in [RPℓℓ]. We omit the details as the arguments are analogous to those in Step 4 in the proof of Theorem 3.

Step 6: Under-experimentation and positive rents for both types

We first show that the solution to [RPℓℓ] does not induce over-experimentation by either type: TLtLT^{L}\leq t^{L} and THtHT^{H}\leq t^{H}. It is useful for our arguments to rewrite the principal’s payoff by substituting with (1); we obtain that the objective in [RPℓℓ] can be rewritten as

μ0{β0t=1TH(1λH)t1λH(1btH)W0H}+(1μ0){β0t=1TL(1λL)t1λL(1btL)W0L}.\mu_{0}\left\{\beta_{0}\sum\limits_{t=1}^{T^{H}}\left(1-\lambda^{H}\right)^{t-% 1}\lambda^{H}\left(1-b_{t}^{H}\right)-W^{H}_{0}\right\}+\left(1-\mu_{0}\right)% \left\{\beta_{0}\sum\limits_{t=1}^{T^{L}}\left(1-\lambda^{L}\right)^{t-1}% \lambda^{L}\left(1-b^{L}_{t}\right)-W_{0}^{L}\right\}. (D.14)

Suppose per contra that a solution to [RPℓℓ] has a menu of connected bonus contracts (𝐂L,𝐂H)(\mathbf{C}^{L},\mathbf{C}^{H}) such that Tθ>tθT^{\theta}>t^{\theta} for some θ{L,H}\theta\in\{L,H\}. Without loss by Step 2, 𝐂θ=(Tθ,W0θ,𝒃θ)\mathbf{C}^{\theta}=(T^{\theta},W^{\theta}_{0},\bm{b}^{\theta}) induces type θ\theta to work in each period t{1,,Tθ}t\in\{1,\ldots,T^{\theta}\}. Note that by the arguments in Step 5, type θ\theta’s incentive constraint for effort binds in each period of contract 𝐂θ\mathbf{C}^{\theta} if and only if btθ=cβ¯Tθθλθb^{\theta}_{t}=\frac{c}{\overline{\beta}^{\theta}_{T^{\theta}}\lambda^{\theta}} for all t{1,,Tθ}t\in\{1,\ldots,T^{\theta}\}; hence, contract 𝐂θ\mathbf{C}^{\theta} must have btθcβ¯Tθθλθb^{\theta}_{t}\geq\frac{c}{\overline{\beta}^{\theta}_{T^{\theta}}\lambda^{% \theta}} for all t{1,,Tθ}t\in\{1,\ldots,T^{\theta}\} and Tθ>tθT^{\theta}>t^{\theta} implies btθ>1b^{\theta}_{t}>1 for all t{1,,Tθ}t\in\{1,\ldots,T^{\theta}\}. Using (D.14), this implies that the principal’s payoff from type θ\theta is strictly negative if Tθ>tθT^{\theta}>t^{\theta}. But then we can show that there exists a menu of connected bonus contracts that satisfies all the constraints in [RPℓℓ] and yields the principal a strictly larger payoff than the original menu (𝐂L,𝐂H)(\mathbf{C}^{L},\mathbf{C}^{H}). This is immediate if the original menu induces both TL>tLT^{L}>t^{L} and TH>tHT^{H}>t^{H}, as the principal gets a strictly negative payoff from each type in this case. Suppose instead that the original menu is (𝐂θ,𝐂θ)(\mathbf{C}^{\theta},\mathbf{C}^{\theta^{\prime}}) with TθtθT^{\theta}\leq t^{\theta} for type θ{L,H}\theta\in\{L,H\} and Tθ>tθT^{\theta^{\prime}}>t^{\theta^{\prime}} for θθ\theta^{\prime}\neq\theta. Then consider a menu (^𝐂θ,^𝐂θ)(\widehat{}\mathbf{C}^{\theta},\widehat{}\mathbf{C}^{\theta^{\prime}}) where ^𝐂θ=^𝐂θ=(Tθ,0,b¯θ(Tθ))\widehat{}\mathbf{C}^{\theta}=\widehat{}\mathbf{C}^{\theta^{\prime}}=(T^{% \theta},0,\overline{b}^{\theta}(T^{\theta})). This menu trivially satisfies all the constraints in the principal’s program. Moreover, compared to the original menu, this menu yields the principal a weakly larger payoff from type θ\theta because it induces this type to work for the same periods as 𝐂θ\mathbf{C}^{\theta} with a (weakly) lower initial transfer and (weakly) lower bonuses in each period t{1,,Tθ}t\in\{1,\ldots,T^{\theta}\}, and it yields the principal a strictly larger payoff from type θ\theta^{\prime} because the payoff from this type under the new menu is non-negative given that the bonus is b¯θ(Tθ)1\overline{b}^{\theta}(T^{\theta})\leq 1 in each period t{1,,Tθ}t\in\{1,\ldots,T^{\theta}\}.

Next, we show that the solution to [RPℓℓ] yields a positive rent to type HH (i.e. U0H(𝐂H,𝟏)>0U^{H}_{0}(\mathbf{C}^{H},\mathbf{1})>0) and it also yields a positive rent to type LL (i.e. U0L(𝐂L,𝟏)>0U^{L}_{0}(\mathbf{C}^{L},\mathbf{1})>0) if type LL is not excluded. By the limited liability constraints (\ell\ellL) and (\ell\ellH), U0L(𝐂L,𝟏)0U^{L}_{0}(\mathbf{C}^{L},\mathbf{1})\geq 0 and U0H(𝐂H,𝟏)0U^{H}_{0}(\mathbf{C}^{H},\mathbf{1})\geq 0. Moreover, given limited liability, U0θ(𝐂θ,𝟏)=0U^{\theta}_{0}(\mathbf{C}^{\theta},\mathbf{1})=0 for a type θ{L,H}\theta\in\{L,H\} implies Tθ=0T^{\theta}=0. Hence, if type θ\theta is not excluded, this type receives a strictly positive rent. All that is left to be shown is that the solution to [RPℓℓ] cannot exclude type HH, and thus it always yields U0H(𝐂H,𝟏)>0U^{H}_{0}(\mathbf{C}^{H},\mathbf{1})>0. First, suppose that U0L(𝐂L,𝟏)>0U^{L}_{0}(\mathbf{C}^{L},\mathbf{1})>0 and U0H(𝐂H,𝟏)=0U^{H}_{0}(\mathbf{C}^{H},\mathbf{1})=0. Then since β¯tHλH>β¯tLλL\overline{\beta}^{H}_{t}\lambda^{H}>\overline{\beta}^{L}_{t}\lambda^{L} for all ttLt\leq t^{L} (by the assumption that tH>tLt^{H}>t^{L}) and TLtLT^{L}\leq t^{L}, it follows that U0H(𝐂L,𝟏)>U0L(𝐂L,𝟏)>0=U0H(𝐂H,𝟏)U^{H}_{0}(\mathbf{C}^{L},\mathbf{1})>U^{L}_{0}(\mathbf{C}^{L},\mathbf{1})>0=U^% {H}_{0}(\mathbf{C}^{H},\mathbf{1}), and thus (Weak-ICHL) is violated. Next, suppose that U0θ(𝐂θ,𝟏)=0U^{\theta}_{0}(\mathbf{C}^{\theta},\mathbf{1})=0 for both types θ{L,H}\theta\in\{L,H\}. Then Tθ=0T^{\theta}=0 for both types θ{L,H}\theta\in\{L,H\} and the principal’s payoff is zero. However, the principal can then strictly improve upon this menu by using a menu of constant-bonus contracts ^𝐂L=^𝐂H=(1,0,b¯H(1))\widehat{}\mathbf{C}^{L}=\widehat{}\mathbf{C}^{H}=(1,0,\overline{b}^{H}(1)), where note that b¯H(1)<1\overline{b}^{H}(1)<1.

Step 7: The high type experiments more than the low type

We show that the solution to [RPℓℓ] must have TLTHT^{L}\leq T^{H}. Suppose per contra that the solution is a menu of connected bonus contracts {𝐂L,𝐂H}\{\mathbf{C}^{L},\mathbf{C}^{H}\} such that TL>THT^{L}>T^{H}. Without loss by Step 5, let 𝐂L=(TL,W0L,b¯L(TL))\mathbf{C}^{L}=(T^{L},W^{L}_{0},\overline{b}^{L}(T^{L})). Note that by (Weak-ICHL), U0H(𝐂H,𝟏)U0H(𝐂L,𝟏)U^{H}_{0}(\mathbf{C}^{H},\mathbf{1})\geq U^{H}_{0}(\mathbf{C}^{L},\mathbf{1}). Moreover, by Step 6, TLtLT^{L}\leq t^{L}, which in turn implies TL<tHT^{L}<t^{H}. But then it is immediate that a menu (~𝐂L,~𝐂H)(\widetilde{}\mathbf{C}^{L},\widetilde{}\mathbf{C}^{H}) where ~𝐂L=~𝐂H=(TL,0,b¯L(TL))\widetilde{}\mathbf{C}^{L}=\widetilde{}\mathbf{C}^{H}=(T^{L},0,\overline{b}^{L% }(T^{L})) yields the same amount of experimentation by type LL, strictly more efficient experimentation by type HH, and payoffs U0L(~𝐂L,𝟏)U0L(𝐂L,𝟏)U^{L}_{0}(\widetilde{}\mathbf{C}^{L},\mathbf{1})\leq U^{L}_{0}(\mathbf{C}^{L},% \mathbf{1}) and U0H(~𝐂H,𝟏)U0H(𝐂H,𝟏)U^{H}_{0}(\widetilde{}\mathbf{C}^{H},\mathbf{1})\leq U^{H}_{0}(\mathbf{C}^{H},% \mathbf{1}), while satisfying all the constraints in [RPℓℓ]. It follows that (~𝐂L,~𝐂H)(\widetilde{}\mathbf{C}^{L},\widetilde{}\mathbf{C}^{H}) yields a strictly larger payoff to the principal than the original menu (𝐂L,𝐂H)(\mathbf{C}^{L},\mathbf{C}^{H}), which therefore cannot be optimal.

Step 8: Back to the original problem

We now show that the solution to the relaxed program [RPℓℓ] is feasible and thus a solution to the original program [Pℓℓ]. Recall that (given Steps 1-3) the only relaxation in program [RPℓℓ] relative to [Pℓℓ] is that [RPℓℓ] imposes (Weak-ICHL) instead of (ICHL). Thus, all we need to show is that given a constant-bonus contract 𝐂L=(TL,W0L,b¯L(TL))\mathbf{C}^{L}=(T^{L},W^{L}_{0},\overline{b}^{L}(T^{L})) with length TLtLT^{L}\leq t^{L}, it would be optimal for type HH to work in each period 1,,TL1,\ldots,T^{L}. The claim follows from Step 6 in the proof of Theorem 3 and the proof of Proposition 1. ∎

D.6 Details for Subsection 7.3

Here we provide details for the discussion in Subsection 7.3 of the paper.

Assume β0=1\beta_{0}=1 and for simplicity that there is some finite time, T¯\overline{T}, at which the game ends. Since β¯tθ=1\overline{\beta}^{\theta}_{t}=1 for all θ{L,H}\theta\in\{L,H\} and t{1,,T¯}t\in\{1,\ldots,\overline{T}\}, the high type always has a higher expected marginal product than the low type, i.e. β¯tHλH=λH>β¯tLλL=λL\overline{\beta}^{H}_{t}\lambda^{H}=\lambda_{H}>\overline{\beta}^{L}_{t}% \lambda^{L}=\lambda_{L} for all tt. Consequently, the methodology used in proving Theorem 3 can be applied, with the conclusions that if the optimal length of experimentation for the low type is some TT (constrained to be no larger than T¯\overline{T}), the optimal penalty contract for the low type is given by the analog of (6) with β¯tL=1\overline{\beta}^{L}_{t}=1 for all tt:

ltL={(1δ)cλL if t<T,cλL if t=T,{l}_{t}^{L}=\begin{cases}-\left(1-\delta\right)\frac{c}{\lambda^{L}}&\text{ if% }t<T,\\ -\frac{c}{\lambda^{L}}&\text{ if }t=T,\end{cases}

and the portion of the principal’s payoff that depends on TT is given by the analog of (D.2) with the simplification of β0=1\beta_{0}=1:

V^(T)\displaystyle\widehat{V}(T) =\displaystyle= (1μ0)t=1Tδt(1λL)t1(λLc)\displaystyle\left(1-\mu_{0}\right)\sum\limits_{t=1}^{T}\delta^{t}\left(1-% \lambda^{L}\right)^{t-1}\left(\lambda^{L}-c\right)
μ0{cλLt=1T1δt(1δ)[(1λH)t(1λL)t]cλLδT[(1λH)T(1λL)T]t=1Tδtc[(1λH)t1(1λL)t1]}.\displaystyle-\mu_{0}\left\{\begin{array}[]{l}-\frac{c}{\lambda^{L}}\sum% \limits_{t=1}^{T-1}\delta^{t}\left(1-\delta\right)\left[\left(1-\lambda^{H}% \right)^{t}-\left(1-\lambda^{L}\right)^{t}\right]-\frac{c}{\lambda^{L}}\delta^% {T}\left[\left(1-\lambda^{H}\right)^{T}-\left(1-\lambda^{L}\right)^{T}\right]% \\ -\sum\limits_{t=1}^{T}\delta^{t}c\left[\left(1-\lambda^{H}\right)^{t-1}-\left(% 1-\lambda^{L}\right)^{t-1}\right]\end{array}\right\}.

Hence, for any T{0,,T¯1}T\in\{0,\ldots,\overline{T}-1\} we have the following analog of (D.2):

V^(T+1)V^(T)=δT+1[(1μ0)(1λL)T(λLc)μ0cλL(1λH)T(λHλL)].\widehat{V}(T+1)-\widehat{V}(T)=\delta^{T+1}\left[\left(1-\mu_{0}\right)\left(% 1-\lambda^{L}\right)^{T}\left(\lambda^{L}-c\right)-\mu_{0}\frac{c}{\lambda^{L}% }\left(1-\lambda^{H}\right)^{T}\left(\lambda^{H}-\lambda^{L}\right)\right].

Clearly, V^(T+1)V^(T)>(<)0\widehat{V}(T+1)-\widehat{V}(T)>(<)0 if and only if

(1λL1λH)T>(<)μ0c(λHλL)(1μ0)(λLc)λL.\left(\frac{1-\lambda^{L}}{1-\lambda^{H}}\right)^{T}>(<)\frac{\mu_{0}c\left(% \lambda^{H}-\lambda^{L}\right)}{\left(1-\mu_{0}\right)\left(\lambda^{L}-c% \right)\lambda^{L}}.

Since the left-hand side above is strictly increasing in TT, it follows that V^(T)\widehat{V}(T) is maximized by t¯L{0,T¯}\overline{t}^{L}\in\{0,\overline{T}\}. Hence, whenever it is optimal to have the low type experiment for any positive amount of time, it is optimal to have the low type experiment until T¯\overline{T}, no matter the value of T¯\overline{T}. Note that whenever exclusion is optimal (i.e. t¯L=0\overline{t}^{L}=0) when β0=1\beta_{0}=1, it would also be optimal for all β01\beta_{0}\leq 1; this follows from the comparative static of t¯L\overline{t}^{L} with respect to β0\beta_{0} in Proposition 2.

HTML from LaTeXML, with custom CSS/JS. The PDF is more accurate.