Optimal Portfolio and Consumption in Continuous Time
Introduction
The intertemporal portfolio choice notebook derived optimal consumption and portfolio rules in discrete time using the Bellman equation. Continuous time offers a cleaner framework: replacing the discrete recursion with a differential equation, the Hamilton-Jacobi-Bellman (HJB) equation, allows us to characterize optimal policies in closed form and to connect the value function directly to the stochastic discount factor (SDF) and equilibrium risk premia.
This notebook solves the consumption-portfolio problem under additive utility: the investor maximizes the expected discounted integral of a felicity function u(c). The key results are Merton’s portfolio separation theorem (Merton 1969, 1971), which decomposes the optimal portfolio into a myopic and a hedging component, and his Intertemporal CAPM (ICAPM) (Merton 1973), which links equilibrium risk premia to covariances with both aggregate wealth and state variables that shift the investment opportunity set.
The analysis proceeds as follows. We first set up the optimization problem and write down the wealth dynamics. The Bellman principle then delivers the HJB equation, whose first-order conditions characterize optimal consumption and portfolio choice. The consumption first-order condition identifies the marginal value of wealth with the marginal utility of consumption, and the envelope condition shows that this marginal value, suitably discounted, is an SDF. Substituting the SDF into the fundamental pricing equation yields the ICAPM. Throughout, the investor takes prices as given, so the analysis is partial equilibrium.
The next notebook specializes the additive-utility problem to a one-factor Gaussian state model and solves explicitly for the value function, the wealth-consumption ratio, and the portfolio rule. The consumption-based asset pricing notebook then moves to general equilibrium with a representative investor. The subsequent recursive utility notebook then replaces additive utility with Epstein-Zin preferences, where a modified HJB equation introduces an additional pricing factor that captures the investor’s concern for news about future investment opportunities. The structure developed here (the HJB, the envelope condition, and the market price of risk) carries over directly.
The Optimization Problem
An investor holds a portfolio of N risky assets and a money-market account. As in Discount Factors in Continuous Time, we track each asset through its dividend-reinvested price P_i, so that dP_i/P_i is the asset’s total return, capital gain plus dividend yield. The vector of total returns follows \frac{d\mathbf{P}}{\mathbf{P}} = \pmb{\mu}(\mathbf{z}) \, dt + \pmb{\sigma}(\mathbf{z}) \, d\mathbf{B}, where the division is componentwise, \pmb{\mu}(\mathbf{z}) is the N \times 1 vector of expected total returns (the \mu_i + q_i of the SDF notebook), \pmb{\sigma}(\mathbf{z}) is an N \times K volatility matrix, and \mathbf{B} is a K-dimensional standard Brownian motion. The risk-free rate is r(\mathbf{z}). All coefficients depend on a vector of state variables \mathbf{z} \in \mathbb{R}^L that summarizes the investment opportunity set and evolves as d\mathbf{z} = \pmb{\mu}^z(\mathbf{z}) \, dt + \pmb{\sigma}^z(\mathbf{z}) \, d\mathbf{B}, \tag{1} where \pmb{\sigma}^z(\mathbf{z}) is an L \times K matrix. The same Brownian motion \mathbf{B} drives both asset returns and state variable innovations, capturing the covariation between portfolio returns and shifts in investment opportunities.
Let \pmb{\alpha}_t denote the N \times 1 vector of portfolio weights in risky assets. The wealth process satisfies \frac{dW}{W} = \left( r(\mathbf{z}) + \pmb{\alpha}'\!\left(\pmb{\mu}(\mathbf{z}) - r(\mathbf{z})\pmb{\iota}\right) - \frac{c}{W} \right) dt + \pmb{\alpha}'\pmb{\sigma}(\mathbf{z}) \, d\mathbf{B}, \tag{2} where c \geq 0 is the consumption rate and \pmb{\iota} is the N \times 1 vector of ones. The drift is the expected portfolio return net of consumption, and the diffusion term carries portfolio-return risk.
The investor maximizes expected discounted utility: \max_{\{c_s,\, \pmb{\alpha}_s\}_{s \geq t}} \operatorname{E}_t\!\left[ \int_t^\infty e^{-\delta(s-t)} u(c_s) \, ds \right], \tag{3} subject to (2). We assume u' > 0, u'' < 0, and Inada conditions u'(0) = \infty, u'(\infty) = 0, which guarantee interior optimal consumption.
The Hamilton-Jacobi-Bellman Equation
The value function V(W, \mathbf{z}) gives the highest lifetime utility achievable from state (W, \mathbf{z}): V(W, \mathbf{z}) = \max_{\{c_s,\, \pmb{\alpha}_s\}_{s \geq t}} \operatorname{E}_t\!\left[ \int_t^\infty e^{-\delta(s-t)} u(c_s) \, ds \right].
To derive its characterizing equation, apply the Bellman principle over an interval of length dt: V(W, \mathbf{z}) = \max_{c,\, \pmb{\alpha}} \left\{ u(c) \, dt + e^{-\delta\, dt} \, \operatorname{E}_t\!\left[ V(W + dW,\, \mathbf{z} + d\mathbf{z}) \right] \right\}. Expanding e^{-\delta dt} \approx 1 - \delta \, dt and applying Ito’s lemma to V(W + dW, \mathbf{z} + d\mathbf{z}), taking expectations (so that the d\mathbf{B} terms vanish), collecting terms of order dt, and rearranging yields the Hamilton-Jacobi-Bellman (HJB) equation: \delta V = \max_{c,\, \pmb{\alpha}} \left\{ u(c) + \mathcal{L}^{c,\pmb{\alpha}} V \right\}, \tag{4} where \mathcal{L}^{c,\pmb{\alpha}} is the Ito generator of the joint process (W_t, \mathbf{z}_t) under controls (c, \pmb{\alpha}): \begin{aligned} \mathcal{L}^{c,\pmb{\alpha}} V &= V_W \!\left[ W\!\left(r + \pmb{\alpha}'(\pmb{\mu} - r\pmb{\iota})\right) - c \right] + \tfrac{1}{2} V_{WW} W^2 \pmb{\alpha}'\pmb{\sigma}\pmb{\sigma}'\pmb{\alpha} \\ &\quad + (\nabla_z V)'\pmb{\mu}^z + \tfrac{1}{2} \operatorname{tr}\!\left(\pmb{\sigma}^z (\pmb{\sigma}^z)' H_z V\right) + W (\nabla_z V_W)'\pmb{\sigma}^z\pmb{\sigma}'\pmb{\alpha}. \end{aligned} \tag{5} Here V_W = \partial V/\partial W, V_{WW} = \partial^2 V / \partial W^2, \nabla_z V is the gradient of V with respect to \mathbf{z}, \nabla_z V_W is the gradient of the marginal value of wealth with respect to \mathbf{z}, and H_z V is the Hessian of V with respect to \mathbf{z}. The last term in (5) captures the Ito cross-variation between wealth and state variables: since dW and d\mathbf{z} share the same Brownian driver, they are instantaneously correlated.
The HJB equation (4) is the continuous-time counterpart of the discrete Bellman equation. The derivation above is heuristic; Duffie (2010) gives the verification conditions under which a solution of the HJB equation is the value function. The left-hand side \delta V is the required return on the value function, the rate the investor demands to postpone utility, and the right-hand side is the maximum flow of utility plus the expected capital gain on V per unit time. At the optimum, the two are equal.
First-Order Conditions
The HJB equation (4) is solved by maximizing the right-hand side over (c, \pmb{\alpha}) pointwise at each state (W, \mathbf{z}). We assume V is strictly concave in wealth, V_{WW} < 0, so that the first-order conditions below characterize a maximum.
Consumption. Differentiating with respect to c: u'(c^*) = V_W(W, \mathbf{z}). \tag{6} The optimal consumption rate equates the marginal utility of consuming today to the shadow value V_W of an extra unit of wealth. At the optimum, a unit of wealth is worth the same whether it is consumed or invested.
Portfolio. Differentiating the right-hand side of (4) with respect to \pmb{\alpha} and setting equal to zero: V_W W (\pmb{\mu} - r\pmb{\iota}) + V_{WW} W^2 \pmb{\sigma}\pmb{\sigma}' \pmb{\alpha}^* + W \pmb{\sigma}(\pmb{\sigma}^z)' \nabla_z V_W = \mathbf{0}. \tag{7} The three terms represent the marginal benefit of tilting toward higher-expected-return assets (first term), the variance cost of bearing more return risk (second term), and the contribution of portfolio choice to the covariance between wealth and state variable innovations (third term).
Merton’s Portfolio Separation
Let \text{rra} = -{W V_{WW}}/{V_W} > 0 denote the coefficient of relative risk aversion implied by the value function. Substituting V_{WW} = -\text{rra}\, V_W / W into (7) and dividing by WV_W gives: (\pmb{\mu} - r\pmb{\iota}) - \text{rra}\,\pmb{\sigma}\pmb{\sigma}'\pmb{\alpha}^* + \pmb{\sigma}(\pmb{\sigma}^z)' \frac{\nabla_z V_W}{V_W} = \mathbf{0}.
Assuming \pmb{\sigma}\pmb{\sigma}' is invertible, the optimal risky-asset portfolio is:
Property 1 (Merton’s Portfolio Separation) \pmb{\alpha}^* = \underbrace{\frac{1}{\text{rra}} (\pmb{\sigma}\pmb{\sigma}')^{-1}(\pmb{\mu} - r\pmb{\iota})}_{\text{myopic demand}} + \underbrace{\frac{1}{\text{rra}} (\pmb{\sigma}\pmb{\sigma}')^{-1} \pmb{\sigma}(\pmb{\sigma}^z)' \frac{\nabla_z V_W}{V_W}}_{\text{hedging demand}}. \tag{8}
The decomposition in (8) separates portfolio choice into two economically distinct motives:
Myopic demand. The first term is proportional to the risky portfolio with the highest instantaneous Sharpe ratio, scaled by the inverse of risk aversion. It is identical to the portfolio a one-period investor would choose and depends only on the current levels of (\pmb{\mu}, \pmb{\sigma}, r), not on their dynamics.
Hedging demand. The second term arises from the investor’s desire to hedge against future shifts in investment opportunities. The matrix (\pmb{\sigma}\pmb{\sigma}')^{-1}\pmb{\sigma}(\pmb{\sigma}^z)' identifies the portfolios that best span the Brownian innovations driving \mathbf{z}. The vector \nabla_z V_W / V_W weights each state variable by the semi-elasticity \partial \ln V_W / \partial z_k of the marginal value of wealth. When \partial^2 V / \partial W \partial z_k > 0, a positive shock to z_k raises V_W, so the investor overweights assets correlated with z_k in order to have more wealth in the states where wealth is most valuable.
Whether a high V_W signals worse or better investment opportunities depends on risk aversion. For \text{rra} > 1 the investor values wealth most when opportunities deteriorate, and hedging means buying assets that pay off in those states. For \text{rra} < 1 the direction reverses, as the one-factor sequel shows explicitly. Campbell and Viceira (1999) show that these hedging demands can be large when expected stock returns are predictable.
The hedging demand vanishes when \nabla_z V_W = \mathbf{0}. This happens when the investment opportunity set is constant (Merton 1969; Samuelson 1969), and also for log utility even when opportunities vary, since then V_W = 1/(\delta W) does not depend on \mathbf{z}. In either case all investors hold the same risky portfolio, scaled by 1/\text{rra}. This is the two-fund separation result of the static CAPM: investors choose between the risk-free asset and a single risky portfolio.
Example 1 Take the two stocks of the SDF notebook, with exposure vectors \pmb{\sigma}_{1} = (0.12,\ 0.16)' and \pmb{\sigma}_{2} = (0.20,\ -0.15)' and risk premiums 0.048 and 0.08. The stocks are uncorrelated, so \pmb{\sigma}\pmb{\sigma}' = \operatorname{diag}(0.04,\ 0.0625). Suppose \text{rra} = 2 and there is one state variable that loads only on the second shock, \pmb{\sigma}^z = (0,\ 0.1), with \partial \ln V_W / \partial z = -1, so wealth is most valuable when z is low.
The myopic demand is \frac{1}{2}\begin{pmatrix} 0.048 / 0.04 \\ 0.08 / 0.0625 \end{pmatrix} = \begin{pmatrix} 0.60 \\ 0.64 \end{pmatrix}. The covariances of the two stocks with the state variable are \pmb{\sigma}(\pmb{\sigma}^z)' = (0.016,\ -0.015)', so the hedging demand is \frac{1}{2}\begin{pmatrix} 0.016 / 0.04 \\ -0.015 / 0.0625 \end{pmatrix}(-1) = \begin{pmatrix} -0.20 \\ 0.12 \end{pmatrix}, and \pmb{\alpha}^* = (0.40,\ 0.76)'. The second stock pays off when B_{2} falls, which is when z falls and wealth is most valuable, so the investor holds more of it. The first stock pays off when z rises, so the investor holds less of it than the myopic demand.
The Stochastic Discount Factor
The consumption FOC (6) establishes u'(c^*) = V_W. Under additive utility, the natural candidate for the SDF is therefore \Lambda_t = e^{-\delta t} V_W(W_t, \mathbf{z}_t) = e^{-\delta t} u'(c_t^*), \tag{9} which is the continuous-time analogue of m_{t+1} = \beta u'(c_{t+1})/u'(c_t) in discrete time.
To confirm that (9) is an SDF, we first check the no-arbitrage condition \operatorname{E}_t(d\Lambda/\Lambda) = -r\,dt. This follows from the envelope condition: differentiate the HJB equation (4) with respect to W at the optimum. By the envelope theorem, the optimal controls can be held fixed, so \begin{aligned} \delta V_W &= \mathcal{L}^{c^*,\pmb{\alpha}^*} V_W + V_W \bigl(r + \pmb{\alpha}^{*\prime}(\pmb{\mu} - r\pmb{\iota})\bigr) + V_{WW} W \pmb{\alpha}^{*\prime}\pmb{\sigma}\pmb{\sigma}'\pmb{\alpha}^* + (\nabla_z V_W)'\pmb{\sigma}^z\pmb{\sigma}'\pmb{\alpha}^* \\ &= \mathcal{L}^{c^*,\pmb{\alpha}^*} V_W + r V_W + \frac{1}{W}\pmb{\alpha}^{*\prime}\Bigl[V_W W (\pmb{\mu} - r\pmb{\iota}) + V_{WW} W^2 \pmb{\sigma}\pmb{\sigma}' \pmb{\alpha}^* + W \pmb{\sigma}(\pmb{\sigma}^z)' \nabla_z V_W\Bigr], \end{aligned} where \mathcal{L}^{c^*,\pmb{\alpha}^*} V_W is the Ito drift of V_W(W_t, \mathbf{z}_t) and the extra terms come from differentiating the coefficients of (5) that depend on W. The bracket is the portfolio FOC (7) and equals zero. Hence the drift of V_W is (\delta - r)V_W, and the drift of \Lambda_t = e^{-\delta t} V_W is -\delta + (\delta - r) = -r per unit of time.
The diffusion component of d\Lambda/\Lambda determines the market price of risk. Applying Ito’s lemma to \Lambda = e^{-\delta t} V_W, the d\mathbf{B} part is: \frac{d\Lambda}{\Lambda}\bigg|_{d\mathbf{B}} = \frac{V_{WW}}{V_W} dW\big|_{d\mathbf{B}} + \frac{(\nabla_z V_W)'}{V_W} d\mathbf{z}\big|_{d\mathbf{B}} = \left( -\text{rra}\,\pmb{\alpha}^{*\prime}\pmb{\sigma} + \frac{(\nabla_z V_W)'}{V_W}\pmb{\sigma}^z \right) d\mathbf{B}. Comparing with the generic SDF form d\Lambda/\Lambda = -r\,dt - \pmb{\lambda}' d\mathbf{B} from Discount Factors in Continuous Time, the market price of risk vector is \pmb{\lambda}' = \text{rra}\,\pmb{\alpha}^{*\prime}\pmb{\sigma} - \frac{(\nabla_z V_W)'}{V_W}\pmb{\sigma}^z. \tag{10} The first term is risk aversion times the exposure of the investor’s wealth to each Brownian shock. The second term prices the exposure of the state variables to each shock, weighted by how strongly each state variable moves the marginal value of wealth.
Substituting the optimal portfolio (8) into (10) shows which of the many admissible SDFs of the SDF notebook the investor’s marginal utility selects: \pmb{\lambda} = \underbrace{\pmb{\sigma}'(\pmb{\sigma}\pmb{\sigma}')^{-1}(\pmb{\mu} - r\pmb{\iota})}_{\pmb{\lambda}_{\min}} \;-\; \underbrace{\bigl[\mathbf{I} - \pmb{\sigma}'(\pmb{\sigma}\pmb{\sigma}')^{-1}\pmb{\sigma}\bigr](\pmb{\sigma}^z)'\frac{\nabla_z V_W}{V_W}}_{\in\ \mathrm{Null}(\pmb{\sigma})}. \tag{11} The first term is the minimum-norm market price of risk. The second term lies in the null space of \pmb{\sigma}, since \pmb{\sigma}\bigl[\mathbf{I} - \pmb{\sigma}'(\pmb{\sigma}\pmb{\sigma}')^{-1}\pmb{\sigma}\bigr] = \mathbf{0}: it is the part of state-variable risk that no portfolio of traded assets can hedge. The investor’s SDF therefore prices exactly that unhedgeable state risk on top of \pmb{\lambda}_{\min}. When markets are complete (N = K), the null space is trivial and \pmb{\lambda} = \pmb{\lambda}_{\min}, whatever the state variables do. He and Pearson (1991) and Karatzas et al. (1991) derive this selection from a dual problem, and Cox and Huang (1989) solve the complete-markets case directly from the unique SDF, without the HJB equation.
Example 2 In Example 1 the market is complete, so (10) must return the unique market price of risk \pmb{\lambda} = (0.4,\ 0)' of the SDF notebook. The exposure of wealth to the two shocks is \pmb{\sigma}'\pmb{\alpha}^* = 0.40\begin{pmatrix} 0.12 \\ 0.16 \end{pmatrix} + 0.76\begin{pmatrix} 0.20 \\ -0.15 \end{pmatrix} = \begin{pmatrix} 0.20 \\ -0.05 \end{pmatrix}, and therefore \pmb{\lambda} = 2\begin{pmatrix} 0.20 \\ -0.05 \end{pmatrix} - (-1)\begin{pmatrix} 0 \\ 0.1 \end{pmatrix} = \begin{pmatrix} 0.4 \\ 0 \end{pmatrix}. The investor’s wealth is exposed to B_{2}, and so is the state variable, but the two exposures offset exactly in marginal utility, so B_{2} remains unpriced.
The Intertemporal CAPM
The fundamental pricing equation of the SDF notebook requires that for any risky asset S paying a dividend yield D/S, \operatorname{E}\!\left(\frac{dS}{S}\right) + \frac{D}{S}\,dt - r\,dt = -\frac{d\Lambda}{\Lambda}\,\frac{dS}{S}. \tag{12} Substituting the SDF dynamics from (10) into the right-hand side, and using (d\mathbf{B})(d\mathbf{B})' = \mathbf{I}\,dt to express \pmb{\alpha}^{*\prime}\pmb{\sigma}\,d\mathbf{B} and \pmb{\sigma}^z d\mathbf{B} as the diffusion parts of dW/W and d\mathbf{z}, yields Merton’s Intertemporal CAPM (Merton 1973).
Property 2 (Merton’s Intertemporal CAPM) The risk premium of any risky asset satisfies \operatorname{E}\left(\frac{dS}{S}\right) + \frac{D}{S} dt - r \, dt = \text{rra} \cdot \frac{dW}{W} \frac{dS}{S} - \frac{(\nabla_z V_W)'}{V_W} \left(d\mathbf{z} \, \frac{dS}{S}\right). \tag{13} Expected excess returns depend on L + 1 factors: the wealth portfolio and L state-variable hedging portfolios, one for each source of time-variation in the investment opportunity set.
Equation (13) holds for every investor who optimizes, with W that investor’s wealth. To turn it into a statement about market aggregates, assume a representative investor, as in the general equilibrium model of Cox et al. (1985), so that W is aggregate wealth and c^* is aggregate consumption. The risk premium then has two economically distinct components:
Wealth risk premium. The term \text{rra} \cdot (dW/W)(dS/S)/dt is the instantaneous covariance of asset returns with wealth growth, scaled by the coefficient of relative risk aversion. Assets that covary positively with aggregate wealth are risky and command higher expected returns.
Hedging demand premium. The term -(\nabla_z V_W)'(d\mathbf{z})(dS/S)/(V_W\,dt) reflects investors’ desire to hedge against shifts in the investment opportunity set. An asset that covaries positively with a state variable z_k for which V_{Wz_k} > 0 pays off in states where wealth is most valuable, acts as a hedge, and therefore commands a lower risk premium.
When there are no state variables, or whenever V_{Wz_k} = 0 for all k, the ICAPM reduces to a single-factor model in which the market (wealth) portfolio is the only priced risk. In that special case, equation (13) recovers the continuous-time CAPM: risk premiums are proportional to covariance with aggregate wealth.
The ICAPM is a statement about any investor who optimizes at given prices. The Consumption-Based Asset Pricing notebook takes the equilibrium view instead: with a representative investor, consumption is aggregate consumption, the ICAPM factors collapse into a single consumption factor (Breeden 1979), and prices are determined by what makes the investor willing to consume the aggregate endowment.