Material adapted from Sergey Levine’s CS294-112 Fall 2018 and Josh Achiam’s Spinning Up With RL
Currently expanding unfinished article into upcoming research paper. Stay tuned!
What is a Policy Gradient and how it is useful?
Arguably the goal in Reinforcement learning is to find an optimal policy. By nature this optimal policy results in relatively highest rewards(relative to other non-optimal policies). With policy gradients, we technically incentivize the distribution of actions to generate a higher reward and likewise deter distributions of actions that generate sub-optimal rewards. Overtime we generate better trajectories, creating the optimal policy.
Derivation
The goal of RL is to optimize the objective function.
definecolorredRGB255,59,48definecolororangeRGB255,149,0definecoloryellowRGB255,204,0definecolorgreenRGB76,217,100definecolortealblueRGB90,200,250definecolorblueRGB0,122,255definecolorpurpleRGB88,86,214definecolorpinkRGB255,45,85
\\pi\_{\\theta}^\\star = \\text{arg}\\underset{\\pi\_{\\theta}}{\\max}\\color{orange}E\_{\\tau\\sim p\_{\\pi\_{\\theta}}(\\tau)}\[\\sum\_{t} r(s\_t,a\_t)\]
What we’re doing here is taking all the state action pairs along the trajectory, (s_t,a_t), summing their rewards up, finding the total reward, and maximizing them with respect to theta, the parameters of the policy.
Let’s denote the expectation of the sum \\color{orange}E\_{\\tau\\sim p\_{\\pi\_{\\theta}}(\\tau)}\[\\sum\_{t} r(s\_t,a\_t)\] as colorredJ(pi_theta).
Now that we have colorredJ(pi_theta) what do we do? Find the gradient to optimize this expectation via gradient ascent.
Expanding J_theta from expectation form (using definition of expectation):
\\color{red}J(\\pi\_{\\theta}) \\color{black}=\\color{orange} E\_{\\tau\\sim p\_{\\pi\_{\\theta}}(\\tau)}\[\\sum\_{t} r(s\_t,a\_t)\] \\color{black}= \\color{green}\\int \\color{blue}P(\\tau|\\pi\_{\\theta})\\color{green}R(\\tau)d\\tau
Finding the gradient:
nabla_thetacolorredJ(pi_theta)colorblack=nabla_thetacolorgreenintcolorblueP(tau∣pi_theta)colorgreenR(tau)dtau
=colorgreenintcolorblacknabla_thetacolorblueP(tau∣pi_theta)colorgreenR(tau)dtau
Using the log derivative trick:
nabla_thetacolorblueP(tau∣pi_theta)colorblack=colorblueP(tau∣pi_theta)fraccolorblacknabla_thetacolorblueP(tau∣pi_theta)P(tau∣pi_theta)=coloryellowP(tau∣pi_theta)colorblacknabla_thetacoloryellowlogP(tau∣pi_theta)
Continuing from our previous step:
=colorgreenintcoloryellowP(tau∣pi_theta)colorblacknabla_thetacolorblacklogP(tau∣pi_theta)colorgreenR(tau)dtau
Going back to expectation form we get:
= \\underset{\\tau\\sim\\pi\_{\\theta}}{E}\[\\nabla\_{\\theta}\\color{yellow}\\log P(\\tau|\\pi\_{\\theta})\\color{green}R(\\tau)\\color{black}\]
But we still don’t have P(tau∣pi_theta) which by the way is the probability of the trajectory
P(tau∣pi_theta)=corr(s_0)prod_t=1TP(s_t+1∣s_t,a_t)pi_theta(a_t∣s_t)
Taking the log of both sides to help us simplify our previous expectation:
logP(tau∣pi_theta)=logcorr(s_o)+sum_t=1T(logP(s_t+1∣s_t,at)+logpi_theta(a_t∣s_t))
nabla_thetalogP(tau∣pi_theta)=nabla_thetalogcorr(s_o)+sum_t=1T(logP(s_t+1∣s_t,at)+logpi_theta(a_t∣s_t))
And since logcorr(s_o) and logpi_theta(a_t∣s_t)) don’t depend on theta, we can simplify the the gradient of the expectation to:
\\nabla\_{\\theta}J(\\pi\_{\\theta}) = \\underset{\\tau\\sim\\pi\_{\\theta}}{E}\[\\cancel{\\nabla\_{\\theta}\\log corr(s\_{o})} + \\nabla\_{\\theta}\\color{yellow} \\sum\_{t=1}^{T} (\\log P(s\_{t+1}|s\_{t},a{t})+ \\log \\pi\_{\\theta}(a\_t|s\_t))\\color{green}R(\\tau)\\color{black}\]
Unfinished derivation (finished the implementation which I will soon update here). Currently working on a paper and will come back to finish this post.