Systems · AI Systems · Core
Policy Gradient
Policy gradient methods optimize a parameterized policy π_θ directly by ascending the gradient of expected return, rather than via an action-value table.
Definition
Policy gradient methods optimize a parameterized policy π_θ directly by ascending the gradient of expected return, rather than via an action-value table.
RelatedProximal Policy OptimizationActivation FunctionAdam OptimizerAgent WorkflowAttention MechanismAutoencoderBackpropagationBatch NormalizationBeam SearchBERTByte Pair EncodingChain of ThoughtCLIPContrastive Learning