Search

Search

Papers, wiki, Option Blackboard, encyclopedia, and cards.

Results for “policy weights” · papers 18 · wiki 1
Academic Papers · 18arXiv q-fin live 0 · desk corpus 145
arXiv · arXiv · 2025

Cryptocurrency Portfolio Management with Reinforcement Learning: Soft Actor--Critic and Deep Deterministic Policy Gradient Algorithms

This paper proposes a reinforcement learning--based framework for cryptocurrency portfolio management using the Soft Actor--Critic (SAC) and Deep Deterministic Policy Gradient (DDPG) algorithms. Traditional portfolio optimization methods often struggle to adapt to the highly volatile and nonlinear dynamics of cryptocurrency markets. To address this, we design an agent that learns continuous trading actions directly f

Kamal Paykan
arXiv · arXiv · 2025

FR-LUX: Friction-Aware, Regime-Conditioned Policy Optimization for Implementable Portfolio Management

Transaction costs and regime shifts are major reasons why paper portfolios fail in live trading. We introduce FR-LUX (Friction-aware, Regime-conditioned Learning under eXecution costs), a reinforcement learning framework that learns after-cost trading policies and remains robust across volatility-liquidity regimes. FR-LUX integrates three ingredients: (i) a microstructure-consistent execution model combining proporti

Jian'an Zhang
arXiv · arXiv · 2021

A Deep Deterministic Policy Gradient-based Strategy for Stocks Portfolio Management

With the improvement of computer performance and the development of GPU-accelerated technology, trading with machine learning algorithms has attracted the attention of many researchers and practitioners. In this research, we propose a novel portfolio management strategy based on the framework of Deep Deterministic Policy Gradient, a policy-based reinforcement learning framework, and compare its performance to that of

Huanming Zhang, Zhengyong Jiang, Jionglong Su
arXiv · arXiv · 2026

Deep Learning of Robust Market Making under Regime-Switching Order Flow

Classical market-making strategies based on stochastic control, such as the Avellaneda-Stoikov and the Guéant-Lehalle-Fernandez-Tapia (GLFT) extension, provide closed-form quoting rules, but rest on assumptions that break down at realistic microstructure timescales. One of them is that order flow is stationary, while empirical evidence points to the existence of regimes, possibly associated with algorithmic execution

Felipe Moret, Fabrizio Lillo
arXiv · arXiv · 2021

Liquidity Stress Testing in Asset Management -- Part 3. Managing the Asset-Liability Liquidity Risk

This article is part of a comprehensive research project on liquidity risk in asset management, which can be divided into three dimensions. The first dimension covers the modeling of the liability liquidity risk (or funding liquidity), the second dimension is dedicated to the modeling of the asset liquidity risk (or market liquidity), whereas the third dimension considers the management of the asset-liability liquidi

Thierry Roncalli
arXiv · arXiv · 2019

Systemic liquidity contagion in the European interbank market

Systemic liquidity risk, defined by the IMF as "the risk of simultaneous liquidity difficulties at multiple financial institutions", is a key topic in macroprudential policy and financial stress analysis. Specialized models to simulate funding liquidity risk and contagion are available but they require not only banks' bilateral exposures data but also balance sheet data with sufficient granularity, which are hardly a

V. Macchiati, G. Brandi, G. Cimini, G. Caldarelli, D. Paolotti
arXiv · arXiv · 2021

WaveCorr: Correlation-savvy Deep Reinforcement Learning for Portfolio Management

The problem of portfolio management represents an important and challenging class of dynamic decision making problems, where rebalancing decisions need to be made over time with the consideration of many factors such as investors preferences, trading environments, and market conditions. In this paper, we present a new portfolio policy network architecture for deep reinforcement learning (DRL)that can exploit more eff

Saeed Marzban, Erick Delage, Jonathan Yumeng Li, Jeremie Desgagne-Bouchard, Carl Dussault
arXiv · arXiv · 2026

When Does Order Flow Matter? State-Dependent L2 Liquidity-State Transitions in Crypto Futures

Building event-conditioned market models requires separating macro-event labels from persistent microstructure state. We study this distinction in Binance BTCUSDT and ETHUSDT futures from 2023-2026, combining top-20 L2 order book data, trade-flow records, and macro-event windows. We define a supervised discrete L2 liquidity-state transition task, distinct from latent-regime detection and price-direction prediction, a

Joohyoung Jeon
arXiv · arXiv · 2026

Regime-Adaptive Continual Learning for Portfolio Management

Financial markets are inherently non-stationary, exhibiting frequent regime shifts and structural changes that render traditional Portfolio Management (PM) approaches ineffective. Existing remedies, such as rolling-window retraining and naive online fine-tuning, are hindered by high computational costs and insufficient knowledge utilization, respectively, resulting in low returns and limited adaptability. Continual l

Chaofan Pan, Lingfei Ren, Linbo Xiong, Yonghao Li, Wei Wei
arXiv · arXiv · 2025

RL-Exec: Impact-Aware Reinforcement Learning for Opportunistic Optimal Liquidation, Outperforms TWAP and a Book-Liquidity VWAP on BTC-USD Replays

We study opportunistic optimal liquidation over fixed deadlines on BTC-USD limit-order books (LOB). We present RL-Exec, a PPO agent trained on historical replays augmented with endogenous transient impact (resilience), partial fills, maker/taker fees, and latency. The policy observes depth-20 LOB features plus microstructure indicators and acts under a sell-only inventory constraint to reach a residual target. Evalua

Enzo Duflot, Stanislas Robineau
arXiv · arXiv · 2025

Myopic Optimality: why reinforcement learning portfolio management strategies lose money

Myopic optimization (MO) outperforms reinforcement learning (RL) in portfolio management: RL yields lower or negative returns, higher variance, larger costs, heavier CVaR, lower profitability, and greater model risk. We model execution/liquidation frictions with mark-to-market accounting. Using Malliavin calculus (Clark-Ocone/BEL), we derive policy gradients and risk shadow price, unifying HJB and KKT. This gives dua

Yuming Ma
arXiv · arXiv · 2025

Statistical modeling of SOFR term structure

SOFR derivatives market remains illiquid and incomplete so it is not amenable to classical risk-neutral term structure models which are based on the assumption of perfect liquidity and completeness. This paper develops a statistical SOFR term structure model that is well-suited for risk management and derivatives pricing within the incomplete markets paradigm. The model incorporates relevant macroeconomic factors tha

Teemu Pennanen, Waleed Taoum
arXiv · arXiv · 2025

Increasing Systemic Resilience to Socioeconomic Challenges: Modeling the Dynamics of Liquidity Flows and Systemic Risks Using Navier-Stokes Equations

Modern economic systems face unprecedented socioeconomic challenges, making systemic resilience and effective liquidity flow management essential. Traditional models such as CAPM, VaR, and GARCH often fail to reflect real market fluctuations and extreme events. This study develops and validates an innovative mathematical model based on the Navier-Stokes equations, aimed at the quantitative assessment, forecasting, an

Davit Gondauri
arXiv · arXiv · 2025

Heterogeneous Trader Responses to Macroeconomic Surprises: Simulating Order Flow Dynamics

Understanding how market participants react to shocks like scheduled macroeconomic news is crucial for both traders and policymakers. We develop a calibrated data generation process DGP that embeds four stylized trader archetypes retail, pension, institutional, and hedge funds into an extended CAPM augmented by CPI surprises. Each agents order size choice is driven by a softmax discrete choice rule over small, medium

Haochuan Wang
arXiv · arXiv · 2025

Optimal Investment in Equity and Credit Default Swaps in the Presence of Default

We consider an equity market subject to risk from both unhedgeable shocks and default. The novelty of our work is that to partially offset default risk, investors may dynamically trade in a credit default swap (CDS) market. Assuming investment opportunities are driven by functions of an underlying diffusive factor process, we identify the certainty equivalent for a constant absolute risk aversion investor with a semi

Zhe Fei, Scott Robertson
arXiv · arXiv · 2025

Improving DeFi Accessibility through Efficient Liquidity Provisioning with Deep Reinforcement Learning

This paper applies deep reinforcement learning (DRL) to optimize liquidity provisioning in Uniswap v3, a decentralized finance (DeFi) protocol implementing an automated market maker (AMM) model with concentrated liquidity. We model the liquidity provision task as a Markov Decision Process (MDP) and train an active liquidity provider (LP) agent using the Proximal Policy Optimization (PPO) algorithm. The agent dynamica

Haonan Xu, Alessio Brini
arXiv · arXiv · 2024

Reinforcement Learning for Optimal Execution when Liquidity is Time-Varying

Optimal execution is an important problem faced by any trader. Most solutions are based on the assumption of constant market impact, while liquidity is known to be dynamic. Moreover, models with time-varying liquidity typically assume that it is observable, despite the fact that, in reality, it is latent and hard to measure in real time. In this paper we show that the use of Double Deep Q-learning, a form of Reinforc

Andrea Macrì, Fabrizio Lillo
arXiv · arXiv · 2023

Deep Reinforcement Learning for ESG financial portfolio management

This paper investigates the application of Deep Reinforcement Learning (DRL) for Environment, Social, and Governance (ESG) financial portfolio management, with a specific focus on the potential benefits of ESG score-based market regulation. We leveraged an Advantage Actor-Critic (A2C) agent and conducted our experiments using environments encoded within the OpenAI Gym, adapted from the FinRL platform. The study inclu

Eduardo C. Garrido-Merchán, Sol Mora-Figueroa-Cruz-Guzmán, María Coronado-Vaca
Wiki Entities · 1
Option Blackboard · 0
No Option Blackboard entries matched.
Encyclopedia · 0
No encyclopedia foundations matched.
Cards · 0
No cards matched.
← Back to Codex