arXiv · arXiv · 2025
This paper proposes a reinforcement learning--based framework for cryptocurrency portfolio management using the Soft Actor--Critic (SAC) and Deep Deterministic Policy Gradient (DDPG) algorithms. Traditional portfolio optimization methods often struggle to adapt to the highly volatile and nonlinear dynamics of cryptocurrency markets. To address this, we design an agent that learns continuous trading actions directly f…
Kamal Paykan
arXiv · arXiv · 2025
Transaction costs and regime shifts are major reasons why paper portfolios fail in live trading. We introduce FR-LUX (Friction-aware, Regime-conditioned Learning under eXecution costs), a reinforcement learning framework that learns after-cost trading policies and remains robust across volatility-liquidity regimes. FR-LUX integrates three ingredients: (i) a microstructure-consistent execution model combining proporti…
Jian'an Zhang
arXiv · arXiv · 2021
With the improvement of computer performance and the development of GPU-accelerated technology, trading with machine learning algorithms has attracted the attention of many researchers and practitioners. In this research, we propose a novel portfolio management strategy based on the framework of Deep Deterministic Policy Gradient, a policy-based reinforcement learning framework, and compare its performance to that of…
Huanming Zhang, Zhengyong Jiang, Jionglong Su
arXiv · arXiv · 2026
Classical market-making strategies based on stochastic control, such as the Avellaneda-Stoikov and the Guéant-Lehalle-Fernandez-Tapia (GLFT) extension, provide closed-form quoting rules, but rest on assumptions that break down at realistic microstructure timescales. One of them is that order flow is stationary, while empirical evidence points to the existence of regimes, possibly associated with algorithmic execution…
Felipe Moret, Fabrizio Lillo
arXiv · arXiv · 2021
This article is part of a comprehensive research project on liquidity risk in asset management, which can be divided into three dimensions. The first dimension covers the modeling of the liability liquidity risk (or funding liquidity), the second dimension is dedicated to the modeling of the asset liquidity risk (or market liquidity), whereas the third dimension considers the management of the asset-liability liquidi…
Thierry Roncalli
arXiv · arXiv · 2019
Systemic liquidity risk, defined by the IMF as "the risk of simultaneous liquidity difficulties at multiple financial institutions", is a key topic in macroprudential policy and financial stress analysis. Specialized models to simulate funding liquidity risk and contagion are available but they require not only banks' bilateral exposures data but also balance sheet data with sufficient granularity, which are hardly a…
V. Macchiati, G. Brandi, G. Cimini, G. Caldarelli, D. Paolotti
arXiv · arXiv · 2021
The problem of portfolio management represents an important and challenging class of dynamic decision making problems, where rebalancing decisions need to be made over time with the consideration of many factors such as investors preferences, trading environments, and market conditions. In this paper, we present a new portfolio policy network architecture for deep reinforcement learning (DRL)that can exploit more eff…
Saeed Marzban, Erick Delage, Jonathan Yumeng Li, Jeremie Desgagne-Bouchard, Carl Dussault
arXiv · arXiv · 2026
Building event-conditioned market models requires separating macro-event labels from persistent microstructure state. We study this distinction in Binance BTCUSDT and ETHUSDT futures from 2023-2026, combining top-20 L2 order book data, trade-flow records, and macro-event windows. We define a supervised discrete L2 liquidity-state transition task, distinct from latent-regime detection and price-direction prediction, a…
Joohyoung Jeon
arXiv · arXiv · 2026
Financial markets are inherently non-stationary, exhibiting frequent regime shifts and structural changes that render traditional Portfolio Management (PM) approaches ineffective. Existing remedies, such as rolling-window retraining and naive online fine-tuning, are hindered by high computational costs and insufficient knowledge utilization, respectively, resulting in low returns and limited adaptability. Continual l…
Chaofan Pan, Lingfei Ren, Linbo Xiong, Yonghao Li, Wei Wei
arXiv · arXiv · 2025
We study opportunistic optimal liquidation over fixed deadlines on BTC-USD limit-order books (LOB). We present RL-Exec, a PPO agent trained on historical replays augmented with endogenous transient impact (resilience), partial fills, maker/taker fees, and latency. The policy observes depth-20 LOB features plus microstructure indicators and acts under a sell-only inventory constraint to reach a residual target. Evalua…
Enzo Duflot, Stanislas Robineau
arXiv · arXiv · 2025
Myopic optimization (MO) outperforms reinforcement learning (RL) in portfolio management: RL yields lower or negative returns, higher variance, larger costs, heavier CVaR, lower profitability, and greater model risk. We model execution/liquidation frictions with mark-to-market accounting. Using Malliavin calculus (Clark-Ocone/BEL), we derive policy gradients and risk shadow price, unifying HJB and KKT. This gives dua…
Yuming Ma
arXiv · arXiv · 2025
SOFR derivatives market remains illiquid and incomplete so it is not amenable to classical risk-neutral term structure models which are based on the assumption of perfect liquidity and completeness. This paper develops a statistical SOFR term structure model that is well-suited for risk management and derivatives pricing within the incomplete markets paradigm. The model incorporates relevant macroeconomic factors tha…
Teemu Pennanen, Waleed Taoum
arXiv · arXiv · 2025
Modern economic systems face unprecedented socioeconomic challenges, making systemic resilience and effective liquidity flow management essential. Traditional models such as CAPM, VaR, and GARCH often fail to reflect real market fluctuations and extreme events. This study develops and validates an innovative mathematical model based on the Navier-Stokes equations, aimed at the quantitative assessment, forecasting, an…
Davit Gondauri
arXiv · arXiv · 2025
Understanding how market participants react to shocks like scheduled macroeconomic news is crucial for both traders and policymakers. We develop a calibrated data generation process DGP that embeds four stylized trader archetypes retail, pension, institutional, and hedge funds into an extended CAPM augmented by CPI surprises. Each agents order size choice is driven by a softmax discrete choice rule over small, medium…
Haochuan Wang
arXiv · arXiv · 2025
We consider an equity market subject to risk from both unhedgeable shocks and default. The novelty of our work is that to partially offset default risk, investors may dynamically trade in a credit default swap (CDS) market. Assuming investment opportunities are driven by functions of an underlying diffusive factor process, we identify the certainty equivalent for a constant absolute risk aversion investor with a semi…
Zhe Fei, Scott Robertson
arXiv · arXiv · 2025
This paper applies deep reinforcement learning (DRL) to optimize liquidity provisioning in Uniswap v3, a decentralized finance (DeFi) protocol implementing an automated market maker (AMM) model with concentrated liquidity. We model the liquidity provision task as a Markov Decision Process (MDP) and train an active liquidity provider (LP) agent using the Proximal Policy Optimization (PPO) algorithm. The agent dynamica…
Haonan Xu, Alessio Brini
arXiv · arXiv · 2024
Optimal execution is an important problem faced by any trader. Most solutions are based on the assumption of constant market impact, while liquidity is known to be dynamic. Moreover, models with time-varying liquidity typically assume that it is observable, despite the fact that, in reality, it is latent and hard to measure in real time. In this paper we show that the use of Double Deep Q-learning, a form of Reinforc…
Andrea Macrì, Fabrizio Lillo
arXiv · arXiv · 2023
This paper investigates the application of Deep Reinforcement Learning (DRL) for Environment, Social, and Governance (ESG) financial portfolio management, with a specific focus on the potential benefits of ESG score-based market regulation. We leveraged an Advantage Actor-Critic (A2C) agent and conducted our experiments using environments encoded within the OpenAI Gym, adapted from the FinRL platform. The study inclu…
Eduardo C. Garrido-Merchán, Sol Mora-Figueroa-Cruz-Guzmán, María Coronado-Vaca