arXiv · arXiv q-fin · 2023
Recently, there are many trials to apply reinforcement learning in asset allocation for earning more stable profits. In this paper, we compare performance between several reinforcement learning algorithms - actor-only, actor-critic and PPO models. Furthermore, we analyze each models' character and then introduce the advanced algorithm, so called Reward clipping model. It seems that the Reward Clipping model is better…
Jiwon Kim, Moon-Ju Kang, KangHun Lee, HyungJun Moon, Bo-Kwan Jeon
arXiv · arXiv q-fin · 2017
In order to reduce signalling, traders may resort to limiting access to dark venues and imposing limits on minimum fill sizes they are willing to trade. However, doing this also restricts the liquidity available to the trader since an ever increasing quantity of orders are traded by algos in clips. An alternative is to attempt to monitor signalling in real time and dynamically make adjustments to the dark liquidity a…
Ilija I. Zovko
arXiv · arXiv q-fin · 2026
Decision-focused learning (DFL) is attractive for portfolio optimization because it trains predictors according to downstream decision quality rather than prediction accuracy alone. However, SPO(Smart, Predict then Optimize surrogate)-based DFL may produce inflated return signals and unstable portfolio reallocations. This study provides a KKT-based interpretation showing that portfolio decisions can be viewed as rank…
Yi Wang, Takashi Hasuike
arXiv · arXiv q-fin · 2026
We study behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments. TradeArena, an auditable trading-agent testbed with risk reports, execution simulation, memory, and replayable trajectories, lets us analyze how rationales, positions, and interventions evolve under market stress. Code and data artifacts are available through the \href{https://github.com/…
Weicheng Xue
arXiv · arXiv q-fin · 2026
The daily return of a stock is often restricted to an exchange-imposed band to curb extreme fluctuations. Any attempted price movement beyond this band is clipped, leaving an unobserved excess. We introduce a minimal stochastic latent-state model in which a fraction of this hidden excess is retained for the next day. This retention generates memory, even though the daily stochastic driving shocks are independent. For…
Debraj Das
arXiv · arXiv q-fin · 2025
We present a white-box, risk-sensitive framework for jointly hedging SPX and VIX exposures under transaction costs and regime shifts. The approach couples an arbitrage-free market teacher with a control layer that enforces safety as constraints. On the market side, we integrate an SSVI-based implied-volatility surface and a Cboe-compliant VIX computation (including wing pruning and 30-day interpolation), and connect …
Jian'an Zhang
arXiv · arXiv q-fin · 2025
Deep neural networks (DNNs) have transformed fields such as computer vision and natural language processing by employing architectures aligned with domain-specific structural patterns. In algorithmic trading, however, there remains a lack of architectures that directly incorporate the logic of traditional technical indicators. This study introduces Technical Indicator Networks (TINs), a structured neural design that …
Longfei Lu
arXiv · arXiv q-fin · 2023
We evaluate benchmark deep reinforcement learning algorithms on the task of portfolio optimisation using simulated data. The simulator to generate the data is based on correlated geometric Brownian motion with the Bertsimas-Lo market impact model. Using the Kelly criterion (log utility) as the objective, we can analytically derive the optimal policy without market impact as an upper bound to measure performance when …
Chung I Lu
arXiv · arXiv q-fin · 2023
We consider the problem of simultaneously approximating the conditional distribution of market prices and their log returns with a single machine learning model. We show that an instance of the GDN model of Kratsios and Papon (2022) solves this problem without having prior assumptions on the market's "clipped" log returns, other than that they follow a generalized Ornstein-Uhlenbeck process with a priori unknown dyna…
Anastasis Kratsios, Cody Hyndman
arXiv · arXiv q-fin · 2023
Neural SDEs are continuous-time generative models for sequential data. State-of-the-art performance for irregular time series generation has been previously obtained by training these models adversarially as GANs. However, as typical for GAN architectures, training is notoriously unstable, often suffers from mode collapse, and requires specialised techniques such as weight clipping and gradient penalty to mitigate th…
Zacharia Issa, Blanka Horvath, Maud Lemercier, Cristopher Salvi
arXiv · arXiv q-fin · 2016
Since governments give stimulus to firms and expect the spillover effect by fiscal policies, it is important to know the effectiveness that they can control the economy. To clarify the controllability of the economy, we investigate a firm production network observed exhaustively in Japan and what firms should be directly or indirectly controlled by using control theory. By control theory, we can classify firms into t…
Hiroyasu Inoue