arXiv · arXiv q-fin · 2025
Myopic optimization (MO) outperforms reinforcement learning (RL) in portfolio management: RL yields lower or negative returns, higher variance, larger costs, heavier CVaR, lower profitability, and greater model risk. We model execution/liquidation frictions with mark-to-market accounting. Using Malliavin calculus (Clark-Ocone/BEL), we derive policy gradients and risk shadow price, unifying HJB and KKT. This gives dua…
Yuming Ma
arXiv · arXiv q-fin · 2026
Decision-focused learning (DFL) is attractive for portfolio optimization because it trains predictors according to downstream decision quality rather than prediction accuracy alone. However, SPO(Smart, Predict then Optimize surrogate)-based DFL may produce inflated return signals and unstable portfolio reallocations. This study provides a KKT-based interpretation showing that portfolio decisions can be viewed as rank…
Yi Wang, Takashi Hasuike
arXiv · arXiv q-fin · 2026
We study continuous-time multi-asset portfolio choice and consumption under smooth pointwise constraints, including state-dependent feasible sets. The method separates dynamic information acquisition from local constrained recovery. A pointwise-feasible neural actor generates reference rollouts; after training, its realized latent outputs are frozen and first- and second-order adjoints are harvested from a fixed-late…
Jaegi Jeon, Jeonggyu Huh, Hyeng Keun Koo, Byung Hwa Lim
arXiv · arXiv q-fin · 2026
This paper proposes an Extended State-Dependent Hawkes Process (ExsdHawkes) to model the intricate dynamics of Limit Order Books (LOBs). Our theoretical contribution lies in relaxing traditional constraints by allowing for state disappearances -- a phenomenon frequently observed in high-frequency trading. We mathematically prove, using Karush--Kuhn--Tucker (KKT) conditions, that the maximum likelihood estimation rema…
Akitoshi Kimura
arXiv · arXiv q-fin · 2026
This paper proposes the certainty-equivalent first-order learning (CEFOL) algorithm, a deep learning algorithm for solving discrete-time dynamic programming problems with recursive utility. Dynamic programming with recursive utility is challenging because nonlinear certainty equivalent appears in the Bellman equation and the first-order optimality conditions but is difficult to evaluate. By introducing a separate neu…
Xianhua Peng, Wu Guo, Songyan Wang, Jianfei Zhu
arXiv · arXiv q-fin · 2025
We present a white-box, risk-sensitive framework for jointly hedging SPX and VIX exposures under transaction costs and regime shifts. The approach couples an arbitrage-free market teacher with a control layer that enforces safety as constraints. On the market side, we integrate an SSVI-based implied-volatility surface and a Cboe-compliant VIX computation (including wing pruning and 30-day interpolation), and connect …
Jian'an Zhang
arXiv · arXiv q-fin · 2025
We propose a scalable, policy-centric framework for continuous-time multi-asset portfolio-consumption optimization under inequality constraints. Our method integrates neural policies with Pontryagin's Maximum Principle (PMP) and enforces feasibility by maximizing a log-barrier-regularized Hamiltonian at each time-state pair, thereby satisfying KKT conditions without value-function grids. Theoretically, we show that t…
Jeonggyu Huh, Jaegi Jeon, Hyeng Keun Koo, Byung Hwa Lim
arXiv · arXiv q-fin · 2025
We study the construction of SPX--VIX (multi\textendash product) option surfaces that are simultaneously free of static arbitrage and dynamically chain\textendash consistent across maturities. Our method unifies \emph{constructive} PCA--Smolyak approximation and a \emph{chain\textendash consistent} diffusion model with a tri\textendash marginal, martingale\textendash constrained entropic OT (c\textendash EMOT) bridge…
Jian'an Zhang
arXiv · arXiv q-fin · 2024
Data-driven decision-making processes increasingly utilize end-to-end learnable deep neural networks to render final decisions. Sometimes, the output of the forward functions in certain layers is determined by the solutions to mathematical optimization problems, leading to the emergence of differentiable optimization layers that permit gradient back-propagation. However, real-world scenarios often involve large-scale…
Jianming Pan, Zeqi Ye, Xiao Yang, Xu Yang, Weiqing Liu
arXiv · arXiv q-fin · 2021
Recent advances in neural-network architecture allow for seamless integration of convex optimization problems as differentiable layers in an end-to-end trainable neural network. Integrating medium and large scale quadratic programs into a deep neural network architecture, however, is challenging as solving quadratic programs exactly by interior-point methods has worst-case cubic complexity in the number of variables.…
Andrew Butler, Roy Kwon
arXiv · arXiv q-fin · 2020
We investigate the optimal portfolio deleveraging (OPD) problem with permanent and temporary price impacts, where the objective is to maximize equity while meeting a prescribed debt/equity requirement. We take the real situation with cross impact among different assets into consideration. The resulting problem is, however, a non-convex quadratic program with a quadratic constraint and a box constraint, which is known…
Hezhi Luo, Yuanyuan Chen, Xianye Zhang, Duan Li, Huixian Wu
arXiv · arXiv q-fin · 2013
In this paper, we propose $\ell_p$-norm regularized models to seek near-optimal sparse portfolios. These sparse solutions reduce the complexity of portfolio implementation and management. Theoretical results are established to guarantee the sparsity of the second-order KKT points of the $\ell_p$-norm regularized models. More interestingly, we present a theory that relates sparsity of the KKT points with Projected cor…
Caihua Chen, Xindan Li, Caleb Tolman, Suyang Wang, Yinyu Ye