arXiv · arXiv · 2026
Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM) remains poorly benchmarked. Existing benchmarks exhibit two gaps: they are often equity-only and ignore cross-asset correlations; they fail to evaluate the complete PM decision pipeline. We introduce PortBench, a benchmark spanning six heterogeneous asset classes from 2015 to 2025. PortBench compri…
Yuxuan Zhao, Sijia Chen, Ningxin Su
arXiv · arXiv · 2026
This paper proposes a public daily-frequency benchmark for post-GFC government-bond CIP deviations. Although CIP deviations are observed daily, the literature lacks a canonical benchmark for daily regressions comparable to standard factor models in asset pricing. Using G10 plus KRW currency-tenor panels, I show that three lagged public state variables-NFCI, the nominal broad U.S. dollar index, and the Treasury 10-yea…
Useong Shin
arXiv · arXiv · 2026
RED-2400 is a public benchmark of 6,660 algorithmically-rejected trading events from a live Solana decentralised-exchange filter stack, observed continuously over 22 calendar days (2026-04-10T21:10Z through 2026-05-02T21:48Z, UTC). Each rejection event is linked to its post-rejection price-and-liquidity trajectory. The deposit contains 169,123 forward-outcome observations and 1,837 graveyard-tracker lifecycle snapsho…
Arati U. Kamat
arXiv · arXiv · 2026
This study introduces a benchmark framework for evaluating the financial decision-making capabilities of large language models (LLMs) through portfolio optimization problems with mathematically explicit solutions. Unlike existing financial benchmarks that emphasize language-processing tasks, the proposed framework directly tests optimization-based reasoning in investment contexts. A large set of multiple-choice quest…
Hanyong Cho, Jang Ho Kim
arXiv · arXiv · 2026
Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible and even numerically correct answer can be grounded in the wrong evidence. Similar facts and disclosures recur across sections of a filing, across reporting periods of the same firm, and across comparable firms. FinRank targets this provenance-sensitive retrieval problem by requiring systems to identify evidence …
Sasan Mansouri, Daniel Saad, Mark Wahrenburg, Manu Weissel, Fabian Woebbeking
arXiv · arXiv · 2026
The Gasoil options market is illiquid, making it difficult to construct its implied volatility surface directly. However, it is closely linked to the highly liquid Brent options market. In this paper, we jointly model Brent and Gasoil futures prices through a correlated Bachelier local volatility model: the Brent factor is described by a normal mixture diffusion model, while the Gasoil-Brent spot volatility spread is…
Federico Aluigi, Lucia Caramellino, Paolo Pigato, Edoardo Scrima
arXiv · arXiv · 2026
Large pretrained time-series models such as TimesFM are attractive for financial forecasting, but raw directional accuracy is a misleading scoreboard in equity markets. An early LoRA adapter in this project appeared to reach roughly 80% directional accuracy; we show this is not evidence of skill. Over a long horizon in a rising market, a trivial "always-up" rule attains comparably high accuracy without using the inpu…
Taizhen Cheung
arXiv · arXiv · 2026
LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading. Yet most still rank agents by returns over a fixed window, a weak proxy: the market path dominates a period's return, and apparent alpha can dissolve once look-ahead leakage is controlled. We introduce CLQT, which reframes closed-loop trading evaluation as diagnosis bef…
Bo Qu, Mingguang Chen
arXiv · arXiv · 2026
Large language models (LLMs) are increasingly used as components of agentic systems that observe, plan, and act. In finance, even "assistive" systems become decision-relevant once their outputs are used to size trades or allocate risk. A key failure mode is the confidence--competence gap: a model that is only slightly better than chance but consistently overconfident will, under typical bet-sizing rules, generate neg…
Rishab Ghosh, Vinay Devarakonda
arXiv · arXiv · 2026
Forecasting benchmarks for retrieval-augmented LLMs routinely confound model capability with information leakage: features labeled with a target's timestamp are often not observable at the system's decision time. We study leakage-controlled equity factor ranking with a retrieval-augmented 7B open-source LLM forecaster. At each month-end from 2023-04 to 2026-03, the forecaster observes only decision-time information: …
Mao Guan, Qian Chen
arXiv · arXiv · 2026
Benchmarking forecasting architectures for daily equity portfolios is not just a prediction exercise. It also asks which model remains usable after preferences, costs, and portfolio constraints are imposed. We build a CRSP daily-stock benchmark for 15 deep and statistical time-series architectures over 2018--2024. The protocol combines common-window decile portfolios, stochastic multi-criteria acceptability analysis,…
Aoxin Zhang, Yuhan Cheng, Kwanting Leung
arXiv · arXiv · 2026
Quantum combinatorial optimization offers theoretical advantages for complex financial modeling, but physical implementation on Noisy Intermediate Scale Quantum (NISQ) devices is severely constrained by hardware topology. This study presents a hardware benchmarking analysis between a Hardware Efficient Variational Quantum Neural Network (HE-VQNN) and the Warm Start Quantum Approximate Optimization Algorithm (WS-QAOA)…
Prashik N. Somkuwar, K. Srinivasan, G. Raghavan
arXiv · arXiv · 2026
Evaluating whether large language model (LLM) agents can profit in capital markets is increasingly framed as end-to-end trading: place an agent in a historical market, let it trade, and measure portfolio returns. This setup is vulnerable to two evaluation failures. First, long backtests often overlap with the knowledge cutoffs of frontier LLMs, allowing memorized tickers, dates, prices, and market narratives to subst…
Taojie Zhu, Wentao Zhao, Rui Sun, Beidi Luan, Jiacheng Lu
arXiv · arXiv · 2026
Structured launch signals on Product Hunt contain statistically significant predictive information for Series A funding outcomes. We construct PHBench from 67,292 featured Product Hunt posts spanning 2019-2025, linked to Crunchbase funding records via deterministic domain matching, identifying 528 verified Series A raises within 18 months of launch (positive rate: 0.78%). Our best-performing model, a three-component …
Yagiz Ihlamur, Ben Griffin, Rick Chen
arXiv · arXiv · 2026
The estimation of marginal loan write-off probabilities is a non-trivial task when modelling the loss given default (LGD) risk parameter in credit risk. We explore two types of survival models in estimating the overall write-off probability over default spell time, where these probabilities form the term-structure of write-off risk in aggregate. These survival models include a discrete-time hazard (DtH) model and a c…
Arno Botha, Mohammed Gabru, Marcel Muller, Janette Larney
arXiv · arXiv · 2026
We present a large scale benchmark of modern deep learning architectures for a financial time series prediction and position sizing task, with a primary focus on Sharpe ratio optimization. Evaluating linear models, recurrent networks, transformer based architectures, state space models, and recent sequence representation approaches, we assess out of sample performance on a daily futures dataset spanning commodities, …
Adir Saly-Kaufmann, Kieran Wood, Jan Peter-Calliess, Stefan Zohren
arXiv · arXiv · 2026
The rapid advancement of Large Language Models (LLMs) has led to a surge of financial benchmarks, evolving from static knowledge evaluation toward interactive trading simulations. However, existing frameworks for evaluating real-time trading largely overlook a critical failure mode: the severe behavioral instability of LLMs in sequential decision-making under financial uncertainty. Through extensive experiments, we s…
Wentao Zhang, Mingxuan Zhao, Jincheng Gao, Jieshun You, Huaiyu Jia
arXiv · arXiv · 2025
Financial institutions face a trade-off between predictive accuracy and interpretability when deploying machine learning models for credit risk. Monotonicity constraints align model behavior with domain knowledge, but their performance cost - the price of monotonicity - is not well quantified. This paper benchmarks monotone-constrained versus unconstrained gradient boosting models for credit probability of default ac…
Petr Koklev