Search

Search

Papers, wiki, Option Blackboard, encyclopedia, and cards.

Results for “benchmark” · papers 18 · wiki 14
Academic Papers · 18arXiv q-fin live 8 · desk corpus 122
arXiv · arXiv · 2026

PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management

Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM) remains poorly benchmarked. Existing benchmarks exhibit two gaps: they are often equity-only and ignore cross-asset correlations; they fail to evaluate the complete PM decision pipeline. We introduce PortBench, a benchmark spanning six heterogeneous asset classes from 2015 to 2025. PortBench compri

Yuxuan Zhao, Sijia Chen, Ningxin Su
arXiv · arXiv · 2026

A Three-Variable Benchmark for Post-GFC Covered Interest Parity Deviations

This paper proposes a public daily-frequency benchmark for post-GFC government-bond CIP deviations. Although CIP deviations are observed daily, the literature lacks a canonical benchmark for daily regressions comparable to standard factor models in asset pricing. Using G10 plus KRW currency-tenor panels, I show that three lagged public state variables-NFCI, the nominal broad U.S. dollar index, and the Treasury 10-yea

Useong Shin
arXiv · arXiv · 2026

RED-2400: A Public Benchmark of Algorithmically-Rejected Trading Events with Outcome Labels

RED-2400 is a public benchmark of 6,660 algorithmically-rejected trading events from a live Solana decentralised-exchange filter stack, observed continuously over 22 calendar days (2026-04-10T21:10Z through 2026-05-02T21:48Z, UTC). Each rejection event is linked to its post-rejection price-and-liquidity trajectory. The deposit contains 169,123 forward-outcome observations and 1,837 graveyard-tracker lifecycle snapsho

Arati U. Kamat
arXiv · arXiv · 2026

Constructing a Portfolio Optimization Benchmark Framework for Evaluating Large Language Models

This study introduces a benchmark framework for evaluating the financial decision-making capabilities of large language models (LLMs) through portfolio optimization problems with mathematically explicit solutions. Unlike existing financial benchmarks that emphasize language-processing tasks, the proposed framework directly tests optimization-based reasoning in investment contexts. A large set of multiple-choice quest

Hanyong Cho, Jang Ho Kim
arXiv · arXiv · 2026

FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings

Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible and even numerically correct answer can be grounded in the wrong evidence. Similar facts and disclosures recur across sections of a filing, across reporting periods of the same firm, and across comparable firms. FinRank targets this provenance-sensitive retrieval problem by requiring systems to identify evidence

Sasan Mansouri, Daniel Saad, Mark Wahrenburg, Manu Weissel, Fabian Woebbeking
arXiv · arXiv · 2026

Pricing options on illiquid assets using liquid market benchmarks: an application to energy markets

The Gasoil options market is illiquid, making it difficult to construct its implied volatility surface directly. However, it is closely linked to the highly liquid Brent options market. In this paper, we jointly model Brent and Gasoil futures prices through a correlated Bachelier local volatility model: the Brent factor is described by a normal mixture diffusion model, while the Gasoil-Brent spot volatility spread is

Federico Aluigi, Lucia Caramellino, Paolo Pigato, Edoardo Scrima
arXiv · arXiv · 2026

When Directional Accuracy Lies: A Base-Rate-Honest Benchmark for LoRA-Adapted TimesFM on Equity Forecasting

Large pretrained time-series models such as TimesFM are attractive for financial forecasting, but raw directional accuracy is a misleading scoreboard in equity markets. An early LoRA adapter in this project appeared to reach roughly 80% directional accuracy; we show this is not evidence of skill. Over a long horizon in a rising market, a trivial "always-up" rule attains comparably high accuracy without using the inpu

Taizhen Cheung
arXiv · arXiv · 2026

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading. Yet most still rank agents by returns over a fixed window, a weak proxy: the market path dominates a period's return, and apparent alpha can dissolve once look-ahead leakage is controlled. We introduce CLQT, which reframes closed-loop trading evaluation as diagnosis bef

Bo Qu, Mingguang Chen
arXiv · arXiv · 2026

FinBench: Time-Gated Calibration and Uncertainty Benchmarking for Agentic Financial Forecasting

Large language models (LLMs) are increasingly used as components of agentic systems that observe, plan, and act. In finance, even "assistive" systems become decision-relevant once their outputs are used to size trades or allocate risk. A key failure mode is the confidence--competence gap: a model that is only slightly better than chance but consistently overconfident will, under typical bet-sizing rules, generate neg

Rishab Ghosh, Vinay Devarakonda
arXiv · arXiv · 2026

Leakage-Aware Benchmarking of LLM Forecasting: Real-Time Nowcasts as the Decision-Time Input for Macro Factor Ranking

Forecasting benchmarks for retrieval-augmented LLMs routinely confound model capability with information leakage: features labeled with a target's timestamp are often not observable at the system's decision time. We study leakage-controlled equity factor ranking with a retrieval-augmented 7B open-source LLM forecaster. At each month-end from 2023-04 to 2026-03, the forecaster observes only decision-time information:

Mao Guan, Qian Chen
arXiv · arXiv · 2026

Benchmarking Deep Time Series Models for Equity Portfolios

Benchmarking forecasting architectures for daily equity portfolios is not just a prediction exercise. It also asks which model remains usable after preferences, costs, and portfolio constraints are imposed. We build a CRSP daily-stock benchmark for 15 deep and statistical time-series architectures over 2018--2024. The protocol combines common-window decile portfolios, stochastic multi-criteria acceptability analysis,

Aoxin Zhang, Yuhan Cheng, Kwanting Leung
arXiv · arXiv · 2026

Benchmarking Quantum Algorithmic Resilience for CVaR Portfolio Optimization: The Expressibility-Coherence Trade-off

Quantum combinatorial optimization offers theoretical advantages for complex financial modeling, but physical implementation on Noisy Intermediate Scale Quantum (NISQ) devices is severely constrained by hardware topology. This study presents a hardware benchmarking analysis between a Hardware Efficient Variational Quantum Neural Network (HE-VQNN) and the Warm Start Quantum Approximate Optimization Algorithm (WS-QAOA)

Prashik N. Somkuwar, K. Srinivasan, G. Raghavan
arXiv · arXiv · 2026

From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets

Evaluating whether large language model (LLM) agents can profit in capital markets is increasingly framed as end-to-end trading: place an agent in a historical market, let it trade, and measure portfolio returns. This setup is vulnerable to two evaluation failures. First, long backtests often overlap with the knowledge cutoffs of frontier LLMs, allowing memorized tickers, dates, prices, and market narratives to subst

Taojie Zhu, Wentao Zhao, Rui Sun, Beidi Luan, Jiacheng Lu
arXiv · arXiv · 2026

PHBench: A Benchmark for Predicting Startup Series A Funding from Product Hunt Launch Signals

Structured launch signals on Product Hunt contain statistically significant predictive information for Series A funding outcomes. We construct PHBench from 67,292 featured Product Hunt posts spanning 2019-2025, linked to Crunchbase funding records via deterministic domain matching, identifying 528 verified Series A raises within 18 months of launch (positive rate: 0.78%). Our best-performing model, a three-component

Yagiz Ihlamur, Ben Griffin, Rick Chen
arXiv · arXiv · 2026

Deriving the term-structure of loan write-off risk under IFRS 9 by using survival analysis: A benchmark study

The estimation of marginal loan write-off probabilities is a non-trivial task when modelling the loss given default (LGD) risk parameter in credit risk. We explore two types of survival models in estimating the overall write-off probability over default spell time, where these probabilities form the term-structure of write-off risk in aggregate. These survival models include a discrete-time hazard (DtH) model and a c

Arno Botha, Mohammed Gabru, Marcel Muller, Janette Larney
arXiv · arXiv · 2026

Deep Learning for Financial Time Series: A Large-Scale Benchmark of Risk-Adjusted Performance

We present a large scale benchmark of modern deep learning architectures for a financial time series prediction and position sizing task, with a primary focus on Sharpe ratio optimization. Evaluating linear models, recurrent networks, transformer based architectures, state space models, and recent sequence representation approaches, we assess out of sample performance on a daily futures dataset spanning commodities,

Adir Saly-Kaufmann, Kieran Wood, Jan Peter-Calliess, Stefan Zohren
arXiv · arXiv · 2026

AlphaForgeBench: Benchmarking End-to-End Trading Strategy Design with Large Language Models

The rapid advancement of Large Language Models (LLMs) has led to a surge of financial benchmarks, evolving from static knowledge evaluation toward interactive trading simulations. However, existing frameworks for evaluating real-time trading largely overlook a critical failure mode: the severe behavioral instability of LLMs in sequential decision-making under financial uncertainty. Through extensive experiments, we s

Wentao Zhang, Mingxuan Zhao, Jincheng Gao, Jieshun You, Huaiyu Jia
arXiv · arXiv · 2025

What's the Price of Monotonicity? A Multi-Dataset Benchmark of Monotone-Constrained Gradient Boosting for Credit PD

Financial institutions face a trade-off between predictive accuracy and interpretability when deploying machine learning models for credit risk. Monotonicity constraints align model behavior with domain knowledge, but their performance cost - the price of monotonicity - is not well quantified. This paper benchmarks monotone-constrained versus unconstrained gradient boosting models for credit probability of default ac

Petr Koklev
Wiki Entities · 14
CTA

SG CTA and SG Trend Indexes

The industry tape: SG CTA Index for a broad managed-futures peer set, SG Trend for the large trend-followers — the benchmarks allocators actually quote.

Derivatives

Move Index

The MOVE Index tracks implied volatility in the U.S. Treasury market and serves as a benchmark for rates uncertainty and macro stress.

Desk Slang

Behind the Curve

Behind the curve means policy (or a book) is too easy or too slow relative to incoming inflation, growth, or a Taylor-type benchmark — the market is already pricing a catch-up.

Desk Slang

On-the-Run vs Off-the-Run

On-the-run is the latest issued Treasury (or benchmark) in a maturity; off-the-runs are older issues. The on-the-run is richer and more liquid; the spread is a liquidity and specials object.

Liquidity

Commercial Paper Spread

Commercial paper spreads track the cost of short-term corporate borrowing relative to safer benchmarks and help identify stress in corporate funding markets.

Liquidity

LIBOR-OIS Spread

LIBOR-OIS spread tracks the gap between unsecured bank funding rates and overnight indexed swap rates, historically serving as a benchmark for banking-system stress.

Liquidity

SOFR

SOFR is the Secured Overnight Financing Rate, a key benchmark for U.S. dollar funding based on overnight Treasury repo transactions.

Macro Policy

Yield Curve Control

Yield Curve Control — Official caps on benchmark yields and the distortions they create in RV and cross-market hedging.

Microstructure

Volume-Weighted Average Price

VWAP is the day’s (or window’s) average price weighted by volume — a benchmark for whether you traded with the tape or against it.

Quant

Efficient Market Hypothesis

EMH says prices reflect available information so that you cannot systematically earn risk-adjusted profits from that information — a benchmark, not a religion.

Quant

Information Ratio

The information ratio is active return over active risk — residual performance per unit of tracking error versus a benchmark.

Quant

Tracking Error

Tracking error is the volatility of active return versus a benchmark — how much the book is allowed to be not-the-index.

Quant

Transaction Cost Analysis

Transaction Cost Analysis — Post-trade measurement of slippage versus benchmarks for alpha decay control.

Rates

US 10-Year Real Yield

US 10-Year Real Yield measures the inflation-adjusted yield on 10-year Treasuries and is a key benchmark for discount rates, financial conditions, and macro asset pricing.

Option Blackboard · 0
No Option Blackboard entries matched.
Encyclopedia · 14
Desk Slang · Foundations

Behind the Curve

Behind the curve means policy (or a book) is too easy or too slow relative to incoming inflation, growth, or a Taylor-type benchmark — the market is already pricing a catch-up.

Liquidity · Foundations

Commercial Paper Spread

Commercial paper spreads track the cost of short-term corporate borrowing relative to safer benchmarks and help identify stress in corporate funding markets.

Quant · Foundations

Efficient Market Hypothesis

EMH says prices reflect available information so that you cannot systematically earn risk-adjusted profits from that information — a benchmark, not a religion.

Quant · Foundations

Information Ratio

The information ratio is active return over active risk — residual performance per unit of tracking error versus a benchmark.

Liquidity · Foundations

LIBOR-OIS Spread

LIBOR-OIS spread tracks the gap between unsecured bank funding rates and overnight indexed swap rates, historically serving as a benchmark for banking-system stress.

Derivatives · Foundations

Move Index

The MOVE Index tracks implied volatility in the U.S. Treasury market and serves as a benchmark for rates uncertainty and macro stress.

Desk Slang · Foundations

On-the-Run vs Off-the-Run

On-the-run is the latest issued Treasury (or benchmark) in a maturity; off-the-runs are older issues. The on-the-run is richer and more liquid; the spread is a liquidity and specials object.

CTA · Foundations

SG CTA and SG Trend Indexes

The industry tape: SG CTA Index for a broad managed-futures peer set, SG Trend for the large trend-followers — the benchmarks allocators actually quote.

Liquidity · Foundations

SOFR

SOFR is the Secured Overnight Financing Rate, a key benchmark for U.S. dollar funding based on overnight Treasury repo transactions.

Quant · Foundations

Tracking Error

Tracking error is the volatility of active return versus a benchmark — how much the book is allowed to be not-the-index.

Quant · Foundations

Transaction Cost Analysis

Transaction Cost Analysis — Post-trade measurement of slippage versus benchmarks for alpha decay control.

Rates · Foundations

US 10-Year Real Yield

US 10-Year Real Yield measures the inflation-adjusted yield on 10-year Treasuries and is a key benchmark for discount rates, financial conditions, and macro asset pricing.

Microstructure · Foundations

Volume-Weighted Average Price

VWAP is the day’s (or window’s) average price weighted by volume — a benchmark for whether you traded with the tape or against it.

Macro Policy · Foundations

Yield Curve Control

Yield Curve Control — Official caps on benchmark yields and the distortions they create in RV and cross-market hedging.

Cards · 0
No cards matched.
← Back to Codex