Search

Search

Papers, wiki, Option Blackboard, encyclopedia, and cards.

Results for “data” · papers 18 · wiki 14
Academic Papers · 18arXiv q-fin live 8 · desk corpus 453
arXiv · arXiv · 2026

Data-Driven Duration Management -- Term Structure Forecasting Using Machine Learning

This paper compares different methods for forecasting the term structure of U.S. and European zero-coupon government bonds using both traditional econometric and Machine Learning (ML) approaches. We compare classical models (e.g., Dynamic Nelson-Siegel (DNS) and Principal Component Analysis (PCA)) with different Neural Network (NN) architectures, including those inspired by the classical models, on the U.S. Treasury

Tobias Lausser, Joao Eduardo Vuolo, Rudi Zagst
arXiv · arXiv · 2024

Market-Neutral Strategies in Mid-Cap Portfolio Management: A Data-Driven Approach to Long-Short Equity

Mid-cap companies, generally valued between \$2 billion and \$10 billion, provide investors with a well-rounded opportunity between the fluctuation of small-cap stocks and the stability of large-cap stocks. This research builds upon the long-short equity approach (e.g., Michaud, 2018; Dimitriu, Alexander, 2002) customized for mid-cap equities, providing steady risk-adjusted returns yielding a significant Sharpe ratio

Saumya Kothari, Harsh Shah, Utkarsh Prajapati, Shrinjay Kaushik
arXiv · arXiv · 2023

Improved Data Generation for Enhanced Asset Allocation: A Synthetic Dataset Approach for the Fixed Income Universe

We present a novel process for generating synthetic datasets tailored to assess asset allocation methods and construct portfolios within the fixed income universe. Our approach begins by enhancing the CorrGAN model to generate synthetic correlation matrices. Subsequently, we propose an Encoder-Decoder model that samples additional data conditioned on a given correlation matrix. The resulting synthetic dataset facilit

Szymon Kubiak, Tillman Weyde, Oleksandr Galkin, Dan Philps, Ram Gopal
arXiv · arXiv · 2023

Handling missing data in Burundian sovereign bond market

Constructing an accurate yield curve is essential for evaluating financial instruments and analyzing market trends in the bond market. However, in the case of the Burundian sovereign bond market, the presence of missing data poses a significant challenge to accurately constructing the yield curve. In this paper, we explore the limitations and data availability constraints specific to the Burundian sovereign market an

Irène Irakoze, Rédempteur Ntawiratsa, David Niyukuri
arXiv · arXiv · 2023

A Scalable Reinforcement Learning-based System Using On-Chain Data for Cryptocurrency Portfolio Management

On-chain data (metrics) of blockchain networks, akin to company fundamentals, provide crucial and comprehensive insights into the networks. Despite their informative nature, on-chain data have not been utilized in reinforcement learning (RL)-based systems for cryptocurrency (crypto) portfolio management (PM). An intriguing subject is the extent to which the utilization of on-chain data can enhance an RL-based system'

Zhenhan Huang, Fumihide Tanaka
arXiv · arXiv · 2023

An Empirical Study of Capital Asset Pricing Model based on Chinese A-share Trading Data

This paper presents an empirical analysis of the capital asset pricing model using trading data for the Chinese A-share market from 2000 to 2019. Firstly, the standard CAPM is tested using a Fama-MacBetch regression and although the results successfully test the three core hypotheses, the resulting beta risk does not have a significant impact on returns. Secondly, the Fama-French three-factor model, which uses a comb

Kai Ren
arXiv · arXiv · 2022

Method of indirect estimation of default probability dynamics for industry-target segments according to the data of Bank of Russia

A direct method for calculating default rates by industry and target corporate segments is not possible given the lack of statistical data. The proposed paper considers a model for filtering the dynamics of the probability of default of corporate companies and other borrowers based on indirect data on the dynamics of overdue debt supplied by the Bank of Russia. The model is based on the equation of the balance of tot

Mikhail Pomazanov
arXiv · arXiv · 2022

Are all Credit Default Swap Databases equal?

We compare the five major sources of corporate Credit Default Swap prices: GFI, Fenics, Reuters, CMA, and Markit, using the most liquid single name 5-year CDS in the iTraxx and CDX indexes from 2004 to 2010. Deviations from the common trend among prices in the different databases are not random but are explained by idiosyncratic factors, financing costs, global risk, and other trading factors. The CMA quotes lead the

Sergio Mayordomo, Juan Ignacio Peña, Eduardo S. Schwartz
arXiv · arXiv · 2016

Reconstruction of Order Flows using Aggregated Data

In this work we investigate tick-by-tick data provided by the TRTH database for several stocks on three different exchanges (Paris - Euronext, London and Frankfurt - Deutsche Börse) and on a 5-year span. We use a simple algorithm that helps the synchronization of the trades and quotes data sources, providing enhancements to the basic procedure that, depending on the time period and the exchange, are shown to be signi

Ioane Muni Toke
arXiv · arXiv · 2014

Liquidity commonality does not imply liquidity resilience commonality: A functional characterisation for ultra-high frequency cross-sectional LOB data

We present a large-scale study of commonality in liquidity and resilience across assets in an ultra high-frequency (millisecond-timestamped) Limit Order Book (LOB) dataset from a pan-European electronic equity trading facility. We first show that extant work in quantifying liquidity commonality through the degree of explanatory power of the dominant modes of variation of liquidity (extracted through Principal Compone

Efstathios Panayi, Gareth Peters, Ioannis Kosmidis
arXiv · arXiv · 2026

dexamine: A Python package for Uniswap event data on Ethereum

Decentralized exchanges record trading and liquidity provision on public blockchains, but empirical analysis requires interpreting these records and linking them to execution metadata. dexamine is a Python package that parses Uniswap v2 and v3 events on Ethereum. It converts transaction receipt logs into observations of trades and liquidity changes, with token quantities, pool state, transaction order, and gas inform

Magnus Hansson
arXiv · arXiv · 2026

Metaorder modelling and identification from public data

Market-order flow in financial markets exhibits long-range correlations. This is a widely known stylised fact of financial markets. A popular hypothesis for this stylised fact comes from the Lillo-Mike-Farmer (LMF) order-splitting theory. However, quantitative tests of this theory have historically relied on proprietary datasets with trader identifiers, limiting reproducibility and cross-market validation. We investi

Ezra Goliath, Tim Gebbie
arXiv · arXiv · 2026

Data-Driven Measures of High-Frequency Trading

Public data do not identify high-frequency trading (HFT), and standard proxies do not separate liquidity-supplying from liquidity-demanding strategies. We overcome this measurement challenge by training machine learning models on proprietary Nasdaq data to map observed HFT activity to public intraday variables. Applying this mapping, we generate daily measures of liquidity-supplying and liquidity-demanding HFT for al

Gbenga Ibikunle, Ben Moews, Dmitriy Muravyev, Khaladdin Rzayev
arXiv · arXiv · 2026

OpenMarket: A Synchronized Polymarket-Binance Dataset for High-Frequency Prediction-Market Research

OpenMarket began as an attempt to trade Polymarket's BTC 15-minute binary markets against Binance BTC/USDT order flow. The attempt did not produce a tradable edge: out-of-sample, a walk-forward logistic model over 43 microstructure features does not beat, and slightly underperforms, the probability already implied by Polymarket's own order book, and simulated trading nets -0.116 normalized payoff units per attempted

Gregory Young
arXiv · arXiv · 2025

Inverse Portfolio Optimization with Synthetic Investor Data: Recovering Risk Preferences under Uncertainty

This study develops an inverse portfolio optimization framework for recovering latent investor preferences including risk aversion, transaction cost sensitivity, and ESG orientation from observed portfolio allocations. Using controlled synthetic data, we assess the estimator's statistical properties such as consistency, coverage, and dynamic regret. The model integrates robust optimization and regret-based inference

Jinho Cha, Long Pham, Thi Le Hoa Vo, Jaeyoung Cho, Jaejin Lee
arXiv · arXiv · 2025

TLOB: A Novel Transformer Model with Dual Attention for Price Trend Prediction with Limit Order Book Data

Price Trend Prediction (PTP) based on Limit Order Book (LOB) data is a fundamental challenge in financial markets. Despite advances in deep learning, existing models fail to generalize across different market conditions and assets. Surprisingly, by adapting a simple MLP-based architecture to LOB, we show that we surpass SoTA performance; thus, challenging the necessity of complex architectures. Unlike past work that

Leonardo Berti, Gjergji Kasneci
arXiv · arXiv · 2023

Adaptive Agents and Data Quality in Agent-Based Financial Markets

We present our Agent-Based Market Microstructure Simulation (ABMMS), an Agent-Based Financial Market (ABFM) that captures much of the complexity present in the US National Market System for equities (NMS). Agent-Based models are a natural choice for understanding financial markets. Financial markets feature a constrained action space that should simplify model creation, produce a wealth of data that should aid model

Colin M. Van Oort, Ethan Ratliff-Crain, Brian F. Tivnan, Safwan Wshah
arXiv · arXiv · 2023

Cost of Implementation of Basel III reforms in Bangladesh -- A Panel data analysis

Inspired by the recent debate on the macroeconomic implications of the new bank regulatory standards known as Basel III, we tried to find out in this study that the impact of Basel III liquidity and capital requirements in Bangladesh proposed by Basel Committee on Banking Supervision (BCBS, 2010a). A small set of macro variables, using a sample of 22 private commercial banks operating in Bangladesh for the period of

Dipti Rani Hazra, Md. Shah Naoaj, Mohammed Mahinur Alam, Abdul Kader
Wiki Entities · 14
AI Systems

Diffusion Model

A diffusion model learns to reverse a gradual noising process. Sampling starts from noise and iteratively denoises toward the data distribution.

AI Systems

Generative Adversarial Network

A GAN trains a generator and a discriminator against each other: the generator maps noise to fake samples, the discriminator learns real vs fake, and the equilibrium is a generator whose samples match the data distribution.

AI Systems

Gradient Descent

Gradient descent updates parameters against the gradient of a loss: θ ← θ − η ∇_θ L. Stochastic and mini-batch variants make the method tractable on large datasets.

AI Systems

Hallucination

Hallucination is fluent generation that is not supported by the source or the world — a likelihood-trained model completing a pattern, not a database lookup.

AI Systems

Regularization

Regularization is any constraint that trades train fit for expected live error: weight decay, dropout, early stopping, data augmentation, or a simpler hypothesis class.

AI Systems

Scaling Laws

Scaling laws are empirical power laws relating language-model loss to parameter count, data, and compute, used to plan pretraining rather than guess.

Crypto

Blockchain

A blockchain is an append-only replicated ledger with a consensus rule — a database with an incentive system, not a price target.

CTA

Systematic Macro CTA

A CTA that trades futures on economic data, not only price — growth, inflation, positioning, and nowcasts as the signal set.

Economy

GDP Nowcast

GDP Nowcast — High-frequency aggregation of activity data to estimate current-quarter growth in real time.

Mathematics

Bayesian Inference

Bayesian inference updates a prior distribution over parameters with data via Bayes’ rule to get a posterior — beliefs as probabilities, not just a point estimate.

Strategies

Alpha Cloning — Following 13F Filings

Copy (with a lag) the disclosed long holdings of selected 13F filers — a delayed clone of someone else’s book.

Strategies

Filing Similarity and Stock Returns

Use how similar this year’s 10-K/10-Q language is to last year’s as a signal — boilerplate vs change as alternative data.

Strategies

Insider Buying Strategy

Overweight names with clustered open-market insider buys and avoid heavy insider sales — a delayed Form-4 signal.

Systems

Feature Store

Feature Store — Centralized repository for model features ensuring consistency between research and production.

Option Blackboard · 0
No Option Blackboard entries matched.
Encyclopedia · 11
Mathematics · Foundations

Bayesian Inference

Bayesian inference updates a prior distribution over parameters with data via Bayes’ rule to get a posterior — beliefs as probabilities, not just a point estimate.

Crypto · Foundations

Blockchain

A blockchain is an append-only replicated ledger with a consensus rule — a database with an incentive system, not a price target.

AI Systems · Foundations

Diffusion Model

A diffusion model learns to reverse a gradual noising process. Sampling starts from noise and iteratively denoises toward the data distribution.

Strategies · Foundations

Filing Similarity and Stock Returns

Use how similar this year’s 10-K/10-Q language is to last year’s as a signal — boilerplate vs change as alternative data.

Economy · Foundations

GDP Nowcast

GDP Nowcast — High-frequency aggregation of activity data to estimate current-quarter growth in real time.

AI Systems · Foundations

Generative Adversarial Network

A GAN trains a generator and a discriminator against each other: the generator maps noise to fake samples, the discriminator learns real vs fake, and the equilibrium is a generator whose samples match the data distribution.

AI Systems · Foundations

Gradient Descent

Gradient descent updates parameters against the gradient of a loss: θ ← θ − η ∇_θ L. Stochastic and mini-batch variants make the method tractable on large datasets.

AI Systems · Foundations

Hallucination

Hallucination is fluent generation that is not supported by the source or the world — a likelihood-trained model completing a pattern, not a database lookup.

AI Systems · Foundations

Regularization

Regularization is any constraint that trades train fit for expected live error: weight decay, dropout, early stopping, data augmentation, or a simpler hypothesis class.

AI Systems · Foundations

Scaling Laws

Scaling laws are empirical power laws relating language-model loss to parameter count, data, and compute, used to plan pretraining rather than guess.

CTA · Foundations

Systematic Macro CTA

A CTA that trades futures on economic data, not only price — growth, inflation, positioning, and nowcasts as the signal set.

Cards · 2
← Back to Codex