Search

Search

Papers, wiki, Option Blackboard, encyclopedia, and cards.

Results for “Transformer” · papers 18 · wiki 10
Academic Papers · 18arXiv q-fin live 8 · desk corpus 32
arXiv · arXiv q-fin · 2025

Trading Under Uncertainty: A Distribution-Based Strategy for Futures Markets Using FutureQuant Transformer

In the complex landscape of traditional futures trading, where vast data and variables like real-time Limit Order Books (LOB) complicate price predictions, we introduce the FutureQuant Transformer model, leveraging attention mechanisms to navigate these challenges. Unlike conventional models focused on point predictions, the FutureQuant model excels in forecasting the range and volatility of future prices, thus offer

Wenhao Guo, Yuda Wang, Zeqiao Huang, Changjiang Zhang, Shumin ma
arXiv · arXiv q-fin · 2024

Dynamic ETF Portfolio Optimization Using enhanced Transformer-Based Models for Covariance and Semi-Covariance Prediction(Work in Progress)

This study explores the use of Transformer-based models to predict both covariance and semi-covariance matrices for ETF portfolio optimization. Traditional portfolio optimization techniques often rely on static covariance estimates or impose strict model assumptions, which may fail to capture the dynamic and non-linear nature of market fluctuations. Our approach leverages the power of Transformer models to generate a

Jiahao Zhu, Hengzhi Wu
arXiv · arXiv · 2025

Asset Pricing in Pre-trained Transformer

This paper proposes an innovative Transformer model, Single-directional representative from Transformer (SERT), for US large capital stock pricing. It also innovatively applies the pre-trained Transformer models under the stock pricing and factor investment context. They are compared with standard Transformer models and encoder-only Transformer models in three periods covering the entire COVID-19 pandemic to examine

Shanyan Lai
arXiv · arXiv · 2025

TLOB: A Novel Transformer Model with Dual Attention for Price Trend Prediction with Limit Order Book Data

Price Trend Prediction (PTP) based on Limit Order Book (LOB) data is a fundamental challenge in financial markets. Despite advances in deep learning, existing models fail to generalize across different market conditions and assets. Surprisingly, by adapting a simple MLP-based architecture to LOB, we show that we surpass SoTA performance; thus, challenging the necessity of complex architectures. Unlike past work that

Leonardo Berti, Gjergji Kasneci
arXiv · arXiv · 2024

MarketGPT: Developing a Pre-trained transformer (GPT) for Modeling Financial Time Series

This work presents a generative pre-trained transformer (GPT) designed for modeling financial time series. The GPT functions as an order generation engine within a discrete event simulator, enabling realistic replication of limit order book dynamics. Our model leverages recent advancements in large language models to produce long sequences of order messages in a steaming manner. Our results demonstrate that the model

Aaron Wheeler, Jeffrey D. Varner
arXiv · arXiv · 2024

IVE: Enhanced Probabilistic Forecasting of Intraday Volume Ratio with Transformers

This paper presents a new approach to volume ratio prediction in financial markets, specifically targeting the execution of Volume-Weighted Average Price (VWAP) strategies. Recognizing the importance of accurate volume profile forecasting, our research leverages the Transformer architecture to predict intraday volume ratio at a one-minute scale. We diverge from prior models that use log-transformed volume or turnover

Hanwool Lee, Heehwan Park
arXiv · arXiv · 2026

Measuring Sentiment News with Transformer-Based Language Models

Measuring sentiment from financial news is a central task in economics and finance, yet most existing indicators rely on dictionary-based approaches that infer sentiment from word counts and only partially capture context, negation, and semantic structure. This paper proposes a framework for constructing daily news mood indices using transformer-based language models and evaluates whether they better represent sentim

Maria Saveria Mavillonio, Stefano Borgioli, Caterina Giannetti, Chiara Ongari, Giampiero M. Gallo
arXiv · arXiv · 2026

Robust Transformer-Based One-Step Stock Index Forecasting via Shifted Data Augmentation

Transformers have shown remarkable success in sequence modeling, yet their direct application to financial time series remains challenging due to noisy signals, short-memory dynamics, and distributional shifts. This paper proposes a modified Transformer architecture for one-step stock index forecasting, combined with advanced learning-rate scheduling and a novel Shifted Data Augmentation (SDA) technique. We evaluate

Tien Thanh Thach
arXiv · arXiv · 2026

From Index to Equity: Pre-Training Transformers for Stock Return Prediction

This research aims to leverage machine learning to improve stock price prediction and support informed investment decisions related to buying, selling, and holding assets. Specifically, this work investigates transformer-based models for stock prediction and examines the impact of pre-training strategies on forecasting performance. A transformer model was first pre-trained on the Toronto Stock Exchange Index (TSX) to

Marie Soehl Coolsaet, Roberto Gallardo, Zhen Gao
arXiv · arXiv · 2026

Stock Market Prediction Using Node Transformer Architecture Integrated with BERT Sentiment Analysis

Stock market prediction presents considerable challenges for investors, financial institutions, and policymakers operating in complex market environments characterized by noise, non-stationarity, and behavioral dynamics. Traditional forecasting methods, including fundamental analysis and technical indicators, often fail to capture the intricate patterns and cross-sectional dependencies inherent in financial markets.

Mohammad Al Ridhawi, Mahtab Haj Ali, Hussein Al Osman
arXiv · arXiv · 2026

A Learnable Wavelet Transformer for Long-Short Equity Trading and Risk-Adjusted Return Optimization

Learning profitable intraday trading policies from financial time series is challenging due to heavy noise, non-stationarity, and strong cross-sectional dependence among related assets. We propose \emph{WaveLSFormer}, a learnable wavelet-based long-short Transformer that jointly performs multi-scale decomposition and return-oriented decision learning. Unlike standard time-series forecasting that optimizes prediction

Shuozhe Li, Du Cheng, Leqi Liu
arXiv · arXiv · 2025

EXFormer: A Multi-Scale Trend-Aware Transformer with Dynamic Variable Selection for Foreign Exchange Returns Prediction

Accurately forecasting daily exchange rate returns represents a longstanding challenge in international finance, as the exchange rate returns are driven by a multitude of correlated market factors and exhibit high-frequency fluctuations. This paper proposes EXFormer, a novel Transformer-based architecture specifically designed for forecasting the daily exchange rate returns. We introduce a multi-scale trend-aware sel

Dinggao Liu, Robert Ślepaczuk, Zhenpeng Tang
arXiv · arXiv · 2025

FinAI-BERT: A Transformer-Based Model for Sentence-Level Detection of AI Disclosures in Financial Reports

The proliferation of artificial intelligence (AI) in financial services has prompted growing demand for tools that can systematically detect AI-related disclosures in corporate filings. While prior approaches often rely on keyword expansion or document-level classification, they fall short in granularity, interpretability, and robustness. This study introduces FinAI-BERT, a domain-adapted transformer-based language m

Muhammad Bilal Zafar
arXiv · arXiv · 2025

An Advanced Ensemble Deep Learning Framework for Stock Price Prediction Using VAE, Transformer, and LSTM Model

This research proposes a cutting-edge ensemble deep learning framework for stock price prediction by combining three advanced neural network architectures: The particular areas of interest for the research include but are not limited to: Variational Autoencoder (VAE), Transformer, and Long Short-Term Memory (LSTM) networks. The presented framework is aimed to substantially utilize the advantages of each model which w

Anindya Sarkar, G. Vadivu
arXiv · arXiv · 2024

Higher Order Transformers: Enhancing Stock Movement Prediction On Multimodal Time-Series Data

In this paper, we tackle the challenge of predicting stock movements in financial markets by introducing Higher Order Transformers, a novel architecture designed for processing multivariate time-series data. We extend the self-attention mechanism and the transformer architecture to a higher order, effectively capturing complex market dynamics across time and variables. To manage computational complexity, we propose a

Soroush Omranpour, Guillaume Rabusseau, Reihaneh Rabbany
arXiv · arXiv · 2024

Pretrained LLM Adapted with LoRA as a Decision Transformer for Offline RL in Quantitative Trading

Developing effective quantitative trading strategies using reinforcement learning (RL) is challenging due to the high risks associated with online interaction with live financial markets. Consequently, offline RL, which leverages historical market data without additional exploration, becomes essential. However, existing offline RL methods often struggle to capture the complex temporal dependencies inherent in financi

Suyeol Yun
arXiv · arXiv · 2024

Comparative Analysis of LSTM, GRU, and Transformer Models for Stock Price Prediction

In recent fast-paced financial markets, investors constantly seek ways to gain an edge and make informed decisions. Although achieving perfect accuracy in stock price predictions remains elusive, artificial intelligence (AI) advancements have significantly enhanced our ability to analyze historical data and identify potential trends. This paper takes AI driven stock price trend prediction as the core research, makes

Jue Xiao, Tingting Deng, Shuochen Bi
arXiv · arXiv · 2024

DiffsFormer: A Diffusion Transformer on Stock Factor Augmentation

Machine learning models have demonstrated remarkable efficacy and efficiency in a wide range of stock forecasting tasks. However, the inherent challenges of data scarcity, including low signal-to-noise ratio (SNR) and data homogeneity, pose significant obstacles to accurate forecasting. To address this issue, we propose a novel approach that utilizes artificial intelligence-generated samples (AIGS) to enhance the tra

Yuan Gao, Haokun Chen, Xiang Wang, Zhicai Wang, Xue Wang
Wiki Entities · 10
AI Systems

Attention Mechanism

Attention builds a weighted average of values, with weights from a compatibility function of queries and keys. It lets a model focus on relevant parts of a context instead of a single fixed vector.

AI Systems

BERT

BERT is a bidirectional Transformer encoder trained with masked language modeling and next-sentence prediction, then fine-tuned on downstream NLP tasks.

AI Systems

Flash Attention

FlashAttention computes exact attention with tiling that keeps softmax stats in SRAM, cutting HBM traffic and unlocking longer contexts at the same FLOP count.

AI Systems

GPT

GPT is a decoder-only Transformer trained to predict the next token. Scale plus this objective produced in-context learning and the current foundation-model product line.

AI Systems

Layer Normalization

Layer normalization standardizes activations across features for each example, not across the batch — the stabilizer that made Transformers trainable.

AI Systems

Mixture of Experts

MoE routes each token (or example) to a sparse subset of specialist feed-forward experts, raising parameter count without paying dense FLOPs on every token.

AI Systems

Positional Encoding

Positional encodings inject order into a permutation-invariant attention mixer so the model knows that token i is not token j.

AI Systems

Self-Attention

Self-attention is attention where queries, keys, and values all come from the same sequence, so each position can mix information from every other position in one layer.

AI Systems

Sequence-to-Sequence

Seq2seq maps an input sequence to an output sequence of possibly different length via an encoder–decoder, originally with RNNs and later with Transformers.

AI Systems

Transformer

The Transformer is a sequence model built only from self-attention and feed-forward blocks, with no recurrence. It is the architecture behind BERT, GPT, T5, and almost every modern foundation model.

Option Blackboard · 0
No Option Blackboard entries matched.
Encyclopedia · 5
Cards · 0
No cards matched.
← Back to Codex