Search

Search

Papers, wiki, Option Blackboard, encyclopedia, and cards.

Results for “off-policy” · papers 11 · wiki 1
Academic Papers · 11arXiv q-fin live 11 · desk corpus 4
arXiv · arXiv q-fin · 2026

Insurance Pricing Optimization via Off-Policy Evaluation

Traditional insurance pricing relies on risk-based principles that ensure actuarial fairness and solvency but do not explicitly account for policyholders' price sensitivity. We formulate insurance pricing as a decision-making problem and study it using tools from off-policy evaluation and stochastic control. We propose a kernelized inverse propensity score estimator that exploits local structure in the action space a

Sascha Günther, Dimitri Semenovich, Mario V. Wüthrich
arXiv · arXiv q-fin · 2025

Off-Policy Evaluation and Counterfactual Methods in Dynamic Auction Environments

Counterfactual estimators are critical for learning and refining policies using logged data, a process known as Off-Policy Evaluation (OPE). OPE allows researchers to assess new policies without costly experiments, speeding up the evaluation process. Online experimental methods, such as A/B tests, are effective but often slow, thus delaying the policy selection and optimization process. In this work, we explore the a

Ritam Guha, Nilavra Pathak
arXiv · arXiv q-fin · 2020

Off-Policy Optimization of Portfolio Allocation Policies under Constraints

The dynamic portfolio optimization problem in finance frequently requires learning policies that adhere to various constraints, driven by investor preferences and risk. We motivate this problem of finding an allocation policy within a sequential decision making framework and study the effects of: (a) using data collected under previously employed policies, which may be sub-optimal and constraint-violating, and (b) im

Nymisha Bandi, Theja Tulabandhula
arXiv · arXiv q-fin · 2022

Model-Free Reinforcement Learning for Asset Allocation

Asset allocation (or portfolio management) is the task of determining how to optimally allocate funds of a finite budget into a range of financial instruments/assets such as stocks. This study investigated the performance of reinforcement learning (RL) when applied to portfolio management using model-free deep RL agents. We trained several RL agents on real-world stock prices to learn how to perform asset allocation.

Adebayo Oshingbesan, Eniola Ajiboye, Peruth Kamashazi, Timothy Mbaka
arXiv · arXiv q-fin · 2026

Hour-Aware Adaptive Risk Management for Autonomous Memecoin Trading on Solana DEXs: Evidence, Theory, and Design Lessons from a 15-Day Deployment

We report a 15-day paper-traded autonomous memecoin trading deployment on Solana decentralised exchanges (DEXs), designed as a controlled measurement of three microstructure questions on which classical equity theory offers well-defined predictions but on which the AMM Solana venue lacks published measurement: (i) time-of-day return patterns on a 24/7 permissionless venue; (ii) whether decision-time filter stacks are

Arati Uday Kamat
arXiv · arXiv q-fin · 2024

Limit Order Book Simulation and Trade Evaluation with $K$-Nearest-Neighbor Resampling

In this paper, we show how $K$-nearest neighbor ($K$-NN) resampling, an off-policy evaluation method proposed in \cite{giegrich2023k}, can be applied to simulate limit order book (LOB) markets and how it can be used to evaluate and calibrate trading strategies. Using historical LOB data, we demonstrate that our simulation method is capable of recreating realistic LOB dynamics and that synthetic trading within the sim

Michael Giegrich, Roel Oomen, Christoph Reisinger
arXiv · arXiv q-fin · 2024

Reinforcement Learning in Non-Markov Market-Making

We develop a deep reinforcement learning (RL) framework for an optimal market-making (MM) trading problem, specifically focusing on price processes with semi-Markov and Hawkes Jump-Diffusion dynamics. We begin by discussing the basics of RL and the deep RL framework used, where we deployed the state-of-the-art Soft Actor-Critic (SAC) algorithm for the deep learning part. The SAC algorithm is an off-policy entropy max

Luca Lalor, Anatoliy Swishchuk
arXiv · arXiv q-fin · 2023

Evaluation of Reinforcement Learning Techniques for Trading on a Diverse Portfolio

This work seeks to answer key research questions regarding the viability of reinforcement learning over the S&P 500 index. The on-policy techniques of Value Iteration (VI) and State-action-reward-state-action (SARSA) are implemented along with the off-policy technique of Q-Learning. The models are trained and tested on a dataset comprising multiple years of stock market data from 2000-2023. The analysis presents the

Ishan S. Khare, Tarun K. Martheswaran, Akshana Dassanaike-Perera
arXiv · arXiv q-fin · 2023

Evaluation of Deep Reinforcement Learning Algorithms for Portfolio Optimisation

We evaluate benchmark deep reinforcement learning algorithms on the task of portfolio optimisation using simulated data. The simulator to generate the data is based on correlated geometric Brownian motion with the Bertsimas-Lo market impact model. Using the Kelly criterion (log utility) as the objective, we can analytically derive the optimal policy without market impact as an upper bound to measure performance when

Chung I Lu
arXiv · arXiv q-fin · 2022

q-Learning in Continuous Time

We study the continuous-time counterpart of Q-learning for reinforcement learning (RL) under the entropy-regularized, exploratory diffusion process formulation introduced by Wang et al. (2020). As the conventional (big) Q-function collapses in continuous time, we consider its first-order approximation and coin the term ``(little) q-function". This function is related to the instantaneous advantage rate function as we

Yanwei Jia, Xun Yu Zhou
arXiv · arXiv q-fin · 2021

Deep Reinforcement Learning for Equal Risk Pricing and Hedging under Dynamic Expectile Risk Measures

Recently equal risk pricing, a framework for fair derivative pricing, was extended to consider dynamic risk measures. However, all current implementations either employ a static risk measure that violates time consistency, or are based on traditional dynamic programming solution schemes that are impracticable in problems with a large number of underlying assets (due to the curse of dimensionality) or with incomplete

Saeed Marzban, Erick Delage, Jonathan Yumeng Li
Wiki Entities · 1
Option Blackboard · 0
No Option Blackboard entries matched.
Encyclopedia · 1
Cards · 0
No cards matched.
← Back to Codex