Search

Search

Papers, wiki, Option Blackboard, encyclopedia, and cards.

Results for “hallucination” · papers 14 · wiki 1
Academic Papers · 14arXiv q-fin live 13 · desk corpus 5
arXiv · arXiv q-fin · 2025

Detecting AI Hallucinations in Finance: An Information-Theoretic Method Cuts Hallucination Rate by 92%

Large language models (LLMs) produce fluent but unsupported answers - hallucinations - limiting safe deployment in high-stakes domains. We propose ECLIPSE, a framework that treats hallucination as a mismatch between a model's semantic entropy and the capacity of available evidence. We combine entropy estimation via multi-sample clustering with a novel perplexity decomposition that measures how models use retrieved ev

Mainak Singha
arXiv · arXiv q-fin · 2025

Prompt-Response Semantic Divergence Metrics for Faithfulness Hallucination and Misalignment Detection in Large Language Models

The proliferation of Large Language Models (LLMs) is challenged by hallucinations, critical failure modes where models generate non-factual, nonsensical or unfaithful text. This paper introduces Semantic Divergence Metrics (SDM), a novel lightweight framework for detecting Faithfulness Hallucinations -- events of severe deviations of LLMs responses from input contexts. We focus on a specific implementation of these L

Igor Halperin
arXiv · arXiv q-fin · 2023

Towards reducing hallucination in extracting information from financial reports using Large Language Models

For a financial analyst, the question and answer (Q\&A) segment of the company financial report is a crucial piece of information for various analysis and investment decisions. However, extracting valuable insights from the Q\&A section has posed considerable challenges as the conventional methods such as detailed reading and note-taking lack scalability and are susceptible to human errors, and Optical Character Reco

Bhaskarjit Sarmah, Tianjie Zhu, Dhagash Mehta, Stefano Pasquali
arXiv · arXiv · 2025

Semantic Faithfulness and Entropy Production Measures to Tame Your LLM Demons and Manage Hallucinations

Evaluating faithfulness of Large Language Models (LLMs) to a given task is a complex challenge. We propose two new unsupervised metrics for faithfulness evaluation using insights from information theory and thermodynamics. Our approach treats an LLM as a bipartite information engine where hidden layers act as a Maxwell demon controlling transformations of context $C $ into answer $A$ via prompt $Q$. We model Question

Igor Halperin
arXiv · arXiv q-fin · 2026

OOM-RL: Out-of-Money Reinforcement Learning Market-Driven Alignment for LLM-Based Multi-Agent Systems

The alignment of Multi-Agent Systems (MAS) for autonomous software engineering is constrained by evaluator epistemic uncertainty. Current paradigms, such as Reinforcement Learning from Human Feedback (RLHF) and AI Feedback (RLAIF), frequently induce model sycophancy, while execution-based environments suffer from adversarial "Test Evasion" by unconstrained agents. In this paper, we introduce an objective alignment pa

Kun Liu, Liqun Chen
arXiv · arXiv q-fin · 2026

PolySwarm: A Multi-Agent Large Language Model Framework for Prediction Market Trading and Latency Arbitrage

This paper presents PolySwarm, a novel multi-agent large language model (LLM) framework designed for real-time prediction market trading and latency arbitrage on decentralized platforms such as Polymarket. PolySwarm deploys a swarm of 50 diverse LLM personas that concurrently evaluate binary outcome markets, aggregating individual probability estimates through confidence-weighted Bayesian combination of swarm consens

Rajat M. Barot, Arjun S. Borkhatariya
arXiv · arXiv q-fin · 2026

AI Trading: Evaluating Large Language Models for Technical Market Analysis

Large Language Models (LLMs) have emerged as powerful tools for processing the heterogeneous information environments of modern financial markets. This paper presents a systematic, comparative evaluation of five prominent LLMs: GPT-4 Turbo, Claude 3 Opus, Gemini 1.5 Pro, Llama 3 70B, and the domain-specialized FinGPT, with respect to their capacity for technical market analysis. The evaluation spans four structured t

Geofrey Ntale
arXiv · arXiv q-fin · 2025

The New Quant: A Survey of Large Language Models in Financial Prediction and Trading

Large language models are reshaping quantitative investing by turning unstructured financial information into evidence-grounded signals and executable decisions. This survey synthesizes research with a focus on equity return prediction and trading, consolidating insights from domain surveys and more than fifty primary studies. We propose a task-centered taxonomy that spans sentiment and event extraction, numerical an

Weilong Fu
arXiv · arXiv q-fin · 2025

Beyond the Black Box: Interpretability of LLMs in Finance

Large Language Models (LLMs) exhibit remarkable capabilities across a spectrum of tasks in financial services, including report generation, chatbots, sentiment analysis, regulatory compliance, investment advisory, financial knowledge retrieval, and summarization. However, their intrinsic complexity and lack of transparency pose significant challenges, especially in the highly regulated financial sector, where interpr

Hariom Tatsat, Ariye Shater
arXiv · arXiv q-fin · 2024

A Survey of Large Language Models in Finance (FinLLMs)

Large Language Models (LLMs) have shown remarkable capabilities across a wide variety of Natural Language Processing (NLP) tasks and have attracted attention from multiple domains, including financial services. Despite the extensive research into general-domain LLMs, and their immense potential in finance, Financial LLM (FinLLM) research remains limited. This survey provides a comprehensive overview of FinLLMs, inclu

Jean Lee, Nicholas Stevens, Soyeon Caren Han, Minseok Song
arXiv · arXiv q-fin · 2023

ChatGPT-based Investment Portfolio Selection

In this paper, we explore potential uses of generative AI models, such as ChatGPT, for investment portfolio selection. Trusting investment advice from Generative Pre-Trained Transformer (GPT) models is a challenge due to model "hallucinations", necessitating careful verification and validation of the output. Therefore, we take an alternative approach. We use ChatGPT to obtain a universe of stocks from S&P500 market i

Oleksandr Romanko, Akhilesh Narayan, Roy H. Kwon
arXiv · arXiv q-fin · 2026

FinBench: Time-Gated Calibration and Uncertainty Benchmarking for Agentic Financial Forecasting

Large language models (LLMs) are increasingly used as components of agentic systems that observe, plan, and act. In finance, even "assistive" systems become decision-relevant once their outputs are used to size trades or allocate risk. A key failure mode is the confidence--competence gap: a model that is only slightly better than chance but consistently overconfident will, under typical bet-sizing rules, generate neg

Rishab Ghosh, Vinay Devarakonda
arXiv · arXiv q-fin · 2026

Large Language Models and Stock Investing: Is the Human Factor Required?

This paper investigates whether large language models (LLMs) can generate reliable stock market predictions. We evaluate four state-of-the-art models - ChatGPT, Gemini, DeepSeek, and Perplexity - across three prompting strategies: a naive query, a structured approach, and chain-of-thought reasoning. Our results show that LLM-generated recommendations are hindered by recurring reasoning failures, including financial m

Ricardo Crisostomo, Diana Mykhalyuk
arXiv · arXiv q-fin · 2025

Standard Benchmarks Fail -- Auditing LLM Agents in Finance Must Prioritize Risk

Standard benchmarks fixate on how well large language model (LLM) agents perform in finance, yet say little about whether they are safe to deploy. We argue that accuracy metrics and return-based scores provide an illusion of reliability, overlooking vulnerabilities such as hallucinated facts, stale data, and adversarial prompt manipulation. We take a firm position: financial LLM agents should be evaluated first and f

Zichen Chen, Jiaao Chen, Jianda Chen, Misha Sra
Wiki Entities · 1
Option Blackboard · 0
No Option Blackboard entries matched.
Encyclopedia · 1
Cards · 0
No cards matched.
← Back to Codex