Search

Search

Papers, wiki, Option Blackboard, encyclopedia, and cards.

Results for “RLHF” · papers 2 · wiki 2
Academic Papers · 2arXiv q-fin live 2 · desk corpus 1
arXiv · arXiv q-fin · 2026

OOM-RL: Out-of-Money Reinforcement Learning Market-Driven Alignment for LLM-Based Multi-Agent Systems

The alignment of Multi-Agent Systems (MAS) for autonomous software engineering is constrained by evaluator epistemic uncertainty. Current paradigms, such as Reinforcement Learning from Human Feedback (RLHF) and AI Feedback (RLAIF), frequently induce model sycophancy, while execution-based environments suffer from adversarial "Test Evasion" by unconstrained agents. In this paper, we introduce an objective alignment pa

Kun Liu, Liqun Chen
arXiv · arXiv q-fin · 2024

NIFTY Financial News Headlines Dataset

We introduce and make publicly available the NIFTY Financial News Headlines dataset, designed to facilitate and advance research in financial market forecasting using large language models (LLMs). This dataset comprises two distinct versions tailored for different modeling approaches: (i) NIFTY-LM, which targets supervised fine-tuning (SFT) of LLMs with an auto-regressive, causal language-modeling objective, and (ii)

Raeid Saqur, Ken Kato, Nicholas Vinden, Frank Rudzicz
Wiki Entities · 2
Option Blackboard · 0
No Option Blackboard entries matched.
Encyclopedia · 1
Cards · 0
No cards matched.
← Back to Codex