Search
Search
Papers, wiki, Option Blackboard, encyclopedia, and cards.
Results for “RLHF” · papers 2 · wiki 2
Academic Papers · 2arXiv q-fin live 2 · desk corpus 1
arXiv · arXiv q-fin · 2026
The alignment of Multi-Agent Systems (MAS) for autonomous software engineering is constrained by evaluator epistemic uncertainty. Current paradigms, such as Reinforcement Learning from Human Feedback (RLHF) and AI Feedback (RLAIF), frequently induce model sycophancy, while execution-based environments suffer from adversarial "Test Evasion" by unconstrained agents. In this paper, we introduce an objective alignment pa…
Kun Liu, Liqun Chen
arXiv · arXiv q-fin · 2024
We introduce and make publicly available the NIFTY Financial News Headlines dataset, designed to facilitate and advance research in financial market forecasting using large language models (LLMs). This dataset comprises two distinct versions tailored for different modeling approaches: (i) NIFTY-LM, which targets supervised fine-tuning (SFT) of LLMs with an auto-regressive, causal language-modeling objective, and (ii)…
Raeid Saqur, Ken Kato, Nicholas Vinden, Frank Rudzicz
Option Blackboard · 0
No Option Blackboard entries matched.
← Back to Codex