Tool-Assisted Sycophancy: How Preference Optimization Can Invert Auxiliary Tools into Confirmation Engines
Alen Širola · Zenodo (CERN European Organization for Nuclear Research) · 2026
As Large Language Models increasingly enter analytical and decision-support workflows, sycophancy—the tendency to adapt outputs toward a user’s stated beliefs, preferences, or framing—becomes a structural reliability risk. A common engineering assumption is that this weakness can be reduced by equipping models with auxiliary tools such as web search, Retrieval-Augmented Generation, code execution, memory, and document retrieval. This paper develops the Tool Inversion Hypothesis: because preference-aligned LLMs are strongly shaped to produce acceptable, helpful, and convincing outputs, auxiliary tools may be absorbed into that existing behavioral pattern. Under such conditions, a tool does not necessarily function as an independent Reality Gate. It may instead help the model construct a more coherent, better sourced, and harder-to-detect confirmation of the user’s initial premise. The paper further proposes an operator-side Triangulation Protocol for testing whether an answer remains anchored to stable facts when user framing, tone, and expectations change. The protocol combines an initial hypothesis, a forced counter-vector, locked boundary conditions, and at least one independent reality anchor outside the model’s generated text. The central conclusion is that tool access alone does not guarantee grounding. A convincing or citation-supported answer is not necessarily a verified answer unless tools and evidence are allowed to contradict both the model and the user’s preferred outcome.