Web Agents and Benchmarks - A Survey Focusing on Algorithmic Strategies and Performance Ratings

Lars Krupp, Daniel Geißler, Paul Lukowicz, Jakob Karolus · 2025

Interactive systems powered by artificial intelligence (AI) are gaining incredible momentum among everyday computing devices. Recent research works on so-called web agents, which are large language models (LLM) capable of interacting with the world wide web on their own, promise the possibility to develop universally applicable systems. These systems have the potential to fundamentally change how we interact with - and seek out - information on the web. In this work, we present a systematic review of 79 state-of-the-art web agents and benchmarks. We extend existing work in this domain through our categorization of components and algorithmic strategies employed by the agents and further present associated benchmarks exposing current shortcomings in global performance comparisons. We conclude with a call for increasing efforts on comparability among web agent performance ratings and future directions for web agent research.

Read the paper · More papers on PaperTik