← Discover MCPs and Agents
G
AgentAI & MLGitHub

GVA-Survey

Official repository of the paper "Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms"

Links

README

From the repo.

Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms

Corresponding author: Juncheng Li (junchengli@zju.edu.cn)

🔥 News

  • [December 10, 2024] We have developed an agent that automatically collects and analyzes the latest papers in the GVA field. It will update the Related Papers daily at 0:30 AM UTC+8.

  • [December 7, 2024] We have released a Chinese version of the survey, please click 中文版综述 to access!

  • [November 17, 2024] Our survey is available on the arXiv platform: https://arxiv.org/abs/2411.10943

📖 Table of Content

🤖 Introduction

Welcome to the GitHub repository for our survey paper titled "Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms". This repository includes all the resources, code, and references related to the paper. Our objective is to provide a comprehensive overview of Generalist Virtual Agents (GVAs), covering their definition, necessity, implementation approaches, evaluation methods, limitations and future directions. We aim to bridge the gap between theory and practice in GVA research, providing a systematic framework for future development in this field.


📚 Cited Papers

Here we list the most important references cited in our survey, organized by different sections. We particularly focus on works that have made substantial impact or proposed innovative methodologies.

Additionally, we note that some papers may be cited across multiple sections. For the convenience of researchers, we provide complete citation information under each section.

Section II: What is GVA?

alt text

Environment - Web

  • World of bits: an open-domain platform for web-based agents

    OpenAI, ICML, 2017 Paper

  • Reinforcement Learning on Web Interfaces using Workflow-Guided Exploration

    Stanford University, ICLR, 2018 Paper Star

  • WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

    Department of Computer Science, Princeton University, NeurIPS, 2022 Paper Website

  • WebArena: A Realistic Web Environment for Building Autonomous Agents

    Carnegie Mellon University, ICLR, 2024 Paper Website

  • VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

    Carnegie Mellon University, arXiv preprint arXiv:2401.13649, 2024 Paper

  • WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

    ServiceNow Research, arXiv preprint arXiv:2403.07718, 2024 Paper Star

  • Mind2Web: Towards a Generalist Agent for the Web

    The Ohio State University, NeurIPS, 2023 Paper Star Website

  • WebVLN: Vision-and-Language Navigation on Websites

    Australian Institute for Machine Learning, The University of Adelaide, AAAI, 2024 Paper Star

Environment - Application

  • AppAgent: Multimodal Agents as Smartphone Users

    Tencent, arXiv preprint arXiv:2312.13771, 2023 Paper Website

  • Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

    Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024 Paper Star

  • A Dataset for Interactive Vision Language Navigation with Unknown Command Feasibility

    Boston University, ECCV, 2022 Paper

  • AndroidInTheWild: A Large-Scale Dataset For Android Device Control

    DeepMind, NeurIPS, 2023 Paper

  • Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning

    University of California, Santa Barbara, EMNLP, 2023 Paper Star

  • Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding

    Harvard University, NeurIPS, 2018 Paper

  • Visual Programming: Compositional visual reasoning without training

    PRIOR @ Allen Institute for AI, CVPR, 2023 Paper

  • Toolformer: Language Models Can Teach Themselves to Use Tools

    Meta AI Research, NeurIPS, 2023 Paper

  • AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production

    DAMO Academy, Alibaba Group, arXiv preprint arXiv:2403.07952, 2024 Paper Website

Environment - Operating System

  • OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

    The University of Hong Kong, arXiv preprint arXiv:2404.07972, 2024 Paper Star Website

  • AgentStudio: A Toolkit for Building General Virtual Agents

    NTU, Singapore, arXiv preprint arXiv:2403.17918, 2024 Paper Website

  • MMAC-Copilot: Multi-modal Agent Collaboration Operating System Copilot

    University of Technology Sydney, arXiv preprint arXiv:2404.18074, 2024 Paper

  • UFO: A UI-Focused Agent for Windows OS Interaction

    Microsoft, arXiv preprint arXiv:2402.07939, 2024 Paper Star

Task - Command Task

  • AppAgent: Multimodal Agents as Smartphone Users

    Tencent, arXiv preprint arXiv:2312.13771, 2023 Paper Website

  • Android In The Wild: A Large-Scale Dataset For Android Device Control

    DeepMind, NeurIPS, 2023 Paper

  • UFO: A UI-Focused Agent for Windows OS Interaction

    Microsoft, arXiv preprint arXiv:2402.07939, 2024 Paper Star

  • From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces

    Google DeepMind, NeurIPS, 2023 Paper Star

  • WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

    Department of Computer Science, Princeton University, NeurIPS, 2022 Paper Website

Task - Query Task

  • ViperGPT: Visual Inference via Python Execution for Reasoning

    Columbia University, ICCV, 2023 Paper

  • Visual Programming: Compositional visual reasoning without training

    PRIOR @ Allen Institute for AI, CVPR, 2023 Paper

  • Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning

    University of California, Santa Barbara, EMNLP, 2023 Paper Star

  • HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

    Zhejiang University, NeurIPS, 2023 Paper Star

  • Mm-react: Prompting chatgpt for multimodal reasoning and action

    Microsoft Azure AI, arXiv preprint arXiv:2303.11381, 2023 Paper Website

  • WebVLN: Vision-and-Language Navigation on Websites

    Australian Institute for Machine Learning, The University of Adelaide, AAAI, 2024 Paper Star

Task - Dialogue Task

  • Windows Copilot Plus for PCs

    Microsoft, 2024 Website

  • Apple Intelligence Overview

    Apple, 2024 Website

  • NICE: Neural Image Commenting with Empathy

    Microsoft, EMNLP, 2021 Paper

  • Is ChatGPT Equipped with Emotional Dialogue Capabilities?

    Harbin Institute of Technology, China, arXiv preprint arXiv:2304.09582, 2023 Paper

Observation Space - Command Line Interface

  • AutoGPT

    OpenAI, 2024 Website

  • OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

    The University of Hong Kong, arXiv preprint arXiv:2404.07972, 2024 Paper Star Website

  • AgentBench: Evaluating LLMs as Agents

    Tsinghua University, ICLR, 2024 Paper Star Website

  • AgentStudio: A Toolkit for Building General Virtual Agents

    NTU, Singapore, arXiv preprint arXiv:2403.17918, 2024 Paper Website

  • InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback

    Department of Computer Science, Princeton University, NeurIPS, 2023 Paper Star Website

Observation Space - Document Object Model

  • World of bits: an open-domain platform for web-based agents

    OpenAI, ICML, 2017 Paper

  • Reinforcement Learning on Web Interfaces using Workflow-Guided Exploration

    Stanford University, ICLR, 2018 Paper Star

  • WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

    Department of Computer Science, Princeton University, NeurIPS, 2022 Paper Website

  • Mind2Web: Towards a Generalist Agent for the Web

    The Ohio State University, NeurIPS, 2023 Paper Star Website

  • AppAgent: Multimodal Agents as Smartphone Users

    Tencent, arXiv preprint arXiv:2312.13771, 2023 Paper Website

  • A data-driven approach for learning to control computers

    DeepMind, London, United Kingdom, ICML, 2022 Paper

Observation Space - Screen

  • You Only Look at Screens: Multimodal Chain-of-Action Agents

    School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University, arXiv preprint arXiv:2309.11436, 2024 Paper Star

  • SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

    Department of Computer Science and Technology, Nanjing University, arXiv preprint arXiv:2401.10935, 2024 Paper Star

  • WebVLN: Vision-and-Language Navigation on Websites

    Australian Institute for Machine Learning, The University of Adelaide, AAAI, 2024 Paper Star

  • From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces

    Google DeepMind, NeurIPS, 2023 Paper Star

  • GPT-4V(ision) is a Generalist Web Agent, if Grounded

    The Ohio State University, arXiv preprint arXiv:2401.01614, 2024 Paper Star Website

  • UFO: A UI-Focused Agent for Windows OS Interaction

    Microsoft, arXiv preprint arXiv:2402.07939, 2024 Paper Star

  • GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

    UC San Diego, arXiv preprint arXiv:2311.07562, 2023 Paper Star

  • Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

    Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024 Paper Star

Action Space - Keyboard

  • World of bits: an open-domain platform for web-based agents

    OpenAI, ICML, 2017 Paper

  • Mind2Web: Towards a Generalist Agent for the Web

    The Ohio State University, NeurIPS, 2023 Paper Star Website

  • WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

    Zhejiang University, arXiv preprint arXiv:2401.13919, 2024 Paper Star

  • WebArena: A Realistic Web Environment for Building Autonomous Agents

    Carnegie Mellon University, ICLR, 2024 Paper Website

  • WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

    ServiceNow Research, arXiv preprint arXiv:2403.07718, 2024 Paper Star

Action Space - Mouse

  • World of bits: an open-domain platform for web-based agents

    OpenAI, ICML, 2017 Paper

  • From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces

    Google DeepMind, NeurIPS, 2023 Paper Star

  • WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

    Zhejiang University, arXiv preprint arXiv:2401.13919, 2024 Paper Star

  • WebArena: A Realistic Web Environment for Building Autonomous Agents

    Carnegie Mellon University, ICLR, 2024 Paper Website

  • Mind2Web: Towards a Generalist Agent for the Web

    The Ohio State University, NeurIPS, 2023 Paper Star Website

  • WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

    ServiceNow Research, arXiv preprint arXiv:2403.07718, 2024 Paper Star

Action Space - Touchscreen

  • AppAgent: Multimodal Agents as Smartphone Users

    Tencent, arXiv preprint arXiv:2312.13771, 2023 Paper Website

  • AndroidInTheWild: A Large-Scale Dataset For Android Device Control

    DeepMind, NeurIPS, 2023 Paper

  • You Only Look at Screens: Multimodal Chain-of-Action Agents

    School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University, arXiv preprint arXiv:2309.11436, 2024 Paper Star

  • Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

    Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024 Paper Star

Action Space - Others

  • HyperPalm: DNN-based hand gesture recognition interface for intelligent communication with quadruped robot in 3D space

    Intelligent Space Robotics Laboratory, Skoltech, SMC, 2022 Paper

Section III: Why we need GVA?

alt text

Perspective of AI and Machine Learning

  • CogAgent: A Visual Language Model for GUI Agents

    Tsinghua University, arXiv preprint arXiv:2312.08914, 2023 Paper Star

  • You Only Look at Screens: Multimodal Chain-of-Action Agents

    School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University, arXiv preprint arXiv:2309.11436, 2024 Paper Star

  • SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

    Department of Computer Science and Technology, Nanjing University, arXiv preprint arXiv:2401.10935, 2024 Paper Star

  • From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces

    Google DeepMind, NeurIPS, 2023 Paper Star

  • GPT-4V(ision) is a Generalist Web Agent, if Grounded

    The Ohio State University, arXiv preprint arXiv:2401.01614, 2024 Paper Star Website

  • Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

    Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024 Paper Star

Perspective of Interaction

  • Inferring Rewards from Language in Context

    University of California, Berkeley, ACL, 2022 Paper

  • Deep Reinforcement Learning from Human Preferences

    OpenAI, NeurIPS, 2017 Paper

  • Apple Intelligence Overview

    Apple, 2024 Website

Perspective of Agent Applications

  • Distributed and reactive query planning in R-MAGIC: an agent-based multimedia retrieval system

    IEEE Transactions on Knowledge and Data Engineering, 2004 Paper

  • Rap: Retrieval-augmented planning with contextual memory for multimodal llm agents

    Panasonic Connect Co., Ltd., Japan, arXiv preprint arXiv:2402.03610, 2024 Paper

  • Plan4mc: Skill reinforcement learning and planning for open-world minecraft tasks

    School of Computer Science, Peking University, arXiv preprint arXiv:2303.16563, 2023 Paper Website

  • Autoact: Automatic agent learning from scratch via self-planning

    Zhejiang University, arXiv preprint arXiv:2401.05268, 2024 Paper Star

  • Self-Refine: Iterative Refinement with Self-Feedback

    Language Technologies Institute, Carnegie Mellon University, NeurIPS, 2023 Paper Star

  • Reflexion: language agents with verbal reinforcement learning

    Northeastern University, NeurIPS, 2023 Paper Star

  • Eureka: Human-Level Reward Design via Coding Large Language Models

    NVIDIA, ICLR, 2024 Paper Star

  • WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

    ServiceNow Research, arXiv preprint arXiv:2403.07718, 2024 Paper Star

  • AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production

    DAMO Academy, Alibaba Group, arXiv preprint arXiv:2403.07952, 2024 Paper Website

Section IV: How to implement GVA?

alt text

Environment - Off-line

  • World of bits: an open-domain platform for web-based agents

    OpenAI, ICML, 2017 Paper

  • Reinforcement Learning on Web Interfaces using Workflow-Guided Exploration

    Stanford University, ICLR, 2018 Paper Star

  • SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

    Department of Computer Science and Technology, Nanjing University, arXiv preprint arXiv:2401.10935, 2024 Paper Star

  • AndroidInTheWild: A Large-Scale Dataset For Android Device Control

    DeepMind, NeurIPS, 2023 Paper

  • Mind2Web: Towards a Generalist Agent for the Web

    The Ohio State University, NeurIPS, 2023 Paper Star Website

  • A Dataset for Interactive Vision Language Navigation with Unknown Command Feasibility

    Boston University, ECCV, 2022 Paper

  • META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI

    Shanghai Jiao Tong University, EMNLP, 2022 Paper Website

Environment - On-line

  • WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

    Department of Computer Science, Princeton University, NeurIPS, 2022 Paper Website

  • WebArena: A Realistic Web Environment for Building Autonomous Agents

    Carnegie Mellon University, ICLR, 2024 Paper Website

  • VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

    Carnegie Mellon University, arXiv preprint arXiv:2401.13649, 2024 Paper

Environment - Updated

  • WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

    ServiceNow Research, arXiv preprint arXiv:2403.07718, 2024 Paper Star

  • WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

    Zhejiang University, arXiv preprint arXiv:2401.13919, 2024 Paper Star

  • AppAgent: Multimodal Agents as Smartphone Users

    Tencent, arXiv preprint arXiv:2312.13771, 2023 Paper Website

  • Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

    Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024 Paper Star

  • OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

    The University of Hong Kong, arXiv preprint arXiv:2404.07972, 2024 Paper Star Website

  • AgentStudio: A Toolkit for Building General Virtual Agents

    NTU, Singapore, arXiv preprint arXiv:2403.17918, 2024 Paper Website

  • UFO: A UI-Focused Agent for Windows OS Interaction

    Microsoft, arXiv preprint arXiv:2402.07939, 2024 Paper Star

Model - Retriever-based Agent

  • Distributed and reactive query planning in R-MAGIC: an agent-based multimedia retrieval system

    nan, IEEE Transactions on Knowledge and Data Engineering, 2004 Paper

  • Mobile agents and their use for information retrieval: a brief overview and an elaborate case study

    nan, IEEE Network, 2002 Paper

  • SAIRE—a scalable agent-based information retrieval engine

    nan, Proceedings of the first international conference on Autonomous agents, 1997 Paper

  • ACQUIRE: agent-based complex query and information retrieval engine

    nan, Proceedings of the first international joint conference on Autonomous agents and multiagent systems: part 2, 2002 Paper

  • Thought-retriever: Don’t just retrieve raw data, retrieve thoughts

    University of Illinois at Urbana-Champaign, ICLR Workshop, 2024 Paper

Model - LLM-based Agent

  • Agent-pro: Learning to evolve via policy-level reflection and optimization

    Zhejiang University, arXiv preprint arXiv:2402.17574, 2024 Paper Star

  • Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning

    University of California, Santa Barbara, EMNLP, 2023 Paper Star

  • HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

    Zhejiang University, NeurIPS, 2023 Paper Star

  • Mm-react: Prompting chatgpt for multimodal reasoning and action

    Microsoft Azure AI, arXiv preprint arXiv:2303.11381, 2023 Paper Website

  • Mind2Web: Towards a Generalist Agent for the Web

    The Ohio State University, NeurIPS, 2023 Paper Star Website

  • AllTogether: Investigating the Efficacy of Spliced Prompt for Web Navigation using Large Language Models

    nan, arXiv preprint arXiv:2310.18331, 2023 Paper

  • Building Cooperative Embodied Agents Modularly with Large Language Models

    University of Massachusetts Amherst, ICLR, 2024 Paper Website

  • RoboGPT: an intelligent agent of making embodied long-term decisions for daily instruction tasks

    State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China, arXiv preprint arXiv:2311.15649, 2024 Paper

  • AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation

    University of Hong Kong, arXiv preprint arXiv:2312.13010, 2024 Paper

  • Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering

    Tokyo Institute of Technology, arXiv preprint arXiv:2407.20563, 2024 Paper

  • Modular Visual Question Answering via Code Generation

    UC Berkeley, ACL, 2023 Paper Star

  • ViperGPT: Visual Inference via Python Execution for Reasoning

    Columbia University, ICCV, 2023 Paper

Model - MLLM-based Agent

  • GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

    UC San Diego, arXiv preprint arXiv:2311.07562, 2023 Paper Star

  • Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

    Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024 Paper Star

  • Autonomous Evaluation and Refinement of Digital Agents

    UC Berkeley, arXiv preprint arXiv:2404.06474, 2024 Paper Star

  • Clova: A closed-loop visual assistant with tool usage and update

    Peking University, CVPR, 2024 Paper Website

  • Agent smith: A single image can jailbreak one million multimodal llm agents exponentially fast

    Sea AI Lab, arXiv preprint arXiv:2402.08567, 2024 Paper Star

  • WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

    Zhejiang University, arXiv preprint arXiv:2401.13919, 2024 Paper Star

  • You Only Look at Screens: Multimodal Chain-of-Action Agents

    School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University, arXiv preprint arXiv:2309.11436, 2024 Paper Star

  • VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

    Carnegie Mellon University, arXiv preprint arXiv:2401.13649, 2024 Paper

  • CogAgent: A Visual Language Model for GUI Agents

    Tsinghua University, arXiv preprint arXiv:2312.08914, 2023 Paper Star

Model - VLA-based Agent

  • Rt-2: Vision-language-action models transfer web knowledge to robotic control

    Google DeepMind, arXiv preprint arXiv:2307.15818, 2023 Paper Website

  • Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action

    UC Berkeley, Conference on robot learning, 2023 Paper

  • Octo: An open-source generalist robot policy

    UC Berkeley, arXiv preprint arXiv:2405.12213, 2024 Paper Website

  • OpenVLA: An Open-Source Vision-Language-Action Model

    Stanford University, arXiv preprint arXiv:2406.09246, 2024 Paper Website

  • 3d-vla: A 3d vision-language-action generative world model

    University of Massachusetts Amherst, arXiv preprint arXiv:2403.09631, 2024 Paper Website

  • Palm-e: An embodied multimodal language model

    Robotics at Google, arXiv preprint arXiv:2303.03378, 2023 Paper Website

  • Actra: Optimized Transformer Architecture for Vision-Language-Action Models in Robot Learning

    The Chinese University of Hong Kong, arXiv preprint arXiv:2408.01147, 2024 Paper

Strategy - Adaption Strategy

  • Mm-react: Prompting chatgpt for multimodal reasoning and action

    Microsoft Azure AI, arXiv preprint arXiv:2303.11381, 2023 Paper Website

  • Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models

    Singapore Management University, arXiv preprint arXiv:2305.04091, 2023 Paper Star

  • Expertprompting: Instructing large language models to be distinguished experts

    University of Science and Technology of China, arXiv preprint arXiv:2305.14688, 2023 Paper Star

  • Self-refine: Iterative refinement with self-feedback

    Language Technologies Institute, Carnegie Mellon University, NeurIPS, 2024 Paper Star

  • Pangu-coder2: Boosting large language models for code with ranking feedback

    Huawei Cloud Co., Ltd., arXiv preprint arXiv:2307.14936, 2023 Paper

  • De-fine: Decomposing and refining visual programs with auto-feedback

    Zhejiang University, arXiv preprint arXiv:2311.12890, 2023 Paper

  • A survey on the memory mechanism of large language model based agents

    Gaoling School of Artificial Intelligence, Renmin University of China, arXiv preprint arXiv:2404.13501, 2024 Paper Star

  • Rap: Retrieval-augmented planning with contextual memory for multimodal llm agents

    Panasonic Connect Co., Ltd., Japan, arXiv preprint arXiv:2402.03610, 2024 Paper

  • Chatdb: Augmenting llms with databases as their symbolic memory

    Tsinghua University, arXiv preprint arXiv:2306.03901, 2023 Paper Website

  • Memorybank: Enhancing large language models with long-term memory

    Sun Yat-Sen University, AAAI, 2024 Paper Star

Strategy - Fine-tuning Strategy

  • CoEvol: Constructing Better Responses for Instruction Finetuning through Multi-Agent Cooperation

    University of Macau, arXiv preprint arXiv:2406.07054, 2024 Paper Star

  • Reflection-Reinforced Self-Training for Language Agents

    University of California, Los Angeles, arXiv preprint arXiv:2406.01495, 2024 Paper Star

  • Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training

    Columbia University, arXiv preprint arXiv:2406.00222, 2024 Paper

  • Enhancing the General Agent Capabilities of Low-Parameter LLMs through Tuning and Multi-Branch Reasoning

    Huazhong University of Science and Technology, arXiv preprint arXiv:2403.19962, 2024 Paper

  • Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models

    Department of Automation, University of Science and Technology of China, arXiv preprint arXiv:2403.12881, 2024 Paper Star

Strategy - Reinforcement Learning Strategy

  • Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning

    UC Berkeley, arXiv preprint arXiv:2405.10292, 2024 Paper Star

  • Search Beyond Queries: Training Smaller Language Models for Web Interactions via Reinforcement Learning

    University of Kentucky, arXiv preprint arXiv:2404.10887, 2024 Paper

  • Juewu-mc: Playing minecraft with sample-efficient hierarchical reinforcement learning

    Tencent AI Lab, Shenzhen, China, arXiv preprint arXiv:2112.04907, 2021 Paper

  • Towards robust and domain agnostic reinforcement learning competitions: MineRL 2020

    CMU, NeurIPS, 2021 Paper

  • Plan4mc: Skill reinforcement learning and planning for open-world minecraft tasks

    School of Computer Science, Peking University, arXiv preprint arXiv:2303.16563, 2023 Paper Website

  • Pok'eLLMon: A Human-Parity Agent for Pok'emon Battles with Large Language Models

    Georgia Institute of Technology, arXiv preprint arXiv:2402.01118, 2024 Paper Star Website

  • Scaling instructable agents across many simulated worlds

    Google DeepMind, arXiv preprint arXiv:2404.10179, 2024 Paper

  • Mental Modeling of Reinforcement Learning Agents by Language Models

    University of Hamburg, arXiv preprint arXiv:2406.18505, 2024 Paper Star

  • LLMSat: A Large Language Model-Based Goal-Oriented Agent for Autonomous Space Exploration

    University of Toronto, arXiv preprint arXiv:2405.01392, 2024 Paper

Strategy - Cooperation or Competition Strategy

  • CMAT: A Multi-Agent Collaboration Tuning Framework for Enhancing Small Language Models

    East China Jiaotong University, arXiv preprint arXiv:2404.01663, 2024 Paper Star

  • Autoact: Automatic agent learning from scratch via self-planning

    Zhejiang University, arXiv preprint arXiv:2401.05268, 2024 Paper Star

  • ProAgent: building proactive cooperative agents with large language models

    The Chinese University of Hong Kong, Shenzhen, AAAI, 2024 Paper Website

  • Chatllm network: More brains, more intelligence

    Beijing University of Posts and Telecommunications, arXiv preprint arXiv:2304.12998, 2023 Paper

  • Encouraging divergent thinking in large language models through multi-agent debate

    Tsinghua University, arXiv preprint arXiv:2305.19118, 2023 Paper Star

  • Improving factuality and reasoning in language models through multiagent debate

    MIT CSAIL, arXiv preprint arXiv:2305.14325, 2023 Paper Website

  • Chateval: Towards better llm-based evaluators through multi-agent debate

    Tsinghua University, arXiv preprint arXiv:2308.07201, 2023 Paper Star

  • How susceptible are llms to logical fallacies?

    Department of Computer Science, George Mason University, arXiv preprint arXiv:2308.09853, 2023 Paper Star

Section V: How to evaluate GVA?

alt text

Overall Evaluation

  • WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

    Department of Computer Science, Princeton University, NeurIPS, 2022 Paper Website

  • Mind2Web: Towards a Generalist Agent for the Web

    The Ohio State University, NeurIPS, 2023 Paper Star Website

  • WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

    ServiceNow Research, arXiv preprint arXiv:2403.07718, 2024 Paper Star

  • GUI Odyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices

    OpenGVLab, Shanghai AI Laboratory, arXiv preprint arXiv:2406.08451, 2024 Paper Star

  • WebSRC: A Dataset for Web-Based Structural Reading Comprehension

    Shanghai Jiao Tong University, EMNLP, 2021 Paper Star

  • WebArena: A Realistic Web Environment for Building Autonomous Agents

    Carnegie Mellon University, ICLR, 2024 Paper Website

  • VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

    Carnegie Mellon University, arXiv preprint arXiv:2401.13649, 2024 Paper

  • Mapping Natural Language Instructions to Mobile UI Action Sequences

    Google Research, Mountain View, CA, 94043, ACL, 2020 Paper Star

  • Grounding Open-Domain Instructions to Automate Web Support Tasks

    Stanford University, NAACL, 2021 Paper Star

Detail Evaluation - Step-wise Evaluation

  • World of bits: an open-domain platform for web-based agents

    OpenAI, ICML, 2017 Paper

  • Mind2Web: Towards a Generalist Agent for the Web

    The Ohio State University, NeurIPS, 2023 Paper Star Website

  • Mapping Natural Language Instructions to Mobile UI Action Sequences

    Google Research, Mountain View, CA, 94043, ACL, 2020 Paper Star

  • META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI

    Shanghai Jiao Tong University, EMNLP, 2022 Paper Website

  • AndroidInTheWild: A Large-Scale Dataset For Android Device Control

    DeepMind, NeurIPS, 2023 Paper

  • Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

    Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024 Paper Star

  • WebSRC: A Dataset for Web-Based Structural Reading Comprehension

    Shanghai Jiao Tong University, EMNLP, 2021 Paper Star

  • WebVLN: Vision-and-Language Navigation on Websites

    Australian Institute for Machine Learning, The University of Adelaide, AAAI, 2024 Paper Star

  • Room2Room: Enabling Life-Size Telepresence in a Projected Augmented Reality Environment

    Microsoft, CSCW, 2016 Paper

  • OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web

    Carnegie Mellon University, arXiv preprint arXiv:2402.17553, 2024 Paper Website

Detail Evaluation - Set Inclusion Evaluation

  • Grounding Open-Domain Instructions to Automate Web Support Tasks

    Stanford University, NAACL, 2021 Paper Star

  • META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI

    Shanghai Jiao Tong University, EMNLP, 2022 Paper Website

  • VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

    Carnegie Mellon University, arXiv preprint arXiv:2401.13649, 2024 Paper

  • WebQA: Multihop and Multimodal QA

    Carnegie Mellon University, CVPR, 2022 Paper Website

  • WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

    Department of Computer Science, Princeton University, NeurIPS, 2022 Paper Website

  • Sequence to sequence learning with neural networks

    Google, NeurIPS, 2014 Paper

  • A Dataset for Interactive Vision Language Navigation with Unknown Command Feasibility

    Boston University, ECCV, 2022 Paper

Detail Evaluation - Multi-dimensional Evaluation

  • OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

    The University of Hong Kong, arXiv preprint arXiv:2404.07972, 2024 Paper Star Website

  • WebVLN: Vision-and-Language Navigation on Websites

    Australian Institute for Machine Learning, The University of Adelaide, AAAI, 2024 Paper Star

  • ChatDev: Communicative Agents for Software Development

    Tsinghua University, arXiv preprint arXiv:2307.07924, 2023 Paper Star

  • Online Adaptation of Language Models with a Memory of Amortized Contexts

    KAIST, arXiv preprint arXiv:2403.04317, 2024 Paper Star

  • Voyager: An Open-Ended Embodied Agent with Large Language Models

    NVIDIA, arXiv preprint arXiv:2305.16291, 2023 Paper Website

  • GUI Odyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices

    OpenGVLab, Shanghai AI Laboratory, arXiv preprint arXiv:2406.08451, 2024 Paper Star

  • Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies

    Olin College of Engineering, ICML, 2023 Paper

  • MemoChat: Tuning LLMs to Use Memos for Consistent Long-Range Open-Domain Conversation

    University of Warwick, arXiv preprint arXiv:2308.08239, 2023 Paper Star

  • Secrets of RLHF in Large Language Models Part I: PPO

    Fudan NLP Group, arXiv preprint arXiv:2307.04964, 2023 Paper Star

  • WebSRC: A Dataset for Web-Based Structural Reading Comprehension

    Shanghai Jiao Tong University, EMNLP, 2021 Paper Star

  • A Dataset for Interactive Vision Language Navigation with Unknown Command Feasibility

    Boston University, ECCV, 2022 Paper

Human Evaluation

  • WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

    Zhejiang University, arXiv preprint arXiv:2401.13919, 2024 Paper Star

  • AndroidInTheWild: A Large-Scale Dataset For Android Device Control

    DeepMind, NeurIPS, 2023 Paper

  • Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

    Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024 Paper Star

MLLM-based Evaluation

  • Gpt-4 technical report

    OpenAI, arXiv preprint arXiv:2303.08774, 2023 Paper

  • VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

    Carnegie Mellon University, arXiv preprint arXiv:2401.13649, 2024 Paper

  • WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

    Zhejiang University, arXiv preprint arXiv:2401.13919, 2024 Paper Star

  • GUI-WORLD: A Dataset for GUI-oriented Multimodal LLM-based Agents

    Huazhong University of Science and Technology, arXiv preprint arXiv:2406.10819, 2024 Paper Website

Section VI: Limitations

alt text

Unrealistic Environment and Dataset

  • OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

    The University of Hong Kong, arXiv preprint arXiv:2404.07972, 2024 Paper Star Website

  • Mapping Natural Language Instructions to Mobile UI Action Sequences

    Google Research, Mountain View, CA, 94043, ACL, 2020 Paper Star

  • AppAgent: Multimodal Agents as Smartphone Users

    Tencent, arXiv preprint arXiv:2312.13771, 2023 Paper Website

Insufficient Transferability

  • GPT-4V(ision) is a Generalist Web Agent, if Grounded

    The Ohio State University, arXiv preprint arXiv:2401.01614, 2024 Paper Star Website

  • Visually Grounded Language Learning: a review of language games, datasets, tasks, and models

    Heriot-Watt University, Journal of Artificial Intelligence Research, 2024 Paper

Limited Long-Sequence Decision-Making

  • Memorybank: Enhancing large language models with long-term memory

    Sun Yat-Sen University, AAAI, 2024 Paper Star

  • Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

    Skywork AI, arXiv preprint arXiv:2406.14283, 2024 Paper

  • Reasoning with language model is planning with world model

    UC San Diego, arXiv preprint arXiv:2305.14992, 2023 Paper Star

Heightened Security Concerns

  • LLaVA-phi: Efficient Multi-Modal Assistant with Small Language Model

    Midea Group, arXiv preprint arXiv:2401.02330, 2024 Paper Star

  • Smoothquant: Accurate and efficient post-training quantization for large language models

    Massachusetts Institute of Technology, ICML, 2023 Paper Star

  • AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration

    MIT, Proceedings of Machine Learning and Systems, 2024 Paper

Section VII: Future

alt text

From Individual to Systematic

  • LLM multi-agent systems: Challenges and open problems

    University of California, Irvine, arXiv preprint arXiv:2402.03578, 2024 Paper

  • Combining multi-agent systems and Artificial Intelligence of Things: Technical challenges and gains

    Belfort Montbeliard University of Technology, Internet of Things, 2024 Paper

  • Multi-agent systems in Peer-to-Peer energy trading: A comprehensive survey

    University of Galway, Engineering Applications of Artificial Intelligence, 2024 Paper

From Virtual to Physical

  • A survey on robotics with foundation models: toward embodied ai

    Midea Group, arXiv preprint arXiv:2402.02385, 2024 Paper

  • Embodied AI with Two Arms: Zero-shot Learning, Safety and Modularity

    Google Deep Mind Robotics, arXiv preprint arXiv:2404.03570, 2024 Paper

🔍 Related Papers

We are committed to offering researchers the latest advancements in the field. By regularly reviewing and evaluating recent research studies, we ensure that the list of papers stays up-to-date.

⚠️ The paper analysis may not be accurate and is for reference only!

DatePaperEnvironmentContributionAvailable Link
Apr 2026Parallax: Why AI Agents That Think Must Never Act





• Affiliation: Independent Researcher
• Agent Name: OpenParallax, Base Model: Claude Haiku 4.5, Strategy: Prompt
• Benchmark Name: Assume-Compromise Evaluation, Task Number: 337, Dataset Source: synthetic
Apr 2026Towards Long-horizon Agentic Multimodal Search



• Affiliation: Renmin University of China
• Agent Name: LMM-Searcher, Base Model: Qwen3-VL-Thinking-30A3B, Strategy: SFT
Apr 2026QuarkMedSearch: A Long-Horizon Deep Search Agent for Exploring Medical Intelligence



• Affiliation: Alibaba
• Agent Name: QuarkMedSearch, Base Model: Tongyi DeepResearch 30B-A3B, Strategy: Two-stage SFT and RLVR
• Benchmark Name: QuarkMedSearch Benchmark, Task Number: 140, Dataset Source: LLM synthesis and open-source benchmark filtering with expert human verification
Apr 2026From Imitation to Discrimination: Progressive Curriculum Learning for Robust Web Navigation


• Affiliation: Beijing Jiaotong University
• Agent Name: Triton-GRPO-32B, Base Model: Qwen2.5-Coder-32B-Instruct, Strategy: Progressive training curriculum (SFT, ORPO for discrimination, and GRPO for consistency)
Apr 2026Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning



• Affiliation: NVIDIA
• Agent Name: Nemotron 3 Super, Base Model: Nemotron 3 Super 120B-A12B, Strategy: SFT, RL (RLVR, SWE-RL, RLHF)
Apr 2026WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents



• Affiliation: National University of Singapore
• Agent Name: WebAgentGuard, Base Model: Qwen3-VL-Instruct, Strategy: reasoning-intensive supervised fine-tuning (SFT) and Group Relative Policy Optimization (GRPO)
• Benchmark Name: synthetic multimodal dataset, Task Number: 1000, Dataset Source: LLM synthesis (GPT-5)
Apr 2026AlphaEval: Evaluating Agents in Production




• Affiliation: SII
• Benchmark Name: AlphaEval, Task Number: 94, Dataset Source: sourced from seven companies deploying AI agents in their core business
• Paper Number: 106
Apr 2026The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break


• Affiliation: University of Wisconsin–Madison
• Benchmark Name: HORIZON, Task Number: 700+, Dataset Source: The benchmark is constructed by systematically extending tasks from existing suites like WebArena, AgentBench, MAC-SQL, and Isaac Sim using controlled horizon extension methods.
Apr 2026AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection



• Affiliation: University of Illinois Urbana-Champaign
• Agent Name: AnyPoC, Base Model: Claude Sonnet 4.5, Claude Sonnet 4.6, Opus 4.5, Opus 4.6, GPT-5.3-Codex, Strategy: Multi-agent framework featuring dedicated Analyzer, Generator, and Checker agents with an iterative synthesis-execution-validation loop and a self-evolving knowledge base.
Apr 2026ProbeLogits: Kernel-Level LLM Inference Primitives for AI-Native Operating Systems




• Affiliation: Independent Researcher
• Agent Name: Anima OS agents, Base Model: Qwen2.5-7B-Instruct, Strategy: ProbeLogits
• Benchmark Name: OS action benchmark, Task Number: 260, Dataset Source: human-annotated
Apr 2026ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection



• Affiliation: Singapore Management University
• Agent Name: CLAWGUARD, Base Model: DeepSeek-V3.2, GLM-5, Kimi-K2.5, MiniMax-M2.5, Qwen3.5-397B-A17B, Strategy: Rule-based enforcement, Sanitization, Prompt-based Rule Induction
Apr 2026ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents






• Affiliation: Zhejiang University
• Agent Name: ClawGUI-2B, Base Model: MAI-UI-2B, Strategy: GiGPO with a Process Reward Model for dense step-level supervision
Apr 2026Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks



• Affiliation: Princeton University
• Agent Name: AggAgent, Base Model: GLM-4.7-Flash, Qwen3.5-122B, MiniMax-M2.5, Strategy: Prompt
Apr 2026Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation


• Affiliation: Beijing Institute of Technology
• Agent Name: MISE, Base Model: LLaMA3-8B-Instruct, Strategy: Reinforcement Learning with Mutual Information Self-Evaluation
• Agent Name: MISE, Base Model: Qwen2-7B-Instruct, Strategy: Reinforcement Learning with Mutual Information Self-Evaluation
• Agent Name: MISE, Base Model: Gemma2-9B-Instruct, Strategy: Reinforcement Learning with Mutual Information Self-Evaluation
Apr 2026From Translation to Superset: Benchmark-Driven Evolution of a Production AI Agent from Rust to Python




• Affiliation: JP Morgan Chase & Co.
• Agent Name: CODEXCLI, Base Model: GPT-5.4, Strategy: Prompt
Apr 2026Escaping the Context Bottleneck: Active Context Curation for LLM Agents via Reinforcement Learning


• Affiliation: Tongji University
• Agent Name: ActiveContext, Base Model: Qwen-2.5-7B-Instruct, Strategy: RL
Apr 2026Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization




• Affiliation: Shandong University
• Agent Name: TIPO, Base Model: Qwen2.5VL-3B, Strategy: Trajectory Induced Preference Optimization (preference-intensity weighting and padding gating)
• Benchmark Name: Privacy Preference Dataset, Task Number: 151, Dataset Source: human demonstration
Apr 2026CocoaBench: Evaluating Unified Digital Agents in the Wild




• Affiliation: UC San Diego
• Agent Name: COCOA-AGENT, Base Model: GPT-5.4, Claude Sonnet 4.6, Gemini-3.1-pro, Gemini-Flash-3.0, Kimi-k2.5, Qwen3.5-397B-A13B, Strategy: ReAct-based scaffold
• Benchmark Name: COCOABENCH, Task Number: 153, Dataset Source: human-authored
Apr 2026Sema Code: Decoupling AI Coding Agents into Programmable, Embeddable Infrastructure






• Affiliation: Midea AIRC
• Agent Name: Sema Code, Base Model: Anthropic, OpenAI-compatible, Code Llama, DeepSeek-Coder, GLM-5, Qwen3-235B, Gemini 2.5 Pro, Strategy: ReAct, multi-agent collaborative scheduling, adaptive context compression, Prompt
Apr 2026WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent Benchmark


• Affiliation: Tencent BAC
• Benchmark Name: WebForge-Bench, Task Number: 934, Dataset Source: LLM synthesis
Apr 2026AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web



• Affiliation: Carnegie Mellon University
• Benchmark Name: AgentWebBench, Task Number: 4, Dataset Source: ClueWeb22-B, MS MARCO, ORBIT, DeepResearchGym
Apr 2026Mem$^2$Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation



• Affiliation: Beihang University
• Agent Name: Mem2Evolve, Base Model: GPT-5-chat, Strategy: Prompt
Apr 2026The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents


• Affiliation: University of Wisconsin–Madison
• Benchmark Name: OS-BLIND, Task Number: 300, Dataset Source: human-crafted
Apr 2026Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?




• Affiliation: Soochow University
• Benchmark Name: Agent2RL-Bench, Task Number: 6, Dataset Source:
Apr 2026ClawVM: Harness-Managed Virtual Memory for Stateful Tool-Using LLM Agents



• Affiliation: Independent Researcher
• Agent Name: ClawVM, Base Model: Claude, Strategy: Prompt
Apr 2026STARS: Skill-Triggered Audit for Request-Conditioned Invocation Safety in Agent Systems



• Affiliation: Shenzhen University
• Benchmark Name: SIA-Bench, Task Number: 3000, Dataset Source: LLM synthesis
Apr 2026The Amazing Agent Race: Strong Tool Users, Weak Navigators




• Affiliation: University of Minnesota Twin Cities
• Benchmark Name: THE AMAZING AGENT RACE (AAR), Task Number: 1400, Dataset Source: Wikipedia
Apr 2026HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks





• Affiliation: Stanford University
• Agent Name: Qwen-3.5-Kinetic-SFT, Base Model: Qwen-3.5-27B, Strategy: SFT
• Benchmark Name: HEALTHADMINBENCH, Task Number: 135, Dataset Source: expert-designed tasks derived from hospital shadowing
Apr 2026EE-MCP: Self-Evolving MCP-GUI Agents via Automated Environment Generation and Experience Learning


• Affiliation: University College London
• Agent Name: EE-MCP, Base Model: Qwen3-VL-8B, Strategy: SFT, Trajectory Distillation, Experience Bank, Rejection Sampling, Profile-Guided Task Generation
Apr 2026Agentic Compilation: Mitigating the LLM Rerun Crisis for Minimized-Inference-Cost Web Automation


• Affiliation: Selfotix
• Agent Name: One-Shot Agentic Compilation, Base Model: Claude Opus, Claude Sonnet, GPT-4, Qwen, Strategy: Prompt
Apr 2026NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration


• Affiliation: Hiroshima University
• Benchmark Name: NetAgentBench, Task Number: 5, Dataset Source: literature-grounded tasks (CCNA, CCNP, and CCIE certification guides)
Apr 2026HearthNet: Edge Multi-Agent Orchestration for Smart Homes




• Affiliation: Imperial College London
• Agent Name: HearthNet, Base Model: Claude Opus 4.6, Gemini 3 Pro, Strategy: Prompt
Apr 2026MobiFlow: Real-World Mobile Agent Benchmarking through Trajectory Fusion



• Affiliation: Shanghai Jiao Tong University
• Benchmark Name: MobiFlow, Task Number: 240, Dataset Source: human demonstration
Apr 2026OpeFlo: Automated UX Evaluation via Simulated Human Web Interaction with GUI Grounding



• Affiliation: University College London
• Agent Name: OpenFlo, Base Model: Gemini-3-Pro, Strategy: Prompt
Apr 2026Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization



• Affiliation: Shanghai Jiao Tong University
• Benchmark Name: Agent Humanization Benchmark (AHB), Task Number: 21, Dataset Source: human demonstration
Apr 2026Tuning Qwen2.5-VL to Improve Its Web Interaction Skills



• Affiliation: Aalto University
• Agent Name: Qwen2.5-VL agent, Base Model: Qwen2.5-VL-32B, Strategy: two-stage fine-tuning, self-distillation
• Benchmark Name: single-click_bench, Task Number: 430, Dataset Source: automatically derived from visible text or metadata
Apr 2026Event-Driven Temporal Graph Networks for Asynchronous Multi-Agent Cyber Defense in NetForge_RL




• Affiliation: None stated
• Agent Name: CT-GMARL, Base Model: all-MiniLM-L6-v2, Strategy: Continuous-Time Graph Multi-Agent Reinforcement Learning with Neural ODE-RNN and Multi-Head Graph Attention
• Benchmark Name: NetForge_RL, Task Number: , Dataset Source: Windows Event XML logs and live Vulhub container payloads
Apr 2026From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models

• Affiliation: Independent Researcher
• Paper Number: 65
Apr 2026Confidence Without Competence in AI-Assisted Knowledge Work





• Affiliation: The University of Hong Kong
• Benchmark Name: OS-World, Task Number: 369, Dataset Source: human annotation
Apr 2026Many-Tier Instruction Hierarchy in LLM Agents






• Affiliation: National University of Singapore
• Agent Name: AssistGUI, Base Model: GPT-4V, Strategy: Prompt
• Benchmark Name: AssistBench, Task Number: 100, Dataset Source: human demonstration
Apr 2026DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math?


• Affiliation: IBM Research
• Benchmark Name: DRBENCHER, Task Number: 354, Dataset Source: LLM synthesis
Apr 2026CORA: Conformal Risk-Controlled Agents for Safeguarded Mobile GUI Automation




• Affiliation: The University of Hong Kong
• Agent Name: CORA, Base Model: AutoGLM-Phone-9B-Multilingual, Strategy: Conformal Risk Control, Selective action execution, LoRA
• Benchmark Name: Phone-Harm, Task Number: 300, Dataset Source: human-authored and human-annotated
Apr 2026SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment



• Affiliation: Fudan University
• Benchmark Name: SEA-Eval, Task Number: 120, Dataset Source: human experts and LLM synthesis
Apr 2026PSI: Shared State as the Missing Layer for Coherent AI-Generated Instruments in Personal AI Agents



• Affiliation: University of Virginia
• Agent Name: Facai, Base Model: , Strategy: Prompt
Apr 2026ClawBench: Can AI Agents Complete Everyday Online Tasks?



• Affiliation: University of British Columbia
• Benchmark Name: CLAWBENCH, Task Number: 153, Dataset Source: human demonstration
Apr 2026MolmoWeb: Open Visual Web Agent and Open Data for the Open Web



• Affiliation: Allen Institute for AI
• Agent Name: MolmoWeb, Base Model: Molmo2 (Qwen3 + SigLIP2), Strategy: SFT
• Benchmark Name: MolmoWebMix, Task Number: 278500, Dataset Source: human demonstration and LLM synthesis
Apr 2026KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation




• Affiliation: Zhejiang University
• Benchmark Name: KnowU-Bench, Task Number: 192, Dataset Source: LLM-driven user simulator grounded in structured profiles
Apr 2026SkillClaw: Let Skills Evolve Collectively with Agentic Evolver




• Affiliation: DreamX Team
• Agent Name: SkillClaw, Base Model: Qwen3-Max, Strategy: Prompt
Apr 2026HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation





• Affiliation: University of Southern California
• Agent Name: SeeClick, Base Model: PaliGemma-3B-224, Strategy: SFT
Apr 2026LogAct: Enabling Agentic Reliability via Shared Logs


• Affiliation: Meta
• Agent Name: LogClaw, Base Model: FrontierModel, Strategy: Prompt
Apr 2026Same Outcomes, Different Journeys: A Trace-Level Framework for Comparing Human and GUI-Agent Behavior in Production Search Systems


• Affiliation: Spotify
• Benchmark Name: trace-level evaluation framework, Task Number: 10, Dataset Source: human demonstration
Apr 2026EigentSearch-Q+: Enhancing Deep Research Agents with Structured Reasoning Tools









• Affiliation: Shanghai Jiao Tong University
• Agent Name: OS-Atlas, Base Model: InternLM-XComposer2-7B, Strategy: SFT
• Benchmark Name: OSW-Bench, Task Number: 450, Dataset Source: human demonstration
Apr 2026An Agentic Evaluation Architecture for Historical Bias Detection in Educational Textbooks






• Affiliation: Tencent
• Agent Name: AppAgent-V2, Base Model: GPT-4o, Strategy: Multi-layer codebase
• Benchmark Name: App-Evaluation-V2, Task Number: 100, Dataset Source: human demonstration
Apr 2026Task-Adaptive Retrieval over Agentic Multi-Modal Web Histories via Learned Graph Memory


• Affiliation: Royal Melbourne Institute of Technology University
• Agent Name: ACGM, Base Model: CLIP, RoBERTa, Strategy: policy-gradient, supervised pre-training
Apr 2026Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection







• Affiliation: University of Chinese Academy of Sciences
• Agent Name: Editor, Base Model: Qwen3-VL-Plus, Strategy: Prompt
• Benchmark Name: Semantic-level UI Element Injection framework, Task Number: 885, Dataset Source: OS-Atlas, SeeClick, AMEX, and ShowUI
Apr 2026Structured Distillation of Web Agent Capabilities Enables Generalization



• Affiliation: Mila – Quebec AI Institute
• Agent Name: A3-Qwen3.5-9B, Base Model: Qwen3.5-9B, Strategy: SFT
Apr 2026RoboAgent: Chaining Basic Capabilities for Embodied Task Planning





• Affiliation: Nanyang Technological University
• Agent Name: LLaVA-OneVision, Base Model: Qwen2-7B, Qwen2-72B, Llama-3-8B, Strategy: SFT
Apr 2026Towards Knowledgeable Deep Research: Framework and Benchmark





• Affiliation: McGill University
• Agent Name: WebLINX agents, Base Model: Llama-2, Flan-T5, Pix2Struct, Strategy: Standard fine-tuning
• Benchmark Name: WebLINX, Task Number: 2337, Dataset Source: human demonstration
Apr 2026From Debate to Decision: Conformal Social Choice for Safe Multi-Agent Deliberation




• Affiliation: Tsinghua University
• Benchmark Name: VisualWebBench, Task Number: 756, Dataset Source: real-world websites
Apr 2026Behavior Latticing: Inferring User Motivations from Unstructured Interactions



• Affiliation: Stanford University
• Agent Name: Dawn, Base Model: Claude Sonnet 4.5, Claude Sonnet 4.6, Gemini 2.0 Flash Lite, Gemini 3 Pro, Strategy: ReAct, tool-calling MCP agent
Apr 2026MCP-DPT: A Defense-Placement Taxonomy and Coverage Analysis for Model Context Protocol Security

• Affiliation: Old Dominion University
• Paper Number: 69
Apr 2026CLEAR: Context Augmentation from Contrastive Learning of Experience via Agentic Reflection



• Affiliation: AWS AI Labs
• Agent Name: CAM, Base Model: Qwen/Qwen3-32B, Strategy: SFT, RL (GRPO), Agentic Reflection, Contrastive Learning
Apr 2026HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents





• Affiliation: Carnegie Mellon University
• Agent Name: Agent-S, Base Model: GPT-4o, Strategy: Experience-augmented hierarchical planning
Apr 2026GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents





• Affiliation: ServiceNow Research
• Agent Name: WebLlama, Base Model: Llama-2-7B, Strategy: SFT
• Benchmark Name: WebLINX, Task Number: 2337, Dataset Source: human demonstration
Apr 2026Android Coach: Improve Online Agentic Training Efficiency with Single State Multiple Actions



• Affiliation: Zhejiang University
• Agent Name: ANDROIDCOACH, Base Model: UI-TARS-1.5-7B, Strategy: Online RL with Single State Multiple Actions (SSMA)
Apr 2026TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories



• Affiliation: CyCraft AI Lab, Taiwan
• Benchmark Name: TRACESAFE-BENCH, Task Number: 1080, Dataset Source: Berkeley Function Calling Leaderboard (BFCL) and LLM-assisted mutation
Apr 2026Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization


• Affiliation: George Washington University
• Agent Name: T-STAR, Base Model: Qwen2.5-3B-Instruct, Phi-4-mini-instruct-3.8B, Strategy: Surgical Policy Optimization, In-Context Thought Grafting, Reinforcement Learning
Apr 2026SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems



• Affiliation: Peking University
• Agent Name: Agent-R, Base Model: Llama-3-8B-Instruct, Strategy: Reflection-Augmented Self-Play
Apr 2026PoC-Adapt: Semantic-Aware Automated Vulnerability Reproduction with LLM Multi-Agents and Reinforcement Learning-Driven Adaptive Policy


• Affiliation: University of Information Technology, Ho Chi Minh City
• Agent Name: PoC-Adapt, Base Model: Gemini-2.5-Pro, Strategy: Prompt, RL
Apr 2026SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills



• Affiliation: Imperial College London
• Benchmark Name: SkillSieve benchmark, Task Number: 49592, Dataset Source: ClawHub archive, Snyk ToxicSkills, ClawHavoc samples, human review
Apr 2026Neural Computers




• Affiliation: Meta AI
• Agent Name: NC CLIGen, Base Model: Wan2.1, Strategy: Action-conditioned video diffusion
• Agent Name: NC GUIWorld, Base Model: Wan2.1, Strategy: Action-conditioned video diffusion
• Benchmark Name: CLIGen, Task Number: 823,989, Dataset Source: asciinema .cast trajectories and vhs scripts
• Benchmark Name: GUIWorld, Task Number: 1500, Dataset Source: Desktop interaction traces and Claude CUA (Anthropic) trajectories
Apr 2026MTA-Agent: An Open Recipe for Multimodal Deep Search Agents




• Affiliation: Salesforce AI Research
• Agent Name: MTA-Agent, Base Model: Qwen3-VL-32B-Instruct, Strategy: RL (DAPO)
• Benchmark Name: MTA-Vision-DeepSearch-test, Task Number: 178, Dataset Source: LLM synthesis
Apr 2026WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks



• Affiliation: University of Wisconsin-Madison
• Agent Name: WebSP-Eval agentic framework, Base Model: Gemini-3-Pro, Gemini-2.5-Pro, Gemini-2.5-Flash, Claude-Sonnet-4.5, Claude-Haiku-4.5, GPT-5.1, GPT-5-mini, Gemma-3-27B, Strategy: Prompt
• Benchmark Name: WebSP-Eval, Task Number: 200, Dataset Source: manually crafted
Apr 2026RAGEN-2: Reasoning Collapse in Agentic RL



• Affiliation: Northwestern University
• Agent Name: RAGEN-2, Base Model: Qwen2.5, Strategy: Reinforcement Learning with SNR-Aware Filtering
Apr 2026The Art of Building Verifiers for Computer Use Agents




• Affiliation: Microsoft Research
• Agent Name: Universal Verifier, Base Model: gpt-5.2, Strategy: Prompt
• Benchmark Name: CUAVerifierBench, Task Number: 246, Dataset Source: human labels
Apr 2026VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics



• Affiliation: Ant Group
• Benchmark Name: VenusBench-Mobile, Task Number: 149 primary tasks and 80 variations, Dataset Source: manual construction and annotation
Apr 2026WebExpert: domain-aware web agents with critic-guided expert experience for high-precision search



• Affiliation: Shanghai Jiao Tong University
• Agent Name: WebExpert, Base Model: QwQ-32B, Strategy: SFT, preference optimization, experience-conditioned planning, retrieval-augmented generation
Apr 2026Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents



• Affiliation: Peking University
• Benchmark Name: Claw-Eval, Task Number: 300, Dataset Source: human-verified tasks
Apr 2026Gym-Anything: Turn any Software into an Agent Environment







• Affiliation: CMU
• Agent Name: Ours (2B distilled), Base Model: Qwen3-VL-2B-Thinking, Strategy: SFT
• Agent Name: Test-Time Auditing (TTA) Agent, Base Model: Gemini 3 Flash, Strategy: Prompt
• Benchmark Name: CUA-World, Task Number: 12,103, Dataset Source: LLM synthesis
Mar 2026The Kitchen Loop: User-Spec-Driven Development for a Self-Evolving Codebase





• Affiliation: 0xAgentKitchen
• Agent Name: The Kitchen Loop, Base Model: Claude, Strategy: Prompt
Mar 2026Is Mathematical Problem-Solving Expertise in Large Language Models Associated with Assessment Performance?




• Affiliation: SenseTime Research
• Agent Name: SDG-Agent, Base Model: Qwen-VL-7B, Strategy: SFT
Mar 2026EcoThink: A Green Adaptive Inference Framework for Sustainable and Accessible Agents





• Affiliation: Alibaba Group
• Agent Name: SeeClick, Base Model: Qwen-VL, Strategy: SFT
• Benchmark Name: Screen-Grounding, Task Number: 127,113, Dataset Source: human demonstration
Mar 2026WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing




• Affiliation: Northeastern University
• Agent Name: WebTester, Base Model: Claude 4.5, GPT-5, GLM-5, Step-3.5, Qwen3, MiMo-V2, Minimax-M2.1, Strategy: Prompt
• Benchmark Name: WebTestBench, Task Number: 100, Dataset Source: LLM synthesis
Mar 2026TopoPilot: Reliable Conversational Workflow Automation for Topological Data Analysis and Visualization


• Affiliation: University of Utah
• Agent Name: TopoPilot, Base Model: ChatGPT 4o, Strategy: Prompt
Mar 2026Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale






• Affiliation: Shanghai AI Laboratory
• Agent Name: Intern-S1-Pro, Base Model: Intern-S1, Strategy: Reinforcement Learning (RL)
Mar 2026Rethinking Failure Attribution in Multi-Agent Systems: A Multi-Perspective Benchmark and Evaluation



• Affiliation: KAIST
• Benchmark Name: MP-Bench, Task Number: 289, Dataset Source: human expert annotation
Mar 2026PII Shield: A Browser-Level Overlay for User-Controlled Personal Identifiable Information (PII) Management in AI Interactions









• Affiliation: Shanghai AI Laboratory
• Agent Name: OS-Atlas, Base Model: Qwen2-VL-7B, Strategy: SFT
• Benchmark Name: OS-Ground, Task Number: 5022, Dataset Source: existing GUI grounding and agent datasets
Mar 2026A Dual-Threshold Probabilistic Knowing Value Logic


• Affiliation: Rutgers University
• Agent Name: UI-Point, Base Model: LayoutLMv3-base, Strategy: VARL (View-Aware Representation Learning)
Mar 2026Chameleon: Episodic Memory for Long-Horizon Robotic Manipulation




• Affiliation: Beijing Institute of Technology
• Agent Name: Mobile-Agent-v2, Base Model: GPT-4o, Strategy: Multi-agent architecture (Planning Agent, Decision Agent, and Self-Reflection Agent)
• Benchmark Name: Mobile-Eval, Task Number: 128, Dataset Source:
Mar 2026AVO: Agentic Variation Operators for Autonomous Evolutionary Search



• Affiliation: ByteDance
• Agent Name: Coflow, Base Model: GPT-4o, Strategy: Prompt
Mar 2026Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs





• Affiliation: McGill University
• Agent Name: WebLINX-Llama, Base Model: Llama-2, Strategy: SFT
• Benchmark Name: WebLINX, Task Number: 2337, Dataset Source: human demonstration
Mar 2026Mechanic: Sorrifier-Driven Formal Decomposition Workflow for Automated Theorem Proving



• Affiliation: Nanyang Technological University
• Agent Name: LASER, Base Model: GPT-4o, GPT-4 (Vision), Gemini-1.5-Pro, Claude-3.5-Sonnet, Strategy: Self-Correction (Error Detection and Error Recovery)
Mar 2026OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning



• Affiliation: Tsinghua University
• Agent Name: WebRL, Base Model: GLM-4-9B, Llama-3-8B, Strategy: Reinforcement Learning, SFT, Self-evolving curriculum
Mar 2026ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers







• Affiliation: Beijing University of Posts and Telecommunications
• Agent Name: ClawKeeper, Base Model: GLM-5, Strategy: Prompt
• Benchmark Name: seven categories of safety tasks, Task Number: 140, Dataset Source:
Mar 2026AI-Supervisor: Autonomous AI Research Supervision via a Persistent Research World Model



• Affiliation: The Hong Kong University of Science and Technology
• Agent Name: FPA, Base Model: Mistral-7B-v0.1, Strategy: SFT, Fine-Grained Preference Alignment
Mar 2026GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents




• Affiliation: Microsoft
• Agent Name: UFO, Base Model: GPT-4V, Strategy: Prompt
• Benchmark Name: Windows-Bench, Task Number: 50, Dataset Source: human demonstration
Mar 2026Large Language Model Guided Incentive Aware Reward Design for Cooperative Multi-Agent Reinforcement Learning





• Affiliation: Beijing University of Posts and Telecommunications
• Agent Name: ScreenAgent, Base Model: GPT-4-Vision-Preview, Strategy: Prompt
• Benchmark Name: ScreenAgent dataset, Task Number: 303, Dataset Source: human demonstration
Mar 2026C-STEP: Continuous Space-Time Empowerment for Physics-informed Safe Reinforcement Learning of Mobile Agents




• Affiliation: Microsoft Research
• Agent Name: UFO, Base Model: GPT-4V, Strategy: Prompt
• Benchmark Name: WindowsBench, Task Number: 50, Dataset Source: human demonstration
Mar 2026Where Do Your Citations Come From? Citation-Constellation: A Free, Open-Source, No-Code, and Auditable Tool for Citation Network Decomposition with Complementary BARON and HEROCON Scores




• Affiliation: McGill University
• Benchmark Name: WebLINX, Task Number: 2337, Dataset Source: human demonstration
Mar 2026On Gossip Algorithms for Machine Learning with Pairwise Objectives







• Affiliation: The University of Hong Kong
• Agent Name: Friday, Base Model: GPT-4, Strategy: Self-Improvement
Mar 2026From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments


• Affiliation: Sun Yat-sen University
• Paper Number: 170
Mar 2026From AI Assistant to AI Scientist: Autonomous Discovery of LLM-RL Algorithms with LLM Agents




• Affiliation: Tsinghua University
• Agent Name: V-PPO, Base Model: LLaVA-1.5-7B, Qwen-VL, Strategy: Reinforcement Learning
Mar 2026AnalogAgent: Self-Improving Analog Circuit Design Automation with LLM Agents





• Affiliation: Google Research
• Agent Name: AndroidWorld Agent, Base Model: Gemini Pro Vision, Strategy: Prompt
• Benchmark Name: AndroidWorld, Task Number: 116, Dataset Source: manually designed
Mar 2026BeliefShift: Benchmarking Temporal Belief Consistency and Opinion Drift in LLM Agents







• Affiliation: Tsinghua University
• Agent Name: SeeClick, Base Model: InternViT-6B-Vicuna-7B (InternVL), Strategy: point-based GUI grounding pre-training
• Benchmark Name: SeeClick-Bench, Task Number: 124, Dataset Source: human annotation
Mar 2026VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents




• Affiliation: The University of Hong Kong
• Agent Name: RLEF, Base Model: Qwen-VL-7B, CogVLM-17B, Strategy: Reinforcement Learning (PPO)
Mar 2026AgentRFC: Security Design Principles and Conformance Testing for Agent Protocols





• Affiliation: University of California, Berkeley
• Agent Name: DigiRL, Base Model: Llama-3-8B-Instruct, Gemma-2-9B-IT, Strategy: Offline-to-Online Reinforcement Learning
Mar 2026BXRL: Behavior-Explainable Reinforcement Learning


• Affiliation: Peking University
• Agent Name: Auto-GUI, Base Model: Llama-2-7b-chat-hf, Strategy: SFT, Chain-of-Thought, Interactive feedback
Mar 2026Environment Maps: Structured Environmental Representations for Long-Horizon Agents


• Affiliation: Distyl AI
• Agent Name: Environment Maps, Base Model: claude-sonnet-4-5, Strategy: Prompt
Mar 2026AscendOptimizer: Episodic Agent for Ascend NPU Operator Optimization



• Affiliation: National University of Singapore
• Agent Name: Agent-R, Base Model: Llama3-8B-Instruct, Strategy: Iterative Self-Reflection and Rewarding
Mar 2026CAPTCHA Solving for Native GUI Agents: Automated Reasoning-Action Data Generation and Self-Corrective Training





• Affiliation: McGill University
• Agent Name: WebLINX-Llama, Base Model: Llama-2-7b, Strategy: SFT
• Agent Name: WebLINX-Mistral, Base Model: Mistral-7B, Strategy: SFT
• Benchmark Name: WebLINX, Task Number: 2337, Dataset Source: human demonstration
Mar 2026Plato's Cave: A Human-Centered Research Verification System



• Affiliation: University of Florida
• Agent Name: Plato’s Cave, Base Model: gpt-5-mini, Strategy: Prompt
Mar 2026MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens





• Affiliation: Microsoft Research
• Agent Name: Navi, Base Model: GPT-4o, Strategy: Prompt
• Benchmark Name: Windows Agent Arena, Task Number: 154, Dataset Source: human demonstration
Mar 2026AgentRAE: Remote Action Execution through Notification-based Visual Backdoors against Screenshots-based Mobile GUI Agents


• Affiliation: Nanjing University of Science and Technology
• Agent Name: AgentRAE, Base Model: Qwen-VL-Chat, Strategy: Supervised Contrastive Learning and Supervised Poisoning Training
Mar 2026PaperVoyager : Building Interactive Web with Visual Language Models



• Affiliation: Vast Intelligence Lab
• Agent Name: PaperVoyager, Base Model: Qwen3-VL-4B-Instruct, Strategy: Prompt, structured generation specification, block-level decomposition, and VLM-based candidate filtering.
• Benchmark Name: PaperVoyager benchmark, Task Number: 19, Dataset Source: expert-authored interactive web systems
Mar 2026SoK: The Attack Surface of Agentic AI -- Tools, and Autonomy

• Affiliation: University of Guelph
• Paper Number: 28
Mar 2026IntentWeave: A Progressive Entry Ladder for Multi-Surface Browser Agents in Cloud Portals


• Affiliation: Alibaba Cloud Computing, Alibaba Group
• Agent Name: IntentWeave, Base Model: , Strategy: progressive entry ladder
Mar 2026The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration

• Affiliation: Harbin Institute of Technology
• Paper Number: 237
Mar 2026Beyond Binary Correctness: Scaling Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks


• Affiliation: metaphi.ai
• Benchmark Name: LH-Bench, Task Number: 216, Dataset Source: Figma Community, enterprise design partners, and expert-curated sources
Mar 2026Synthetic or Authentic? Building Mental Patient Simulators from Longitudinal Evidence





• Affiliation: Peking University
• Agent Name: RL-VLM-F, Base Model: Qwen-VL-Chat, CogVLM-17B, Strategy: Reinforcement Learning (PPO)
Mar 2026Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos




• Affiliation: Google Research
• Agent Name: AndroidWorld Agent, Base Model: GPT-4V, Gemini Pro, Strategy: Prompt
• Benchmark Name: AndroidWorld, Task Number: 116, Dataset Source: human-curated
Mar 2026OrgForge-IT: A Verifiable Synthetic Benchmark for LLM-Based Insider Threat Detection




• Affiliation: Lehigh University
• Agent Name: Agent-S, Base Model: GPT-4o, Strategy: Experience Search, Knowledge Retrieval, Task Planner, and Sub-task Executor
Mar 2026From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents


• Affiliation: Rensselaer Polytechnic Institute

…(truncated)

Collected info

  • 86 stars
  • 2 forks
  • Source updated: 7/24/2026