GVA-Survey
Official repository of the paper "Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms"
Links
README
From the repo.
Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms
Minghe Gao1,
Wendong Bu1,
Bingchen Miao1,
Yang Wu2,
Yunfei Li2,
Juncheng Li1†,
Siliang Tang1,
Qi Wu3,
Yueting Zhuang1,
Meng Wang4
1Zhejiang University, Hangzhou, China
2Antgroup, China
3The University of Adelaide, Adelaide, Australia
4Hefei University of Technology, Hefei, China
🔥 News
-
[December 10, 2024] We have developed an agent that automatically collects and analyzes the latest papers in the GVA field. It will update the Related Papers daily at 0:30 AM UTC+8.
-
[December 7, 2024] We have released a Chinese version of the survey, please click 中文版综述 to access!
-
[November 17, 2024] Our survey is available on the arXiv platform: https://arxiv.org/abs/2411.10943
📖 Table of Content
🤖 Introduction
Welcome to the GitHub repository for our survey paper titled "Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms". This repository includes all the resources, code, and references related to the paper. Our objective is to provide a comprehensive overview of Generalist Virtual Agents (GVAs), covering their definition, necessity, implementation approaches, evaluation methods, limitations and future directions. We aim to bridge the gap between theory and practice in GVA research, providing a systematic framework for future development in this field.
📚 Cited Papers
Here we list the most important references cited in our survey, organized by different sections. We particularly focus on works that have made substantial impact or proposed innovative methodologies.
Additionally, we note that some papers may be cited across multiple sections. For the convenience of researchers, we provide complete citation information under each section.
Section II: What is GVA?

Environment - Web
-
World of bits: an open-domain platform for web-based agents
-
Reinforcement Learning on Web Interfaces using Workflow-Guided Exploration
-
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Department of Computer Science, Princeton University, NeurIPS, 2022
-
WebArena: A Realistic Web Environment for Building Autonomous Agents
-
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Carnegie Mellon University, arXiv preprint arXiv:2401.13649, 2024
-
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
-
Mind2Web: Towards a Generalist Agent for the Web
-
WebVLN: Vision-and-Language Navigation on Websites
Australian Institute for Machine Learning, The University of Adelaide, AAAI, 2024
Environment - Application
-
AppAgent: Multimodal Agents as Smartphone Users
-
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024
-
A Dataset for Interactive Vision Language Navigation with Unknown Command Feasibility
-
AndroidInTheWild: A Large-Scale Dataset For Android Device Control
-
Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning
-
Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding
-
Visual Programming: Compositional visual reasoning without training
-
Toolformer: Language Models Can Teach Themselves to Use Tools
-
AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production
DAMO Academy, Alibaba Group, arXiv preprint arXiv:2403.07952, 2024
Environment - Operating System
-
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
The University of Hong Kong, arXiv preprint arXiv:2404.07972, 2024
-
AgentStudio: A Toolkit for Building General Virtual Agents
-
MMAC-Copilot: Multi-modal Agent Collaboration Operating System Copilot
University of Technology Sydney, arXiv preprint arXiv:2404.18074, 2024
-
UFO: A UI-Focused Agent for Windows OS Interaction
Task - Command Task
-
AppAgent: Multimodal Agents as Smartphone Users
-
Android In The Wild: A Large-Scale Dataset For Android Device Control
-
UFO: A UI-Focused Agent for Windows OS Interaction
-
From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces
-
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Department of Computer Science, Princeton University, NeurIPS, 2022
Task - Query Task
-
ViperGPT: Visual Inference via Python Execution for Reasoning
-
Visual Programming: Compositional visual reasoning without training
-
Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning
-
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
-
Mm-react: Prompting chatgpt for multimodal reasoning and action
-
WebVLN: Vision-and-Language Navigation on Websites
Australian Institute for Machine Learning, The University of Adelaide, AAAI, 2024
Task - Dialogue Task
-
Windows Copilot Plus for PCs
-
Apple Intelligence Overview
-
NICE: Neural Image Commenting with Empathy
-
Is ChatGPT Equipped with Emotional Dialogue Capabilities?
Harbin Institute of Technology, China, arXiv preprint arXiv:2304.09582, 2023
Observation Space - Command Line Interface
-
AutoGPT
-
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
The University of Hong Kong, arXiv preprint arXiv:2404.07972, 2024
-
AgentBench: Evaluating LLMs as Agents
-
AgentStudio: A Toolkit for Building General Virtual Agents
-
InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback
Department of Computer Science, Princeton University, NeurIPS, 2023
Observation Space - Document Object Model
-
World of bits: an open-domain platform for web-based agents
-
Reinforcement Learning on Web Interfaces using Workflow-Guided Exploration
-
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Department of Computer Science, Princeton University, NeurIPS, 2022
-
Mind2Web: Towards a Generalist Agent for the Web
-
AppAgent: Multimodal Agents as Smartphone Users
-
A data-driven approach for learning to control computers
Observation Space - Screen
-
You Only Look at Screens: Multimodal Chain-of-Action Agents
School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University, arXiv preprint arXiv:2309.11436, 2024
-
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Department of Computer Science and Technology, Nanjing University, arXiv preprint arXiv:2401.10935, 2024
-
WebVLN: Vision-and-Language Navigation on Websites
Australian Institute for Machine Learning, The University of Adelaide, AAAI, 2024
-
From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces
-
GPT-4V(ision) is a Generalist Web Agent, if Grounded
The Ohio State University, arXiv preprint arXiv:2401.01614, 2024
-
UFO: A UI-Focused Agent for Windows OS Interaction
-
GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation
-
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024
Action Space - Keyboard
-
World of bits: an open-domain platform for web-based agents
-
Mind2Web: Towards a Generalist Agent for the Web
-
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
-
WebArena: A Realistic Web Environment for Building Autonomous Agents
-
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
Action Space - Mouse
-
World of bits: an open-domain platform for web-based agents
-
From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces
-
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
-
WebArena: A Realistic Web Environment for Building Autonomous Agents
-
Mind2Web: Towards a Generalist Agent for the Web
-
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
Action Space - Touchscreen
-
AppAgent: Multimodal Agents as Smartphone Users
-
AndroidInTheWild: A Large-Scale Dataset For Android Device Control
-
You Only Look at Screens: Multimodal Chain-of-Action Agents
School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University, arXiv preprint arXiv:2309.11436, 2024
-
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024
Action Space - Others
-
HyperPalm: DNN-based hand gesture recognition interface for intelligent communication with quadruped robot in 3D space
Section III: Why we need GVA?

Perspective of AI and Machine Learning
-
CogAgent: A Visual Language Model for GUI Agents
-
You Only Look at Screens: Multimodal Chain-of-Action Agents
School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University, arXiv preprint arXiv:2309.11436, 2024
-
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Department of Computer Science and Technology, Nanjing University, arXiv preprint arXiv:2401.10935, 2024
-
From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces
-
GPT-4V(ision) is a Generalist Web Agent, if Grounded
The Ohio State University, arXiv preprint arXiv:2401.01614, 2024
-
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024
Perspective of Interaction
-
Inferring Rewards from Language in Context
-
Deep Reinforcement Learning from Human Preferences
-
Apple Intelligence Overview
Perspective of Agent Applications
-
Distributed and reactive query planning in R-MAGIC: an agent-based multimedia retrieval system
-
Rap: Retrieval-augmented planning with contextual memory for multimodal llm agents
Panasonic Connect Co., Ltd., Japan, arXiv preprint arXiv:2402.03610, 2024
-
Plan4mc: Skill reinforcement learning and planning for open-world minecraft tasks
School of Computer Science, Peking University, arXiv preprint arXiv:2303.16563, 2023
-
Autoact: Automatic agent learning from scratch via self-planning
-
Self-Refine: Iterative Refinement with Self-Feedback
Language Technologies Institute, Carnegie Mellon University, NeurIPS, 2023
-
Reflexion: language agents with verbal reinforcement learning
-
Eureka: Human-Level Reward Design via Coding Large Language Models
-
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
-
AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production
DAMO Academy, Alibaba Group, arXiv preprint arXiv:2403.07952, 2024
Section IV: How to implement GVA?

Environment - Off-line
-
World of bits: an open-domain platform for web-based agents
-
Reinforcement Learning on Web Interfaces using Workflow-Guided Exploration
-
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Department of Computer Science and Technology, Nanjing University, arXiv preprint arXiv:2401.10935, 2024
-
AndroidInTheWild: A Large-Scale Dataset For Android Device Control
-
Mind2Web: Towards a Generalist Agent for the Web
-
A Dataset for Interactive Vision Language Navigation with Unknown Command Feasibility
-
META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
Environment - On-line
-
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Department of Computer Science, Princeton University, NeurIPS, 2022
-
WebArena: A Realistic Web Environment for Building Autonomous Agents
-
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Carnegie Mellon University, arXiv preprint arXiv:2401.13649, 2024
Environment - Updated
-
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
-
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
-
AppAgent: Multimodal Agents as Smartphone Users
-
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024
-
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
The University of Hong Kong, arXiv preprint arXiv:2404.07972, 2024
-
AgentStudio: A Toolkit for Building General Virtual Agents
-
UFO: A UI-Focused Agent for Windows OS Interaction
Model - Retriever-based Agent
-
Distributed and reactive query planning in R-MAGIC: an agent-based multimedia retrieval system
nan, IEEE Transactions on Knowledge and Data Engineering, 2004
-
Mobile agents and their use for information retrieval: a brief overview and an elaborate case study
-
SAIRE—a scalable agent-based information retrieval engine
nan, Proceedings of the first international conference on Autonomous agents, 1997
-
ACQUIRE: agent-based complex query and information retrieval engine
nan, Proceedings of the first international joint conference on Autonomous agents and multiagent systems: part 2, 2002
-
Thought-retriever: Don’t just retrieve raw data, retrieve thoughts
University of Illinois at Urbana-Champaign, ICLR Workshop, 2024
Model - LLM-based Agent
-
Agent-pro: Learning to evolve via policy-level reflection and optimization
-
Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning
-
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
-
Mm-react: Prompting chatgpt for multimodal reasoning and action
-
Mind2Web: Towards a Generalist Agent for the Web
-
AllTogether: Investigating the Efficacy of Spliced Prompt for Web Navigation using Large Language Models
-
Building Cooperative Embodied Agents Modularly with Large Language Models
-
RoboGPT: an intelligent agent of making embodied long-term decisions for daily instruction tasks
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China, arXiv preprint arXiv:2311.15649, 2024
-
AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation
University of Hong Kong, arXiv preprint arXiv:2312.13010, 2024
-
Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering
Tokyo Institute of Technology, arXiv preprint arXiv:2407.20563, 2024
-
Modular Visual Question Answering via Code Generation
-
ViperGPT: Visual Inference via Python Execution for Reasoning
Model - MLLM-based Agent
-
GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation
-
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024
-
Autonomous Evaluation and Refinement of Digital Agents
-
Clova: A closed-loop visual assistant with tool usage and update
-
Agent smith: A single image can jailbreak one million multimodal llm agents exponentially fast
-
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
-
You Only Look at Screens: Multimodal Chain-of-Action Agents
School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University, arXiv preprint arXiv:2309.11436, 2024
-
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Carnegie Mellon University, arXiv preprint arXiv:2401.13649, 2024
-
CogAgent: A Visual Language Model for GUI Agents
Model - VLA-based Agent
-
Rt-2: Vision-language-action models transfer web knowledge to robotic control
-
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
-
Octo: An open-source generalist robot policy
-
OpenVLA: An Open-Source Vision-Language-Action Model
-
3d-vla: A 3d vision-language-action generative world model
University of Massachusetts Amherst, arXiv preprint arXiv:2403.09631, 2024
-
Palm-e: An embodied multimodal language model
-
Actra: Optimized Transformer Architecture for Vision-Language-Action Models in Robot Learning
The Chinese University of Hong Kong, arXiv preprint arXiv:2408.01147, 2024
Strategy - Adaption Strategy
-
Mm-react: Prompting chatgpt for multimodal reasoning and action
-
Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models
Singapore Management University, arXiv preprint arXiv:2305.04091, 2023
-
Expertprompting: Instructing large language models to be distinguished experts
University of Science and Technology of China, arXiv preprint arXiv:2305.14688, 2023
-
Self-refine: Iterative refinement with self-feedback
Language Technologies Institute, Carnegie Mellon University, NeurIPS, 2024
-
Pangu-coder2: Boosting large language models for code with ranking feedback
Huawei Cloud Co., Ltd., arXiv preprint arXiv:2307.14936, 2023
-
De-fine: Decomposing and refining visual programs with auto-feedback
-
A survey on the memory mechanism of large language model based agents
Gaoling School of Artificial Intelligence, Renmin University of China, arXiv preprint arXiv:2404.13501, 2024
-
Rap: Retrieval-augmented planning with contextual memory for multimodal llm agents
Panasonic Connect Co., Ltd., Japan, arXiv preprint arXiv:2402.03610, 2024
-
Chatdb: Augmenting llms with databases as their symbolic memory
-
Memorybank: Enhancing large language models with long-term memory
Strategy - Fine-tuning Strategy
-
CoEvol: Constructing Better Responses for Instruction Finetuning through Multi-Agent Cooperation
-
Reflection-Reinforced Self-Training for Language Agents
University of California, Los Angeles, arXiv preprint arXiv:2406.01495, 2024
-
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training
-
Enhancing the General Agent Capabilities of Low-Parameter LLMs through Tuning and Multi-Branch Reasoning
Huazhong University of Science and Technology, arXiv preprint arXiv:2403.19962, 2024
-
Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models
Department of Automation, University of Science and Technology of China, arXiv preprint arXiv:2403.12881, 2024
Strategy - Reinforcement Learning Strategy
-
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
-
Search Beyond Queries: Training Smaller Language Models for Web Interactions via Reinforcement Learning
University of Kentucky, arXiv preprint arXiv:2404.10887, 2024
-
Juewu-mc: Playing minecraft with sample-efficient hierarchical reinforcement learning
Tencent AI Lab, Shenzhen, China, arXiv preprint arXiv:2112.04907, 2021
-
Towards robust and domain agnostic reinforcement learning competitions: MineRL 2020
-
Plan4mc: Skill reinforcement learning and planning for open-world minecraft tasks
School of Computer Science, Peking University, arXiv preprint arXiv:2303.16563, 2023
-
Pok'eLLMon: A Human-Parity Agent for Pok'emon Battles with Large Language Models
Georgia Institute of Technology, arXiv preprint arXiv:2402.01118, 2024
-
Scaling instructable agents across many simulated worlds
-
Mental Modeling of Reinforcement Learning Agents by Language Models
University of Hamburg, arXiv preprint arXiv:2406.18505, 2024
-
LLMSat: A Large Language Model-Based Goal-Oriented Agent for Autonomous Space Exploration
University of Toronto, arXiv preprint arXiv:2405.01392, 2024
Strategy - Cooperation or Competition Strategy
-
CMAT: A Multi-Agent Collaboration Tuning Framework for Enhancing Small Language Models
East China Jiaotong University, arXiv preprint arXiv:2404.01663, 2024
-
Autoact: Automatic agent learning from scratch via self-planning
-
ProAgent: building proactive cooperative agents with large language models
-
Chatllm network: More brains, more intelligence
Beijing University of Posts and Telecommunications, arXiv preprint arXiv:2304.12998, 2023
-
Encouraging divergent thinking in large language models through multi-agent debate
-
Improving factuality and reasoning in language models through multiagent debate
-
Chateval: Towards better llm-based evaluators through multi-agent debate
-
How susceptible are llms to logical fallacies?
Department of Computer Science, George Mason University, arXiv preprint arXiv:2308.09853, 2023
Section V: How to evaluate GVA?

Overall Evaluation
-
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Department of Computer Science, Princeton University, NeurIPS, 2022
-
Mind2Web: Towards a Generalist Agent for the Web
-
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
-
GUI Odyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices
OpenGVLab, Shanghai AI Laboratory, arXiv preprint arXiv:2406.08451, 2024
-
WebSRC: A Dataset for Web-Based Structural Reading Comprehension
-
WebArena: A Realistic Web Environment for Building Autonomous Agents
-
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Carnegie Mellon University, arXiv preprint arXiv:2401.13649, 2024
-
Mapping Natural Language Instructions to Mobile UI Action Sequences
-
Grounding Open-Domain Instructions to Automate Web Support Tasks
Detail Evaluation - Step-wise Evaluation
-
World of bits: an open-domain platform for web-based agents
-
Mind2Web: Towards a Generalist Agent for the Web
-
Mapping Natural Language Instructions to Mobile UI Action Sequences
-
META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
-
AndroidInTheWild: A Large-Scale Dataset For Android Device Control
-
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024
-
WebSRC: A Dataset for Web-Based Structural Reading Comprehension
-
WebVLN: Vision-and-Language Navigation on Websites
Australian Institute for Machine Learning, The University of Adelaide, AAAI, 2024
-
Room2Room: Enabling Life-Size Telepresence in a Projected Augmented Reality Environment
-
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
Carnegie Mellon University, arXiv preprint arXiv:2402.17553, 2024
Detail Evaluation - Set Inclusion Evaluation
-
Grounding Open-Domain Instructions to Automate Web Support Tasks
-
META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
-
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Carnegie Mellon University, arXiv preprint arXiv:2401.13649, 2024
-
WebQA: Multihop and Multimodal QA
-
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Department of Computer Science, Princeton University, NeurIPS, 2022
-
Sequence to sequence learning with neural networks
-
A Dataset for Interactive Vision Language Navigation with Unknown Command Feasibility
Detail Evaluation - Multi-dimensional Evaluation
-
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
The University of Hong Kong, arXiv preprint arXiv:2404.07972, 2024
-
WebVLN: Vision-and-Language Navigation on Websites
Australian Institute for Machine Learning, The University of Adelaide, AAAI, 2024
-
ChatDev: Communicative Agents for Software Development
-
Online Adaptation of Language Models with a Memory of Amortized Contexts
-
Voyager: An Open-Ended Embodied Agent with Large Language Models
-
GUI Odyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices
OpenGVLab, Shanghai AI Laboratory, arXiv preprint arXiv:2406.08451, 2024
-
Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies
-
MemoChat: Tuning LLMs to Use Memos for Consistent Long-Range Open-Domain Conversation
University of Warwick, arXiv preprint arXiv:2308.08239, 2023
-
Secrets of RLHF in Large Language Models Part I: PPO
-
WebSRC: A Dataset for Web-Based Structural Reading Comprehension
-
A Dataset for Interactive Vision Language Navigation with Unknown Command Feasibility
Human Evaluation
-
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
-
AndroidInTheWild: A Large-Scale Dataset For Android Device Control
-
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
Beijing Jiaotong University, arXiv preprint arXiv:2401.16158, 2024
MLLM-based Evaluation
-
Gpt-4 technical report
-
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Carnegie Mellon University, arXiv preprint arXiv:2401.13649, 2024
-
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
-
GUI-WORLD: A Dataset for GUI-oriented Multimodal LLM-based Agents
Huazhong University of Science and Technology, arXiv preprint arXiv:2406.10819, 2024
Section VI: Limitations

Unrealistic Environment and Dataset
-
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
The University of Hong Kong, arXiv preprint arXiv:2404.07972, 2024
-
Mapping Natural Language Instructions to Mobile UI Action Sequences
-
AppAgent: Multimodal Agents as Smartphone Users
Insufficient Transferability
-
GPT-4V(ision) is a Generalist Web Agent, if Grounded
The Ohio State University, arXiv preprint arXiv:2401.01614, 2024
-
Visually Grounded Language Learning: a review of language games, datasets, tasks, and models
Heriot-Watt University, Journal of Artificial Intelligence Research, 2024
Limited Long-Sequence Decision-Making
-
Memorybank: Enhancing large language models with long-term memory
-
Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning
-
Reasoning with language model is planning with world model
Heightened Security Concerns
-
LLaVA-phi: Efficient Multi-Modal Assistant with Small Language Model
-
Smoothquant: Accurate and efficient post-training quantization for large language models
-
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
Section VII: Future

From Individual to Systematic
-
LLM multi-agent systems: Challenges and open problems
University of California, Irvine, arXiv preprint arXiv:2402.03578, 2024
-
Combining multi-agent systems and Artificial Intelligence of Things: Technical challenges and gains
Belfort Montbeliard University of Technology, Internet of Things, 2024
-
Multi-agent systems in Peer-to-Peer energy trading: A comprehensive survey
University of Galway, Engineering Applications of Artificial Intelligence, 2024
From Virtual to Physical
-
A survey on robotics with foundation models: toward embodied ai
-
Embodied AI with Two Arms: Zero-shot Learning, Safety and Modularity
Google Deep Mind Robotics, arXiv preprint arXiv:2404.03570, 2024
🔍 Related Papers
We are committed to offering researchers the latest advancements in the field. By regularly reviewing and evaluating recent research studies, we ensure that the list of papers stays up-to-date.
⚠️ The paper analysis may not be accurate and is for reference only!
Collected info
- ★ 86 stars
- ⎇ 2 forks
- Source updated: 7/24/2026