← Discover MCPs and Agents
A
AgentAI & MLGitHub

Awesome-LLM-Papers-Comprehensive-Topics

Awesome LLM Papers and repos on very comprehensive topics.

Links

README

From the repo.

Awesome-LLM-Related-Papers-Comprehensive-Topics

Static Badge Static Badge GitHub Repo stars

We provide awesome papers and repos on very comprehensive topics as follows.

CoT / VLM / Quantization / Grounding / Text2IMG&VID / Prompt Engineering / Prompt Tuning / Reasoning / Robot / Agent / Planning / Reinforcement-Learning / Feedback / In-Context-Learning / Few-Shot / Zero-Shot / Instruction Tuning / PEFT / RLHF / RAG / Embodied / VQA / Hallucination / Diffusion / Scaling / Context-Window / WorldModel / Memory / Zero-Shot / RoPE / Speech / Perception / Survey / Segmentation / Learge Action Model / Foundation / RoPE / LoRA / PPO / DPO

We strongly recommend checking our Notion table for an interactive experience.

drawing

Number of papers and repos in total: 516

CategoryTitleLinksDate
Zero-shotCan Foundation Models Perform Zero-Shot Task Speci
fication For Robot Manipulation?
World-modelLeveraging Pre-trained Large Language Models to Co
nstruct and Utilize World Models for Model-based Task Planning
ArXiv2023/05/07
World-modelLearning and Leveraging World Models in Visual Rep
resentation Learning
World-modelLanguage Models Meet World ModelsArXiv
World-modelLearning to Model the World with LanguageArXiv
World-modelDiffusion World ModelArXiv
World-modelLearning to Model the World with LanguageArXiv
VisualPromptFerret-v2: An Improved Baseline for Referring and
Grounding with Large Language Models
ArXiv
VisualPromptMaking Large Multimodal Models Understand Arbitrar
y Visual Prompts
ArXiv
VisualPromptWhat does CLIP know about a red circle? Visual pro
mpt engineering for VLMs
ArXiv
VisualPromptMOKA: Open-Vocabulary Robotic Manipulation through
Mark-Based Visual Prompting
ArXiv
VisualPromptSoM : Set-of-Mark PromptingUnleashes Extraordinary
Visual Grounding in GPT-4V
ArXiv
VisualPromptSet-of-Mark Prompting Unleashes Extraordinary Visu
al Grounding in GPT-4V
ArXiv, GitHub
VideoMA-LMM: Memory-Augmented Large Multimodal Model fo
r Long-Term Video Understanding
ArXiv
ViFM, VideoInternVideo2: Scaling Video Foundation Models for
Multimodal Video Understanding
ArXiv, GitHub
VLM, World-modelLarge World ModelArXiv
VLM, VQACogVLM: Visual Expert for Pretrained Language Mode
ls
ArXiv2023/11/06
VLM, VQAChameleon: Plug-and-Play Compositional Reasoning w
ith Large Language Models
ArXiv2023/04/19
VLM, VQADeepSeek-VL: Towards Real-World Vision-Language Un
derstanding01
VLMPaLM: Scaling Language Modeling with PathwaysArXiv2022/04/05
VLMScreenAI: A Vision-Language Model for UI and Infog
raphics Understanding
VLMMoE-LLaVA: Mixture of Experts for Large Vision-Lan
guage Models
ArXiv, GitHub
VLMLLaVA-NeXT: Improved reasoning, OCR, and world kno
wledge
GitHub
VLMMini-Gemini: Mining the Potential of Multi-modalit
y Vision Language Models
ArXiv
Text-to-Image, World-modelWorld Model on Million-Length Video And Language W
ith RingAttention
ArXiv
Tex2ImgBe Yourself: Bounded Attention for Multi-Subject T
ext-to-Image Generation
ArXiv
TemporalExplorative Inbetweening of Time and SpaceArXiv
Survey, VideoVideo Understanding with Large Language Models: A
Survey
ArXiv
Survey, VLMMM-LLMs: Recent Advances in MultiModal Large Langu
age Models
Survey, TrainingUnderstanding LLMs: A Comprehensive Overview from
Training to Inference
ArXiv
Survey, TimeSeriesLarge Models for Time Series and Spatio-Temporal D
ata: A Survey and Outlook
SurveyEfficient Large Language Models: A SurveyArXiv, GitHub
Sora, Text-to-VideoSora: A Review on Background, Technology, Limitati
ons, and Opportunities of Large Vision Models
Sora, Text-to-VideoMora: Enabling Generalist Video Generation via A M
ulti-Agent Framework
ArXiv
SegmentationLISA: Reasoning Segmentation via Large Language Mo
del
ArXiv
SegmentationGRES: Generalized Referring Expression Segmentatio
n
SegmentationGeneralized Decoding for Pixel, Image, and Languag
e
ArXiv
SegmentationSEEM: Segment Everything Everywhere All at OnceArXiv, GitHub
SegmentationSegGPT: Segmenting Everything In ContextArXiv
SegmentationGrounded SAM: Assembling Open-World Models for Div
erse Visual Tasks
ArXiv
ScalingLeave No Context Behind: Efficient Infinite Contex
t Transformers with Infini-attention
ArXiv
SLM, ScalingTextbooks Are All You NeedArXiv
Robot, Zero-shotBC-Z: Zero-Shot Task Generalization with Robotic I
mitation Learning
ArXiv
Robot, Zero-shotUniversal Manipulation Interface: In-The-Wild Robo
t Teaching Without In-The-Wild Robots
ArXiv
Robot, Zero-shotMirage: Cross-Embodiment Zero-Shot Policy Transfer
with Cross-Painting
ArXiv
Robot, Task-Decompose, Zero-shotLanguage Models as Zero-Shot Planners: Extracting
Actionable Knowledge for Embodied Agents
ArXiv2022/01/18
Robot, Task-DecomposeSayPlan: Grounding Large Language Models using 3D
Scene Graphs for Scalable Robot Task Planning
ArXiv2023/07/12
Robot, TAMPLLM3:Large Language Model-based Task and Motion Pl
anning with Motion Failure Reasoning
ArXiv
Robot, TAMPTask and Motion Planning with Large Language Model
s for Object Rearrangement
ArXiv
Robot, SurveyToward General-Purpose Robots via Foundation Model
s: A Survey and Meta-Analysis
ArXiv2023/12/14
Robot, SurveyLanguage-conditioned Learning for Robotic Manipula
tion: A Survey
ArXiv2023/12/17
Robot, SurveyRobot Learning in the Era of Foundation Models: A
Survey
ArXiv2023/11/24
Robot, SurveyReal-World Robot Applications of Foundation Models
: A Review
ArXiv
RobotOK-Robot: What Really Matters in Integrating Open-
Knowledge Models for Robotics
ArXiv
RobotRoCo: Dialectic Multi-Robot Collaboration with Lar
ge Language Models
ArXiv
RobotInteractive Language: Talking to Robots in Real Ti
me
ArXiv
RobotReflexion: Language Agents with Verbal Reinforceme
nt Learning
ArXiv2023/03/20
RobotGenerative Expressive Robot Behaviors using Large
Language Models
ArXiv
RobotRoboCat: A Self-Improving Generalist Agent for Rob
otic Manipulation
RobotIntrospective Tips: Large Language Model for In-Co
ntext Decision Making
ArXiv
RobotPIVOT: Iterative Visual Prompting Elicits Actionab
le Knowledge for VLMs
ArXiv
RobotOCI-Robotics: Object-Centric Instruction Augmentat
ion for Robotic Manipulation
ArXiv
RobotDeliGrasp: Inferring Object Mass, Friction, and Co
mpliance with LLMs for Adaptive and Minimally Deforming Grasp Policies
ArXiv
RobotVoxPoser: Composable 3D Value Maps for Robotic Man
ipulation with Language Models
ArXiv
RobotCreative Robot Tool Use with Large Language ModelsArXiv
RobotAutoTAMP: Autoregressive Task and Motion Planning
with LLMs as Translators and Checkers
ArXiv
RoPERoFormer: Enhanced Transformer with Rotary Positio
n Embedding
ArXiv
Resource[Resource] PaperswithcodeArXiv
Resource[Resource] huggingfaceArXiv
Resource[Resource] dailyarxivArXiv
Resource[Resource] ConnectedpapersArXiv
Resource[Resource] SemanticscholarArXiv
Resource[Resource] AlphaSignalArXiv
Resource[Resource] arxiv-sanityArXiv
Reinforcement-Learning, VIMAFoMo Rewards: Can we cast foundation models as rew
ard functions?
ArXiv
Reinforcement-Learning, RobotTowards A Unified Agent with Foundation ModelsArXiv
Reinforcement-LearningLarge Language Models Are Semi-Parametric Reinforc
ement Learning Agents
ArXiv
Reinforcement-LearningRLang: A Declarative Language for Describing Parti
al World Knowledge to Reinforcement Learning Agents
ArXiv
Reasoning, Zero-shotLarge Language Models are Zero-Shot ReasonersArXiv
Reasoning, VLM, VQAMM-REACT: Prompting ChatGPT for Multimodal Reasoni
ng and Action
ArXiv2023/03/20
Reasoning, TableLarge Language Models are few(1)-shot Table Reason
ers
ArXiv
Reasoning, SymbolicSymbol-LLM: Leverage Language Models for Symbolic
System in Visual Human Activity Reasoning
ArXiv
Reasoning, SurveyReasoning with Language Model Prompting: A SurveyArXiv
Reasoning, RobotAlphaBlock: Embodied Finetuning for Vision-Languag
e Reasoning in Robot Manipulation
ArXiv
Reasoning, RewardLET’S REWARD STEP BY STEP: STEP-LEVEL REWARD MODEL
AS THE NAVIGATORS FOR REASONING
ArXiv
Reasoning, Reinforcement-LearningReFT: Reasoning with Reinforced Fine-Tuning
ReasoningSelection-Inference: Exploiting Large Language Mod
els for Interpretable Logical Reasoning
ArXiv
ReasoningReConcile: Round-Table Conference Improves Reasoni
ng via Consensus among Diverse LLMs.
ArXiv
ReasoningSelf-Discover: Large Language Models Self-Compose
Reasoning Structures
ArXiv
ReasoningChain-of-Thought Reasoning Without PromptingArXiv
ReasoningContrastive Chain-of-Thought PromptingArXiv
ReasoningRephrase and Respond(RaR)
ReasoningTake a Step Back: Evoking Reasoning via Abstractio
n in Large Language Models
ArXiv
ReasoningSTaR: Bootstrapping Reasoning With ReasoningArXiv2022/05/28
ReasoningThe Impact of Reasoning Step Length on Large Langu
age Models
ArXiv
ReasoningBeyond Natural Language: LLMs Leveraging Alternati
ve Formats for Enhanced Reasoning and Communication
ArXiv
ReasoningLarge Language Models as General Pattern MachinesArXiv
RLHF, Reinforcement-Learning, SurveyA Survey of Reinforcement Learning from Human Feed
back
RLHFSecrets of RLHF in Large Language Models Part II:
Reward Modeling
ArXiv
RAG, Temporal LogicsFreshLLMs: Refreshing Large Language Models with S
earch Engine Augmentation
ArXiv
RAG, SurveyLarge Language Models for Information Retrieval: A
Survey
RAG, SurveyRetrieval-Augmented Generation for Large Language
RAG, SurveyRetrieval-Augmented Generation for Large Language
Models: A Survey
ArXiv
RAGTraining Language Models with Memory Augmentation
RAGSelf-RAG: Learning to Retrieve, Generate, and Crit
ique through Self-Reflection
RAGRAG-Fusion: a New Take on Retrieval-Augmented Gene
ration
ArXiv
RAGRAFT: Adapting Language Model to Domain Specific R
AG
ArXiv
RAGAdaptive-RAG: Learning to Adapt Retrieval-Augmente
d Large Language Models through Question Complexity
ArXiv
RAGRAG vs Fine-tuning: Pipelines, Tradeoffs, and a Ca
se Study on Agriculture
ArXiv
RAGFine-Tuning or Retrieval? Comparing Knowledge Inje
ction in LLMs
ArXiv
Quantization, ScalingSliceGPT: Compress Large Language Models by Deleti
ng Rows and Columns
ArXiv
Prompting, SurveyA Systematic Survey of Prompt Engineering in Large
Language Models: Techniques and Applications
ArXiv
Prompting, Robot, Zero-shotZero-Shot Task Generalization with Multi-Task Deep
Reinforcement Learning
ArXiv
PromptingContrastive Chain-of-Thought PromptingArXiv
PersonalCitation, RobotText2Motion: From Natural Language Instructions to
Feasible Plans
ArXiv
Perception, Video, VisionCLIP4Clip: An Empirical Study of CLIP for End to E
nd Video Clip Retrieval
ArXiv
Perception, Task-DecomposeDoReMi: Grounding Language Model by Detecting and
Recovering from Plan-Execution Misalignment
ArXiv2023/07/01
Perception, Robot, SegmentationLanguage Segment-Anything
Perception, RobotLiDAR-LLM: Exploring the Potential of Large Langua
ge Models for 3D LiDAR Understanding
ArXiv2023/12/21
Perception, Reasoning, RobotReasoning Grasping via Multimodal Large Language M
odel
ArXiv
Perception, ReasoningDetGPT: Detect What You Need via ReasoningArXiv
Perception, ReasoningLenna: Language Enhanced Reasoning Detection Assis
tant
ArXiv
PerceptionSimple Open-Vocabulary Object Detection with Visio
n Transformers
ArXiv2022/05/12
PerceptionGrounded Language-Image Pre-trainingArXiv2021/12/07
PerceptionGrounding DINO: Marrying DINO with Grounded Pre-Tr
aining for Open-Set Object Detection
ArXiv2023/03/09
PerceptionPointCLIP: Point Cloud Understanding by CLIPArXiv2021/12/04
PerceptionDINO: DETR with Improved DeNoising Anchor Boxes fo
r End-to-End Object Detection
ArXiv
PerceptionRecognize Anything: A Strong Image Tagging ModelArXiv
PerceptionSimple Open-Vocabulary Object Detection with Visio
n Transformers
ArXiv
PerceptionSigmoid Loss for Language Image Pre-TrainingArXiv
PackageLlamaIndexGitHub
PackageLangChainGitHub
Packageh2oGPTGitHub
PackageDifyGitHub
PackageAlpaca-LoRAGitHub
PackagePromptlayerGitHub
PackageunslothGitHub
PackageInstructor: Structured LLM OutputsGitHub
PRMLet's Verify Step by StepArXiv
PRMLet's reward step by step: Step-Level reward model
as the Navigators for Reasoning
ArXiv
PPO, RLHF, Reinforcement-LearningSecrets of RLHF in Large Language Models Part I: P
PO
ArXiv2024/02/01
Open-source, VLMOpenFlamingo: An Open-Source Framework for Trainin
g Large Autoregressive Vision-Language Models
ArXiv2023/08/02
Open-source, SLMRecurrentGemma: Moving Past Transformers for Effic
ient Open Language Models
ArXiv
Open-source, PerceptionGrounding DINO: Marrying DINO with Grounded Pre-Tr
aining for Open-Set Object Detection
ArXiv
Open-sourceGemma: Introducing new state-of-the-art open model
s
ArXiv
Open-sourceMistral 7BArXiv
Open-sourceQwen Technical ReportArXiv
Navigation, Reasoning, VisionNavGPT: Explicit Reasoning in Vision-and-Language
Navigation with Large Language Models
ArXiv
Natural-Language-as-Polices, RobotRT-H: Action Hierarchies Using LanguageArXiv
Multimodal, Robot, VLMOpen-World Object Manipulation using Pre-trained V
ision-Language Models
ArXiv2023/03/02
Multimodal, RobotMOMA-Force: Visual-Force Imitation for Real-World
Mobile Manipulation
ArXiv2023/08/07
Multimodal, RobotFlamingo: a Visual Language Model for Few-Shot Lea
rning
ArXiv2022/04/29
Multi-Images, VLMMantis: Multi-Image Instruction TuningArXiv, GitHub
MoESwitch Transformers: Scaling to Trillion Parameter
Models with Simple and Efficient Sparsity
ArXiv
MoESparse MoE as the New Dropout: Scaling Dense and S
elf-Slimmable Transformers
ArXiv
Mixtral, MoEMixtral of ExpertsArXiv
Memory, RobotLLM as A Robotic Brain: Unifying Egocentric Memory
and Control
ArXiv2023/04/19
Memory, Reinforcement-LearningSemantic HELM: A Human-Readable Memory for Reinfor
cement Learning
Math, ReasoningDeepSeekMath: Pushing the Limits of Mathematical R
easoning in Open Language Models
ArXiv, GitHub
Math, PRMMath-Shepherd: Verify and Reinforce LLMs Step-by-s
tep without Human Annotations
ArXiv
MathWizardMath: Empowering Mathematical Reasoning for
Large Language Models via Reinforced Evol-Instruct
ArXiv
MathLlemma: An Open Language Model For MathematicsArXiv
Low-level-action, RobotSayTap: Language to Quadrupedal LocomotionArXiv2023/06/13
Low-level-action, RobotPrompt a Robot to Walk with Large Language ModelsArXiv2023/09/18
LoRA, ScalingLoRA: Low-Rank Adaptation of Large Language ModelsArXiv
LoRA, ScalingVera: A General-Purpose Plausibility Estimation Mo
del for Commonsense Statements
ArXiv
LoRALoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A
Technical Report
ArXiv
LabTencent AI Lab - AppAgent, WebVoyager
LabDeepWisdom - MetaGPT
LabReworkd AI - AgentGPT
LabOpenBMB - ChatDev, XAgent, AgentVerse
LabXLANG NLP Lab - OpenAgents
LabRutgers University, AGI Research - OpenAGI
LabKnowledge Engineering Group (KEG) & Data Mining at
Tsinghua University - CogVLM
LabOpenGVLabGitHub
LabImperial College London - Zeroshot trajectory
Labsensetime
Labtsinghua
LabFudan NLP Group
LabPenn State University
LLaVA, VLMTinyLLaVA: A Framework of Small-scale Large Multim
odal Models
ArXiv
LLaVA, MoE, VLMMoE-LLaVA: Mixture of Experts for Large Vision-Lan
guage Models
LLaMA, Lightweight, Open-sourceMobiLlama: Towards Accurate and Lightweight Fully
Transparent GPT
LLM, Zero-shotGPT4Vis: What Can GPT-4 Do for Zero-shot Visual Re
cognition?
ArXiv2023/11/27
LLM, Temporal LogicsNL2TL: Transforming Natural Languages to Temporal
Logics using Large Language Models
ArXiv2023/05/12
LLM, SurveyA Survey of Large Language ModelsArXiv2023/03/31
LLM, SpacialCan Large Language Models be Good Path Planners? A
Benchmark and Investigation on Spatial-temporal Reasoning
ArXiv2023/10/05
LLM, ScalingBitNet: Scaling 1-bit Transformers for Large Langu
age Models
ArXiv
LLM, Robot, Task-DecomposeDo As I Can, Not As I Say: Grounding Language in R
obotic Affordances
ArXiv2022/04/04
LLM, Robot, SurveyLarge Language Models for Robotics: A SurveyArXiv
LLM, Reasoning, SurveyTowards Reasoning in Large Language Models: A Surv
ey
ArXiv2022/12/20
LLM, QuantizationThe Era of 1-bit LLMs: All Large Language Models a
re in 1.58 Bits
ArXiv
LLM, PersonalCitation, Robot, Zero-shotLanguage Models as Zero-Shot Trajectory GeneratorsArXiv
LLM, PersonalCitation, RobotTree-Planner: Efficient Close-loop Task Planning w
ith Large Language Models01
LLM, Open-sourceA self-hosted, offline, ChatGPT-like chatbot, powe
red by Llama 2. 100% private, with no data leaving your device.
GitHub
LLM, Open-sourceOpenFlamingo: An Open-Source Framework for Trainin
g Large Autoregressive Vision-Language Models
ArXiv2023/08/02
LLM, Open-sourceInstructBLIP: Towards General-purpose Vision-Langu
age Models with Instruction Tuning
ArXiv2023/05/11
LLM, Open-sourceChatBridge: Bridging Modalities with Large Languag
e Model as a Language Catalyst
ArXiv2023/05/25
LLM, MemoryMemoryBank: Enhancing Large Language Models with L
ong-Term Memory
ArXiv2023/05/17
LLM, LeaderboardLMSYS Chatbot Arena Leaderboard
LLMLanguage Models are Few-Shot LearnersArXiv2020/05/28
Intaractive, OpenGVLab, VLMInternGPT: Solving Vision-Centric Tasks by Interac
ting with ChatGPT Beyond Language
ArXiv2023/05/09
Instruction-Turning, SurveyIs Prompt All You Need? No. A Comprehensive and Br
oader View of Instruction Learning
Instruction-Turning, SurveyVision-Language Instruction Tuning: A Review and A
nalysis
ArXiv
Instruction-Turning, SurveyA Closer Look at the Limitations of Instruction Tu
ning
ArXiv
Instruction-Turning, SurveyA Survey on Data Selection for LLM Instruction Tun
ing
ArXiv
Instruction-Turning, SelfSelf-Instruct: Aligning Language Models with Self-
Generated Instructions
ArXiv
Instruction-Turning, LLM, Zero-shotFinetuned Language Models Are Zero-Shot LearnersArXiv2021/09/03
Instruction-Turning, LLM, SurveyInstruction Tuning for Large Language Models: A Su
rvey
Instruction-Turning, LLM, PEFTVisual Instruction TuningArXiv2023/04/17
Instruction-Turning, LLM, PEFTLLaMA-Adapter: Efficient Fine-tuning of Language M
odels with Zero-init Attention
ArXiv2023/03/28
Instruction-Turning, LLMTraining language models to follow instructions wi
th human feedback
ArXiv2022/03/04
Instruction-Turning, LLMMiniGPT-4: Enhancing Vision-Language Understanding
with Advanced Large Language Models
ArXiv2023/04/20
Instruction-Turning, LLMSelf-Instruct: Aligning Language Models with Self-
Generated Instructions
ArXiv2022/12/20
Instruction-TurningA Closer Look at the Limitations of Instruction Tu
ning
Instruction-TurningExploring Format Consistency for Instruction Tunin
g
Instruction-TurningExploring the Benefits of Training Expert Language
Models over Instruction Tuning
ArXiv2023/02/06
Instruction-TurningTuna: Instruction Tuning using Feedback from Large
Language Models
ArXiv2023/03/06
In-Context-Learning, VisionWhat Makes Good Examples for Visual In-Context Lea
rning?
In-Context-Learning, VisionVisual Prompting via Image InpaintingArXiv
In-Context-Learning, VideoPrompting Visual-Language Models for Efficient Vid
eo Understanding
In-Context-Learning, VQAVisualCOMET: Reasoning about the Dynamic Context o
f a Still Image
ArXiv2020/04/22
In-Context-Learning, VQASINC: Self-Supervised In-Context Learning for Visi
on-Language Tasks
ArXiv2023/07/15
In-Context-Learning, SurveyA Survey on In-context LearningArXiv
In-Context-Learning, ScalingStructured Prompting: Scaling In-Context Learning
to 1,000 Examples
ArXiv2020/03/06
In-Context-Learning, ScalingRethinking the Role of Scale for In-Context Learni
ng: An Interpretability-based Case Study at 66 Billion Scale
ArXiv2022/03/06
In-Context-Learning, Reinforcement-LearningAMAGO: Scalable In-Context Reinforcement Learning
for Adaptive Agents
ArXiv
In-Context-Learning, Prompt-TuningVisual Prompt TuningArXiv
In-Context-Learning, Perception, VisionVisual In-Context PromptingArXiv
In-Context-Learning, Many-Shot, ReasoningMany-Shot In-Context LearningArXiv
In-Context-Learning, Instruction-TurningIn-Context Instruction Learning
In-Context-LearningReAct: Synergizing Reasoning and Acting in Languag
e Models
ArXiv2023/03/20
In-Context-LearningSmall Models are Valuable Plug-ins for Large Langu
age Models
ArXiv2023/05/15
In-Context-LearningGenerative Agents: Interactive Simulacra of Human
Behavior
ArXiv2023/04/07
In-Context-LearningBeyond the Imitation Game: Quantifying and extrapo
lating the capabilities of language models
ArXiv2022/06/09
In-Context-LearningWhat does CLIP know about a red circle? Visual pro
mpt engineering for VLMs
ArXiv
In-Context-LearningCan large language models explore in-context?ArXiv
Image, LLaMA, PerceptionLLaMA-VID: An Image is Worth 2 Tokens in Large Lan
guage Models
ArXiv
Hallucination, SurveyCombating Misinformation in the Age of LLMs: Oppor
tunities and Challenges
ArXiv
Gym, PPO, Reinforcement-Learning, SurveyCan Language Agents Approach the Performance of RL
? An Empirical Study On OpenAI Gym
ArXiv
Grounding, Reinforcement-LearningGrounding Large Language Models in Interactive Env
ironments with Online Reinforcement Learning
ArXiv
Grounding, ReasoningVisually Grounded Reasoning across Languages and C
ultures
ArXiv
GroundingV-IRL: Grounding Virtual Intelligence in Real Life
Google, GroundingGLaMM: Pixel Grounding Large Multimodal ModelArXiv
Generation, SurveyAdvances in 3D Generation: A Survey
Generation, Robot, Zero-shotZero-Shot Robotic Manipulation with Pretrained Ima
ge-Editing Diffusion Models
ArXiv
Generation, Robot, Zero-shotTowards Generalizable Zero-Shot Manipulationvia Tr
anslating Human Interaction Plans
GPT4V, Robot, VLMClosed-Loop Open-Vocabulary Mobile Manipulation wi
th GPT-4V
ArXiv
GPT4, LLMGPT-4 Technical ReportArXiv2023/03/15
GPT4, Instruction-TurningINSTRUCTION TUNING WITH GPT-4ArXiv
GPT4, Gemini, LLMGemini vs GPT-4V: A Preliminary Comparison and Com
bination of Vision-Language Models Through Qualitative Cases
ArXiv2023/12/22
Foundation, Robot, SurveyFoundation Models in Robotics: Applications, Chall
enges, and the Future
ArXiv2023/12/13
Foundation, LLaMA, VisionVisionLLaMA: A Unified LLaMA Interface for Vision
Tasks
ArXiv
Foundation, LLM, Open-sourceCode Llama: Open Foundation Models for CodeArXiv
Foundation, LLM, Open-sourceLLaMA: Open and Efficient Foundation Language Mode
ls
ArXiv2023/02/27
Feedback, RobotCorrecting Robot Plans with Natural Language Feedb
ack
ArXiv
Feedback, RobotLearning to Learn Faster from Human Feedback with
Language Model Predictive Control
ArXiv
Feedback, RobotREFLECT: Summarizing Robot Experiences for Failure
Explanation and Correction
ArXiv2023/06/27
Feedback, In-Context-Learning, RobotInCoRo: In-Context Learning for Robotics Control w
ith Feedback Loops
ArXiv
Evaluation, LLM, SurveyA Survey on Evaluation of Large Language ModelsArXiv
Evaluationsimple-evalsGitHub
End2End, Multimodal, RobotVIMA: General Robot Manipulation with Multimodal P
rompts
ArXiv2022/10/06
End2End, Multimodal, RobotPaLM-E: An Embodied Multimodal Language ModelArXiv2023/03/06
End2End, Multimodal, RobotPhysically Grounded Vision-Language Models for Rob
otic Manipulation
ArXiv2023/09/05
EnbodiedEmbodied Question AnsweringArXiv
Embodied, World-modelLanguage Models Meet World Models: Embodied Experi
ences Enhance Language Models
Embodied, Robot, Task-DecomposeEmbodied Task Planning with Large Language ModelsArXiv2023/07/04
Embodied, RobotLarge Language Models as Generalizable Policies fo
r Embodied Tasks
ArXiv
Embodied, Reasoning, RobotNatural Language as Policies: Reasoning for Coordi
nate-Level Embodied Control with LLMs
ArXiv, GitHub2024/03/20
Embodied, LLM, Robot, SurveyThe Development of LLMs for Embodied NavigationArXiv2023/11/01
Driving, SpacialGPT-Driver: Learning to Drive with GPTArXiv2023/10/02
Drive, SurveyA Survey on Multimodal Large Language Models for A
utonomous Driving
ArXiv
Distilling, SurveyA Survey on Knowledge Distillation of Large Langua
ge Models
DistillingDistilling Step-by-Step! Outperforming Larger Lang
uage Models with Less Training Data and Smaller Model Sizes01
ArXiv
Diffusion, Text-to-ImageMastering Text-to-Image Diffusion: Recaptioning, P
lanning, and Generating with Multimodal LLMs
ArXiv
Diffusion, SurveyOn the Design Fundamentals of Diffusion Models: A
Survey
ArXiv
Diffusion, Robot3D Diffusion PolicyArXiv
DiffusionA latent text-to-image diffusion model
Demonstration, GPT4, PersonalCitation, Robot, VLMGPT-4V(ision) for Robotics: Multimodal Task Planni
ng from Human Demonstration
Datatset, LLM, SurveyA Survey on Data Selection for Language ModelsArXiv
Datatset, Instruction-TurningREVO-LION: EVALUATING AND REFINING VISION LANGUAGE
INSTRUCTION TUNING DATASETS
Datatset, Instruction-TurningSynthetic Data (Almost) from Scratch: Generalized
Instruction Tuning for Language Models
DatatsetPRM800K: A Process Supervision DatasetGitHub
Data-generation, RobotGenSim: Generating Robotic Simulation Tasks via La
rge Language Models
ArXiv2023/10/02
Data-generation, RobotRoboGen: Towards Unleashing Infinite Data for Auto
mated Robot Learning via Generative Simulation
ArXiv2023/11/02
DPO, PPO, RLHFA Comprehensive Survey of LLM Alignment Techniques
: RLHF, RLAIF, PPO, DPO and More
ArXiv
DPOIs DPO Superior to PPO for LLM Alignment? A Compre
hensive Study
ArXiv
Context-Window, ScalingInfini-gram: Scaling Unbounded n-gram Language Mod
els to a Trillion Tokens
Context-Window, ScalingLONGNET: Scaling Transformers to 1,000,000,000 Tok
ens
ArXiv2023/07/01
Context-Window, Reasoning, RoPE, ScalingResonance RoPE: Improving Context Length Generaliz
ation of Large Language Models
ArXiv
Context-Window, LLM, RoPE, ScalingLongRoPE: Extending LLM Context Window Beyond 2 Mi
llion Tokens
ArXiv
Context-Window, Foundation, Gemini, LLM, ScalingGemini 1.5: Unlocking multimodal understanding acr
oss millions of tokens of context
Context-Window, FoundationMamba: Linear-Time Sequence Modeling with Selectiv
e State Spaces
ArXiv
Context-WindowRoFormer: Enhanced Transformer with Rotary Positio
n Embedding
ArXiv
Context-Awere, Context-WindowDynaCon: Dynamic Robot Planner with Contextual Awa
reness via LLMs
ArXiv
Computer-Resource, ScalingFlashAttention: Fast and Memory-Efficient Exact At
tention with IO-Awareness
ArXiv
Compress, Scaling(Long)LLMLingua: Enhancing Large Language Model In
ference via Prompt Compression
ArXiv
Compress, Quantization, SurveyA Survey on Model Compression for Large Language M
odels
ArXiv
Compress, PromptingLearning to Compress Prompts with Gist TokensArXiv
Code-as-Policies, VLM, VQAVisual Programming: Compositional visual reasoning
without training
ArXiv2022/11/18
Code-as-Policies, RobotSMART-LLM: Smart Multi-Agent Robot Task Planning u
sing Large Language Models
ArXiv2023/09/18
Code-as-Policies, RobotRoboScript: Code Generation for Free-Form Manipula
tion Tasks across Real and Simulation
ArXiv
Code-as-Policies, RobotCreative Robot Tool Use with Large Language ModelsArXiv
Code-as-Policies, Reinforcement-Learning, RewardCode as Reward: Empowering Reinforcement Learning
with VLMs
ArXiv
Code-as-Policies, Reasoning, VLM, VQAViperGPT: Visual Inference via Python Execution fo
r Reasoning
ArXiv2023/03/14
Code-as-Policies, ReasoningChain of Code: Reasoning with a Language Model-Aug
mented Code Emulator
ArXiv
Code-as-Policies, PersonalCitation, Robot, Zero-shotSocratic Models: Composing Zero-Shot Multimodal Re
asoning with Language
ArXiv2022/04/01
Code-as-Policies, PersonalCitation, Robot, State-ManageStatler: State-Maintaining Language Models for Emb
odied Reasoning
ArXiv2023/06/30
Code-as-Policies, PersonalCitation, RobotProgPrompt: Generating Situated Robot Task Plans u
sing Large Language Models
ArXiv2022/09/22
Code-as-Policies, PersonalCitation, RobotRoboCodeX:Multi-modal Code Generation forRobotic B
ehavior Synthesis
ArXiv
Code-as-Policies, PersonalCitation, RobotRoboGPT: an intelligent agent of making embodied l
ong-term decisions for daily instruction tasks
Code-as-Policies, PersonalCitation, RobotChatGPT for Robotics: Design Principles and Model
Abilities
Code-as-Policies, Multimodal, OpenGVLab, PersonalCitation, RobotInstruct2Act: Mapping Multi-modality Instructions
to Robotic Actions with Large Language Model
ArXiv2023/05/18
Code-as-Policies, Embodied, PersonalCitation, RobotCode as Policies: Language Model Programs for Embo
died Control
ArXiv2022/09/16
Code-as-Policies, Embodied, PersonalCitation, Reasoning, Robot, Task-DecomposeInner Monologue: Embodied Reasoning through Planni
ng with Language Models
ArXiv
Code-LLM, Front-EndDesign2Code: How Far Are We From Automating Front-
End Engineering?
ArXiv
Code-LLMStarCoder 2 and The Stack v2: The Next Generation
Chain-of-Thought, Reasoning, TableChain-of-table: Evolving tables in the reasoning c
hain for table understanding
ArXiv
Chain-of-Thought, Reasoning, SurveyTowards Understanding Chain-of-Thought Prompting:
An Empirical Study of What Matters
ArXiv2023/12/20
Chain-of-Thought, Reasoning, SurveyA Survey of Chain of Thought Reasoning: Advances,
Frontiers and Future
ArXiv2023/09/27
Chain-of-Thought, ReasoningChain-of-Thought Prompting Elicits Reasoning in La
rge Language Models
ArXiv2022/01/28
Chain-of-Thought, ReasoningTree of Thoughts: Deliberate Problem Solving with
Large Language Models
ArXiv2023/05/17
Chain-of-Thought, ReasoningMultimodal Chain-of-Thought Reasoning in Language
Models
ArXiv2023/02/02
Chain-of-Thought, ReasoningVerify-and-Edit: A Knowledge-Enhanced Chain-of-Tho
ught Framework
ArXiv2023/05/05
Chain-of-Thought, ReasoningSkeleton-of-Thought: Large Language Models Can Do
Parallel Decoding
ArXiv2023/07/28
Chain-of-Thought, ReasoningRethinking with Retrieval: Faithful Large Language
Model Inference
ArXiv2022/12/31
Chain-of-Thought, ReasoningSelf-Consistency Improves Chain of Thought Reasoni
ng in Language Models
ArXiv2022/03/21
Chain-of-Thought, ReasoningChain-of-Thought Hub: A Continuous Effort to Measu
re Large Language Models' Reasoning Performance
ArXiv2023/05/26
Chain-of-Thought, ReasoningSkeleton-of-Thought: Prompting LLMs for Efficient
Parallel Generation
ArXiv
Chain-of-Thought, PromptingChain-of-Thought Reasoning Without PromptingArXiv
Chain-of-Thought, Planning, ReasoningSelfCheck: Using LLMs to Zero-Shot Check Their Own
Step-by-Step Reasoning
ArXiv2023/08/01
Chain-of-Thought, In-Context-Learning, SelfMeasuring and Narrowing the Compositionality Gap i
n Language Models
ArXiv2022/10/07
Chain-of-Thought, In-Context-Learning, SelfSelf-Polish: Enhance Reasoning in Large Language M
odels via Problem Refinement
ArXiv2023/05/23
Chain-of-Thought, In-Context-LearningChain-of-Table: Evolving Tables in the Reasoning C
hain for Table Understanding
ArXiv
Chain-of-Thought, In-Context-LearningSelf-Refine: Iterative Refinement with Self-Feedba
ck
ArXiv2023/03/30
Chain-of-Thought, In-Context-LearningPlan-and-Solve Prompting: Improving Zero-Shot Chai
n-of-Thought Reasoning by Large Language Models
ArXiv2023/05/06
Chain-of-Thought, In-Context-LearningPAL: Program-aided Language ModelsArXiv2022/11/18
Chain-of-Thought, In-Context-LearningReasoning with Language Model is Planning with Wor
ld Model
ArXiv2023/05/24
Chain-of-Thought, In-Context-LearningLeast-to-Most Prompting Enables Complex Reasoning
in Large Language Models
ArXiv2022/05/21
Chain-of-Thought, In-Context-LearningComplexity-Based Prompting for Multi-Step Reasonin
g
ArXiv2022/10/03
Chain-of-Thought, In-Context-LearningMaieutic Prompting: Logically Consistent Reasoning
with Recursive Explanations
ArXiv2022/05/24
Chain-of-Thought, In-Context-LearningAlgorithm of Thoughts: Enhancing Exploration of Id
eas in Large Language Models
ArXiv2023/08/20
Chain-of-Thought, GPT4, Reasoning, RobotLook Before You Leap: Unveiling the Power ofGPT-4V
in Robotic Vision-Language Planning
ArXiv2023/11/29
Chain-of-Thought, Embodied, RobotEgoCOT: Embodied Chain-of-Thought Dataset for Visi
on Language Pre-training
ArXiv
Chain-of-Thought, Embodied, PersonalCitation, Robot, Task-DecomposeEmbodiedGPT: Vision-Language Pre-Training via Embo
died Chain of Thought
ArXiv2023/05/24
Chain-of-Thought, Code-as-Policies, PersonalCitation, RobotDemo2Code: From Summarizing Demonstrations to Synt
hesizing Code via Extended Chain-of-Thought
ArXiv
Chain-of-Thought, Code-as-PoliciesChain of Code: Reasoning with a Language Model-Aug
mented Code Emulator
ArXiv
Caption, VideoPLLaVA : Parameter-free LLaVA Extension from Image
s to Videos for Video Dense Captioning
ArXiv
Caption, VLM, VQACaption Anything: Interactive Image Description wi
th Diverse Multimodal Controls
ArXiv2023/05/04
CRAG, RAGCorrective Retrieval Augmented GenerationArXiv
Brain, Instruction-TurningInstruction-tuning Aligns LLMs to the Human BrainArXiv
Brain, ConsciousCould a Large Language Model be Conscious?
BrainLLM-BRAIn: AI-driven Fast Generation of Robot Beha
viour Tree based on Large Language Model
ArXiv
BrainA Neuro-Mimetic Realization of the Common Model of
Cognition via Hebbian Learning and Free Energy Minimization
Benchmark, Sora, Text-to-VideoLIDA: A Tool for Automatic Generation of Grammar-A
gnostic Visualizations and Infographics using Large Language Models01
Benchmark, In-Context-LearningARB: Advanced Reasoning Benchmark for Large Langua
ge Models
ArXiv2023/07/25
Benchmark, In-Context-LearningPlanBench: An Extensible Benchmark for Evaluating
Large Language Models on Planning and Reasoning about Change
ArXiv2022/06/21
Benchmark, GPT4Sparks of Artificial General Intelligence: Early e
xperiments with GPT-4
Awesome Repo, VLMawesome-vlm-architecturesGitHub
Awesome Repo, SurveyLLMSurveyGitHub
Awesome Repo, RobotAwesome-LLM-RoboticsGitHub
Awesome Repo, ReasoningAwesome LLM ReasoningGitHub
Awesome Repo, ReasoningAwesome-Reasoning-Foundation-ModelsGitHub
Awesome Repo, RLHF, Reinforcement-LearningAwesome RLHF (RL with Human Feedback)GitHub
Awesome Repo, Perception, VLMAwesome Vision-Language NavigationGitHub
Awesome Repo, PackageAwesome LLMOpsGitHub
Awesome Repo, MultimodalAwesome-Multimodal-LLMGitHub
Awesome Repo, MultimodalAwesome-Multimodal-Large-Language-ModelsGitHub
Awesome Repo, Math, ScienceAwesome Scientific Language ModelsGitHub
Awesome Repo, LLM, VisionLLM-in-VisionGitHub
Awesome Repo, LLM, VLMMultimodal & Large Language ModelsGitHub
Awesome Repo, LLM, SurveyAwesome-LLM-SurveyGitHub
Awesome Repo, LLM, RobotEverything-LLMs-And-RoboticsGitHub
Awesome Repo, LLM, LeaderboardLLM-LeaderboardGitHub
Awesome Repo, LLMAwesome-LLMGitHub
Awesome Repo, Koreanawesome-korean-llmGitHub
Awesome Repo, Japanese, LLM日本語LLMまとめGitHub
Awesome Repo, In-Context-LearningPaper List for In-context LearningGitHub
Awesome Repo, IROS, RobotIROS2023PaperListGitHub
Awesome Repo, Hallucination, SurveyA Survey on Hallucination in Large Language Models
: Principles, Taxonomy, Challenges, and Open Questions
ArXiv, GitHub
Awesome Repo, EmbodiedAwesome Embodied VisionGitHub
Awesome Repo, DiffusionAwesome-Diffusion-ModelsGitHub
Awesome Repo, CompressAwesome LLM CompressionGitHub
Awesome Repo, ChineseAwesome-Chinese-LLMGitHub
Awesome Repo, Chain-of-ThoughtChain-of-ThoughtsPapersGitHub
Automate, PromptingLarge Language Models Are Human-Level Prompt Engin
eers
ArXiv2022/11/03
Automate, Chain-of-Thought, ReasoningAutomatic Chain of Thought Prompting in Large Lang
uage Models
ArXiv2022/10/07
Audio2Video, Diffusion, Generation, VideoEMO: Emote Portrait Alive - Generating Expressive
Portrait Videos with Audio2Video Diffusion Model under Weak Conditions
ArXiv
AudioRobust Speech Recognition via Large-Scale Weak Sup
ervision
Apple, VLMMM1: Methods, Analysis & Insights from Multimodal
LLM Pre-training
ArXiv
Apple, VLMGuiding Instruction-based Image Editing via Multim
odal Large Language Models
ArXiv
Apple, RobotLarge Language Models as Generalizable Policies fo
r Embodied Tasks
ArXiv
Apple, LLM, Open-sourceOpenELM: An Efficient Language Model Family with O
pen Training and Inference Framework
ArXiv
Apple, LLMReALM: Reference Resolution As Language ModelingArXiv
Apple, LLMLLM in a flash: Efficient Large Language Model Inf
erence with Limited Memory
ArXiv
Apple, In-Context-Learning, PerceptionSAM-CLIP: Merging Vision Foundation Models towards
Semantic and Spatial Understanding
ArXiv
Apple, Code-as-Policies, RobotExecutable Code Actions Elicit Better LLM AgentsArXiv
AppleFerret-v2: An Improved Baseline for Referring and
Grounding with Large Language Models
ArXiv
AppleFerret-UI: Grounded Mobile UI Understanding with M
ultimodal LLMs
ArXiv
AppleFerret: Refer and Ground Anything Anywhere at Any
Granularity
ArXiv
Anything, LLM, Open-source, Perception, SegmentationSegment AnythingArXiv2023/04/05
Anything, DepthDepth Anything: Unleashing the Power of Large-Scal
e Unlabeled Data
ArXiv
Anything, Caption, Perception, SegmentationSegment and Caption AnythingArXiv
Anything, CLIP, PerceptionSAM-CLIP: Merging Vision Foundation Models towards
Semantic and Spatial Understanding
ArXiv
Agent-Project, Code-LLMopen-interpreterGitHub
Agent, WebWebLINX: Real-World Website Navigation with Multi-
Turn Dialogue
ArXiv
Agent, WebWebVoyager: Building an End-to-End Web Agent with
Large Multimodal Models
ArXiv
Agent, WebOmniACT: A Dataset and Benchmark for Enabling Mult
imodal Generalist Autonomous Agents for Desktop and Web
ArXiv
Agent, WebOS-Copilot: Towards Generalist Computer Agents wit
h Self-Improvement
ArXiv
Agent, Video-for-AgentVideo as the New Language for Real-World Decision
Making
Agent, VLMAssistGPT: A General Multi-modal Assistant that ca
n Plan, Execute, Inspect, and Learn
ArXiv
Agent, ToolGorilla: Large Language Model Connected with Massi
ve APIs
ArXiv
Agent, ToolToolLLM: Facilitating Large Language Models to Mas
ter 16000+ Real-world APIs
ArXiv
Agent, SurveyA Survey on Large Language Model based Autonomous
Agents
ArXiv2023/08/22
Agent, SurveyThe Rise and Potential of Large Language Model Bas
ed Agents: A Survey
ArXiv2023/09/14
Agent, SurveyAgent AI: Surveying the Horizons of Multimodal Int
eraction
ArXiv
Agent, SurveyLarge Multimodal Agents: A SurveyArXiv
Agent, Soft-DevCommunicative Agents for Software DevelopmentGitHub
Agent, Soft-DevMetaGPT: Meta Programming for A Multi-Agent Collab
orative Framework
ArXiv
Agent, Robot, SurveyA Survey on LLM-based Autonomous AgentsGitHub
Agent, Reinforcement-Learning, RewardReward Design with Language ModelsArXiv2023/02/27
Agent, Reinforcement-Learning, RewardEAGER: Asking and Answering Questions for Automati
c Reward Shaping in Language-guided RL
ArXiv2022/06/20
Agent, Reinforcement-Learning, RewardText2Reward: Automated Dense Reward Function Gener
ation for Reinforcement Learning
ArXiv2023/09/20
Agent, Reinforcement-LearningEureka: Human-Level Reward Design via Coding Large
Language Models
ArXiv2023/10/19
Agent, Reinforcement-LearningLanguage to Rewards for Robotic Skill SynthesisArXiv2023/06/14
Agent, Reinforcement-LearningLanguage Instructed Reinforcement Learning for Hum
an-AI Coordination
ArXiv2023/04/13
Agent, Reinforcement-LearningGuiding Pretraining in Reinforcement Learning with
Large Language Models
ArXiv2023/02/13
Agent, Reinforcement-LearningSTARLING: SELF-SUPERVISED TRAINING OF TEXTBASED RE
INFORCEMENT LEARNING AGENT WITH LARGE LANGUAGE MODELS
Agent, Reasoning, Zero-shotAgent Instructs Large Language Models to be Genera
l Zero-Shot Reasoners
ArXiv2023/10/05
Agent, ReasoningPangu-Agent: A Fine-Tunable Generalist Agent with
Structured Reasoning
ArXiv
Agent, ReasoningAGENT INSTRUCTS LARGE LANGUAGE MODELS TO BE GENERA
L ZERO-SHOT REASONERS
ArXiv
Agent, Multimodal, RobotA Generalist AgentArXiv2022/05/12
Agent, MultiWar and Peace (WarAgent): Large Language Model-bas
ed Multi-Agent Simulation of World Wars
ArXiv
Agent, MobileAppYou Only Look at Screens: Multimodal Chain-of-Acti
on Agents
ArXiv, GitHub
Agent, Minecraft, Reinforcement-LearningRLAdapter: Bridging Large Language Models to Reinf
orcement Learning in Open Worlds
Agent, MinecraftVoyager: An Open-Ended Embodied Agent with Large L
anguage Models
ArXiv2023/05/25
Agent, MinecraftDescribe, Explain, Plan and Select: Interactive Pl
anning with Large Language Models Enables Open-World Multi-Task Agents
ArXiv2023/02/03
Agent, MinecraftLARP: Language-Agent Role Play for Open-World Game
s
ArXiv
Agent, MinecraftSteve-Eye: Equipping LLM-based Embodied Agents wit
h Visual Perception in Open Worlds
ArXiv
Agent, MinecraftS-Agents: Self-organizing Agents in Open-ended Env
ironment
ArXiv
Agent, MinecraftGhost in the Minecraft: Generally Capable Agents f
or Open-World Environments via Large Language Models with Text-based Knowledge and Memory01
ArXiv
Agent, Memory, RAG, RobotRAP: Retrieval-Augmented Planning with Contextual
Memory for Multimodal LLM Agents
ArXiv2024/02/06
Agent, Memory, MinecraftJARVIS-1: Open-World Multi-task Agents with Memory
-Augmented Multimodal Language Models
ArXiv2023/11/10
Agent, LLM, PlanningLLM-Planner: Few-Shot Grounded Planning for Embodi
ed Agents with Large Language Models
ArXiv
Agent, Instruction-TurningAgentTuning: Enabling Generalized Agent Abilities
For LLMs
ArXiv
Agent, GameLEARNING EMBODIED VISION-LANGUAGE PRO- GRAMMING FR
OM INSTRUCTION, EXPLORATION, AND ENVIRONMENTAL FEEDBACK
ArXiv
Agent, GUI, Web"What’s important here?": Opportunities and Challe
nges of Using LLMs in Retrieving Informatio from Web Interfaces
ArXiv
Agent, GUI, MobileAppMobile-Agent: Autonomous Multi-Modal Mobile Device
Agent with Visual Perception
Agent, GUI, MobileAppAppAgent: Multimodal Agents as Smartphone UsersArXiv
Agent, GUI, MobileAppYou Only Look at Screens: Multimodal Chain-of-Acti
on Agents
Agent, GUICogAgent: A Visual Language Model for GUI AgentsArXiv
Agent, GUIScreenAgent: A Computer Control Agent Driven by Vi
sual Language Large Model
GitHub
Agent, GUISeeClick: Harnessing GUI Grounding for Advanced Vi
sual GUI Agents
ArXiv
Agent, GPT4, WebGPT-4V(ision) is a Generalist Web Agent, if Ground
ed
ArXiv
Agent, Feedback, Reinforcement-Learning, RobotAccelerating Reinforcement Learning of Robotic Man
ipulations via Feedback from Large Language Models
ArXiv2023/11/04
Agent, Feedback, Reinforcement-LearningAdaRefiner: Refining Decisions of Language Models
with Adaptive Feedback
ArXiv2023/09/29
Agent, End2End, Game, RobotAn Interactive Agent Foundation ModelArXiv
Agent, Embodied, SurveyApplication of Pretrained Large Language Models in
Embodied Artificial Intelligence
ArXiv
Agent, Embodied, RobotAutoRT: Embodied Foundation Models for Large Scale
Orchestration of Robotic Agents
ArXiv
Agent, Embodied, RobotOPEx: A Component-Wise Analysis of LLM-Centric Age
nts in Embodied Instruction Following
ArXiv
Agent, EmbodiedOpenAgents: An Open Platform for Language Agents i
n the Wild
ArXiv, GitHub
Agent, EmbodiedLLM-Planner: Few-Shot Grounded Planning for Embodi
ed Agents with Large Language Models
ArXiv
Agent, EmbodiedEmbodied Multi-Modal Agent trained by an LLM from
a Parallel TextWorld
ArXiv
Agent, EmbodiedOctopus: Embodied Vision-Language Programmer from
Environmental Feedback
Agent, EmbodiedEmbodied Task Planning with Large Language ModelsArXiv
Agent, Diffusion, SpeechNaturalSpeech 3: Zero-Shot Speech Synthesis with F
actorized Codec and Diffusion Models
ArXiv
Agent, Code-as-PoliciesExecutable Code Actions Elicit Better LLM AgentsArXiv2024/01/24
Agent, Code-LLM, Code-as-Policies, SurveyIf LLM Is the Wizard, Then Code Is the Wand: A Sur
vey on How Code Empowers Large Language Models to Serve as Intelligent Agents
ArXiv
Agent, Code-LLMTaskWeaver: A Code-First Agent Framework
Agent, BlogLLM Powered Autonomous AgentsArXiv
Agent, Awesome Repo, LLMAwesome-Embodied-Agent-with-LLMsGitHub
Agent, Awesome Repo, LLMCoALA: Awesome Language AgentsArXiv, GitHub
Agent, Awesome Repo, Embodied, GroundingXLang Paper ReadingGitHub
Agent, Awesome RepoAwesome AI AgentsGitHub
Agent, Awesome RepoAutonomous AgentsGitHub
Agent, Awesome RepoAwesome-Papers-Autonomous-AgentGitHub
Agent, Awesome RepoAwesome Large Multimodal AgentsGitHub
Agent, Awesome RepoLLM Agents PapersGitHub
Agent, Awesome RepoAwesome LLM-Powered AgentGitHub
AgentXAgent: An Autonomous Agent for Complex Task Solvi
ng
AgentLLM-Powered Hierarchical Language Agent for Real-t
ime Human-AI Coordination
ArXiv
AgentAgentVerse: Facilitating Multi-Agent Collaboration
and Exploring Emergent Behaviors
ArXiv
AgentAgents: An Open-source Framework for Autonomous La
nguage Agents
ArXiv, GitHub
AgentAutoAgents: A Framework for Automatic Agent Genera
tion
GitHub
AgentDSPy: Compiling Declarative Language Model Calls i
nto Self-Improving Pipelines
ArXiv
AgentAutoGen: Enabling Next-Gen LLM Applications via Mu
lti-Agent Conversation
ArXiv
AgentCAMEL: Communicative Agents for “Mind” Exploration
of Large Language Model Society
ArXiv
AgentXAgent: An Autonomous Agent for Complex Task Solvi
ng
ArXiv
AgentGenerative Agents: Interactive Simulacra of Human
Behavior
ArXiv
AgentLLM+P: Empowering Large Language Models with Optim
al Planning Proficiency
ArXiv2023/04/22
AgentAgentSims: An Open-Source Sandbox for Large Langua
ge Model Evaluation
ArXiv2023/08/08
AgentAgents: An Open-source Framework for Autonomous La
nguage Agents
ArXiv
AgentMindAgent: Emergent Gaming InteractionArXiv
AgentInfiAgent: A Multi-Tool Agent for AI Operating Sys
tems
AgentPredictive Minds: LLMs As Atypical Active Inferenc
e Agents
AgentswarmsGitHub
AgentScreenAgent: A Vision Language Model-driven Comput
er Control Agent
ArXiv
AgentAssistGPT: A General Multi-modal Assistant that ca
n Plan, Execute, Inspect, and Learn
ArXiv
AgentPromptAgent: Strategic Planning with Language Mode
ls Enables Expert-level Prompt Optimization
ArXiv
AgentCognitive Architectures for Language AgentsArXiv
AgentAIOS: LLM Agent Operating SystemArXiv
AgentLLM as OS, Agents as Apps: Envisioning AIOS, Agent
s and the AIOS-Agent Ecosystem
ArXiv
AgentTowards General Computer Control: A Multimodal Age
nt for Red Dead Redemption II as a Case Study
Affordance, SegmentationManipVQA: Injecting Robotic Affordance and Physica
lly Grounded Information into Multi-Modal Large Language Models
ArXiv
Action-Model, Agent, LAMLaVagueGitHub
Action-Generation, Generation, PromptingPrompt a Robot to Walk with Large Language ModelsArXiv
APIs, Agent, ToolGorilla: Large Language Model Connected with Massi
ve APIs
ArXiv
AGI, SurveyLevels of AGI: Operationalizing Progress on the Pa
th to AGI
ArXiv
AGI, BrainWhen Brain-inspired AI Meets AGIArXiv
AGI, BrainDivergences between Language Models and Human Brai
ns
ArXiv
AGI, Awesome Repo, SurveyAwesome-LLM-Papers-Toward-AGIGitHub
AGI, AgentOpenAGI: When LLM Meets Domain Experts
3D, Open-source, Perception, Robot3D-LLM: Injecting the 3D World into Large Language
Models
ArXiv2023/07/24
3D, GPT4, VLMGPT-4V(ision) is a Human-Aligned Evaluator for Tex
t-to-3D Generation
ArXiv
ChatEval: Towards Better LLM-based Evaluators thro
ugh Multi-Agent Debate
ArXiv2023/08/14

Collected info

  • 223 stars
  • 23 forks
  • Source updated: 6/26/2026