メインナビゲーションにスキップ 検索にスキップ メインコンテンツにスキップ

Evaluating Large Language Models with RAG Capability: A Perspective from Robot Behavior Planning and Execution

  • Fujitsu Research of America

研究成果: 書籍の章/レポート/Proceedings会議への寄与査読

4 被引用数 (Scopus)

抄録

After the significant performance of Large Language Models (LLMs) was revealed, their capabilities were rapidly expanded with techniques such as Retrieval Augmented Generation (RAG). Given their broad applicability and fast development, it's crucial to consider their impact on social systems. On the other hand, assessing these advanced LLMs poses challenges due to their extensive capabilities and the complex nature of social systems. In this study, we pay attention to the similarity between LLMs in social systems and humanoid robots in open environments. We enumerate the essential components required for controlling humanoids in problem solving which help us explore the core capabilities of LLMs and assess the effects of any deficiencies within these components. This approach is justified because the effectiveness of humanoid systems has been thoroughly proven and acknowledged. To identify needed components for humanoids in problem-solving tasks, we create an extensive component framework for planning and controlling humanoid robots in an open environment. Then assess the impacts and risks of LLMs for each component, referencing the latest benchmarks to evaluate their current strengths and weaknesses. Following the assessment guided by our framework, we identified certain capabilities that LLMs lack and concerns in social systems.

本文言語英語
ホスト出版物のタイトルAAAI Spring Symposium - Technical Report
編集者Ron Petrick, Christopher Geib
出版社Association for the Advancement of Artificial Intelligence
ページ452-456
ページ数5
1
ISBN(電子版)9781577358886
DOI
出版ステータス出版済み - 21 5月 2024
イベント2024 AAAI Spring Symposium Series, SSS 2024 - Stanford, 米国
継続期間: 25 3月 202427 3月 2024

出版物シリーズ

名前AAAI Spring Symposium - Technical Report
番号1
3

会議

会議2024 AAAI Spring Symposium Series, SSS 2024
国/地域米国
CityStanford
Period25/03/2427/03/24

フィンガープリント

「Evaluating Large Language Models with RAG Capability: A Perspective from Robot Behavior Planning and Execution」の研究トピックを掘り下げます。これらがまとまってユニークなフィンガープリントを構成します。

引用スタイル