Profit Sharing 法における強化関数に関する一考察

Abstract

In this paper, we consider profit sharing that is one of the reinforcement learning methods. An agent learns a candidate solution of a problem from the reward that is received from the environment if and only if it reaches the destination state. A function that distributes the received reward to each action of the candidate solution is called the reinforcement function. On this learning system, the agent can reinforce the set of selected actions when it gets the reward. And the agent should not reinforce the detour actions. First, we will propose a new constraint equation about reinforcement functions to distribute the reinforcement values on the non-detour actions. If we use the reinforcement function to satisfy the constraint equation, the agent can select the non-detour actions directing to the destination state. Next, it is shown that the reinforcement function can be constant after learning process to suppress the selection of detour actions. Lastly, in computer simulations for maze problems, we show that the learning performance of agents does not depend on the size of environment.

Other Versions

No versions found

Links

PhilArchive

External links

Setup an account with your affiliations in order to access resources via your University's proxy server

Through your library

Similar books and articles

経験に固執しない Profit Sharing 法.Ueno Atsushi Uemura Wataru - 2006 - Transactions of the Japanese Society for Artificial Intelligence 21:81-93.
罰を回避する合理的政策の学習.宮崎 和光, 坪井 創吾 & 小林 重信 - 2001 - Transactions of the Japanese Society for Artificial Intelligence 16 (2):185-192.
強化学習エージェントへの階層化意志決定法の導入―追跡問題を例に―.片山 謙吾, 輿石 尚宏 & 成久 洋之 - 2004 - Transactions of the Japanese Society for Artificial Intelligence 19 (4):279-291.
尤度情報に基づく温度分布を用いた強化学習法.小堀 訓成, 鈴木 健嗣, ハルトノ, ピトヨ. & 橋本 周司 - 2005 - Transactions of the Japanese Society for Artificial Intelligence 20 (4):297-305.
不完全知覚判定法を導入した Profit Sharing.Masuda Shiro Saito Ken - 2004 - Transactions of the Japanese Society for Artificial Intelligence 19:379-388.
環境状況に応じて自己の報酬を操作する学習エージェントの構築.森山 甲一 & 沼尾 正行 - 2002 - Transactions of the Japanese Society for Artificial Intelligence 17 (6):676-683.
強化学習を用いた自律移動型ロボットの行動計画法の提案.五十嵐 治一 - 2001 - Transactions of the Japanese Society for Artificial Intelligence 16 (6):501-509.
合理的政策形成アルゴリズムの連続値入力への拡張.宮崎 和光, 木村 元 & 小林 重信 - 2007 - Transactions of the Japanese Society for Artificial Intelligence 22 (3):332-341.

Analytics

Added to PP
2014-03-21

Downloads
54 (#1,105,913)

6 months
16 (#751,944)

Historical graph of downloads
How can I increase my downloads?

Citations of this work

No citations found.

Add more citations

References found in this work

No references found.

Add more references