罰を回避する合理的政策の学習

Transactions of the Japanese Society for Artificial Intelligence 16 (2):185-192 (2001)
  Copy   BIBTEX

Abstract

Reinforcement learning is a kind of machine learning. It aims to adapt an agent to a given environment with a clue to rewards. In general, the purpose of reinforcement learning system is to acquire an optimum policy that can maximize expected reward per an action. However, it is not always important for any environment. Especially, if we apply reinforcement learning system to engineering, environments, we expect the agent to avoid all penalties. In Markov Decision Processes, a pair of a sensory input and an action is called rule. We call a rule penalty if and only if it has a penalty or it can transit to a penalty state where it does not contribute to get any reward. After suppressing all penalty rules, we aim to make a rational policy whose expected reward per an action is larger than zero. In this paper, we propose a suppressing penalty algorithm that can suppress any penalty and get a reward constantly. By applying the algorithm to the tick-tack-toe, its effectiveness is shown.

Other Versions

No versions found

Links

PhilArchive

External links

Setup an account with your affiliations in order to access resources via your University's proxy server

Through your library

Similar books and articles

罰回避政策形成アルゴリズムの改良とオセロゲームへの応用.宮崎 和光, 坪井 創吾 & 小林 重信 - 2002 - Transactions of the Japanese Society for Artificial Intelligence 17 (5):548-556.
合理的政策形成アルゴリズムの連続値入力への拡張.宮崎 和光, 木村 元 & 小林 重信 - 2007 - Transactions of the Japanese Society for Artificial Intelligence 22 (3):332-341.
Profit Sharing 法における強化関数に関する一考察.Tatsumi Shoji Uemura Wataru - 2004 - Transactions of the Japanese Society for Artificial Intelligence 19:197-203.
環境状況に応じて自己の報酬を操作する学習エージェントの構築.森山 甲一 & 沼尾 正行 - 2002 - Transactions of the Japanese Society for Artificial Intelligence 17 (6):676-683.
経験に固執しない Profit Sharing 法.Ueno Atsushi Uemura Wataru - 2006 - Transactions of the Japanese Society for Artificial Intelligence 21:81-93.
強化学習を用いた自律移動型ロボットの行動計画法の提案.五十嵐 治一 - 2001 - Transactions of the Japanese Society for Artificial Intelligence 16 (6):501-509.
Ga により探索空間の動的生成を行う Q 学習.Matsuno Fumitoshi Ito Kazuyuki - 2001 - Transactions of the Japanese Society for Artificial Intelligence 16:510-520.

Analytics

Added to PP
2014-03-25

Downloads
44 (#1,294,647)

6 months
9 (#1,255,287)

Historical graph of downloads
How can I increase my downloads?

Citations of this work

No citations found.

Add more citations

References found in this work

No references found.

Add more references