首页 | 本学科首页   官方微博 | 高级检索  
     

基于杂合标准的POMDP值迭代求解算法*
引用本文:刘峰.基于杂合标准的POMDP值迭代求解算法*[J].模式识别与人工智能,2016,29(11):961-968.
作者姓名:刘峰
作者单位:南京大学 软件学院 南京 210093
南京大学 计算机软件新技术国家重点实验室 南京 210093
基金项目:计算机软件新技术国家重点实验室面上项目(No.ZZKT2016B07)资助
摘    要:基于点的值迭代方法是求解部分可观测马尔科夫决策过程(POMDP)问题的一类有效算法.目前基于点的值迭代算法大都基于单一启发式标准探索信念点集,从而限制算法效果.基于此种情况,文中提出基于杂合标准探索信念点集的值迭代算法(HHVI),可以同时维持值函数的上界和下界.在扩展探索点集时,选取值函数上下界差值大于阈值的信念点进行扩展,并且在值函数上下界差值大于阈值的后继信念点中选择与已探索点集距离最远的信念点进行探索,保证探索点集尽量有效分布于可达信念空间内.在4个基准问题上的实验表明,HHVI能保证收敛效率,并能收敛到更好的全局最优解.

关 键 词:部分可观测马尔科夫决策过程(POMDP)    杂合启发式值迭代    可达信念空间    探索价值  
收稿时间:2016-05-04

Hybrid Heuristic Value Iteration POMDP Algorithm
LIU Feng.Hybrid Heuristic Value Iteration POMDP Algorithm[J].Pattern Recognition and Artificial Intelligence,2016,29(11):961-968.
Authors:LIU Feng
Affiliation:Software Institute, Nanjing University, Nanjing 210093
State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing 210093
Abstract:Point-based value iteration methods are a kind of algorithms for effectively solving partially observable Markov decision process (POMDP) model. However, the algorithm efficiency is limited by the belief point set explored in most of the algorithms by single heuristic criterion. A hybrid heuristic value iteration algorithm (HHVI) for exploring belief point set is presented in this paper. The upper and lower bounds on the value function are maintained and only the belief points with its value function bounds difference greater than the threshold are selected to expand. Furthermore, the furthest belief point away from the explored point set among the subsequent belief points with the above difference also greater than the threshold is explored. The convergence effect of HHVI is guaranteed by making the explored point set fully distributed in the reachable belief space. Experimental results of four benchmarks show that HHVI can guarantee the convergence efficiency and obtain better global optimal solution.
Keywords:Partially Observable Markov Decision Process(POMDP)  Hybrid Heuristic Value Iteration  Reachable Belief Space  Exploration Value  
点击此处可从《模式识别与人工智能》浏览原始摘要信息
点击此处可从《模式识别与人工智能》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号