The first state transition probability and prescribed first reward are used to obtain a state value function V^π(s) on the basis of dynamic programming in a Markov decision process. 第一状態遷移確率及び所定の第一報酬を用いて、マルコフ決定過程における動的計画法に基づき、状態価値関数V^π(s)を求める。 - 特許庁
The second state transition probability and the second reward are used to obtain action value function Q^π(s, a) and the state value function V^π(s) on the basis of the dynamic programming in the Markov decision process. 第二状態遷移確率及び第二報酬を用いて、マルコフ決定過程における動的計画法に基づき、行動価値関数Q^π(s,a)及び状態価値関数V^π(s)を求める。 - 特許庁