此參考說明所註明的目錄。請向該機構確認當前課程供給與適用於你入學的條件。
課程說明
強化學習(亦開設為 COSC 4P83) 介紹強化學習的基本概念,包括多臂賭徒問題、馬可夫決策過程、基於模型與無模型的方法(如動態規劃、蒙地卡羅方法與時序差分方法)用以學習價值函數與策略函數。近似解法包括深度強化學習。
原文參考文本
Reinforcement Learning (also offered as COSC 4P83) Introduction to fundamental reinforcement learning concepts including multi-armed bandits, Markov decision processes, model-based and model-free methods (such as dynamic programming, Monte Carlo methods, and temporal-difference methods) for learning value and policy functions. Approximation solutions including deep reinforcement learning.
來源與參考
保留日期與來源以協助你核實資料。為便於閱讀提供譯文;官方來源為條件與要求的參照。
來源參考 : https://brocku.ca/webcal/2024/graduate/cosc.html