이 참고자료는 표시된 카탈로그를 설명합니다. 최신 제공 및 적용 조건은 입학 전에 기관에 확인하세요.
설명
강화학습 다중무장 도박(Multi-armed bandits), 마르코프 결정 과정, 가치 함수와 정책 함수를 학습하기 위한 모델 기반 및 모델 자유 방법(동적 계획법, 몬테카를로 방법, 시차학습 방법 등). 딥 강화학습을 포함한 근사 해법. 강의, 주당 3시간. 제한: COSC 단독 또는 복합, GAME, NEUR 및 Data Science 프로그램에 개설. 선수과목: COSC 3P71 (최소 60퍼센트). 참고: 이 과목은 여러 교육 방식으로 개설될 수 있음. 교육 방식은 해당 학기 학사 시간표에 기재됨.
선수 과목
- 선수과목: COSC 3P71(최소 60퍼센트).
조건 및 세부사항
- 제한: COSC 단독 또는 복합, GAME, NEUR 및 Data Science 프로그램에 개설.
- 선수과목: COSC 3P71(최소 60퍼센트).
- 참고: 이 강좌는 여러 전달 방식으로 제공될 수 있습니다. 전달 방식은 해당 학기의 학사 시간표에 기재됩니다.
원문 참조 텍스트
Reinforcement Learning Multi-armed bandits, Markov decision processes, model-based and model-free methods (such as dynamic programming, Monte Carlo methods, and temporal-difference methods) for learning value and policy functions. Approximation solutions including deep reinforcement learning. Lectures, 3 hours per week. Restriction: open to COSC single or combined, GAME, NEUR , and Data Science programs. Prerequisite(s): COSC 3P71 (minimum 60 percent). Note: this course may be offered in multiple modes of delivery. The method of delivery will be listed on the academic timetable, in the applicable term.
- Prerequisite(s): COSC 3P71 (minimum 60 percent).
- Restriction: open to COSC single or combined, GAME, NEUR , and Data Science programs.
- Note: this course may be offered in multiple modes of delivery. The method of delivery will be listed on the academic timetable, in the applicable term.
출처 및 참고문헌
날짜와 출처는 정보를 확인하는 데 도움이 되도록 보관됩니다. 번역은 읽기 편의를 위해 제공되며 조건과 요건은 공식 출처가 기준입니다.
출처 참조 : https://brocku.ca/webcal/2024/undergrad/cosc.html