<aside> 🤚🏻

Welcome to the Week 3 of Reinforcement Learning (RL) Track of CSOC'26 — we’re excited to dive deeper into this journey with you!

</aside>

Introduction

This week moves into model-free learning — methods where the agent learns purely from experience, without any prior knowledge of the environment's dynamics.

We cover two major families:

What you'll learn

Monte Carlo methods

Unlike dynamic programming, Monte Carlo methods do not assume full knowledge of the environment. Instead, the agent learns by sampling complete episodes and averaging observed returns. This makes them practical for real-world problems where a model is unavailable.

Temporal-Difference learning — SARSA & Q-Learning

TD learning updates value estimates after each step, rather than waiting for an episode to finish. Two key algorithms:


Communities

All announcements, resources, and updates will be disseminated through the CSOC Website and the official Discord Community. Please ensure you have registered on the website and joined the Discord server to stay informed.

If you do not receive a timely response on Discord, feel free to reach out directly via WhatsApp: