<aside> 🤚🏻
Welcome to the Week 3 of Reinforcement Learning (RL) Track of CSOC'26 — we’re excited to dive deeper into this journey with you!
</aside>
This week moves into model-free learning — methods where the agent learns purely from experience, without any prior knowledge of the environment's dynamics.
We cover two major families:
Unlike dynamic programming, Monte Carlo methods do not assume full knowledge of the environment. Instead, the agent learns by sampling complete episodes and averaging observed returns. This makes them practical for real-world problems where a model is unavailable.
TD learning updates value estimates after each step, rather than waiting for an episode to finish. Two key algorithms:
All announcements, resources, and updates will be disseminated through the CSOC Website and the official Discord Community. Please ensure you have registered on the website and joined the Discord server to stay informed.
If you do not receive a timely response on Discord, feel free to reach out directly via WhatsApp: