some notes of my RL learning, some exercises
- Non-stationary 10-armed bandit problem
- Optimistic initial values
- QLearning trial
- UCB action selection
- Compare to others
- Trying some trainable parameters in AS and UCB methods
- Calc the cumulative rewards
- Policy Gradient
- Pole Balancing
- DQN
- A2C
- PPO
- Comparison
- Lunar Landing(PPO)
- Grid World
- Q-learning
- Jack's Car Rental
- Monte Carlo Method
- Estimate PI