Observation: If the reward is negative at the last timestep, then the longer the environment remains non-negative, the better. The agent would be correctly learning to maximize the reward. See https://www.reddit.com/r/reinforcementlearning/comments/k27lnv/do_strictly_negative_rewards_work_with_discounting/
Observation: If the reward is negative at the last timestep, then the longer the environment remains non-negative, the better. The agent would be correctly learning to maximize the reward. See https://www.reddit.com/r/reinforcementlearning/comments/k27lnv/do_strictly_negative_rewards_work_with_discounting/