Skip to content

Reward shaping #5

Description

@Tuxliri

Observation: If the reward is negative at the last timestep, then the longer the environment remains non-negative, the better. The agent would be correctly learning to maximize the reward. See https://www.reddit.com/r/reinforcementlearning/comments/k27lnv/do_strictly_negative_rewards_work_with_discounting/

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions