Between MDPs And semi-MDPs: A Framework For Temporal Abstraction In Reinforcement Learning | Manumap