Deep Reinforcement Learning with pytorch & visdom

Jingwei Zhang

Last update: Jan 4, 2023

Related tags

Deep Learning reinforcement-learning deep-learning deep-reinforcement-learning pytorch dqn a3c actor-critic pytorch-a3c acer trpo visdom

Overview

Deep Reinforcement Learning with

pytorch & visdom

Sample testings of trained agents (DQN on Breakout, A3C on Pong, DoubleDQN on CartPole, continuous A3C on InvertedPendulum(MuJoCo)):

Sample on-line plotting while training an A3C agent on Pong (with 16 learner processes):
Sample loggings while training a DQN agent on CartPole (we use WARNING as the logging level currently to get rid of the INFO printouts from visdom):

[WARNING ] (MainProcess) <===================================>
[WARNING ] (MainProcess) bash$: python -m visdom.server
[WARNING ] (MainProcess) http://localhost:8097/env/daim_17040900
[WARNING ] (MainProcess) <===================================> DQN
[WARNING ] (MainProcess) <-----------------------------------> Env
[WARNING ] (MainProcess) Creating {gym | CartPole-v0} w/ Seed: 123
[INFO    ] (MainProcess) Making new env: CartPole-v0
[WARNING ] (MainProcess) Action Space: [0, 1]
[WARNING ] (MainProcess) State  Space: 4
[WARNING ] (MainProcess) <-----------------------------------> Model
[WARNING ] (MainProcess) MlpModel (
  (fc1): Linear (4 -> 16)
  (rl1): ReLU ()
  (fc2): Linear (16 -> 16)
  (rl2): ReLU ()
  (fc3): Linear (16 -> 16)
  (rl3): ReLU ()
  (fc4): Linear (16 -> 2)
)
[WARNING ] (MainProcess) No Pretrained Model. Will Train From Scratch.
[WARNING ] (MainProcess) <===================================> Training ...
[WARNING ] (MainProcess) Validation Data @ Step: 501
[WARNING ] (MainProcess) Start  Training @ Step: 501
[WARNING ] (MainProcess) Reporting       @ Step: 2500 | Elapsed Time: 5.32397913933
[WARNING ] (MainProcess) Training Stats:   epsilon:          0.972
[WARNING ] (MainProcess) Training Stats:   total_reward:     2500.0
[WARNING ] (MainProcess) Training Stats:   avg_reward:       21.7391304348
[WARNING ] (MainProcess) Training Stats:   nepisodes:        115
[WARNING ] (MainProcess) Training Stats:   nepisodes_solved: 114
[WARNING ] (MainProcess) Training Stats:   repisodes_solved: 0.991304347826
[WARNING ] (MainProcess) Evaluating      @ Step: 2500
[WARNING ] (MainProcess) Iteration: 2500; v_avg: 1.73136949539
[WARNING ] (MainProcess) Iteration: 2500; tderr_avg: 0.0964358523488
[WARNING ] (MainProcess) Iteration: 2500; steps_avg: 9.34579439252
[WARNING ] (MainProcess) Iteration: 2500; steps_std: 0.798395631184
[WARNING ] (MainProcess) Iteration: 2500; reward_avg: 9.34579439252
[WARNING ] (MainProcess) Iteration: 2500; reward_std: 0.798395631184
[WARNING ] (MainProcess) Iteration: 2500; nepisodes: 107
[WARNING ] (MainProcess) Iteration: 2500; nepisodes_solved: 106
[WARNING ] (MainProcess) Iteration: 2500; repisodes_solved: 0.990654205607
[WARNING ] (MainProcess) Saving Model    @ Step: 2500: /home/zhang/ws/17_ws/pytorch-rl/models/daim_17040900.pth ...
[WARNING ] (MainProcess) Saved  Model    @ Step: 2500: /home/zhang/ws/17_ws/pytorch-rl/models/daim_17040900.pth.
[WARNING ] (MainProcess) Resume Training @ Step: 2500
...

What is included?

This repo currently contains the following agents:

Deep Q Learning (DQN) [1], [2]
Double DQN [3]
Dueling network DQN (Dueling DQN) [4]
Asynchronous Advantage Actor-Critic (A3C) (w/ both discrete/continuous action space support) [5], [6]
Sample Efficient Actor-Critic with Experience Replay (ACER) (currently w/ discrete action space support (Truncated Importance Sampling, 1st Order TRPO)) [7], [8]

Work in progress:

Testing ACER

Future Plans:

Deep Deterministic Policy Gradient (DDPG) [9], [10]
Continuous DQN (CDQN or NAF) [11]

Code structure & Naming conventions:

NOTE: we follow the exact code structure as pytorch-dnc so as to make the code easily transplantable.

./utils/factory.py

We suggest the users refer to ./utils/factory.py, where we list all the integrated Env, Model, Memory, Agent into Dict's. All of those four core classes are implemented in ./core/. The factory pattern in ./utils/factory.py makes the code super clean, as no matter what type of Agent you want to train, or which type of Env you want to train on, all you need to do is to simply modify some parameters in ./utils/options.py, then the ./main.py will do it all (NOTE: this ./main.py file never needs to be modified).

namings

To make the code more clean and readable, we name the variables using the following pattern (mainly in inherited Agent's):

*_vb: torch.autograd.Variable's or a list of such objects

*_ts: torch.Tensor's or a list of such objects

otherwise: normal python datatypes

Dependencies

How to run:

You only need to modify some parameters in ./utils/options.py to train a new configuration.

Configure your training in ./utils/options.py:

line 14: add an entry into CONFIGS to define your training (agent_type, env_type, game, model_type, memory_type)

line 33: choose the entry you just added

line 29-30: fill in your machine/cluster ID (MACHINE) and timestamp (TIMESTAMP) to define your training signature (MACHINE_TIMESTAMP), the corresponding model file and the log file of this training will be saved under this signature (./models/MACHINE_TIMESTAMP.pth & ./logs/MACHINE_TIMESTAMP.log respectively). Also the visdom visualization will be displayed under this signature (first activate the visdom server by type in bash: python -m visdom.server &, then open this address in your browser: http://localhost:8097/env/MACHINE_TIMESTAMP)

line 32: to train a model, set mode=1 (training visualization will be under http://localhost:8097/env/MACHINE_TIMESTAMP); to test the model of this current training, all you need to do is to set mode=2 (testing visualization will be under http://localhost:8097/env/MACHINE_TIMESTAMP_test).

Run:

python main.py

Bonus Scripts :)

We also provide 2 additional scripts for quickly evaluating your results after training. (Dependecies: lmj-plot)

plot.sh (e.g., plot from log file: logs/machine1_17080801.log)

./plot.sh machine1 17080801

the generated figures will be saved into figs/machine1_17080801/

plot_compare.sh (e.g., compare log files: logs/machine1_17080801.log,logs/machine2_17080802.log)

./plot.sh 00 machine1 17080801 machine2 17080802

the generated figures will be saved into figs/compare_00/

the color coding will be in the order of: red green blue magenta yellow cyan

Repos we referred to during the development of this repo:

matthiasplappert/keras-rl
transedward/pytorch-dqn
ikostrikov/pytorch-a3c
onlytailei/A3C-PyTorch
Kaixhin/ACER
And a private implementation of A3C from @stokasto

Citation

If you find this library useful and would like to cite it, the following would be appropriate:

@misc{pytorch-rl,
  author = {Zhang, Jingwei and Tai, Lei},
  title = {jingweiz/pytorch-rl},
  url = {https://github.com/jingweiz/pytorch-rl},
  year = {2017}
}

Comments

Added opensim-rl environment and a sample configuration and options for a continuous DQN agent to learn in that environment
opensim-rl Is an environment introduced by the NIPS 2017 Learning to run challenge. In this environment, an agent is tasked with learning how to run while avoiding obstacles on the ground. The environment provides a human musculoskeletal model and a physics-based simulation environment which are pretty good. This environment will be useful for training agents that can handle much more complex control tasks even after the NIPS challenge ends. Can be seen as a good alternative or as a complementary environment to Mujoco based environments.

Contributions:

[x] Added the opensim-rl environment as a stand-alone environment into the existing framework

[x] Added sample factory configurations for the opensim-rl environment

[x] Added sample options for training a DQN agent in the opensim-rl environment

[x] Modified the dqn agent to produce a list of actions instead of an action index so that it can be extended to be used for action dimensions >1

[x] Updated README to point to the opensim-rl to set up the dependencies
opened by praveen-palanisamy 7
Can this code reproduce the performance of the baseline?

Can this code reproduce the performance of the baseline without any change on the hyperparameter? I try to run the code with the dqn mode on other Atari games such as "Qbert", "WizardOfWor", but I can not get the result reported by other papers.

opened by xfdywy 3
In python3, change xrange to range to remove python NameError
I run the main.py in python3 environment. It is better to make this clear usage. :)

some error message as follows:

policy_log_vb = [torch.log(policy_vb[i]) for i in xrange(rollout_steps)] NameError: name 'xrange' is not defined Process Process-14:
opened by tigerneil 1
* add agent detection in gym env

Otherwise, config 1 will give an error. Cause there is an enable_continuous detection in step(), GymEnv object should always have an enable_continuous attribute. But in options.py, this is an attribute of a3c agent.

opened by onlytailei 0
pytorch-rl-master

Hello! Excuse me,I didn't find this 'map_server_medium_rooms_simpler.launch' file in your pytorch-rl package? Where can I find it, please? Another question is the video on the website (https://goo.gl/pWbpcF) in your paper. Does the pytorch-rl package have its program? Thank you very much！

opened by H0803 0
Added opensim-rl environment, extended dqn agent for multi-dimensional action space, and a sample configuration and options to config an agent to learn in opensim-rl
opensim-rl Is an environment introduced by the NIPS 2017 Learning to run challenge. In this environment, an agent is tasked with learning how to run while avoiding obstacles on the ground. The environment provides a human musculoskeletal model and a physics-based simulation environment which are pretty good. This environment will be useful for training agents that can handle much more complex control tasks even after the NIPS challenge ends. Can be seen as a good alternative or as a complementary environment to Mujoco based environments.

Contributions:

[x] Added the opensim-rl environment as a stand-alone environment into the existing framework

[x] Updated README to point to the opensim-rl to set up the dependencies

[x] Modified the dqn agent to produce a list of actions instead of an action index so that it can be extended to be used for multiple action dimensions >1

[x] Added sample factory configurations for the opensim-rl environment

[x] Added sample options for training an agent (dqn) in the opensim-rl environment
opened by praveen-palanisamy 0

Deep Reinforcement Learning with pytorch & visdom

Related tags

Overview

Deep Reinforcement Learning with

pytorch & visdom

What is included?

Code structure & Naming conventions:

Dependencies

How to run:

Bonus Scripts :)

Repos we referred to during the development of this repo:

Citation

Comments

Added opensim-rl environment and a sample configuration and options for a continuous DQN agent to learn in that environment

Can this code reproduce the performance of the baseline?

In python3, change xrange to range to remove python NameError

* add agent detection in gym env

pytorch-rl-master

Added opensim-rl environment, extended dqn agent for multi-dimensional action space, and a sample configuration and options to config an agent to learn in opensim-rl

Owner

Jingwei Zhang

PyTorch implementation of Value Iteration Networks (VIN): Clean, Simple and Modular. Visualization in Visdom.

Conservative Q Learning for Offline Reinforcement Reinforcement Learning in JAX

Reinforcement-learning - Repository of the class assignment questions for the course on reinforcement learning

Learning to Communicate with Deep Multi-Agent Reinforcement Learning in PyTorch

PyTorch implementation of Advantage Actor Critic (A2C), Proximal Policy Optimization (PPO), Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation (ACKTR) and Generative Adversarial Imitation Learning (GAIL).

PyTorch implementation of Advantage Actor Critic (A2C), Proximal Policy Optimization (PPO), Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation (ACKTR) and Generative Adversarial Imitation Learning (GAIL).

A resource for learning about deep learning techniques from regression to LSTM and Reinforcement Learning using financial data and the fitness functions of algorithmic trading

PyTorch implementations of deep reinforcement learning algorithms and environments

Deep Learning and Reinforcement Learning Library for Scientists and Engineers 🔥

Deep Learning and Reinforcement Learning Library for Scientists and Engineers 🔥

deep-table implements various state-of-the-art deep learning and self-supervised learning algorithms for tabular data using PyTorch.

DRLib：A concise deep reinforcement learning library, integrating HER and PER for almost off policy RL algos.

MazeRL is an application oriented Deep Reinforcement Learning (RL) framework

PGPortfolio: Policy Gradient Portfolio, the source code of "A Deep Reinforcement Learning Framework for the Financial Portfolio Management Problem"(https://arxiv.org/pdf/1706.10059.pdf).

Deep Reinforcement Learning based Trading Agent for Bitcoin

A list of papers regarding generalization in (deep) reinforcement learning

AutoPentest-DRL: Automated Penetration Testing Using Deep Reinforcement Learning

[ICML 2021] DouZero: Mastering DouDizhu with Self-Play Deep Reinforcement Learning | 斗地主AI

Thermal Control of Laser Powder Bed Fusion using Deep Reinforcement Learning