ebuehrle
|
59681cb16f
|
Merge pull request #3 from sisl/refactor-sampling
Refactor sampling
|
2021-09-15 09:08:13 +02:00 |
|
ebuehrle
|
4928458e08
|
Fix discriminator reward
Had wrong sign.
|
2021-09-15 07:15:43 +02:00 |
|
ebuehrle
|
40da84393c
|
Update imitation version
|
2021-09-14 18:04:53 +02:00 |
|
ebuehrle
|
183657dc36
|
Fix buffer bug, add test
Buffer was not being cleared between plan rollouts
|
2021-09-14 17:55:02 +02:00 |
|
ebuehrle
|
d2932374d9
|
Refactor sampling
|
2021-09-14 13:42:25 +02:00 |
|
Johannes Fischer
|
763a7bb0d3
|
Merge branch 'main' of github.com:sisl/InteractionImitation
|
2021-09-14 11:01:24 +02:00 |
|
ebuehrle
|
deaef45943
|
Add test for discriminator
|
2021-09-14 07:19:58 +02:00 |
|
Johannes Fischer
|
2280597db6
|
Add GAIL with random agent data
|
2021-09-13 19:42:41 +02:00 |
|
ebuehrle
|
7eae74a7d8
|
Switch back to additive reward
|
2021-09-13 17:55:50 +02:00 |
|
johannes-fischer
|
b0b358544f
|
Merge pull request #2 from sisl/fischer/fusion_sample_methods
Merge different methods to sample the policy and collect transitions
|
2021-09-13 16:51:16 +02:00 |
|
Johannes Fischer
|
9c7e6cef3a
|
Rename action to option
|
2021-09-13 16:43:45 +02:00 |
|
Johannes Fischer
|
9b8ceed9c9
|
Render only one episode
|
2021-09-13 16:31:24 +02:00 |
|
Johannes Fischer
|
826c0fa219
|
Merge different methods to sample the policy and collect transitions
|
2021-09-13 15:50:49 +02:00 |
|
ebuehrle
|
f1ece358d7
|
Speed up collision check, assume 0 is fallback option
Due to the conservative approximation of the collision check,
no option might be feasible, thus the necessity of a guaranteed fallback.
|
2021-09-13 09:35:49 +02:00 |
|
ebuehrle
|
87ff3dbb93
|
Add cuda support, normalize actions
|
2021-09-11 21:32:30 +02:00 |
|
ebuehrle
|
e7b0aea427
|
run options gail
|
2021-09-11 21:32:25 +02:00 |
|
Johannes Fischer
|
8ad7457159
|
Merge branch 'main' of github.com:sisl/InteractionImitation
|
2021-09-10 11:38:31 +02:00 |
|
Johannes Fischer
|
544ea4d15a
|
Allow variable horizon trajectories
|
2021-09-10 11:37:34 +02:00 |
|
Johannes Fischer
|
9a107b165a
|
Allow variable horizon trajectories
|
2021-09-10 11:34:36 +02:00 |
|
Johannes Fischer
|
b2b2abafa2
|
Fix typo
|
2021-09-10 11:34:13 +02:00 |
|
ebuehrle
|
5ff4b42c0e
|
Fix imports
|
2021-09-08 20:45:45 +02:00 |
|
ebuehrle
|
3a6139286d
|
Update .gitignore
|
2021-09-08 20:31:45 +02:00 |
|
ebuehrle
|
50916aec05
|
Vanilla GAIL on rasterized observation
|
2021-09-08 20:30:11 +02:00 |
|
ebuehrle
|
802d4a4301
|
Try lower image resolution
|
2021-09-08 20:26:54 +02:00 |
|
ebuehrle
|
f94ec9a4dc
|
Add test for discriminator
|
2021-09-08 20:23:39 +02:00 |
|
Arec
|
de5877aaad
|
filling in available_actions, generate_plan, and feasible helpers
|
2021-09-08 05:14:15 -07:00 |
|
ebuehrle
|
1fd0a71646
|
Draft Options GAIL
|
2021-09-08 11:21:25 +02:00 |
|
ebuehrle
|
d89e491b92
|
Remove debug print statement
|
2021-09-07 19:20:28 +02:00 |
|
ebuehrle
|
e9f09cacb7
|
Add option to render expert rollout
|
2021-09-07 19:19:53 +02:00 |
|
ebuehrle
|
88e0b99d7e
|
Add script for data generation
|
2021-09-07 19:18:59 +02:00 |
|
ebuehrle
|
a70907c0fd
|
Copy over experiments
|
2021-09-01 16:07:15 +02:00 |
|
Johannes Fischer
|
317d329765
|
Add shell script for value dice training
|
2021-08-06 18:53:32 +02:00 |
|
Johannes Fischer
|
88b4466e57
|
MInor change in value dice loss, activate print statements, only do EITHER value OR policy update for each batch
|
2021-08-06 18:52:17 +02:00 |
|
Johannes Fischer
|
025c71767f
|
Change final value network activation to identity
|
2021-08-06 18:48:53 +02:00 |
|
Johannes Fischer
|
66bfba3986
|
Minor formatting
|
2021-08-05 18:37:54 +02:00 |
|
Johannes Fischer
|
6afb112277
|
Bugfix in value dice
FIRST backward() has to be called on both, policy and value, before step() is called for either of them
|
2021-08-05 18:35:14 +02:00 |
|
Johannes Fischer
|
bf4c19a4d0
|
Add value dice ray config
|
2021-08-05 18:33:34 +02:00 |
|
Johannes Fischer
|
8bce5d15f6
|
Restore train_epochs to 200 instead of 8
|
2021-08-05 18:32:46 +02:00 |
|
Johannes Fischer
|
40c55478f3
|
Bugfixes in valuedice
|
2021-08-04 20:53:03 +02:00 |
|
Johannes Fischer
|
2224e2cd14
|
Merge branch 'main' of github.com:sisl/InteractionImitation
|
2021-08-04 19:44:55 +02:00 |
|
Johannes Fischer
|
5c40de66fa
|
Implement ValueDICE and some restructuring
|
2021-08-04 19:34:49 +02:00 |
|
Arec
|
f9729b0a9d
|
making expert data save s, a, sp. making dataloader also load batches thisway. renaming state to ego_state. converting path_x and path_y to single path variable. making number of samples for ray an argument. adjusting metrics, policy, and other functions to be able to handle this
|
2021-08-04 09:45:36 -07:00 |
|
Etienne Buehrle
|
7ae01f73a2
|
AdVIL tests
|
2021-08-04 16:41:18 +02:00 |
|
Johannes Fischer
|
cba42c6e4d
|
Set default divergence to histogram based
|
2021-08-03 17:18:30 +02:00 |
|
Johannes Fischer
|
5919a4e439
|
Use JS divergence in metrics
|
2021-08-03 17:04:47 +02:00 |
|
Johannes Fischer
|
1ee46214a7
|
Implement jenson shannon divergence
|
2021-08-03 17:03:37 +02:00 |
|
Johannes Fischer
|
1916a8fe69
|
Implement metrics and write to tensorboard summary at test time
|
2021-08-03 15:24:29 +02:00 |
|
Arec
|
98294e0c95
|
Merge branch 'main' of https://github.com/sisl/InteractionImitation into main
|
2021-08-03 03:28:23 -07:00 |
|
Arec
|
943e8cda26
|
adding output directory to parse arguments
|
2021-08-03 03:28:11 -07:00 |
|
Arec
|
6e524cf4b5
|
adding options for regularization and relative state masking via interaction graphs during data processing and experiment running. found 0.002 regularization on actions gives up to 3m of deviation with no collisions. added shell script to run ray experiments overnight
|
2021-08-02 14:38:06 -07:00 |
|