Cool! On a quick glance, it doesn't seem like the group/layer index is provided to the compression model. That might help a bit with fidelity at pretty low additional cost.
There is a surge of interest in learning control policies end-to-end. Many of Sergey Levine's recent papers are relevant: http://homes.cs.washington.edu/~svlevine/ (there are also some talks linked there).