This means that we should have - a single optimizer for the generator and discriminator - a single, initial embedding layer in the generator - a single, final layer in the top-level discriminator before predictions are made
This means that we should have