During the MAE Decoder process, why are the mask tokens set to be shared? Can different mask tokens be set as different learnable vectors?
During the MAE Decoder process, why are the mask tokens set to be shared? Can different mask tokens be set as different learnable vectors?