A study of predicting which experts will be activated so model weights can be staged just ahead of demand instead of remaining resident in VRAM.

Full article is a work in progres