An exploration of delaying selected visual tokens until they are needed, reducing wasted computation while preserving the model’s useful context.

Full article is a work in progres