lucidrains/vit-pytorch ? reverse-engineered prompt

Reverse engineered prompt

Build me a PyTorch library for image classification with Vision Transformers, like a clean, reusable package I can import in my own projects.

I want the main ViT model to take an image, split it into patches, run them through transformer blocks, and return class predictions. Please also include a few useful variants from the README, like a simpler ViT version, a distillation setup, and support for variable size images with the NaViT style approach. It would be great if the code is easy to read, with small example scripts showing how to create a model and run a random image through it.

Please make sure the package can be installed normally, includes basic tests, and has a couple of working usage examples. If you need to check the current PyTorch docs for anything newer, go ahead and look it up online.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab