tile-ai/tilelang ? reverse-engineered prompt

Reverse engineered prompt

Build me a Python based toolkit for writing high performance GPU, CPU, and accelerator kernels in a simple, Pythonic way.

I want to be able to describe things like matrix multiply, dequantized matrix multiply, and attention style kernels without writing a lot of low level code, and have it compile down to efficient device code through an internal compiler pipeline. It should feel easy to use from Python, but still let advanced users tune layouts, schedules, and performance details when needed.

Please include a few working examples, especially GEMM and an attention example, plus clear install and getting started docs. Make it usable on modern NVIDIA, AMD, and Apple style targets if possible, and keep the code organized so it can grow into a real compiler project. Add tests and a basic benchmark path too. If you need to check current docs or APIs for anything, look them up online first.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab