Bias92/flashattn-cuda ? reverse-engineered prompt

Reverse engineered prompt

Build me a small CUDA project that implements fast attention for inference on an RTX 4060 Ti.

I want it to support the forward pass for both dense and causal attention, and also handle grouped query attention and different query and key value lengths. It should be written in CUDA C++ with inline PTX, and use the usual speed tricks like Tensor Cores, online softmax, register accumulation, and async tile loading where it makes sense.

Please make it easy to build and run, with a clear Python install path, a few benchmark scripts, and tests that compare the custom kernel against PyTorch attention and cuDNN so I can check correctness and speed. If you need to, look up current CUDA and PyTorch docs online while wiring it up. Also include simple docs that explain how to run the benchmarks and what kind of performance to expect on consumer GPUs.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab