akshay919191/flashattn-V1 ? reverse-engineered prompt

Reverse engineered prompt

Build me a from scratch CUDA FlashAttention project in Python that I can install and import like a normal package.

I want a custom flash_attn(q, k, v, causal=False) function that works with CUDA tensors, supports FP16 inputs, autograd, causal and non causal attention, and also handles cross attention when the query and key lengths are different. It should return the attention output and have a backward pass that gives correct gradients for Q, K, and V. Please include a small test suite that compares the results to a PyTorch reference on a few shapes, plus simple benchmark scripts so I can measure forward and backward speed. If the compiled extension is not already built, it should build on first import or at install time and work from the repo root too. Keep it focused on learning and correctness, and look up current CUDA and PyTorch extension docs online if you need to.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab