rocm/aiter ? reverse-engineered prompt
Reverse engineered prompt
Build me a Python package for AMD ROCm that gives fast GPU kernels for AI workloads, especially attention, MoE, GEMM, normalization, quantization, and a few fused communication ops. I want it to work from both Python and C++, with a clean way to call the operators directly and a set of runnable tests for each kernel so I can verify things like MHA, MLA, MoE, GEMM, and RMSNorm on supported AMD GPUs. It should be set up like a real production library, so it can be installed, built, and packaged, and it should support different kernel backends where it makes sense, including Triton, Composable Kernel, and hand tuned HIP or assembly kernels. If you need to, look up the current ROCm and Triton docs online so the setup matches modern versions. Keep it framework agnostic, but make it easy for serving stacks to plug in later.
Are you gonna build this?
make sure you review the code using coderabbit