@torch.library.custom_op
This can happen on Pytorch versions older than 2.4.
* Add --use-flash-attention flag. This is useful on AMD systems, as FA builds are still 10% faster than Pytorch cross-attention.