pub fn cuda_mla_enabled() -> bool
FullAttention BoundLayer uses CUDA mla_decode unless K3_CUDA_MLA=0. Same polarity as K3_CUDA_KDA. Projections stay on the host either way.
mla_decode
K3_CUDA_MLA=0
K3_CUDA_KDA