pub fn cuda_kda_enabled() -> bool
LinearAttention BoundLayer uses CUDA kda_decode unless K3_CUDA_KDA=0. Projections, AttnRes, and MLP stay on the host either way.
kda_decode
K3_CUDA_KDA=0