cuda_kda_enabled

Function cuda_kda_enabled 

Source
pub fn cuda_kda_enabled() -> bool
Expand description

LinearAttention BoundLayer uses CUDA kda_decode unless K3_CUDA_KDA=0. Projections, AttnRes, and MLP stay on the host either way.