cuda_mla_enabled

Function cuda_mla_enabled 

Source
pub fn cuda_mla_enabled() -> bool
Expand description

FullAttention BoundLayer uses CUDA mla_decode unless K3_CUDA_MLA=0. Same polarity as K3_CUDA_KDA. Projections stay on the host either way.