Expand description
Host launch for K3 CUDA KDA decode (kda_decode PTX module).
Two kernels, one token: conv-4 + SiLU, then L2 q/k + delta-rule.
CPU oracle: atlas_core::kimi_k3::kda_decode_token. BoundLayer serve
LinearAttention default is this launch (K3_CUDA_KDA=0 keeps CPU).
Structs§
- K3Kda
Decode Kernels - KdaDevice
State - Device-resident conv + recurrent. Seed once; decode does not D2H them.
Constants§
- CONV_
ENTRY - MODULE
- PTX module stem =
kernels/gb10/kimi-k3/bf16/kda_decode.cu. - RECURRENT_
ENTRY
Functions§
- launch_
k3_ kda_ decode_ token - One-token CUDA KDA decode. Host-state round-trip (oracle / CPU escape).
- launch_
k3_ kda_ decode_ token_ on_ device - One-token CUDA KDA on device-resident conv/recurrent. Does not D2H state.