Module moe_cuda

Module moe_cuda 

Source
Expand description

Host launch for K3 packed LatentMoE experts (moe_w4a16 E8M0 ptrtable).

Required handle: PTRTABLE_E8M0 from DSV4 extra_cu. Lookup-fail bails; packed tensors must not silently dequant on the host F32 twin path.

Structs§

K3MoeGemmKernels

Constants§

E8M0_ENTRY
MODULE
PTX module from kimi-k3 extra_cu of DSV4 moe_w4a16_grouped_gemm.cu.
PTRTABLE_E8M0

Functions§

launch_k3_latent_moe_experts
Packed w1/w2/w3 SiTU mix. Same contract as atlas_core::kimi_k3::mix_routed_experts.
launch_k3_moe_e8m0_ptrtable
One grouped E8M0 GEMM: C[M, N] = A[M, K] @ B_packed[N, K/2] via ptrtable.