- Sources: primary, discussion
- Summary: The repository documents a single-MI300X deployment of DeepSeek V4 Flash, pins every runtime artifact by SHA-256, and records upstream base revisions for each diff. The first of two named fixes is correctness rather than throughput: an MXFP4 mixture-of-experts kernel masked its bitmatrix padding lanes against the global tensor bound instead of the logical block size, which corrupted routing under load and produced near-match tool names and forgotten schemas on long prompts. The second is that MI300X implements the AMD FNUZ variant of E4M3 while MI325X and newer use OCP-standard FP8, so a kernel that assumes OCP semantics can be wrong by a factor of two in the scale domain.
- Why it matters: Both traps degrade output quality instead of failing loudly, which is the failure mode anyone porting an FP8 model to CDNA3 is least likely to catch.
send feedback on this story