Repository navigation
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
为 GDN 和 GLM-5.3-FLASH KDA 增加可选的
compact/replay状态存储,默认native。KDA 按 key 维衰减,inline/precompute 均使用正向衰减的低秩历史重建;历史 key 递推更新,避免每 token 重算整段指数。复用请求接受、fold、CPU checkpoint 和 PD 生命周期,接入真实 prefill、普通 decode、固定/动态 MTP,并保留 GLM indexer tail。调参只用独立 scratch,区分 recurrence/lower_bound;按 token 数和 sequence 容量分别选配置,在各 CUDA Graph 捕获前固定;模型输出携带 verify 配置,去 padding 和双 microbatch 合并 accept 后仍使用对应布局。Replay checkpoint 使用固定布局,后续 batch 调优不会改变已捕获 Graph。缩短 KDA 调参缓存文件名,支持 TP1/2/4/8 的持久化。实际模型还复现并修复了 native Graph 的大批量 empty/HOLD DeepGEMM 调度问题:初始化零值 dummy pool,真实长度继续屏蔽输出;独立 paged/nonpaged logits 与严格 scratch 越界检查通过。
exp记录。GLM 实际34层KDA,TP8/BF16 state/MTP2:native每请求25.5 MiB/卡,replay L8 inline约17.02、precompute约14.89 MiB,含历史/游标,不含其他缓存。没有MTP时 replay 比单 native checkpoint 多占历史空间。
同 H100 的34层完整更新周期(prepare/forward/accept/fold/export/merge,5次重复,8轮,27候选):BF16 precompute MTP2 在batch64/192更新时间减少约8.5–24.8%,batch16增加约16–29%。这是状态更新组件结果,排除conv、模型/MoE与HTTP;不承诺端到端吞吐收益。
实际GLM纯文本服务(TP8/BF16state/MTP2dynamic/L8,同源码与固定容量):四profile各1050测量请求,共4200,无HTTP失败。Replay inline相对native吞吐中位−1.31%至+3.06%,TPOT−1.48%至+6.64%;precompute吞吐−2.94%至+0.62%。短TTFT波动、FP8/BF16生成差异及高draft接受率的synthetic负载限制外推,不声明普遍端到端加速或正式任务质量无回退。当前推荐inline作为GLM验证配置;raw token/时间和4条自然前缀对照完整保存。视觉峰值探测的现成镜像ABI问题另存,使用既有
--disable_vision --disable_audio完成文本对照。