Skip to content

Pull requests: SemiAnalysisAI/InferenceX

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

ci: use Fable 5.1 for all Claude Code invocations
#2805 opened Sep 2, 2026 by cquil11 Collaborator Loading…
perf(agentx): refresh Kimi-K3 MI355X LMCache curve / 刷新 Kimi-K3 MI355X LMCache 曲线 agentx AgentX benchmarks, recipes, and infrastructure AMD full-sweep-enabled
#2804 opened Sep 2, 2026 by hyukjlee Collaborator Loading…
4 of 9 tasks
Validate vLLM P/D cache-source metrics on GB300 full-sweep-enabled
#2797 opened Sep 1, 2026 by cquil11 Collaborator Loading…
[AMD] [AGENTX] Kimi Perf Tuning agentx AgentX benchmarks, recipes, and infrastructure AMD
#2795 opened Sep 1, 2026 by ajith-sirra-amd Collaborator Loading…
[AgentX] Validate vLLM cached-token tier metrics
#2766 opened Aug 28, 2026 by cquil11 Collaborator Draft
Align CI business priority across node counts and model families
#2765 opened Aug 27, 2026 by cquil11 Collaborator Loading…
test: validate upstream-native vLLM Router topologies
#2731 opened Aug 25, 2026 by cquil11 Collaborator Loading…
Remove FlashInfer benchmark backends
#2727 opened Aug 25, 2026 by hbarclay Collaborator Loading…
[WIP][AMD][AgentX] Add Qwen3.8 FP8 MI355X two-node vLLM agentx-fast Run AgentX throughput with 1 warmup request per lane and a 20-minute profile; not reusable sweep-enabled
#2724 opened Aug 25, 2026 by haic0 Collaborator Loading…
6 of 7 tasks
ProTip! Follow long discussions with comments:>50.