Skip to content

[Klaud Cold] Update dsv4-fp4-b200-vllm-agentic-mtp vLLM image to nightly-dev-x86_64-cu130-ac9126e58aa7 / [Klaud Cold] 将 dsv4-fp4-b200-vllm-agentic-mtp 的 vLLM 镜像更新至 nightly-dev-x86_64-cu130-ac9126e58aa7 - #3700

Closed
Klaud-Cold wants to merge 1 commit into
mainfrom
klaud/auto-b7e2ccf9e45c74b9-a1d975edc8e792ac

Conversation

@Klaud-Cold

@Klaud-Cold Klaud-Cold commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Goal: Update vLLM image from vllm/vllm-openai:nightly-dev-x86_64-cu13.0.1-591bb95 to vllm/vllm-openai:nightly-dev-x86_64-cu130-ac9126e58aa7@sha256:3d17635c5aa0340e2f2502c067ce4b587d8e7dc678195e90b71142490dc6d71c.
Baseline: 2026-09-19 · vllm/vllm-openai:nightly-dev-x86_64-cu13.0.1-591bb95
Mean latency · Sources: API 1, API 2

Point Total tok/s/GPU ↑ Output tok/s/GPU ↑ TTFT ms ↓ TPOT ms ↓
AgentX c1 TP8 EP1 9c1d23 N/A N/A N/A N/A
AgentX c1 TP8 EP1 cce23c 2,201.3 16.41 1,860.04 3.48
AgentX c4 TP8 EP1 929804 3,874.04 27.14 764.45 5.03
AgentX c4 TP8 EP1 f81c5a N/A N/A N/A N/A
AgentX c6 TP8 EP1 0d8f91 N/A N/A N/A N/A
AgentX c6 TP8 EP1 8fbfb3 4,483.15 35.42 855.8 4.76
AgentX c10 TP8 EP1 da6c15 N/A N/A N/A N/A
AgentX c10 TP8 EP1 de1d5b 8,029.38 59.46 736.62 6.12
AgentX c14 TP8 EP1 2dc066 9,418.59 74.89 803.57 6.91
AgentX c14 TP8 EP1 68c217 N/A N/A N/A N/A
AgentX c16 TP8 EP1 eb2826 N/A N/A N/A N/A
AgentX c16 TP8 EP1 edc6c7 11,719.42 78.68 932.12 8.98
AgentX c32 TP8 EP8 5f42d6 17,480.3 113.95 6,331.27 13.31
AgentX c32 TP8 EP8 658e6f N/A N/A N/A N/A
AgentX c64 TP8 EP8 202185 34,561.72 261.21 7,683.01 11.9
AgentX c64 TP8 EP8 ac9a9a N/A N/A N/A N/A
AgentX c96 TP8 EP8 109868 47,863.05 385.49 9,231 13.57
AgentX c96 TP8 EP8 a7529d N/A N/A N/A N/A
AgentX c128 TP8 EP8 5f0930 N/A N/A N/A N/A
AgentX c128 TP8 EP8 609b94 54,814.79 455.61 10,673.35 14.83
AgentX c160 TP8 EP8 51422e N/A N/A N/A N/A
AgentX c160 TP8 EP8 794fe9 64,489.73 499.22 12,379.88 17.17
AgentX c192 TP8 EP8 68a170 N/A N/A N/A N/A
AgentX c192 TP8 EP8 ba0ad2 69,323.46 522.64 14,684.94 22.13

Note: AgentX c1 TP8 EP1 9c1d23, AgentX c4 TP8 EP1 f81c5a, AgentX c6 TP8 EP1 0d8f91, AgentX c10 TP8 EP1 da6c15, AgentX c14 TP8 EP1 68c217, AgentX c16 TP8 EP1 eb2826, AgentX c32 TP8 EP8 658e6f, AgentX c64 TP8 EP8 ac9a9a, AgentX c96 TP8 EP8 a7529d, AgentX c128 TP8 EP8 5f0930, AgentX c160 TP8 EP8 51422e, AgentX c192 TP8 EP8 68a170: unavailable.

Eval: N/A

中文

**目标:**将 vLLM 镜像从 vllm/vllm-openai:nightly-dev-x86_64-cu13.0.1-591bb95 更新为 vllm/vllm-openai:nightly-dev-x86_64-cu130-ac9126e58aa7@sha256:3d17635c5aa0340e2f2502c067ce4b587d8e7dc678195e90b71142490dc6d71c。
**基线:**2026-09-19 · vllm/vllm-openai:nightly-dev-x86_64-cu13.0.1-591bb95
平均延迟 · 来源: API 1, API 2;数值及异常说明见上表。

…c9126e5

Move the B200 DeepSeek-V4-Pro vLLM AgentX family from the local dev build
nightly-dev-x86_64-cu13.0.1-591bb95 to the official
nightly-dev-x86_64-cu130-ac9126e58aa7 image, pinned by digest. Upstream vLLM
at ac9126e5 has no deep_gemm_amxf4_mega_moe backend or
VLLM_DSV4_MEGA_FP8_COMBINE variable, so DEP8 uses deep_gemm_mega_moe and the
unread variable is dropped.

将 B200 DeepSeek-V4-Pro vLLM AgentX 配置从本地 dev 构建
nightly-dev-x86_64-cu13.0.1-591bb95 更新为官方
nightly-dev-x86_64-cu130-ac9126e58aa7 镜像,并按 digest 固定。上游 vLLM
ac9126e5 不提供 deep_gemm_amxf4_mega_moe 后端和 VLLM_DSV4_MEGA_FP8_COMBINE
变量,因此 DEP8 改用 deep_gemm_mega_moe,并移除不再读取的变量。

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps run only when the PR appends an inferencex-e2e/perf-changelog.yaml entry and carries exactly one primary label: full-sweep-fail-fast (strongly recommended; canary plus per-matrix fail-fast), full-sweep-enabled (canary; matrix jobs continue after a failure), or non-canary-full-sweep-enabled (no canary or fail-fast). The modifiers all-evals, evals-only, and agentx-fast require a primary label. On fork PRs, a maintainer applies the label. See sweep labels and reuse.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**只有当 PR 在 inferencex-e2e/perf-changelog.yaml 末尾追加了条目,并且恰好带有一个主标签时,才会运行扫描:full-sweep-fail-fast(强烈推荐;canary 加逐矩阵 fail-fast)、full-sweep-enabled(有 canary;矩阵任务在失败后继续运行)或 non-canary-full-sweep-enabled(无 canary,也无 fail-fast)。修饰标签 all-evals、evals-only 和 agentx-fast 必须与主标签一起使用。fork PR 的标签由维护者添加。参见扫描标签与复用。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

failed · Cleanup pending. Stop and confirm owned runs before closing.

中文

failed · 清理待完成。先停止并确认自有运行结束,再关闭 PR。

@Klaud-Cold Klaud-Cold closed this Oct 3, 2026
@Klaud-Cold
Klaud-Cold deleted the klaud/auto-b7e2ccf9e45c74b9-a1d975edc8e792ac branch October 3, 2026 07:01
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

failed · Reason: baseline-point-mismatch · Repairs: 0 · Runs: —
All owned runs ended. PR closed; branch deleted for retry.

中文

failed · 原因:baseline-point-mismatch · 修复次数:0 · 运行:—
所有自有运行均已结束。PR 已关闭;分支已删除,可重新尝试。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant