[LMCache] Add LMCache arm to official DSV4 FP4 B200 vLLM AgentX sweep / 将 LMCache 分支并入官方 DSV4 FP4 B200 vLLM AgentX 扫描#2231
[LMCache] Add LMCache arm to official DSV4 FP4 B200 vLLM AgentX sweep / 将 LMCache 分支并入官方 DSV4 FP4 B200 vLLM AgentX 扫描#2231ApostaC wants to merge 3 commits into
Conversation
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
中文:将 perf-changelog 条目的 pr-link 填写为 PR #2231。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
| --trust-remote-code | ||
| --kv-cache-dtype fp8 | ||
| --block-size 256 | ||
| "${PARALLEL_ARGS[@]}" | ||
| "${VLLM_CP_ARGS[@]}" | ||
| "${EP_ARGS[@]}" | ||
| --compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE","custom_ops":["all"]}' | ||
| --attention_config.use_fp4_indexer_cache=True | ||
| --max-model-len 1048576 | ||
| --gpu-memory-utilization 0.92 | ||
| --numa-bind | ||
| "${CUMEM_ARGS[@]}" | ||
| --no-enable-flashinfer-autotune | ||
| --tokenizer-mode deepseek_v4 | ||
| --tool-call-parser deepseek_v4 | ||
| --enable-auto-tool-choice | ||
| --reasoning-parser deepseek_v4 | ||
| --enable-prefix-caching | ||
| --attention-config '{"backend":"FLASHINFER_MLA_SPARSE_DSV4","use_prefill_query_quantization":true,"use_fp4_indexer_cache":true}' | ||
| --no-disable-hybrid-kv-cache-manager | ||
| --disable-uvicorn-access-log | ||
| --compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY","mode":0}' | ||
| --max-num-seqs "$MAX_NUM_SEQS" | ||
| --max-cudagraph-capture-size "$MAX_NUM_SEQS" | ||
| "${PARALLEL_ARGS[@]}" | ||
| "${VLLM_CP_ARGS[@]}" | ||
| "${EP_ARGS[@]}" | ||
| "${FAST_MOE_ARGS[@]}" | ||
| "${OFFLOAD_ARGS[@]}" | ||
| ) | ||
| printf '%q ' "${VLLM_CMD[@]}" | tee "$RESULT_DIR/vllm_command.txt" |
There was a problem hiding this comment.
🔴 The rewritten VLLM_CMD in this file drops --tool-call-parser deepseek_v4 and --enable-auto-tool-choice, which were present in the pre-PR version of this exact script and are set in every other agentic vLLM recipe (dsv4_fp4_b300_vllm.sh, dsv4_fp4_mi355x_vllm.sh, dsv4_fp8_h200.sh, minimax/kimi variants). This looks like an accidental drop during the #2224 recipe rewrite rather than an intentional change, and since AgentX trace replay relies on tool_choice=auto, vLLM will reject those requests without --enable-auto-tool-choice, likely failing the sweep across all offload arms (none/vllm-simple/mooncake/lmcache).
Extended reasoning...
What changed: In the VLLM_CMD array construction (around lines 328-350 of benchmarks/single_node/agentic/dsv4_fp4_b200_vllm.sh), the diff removes --tool-call-parser deepseek_v4 and --enable-auto-tool-choice while keeping the adjacent --reasoning-parser deepseek_v4 flag. The pre-PR version of this exact file had both flags (visible as removed lines in the diff), so this is not a case of the recipe never having them — they were explicitly dropped during the VLLM_CMD rewrite that pulled in the #2224 tuned recipe.
Why this is inconsistent with the rest of the repo: Every other agentic-coding vLLM recipe sets both flags together: dsv4_fp4_b300_vllm.sh (lines 209-210), dsv4_fp4_mi355x_vllm.sh (lines 396/398), and dsv4_fp8_h200.sh (lines 63-64) — confirmed directly by grepping those files. The minimax and kimi vLLM agentic recipes follow the same pattern. The PR description states the b200 script 'adopts the #2224 tuned recipe verbatim,' but #2224's own b300 sibling keeps both flags, so this omission is specific to the b200 rewrite and not an intentional nightly-vLLM behavior change.
Why nothing else in the script prevents this from mattering: The workload driven by build_replay_cmd/run_agentic_replay_and_write_outputs is the inferencex-agentx-mvp coding-agent trace replay, which is fundamentally tool-calling — requests are reconstructed with a tools schema and tool_choice. vLLM does not enable auto tool-choice/parsing by default; when a request specifies tool_choice=auto and the server wasn't started with --enable-auto-tool-choice (paired with a --tool-call-parser), vLLM returns an HTTP 400 ('auto tool choice requires --enable-auto-tool-choice and --tool-call-parser'). This is a request-level rejection, independent of whether the replay's default pre-canned mode discards live assistant responses — the request still needs to succeed for aiperf to count it as a valid data point.
Step-by-step proof of impact:
- The lmcache-arm (and every other arm: none/vllm-simple/mooncake) all share this single
VLLM_CMD— none of the offload-backendcasebranches add tool-call flags. - vLLM starts successfully (these flags don't affect startup), so
wait_for_server_readypasses and the sweep proceeds to the replay phase. build_replay_cmdissues chat-completion requests for the AgentX coding-agent trace with atoolsarray andtool_choice=auto(required to reproduce realistic input-token counts for tool-augmented turns).- Because
--enable-auto-tool-choice/--tool-call-parserare absent, vLLM rejects essentially all such requests with a 400. - aiperf's
--failed-request-threshold(10%, perAIPERF_FAILED_REQUEST_THRESHOLD) is exceeded almost immediately, failing the run for all 14 lmcache sweep points plus the parent vllm-simple/mooncake points that share the same script.
Fix: Re-add --tool-call-parser deepseek_v4 and --enable-auto-tool-choice to the shared VLLM_CMD array (near --reasoning-parser deepseek_v4), matching the pre-PR file and every sibling recipe.
| unset VLLM_USE_SIMPLE_KV_OFFLOAD | ||
| OFFLOAD_ARGS=( | ||
| --kv-transfer-config | ||
| "{\"kv_connector\":\"LMCacheMPConnector\",\"kv_role\":\"kv_both\",\"kv_connector_extra_config\":{\"lmcache.mp.host\":\"$LMCACHE_CONNECT_HOST\",\"lmcache.mp.port\":$LMCACHE_PORT,\"lmcache.mp.mq_timeout\":$LMCACHE_MQ_TIMEOUT}}" | ||
| ) |
There was a problem hiding this comment.
🔴 The new lmcache arm's --kv-transfer-config sets kv_connector":"LMCacheMPConnector" but omits kv_connector_module_path, which both sibling scripts wiring the same connector (kimik2.5_fp4_b200.sh:150, dsv4_fp4_mi355x_vllm.sh:344) include. LMCacheMPConnector is an out-of-tree class that vLLM can only resolve via kv_connector_module_path + kv_connector; without it, vllm serve will fail to construct the KV connector at startup, so every one of the 14 lmcache sweep points will crash before producing a result. Fix by adding "kv_connector_module_path":"lmcache.integration.vllm.lmcache_mp_connector" to the JSON at line 288.
Extended reasoning...
The bug: The new lmcache case in dsv4_fp4_b200_vllm.sh (lines 285-289) builds OFFLOAD_ARGS as:
"{\"kv_connector\":\"LMCacheMPConnector\",\"kv_role\":\"kv_both\",\"kv_connector_extra_config\":{...}}"
This is missing the kv_connector_module_path key. LMCacheMPConnector is shipped by the lmcache pip package (installed a few lines earlier via agentic_pip_install ... lmcache==$LMCACHE_VERSION), not by vLLM itself. vLLM'''s built-in KVConnectorFactory only knows connector classes it ships in-tree by bare name; to resolve an out-of-tree connector class, it needs to dynamically import the module named by kv_connector_module_path and then look up kv_connector as a class name inside that module. Without the module path, vLLM falls back to its built-in name registry, which has no entry for LMCacheMPConnector, and connector construction fails at vllm serve startup.
Why the existing sanity check doesn'''t catch it: The script runs python3 -c \"import lmcache.integration.vllm.lmcache_mp_connector\" >/dev/null a few lines before launching the server. That only proves the module is importable in a throwaway Python process — it does not register the connector with vLLM'''s factory in the actual vllm serve process, so it provides no protection against this specific misconfiguration.
Cross-script evidence: Both other scripts in this repo that wire up the identical LMCacheMPConnector include the field explicitly and consistently:
benchmarks/single_node/agentic/kimik2.5_fp4_b200.sh:150:\"kv_connector\":\"LMCacheMPConnector\",\"kv_connector_module_path\":\"lmcache.integration.vllm.lmcache_mp_connector\",...benchmarks/single_node/agentic/dsv4_fp4_mi355x_vllm.sh:344: same field, same value, otherwise near byte-for-byte the same JSON shape (samelmcache.mp.host/lmcache.mp.portextra_config keys).
The new B200 arm is the only place in the codebase wiring LMCacheMPConnector without this field, strongly suggesting it was simply dropped rather than being some new supported invocation.
Impact: This is not a subtle perf regression — it'''s a hard startup crash. Every one of the 14 lmcache sweep points defined in configs/nvidia-master.yaml (dsv4-fp4-b200-vllm-agentic-lmcache: 3 TP8 points + 11 TP8-DEP8 points) launches vllm serve with this malformed --kv-transfer-config, so the server process would fail while constructing the KV connector, before any benchmark traffic is served. This defeats the entire purpose of the PR, which is to add this lmcache arm for comparison against vllm-simple and Mooncake.
Step-by-step proof:
KV_OFFLOAD_BACKEND=lmcacheis set for a sweep point; the script enters thelmcache)case at line ~230.- The LMCache MP server starts and becomes healthy (this part is unaffected — it'''s a separate process from vLLM).
OFFLOAD_ARGSis built at line 285-289 as--kv-transfer-config '\''{\"kv_connector\":\"LMCacheMPConnector\",\"kv_role\":\"kv_both\",...}'\''— nokv_connector_module_path.vllm serveis launched with this flag (line ~372, via\"${OFFLOAD_ARGS[@]}\").- vLLM parses
--kv-transfer-config, seeskv_connector=\"LMCacheMPConnector\"with no module path, and attempts to resolve it against its built-in registry (which contains e.g.LMCacheConnectorV1but notLMCacheMPConnector). - Resolution fails, and
vllm serveexits/crashes during connector construction, before the health-check endpoint ever comes up —wait_for_server_readyat the bottom of the script will time out and the whole benchmark point fails.
Why this wasn'''t caught: The PR description explicitly says validation was only bash -n (syntax check) and generate_sweep_configs.py full-sweep (config generation), neither of which launches an actual vLLM server, so this startup failure was never exercised.
Fix: add \"kv_connector_module_path\":\"lmcache.integration.vllm.lmcache_mp_connector\" to the JSON at line 288, matching the sibling scripts exactly.
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=29460892084 |
…sweep Rebased onto main after PR #2224 merged the tuned B200 recipe. The LMCache points now live directly in the official dsv4-fp4-b200-vllm-agentic search space alongside the vllm-simple and Mooncake arms (no standalone config section). LMCache 0.5.1 MP server (lmcache_driven transfer mode) + LMCacheMPConnector per PR #2153; the lmcache arm drops --enable-cumem-allocator because cuMem/VMM allocations cannot be CUDA-IPC-exported to the LMCache server. Ladder validated in the PR #2231 bring-up sweep: peak ~25.7k total tok/s/GPU at DEP8 conc 72 with 96-98% cache hit. 中文:在 PR #2224 的调优配方合并后 rebase 到 main。LMCache 测试点直接并入 官方 dsv4-fp4-b200-vllm-agentic 搜索空间,与 vllm-simple、Mooncake 分支 并列(不再使用独立配置段)。LMCache 0.5.1 MP server(lmcache_driven 传输模式)+ LMCacheMPConnector(沿用 PR #2153);lmcache 分支去掉 --enable-cumem-allocator(cuMem/VMM 分配无法通过 CUDA IPC 导出给 LMCache server)。测试点阶梯已在 PR #2231 调试扫描中验证:DEP8 并发 72 达到峰值约 25.7k total tok/s/GPU,缓存命中率 96–98%。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
cd7b86f to
14bd83f
Compare
Change the LMCache DEP8 points from the vllm-simple ladder to the Mooncake ladder ([12, 20, 28, 36, 44, 52, 60, 68, 76]) so the two external KV stores are benchmarked at identical concurrency points. TP8 stays on [8, 12, 16] (Mooncake has no TP8 arm). 中文:将 LMCache DEP8 测试点从 vllm-simple 阶梯改为 Mooncake 阶梯 ([12, 20, 28, 36, 44, 52, 60, 68, 76]),使两个外部 KV 存储在完全相同的 并发点上进行基准测试。TP8 保持 [8, 12, 16](Mooncake 无 TP8 分支)。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=29532605829 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=29533029190 |
# Conflicts: # perf-changelog.yaml
Description
Adds an LMCache 0.5.1 DRAM KV-offload arm to the official
dsv4-fp4-b200-vllm-agenticsweep, alongside the vllm-simple and Mooncake arms, on the tuned recipe and image that #2224 merged (vllm/vllm-openai:nightly-dev-x86_64-cu13.0.1-904e4ec). Rebased onto main after #2224 merged; the LMCache points now live directly in the official config's search space (the earlier standalonedsv4-fp4-b200-vllm-agentic-lmcachesection is gone).benchmarks/single_node/agentic/dsv4_fp4_b200_vllm.shadds anlmcacheoffload-backend case (from [LMCache] Add LMCache configs for dsv4 vllm b200/b300 agentix setups #2153): LMCache MP server (lmcache_driventransfer mode) +LMCacheMPConnector, L1 pool derated to 75% ofTOTAL_CPU_DRAM_GB.--enable-cumem-allocatoris dropped, becauseLMCacheMPConnectorexports the KV cache through legacy CUDA IPC handles and cuMem/VMM allocations cannot be exported that way (register_kv_cachesfails withcudaErrorInvalidValue). All other flags are identical across arms, so backends are directly comparable.Already validated on hardware in this PR's earlier bring-up sweep (run 29460892084, pre-rebase, standalone-section structure — same script/recipe): 12/12 points that received nodes passed (2 points lost to SLURM allocation revocations, never reached the container). Peak ~25.7k total tok/s/GPU and 194.5 output tok/s/GPU at DEP8 conc 72 (14.2 tok/s/user), 96–98% LMCache hit rate across the ladder.
中文说明
在官方
dsv4-fp4-b200-vllm-agentic扫描中新增 LMCache 0.5.1 DRAM KV 卸载分支,与 vllm-simple、Mooncake 分支并列,使用 #2224 合并的调优配方与镜像(vllm/vllm-openai:nightly-dev-x86_64-cu13.0.1-904e4ec)。已在 #2224 合并后 rebase 到 main;LMCache 测试点直接并入官方配置的搜索空间(不再使用之前独立的dsv4-fp4-b200-vllm-agentic-lmcache配置段)。lmcache卸载后端分支(来自 [LMCache] Add LMCache configs for dsv4 vllm b200/b300 agentix setups #2153):LMCache MP server(lmcache_driven传输模式)+LMCacheMPConnector,L1 池按TOTAL_CPU_DRAM_GB的 75% 降额。--enable-cumem-allocator:LMCacheMPConnector通过传统 CUDA IPC 句柄导出 KV cache,cuMem/VMM 分配无法以这种方式导出(register_kv_caches报cudaErrorInvalidValue)。其余参数在各分支间完全一致,便于卸载后端直接对比。本 PR 早期的调试扫描(run 29460892084,rebase 前的独立配置段结构,脚本/配方相同)已在硬件上验证:拿到节点的 12 个点全部通过(2 个点因 SLURM 分配被撤销未能启动容器)。峰值约 25.7k total tok/s/GPU、194.5 output tok/s/GPU(DEP8 并发 72,14.2 tok/s/user),全阶梯 LMCache 命中率 96–98%。
Related Issue
Builds on #2224 (merged tuned B200 recipe) and #2153 (LMCache backend). / 基于 #2224(已合并的 B200 调优配方)与 #2153(LMCache 后端)。
Type of Change
Checklist
perf-changelog.yamlperf-changelog.yamlentries are appended to the end of the file (the file is chronological: oldest at top, newest at bottom)OWNER/MEMBER/COLLABORATOR) has commented/reuse-sweep-runon this PR — do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps (remove and re-add the label to force one)🤖 Generated with Claude Code