revise doc #4185

Angazenn · 2025-11-13T12:32:15Z

What this PR does / why we need it?

Does this PR introduce any user-facing change?

How was this patch tested?

vLLM version: v0.11.0
vLLM main: vllm-project/vllm@2918c1b

Signed-off-by: Angazenn <[email protected]>

gemini-code-assist

Code Review

This pull request updates the documentation for ACL Graphs. I've provided one comment to improve clarity by addressing a potentially confusing parameter name (cudagraph_capture_sizes in an ACL graph context) and an undefined acronym ("ep"). These changes will help ensure the documentation is clear and unambiguous for developers.

gemini-code-assist · 2025-11-13T12:33:13Z

docs/source/developer_guide/feature_guide/ACL_Graph.md


 ```

+In vLLM, these thresholds are set by `cudagraph_capture_sizes`. The default capture sizes are like `[1,2,4,8,16,24,32,...,max_capture_size]`. You can customize capture sizes to get fine-grained control over performance. For example, we can set `cudagraph_capture_sizes` as `[1,2,4,6,12,18]` when running Qwen3-235B on decode node in large ep.


This is a helpful addition. However, to prevent potential confusion for developers, I suggest a couple of clarifications:

Confusing Parameter Name: This document is about ACL Graphs for Ascend NPUs, but the parameter is named cudagraph_capture_sizes. This is potentially misleading. It would be good to add a note clarifying that this parameter applies to both CUDA and ACL graphs.

Undefined Acronym: The term "ep" in "large ep" is ambiguous. Please spell it out (e.g., "expert parallelism") for clarity.

Given this is developer-facing documentation where clarity is crucial to prevent configuration errors, I've provided a suggestion to improve it.

Suggested change

In vLLM, these thresholds are set by `cudagraph_capture_sizes`. The default capture sizes are like `[1,2,4,8,16,24,32,...,max_capture_size]`. You can customize capture sizes to get fine-grained control over performance. For example, we can set `cudagraph_capture_sizes` as `[1,2,4,6,12,18]` when running Qwen3-235B on decode node in large ep.

In vLLM, these thresholds are set by `cudagraph_capture_sizes`. Note that this parameter applies to both CUDA graphs and Ascend's ACL graphs. The default capture sizes are like `[1,2,4,8,16,24,32,...,max_capture_size]`. You can customize capture sizes to get fine-grained control over performance. For example, we can set `cudagraph_capture_sizes` as `[1,2,4,6,12,18]` when running Qwen3-235B on a decode node in a large expert parallelism (EP) setup.

github-actions · 2025-11-13T12:33:20Z

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

A PR should do only one thing, smaller PRs enable faster reviews.
Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

revise doc

7471c95

Signed-off-by: Angazenn <[email protected]>

gemini-code-assist bot reviewed Nov 13, 2025

View reviewed changes

github-actions bot added the documentation Improvements or additions to documentation label Nov 13, 2025

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

revise doc #4185

revise doc #4185

Angazenn commented Nov 13, 2025 •

edited by github-actions bot

Loading

Uh oh!

gemini-code-assist bot left a comment

Uh oh!

gemini-code-assist bot Nov 13, 2025

Uh oh!

github-actions bot commented Nov 13, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

1 participant


		```

		In vLLM, these thresholds are set by `cudagraph_capture_sizes`. The default capture sizes are like `[1,2,4,8,16,24,32,...,max_capture_size]`. You can customize capture sizes to get fine-grained control over performance. For example, we can set `cudagraph_capture_sizes` as `[1,2,4,6,12,18]` when running Qwen3-235B on decode node in large ep.

revise doc #4185

Are you sure you want to change the base?

revise doc #4185

Conversation

Angazenn commented Nov 13, 2025 • edited by github-actions bot Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

What this PR does / why we need it?

Does this PR introduce any user-facing change?

How was this patch tested?

Uh oh!

gemini-code-assist bot left a comment

Choose a reason for hiding this comment

Code Review

Uh oh!

gemini-code-assist bot Nov 13, 2025

Choose a reason for hiding this comment

Uh oh!

github-actions bot commented Nov 13, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

1 participant

Angazenn commented Nov 13, 2025 •

edited by github-actions bot

Loading