-
Notifications
You must be signed in to change notification settings - Fork 305
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
[Bug]: KV cache size estimate is ~9.7 MB/token on TPU backend — far exceeds both hybrid and dense full-attention expectations
bugSomething isn't workingSomething isn't workingStatus: Open.#3483 In vllm-project/tpu-inference;[Bug]:vllm_model_wrapper does not honor language_model_only when is_multimodal_model is True`
bugSomething isn't workingSomething isn't workingStatus: Open.#3469 In vllm-project/tpu-inference;[Bug]: TPUConnectorHMA.__init__ does not call super().__init__, causing AttributeError on _kv_transfer_config and _vllm_config`
bugSomething isn't workingSomething isn't workingStatus: Open.#3468 In vllm-project/tpu-inference;- Status: Open.#3467 In vllm-project/tpu-inference;
- Status: Open.#3456 In vllm-project/tpu-inference;
- Status: Open.#3455 In vllm-project/tpu-inference;
- Status: Open.#3454 In vllm-project/tpu-inference;
- Status: Open.#3400 In vllm-project/tpu-inference;
- Status: Open.#3399 In vllm-project/tpu-inference;
[Bug]: Default memory config cannot boot a 256K-vocab model on v5e-1: the [512]-bucket structured_decode program has no room left to load
bugSomething isn't workingSomething isn't workingStatus: Open.#3360 In vllm-project/tpu-inference;[Bug]: TPU GDN batched decode hangs when padded request metadata is narrower than decode_tile_size
bugSomething isn't workingSomething isn't workingStatus: Open.#3356 In vllm-project/tpu-inference;- Status: Open.#3350 In vllm-project/tpu-inference;