-
Notifications
You must be signed in to change notification settings - Fork 2.8k
All issues
Issue creation is restricted in this repository
- #15044 · laikhtewari opened
on Jun 6, 2026 2 - #3148 · juney-nvidia opened
on Mar 29, 2025 5 - #3124 · juney-nvidia opened
on Mar 27, 2025 11
Issues
is:issue state:open
is:issue state:open
Search results
[Bug]: multi-node MPI-orchestrator serve hangs at IPC init (queues bind loopback 127.0.0.1, no routable-IP override)
Disaggregated serving<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.Inference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Status: Open.#19634 In NVIDIA/TensorRT-LLM;[Bug]: ~6 s fixed per-request latency on sm120 (4x RTX 5060 Ti TP4) with Qwen3.5 hybrid-GDN NVFP4 27B — 1.3.0rc29
Inference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Pytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#19623 In NVIDIA/TensorRT-LLM;[Enhancement] Attention backend selection: name every fallback, reject invalid configs early
feature requestNew feature or request. This includes new model, dtype, functionality supportNew feature or request. This includes new model, dtype, functionality supportInference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Status: Open.#19615 In NVIDIA/TensorRT-LLM;[Bug]: 1.3.0rc27/rc28 containers: singleton MPI_Comm_spawn fails with MPI_ERR_UNKNOWN on hosts with long names (Open MPI 5 truncates the singleton name)
Infra<NV>automated tests, build checks, github actions, system stability & efficiency.<NV>automated tests, build checks, github actions, system stability & efficiency.Status: Open.#19607 In NVIDIA/TensorRT-LLM;[Performance]: sm89 ada blockwise FP8 GEMM copies block-scale factors to shared memory 4x/128x redundantly (stride-0 scale TV layouts)
General perf<NV>Broad performance issues not specific to a particular component<NV>Broad performance issues not specific to a particular componentStatus: Open.#19566 In NVIDIA/TensorRT-LLM;[Bug] _load_bin_or_path_file hides the real torch.load error behind UnboundLocalError
Customized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Status: Open.#19564 In NVIDIA/TensorRT-LLM;[Bug]: DisaggClusterManager drops remaining watch events when one event raises
bugSomething isn't workingSomething isn't workingDisaggregated serving<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.Status: Open.#19551 In NVIDIA/TensorRT-LLM;[Bug]: Beam search with a small KV cache pool fails CUDA graph warmup with "No free block found"
Customized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.KV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferenceStatus: Open.#19527 In NVIDIA/TensorRT-LLM;[Bug]: Whisper (PyTorch backend) fails on any repeated request when KV block reuse is on: "Request requires multimodal_embed_mask_cumsum for chunked prefill or KV-cache reuse"
bugSomething isn't workingSomething isn't workingCustomized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.KV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferencePytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#19515 In NVIDIA/TensorRT-LLM;[Bug]: cross Kv sizing with whisper
bugSomething isn't workingSomething isn't workingKV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferencePytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#19514 In NVIDIA/TensorRT-LLM;[Feature]: Add native text reranking support to the PyTorch backend
feature requestNew feature or request. This includes new model, dtype, functionality supportNew feature or request. This includes new model, dtype, functionality supportPytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#19510 In NVIDIA/TensorRT-LLM;- Status: Open.#19501 In NVIDIA/TensorRT-LLM;