Advisory · CVE-2026-103042
Unauthenticated memory exhaustion in LightLLM NCCL control channel
LightLLM through 1.2.0 lets unauthenticated attackers exhaust KV-transfer worker memory through the NCCL control channel's exposed_set_value method when started with --pd_trans_mode nccl, crashing workers and triggering node failure.
- Vendor
- ModelTC
- Product
- LightLLM
- Identifier / CWE
- CVE-2026-103042
CWE-770 - Action timing
- Immediate
Explain it like I’m five
LightLLM's NCCL channel is a storage room with no limit on what gets shelved. A stranger keeps shoving junk in until the room bursts and the whole worker collapses.
- 01NCCL transfer mode
The LightLLM server is started with --pd_trans_mode nccl, enabling the NCCL control channel.
- 02Method exposed
The channel exposes the set_value method to unauthenticated network clients.
- 03Unbounded storage
The attacker calls the method repeatedly to store unbounded key-value pairs with no size limits.
- 04Memory exhaustion
KV-transfer worker memory fills until the worker crashes, triggering node failure.
What happened
LightLLM through 1.2.0 contains a memory exhaustion vulnerability in the NCCL control channel when the server is started with --pd_trans_mode nccl. Unauthenticated attackers can call the exposed exposed_set_value method to store unbounded key-value pairs with no size limits, exhausting KV-transfer worker memory until the worker crashes and triggers node failure.
No patch release is documented in the disclosure at the time of writing; the advisory states only that LightLLM through 1.2.0 is affected.
What to do
- Inventory LightLLM deployments running with
--pd_trans_mode nccl. - Restrict network access to the NCCL control channel so it is not reachable from untrusted clients.
- Switch away from
nccltransfer mode where operationally feasible until a fixed release is published. - Upgrade to a fixed LightLLM release as soon as the vendor publishes one.
- Monitor KV-transfer worker memory and set alerts on abnormal growth as an early warning.
Management note
This is an unauthenticated availability attack against GPU inference capacity: no code execution, but repeated crashes of KV-transfer workers take nodes out of service. Network segmentation plus a mode change is a complete interim mitigation.