FluidPD Released: Revolutionizing LLM Serving with SLO-Aware In-Place Elasticity
By Mr.Xu
Published:
Summary:FluidPD is a novel P/D (prefill/decode) disaggregated LLM serving system designed to address latency issues caused by workload fluctuations in existing systems. It introduces two complementary mechanisms: FluidToken handles transient imbalances by offloading a portion of prefill computation to decode workers when decode-side slack is available, while FluidRole addresses sustained imbalances by reassigning running workers between prefill and decode roles in place, avoiding model reload and engine
Background and Challenges
In the architecture of Large Language Model (LLM) serving systems, the prefill and decode phases are typically separated due to their distinct execution patterns and Service Level Objective (SLO) targets. Existing systems usually employ a fixed prefill/decode worker ratio and distribute tasks through request routing. However, real-world workloads exhibit both short bursts and sustained shifts in the prefill-to-decode demand ratio, leading to configuration mismatches and latency SLO violations even when idle capacity exists elsewhere. While existing autoscaling mechanisms can add capacity, they react slowly, require spare GPUs, and do not directly address short-timescale phase imbalance.
FluidPD's Innovative Solution
FluidPD introduces a SLO-aware in-place elasticity scheme with two complementary mechanisms:
- FluidToken: Dynamically offloads a portion of prefill computation to decode workers when decode-side slack is available, addressing transient imbalances.
- FluidRole: Reassigns running workers between prefill and decode roles in place, avoiding model reload and engine restart, to handle sustained imbalances.
These mechanisms are guided by lightweight pressure indices that expose resource pressure on the prefill and decode sides before SLO violations occur.
Experiments and Results
Experiments on Azure production workloads demonstrate that FluidPD improves SLO attainment by up to 94.6 percentage points over static SGLang, showcasing its potential to enhance service quality without additional resource provisioning.
Industry Impact and Developer Recommendations
- Industry Impact: FluidPD provides a new approach to optimizing LLM serving architectures, particularly in handling dynamic load changes. Its efficient resource utilization and SLO-aware mechanisms make it an ideal choice for high-load, high-demand scenarios.
- Developer Recommendations: For developers building or optimizing LLM services, FluidPD offers a viable solution. It is recommended to consider adopting similar elasticity strategies in existing systems to improve service quality and resource utilization.
Future Outlook
The release of FluidPD marks a significant advancement in LLM serving architectures. As AI technology continues to evolve, similar elasticity mechanisms are expected to find applications in more fields, further driving the普及 and optimization of AI services.
— END —Source: ArXiv AI (cs.AI) (2026-10-07)
Tags: #FluidPD #LLMs & Foundation Models #SLO-aware #Elasticity #AI Services
Community Comments