Paper reading: Aegaeon - Efficient GPU pooling technology for concurrent LLM services on the market
The Aegaeon system deployed by Alibaba Cloud Model Market achieves 82% GPU resource savings through token-level automatic expansion, and supports a single GPU to serve up to 7 models at the same time.
Oct 22, 2025