Coekjan/CrossPool
Efficient GPU Memory Pooling for Multi-LLM Serving via KV Cache and Weight Disaggregation
GitHub repository with 9 stars and 1 forks.
Language: Python
Topics: colocation, disaggregation, kv-cache, llm-inference, llm-serving, memory-disaggregation, memory-pooling