B-A-M-N/SOLLOL
Super Ollama Load Balancer - Performance-aware routing for distributed Ollama deployments with Ray, Dask, and adaptive metrics
GitHub repository with 6 stars and 3 forks.
Language: Python
Topics: ai, asyncio, caching, distributed-inference, gpu-monitoring, http2, hybrid-routing, inference-optimization, llama-cpp, llm