VecSzn/lobes
An experimental local AI runtime. It relays one request through five small models with one job each, on a single 8 GB consumer GPU.
GitHub repository with 13 stars and 5 forks.
Language: Python
Topics: consumer-gpu, gguf, llama-cpp, llm-inference, local-ai, local-llm, model-routing, on-device-ai, openai-compatible-api