123123213weqw/dual-v100-llama.cpp
Reproducible llama.cpp kernel and runtime optimization lab for dual NVIDIA Tesla V100 GPUs (SM70)
GitHub repository with 6 stars and 0 forks.
Language: Python
Topics: cuda, inference, llama-cpp, sm70, speculative-decoding, tensor-parallel, tesla-v100