ML model optimization product to accelerate inference.
-
Updated
Jun 2, 2025 - Python
ML model optimization product to accelerate inference.
Benchmark llama.cpp on your own GPU/CPU boxes and compare against everyone else's — sweep runner, an sha256-keyed index of ~4M GGUF files, and a shared results database you query in plain language. Self-host it, or plug a machine into llamatoaster.com.
A simple tensorflow C++ REST API server
Inference Time Performance stats for various backbone networks.
Benchmarking the impact of compiler-level graph optimizations on NPU inference performance and memory wall bottlenecks.
Optimising train, inference and throughput of expensive ML models
Linear and Multiple Regression with data manipulation using SQL and R functions.
Check the fastText's inference performance for OOV.
To associate your repository with the inference-performance topic, visit your repo's landing page and select "manage topics."