AI inference
AI inference

Projects Tagged ai-inference: Model Serving, Edge & Cloud Inference Optimization for ML Applications

Explore projects filtered by the tags pillar for ai-inference: a curated list of ML and software projects that implement model serving, on-device and edge inference, and scalable cloud inference pipelines. Discover implementation details and long-tail techniques such as low-latency model serving, batch and real-time inference, quantization, pruning, distillation, ONNX and model format conversion, and acceleration on GPU/TPU/NPU hardware; compare projects using TensorFlow Serving, TorchServe, NVIDIA Triton, ONNX Runtime, or custom edge runtimes. Use the filtering UI to narrow results by framework, hardware accelerator, latency/throughput benchmarks, and deployment environment, then review code, benchmarks, and architecture patterns to adopt best practices. Start exploring projects, compare implementations, and contribute or fork repositories to accelerate production-ready inference deployments.
Other Filters