Organizations Tagged with tensor-parallelism for Scalable Model Parallelism and Distributed Training
Discover organizations tagged with tensor-parallelism that implement model sharding, tensor slicing, and inter-GPU communication to scale large language model training and high-throughput inference. This curated list highlights companies, research labs, and open-source projects using frameworks like PyTorch, TensorFlow, DeepSpeed, and Megatron-LM, and showcases implementations such as ZeRO, pipeline/tensor parallelism, NCCL-backed all-reduce, mixed-precision training, and GPU memory optimization. Use the filtering UI below to narrow results by deployment (on-prem, AWS, GCP, Azure), cluster scale (multi-node GPU/TPU), framework integration, and performance metrics, then compare case studies, GitHub repos, and benchmark results to identify best-fit organizations. Explore the list now to evaluate real-world tensor-parallelism adoption, implementation patterns, and integration guides—filter, compare, and dive into technical details to accelerate your model-parallel architecture decisions.