SYNTHETIC-2 logo

SYNTHETIC-2

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Vestibulum luctus felis in nisi tincidunt, vitae facilisis enim tempor.
Category: Compute Network

Description

SYNTHETIC-2 is a large-scale open dataset released by Prime Intellect on July 10, 2025, consisting of four million collaboratively generated and verified reasoning traces, produced using frontier-size models such as DeepSeek-R1-0528 sharded across distributed compute workers via pipeline parallelism. Over 1,250 GPUs (ranging from consumer 4090s to H200 clusters) joined the three-day run, with contributions verified through TOPLOC v2 proofs of computation to ensure honest, tamper-resistant inference. Prime Intellect's global orchestration infrastructure matches GPU nodes into geographically optimized groups, manages node heartbeats, task assignments, and failure recovery, and exposes an orchestrator API for deploying distributed inference workloads. The dataset covers a comprehensive set of new and existing verifiable reasoning tasks (math, coding, code output prediction, JSON/schema adherence, sentence unscrambling, ASCII tree formatting, and precise text extraction), split into SFT and RL subsets with difficulty annotations, and is released as four Hugging Face dataset splits (SYNTHETIC-2, SYNTHETIC-2-SFT-verified, SYNTHETIC-2-SFT-unverified, SYNTHETIC-2-RL). It is intended as the foundation for Prime Intellect's next distributed reinforcement learning runs and future INTELLECT models, serving AI researchers, model trainers, and developers looking to distill or train reasoning-capable language models using open, verifiable data.

Technology & Skills

Uncover the hard and soft skills and tools employed by the organization, and gain insight into the technologies that drive their success