Back to Projects
INTELLECT-3
Lorem ipsum dolor sit amet, consectetur adipiscing elit. Vestibulum luctus felis in nisi tincidunt, vitae facilisis enim tempor.
Category: Compute Network
Description
INTELLECT-3 is a 106B parameter Mixture-of-Experts model trained by Prime Intellect using a two-stage recipe (supervised fine-tuning followed by large-scale reinforcement learning) on a 512 NVIDIA H200 GPU cluster over roughly two months. It was trained end-to-end with prime-rl, Prime Intellect's asynchronous, production-scale RL framework, using agentic RL environments built with the verifiers library and hosted on the community Environments Hub, spanning math, code, science, logic, deep research, and software engineering tasks. Training also relied on Prime Sandboxes, a high-throughput, secure code execution layer for agentic coding rollouts, and a custom compute orchestration stack (Ansible provisioning, Slurm/Cgroup v2, Lustre storage, DCGM/Prometheus observability) managing the H200 cluster. The full recipe — model weights, training frameworks, datasets, RL environments, and evaluations — is open-sourced, and the model can be used via a hosted chat interface (chat.primeintellect.ai) or Prime Intellect's Inference API, with Parasail and Nebius serving as inference providers.Technology & Skills
MACHINE LEARNING
REINFORCEMENT LEARNING
KUBERNETES
STAGEHAND
LLM
REWARD MODELING
AGENT FRAMEWORK
POST-TRAINING
SYNTHETIC DATA
RLVR
BROWSER AUTOMATION
SANDBOX
DSPY
GRAFANA
CODE EXECUTION
BENCHMARK
SFT
RL
PROMETHEUS
INFERENCE
ACCELERATE
HUGGING FACE
EVALUATION
REACT
TERRAFORM
LANGGRAPH
MCP
SGLANG
SWE-BENCH
NEXT.JS
PYTHON
RLHF
VLLM
DISTRIBUTED TRAINING
MODEL GRADING
OBSERVABILITY
DOCKER
TRACING
TYPESCRIPT
HELM
RAY
EVALFLOW
GRPO
AGENT