GR00T N1.7
NVIDIA·
Overview
Open vision-language-action model for humanoid robots, released in early access under Apache 2.0 — the first fully commercially licensed model of the GR00T line. A vision-language backbone is paired with a diffusion transformer head that denoises continuous actions via flow matching. NVIDIA credits the generalization and language-following gains over N1.6 to 20,000 hours of EgoScale human egocentric video in pretraining. At 3B parameters it is small enough to run on local hardware, which is why it also appears in the Local tab.
Capabilities and innovations
Benchmarks
No tracked academic benchmark applies to robotics models; the timeline carries no scores and no leadership tracking for them.
- GPQA Diamond: not reported by the vendor
- Humanity’s Last Exam (no tools): not reported by the vendor
- SWE-bench Verified: not reported by the vendor
- SWE-bench Pro: not reported by the vendor
- MMLU: not reported by the vendor
- MMLU-Pro: not reported by the vendor
- Terminal-Bench 2.x: not reported by the vendor
- MMMU-Pro: not reported by the vendor
Architecture and hardware
- Parameters
- 3B total (Dense)
- Estimated VRAM at Q4
- ~2 GB, Mobile / NPU class
Links
More from NVIDIA
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.