Ornith-1.5-397B
Ornith·
Overview
MIT-licensed flagship of the Ornith-1.5 family from the research lab DeepReinforce. 397B total / 17B active MoE, post-trained from Qwen3.5-397B-A17B and keeping the base model's parameter layout unchanged — the first entry in this catalog to carry a base_model block, because it is a derivative of another lab's open weights rather than an in-house pretraining run. Ornith-1.5 turns the self-scaffolding of Ornith-1.0 into a closed self-improvement loop: the model generates its own tasks, scaffolds and rollouts and uses them as reinforcement-learning signal. Vendor-reported at launch: Terminal-Bench 2.1 86.1 (Harbor/Terminus-2, 128K context, averaged over 5 runs), SWE-bench Verified 86.0 (OpenHands harness, 256K context) and DeepSWE 56.0, up from 8.0 for Ornith-1.0. No independent reproduction published at intake. The MIT grant covers the weights Ornith publishes; the upstream Qwen3.5 terms still govern the lineage.
Post-trained from Qwen3.5-397B-A17B (Alibaba). Ornith-1.0 was built on Qwen3.5 and Gemma 4 with continued pretraining, mid-training and post-training, mapped per variant as 9B on Qwen3.5 9B, 31B on Gemma 4 31B, 35B MoE on Qwen3.5 35B and 397B MoE on Qwen3.5 397B. Ornith-1.5 continues that line without the 31B — the one Gemma-based member — so the 1.5 lineup traces to Qwen3.5. The 397B is described as derived from Qwen3.5-397B-A17B, preserving its 397B/17B layout; the cloud entry alibaba-8 carries exactly that shape.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| SWE-bench Verified | 86% | vendor SWE-bench Verified 86.0 under the OpenHands harness (temp 1.0, top_p 0.95, 256K context). Era-1 benchmark, recorded only because Ornith reports it. Same transcription caveat as Terminal-Bench. |
| Terminal-Bench 2.x | 86.1% | vendor Terminal-Bench 2.1 86.1 via the Harbor/Terminus-2 harness, 128K context, averaged over 5 runs. One secondary transcription reports 85.1 instead; 86.1 is the figure repeated across the majority of launch coverage. The blogpost and the Hugging Face card are not fetchable from this environment (egress), so the value comes from consistent secondary transcriptions and was not independently reproduced. |
- GPQA Diamond: independent evaluation pending — Secondary transcriptions of the launch table cite GPQA Diamond 92.8 for the 397B but contradict each other across variants (the same sources report the 35B at 89.2, which is exactly the Qwen3.8-27B value in this catalog). The only reachable carrier is an aggregator. No value recorded until the model card is verifiable.
- SWE-bench Pro: independent evaluation pending — The launch table runs SWE-bench Verified, Pro and Multilingual under the same OpenHands harness, but the Pro figure (65.1 for the 397B, 59.6 for the 35B) surfaces only through an aggregator, not through the vendor table. To be added once the model card can be read directly.
- Humanity’s Last Exam (no tools): independent evaluation pending — An HLE figure of 44.6 (text-only, no tools) circulates in launch coverage, but the launch table itself is coding- and agent-focused (Terminal-Bench, SWE-bench family, DeepSWE, WideSearch) and no vendor-published HLE row could be confirmed. Recording it would move the open-weight LLM crown on a 1.1-point margin, so it stays out until the card is verifiable.
- MMLU-Pro: not reported by the vendor — Same coding-focused launch table.
- Terminal-Bench 3.0: independent evaluation pending — Coverage refers to a Terminal-Bench 3.0 row next to the 2.1 one, but no 3.0 number could be confirmed. Left null rather than guessed — a 3.0 value in terminal_bench would read as a collapse from 86.1.
Architecture and hardware
- Parameters
- 397B total, 17B active per token (MoE)
- Estimated VRAM at Q4
- ~229 GB, Multi-GPU Self-Host class
- Quantization formats
- GGUF, FP8, NVFP4, BF16
- Recommended runtime
- vLLM
- License
- MIT
Links
More from Ornith
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.