AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Ornith-1.5-397B

Ornith·

LLMsopen-weightMulti-GPU Self-Host

Overview

MIT-licensed flagship of the Ornith-1.5 family from the research lab DeepReinforce. 397B total / 17B active MoE, post-trained from Qwen3.5-397B-A17B and keeping the base model's parameter layout unchanged — the first entry in this catalog to carry a base_model block, because it is a derivative of another lab's open weights rather than an in-house pretraining run. Ornith-1.5 turns the self-scaffolding of Ornith-1.0 into a closed self-improvement loop: the model generates its own tasks, scaffolds and rollouts and uses them as reinforcement-learning signal. Vendor-reported at launch: Terminal-Bench 2.1 86.1 (Harbor/Terminus-2, 128K context, averaged over 5 runs), SWE-bench Verified 86.0 (OpenHands harness, 256K context) and DeepSWE 56.0, up from 8.0 for Ornith-1.0. No independent reproduction published at intake. The MIT grant covers the weights Ornith publishes; the upstream Qwen3.5 terms still govern the lineage.

Post-trained from Qwen3.5-397B-A17B (Alibaba). Ornith-1.0 was built on Qwen3.5 and Gemma 4 with continued pretraining, mid-training and post-training, mapped per variant as 9B on Qwen3.5 9B, 31B on Gemma 4 31B, 35B MoE on Qwen3.5 35B and 397B MoE on Qwen3.5 397B. Ornith-1.5 continues that line without the 31B — the one Gemma-based member — so the 1.5 lineup traces to Qwen3.5. The 397B is described as derived from Qwen3.5-397B-A17B, preserving its 397B/17B layout; the cloud entry alibaba-8 carries exactly that shape.

Capabilities and innovations

397B total / 17B active parameters (MoE)Agentic coding and terminal tasksEvaluated at 256K context (SWE-bench harness)GGUF, FP8 and NVFP4 buildsMIT license without regional restrictionsSelf-improvement loop: the model writes its own tasks, scaffolds and RL rolloutsPost-trained derivative of open weights (Qwen3.5-397B-A17B) at unchanged parameter countDeepSWE jump from 8.0 (Ornith-1.0) to 56.0

Benchmarks

BenchmarkScoreSource
SWE-bench Verified86%vendor
SWE-bench Verified 86.0 under the OpenHands harness (temp 1.0, top_p 0.95, 256K context). Era-1 benchmark, recorded only because Ornith reports it. Same transcription caveat as Terminal-Bench.
Terminal-Bench 2.x86.1%vendor
Terminal-Bench 2.1 86.1 via the Harbor/Terminus-2 harness, 128K context, averaged over 5 runs. One secondary transcription reports 85.1 instead; 86.1 is the figure repeated across the majority of launch coverage. The blogpost and the Hugging Face card are not fetchable from this environment (egress), so the value comes from consistent secondary transcriptions and was not independently reproduced.
  • GPQA Diamond: independent evaluation pending — Secondary transcriptions of the launch table cite GPQA Diamond 92.8 for the 397B but contradict each other across variants (the same sources report the 35B at 89.2, which is exactly the Qwen3.8-27B value in this catalog). The only reachable carrier is an aggregator. No value recorded until the model card is verifiable.
  • SWE-bench Pro: independent evaluation pending — The launch table runs SWE-bench Verified, Pro and Multilingual under the same OpenHands harness, but the Pro figure (65.1 for the 397B, 59.6 for the 35B) surfaces only through an aggregator, not through the vendor table. To be added once the model card can be read directly.
  • Humanity’s Last Exam (no tools): independent evaluation pending — An HLE figure of 44.6 (text-only, no tools) circulates in launch coverage, but the launch table itself is coding- and agent-focused (Terminal-Bench, SWE-bench family, DeepSWE, WideSearch) and no vendor-published HLE row could be confirmed. Recording it would move the open-weight LLM crown on a 1.1-point margin, so it stays out until the card is verifiable.
  • MMLU-Pro: not reported by the vendor — Same coding-focused launch table.
  • Terminal-Bench 3.0: independent evaluation pending — Coverage refers to a Terminal-Bench 3.0 row next to the 2.1 one, but no 3.0 number could be confirmed. Left null rather than guessed — a 3.0 value in terminal_bench would read as a collapse from 86.1.

Architecture and hardware

Parameters
397B total, 17B active per token (MoE)
Estimated VRAM at Q4
~229 GB, Multi-GPU Self-Host class
Quantization formats
GGUF, FP8, NVFP4, BF16
Recommended runtime
vLLM
License
MIT

Links

More from Ornith

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.