AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Ornith-1.5-35B-A3B

Ornith·

LLMsopen-weightEdge / Consumer

Overview

Mid-size member of the Ornith-1.5 family: 35B total / 3B active MoE under MIT, trained with the same self-improvement loop as the 397B flagship and shipped with GGUF, MLX (including a 4-bit build) and NVFP4 checkpoints, so a quantised build fits a 24 GB consumer GPU. Vendor-reported at launch: Terminal-Bench 2.1 67.8 (Harbor/Terminus-2, 128K context, averaged over 5 runs) and SWE-bench Verified 79.0 (OpenHands harness, 256K context). Like the rest of the family it is a post-trained derivative of open weights, not an in-house pretraining run; the MIT grant covers Ornith's own weights, the upstream base terms govern the lineage.

Post-trained from Qwen3.5 (Alibaba, family-level attribution). Ornith-1.0 was built on Qwen3.5 and Gemma 4 with continued pretraining, mid-training and post-training. Third-party reporting maps the 1.0 35B MoE onto Qwen3.5 35B, but Ornith publishes no per-variant mapping of its own, so the lineage is recorded at family level.

Capabilities and innovations

35B total / 3B active parameters (MoE)Agentic coding and terminal tasksGGUF, MLX and NVFP4 buildsRuns quantised on a single 24 GB GPUMIT license without regional restrictionsSelf-improvement loop at 3B active parametersMLX 4-bit build for Apple Silicon at release

Benchmarks

BenchmarkScoreSource
SWE-bench Verified79%vendor
SWE-bench Verified 79.0 under the OpenHands harness (temp 1.0, top_p 0.95, 256K context). Era-1 benchmark, recorded only because Ornith reports it. Same transcription caveat.
Terminal-Bench 2.x67.8%vendor
Terminal-Bench 2.1 67.8 (Harbor/Terminus-2, 128K context, 4-hour timeout, averaged over 5 runs) from the model card. One secondary summary reports 68.5. Card not fetchable from this environment; value taken from consistent secondary transcriptions.
  • GPQA Diamond: independent evaluation pending — Secondary transcriptions cite 89.2 for this variant — exactly the Qwen3.8-27B value in this catalog, which reads as a transcription conflation. No value recorded until the model card is verifiable.
  • SWE-bench Pro: independent evaluation pending — Reported in the launch table under the OpenHands harness; an aggregator lists 59.6 for this variant, but the vendor table itself is not reachable from this environment.
  • Humanity’s Last Exam (no tools): not reported by the vendor — Coding- and agent-focused launch table only.
  • MMLU-Pro: not reported by the vendor — Coding- and agent-focused launch table only.

Architecture and hardware

Parameters
35B total, 3B active per token (MoE)
Estimated VRAM at Q4
~21 GB, Edge / Consumer class
Quantization formats
GGUF, MLX, MLX-4bit, NVFP4, BF16
Recommended runtime
Ollama
License
MIT

Links

More from Ornith

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.