AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Ornith-1.5-9B

Ornith·

LLMsopen-weightEdge / Consumer

Overview

Smallest member of the Ornith-1.5 family: a dense 9B under MIT with a 262K native context and a quantised mobile build for iPhone and Android. GGUF checkpoints range from about 2.8 GB to 17.9 GB, so an 8 GB card runs a Q4/Q5 build and 24 GB runs BF16; on Apple Silicon the same builds run through Metal via MLX. Vendor-reported at launch: Terminal-Bench 2.1 47.0 (Harbor/Terminus-2, 128K context) and SWE-bench Verified 70.6 (OpenHands harness, 256K context). Ornith's size-comparison claims are size-adjusted — the 35B MoE of the same family still scores higher on both figures.

Post-trained from Qwen3.5 (Alibaba, family-level attribution). Ornith-1.0 was built on Qwen3.5 and Gemma 4 with continued pretraining, mid-training and post-training. Third-party reporting maps the 1.0 9B onto Qwen3.5 9B and puts the Gemma 4 lineage in the 31B, which Ornith-1.5 does not continue; Ornith publishes no per-variant mapping of its own, so the lineage is recorded at family level.

Capabilities and innovations

9B dense parameters262K native contextAgentic coding and terminal tasksGGUF and MLX builds from ~2.8 GBQuantised mobile build (iOS/Android)MIT license without regional restrictionsSelf-improvement loop applied at 9B denseMobile-quantised build shipped at release

Benchmarks

BenchmarkScoreSource
SWE-bench Verified70.6%vendor
SWE-bench Verified 70.6 under the OpenHands harness (temp 1.0, top_p 0.95, 256K context). Era-1 benchmark, recorded only because Ornith reports it. Same transcription caveat.
Terminal-Bench 2.x47%vendor
Terminal-Bench 2.1 47.0 (Harbor/Terminus-2, 128K context) from the model card. Card not fetchable from this environment; value taken from consistent secondary transcriptions and not independently reproduced.
  • SWE-bench Pro: independent evaluation pending — Reported in the launch table under the OpenHands harness but not transcribed in accessible coverage.
  • GPQA Diamond: not reported by the vendor — Coding- and agent-focused launch table only.
  • Humanity’s Last Exam (no tools): not reported by the vendor — Coding- and agent-focused launch table only.
  • MMLU-Pro: not reported by the vendor — Coding- and agent-focused launch table only.

Architecture and hardware

Parameters
9B total (Dense)
Estimated VRAM at Q4
~6 GB, Edge / Consumer class
Quantization formats
GGUF, MLX, BF16
Recommended runtime
Ollama
License
MIT

Links

More from Ornith

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.