GPT-5.4 Thinking
OpenAI·
LLMsreasoningcloud
Overview
Multi-step reasoning model with native computer-use and 1M token context.
Capabilities and innovations
Multi-step reasoningNative computer use1M token context (API)Tool-heavy workflowsAgentic tasksNative Computer UseTool SearchEfficient Token Usage28-point OSWorld Jump
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 92.8% | unsourced |
| Humanity’s Last Exam (no tools) | 42% | unsourced |
| SWE-bench Pro | 57.7% | unsourced |
| MMLU-Pro | 87.48% | unsourced |
| Terminal-Bench 2.x | 75.1% | unsourced |
| MMMU-Pro | 81.2% | unsourced |
- SWE-bench Verified: not reported by the vendor
- MMLU: not reported by the vendor
Links
More from OpenAI
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.