Benchmark Overview
This Haiku 5.5 benchmark guide covers published performance evidence, not our own inference tests. A score is meaningful only with its task set, model variant and evaluation conditions.
Officially Reported Results
Developer-reported results below are from Anthropic’s announcement. Competitor scores in that table are also Anthropic-reported, not an independent comparison by AI Model Brief.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna |
|---|---|---|---|
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% |
| FrontierCode 1.1 Main | 46.4% | Not reported | 42.4% |
| Humanity’s Last Exam, no tools | 45.9% | 10.2% | Not reported |
| Humanity’s Last Exam, with tools | 57.4% | 18.7% | Not reported |
Coding Performance
Terminal and repository tasks test different workflows. These published coding results do not predict success on your own codebase; use held-out tasks, record tool permissions and review patches.
Reasoning Performance
Keep tool-assisted and unaided results separate. Access to tools changes the evaluation setup, so the two Humanity’s Last Exam rows should not be treated as the same test.
Speed and Latency
Artificial Analysis reports 243.4 output tokens per second for its Haiku 5.5 (Max) page at review time. This is third-party output throughput, not end-to-end application latency. Prompt processing, thinking, queueing and tool calls can change the time a user waits.
Comparison With Previous Haiku Models
The official table above supplies previous-generation values where published. “Not reported” is a missing value, never zero. Keep the benchmark version and provider-reported conditions attached to any comparison.
Independent Benchmark Results
Artificial Analysis labels the following variants separately. These are its Intelligence Index v4.3.2 scores, not percentages and not this site’s measurements. The page-level settings matter; do not merge these into a single model score.
| Third-party model variant | Intelligence Index score |
|---|---|
| Claude Haiku 5.5 (Max) | 43 |
| Claude Haiku 5.5 (High) | 38 |
Methodology and Limitations
Research date: October 8, 2026. We read published results and did not run inference, latency or coding evaluations. Cross-provider ranking remains sensitive to reasoning budget, harness, prompts, tools and test revisions. This page does not establish matched experimental conditions across every row.
Frequently asked questions
Are these AI Model Brief benchmark measurements?
No. Official rows are developer-reported; Artificial Analysis rows are independent third-party results. We did not execute an inference benchmark.
Is there a separate Haiku 5.5 benchmark page for singular searches?
No. Benchmark and benchmarks refer to the same guide and evidence here.
Go straight to the source
Reviewed Oct 8, 2026. Official statements, third-party results and arithmetic estimates are attributed separately. Unknowns reflect the sources reviewed on this date. No inference benchmarks were run by this site.