Benchmarks
Measured, not promised
Run on WebArena with the same models and the same tasks. Benchmarked with GPT-5-mini and Qwen3-30B — no frontier model required.
| Metric | AlohaJet | Chrome DevTools MCP | Playwright MCP |
|---|---|---|---|
| Task success rate (WebArena) | 88% | 51% | 11% |
| Average time per task | 39s | 73s | 84s |
| Average tokens per task | 158k | 340k | — |
All numbers are from the WebArena benchmark, pooled across GPT-5-mini and Qwen3-30B runs — 10 repetitions per task. Want us to benchmark AlohaJet on your workloads? Request a demo →