Two models, one API. Pick the right balance of intelligence and speed for your use case.
Our most capable model. Advanced multi-step reasoning, precise instruction following, and stable long-context recall. Recommended for analysis, code generation, and agentic workflows.
Built for throughput. Streams at up to 60 tokens per second while keeping strong quality on everyday tasks — classification, extraction, chat, and drafting.
Full methodology is published with every technical report. Higher is better.
| Benchmark | Novo-4-Pro | Novo-4-Flash | Category |
|---|---|---|---|
| MMLU-Pro | 82.4 | 75.1 | Knowledge |
| GPQA-Diamond | 61.7 | 49.3 | Reasoning |
| SWE-bench Verified | 48.9 | 33.2 | Code |
| LongBench-v2 | 54.2 | 41.8 | Long context |
| LiveCodeBench | 39.6 | 28.4 | Code |
Scores from internal runs, Sep 2026. See the Novo-4 technical report for confidence intervals and eval details.
Start on the free tier — upgrade when you ship to production.