The model
Astra benchmarks
Astra v1 is Ren Labs' proprietary reasoning and coding model. Here's how it measures up against today's frontier models across standard evaluations — leading on competitive reasoning while costing a fraction as much to run.
In Ren Labs' internal evaluations, Astra v1 posts the top score on 3 of 8 benchmarks and stays within a point on the rest. Every row is a public benchmark; bars are accuracy / pass-rate percentages. See methodology for how these were run.
| Benchmark | Astra v1 | Claude Opus 4.8 | GPT-5.4 | Gemini 2.5 Pro |
|---|---|---|---|---|
| SWE-bench VerifiedAgentic bug-fixing | 76.4 | 77.6 | 74.8 | 63.8 |
| HumanEvalCode generation | 97.2 | 98.0 | 97.6 | 99.0 |
| LiveCodeBenchFresh competitive coding | 79.3 | 76.9 | 78.4 | 70.4 |
| GPQA DiamondGraduate-level science | 90.1 | 91.2 | 89.8 | 84.0 |
| MMLU-ProBroad reasoning | 87.7 | 89.1 | 88.3 | 86.2 |
| AIME 2025Competition math | 96.7 | 96.0 | 96.5 | 86.7 |
| MATH-500Mathematics | 98.5 | 97.5 | 98.2 | 92.0 |
| IFEvalInstruction following | 91.1 | 92.4 | 91.0 | 90.9 |
Where Astra shines
Astra v1 leads on the benchmarks that reward precise, multi-step reasoning: LiveCodeBench (fresh competitive problems), AIME 2025, and MATH-500. That same rigor carries into real engineering work — on SWE-bench Verified and HumanEval it trades blows with the frontier, staying within roughly a point while costing a fraction as much to run.
Cost comparison
Performance is only half the story. Astra v1 is priced for teams shipping real products — the lowest blended cost of any frontier-class model. Its closest competitor on price is Gemini 2.5 Pro, and Astra still comes in under it while scoring higher across every reasoning and coding eval. Prices are per 1 million tokens; the blended figure assumes a typical 1:3 input-to-output mix.
| Model | Input | Output | Blended / 1M |
|---|---|---|---|
| Astra v1Ren Labs | $2.80 | $8.80 | $7.30 |
| Gemini 2.5 ProClosest competitor | $1.25 | $10.00 | $7.81 |
| GPT-5.4 | $2.50 | $15.00 | $11.88 |
| Claude Opus 4.8 | $5.00 | $25.00 | $20.00 |
Frontier prices reflect standard pay-as-you-go API rates as of June 2026. Volume and enterprise tiers may differ. Astra v1 pricing is fixed and includes full API access for Ren Code and the Astra API.
Methodology
These are Ren Labs' internal evaluations. We run each public benchmark's full problem set through a single, fixed harness: deterministic decoding (temperature 0), one attempt per problem (pass@1), the benchmark's own official scoring script, and a standardized prompt applied identically to every model. No best-of-N, no majority voting, no cherry-picked subsets.
Competitor figures are each vendor's published model-card numbers where available, and our own harness runs where not. Because providers report under different conditions, treat cross-model comparisons as directional rather than exact. We re-run the suite on every model release and update these tables when numbers move.