Independent benchmark · Published 48h after release

DeepSeek V4 Pro 0813: 87% cheaper, but how much slower?

We ran 5 standard tasks + 20 latency tests against GPT-5.2 and Claude 4.5. Full data and code included.

No spam. Unsubscribe anytime.

Why DeepSeek V4 Pro Matters Right Now

When DeepSeek V4 Pro 0813 dropped on August 13, 2026, the Hacker News thread exploded to 671 upvotes and 236 comments within hours. The immediate reaction wasn't excitement — it was skepticism. "Benchmarks?" "How does it compare to GPT-5.2?" "Is the price real?" Developers know that a new model announcement means little without reproducible data. The official docs only show cherry-picked metrics, and independent platforms like Artificial Analysis take 3-7 days to update their leaderboards.

That gap is exactly where we step in. Existing solutions fail because they're either too slow (Artificial Analysis), too subjective (LMArena voting), or too biased (vendor blogs). We built a repeatable benchmark pipeline that runs 5 core tasks — code generation, math reasoning, long-context understanding, JSON output stability, and Chinese language capability — plus 20 latency tests. Every result is published with the exact code and raw data, so you can verify our claims in minutes.

Why now? Because the 48-hour window after a major release is the only time when this information is genuinely scarce. The HN thread shows 807 cross-platform discussions in one day. Teams are making routing decisions today, not next week. Our report gives you the data you need before you commit to a switch — or before you confidently stay put.

How It Works

01

Run Standardized Benchmarks

We execute 5 identical tasks across DeepSeek V4 Pro, GPT-5.2, and Claude 4.5 via the OpenRouter API. Each task is repeated 20 times to measure variance and ensure statistical reliability.

02

Measure Cost & Latency

We record per-million-token pricing and time-to-first-token for every call. This gives you a real cost-performance ratio, not just a single score.

03

Publish Full Data + Code

All raw results, the benchmark script, and the analysis are open-sourced. You can reproduce every number in this report or adapt the code for your own workloads.

What You Get

Reproducible Data

Every benchmark result is backed by the exact Python script and raw JSON output. No black boxes, no cherry-picked numbers. You can re-run the entire evaluation on your own machine in under an hour.

48-Hour Advantage

Published within 48 hours of the official release. You're seeing the first independent analysis — before the big platforms catch up. This is the earliest reliable signal available anywhere.

Actionable Switch Guide

We cut through the noise: specific scenarios where DeepSeek V4 Pro is a clear win, and where you should stick with GPT-5.2 or Claude 4.5. Practical advice based on cost, latency, and task type.

Full code open-sourced 20 latency tests per model No vendor sponsorship