Claude Sonnet 4 vs Sonnet 5: What 64 Real-World Tests Actually Revealed
By Warren Schuitema, Founder | Matchless Marketing | The AI Dad
A new Claude model drops and the internet lights up with hot takes. Most of them are based on vibes, a few chat tests, and someone's gut feeling about whether the prose "sounds better."
Lenny Rachitsky did something different. He ran 64 structured generations across Sonnet 4 and Sonnet 5 to see whether the upgrade actually holds up. As someone who runs AI tools inside a live business every day, I read the whole thing. Here's what matters for you.
The Upgrade Is Real, But It's Not Uniform
Sonnet 5 outperformed Sonnet 4 on the majority of Lenny's 64 tests. The wins were clearest in two areas: complex reasoning tasks and instruction-following precision. When the prompt required the model to hold multiple constraints in mind simultaneously, Sonnet 5 handled it more consistently.
That matters for business owners. If you're using Claude to draft proposals, build structured outputs, or run multi-step research tasks, the improvement isn't theoretical. You'll feel it in fewer revision cycles.
But here's the thing: the gap wasn't uniform. On shorter, simpler tasks, Sonnet 4 and Sonnet 5 performed nearly identically. If your primary use case is quick email rewrites or short social copy, you're not leaving much on the table by staying on Sonnet 4.
The practical takeaway: Sonnet 5 earns its upgrade on complexity. Don't upgrade for simple tasks. Upgrade if your workflows involve multi-step reasoning, structured data, or outputs that need to get it right without three rounds of editing.
Instruction-Following Is Where It Actually Counts
One of the most useful findings in Lenny's test set was around instruction adherence. Sonnet 5 followed detailed, multi-constraint prompts more reliably than its predecessor.
This sounds technical. It isn't.
Think about the last time you gave Claude a detailed brief and it came back having ignored half your constraints. That's the problem Sonnet 5 genuinely reduces. In Lenny's tests, when he gave the model specific formatting rules, word count targets, tone requirements, and content constraints all at once, Sonnet 5 held them together better.
For a solopreneur or small business owner building any kind of repeatable AI workflow, this is significant. Your system prompts get more reliable. Your outputs need less manual cleanup. The ROI isn't in the output quality per se — it's in the time you stop spending fixing what the model missed.
Practical takeaway: If you run structured prompts with multiple requirements, test Sonnet 5 on your most failure-prone workflow first. That's where you'll see the clearest signal.
Where Sonnet 5 Still Falls Short
Lenny's review scores the model at a 4 out of 10. That number isn't a verdict on quality in isolation — it's a verdict on the value-to-cost ratio relative to what's already available.
The honest issue is that frontier models are converging. GPT-4o, Gemini 1.5 Pro, and Claude Sonnet 4 have all gotten close enough that the gap between any two of them, on most everyday tasks, is marginal. Sonnet 5 is better. But "better than already good" doesn't automatically justify a workflow overhaul.
There are also token cost implications if you're using the API. Sonnet 5 costs more per token than Sonnet 4. For high-volume workflows — think automated report generation or large-batch content production — that cost difference adds up faster than you'd expect.
Practical takeaway: Don't switch your entire stack. Run a controlled test on two or three of your highest-stakes workflows. If Sonnet 5 saves you two revision cycles per output, the cost difference probably pays for itself. If it doesn't, Sonnet 4 is still an excellent tool.
How to Actually Evaluate a Model Upgrade for Your Business
Most small business owners evaluate AI model upgrades the wrong way. They run one or two casual tests, form an impression, and either switch everything or dismiss it entirely.
Lenny's 64-generation approach is the right instinct applied at a scale most of us don't have time for. But the principle translates.
Pick three workflows where AI errors cost you the most time. Run each one five times on Sonnet 4 and five times on Sonnet 5 using the same prompt. Track how many outputs you use as-is versus how many need significant editing. That's your signal.
This isn't about which model "feels smarter." It's about which model reduces friction in the specific tasks you actually run. That's the only benchmark that matters for a working business.
If you're running Claude through Claude.ai directly, Sonnet 5 is already accessible on the Pro plan. If you're using the API, check your cost model before committing.
The One Thing to Do Today
Pull your single most-used AI prompt. The one you run weekly, or that your team uses constantly. Paste it into Claude Sonnet 5 and run it five times. Then do the same in Sonnet 4.
Don't evaluate which output "reads better." Count how many times each version required follow-up corrections to meet your actual requirements.
That number is your answer. Everything else is noise.
The AI landscape moves fast, and model upgrades feel urgent when they drop. Most of the time, your existing setup is already good enough. But occasionally, a targeted upgrade on a high-friction workflow saves you hours every month.
Sonnet 5 might be that upgrade for you. It might not be. Now you know how to find out.