AI Models Are About to Get a Lot Cheaper to Run. Here's What That Means for Your Business.
By Warren Schuitema, Founder | Matchless Marketing | The AI Dad
The AI tools you're using every day cost what they cost because the models behind them are enormous. Running a large language model isn't free. Somebody pays for that compute, and right now, a lot of that cost gets passed to you through API pricing, subscription tiers, and rate limits that cut you off right when you're in the middle of something useful.
That's starting to change. And the reason it's changing is worth understanding, even if you never intend to touch a line of code.
What Researchers Just Published (And Why It's Not Just for Tech People)
A company called Multiverse Computing published research on Hugging Face today, September 21, 2026, that describes a new way to shrink large AI models without wrecking what makes them useful.
The technique is called block removal. Instead of trimming weights inside a model the way you'd thin out a hedge, it deletes entire structural sections called transformer blocks. Think of it like pulling floors out of a skyscraper rather than chipping away at the walls. Done wrong, the building falls. Done right, you've got a leaner structure that still holds.
The smart part is how they decide which floors to pull. The researchers, led by physicist Román Orús, mapped the problem onto a framework from statistical physics called an Ising model. It treats each block as a binary choice: keep it or delete it. Then it evaluates thousands of combinations to find the configurations where the model stays sharp.
The results on one benchmark were striking. At 50% compression of Llama-3.3-70B-Instruct, their method outperformed other block-removal approaches by nearly 23 percentage points on the MMLU benchmark. They also tested it on Llama-3.1-8B-Instruct and Qwen3-14B.
To be fair, there are caveats. The candidate solutions were still selected manually in this round, and the method's advantage was measured on MMLU specifically. Other benchmarks told a more mixed story. This is research, not a shipping product.
But the direction is clear.
The Real Story: Inference Costs Are the Problem You're Already Paying For
You don't have to care about physics to care about what this research is pointing toward.
Right now, running a capable AI model costs real money. Multiverse Computing's broader platform, CompactifAI, already claims 50–80% reductions in inference costs for compressed models, with speed improvements of 4x–12x. That's not a rounding error. That's the difference between a tool you can run continuously in your workflow and one you ration because it's too expensive to leave on.
If you've ever hit a token limit mid-task, downgraded to a cheaper model tier because the good one was too expensive, or avoided connecting an AI agent to a data-heavy process because the API costs got unpredictable, you've already felt this problem firsthand.
Cheaper inference doesn't just mean lower bills. It means more capable tools become viable for small business operators who don't have an IT budget.
What This Means for How You Build Your AI Stack Today
You don't need to wait for compression research to mature before making smarter decisions. This news is a signal, and signals are actionable.
First, build your workflows on model behavior, not model names. The specific model running your automation today might get swapped for a smaller, faster, cheaper version in six months. If you've hard-coded your prompts and processes around "ChatGPT-4o" or "Claude Sonnet," you're going to feel that transition. Build around what the model needs to do, and test with alternatives regularly. I do this across every workflow I run at Matchless Marketing.
Second, watch the open-source model space closely. The models being compressed in this research, Llama-3.1-8B and Llama-3.3-70B, are open-source. Smaller, cheaper, locally-runnable versions of powerful models are exactly what makes self-hosted or private AI setups viable for businesses that handle sensitive client data. If you've been waiting for a reason to explore local model deployment, the window is opening.
Third, don't over-invest in API-dependent infrastructure right now. If you're building an n8n workflow or a Make automation that hits a paid API on every run, understand that the pricing landscape for those calls could shift significantly over the next 12–18 months. Build flexibility in. Keep your prompts modular. Document which parts of your workflow depend on specific model capabilities so you can swap without rebuilding from scratch.
None of this requires a developer. It requires thinking about your AI stack the way you'd think about any operational cost: with an eye on where the price is heading, not just where it is today.
The Takeaway for Small Business Owners
The physics in this research is genuinely interesting if you're into that kind of thing. The Ising model. Constrained binary optimization. A co-founding quantum physicist applying particle physics to AI infrastructure. It's a good story.
But here's what actually matters to you at 6am: the gap between "AI tools enterprises use" and "AI tools you can afford" is shrinking. The research coming out today is part of what's closing it.
Smarter compression means smaller models with less accuracy loss. Smaller models with less accuracy loss means lower compute costs. Lower compute costs means pricing pressure on the tools you already subscribe to, and new options that weren't cost-effective six months ago.
You can't control when these models ship. You can control whether you've built your workflows to adapt when they do.
One thing you can do today: open whatever AI workflow you rely on most and write down exactly which model it depends on, what it costs per month, and what would break if you swapped to a smaller or cheaper model. That inventory takes 15 minutes. Most people don't have it. The ones who do will move faster when the landscape shifts.
That's the job. Not to predict the future, but to be ready for it.
Warren Schuitema is the founder of Matchless Marketing and the creator of The AI Dad, a brand and platform helping small business owners implement AI tools without hype, overwhelm, or a developer on retainer. He builds, tests, and documents real AI systems live so his audience can follow what actually works inside a running business. He is the operator behind a fully automated AI agent workforce managing content, leads, research, and client onboarding at Matchless Marketing. Warren's perspective on AI compression research is rooted in the practical question every operator should be asking: what does this cost to run, and when does it get cheaper?
Keep Learning
- How to Build an n8n Workflow That Doesn't Break When Models Update, Why model-agnostic prompt design saves you rebuild time
- ChatGPT vs. Local Models: What Small Business Owners Actually Need to Know, The real trade-offs without the hype
- AI Agents for Small Business: Where to Start, A plain-language primer on building your first automated workflow
PUBLISHER CHECKLIST
- Set the SEO title tag (provided below)
- Set the meta description (provided below)
- Apply FAQPage schema if FAQ section is added at publish
- Add "Last Updated: September 2026" to the byline
- Source a header image: suggested file name
ai-model-compression-small-business.jpg, alt text: "Diagram showing AI model compression and cost reduction for small business AI tools" - Confirm all internal links resolve; add fresh links to newer related content
- Schedule companion social posts
- Set a 12-month refresh reminder