Dapols
PlansAI employeesAI integrationPricingToolsBenchmarksBlog
Sign inFind my AI workflow
Dapols

The job big companies pay a forward-deployed engineer six figures to do — as a tool, for two figures.

AI price watch — free weekly email

Which AI tools changed price, and what launched — for small businesses. Verified from vendor pages. No spam, unsubscribe anytime.

Product

  • AI Deployment Plans
  • AI Plan Finder
  • AI skills library
  • AI tools
  • Pricing

Offerings

  • AI Deployment Plans
  • AI integration
  • Bigger or more complex? Tell us.
  • Bigger or more complex?

Company

  • Blog
  • AI model tracker
  • Submit a tool
  • Contact
  • Privacy
  • Terms of Service

Trust & methodology

  • About
  • Methodology
  • How we rank AI tools
  • Forward-deployed AI
  • Affiliate disclosure
  • AI tool pricing updates
  • Security & data privacy

© 2026 Dapols. All rights reserved.

support@dapols.comX

Put AI to work—one measurable workflow at a time.

Dapols
All articles

English · 7 min read

Kimi K3 costs half what GPT-5.6 Sol does — should your business switch?

Moonshot's Kimi K3 is a frontier-class open-weight model at $3/$15 per million tokens. What that means if you pay per token, what "open weights" actually buys you (licence caveat included), and when switching is the wrong move.

The short answer: if you pay per token for an API in your automations, K3 is the first open model worth pricing out seriously — roughly half the cost of GPT-5.6 Sol and about a third of Claude Fable 5, at close-enough quality for most business work. If you pay a flat monthly seat price for ChatGPT, Claude, or a tool that bundles AI, this changes nothing for you today.

Facts below verified 29 July 2026 against Moonshot's own launch post, VentureBeat's launch coverage, VentureBeat's follow-up on the weights release, and Business Insider. Model prices move; check before you commit.

What launched

On 16 July 2026, Chinese lab Moonshot AI released Kimi K3: a 2.8-trillion-parameter model with a 1-million-token context window and native visual understanding (VentureBeat). It is a sparse mixture-of-experts design that activates just 16 of its 896 experts on any given pass (Moonshot). Moonshot says it is the largest open-source model in the world; Business Insider describes it as the largest open-weight model announced to date (Business Insider).

The full weights landed on 27 July, on schedule — with a licence caveat we'll come back to.

Every outlet framed this as "China catches up." That framing is fine and mostly useless if you run a small business. The number that matters to you is the price.

The price, plainly

ModelInput / 1M tokensOutput / 1M tokens
Kimi K3$3$15
GPT-5.6 Sol$5$30
Claude Fable 5~$10~$50

Prices per Business Insider's launch coverage, which gives Anthropic's figures as approximate. Read the ratios carefully, because the headlines round them badly: against Sol, K3 is about half the cost, not a third. Against Fable 5 it is roughly a third. Two further details push it cheaper still:

  • Cached input drops to $0.30 per million — a tenth of the cache-miss rate. If your automation sends the same long instructions, price list, or policy document on every call (most do), that repeated chunk is billed at $0.30 instead of $3.00 (Moonshot, VentureBeat). This is where the real saving lives, not in the headline rate.
  • Caching is automatic — no cache ID, no TTL, no extra parameter, unlike competitors that make you manage it explicitly (VentureBeat). One less thing for whoever wired up your automation to get wrong.

And K3 is Moonshot's expensive model. Their K2.7 Code and K2.6 both run $0.95 in / $4 out — worth knowing, because most small-business jobs (summarising, drafting, classifying, tagging) do not need a frontier model at all.

Is it actually good, or just cheap?

Two separate questions, and the honest answer differs.

Moonshot's own claims — treat these as the vendor's marketing, not settled fact. In its launch post the company reports frontier-level results across its evaluation suite while conceding that K3's overall performance "still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol," and that it shows "a noticeable gap in user experience" against both. A vendor volunteering that is worth more than the benchmark chart above it.

Third-party numbers are more interesting, and more mixed — all via VentureBeat, drawn from public leaderboards and an evaluation by analytics firm Artificial Analysis:

  • On Arena.ai's Frontend Code Arena, where humans blind-compare outputs without knowing which model produced them, K3 took #1 with a score of 1,679, ahead of Fable 5 and GPT-5.6 Sol. Blind human preference is the hardest number to game.
  • On Artificial Analysis's GDPval-AA, which covers real-world tasks across 44 occupations, K3 scored 1,687 — third, behind Claude Fable 5 Max (1,815) and GPT-5.6 Sol Max (1,747.8).
  • On AA-Briefcase, a private long-horizon knowledge-work benchmark, it came second at 1,527, ahead of Sol Max (1,495) and behind Fable 5 Max (1,587).

So: genuinely at the frontier on some work, a clear step behind on the hardest reasoning. Which is roughly what "half the price" should buy.

Worth hearing the sceptics, too, both via Business Insider. Vercel CEO Guillermo Rauch called the Arena result "the first time that an open model is ahead of all proprietary ones for this comprehensive web engineering benchmark" — but cautioned that "benchmarks don't always tell the full story." Wharton professor Ethan Mollick called K3 the "closest to the frontier yet" while likewise advising against relying on headline scores alone.

"Open weights" — do you actually care?

Almost certainly you will never self-host this. Running a 2.8T-parameter model needs a rack of accelerators and someone to babysit it; that is not a six-person business, and it isn't a fifty-person one either. So the honest answer to "should I care that the weights are open?" is: not for the reason the announcement implies.

Here is what it does get you:

  1. Price pressure on the model you already use. A frontier-class open model at $3/$15 is a number every paid lab now has to answer to. You benefit from that whether or not you ever send K3 a single request.
  2. Provider choice for the same model. When weights are public, multiple hosts serve them — you can move between providers without changing the model, so a price hike or an outage at one vendor isn't a rewrite.
  3. A ceiling on your worst case. If the model is public, no vendor can retire it out from under an automation you depend on.

And here is the catch. "Open" is not the same as "unrestricted": Moonshot's published licence requires companies above roughly $20 million in annual revenue to negotiate separate terms, and is materially more permissive for purely internal use than for building a product on top (VentureBeat). Most readers of this post sit well under that line — but if you plan to resell anything built on K3, that licence is a lawyer question, not a blog question.

What open weights also do not get you is "free." The weights cost nothing; running them costs plenty. That trap is the same one we cover in when cheap AI models beat ChatGPT — free-to-download is not free-to-operate.

Should you switch? A straight answer

This matters to you if you're paying per token through an API — an automation that drafts replies, a script that tags support tickets, a workflow in Make or n8n, a Zapier AI step, anything a developer wired up for you that bills by usage. Your bill scales with volume, so halving the rate roughly halves that line item.

This does not matter to you today if you pay a flat monthly seat price — ChatGPT Business, Claude seats, or AI baked into your CRM or helpdesk. You do not buy tokens; you buy seats. A cheaper model beneath the surface may eventually lower that price or raise the limits, but there is no switch for you to flip this month. Ignore the news and get on with your work.

If you are in the first group, do this:

  1. Find your actual token spend before anything else. Pull last month's API invoice. If it's under about the price of a couple of lunches, the saving isn't worth an afternoon of migration risk — go work on something that moves revenue.
  2. Move one low-stakes job first. Ticket tagging, first-draft summaries, categorisation. Not anything customer-facing, not anything that touches money.
  3. Compare outputs side by side on your own inputs. Your data is the only benchmark that counts — Arena scores don't know your product catalogue.
  4. Check where the data goes. K3 served from Moonshot's API means requests to a Chinese provider. Depending on your sector, clients, and jurisdiction, that may be a straightforward no regardless of price — or a non-issue. Decide it deliberately, in writing, before you migrate anything.
  5. Keep the expensive model for the hard 10%. Cheap model for volume, frontier model for the work that must not be wrong. That split usually saves more than any single switch.

The takeaway

The interesting thing about K3 is not that a Chinese lab reached the frontier. It's that frontier-adjacent quality now costs roughly half of frontier prices, and the gap keeps closing. If you locked in an AI vendor eighteen months ago and haven't re-priced since, you are quite likely overpaying — not because you chose wrong, but because the floor keeps dropping. Re-pricing your AI stack is now just part of running the business, the same way you'd re-quote insurance. Our guide to auditing your AI subscriptions walks through how.

If you'd rather not track model launches at all: our free 2-minute AI plan finder and our business AI plans start from what you already run and what you actually spend, and tell you where a cheaper model would help — and where switching would just cost you a weekend.

Get the weekly AI price watch

One short email a week on AI tool pricing changes for small businesses.

Get your AI plan

Your best tools, quick wins, and budget — in two minutes.

Take the quiz