How WhatsApp service message pricing hits your chatbot bill

From 1 October 2026 your chat agent's replies are billable, and the free tier counts message bubbles, not conversations.

An abstract graphic of stacked message bubbles drawing down a fixed monthly allowance

Say you run a three-clinic physiotherapy group in Abu Dhabi. Most new patients arrive on WhatsApp, at 11pm, from a phone in bed. Your chat agent asks which clinic, asks what hurts, offers three slots, confirms one, then says thanks.

Six messages out. One booking. Nobody has ever costed those six, because free-form replies inside the 24 hour window have been free since November 2024.

WhatsApp service message pricing changes that on 1 October 2026: the free-form replies your chat agent sends inside WhatsApp's 24 hour customer service window become billable, and so do utility templates sent inside that window. Every business phone number gets 1,000 free service messages a month, and the meter counts individual messages, not conversations. An agent that answers in four bubbles spends four of them.

How WhatsApp service message pricing works from 1 October 2026

A service message is any non-template message you send a customer inside the 24 hour customer service window. Plain text, an image, a PDF, a voice note, a menu. The window opens the moment a customer messages you, and it resets every time they message again.

Three dates matter. Free-form replies became free in November 2024. On 1 July 2025 Meta moved template billing from whole conversations to individual messages. On 1 October 2026 service messages join them, priced at the same per-message rate as utility and authentication templates in the recipient's market, with no volume tiers to grow into. Utility templates sent inside the window lose their free status on the same day.

Rates are set by the country of the customer's phone number, not yours, and they differ sharply between markets. Meta publishes the per-market rate card on its WhatsApp Business Platform pricing page. Look up the countries your customers actually message from before you do any arithmetic.

One clarification, because it catches people out. All of this applies to the WhatsApp Business Platform, the API your chat agent runs on. If you and two colleagues answer messages by hand in the free WhatsApp Business app, none of it touches you.

How far do 1,000 free service messages actually go?

Not nearly as far as the number sounds. A thousand a month is about 33 messages a day per number, and a chat agent that takes six messages to book an appointment burns through that in five or six conversations.

Run it against three bot styles:

  • A tight agent averaging 4 outbound messages per conversation gets roughly 8 conversations a day inside the free tier.
  • A typical agent at 6 messages gets about 5.
  • A chatty agent at 10, with a greeting bubble, an acknowledgement, then an answer split across three bubbles, gets 3.

Three conversations a day is a quiet Tuesday at a single dental practice, not a busy operation. The allowance resets monthly and does not roll over, so a slow August buys you nothing in September.

Two things count against it that people forget. Replies typed by a human after handoff are service messages too, so everything your front desk sends inside the window draws on the same 1,000. And a group send draws one unit per delivered recipient, not one for the send.

A diagram showing a single customer message opening a 24 hour window, and the business replies inside it drawing down a monthly allowance of 1,000 free service messages.
A diagram showing a single customer message opening a 24 hour window, and the business replies inside it drawing down a monthly allowance of 1,000 free service messages.

Back to the physiotherapy group. Share one WhatsApp number across all three clinics and every reply, from every site, drains one allowance. Run a number per clinic and each gets its own 1,000.

That is a genuine lever and also a trap: three numbers means three inboxes, three sets of history, and a patient whose conversation from last year sits on a number nobody is watching this year. Split numbers when the sites really do operate separately. Not to game an allowance.

Which of your chat agent's messages are billable?

Every message your side delivers inside the window, including several you have probably never counted as answers.

  • The greeting bubble. "Hi, thanks for messaging the clinic."
  • The AI disclosure line, if it goes out on its own.
  • An answer deliberately split into three bubbles to feel more human.
  • The "are you still there?" nudge after two minutes of silence.
  • "Got it, one moment while I check."
  • Every reply a person sends after the handoff.

The nudge is worth singling out. Say you run a six-van boiler repair firm in Sheffield, and your agent pings anyone who goes quiet mid-booking. Half of those people put the phone down because the kettle boiled. They come back. The nudge bought you nothing and cost a message, twice a conversation, all winter.

The disclosure line is a different problem, because it sits on top of a legal duty. Article 50 of the EU AI Act has applied since 2 August 2026, and it requires that people interacting with an AI system are told so, clearly, and no later than the first interaction. The European Commission's guidance on those transparency obligations sets out the duty.

The compliance answer and the cost answer happen to agree: fold the disclosure into the first substantive reply instead of sending it alone. "You're chatting with our AI assistant. Which clinic is closest to you?" is one message and does both jobs.

That duty binds you for users in the EU. We tell clients to disclose everywhere anyway. A customer who works out halfway through that the friendly agent was software trusts the next thing you say a good deal less.

How do you cut billable messages without making the bot worse?

Change the shape of the conversation, not the quality of the answers. The work runs in roughly this order:

  1. Count before you change anything. Pull 30 days of outbound messages per number from your provider's dashboard and divide by the number of conversations. If the average is under four, stop here. You do not have a problem.
  2. Kill the acknowledgement bubbles. "Got it, one moment" is a message you pay for that carries no information. Answer, then ask the next question, in one reply.
  3. Replace the question ladder with a form. WhatsApp Flows render an interactive multi-step form inside the chat, so name, clinic, injury and preferred slot arrive in one interaction instead of eight messages of ping-pong.
  4. Move qualifying questions to where they are free. A conversation started from a Click to WhatsApp ad keeps a 72 hour free entry point window. A form on your own site costs nothing at all, and if those submissions already land in your CRM (the customer record system your team books from), the agent can open with what it knows instead of asking again. That is the same plumbing we describe in the first process worth automating in a small team.
  5. Hand off earlier when the bot is clearly stuck. Two more rounds of deflection is four more messages and a worse outcome. Speed of reply decides who gets the job anyway, which is the argument in why lead response time decides who gets the job.
  6. Send one confirmation, not three. Date, time, clinic, address and cancellation link, in a single message.

Steps 2 and 3 do most of the work, and neither makes a single answer shorter. The bot says the same things. It just stops saying them in instalments.

When is this not worth fixing?

If you send fewer than 1,000 service messages a month per number, this changes nothing for you, and you should do nothing about it.

Here is the part that costs us money to say. Rebuilding a conversation flow properly is not a quick job. A Flow that collects five fields, validates them, posts to a booking system and handles the failure cases is a couple of days of build plus real testing on real handsets, Android and iOS, because Flows do not render identically on both.

If your bill after 1 October is going to be a few dollars a month, do not hire anyone to save it. Not us, not anyone. Pay the few dollars.

There is a softer limit too. Fewer bubbles can read colder. When someone messages at midnight because they are in pain, or furious about a missed appointment, one dense paragraph is cheaper and worse than two warm ones. Spend the extra message there on purpose.

What is likely to change after 1 October 2026?

Meta has started pricing the work behind the message rather than the message itself. On 1 August 2026 its own Business Agent, the AI that answers on a business's behalf, moved from per-message billing to token-based billing, which bundles the AI processing and the delivery into one rate. That is a different unit of account from everything above, and it points at where the rest is heading.

Two practical consequences. Per-market rate cards get republished, so the country mix of your customers matters more than it used to. A Sheffield firm that starts pulling enquiries from the Gulf will watch its number move for reasons that have nothing to do with its bot. And whatever the next change bills, it will bill something you are not currently measuring.

So measure now. A monthly record of messages per conversation, per number, turns the next pricing announcement into a five minute calculation instead of a fortnight of guessing.

What to do before 1 October

Open your provider's dashboard. Pull last month's outbound message count for each WhatsApp number, divide by 30, and see how close to 33 you already sit. That one number tells you whether 1 October is an accounting footnote or a real cost.

If it is a real cost and you would rather not redesign the conversation yourself, our chat agent work starts exactly there, with the message count. Tell us what your agent currently says and we will tell you which of those messages you can stop paying for.

Common questions

Still wondering

Do messages customers send me count against the 1,000?

No. The allowance and the charge both apply to messages your business delivers to the customer inside the 24 hour window. Inbound messages are free and always have been, which is why a bot that asks lots of short questions looks cheap until you count your own replies. Count outbound only, per phone number, per calendar month.

Does the free allowance apply per business or per phone number?

Per business phone number. Each number registered on the WhatsApp Business Platform gets its own 1,000 free service messages each month, the count resets at the start of the month, and unused messages do not roll over. Running several numbers gives you several allowances, but it also splits your conversation history and your team's inbox across them.

Can I dodge the charge by replying outside the 24 hour window?

No, and trying costs more. Outside the window you cannot send a free-form reply at all. You have to send an approved template, which is billed per message in the marketing, utility or authentication category, and marketing templates are the dearest of the three. Replying quickly inside the window stays the cheaper path.

What about templates I send outside the window?

They are billed per delivered message, as they have been since 1 July 2025, at the rate for the category and the country of the customer's phone number. What changes on 1 October 2026 is that utility templates sent inside the 24 hour window stop being free, so the in-window discount that existed for fifteen months disappears.