Bakbak Engineering · M1

How we engineered M1 for affordable, natural conversations

The architecture and inference engineering behind Bakbak M1, our text-to-speech model built for Indian voice agents.

Bakbak Engineering3 min read

Voice agents can only scale if the speech behind them is affordable at production volume. Yet high-quality TTS APIs are usually expensive, slow, or both, while the cheaper and faster alternatives tend to sound less natural and have inconsistencies in delivery.

We built M1 to remove that trade-off: natural, stable delivery in Indian languages, speech fast enough for a live conversation, and a cost that holds up at scale.

Hear M1 in a conversation

A voice has to hold up across an exchange: the opening question, a follow-up, a change of language, and a precise number read back. Hearing that sequence tells you more about the experience than a single polished sentence.

Talk to the Bakbak agent below. Ask what M1 costs, how quickly the first audio arrives, or switch to Hindi halfway through a question. Every reply is spoken by M1.

Prefer to listen first, or can't use a microphone right now? Switch to the recorded conversation and follow the transcript as it plays.

Talk to the Bakbak agent in your browser. Uses your microphone.

Shaped by nearly 20 million voice agent calls

Before M1, we built and ran voice agents. Across nearly 20 million calls, we worked with commercial TTS APIs and open-source models, and none of them met everything a production voice agent needs at once. Those calls defined the four things M1 had to get right.

  • Natural phrasingPauses, emphasis, and pacing that fit the reply.
  • Faithful deliveryThe intended words, names, and amounts.
  • Indian code-mixingSmooth delivery as languages change within a sentence.
  • Efficient generationSpeech that works economically at scale.

Compact by design

Saying a sentence well takes two separate jobs. The first is deciding how it should sound: where to pause, which words to stress, how quickly to speak, and how to move between languages. The second is producing the sound itself.

M1 gives each job its own component. The first reads the agent's reply and works out a compact plan for how to deliver it. The second turns that plan into speech.

Conceptual M1 architecture and runtime.

Splitting the work keeps the model small. Each component only needs to be as large as its own job requires, so no capacity is spent where it isn't needed, and less work per reply means lower cost.

It also makes the voice easier to get right. Because the delivery plan exists before any audio is generated, the synthesis step already knows the intended phrasing and emphasis. It doesn't have to work them out while producing sound. That gives us one clear place to improve how natural M1 sounds and how faithfully it says the words.

Both components run inside the Bakbak inference runtime. A compact model only stays cheap if the path through it is efficient too, so we designed execution, memory handling, and streaming alongside the model.

Fewer surprises in the middle of a reply

Sounding natural isn't enough if the words come out wrong. A dropped time, a repeated phrase, or an odd pause can derail a simple exchange. M1 is designed to keep the intended words intact while phrasing them naturally, including in sentences that mix Indian languages with English.

Below are four failure patterns we watch for. Play the M1 response to hear how it handles each one.

Previous user turn

पाँच बजे नहीं, साढ़े पाँच बजे चाहिए।

M1 synthesis text

जी, Saturday शाम 5:30 के लिए confirm कर देती हूँ।

Illustrated failure pattern

जी,Saturday शामMisplaced pause5:30 के लिएconfirm कर देती हूँ।

These examples explain the issues. Broader evaluation tells us how often they happen. We score content errors separately from pronunciation and prosody, so we can see which part of a reply needs work.

Making every inference count

The architecture sets the amount of work required to produce speech. The serving stack determines how efficiently that work runs on hardware.

We built Bakbak's inference runtime around short conversational replies, variable input lengths, and requests arriving throughout an active call. Our focus is useful throughput: serving more speech while keeping the time to first playable audio within a conversational latency budget.

From reference execution to the optimised runtime

Compare the same M1 checkpoint, inputs, and hardware across runtime configurations. Latency and throughput describe different parts of the serving result.

Runtime configurations · same checkpoint

Same model. Different runtime.

Configurations are cumulative: each row includes the changes above it.

Time to first playable audio

Milliseconds, p50 and p95. Lower is better.

  • p50
  • p95
  • Reference runtime160 / 310 ms
  • + compiled execution120 / 240 ms
  • + buffer reuse108 / 210 ms
  • + length bucketing90 / 170 ms
  • Optimised runtime82 / 135 ms

Throughput

Generated audio seconds per wall-clock second. Higher is better.

  • Audio s per s
  • Reference runtime6.8 s/s
  • + compiled execution8.7 s/s
  • + buffer reuse10.0 s/s
  • + length bucketing12.4 s/s
  • Optimised runtime15.0 s/s

Built for conversations at production scale

A voice that works for one call also has to work when many calls run at once. As traffic grows, every caller still expects the reply to start quickly and play without gaps.

So we test M1 under load. As concurrent requests increase, we track how long the first playable audio takes, whether playback stays continuous, and how much speech the same hardware can produce. We watch the slowest replies, not just the average, because those are the ones callers notice.

Load and useful throughput

More throughput within a conversational latency budget

  • Reference runtime
  • Optimised runtime
  • Hollow marker: over the latency budget

p95 time to first playable audio

Server timing, log scale. Lower is better.

Latency budget: 250 ms

501002505001,0002,500148163264Concurrent synthesis requestsmsLatency budget: 250 ms

Throughput

Generated audio seconds per wall-clock second. Higher is better.

010203040148163264Concurrent synthesis requestsaudio s per s

Written by Bakbak Engineering

Bakbak M1 · Pricing

Stop overpaying for voice

M1 does less work for every reply, so natural speech costs less at scale.

₹0.39

per 1,000 characters

₹390 per million characters

What would M1 change in your TTS bill?

Example inputs
M1 rate₹0.39 per 1,000 characters
Your current TTS cost₹19,500/month
Estimated M1 TTS cost₹3,900/month

Estimated monthly TTS savings

80% lower

₹15,600/month

Estimate based on the rates and volume entered. Actual billing depends on your plan and character-counting rules. Compare rates on the same tax basis.

Try M1 with your own text

Put M1 into your conversation

Try it on the replies your agent already says, and hear how it fits your product.

Talk to us about your workload