Sarvam AI: Here is all that was announced, claimed and debated

Is Sarvam a foundational model company or an applied AI that also trains models?

Idea 02 of 04

Idea 02 of 04

Speech is where people actually perked up.

All ideas

Introduction

Bengaluru got another big AI day this week. Sarvam used Epoch to push its full-stack story harder: trillion-parameter model in the works, upgrades to the 105B, better speech models, Vision 2.0, coding agents, India-hosted inference, even more smartglasses demos. Ambitious, loud, and very on-brand for a company that has positioned itself as India’s main sovereign AI bet.

The reaction has been a mix of real interest and quiet eye-rolling.

The claims in short: they’re building a trillion-plus parameter model from scratch in India, focused on coding, cybersecurity, science and simulation. Roughly six-month timeline according to the messaging.

Current 105B is being sold as roughly $0.80 per million blended tokens, which they say is 5.5 times cheaper than GPT-5.4 Mini and about 11 times cheaper than Gemini 3.5 Flash. Voice is the big flex. They keep saying if you’re building voice products for India right now, nothing is cheaper or more scalable (we do have questions on latency).

Speech is where people actually perked up.

Saras V4 (speech-to-text) claims better coverage of lower-resource Indian languages and competitive English numbers. Bulbul V4 (text-to-speech) adds emotion and naturalness; several builders said the Hindi output is among the best they’ve heard. Vision 2.0 improves OCR on Indian handwriting and documents. There’s also Sarvam Code, local inference options, telephony tools, and talk of scaling Blackwell clusters plus a San Francisco office.

Pricing for people who actually ship:

  • 105B: ₹4 input / ₹2.5 cached / ₹16 output per million tokens
  • 30B is cheaper
  • Speech-to-text: ₹30 per hour (₹45 with diarization)
  • Text-to-speech: ₹15-30 per 10k characters depending on version
  • Vision: ₹0.5 per page

If the quality holds in production, the economics for call centres, BFSI and government work look interesting. That’s the real hook.

What people are actually questioning:

How much of the coding and agent stuff is real model progress versus a polished harness around other models? GLM keeps coming up in the side conversations. Were the benchmark slides selective? A few people noticed missing or conveniently ranked competitors.

Is Sarvam still trying to be a frontier model lab, or has it become an applied AI and infrastructure company that also trains models? The trillion-parameter plan sounds good on stage. Can they actually train and serve something competitive on the timelines and hardware they have, or does this become another ambitious slide?

Why aren’t more Indian product companies already deep on the 105B if the cost story is this strong? Latency, reliability at scale and basic ecosystem maturity still come up in private chats. After the capital raised and the government proximity, is the delivery matching the narrative?

One post that landed with people: the speech work is solid and didn’t need the questionable comparison graphs. Don’t spend the goodwill on theatre.

Sarvam occupies an important spot.

After other Indian efforts shifted focus, a lot of builders still want this one to work. The full-stack bet (models + speech + vision + inference + agents, India-first) is coherent.

Cost advantages in voice and local languages are not trivial. Sovereignty messaging hits differently when the alternative is shipping every conversation overseas.

At the same time the Indian AI conversation has grown up. People now ask harder questions about evaluation honesty, actual capability versus packaging, and whether capital is turning into durable technical edge. That’s healthy.

Epoch was neither a disaster nor a coronation. It was a serious company showing its current hand while reaching for a much bigger one. The next independent evaluations of the 105B, real production use of the voice stack, and visible progress on the trillion-parameter effort will matter more than any conference day.

Until then the questions stay open. As they should.

All ideas

  1. 01Introduction
  2. 02Speech is where people actually perked up.
  3. 03What people are actually questioning:
  4. 04Sarvam occupies an important spot.

Showing Speech is where people actually perked up., idea 2 of 4.