Earnestness, often mistaken for naivety, is actually a sign of maturity and courage. It involves trusting your own experiences and insights, even when they contradict popular opinion. Founders should cultivate this mindset to navigate the startup landscape with confidence and authenticity, leading to more genuine and innovative solutions.
You become the ideas you spend time with.
Ideas worth knowing, from across business, technology and the world.
What are you learning today?
India’s Voice AI ecosystem at a glance
Several distinct layers power the Voice AI data value chain, from generating raw audio to validating model performance.
Where you sit in this chain, and how hard your position is to copy, shapes both revenue potential and long-term value.
- Sellers: Raw data originators
- Marketplaces: The Commoditised Middle
- Data Processors: The Jamnagar of AI
- Annotation & QA: The Truth-Making Layer
- Synthetic Data: Global Tailwind, Indian Caution
- Eval Providers: The McKinsey of Data
- Buyers & Deployers: The Infosys Layer
Embrace the Journey, Not Just the Goal
Many successful individuals regret not enjoying the journey. The process, with its challenges and growth, is often more rewarding than the end goal. By focusing on the present and finding joy in daily tasks, builders can enhance both their happiness and effectiveness, turning the journey into a fulfilling experience.
Synthetic answers lack real human evidence.
Most synthetic research systems combine large language models with demographic information, behavioral datasets, social media content, transaction data, or prior research. The model then produces responses that mimic what someone from a specific segment might say.
However, the answer still originates from a model.
There is no real customer behind the statement. There is no respondent whose situation can be examined. No interview recording exists to revisit, no behavior to observe, and no person for the researcher to question further.
A synthetic respondent might say:
“I would pay ₹2,499 for a premium protein supplement because I value clean ingredients.”
That sounds helpful. But what does it prove?
It does not show whether a real buyer will pay ₹2,499 when a competing product is available for ₹1,799. It does not reveal if the buyer will abandon the cart after seeing the delivery fee. It does not take into account advice from a gym trainer, distrust of an unfamiliar brand, concerns about taste, or a spouse questioning the monthly expense.
The model has created a plausible explanation, but it has not observed a purchase.
This makes synthetic data hard to use as evidence for high-stakes decisions. A company cannot confidently tell its board, product team, or investors that customers demanded a feature when no actual customers contributed to that finding.
Speech is where people actually perked up.
Saras V4 (speech-to-text) claims better coverage of lower-resource Indian languages and competitive English numbers. Bulbul V4 (text-to-speech) adds emotion and naturalness; several builders said the Hindi output is among the best they’ve heard. Vision 2.0 improves OCR on Indian handwriting and documents. There’s also Sarvam Code, local inference options, telephony tools, and talk of scaling Blackwell clusters plus a San Francisco office.
Pricing for people who actually ship:
- 105B: ₹4 input / ₹2.5 cached / ₹16 output per million tokens
- 30B is cheaper
- Speech-to-text: ₹30 per hour (₹45 with diarization)
- Text-to-speech: ₹15-30 per 10k characters depending on version
- Vision: ₹0.5 per page
If the quality holds in production, the economics for call centres, BFSI and government work look interesting. That’s the real hook.
Hobbling prevents AI from realizing its full potential
Many AI products and prompts inadvertently 'hobble' models by providing overly specific, step-by-step instructions, preventing them from utilizing their advanced problem-solving abilities. This creates 'product overhang,' where the model's true capabilities are not realized.
The solution is to give models harder, higher-level tasks, defining the desired outcome and constraints (guardrails and exit criteria) rather than dictating the process. This allows the model to 'cook' and find optimal solutions.
This framework is crucial for founders and builders because it shifts the paradigm from micro-managing AI to empowering it. By understanding and applying 'unhobbling,' they can design products that truly leverage the cutting-edge capabilities of models, leading to more innovative, efficient, and powerful solutions that might otherwise remain undiscovered.
Founders should rethink their prompt engineering strategies, moving away from overly prescriptive instructions towards defining clear objectives, constraints, and success metrics. This allows models to autonomously explore and execute complex tasks, potentially yielding superior and more creative solutions.
Leverage Unique Knowledge
Your unique knowledge and interests are your greatest assets in building a successful startup. Instead of asking what's hot, focus on what you know deeply and can uniquely contribute to. This approach can lead to groundbreaking innovations that are not immediately obvious to the broader market.
Sellers: Raw data originators
This is the supply side of the ecosystem. These players hold the most valuable raw material: real-world, domain-rich audio. Most, however, lack the infrastructure to monetise it.
- Enterprises & Call Centres
Millions of hours of transactional, support, and sales audio across industries - Vertical Players
Hospitals, BFSI firms, and retailers with rich operational audio data - Communities & NGOs
Agriculture, rural health, and grassroots organisations with low-resource language data - Domain Experts
Specialised contributors, similar to Mercor, bringing curated and labelled voice samples
- Enterprises & Call Centres
Freedom Fuels Productivity
Contrary to popular belief, freedom and productivity are not opposing forces. By optimizing for freedom, you can allocate time more effectively, focusing on what truly matters. This approach not only boosts happiness but also enhances productivity, as you're more engaged and present in tasks that align with your goals.
Plausibility is not the same as prediction.
Large language models are good at generating answers that a reasonable person might give.
Unfortunately, humans are not always reasonable.
People contradict themselves. They forget why they bought something. They claim to care about sustainability but then choose the product that arrives tomorrow. They say they want fewer notifications but continue opening apps designed for notifications. They demand privacy yet trade personal information for a small discount.
Research has shown that synthetic respondents may show less variation than real consumers and can exaggerate the relationship between demographics and attitudes. In other words, they can make consumer segments seem more consistent internally and more different from each other than real people actually are.
What people are actually questioning:
How much of the coding and agent stuff is real model progress versus a polished harness around other models? GLM keeps coming up in the side conversations. Were the benchmark slides selective? A few people noticed missing or conveniently ranked competitors.

Is Sarvam still trying to be a frontier model lab, or has it become an applied AI and infrastructure company that also trains models? The trillion-parameter plan sounds good on stage. Can they actually train and serve something competitive on the timelines and hardware they have, or does this become another ambitious slide?
Why aren’t more Indian product companies already deep on the 105B if the cost story is this strong? Latency, reliability at scale and basic ecosystem maturity still come up in private chats. After the capital raised and the government proximity, is the delivery matching the narrative?
One post that landed with people: the speech work is solid and didn’t need the questionable comparison graphs. Don’t spend the goodwill on theatre.
Enhanced AI safety enables more robust agent design
Opus 5 has largely overcome prompt injection vulnerabilities, a critical security flaw that previously allowed malicious external instructions to override a model's intended behavior. This enhanced safety means developers can build more reliable and secure AI agents.
Concurrently, models like Opus 5 demonstrate remarkable capabilities, such as rewriting entire codebases between languages in days, a task that previously required years of human effort.
For founders and developers, improved prompt injection resistance drastically reduces security risks, enabling the creation of more trustworthy and scalable AI products. The ability to automate complex engineering tasks like code migration offers immense value, accelerating product development and modernization efforts, and opening doors to previously unfeasible projects.
Developers should prioritize leveraging Opus 5's enhanced security for agent design and explore its capabilities for large-scale code transformation projects. This allows for faster iteration, reduced technical debt, and the ability to tackle ambitious engineering challenges with unprecedented efficiency.
The Importance of Co-Founders
Having co-founders can significantly enhance the startup journey. A single person might be seen as a lone visionary, but having a team validates and amplifies your vision. Co-founders bring diverse perspectives and can help navigate challenges more effectively, making the startup more resilient and adaptable.
Marketplaces: The commoditised middle
Most data marketplaces compete on volume and price. That quickly becomes a race to the bottom.
The ones that survive will need to build around three things:
- Eval Integration
Embed evaluation benchmarks directly into marketplace listings - Trust & Provenance Layer
Verified sourcing, consent trails, and licensing metadata - QA & Metadata Richness
Structured tags covering speaker demographics, noise levels, and domain labels
- Eval Integration
Resolve Conflicting Desires to Reduce Stress
Stress often arises from having conflicting desires. Identifying and resolving these conflicts can alleviate stress. Builders should prioritize one desire over another or decide to address it later, thus clearing mental clutter and enhancing focus. This clarity allows for more effective decision-making and reduces unnecessary anxiety.
This is a structural problem, not just an accuracy problem.
A synthetic consumer is built from patterns. A real consumer is shaped by constraints, relationships, habits, contradictions, and moments of irrationality. These elements often drive the purchase.
The melody incident serves as a cautionary tale.
In May 2026, a video featuring Indian Prime Minister Narendra Modi, Italian Prime Minister Giorgia Meloni, and Melody toffees gained attention on social media.
Retail investors then bought shares of Parle Industries, mistakenly associating the listed company with Melody. However, Melody is made by Parle Products, a separate and privately held company. Parle Industries had no ties to the chocolate.
The sequence was irrational but recognizable as human behavior:
- PM Modi gifts Melody to Italian PM.
- Parle makes Melody.
- A listed company has “Parle” in its name.
- Buy the stock.
Similar errors have occurred elsewhere. After Elon Musk tweeted “Use Signal,” investors inflated shares of Signal Advance, an unrelated medical device company. Investors also repeatedly confused Zoom Video Communications with the unrelated Zoom Technologies.
These are extreme cases, but they highlight the core problem.
Human decisions are influenced by availability bias, mistaken connections, social proof, the fear of missing out, and whatever stands out at that moment.
A model trained to create coherent behavior may consistently overlook incoherent behavior.
Consider a protein brand deciding on its next flavor.
Suppose a protein company must decide whether to launch mango, chocolate hazelnut, or unflavored whey.
A synthetic panel can analyze category trends, reviews, demographic preferences, and social media discussions. It might conclude that mango will attract young Indian consumers because it is familiar, culturally relevant, and different from existing chocolate products.
That is a reasonable guess.
But the actual purchase may rely on factors the model cannot grasp:
- Does mango whey taste artificial when mixed with water?
- Does its smell become unpleasant after resting in a shaker for an hour?
- Do consumers think of mango as a refreshing drink instead of a heavy protein product?
- Will gym trainers recommend it?
- Does the bright packaging make it appear less serious than competing products?
- Will customers enjoy the first serving but tire of the flavor after ten days?
These are not just data points. They are experiences.
A synthetic consumer can describe what consuming mango protein might feel like. It cannot actually taste it repeatedly, grow bored with it, regret buying a one-kilogram pack, or leave the half-used container at the back of a kitchen shelf.
This distinction matters because the company is not deciding which idea sounds best. It is deciding which product to manufacture, stock, distribute, and promote.
Sarvam occupies an important spot.
After other Indian efforts shifted focus, a lot of builders still want this one to work. The full-stack bet (models + speech + vision + inference + agents, India-first) is coherent.
Cost advantages in voice and local languages are not trivial. Sovereignty messaging hits differently when the alternative is shipping every conversation overseas.
At the same time the Indian AI conversation has grown up. People now ask harder questions about evaluation honesty, actual capability versus packaging, and whether capital is turning into durable technical edge. That’s healthy.
Epoch was neither a disaster nor a coronation. It was a serious company showing its current hand while reaching for a much bigger one. The next independent evaluations of the 105B, real production use of the voice stack, and visible progress on the trillion-parameter effort will matter more than any conference day.
Until then the questions stay open. As they should.
Playful experimentation uncovers AI's hidden talents
Models often possess capabilities that are not immediately obvious or were not explicitly part of their training data. Through creative experimentation and a willingness to 'play' with the model, users can discover these hidden talents.
This 'solicitation gap' highlights that the model's ability to perform a task depends heavily on how it is asked, not just what it was trained on.
For innovators, this means that the potential of AI models extends beyond their documented features. Fostering a culture of playful exploration can lead to groundbreaking discoveries and novel applications, giving early adopters a significant competitive advantage by unlocking unique functionalities.
Organizations should allocate resources and time for 'AI play days' or hackathons, encouraging engineers and product teams to experiment freely with models. This can lead to unexpected product features, efficiency gains, or entirely new business opportunities.
