Notes · Journal शोध पत्र

Notes from the work.

Thoughts on language intelligence, learning and building AI that feels useful in the many contexts of India.

01 · Language intelligence

Building AI for India’s many voices

Why useful intelligence begins with context, culture and the freedom to speak naturally.

21 July 2026 · 5 min read

India is home to 22 scheduled languages, hundreds of dialects and a daily reality where most conversations blend two or three tongues in a single breath. When someone in Lucknow asks a question, they might start in Hindi, drop in an English term, reference a local government scheme by its colloquial name, and expect the answer to understand all of it. That is not an edge case — it is the norm.

Building AI that works for this reality means rethinking what “understanding language” actually requires. It is not enough to translate. The system must read meaning, context and intent as a single layer — and respond in a way that feels native to the person asking.

Meaning lives beyond words

Language is not a simple exchange of one word for another. Meaning also comes from region, rhythm, shared references and the situation around a conversation. The Hindi spoken in Varanasi carries different registers than the Hindi of Delhi. A Tamil question about “ration card” invokes a specific bureaucratic context that an English translation strips away.

AI made for India has to begin there — not with syntax, but with the web of cultural and situational knowledge that gives words their weight. When a farmer in Maharashtra asks about crop insurance in Marathi mixed with English acronyms, the system needs to recognise the scheme, the season, the region and the crop — not just parse the sentence.

This is why BhaaratMind treats context as a first-class input, not a post-processing step. Every query carries implicit information — geography, season, institutional knowledge, conversational history — and the model is trained to read all of it before it answers.

Natural expression matters

People should not have to reshape a thought for a machine. They should be able to ask naturally, move between languages mid-sentence and still receive an answer that respects the original intent. This is how India already communicates — a fluid, code-mixed exchange where the speaker chooses whatever word fits best, regardless of which language it belongs to.

Most AI systems today force a choice: pick one language, type it correctly, keep your query clean. That friction is invisible to English-first users, but it excludes hundreds of millions of people who think and speak in patterns that do not map to a single language dropdown.

BhaaratMind follows meaning across 22+ Indian languages and code-mixed speech. The goal is not to detect which language someone is using — it is to understand what they are saying, however they choose to say it. The system treats mixed input as natural input, not as noise to be cleaned.

Understanding should lead to action

Context and voice matter because they make outcomes more useful. The goal is not only to recognise language, but to help a person learn, decide and move forward with confidence. An answer that is technically correct but requires three more searches to act on has not really answered the question.

Every response BhaaratMind generates ends with a clear next step. For a student, that might be a focused quiz on the chapter they just asked about. For a parent navigating school admissions, it might be a checklist of documents with the deadlines that apply to their specific board and state. The intelligence is only as useful as the action it enables.

This principle shapes everything from how we rank information to how we structure responses. Clarity and usefulness are not features we add at the end — they are constraints we design around from the start.

02 · Product notes

Designing calm AI for learning

What changes when a study companion turns progress into the next clear action instead of more noise.

21 July 2026 · 4 min read

Most education technology adds complexity. More features, more notifications, more dashboards with numbers that a student has to interpret on their own. The result is a tool that demands attention instead of reducing the effort required to learn.

Manthan AI is built on a different premise: a study companion should make the next step obvious, keep the interface calm, and never make a student feel behind. These are not cosmetic choices — they are structural decisions about how information is organised, what is shown and what is deliberately left out.

Clarity before complexity

A student already has enough to process — textbooks, class notes, exam schedules, the anxiety of a subject they find difficult. A learning product should reduce the cognitive load required to understand what is complete, what needs attention and what to do next.

In Manthan AI, every screen answers one question at a time. The chapter view shows where you stand. The quiz view focuses on one concept. The progress snapshot tells you what improved and what did not. There is no dashboard that tries to show everything at once, because “everything at once” is the opposite of clarity.

This extends to the way answers are structured. When a student asks a question, the response leads with the direct answer, follows with the reasoning, and ends with what to do next. No preamble, no hedging, no walls of text that bury the useful part three paragraphs deep.

Progress should be honest

Useful feedback is specific without being overwhelming. A percentage score after a quiz tells you very little. What matters is which concepts you understood, which ones you guessed on, and what to revisit before the exam.

Manthan AI’s progress snapshots break performance down by topic, not by test. If you scored 70% on a chapter quiz, the snapshot shows you that you have a strong grip on chemical equations but are consistently getting atomic structure questions wrong. That specificity makes the next step actionable — you know exactly what to study, not just that you need to “study more.”

Honest progress also means not inflating scores or celebrating streaks for their own sake. The interface does not gamify learning with badges or points. When you are behind, it says so plainly and offers a catch-up plan. When you are ahead, it moves on to the next chapter without fanfare. The goal is trust, not engagement metrics.

Guidance should feel human

Catch-up guidance works best when it feels supportive rather than punitive. A student who has fallen behind does not need a red warning bar — they need a plan. Calm interfaces, familiar language and clear recommendations help effort turn into insight.

Manthan AI speaks in the tone of a patient tutor, not an automated system. When it suggests reviewing a chapter, it explains why in one sentence. When it generates a quiz, it frames it as practice, not as a test. These are small choices in language and design that compound over weeks of use.

The bilingual interface — English and Hindi, with more languages coming — is part of this philosophy. A student in a Hindi-medium school should not have to mentally translate instructions written in English before they can start learning. The tool should meet them where they are, in the language they think in.

03 · Evaluations

Measuring answers in the language asked

Accuracy in English rarely predicts accuracy in Hindi or Tamil — so we score every language on its own terms.

21 July 2026 · 6 min read

When an AI system reports 90% accuracy, the natural question is: accurate for whom? Most benchmarks are built in English, tested on English queries, and scored against English ground truth. A model that performs well on those benchmarks may still fail on common Hindi idioms, Tamil-specific references, or the kind of code-mixing that happens in everyday Indian conversation.

At BhaaratMind, we evaluate every language on its own terms. Not by translating English test sets, but by building evaluation frameworks rooted in how each language is actually used — with its own idioms, cultural references and structural patterns.

One benchmark does not fit all

The standard approach to multilingual evaluation is to translate an English benchmark into other languages, run the model, and compare scores. This method is fast, but it hides a fundamental problem: translated questions do not test the same knowledge. A question about property tax in English carries different assumptions than the same question asked in Telugu by someone navigating a municipal corporation office.

Translation preserves words but not context. A benchmark that asks “What documents are needed for school admission?” in translated Hindi does not capture the real query, which might reference a specific state board, use the colloquial name for the RTE quota, or mix in English terms like “transfer certificate.” Real accuracy requires test sets built from real queries in each language.

We build evaluation sets by sampling actual questions from Indian users, in the language and style they were originally asked. These sets include code-mixed queries, regional variations, and questions that reference local institutions, schemes and cultural knowledge.

Evaluating in context

A correct answer depends on where the question comes from. The same words carry different weight in a classroom, a government office, or a family kitchen. “How do I apply?” means something entirely different when asked about a college admission, a ration card, or a passport renewal — and the right answer for each depends on the state, the year, and the person’s eligibility.

Our evaluation framework considers context as part of correctness. A response is scored not just on whether the facts are right, but on whether the answer is complete enough for the person asking to act on it. An answer about CBSE exam dates that does not mention the relevant circular or where to check for updates is technically correct but practically incomplete.

We also test for cultural appropriateness — whether the tone, formality and assumptions in the response match what a user in that context would expect. A system that answers a grandmother’s health question in clinical English jargon has not really served her, even if the medical facts are accurate.

Closing the gap language by language

We publish per-language scores openly and use them to decide where to focus next. The goal is not a single aggregate number but honest visibility into how well the system serves each voice. If Hindi accuracy is strong but Kannada lags behind, that gap drives the next round of data collection, fine-tuning and evaluation.

This approach is more expensive and slower than a single translated benchmark, but it produces a system that actually works for the people it is meant to serve. A 92% score averaged across languages hides the fact that some communities are getting 98% accuracy while others get 78%. Per-language scoring makes those gaps visible and actionable.

Our evaluation reports are structured by language, domain and query type. Each report shows where the model excels, where it struggles, and what we are doing about the gaps. This transparency is part of the product — we believe the people using the system deserve to know how well it works in their language, not just in aggregate.

04 · Field notes

Built for the everyday phone

Latency, data and small screens shape a useful answer as much as the model does. Notes from building for Bharat.

21 July 2026 · 3 min read

The average Indian smartphone costs under ₹10,000, runs on a mobile data plan where every megabyte is budgeted, and connects through networks that range from solid 4G in cities to patchy 2G in rural districts. This is not a constraint to work around — it is the design brief.

Building AI for India means the product has to work on the phone people already own, on the network they already have, within the data budget they already manage. If the experience degrades on a Redmi Note on Airtel 3G, it does not work for India.

Designed for real networks

Most of India is not on fast broadband. Answers need to arrive quickly on 2G and 3G connections, which means smaller payloads, progressive rendering, and responses that are useful even before they finish loading. A streaming response that shows the first sentence in 800 milliseconds is more useful than a complete answer that takes four seconds to appear.

We measure performance not from a data centre in Mumbai, but from real devices on real networks across tier-2 and tier-3 cities. Our latency targets are set against 3G connections in Jhansi, not fibre connections in Bangalore. If the product feels slow there, it is slow.

This constraint also shapes how we build the backend. Model inference is optimised for time-to-first-token, responses are streamed progressively, and heavy assets are deferred or eliminated entirely. The interface loads with the minimum viable payload and hydrates as bandwidth allows.

Small screens, full answers

A six-inch screen is the primary window to information for hundreds of millions of people. Layout, typography and interaction patterns must work at phone scale first, not as an afterthought. This means answers are structured vertically, touch targets are generous, and navigation never requires horizontal scrolling or pinch-zooming.

Typography matters more on small screens, not less. We use type sizes that remain readable without zooming, line lengths that prevent re-reading, and spacing that separates content blocks without wasting vertical space. Every pixel of a 720p display is accounted for.

Interactive elements — quiz options, chapter navigation, progress cards — are designed as tappable blocks with clear boundaries, not as text links that require precision. On a crowded bus or in a dimly lit room, the difference between a 44px touch target and a 28px one is the difference between a usable product and a frustrating one.

Respecting the data budget

Data costs money. Every unnecessary byte — an oversized image, a redundant script, a response that rambles — is a cost the user pays. Building for Bharat means treating bandwidth as a design constraint, not as a resource you can assume is abundant.

Manthan AI’s web client loads in under 200KB on first visit. Images are served in modern formats at the exact resolution the device needs. Fonts are subsetted to the characters actually used. Analytics are minimal and batched. Every dependency is questioned: if a library saves development time but adds 40KB to the bundle, we write the feature by hand instead.

This discipline extends to the AI responses themselves. A concise, well-structured answer uses less data to transmit, loads faster, and is easier to read on a small screen. Brevity is not a style choice — it is a performance feature.