Why India’s AI mission is funding the wrong half

Compute take over 40 per cent of the IndiaAI budget. It is the one input the world was always going to make cheap for us
Think of building an artificial intelligence system as cooking a meal.
You need two things. A stove, and ingredients.
The stove is computing power — the specialised chips, called GPUs, that do the arithmetic. The ingredients are data: the text, speech, images and records that the machine learns from. Without the stove, nothing cooks. Without ingredients, there is nothing to cook.
India’s national AI programme, the IndiaAI Mission, was approved by the Cabinet in March 2024 with an outlay of Rs 10,371.92 crore over five years. It has seven pillars. It has bought a great many stoves. More than 38,000 GPUs have been brought on board, and they are offered to Indian startups and researchers at a subsidised rate of about Rs 65 an hour — a genuinely low price, and one that has already been used by Indian teams building their own models.
Now here is the number that should make us stop and think. An analysis by Bharath Reddy of the Takshashila Institution, using the pillar-wise split placed before the Lok Sabha in December 2024, puts compute subsidies at over 40 per cent of the mission’s total allocation. The other pillars, he notes, do not appear to be getting much funding attention at all. The single largest share of India’s AI money is going into the stove.
I want to make an argument that runs against almost everything written about this programme, including most of the praise and most of the criticism. The problem is not that the mission spends too little. The problem is what it spends on.
One ingredient gets cheaper on its own
Consider what has happened to the price of computing.
Chips get faster. Software gets more efficient. Competitors enter. Older equipment becomes cheap to rent. Every year, the amount of money needed to reach a given level of AI capability falls - and it falls sharply. This is not a prediction or an opinion. It is the most reliably observed trend in the entire industry, and it has held for decades across computing generally.
Governments have a standard test for whether to subsidise something. Would this happen anyway, without public money?
If yes, the subsidy is mostly buying you a few years of speed.
Compute passes that test in the wrong direction. It is the one input in this business whose price is already collapsing without anyone’s help. There is a global industry of enormous companies competing furiously to make it cheaper. They will do this whether or not India pays them.
Meanwhile the GPUs themselves age. A chip bought today is a depreciating asset. Within a few years it is outdated, within eight or so it is close to scrap. We are buying a perishable good whose replacement will cost less than the original.
The other ingredient is running out
Now look at the ingredients.
India’s real advantage in this field was never chips. We do not make them. Our advantage is that we are the most linguistically rich large country on earth. Twenty-two scheduled languages. Hundreds more that are spoken but not scheduled. Dialects that change every few hundred kilometres. Centuries of legal, medical, agricultural and literary tradition in scripts most models have barely seen.
If an Indian AI is ever going to be better than a foreign one at anything, it will be here.
And this ingredient is not getting cheaper. It is getting scarcer, for three reasons, and each of them is a one-way door.
The first is the internet itself. The open web is now filling rapidly with text written by machines. That matters more than it sounds. When a model learns from text produced by other models, the unusual gets thinned out first — the rare word, the regional usage, the minority language, the local remedy known in one district. What survives is the average. So the value of scraping the web is falling at exactly the moment we are scaling up to scrape it. Human-written Indian text is becoming a smaller and smaller fraction of what is out there, and the honest way to describe this is that we are eating our seed corn.
The second is demography. A great deal of India’s knowledge has never been written down at all. It lives in speech, in people who are old, in villages where the young are switching to a bigger language. Every year some of that goes quietly out of existence. You cannot buy it back later. There is no price at which a language returns once its last fluent speakers are gone.
The third is legal. As data protection rules tighten worldwide - properly so - the cost and difficulty of assembling large collections of human material rises. What was free to gather in 2015 requires consent, contracts and compliance in 2026.
So: one input falls in price every year and will keep falling without us. The other rises in price every year and is disappearing permanently.
We have put over 40 per cent of the mission behind the first.
What this means in practice
Let me be fair. The mission is not blind to data. It runs a datasets platform, AIKosh, meant to pool public and anonymised datasets so Indian builders do not have to scrape everything themselves. That is the right instinct, and it exists.
But instinct is not allocation. The compute pillar has the money and the headlines. The data work has neither. And the mission’s yearly funding has itself been squeezed — last year’s Rs 2,000 crore budget saw actual utilisation nearer Rs 800 crore, and the allocation for the coming year stands at around Rs 1,000 crore. When the pot shrinks, the biggest line item defends itself. It is the small, slow, unglamorous work that gets postponed.
There is a further irony worth naming. Bhashini, the Government’s own language platform, covering all 22 scheduled languages, reportedly handled around 2.5 billion API calls in 2026. Every one of those calls is a human being bringing an Indian-language sentence to a machine. That is one of the largest live streams of Indian-language interaction anywhere in the world. It is being run as a service. Its most durable value is as a record.
What to do instead
I am not arguing that India should stop subsidising compute. Researchers need machines, and a startup that cannot afford a GPU cannot begin.
I am arguing about the marginal rupee — the next one we spend. Send it to the ingredient that will not be there later.
That means paying people to record: fieldwork in languages with no written corpus, with consent, with payment, with the recordings held publicly. It means digitising the archives we already own — court records, land records, agricultural extension notes, All India Radio’s tapes, university theses — which cost nothing to create and are rotting for want of scanning. It means settling, clearly and fairly, who owns the data trail that public platforms generate and under what consent it may be used. And it means treating all of this as a national reserve, in the way we treat foodgrain or foreign exchange: a stock we build in good years because it cannot be conjured in bad ones.
None of this requires frontier technology. It requires clerks, microphones, scanners, lawyers and patience. It is the least exciting proposal in Indian technology policy, and that is precisely why nobody is making it.
Here is the sentence I would like to leave with the reader. A GPU bought in 2026 will be scrap by the middle of the next decade. A recording of an eighty-year-old speaking her mother tongue in 2026 will be beyond price in 2086, and there will be no way to make another one.
We are spending our largest share on the thing the world was going to hand us cheaply, and leaving to chance the one thing only India could ever have collected.
The author is a physicist at the University of North Carolina at Chapel Hill and a columnist on AI, infrastructure and global systems; Views presented are personal.















