The rise of self-hosted AI: Global shift

As AI decentralises, organisations seek control, sovereignty, and efficiency beyond cloud dependence and centralised platforms
A quiet but significant shift is underway in the world of Artificial Intelligence. After years of relying on large, cloud-based models controlled by a handful of global technology companies, businesses, Governments, and academic institutions are increasingly turning towards self-hosted Large Language Models (LLMs).
This transition is not merely technological. It is economic, strategic, and geopolitical. Concerns over data sovereignty, rising operational costs, and global supply constraints are pushing organisations to rethink where and how AI should be deployed.
For years, AI was largely accessed through the cloud. Organisations depended on powerful remote servers owned by global technology companies (like AWS, Microsoft Azure and Google Cloud). While this model made AI widely accessible, it also created a quiet but significant dependency. Today, that dependency is increasingly being questioned.
Why is the Shift Taking Place
Several forces are driving this transition towards self-hosted AI systems.
While cloud-based AI platforms initially appeared economical, their long-term usage costs, especially through API-based pricing, have steadily risen. For universities, startups and even mid-sized enterprises, these recurring expenses are becoming difficult to sustain. Running models locally, despite upfront investment, often proves more predictable and cost-efficient.
Privacy, control and data sovereignty are also becoming important. Sensitive sectors such as defence, finance, healthcare, governance, and education are increasingly uncomfortable with sending data to external servers, often located outside national boundaries. Self-hosted models allow organisations to keep their data within their own infrastructure, ensuring greater security.
Technology is Catching Up
Until recently, self-hosting LLMs was impractical. These systems demanded enormous computing power, specialised hardware, and deep technical expertise.
That reality is now changing rapidly. A new generation of models is emerging that is more efficient, modular, and adaptable. This shift is being enabled by a growing layer of open tools, shared models, and collaborative platforms. Platforms such as Hugging Face have made thousands of models accessible to developers, researchers, and organisations, turning complex AI systems into downloadable, customisable building blocks, accelerating the move from cloud dependence to self-hosted intelligence. Models like Alibaba’s Qwen are demonstrating that high performance no longer requires exclusive access to hyperscale cloud infrastructure. At the same time, there is growing interest in smaller, more efficient models, often referred to as Small Language Models (SLMs), that can run on local servers or even high-end personal computers.
Closely related are advances in Vision-Language Models (VLMs), including compact variants called “Smol VLMs” which combine text and image understanding in lightweight formats. These models expand the scope of self-hosted AI beyond text, enabling document analysis, visual inspection, and multimodal applications.
The Role of Open Ecosystems
Unlike earlier phases of innovation dominated by proprietary systems, today’s AI landscape is more collaborative and accessible. Model repositories, shared datasets, and open research have enabled even smaller organisations to experiment with and deploy advanced capabilities.
This shift is not merely technical but strategic. By reducing dependence on a few dominant platforms, open ecosystems foster a more competitive and innovative environment and allow countries like India to participate not just as users of AI, but as creators.
Of course, organisations must still invest in talent, infrastructure, and governance frameworks to use these tools effectively. In that sense, the barriers to entry may be lower, but the responsibilities of ownership remain.
India’s Push for Sovereign AI
In India, this global shift is taking on a distinctly national character. The Government, through the IndiaAI Mission (led by MeitY), is actively supporting the development of a sovereign and proprietary AI ecosystem. The goal is clear: reduce dependence on foreign platforms, ensure data sovereignty, and build systems that reflect India’s unique linguistic and social realities.
Under this initiative, more than 20 indigenous foundational model proposals are being supported, including 12 LLMs and 8 SLMs. This represents one of the most ambitious national AI efforts globally.
Companies such as Sarvam AI have developed full-stack capabilities, including large conversational models, document-processing systems, and multi-dialect speech technologies.
Unlike Western models that prioritise English, Indian systems are being built to support all 22 scheduled languages. This linguistic depth is critical in India, where language remains one of the biggest barriers to digital access.
Another defining feature of India’s approach is its voice-first orientation. A large segment of the population prefers speaking over typing, especially in regional languages. As a result, AI systems are being optimised for low-latency voice interaction, making them more accessible and inclusive.
There is also a growing shift towards SLMs in India, driven by practical constraints. Limited access to high-end GPUs and the need for cost efficiency are encouraging organisations, especially in regulated sectors to adopt smaller, domain-specific models that can be deployed locally.
Importantly, these AI systems are being integrated with India’s Digital Public Infrastructure (DPI), enabling applications in governance, agriculture, healthcare, and public service delivery. This integration has the potential to transform how citizens interact with the state.
The Hardware Reality: Silicon Shortages and “Memflation”
While software capabilities are advancing rapidly, the hardware side tells a more complex story.
Says Pankaj Gulati, President, CDIL Semiconductors: ‘India is highly vulnerable to global semiconductor supply chains. It imports nearly all of its memory components and over 90% of its semiconductor requirements to satisfy its domestic demand. Lacking domestic wafer fabrication, global silicon shortages don’t hit local manufacturing directly. Instead, they choke the availability and drive up the costs of our critical imports.’
As global manufacturing increasingly shifts toward advanced AI-oriented processors, particularly GPUs, capacity for standard consumer and industrial chips has tightened.
One of the most visible outcomes is the rise in memory prices, particularly dynamic random-access memory (DRAM), resulting in “memflation”, a localized inflationary pressure driven by rising memory costs.
For Indian consumers and institutions, the impact is direct. Prices of laptops and tablets have risen consistently over recent months, while mobile devices have also become more expensive. High-performance computing systems, essential for running self-hosted AI models, are becoming costlier and harder to access.
Ironically, as organizations move towards self-hosted AI to reduce long-term costs and dependencies, they face short-term capital pressures due to rising hardware prices.
Implications for Education and Universities
The shift towards self-hosted AI has particularly important implications for educational institutions.
Universities, long dependent on cloud-based tools, now have an opportunity to build in-house AI capabilities. This not only reduces recurring costs but also enables greater experimentation, customization, and research independence.
Self-hosted models allow institutions to develop domain-specific applications, work with sensitive datasets without external exposure, train students on real-world AI deployment and encourage innovation in local languages.
However, this transition also demands investment in hardware, technical expertise, and infrastructure. Institutions will need to rethink their digital strategies, moving from being passive users of AI tools to active builders of AI systems.
A More Distributed AI Future
The broader implication of this shift is the emergence of a more distributed AI ecosystem.
Instead of intelligence being concentrated in a few global data centers, it is gradually spreading across organizations, institutions, and even individual devices. AI is moving closer to the “edge”, where data is generated and decisions are made.
This does not mean that cloud-based AI will disappear. Rather, a hybrid model is likely to emerge, where organizations balance cloud capabilities with local deployments depending on their needs.
For India, this hybrid future presents both an opportunity and a challenge. The opportunity lies in building systems that are locally relevant, inclusive, and sovereign. The challenge lies in overcoming infrastructure gaps, supply chain vulnerabilities, and skill shortages.
The shift towards self-hosted LLMs represents more than a technological trend, it signals a rethinking of how AI is built, controlled, and deployed.
For India, this moment is particularly significant. With strong policy backing, a vibrant startup ecosystem, and a clear focus on inclusivity, the country is well positioned to shape its own AI future. Yet, success will depend on balancing ambition with realism, investing in hardware, nurturing talent, and building resilient systems.
As AI becomes deeply embedded in everyday life, one thing is becoming clear: the question is no longer just what AI can do, but who owns it, where it runs, and whose needs it ultimately serves.
The author is an alumnus of IIM Ahmedabad and Professor of Practice at IILM University, Gurugram with an interest in AI, Technology and Strategy; Views presented are personal.















