Together AI Builds India's Largest 10,000-GPU Open-Source AI Factory

L&T and Together AI are deploying 10,000 NVIDIA B300 GPUs in Chennai, marking India's largest single-cluster AI factory in a deal worth up to $1.8B.

·
·
Together AI Builds India's Largest 10,000-GPU Open-Source AI Factory
Read7 min
TopicGpus · Business
SubtopicAccelerators
  • Largest AI cluster in India: L&T and Together AI are deploying 10,000 NVIDIA B300 GPUs at Vyoma's Chennai data centre, India's biggest single-cluster AI infrastructure.
  • Deal value: Classified as a "Mega" order by L&T, valued between Rs 10,000–15,000 crore (~$1.2B–$1.8B USD).
  • Together AI's momentum: The company closed an $800M Series C at an $8.3B valuation and is now serving 400 trillion tokens monthly.
  • Second major B300 deal in a week: IBM signed a $240M multi-year agreement with Together AI for a B300 cluster on IBM Cloud just days prior.
  • Geopolitical tailwind: India's Tier 1 US export control status gives it unrestricted access to NVIDIA's latest chips, unlike China which is blocked from B300-class hardware.
  • Open-source focus: The factory will power inference, fine-tuning, and training for open-source models (Llama, DeepSeek, Qwen, Mistral), with sovereign data residency for Indian enterprises.

India just got its largest AI factory. Larsen & Toubro (L&T), the Indian engineering and infrastructure giant, has partnered with US-based AI cloud company Together AI to build a 10,000-GPU NVIDIA B300 cluster in Chennai. It is the largest single-cluster AI infrastructure deployment in the country's history, and it signals that the global race for sovereign AI compute has a new, serious contender.

The deal, by the numbers

Vyoma.AI, an L&T company, through its AI infrastructure subsidiary LTN Compute, has secured India's largest single-cluster AI infrastructure , an NVIDIA B300 AI Factory. According to L&T's project classification, the deal is valued between Rs 10,000 crore and Rs 15,000 crore , roughly $1.2B to $1.8B USD. L&T's share price rose 1.26% to Rs 4,070.70 on the announcement.

The hardware at the center of this is worth understanding. The B300, announced at GTC March 2025, delivers 15 petaFLOPS of dense FP4 compute, 288 GB of HBM3e memory, 8 TB/s bandwidth, and a 1,400W TDP. That FP4 compute figure is 67% more than the B200 and roughly 18x more than the H100. For inference-heavy workloads , which is exactly what Together AI runs , the memory headroom matters enormously: the 288 GB VRAM gives headroom to maintain large context windows without evicting the KV cache, which directly impacts reasoning quality and latency.

Who is actually building this

The key players span two continents:

  • L&T / Vyoma.AI / LTN Compute , L&T is one of India's largest conglomerates, historically known for civil engineering and heavy infrastructure. Vyoma.AI is its AI infrastructure arm, and LTN Compute is the subsidiary executing this build.
  • Together AI , a cloud platform for open and custom AI, offering fast inference for open-source models like Llama, Mixtral, DeepSeek, and Qwen, plus fine-tuning infrastructure. Founded in 2022 by Vipul Ved Prakash and Ce Zhang, based on research work at Stanford.
  • NVIDIA , supplying the B300 hardware and, per the broader L&T-NVIDIA partnership, the networking, storage, and AI software stack.

Together AI recently raised an $800M Series C financing round at an $8.3B valuation to expand its AI Native Cloud. The company reports it is now serving 400 trillion tokens monthly. This is not a startup dabbling in infrastructure , it is a company at production scale, looking for more compute to feed a growing pipeline.

Why India, why now

The timing is not accidental. India is in the middle of a government-backed compute buildout unlike anything it has attempted before. India is poised to quintuple its installed GPU capacity from 38,000 to 100,000 by the end of 2026, driven by the IndiaAI Mission. The IndiaAI Mission is a $1.25B (Rs 10,372 crore) program approved in March 2024 and scaling through 2026.

Crucially, India has a geopolitical advantage that its neighbors do not. India is Tier 1 under US export controls , unrestricted access , giving it a structural advantage over Tier 2 competitors in AI infrastructure buildout. While China is locked out of H100s, B200s, and their successors, India can freely import the most advanced NVIDIA silicon available. That is a window that L&T and Together AI are sprinting through.

For Together AI specifically, the India bet is also a diversification play. Just days before this announcement, IBM signed a multi-year $240M agreement with Together AI to deploy a large cluster of NVIDIA HGX B300 systems on IBM Cloud. The Chennai factory is the second major B300 cluster commitment in a week , Together AI is building a geographically distributed, open-source inference backbone at a pace that would have seemed implausible 18 months ago.

What is actually being built

The integrated AI Factory, with capacity of 10,000 B300 NVIDIA GPUs, will be hosted at Vyoma's Chennai data centre. The platform combines hyperscale data centre infrastructure, accelerated computing, high-performance networking, ultra-low-latency interconnects, high-throughput parallel storage and AI infrastructure operations.

Vyoma's Chennai data centre campus is a gigawatt-scale AI infrastructure site, with Phase 1 designed for 250 MW and power infrastructure readiness of 150 MVA, providing a scalable foundation for future AI Factory expansion. The roadmap does not stop at Chennai either: L&T's plans include initial expansions in Chennai to 30 megawatts as well as a new 40-megawatt facility in Mumbai.

The workloads this cluster will serve cover the full AI development lifecycle:

  • Inference , serving open-source models (Llama, DeepSeek, Qwen, Mistral) at low latency and high throughput for developers and enterprises
  • Fine-tuning , letting teams adapt base models on proprietary datasets without managing their own GPU clusters
  • Training , full pre-training runs for teams building new models from scratch

The real story: open-source AI gets sovereign infrastructure

The headline is a GPU count. The actual story is about what kind of AI ecosystem gets built around it. Together AI is built on the principle that open-source models are essential for the future of AI, and that developers should be able to build using open, modular stacks. Planting 10,000 of the world's most powerful GPUs in India , dedicated to open-source inference and training , means Indian startups, researchers, and enterprises get access to frontier compute without routing every workload through a US hyperscaler.

The venture aims to build sovereign AI infrastructure that retains data, models and workloads within India while maintaining compatibility with global systems. For regulated industries like banking, healthcare, and government services, that data residency guarantee is often the difference between adoption and a hard no.

The competitive pressure this creates is real. Gartner forecasts that worldwide spending on AI-optimized infrastructure as a service is projected to grow 96% through 2026, reaching $42 billion. The rise of agentic AI amplifies compute intensity through multistep, autonomous execution, making inference the dominant consumption model. A cluster of this size, running the latest Blackwell Ultra silicon, positions L&T and Together AI to capture a significant slice of that demand before AWS, Azure, or Google Cloud can build comparable in-country capacity.

Who wins, who watches nervously

The clear winners are Indian AI-native companies that need serious compute without dollar-denominated cloud bills. India's policy isn't a compute arms race , it's a subsidy that lets domestic teams reach the starting line without dollar-denominated cloud bills crushing them. A dedicated open-source inference cluster in Chennai accelerates that access significantly.

The more complicated picture is for the hyperscalers. AWS, Google Cloud, and Azure all have India regions, but none of them are running 10,000 B300 GPUs dedicated to open-source model serving. If Together AI can deliver competitive token economics from Chennai, it becomes a credible alternative for any Indian enterprise currently paying hyperscaler inference rates.

The risk worth watching is hardware concentration. India has built its AI progress on a foundation of American chips from a company that could, in a different geopolitical moment, cut off supply the way the US restricted Chinese access to advanced semiconductors. A 10,000-GPU cluster is also a 10,000-GPU single point of dependency. That said, India's Tier 1 export control status makes this a theoretical risk rather than an imminent one.

For the open-source AI ecosystem globally, this is a meaningful infrastructure commitment. Together AI provides a serverless inference API supporting more than 200 open-source models including Llama, Mistral, DeepSeek, and Qwen. More B300 capacity means lower token costs, higher throughput, and the ability to serve larger models , which directly benefits every developer building on those APIs. The India factory is not just an Indian story. It is a capacity expansion for the open-source inference layer that a significant portion of the global AI developer community already depends on.

Comments

avatar