Last week, we wrote about how the AI boom has led to an AI arms race between countries, but the arms race between big tech companies has been heating up as well.
Compute is one of the primary inputs into AI performance, both in training AI models and in AI inference. NVIDIA, which now has a $3.5 trillion market cap, has until recently had a virtual monopoly over supplying foundation model companies like OpenAI and Anthropic, or big tech companies developing competing models, with the necessary compute.
Some of the biggest tech companies — especially Microsoft (with its $14 billion investment in OpenAI), Amazon (which has invested a total of $8 billion in Anthropic), and Google — have been competing on AI CapEx, with NVIDIA benefiting from the competition as the principal provider of their GPUs. NVIDIA, which owns 95% of the AI chip market, has set benchmarks in AI training and inference through GPUs like the A100 and H100.
But things have lately taken a turn. This week, the Wall Street Journal reported that Amazon is building the biggest supercomputer in the world. At its annual re:Invent conference this past week, AWS unveiled Project Rainier, a supercomputer built on hundreds of thousands of its Trainium chips, which it claims will be one of the largest AI training clusters globally when it launches in 2025. Anthropic, an AI startup that received $4 billion in funding from Amazon last month, will use the supercomputer, which is five times larger by exaflops (a unit of measurement for supercomputing performance) than Anthropic’s existing training infrastructure.
Amazon is also claiming that early test results show its in-house AI chips are 30% to 40% cheaper than NVIDIA’s flagship H100 chips and is making a bid to convince AWS customers to use servers powered by Amazon’s own AI chip. Additionally, Amazon unveiled "Nova" foundation models for AI-powered text, image, and video generation, aiming to compete with OpenAI and Meta.
Last year, AWS debuted its Trainium2 chip and is now developing the Trainium3, expected to deliver four times the performance of Trainium2. By building its own chips and infrastructure, AWS aims to reduce dependency on NVIDIA GPUs and offer customers more cost-effective and scalable AI solutions. Other companies like Cerebras, Grok, and SambaNova Systems are also vying for a share of the market. Amazon’s cloud rivals like Microsoft and Google are also investing in developing their own AI chips.
Amazon is already seeing some traction: AWS announced this week that Apple is one of its newest chip customers, and companies like Databricks and Adobe (along with Anthropic) have reportedly seen promising results in their own recent tests of Amazon’s AI chips.
For most businesses, the choice between NVIDIA and Amazon for AI hardware isn’t a critical issue, according to analysts. Large companies are more focused on leveraging AI models for business value rather than the technical specifics of training them. This could work in Amazon’s favor, as it can integrate its Trainium chips through partnerships without customers needing to be aware of the underlying hardware. On the other hand, however, software developers are accustomed to writing software for NVIDIA’s chips.
It’s too early to say how it’ll all play out. But one thing is clear: while NVIDIA’s dominant market share has seemed unassailable to this point, the competitive landscape for AI hardware is starting to shift. Meanwhile, AWS is quietly becoming one of the biggest AI competitors by trying to be a one-stop shop for AI: users can deploy GenAI tools, build proprietary models, and now get compute directly from Amazon.



