
On 26 August 2026 Nvidia reported revenue of $96.2 billion for the quarter that ended on 26 July, up 106% on a year earlier, and told investors to expect about $108 billion in the quarter that follows. Thirteen days later, on 8 September, Qualcomm announced a multi-generation collaboration with Amazon to build customised chips for running AI models in Amazon Web Services data centres.
Read together, the two announcements describe a market moving in two directions at once. Demand for general-purpose AI accelerators is still rising steeply. At the same time, the largest buyers are putting money behind alternatives designed for one job: serving output from models that have already been trained, at the lowest possible cost per token.
For organisations that pay for AI by the call, the second trend may matter more than the first. This article separates what the companies disclosed in their own releases and regulatory filings from what was added by press coverage, because in the Qualcomm case the headline number did not appear in the announcement at all.
What Nvidia reported, and what it assumed
Nvidia's release puts data-centre revenue at $89.0 billion, up 117% year on year. Edge computing, the only other revenue line the release breaks out, brought in $7.2 billion. Total revenue grew 18% on the previous quarter, and gross margin was 75.0% on both GAAP and non-GAAP measures.
The outlook of $108.0 billion, plus or minus 2%, came with an explicit condition: Nvidia is not assuming any data-centre compute revenue from China. The release does not say why, but the quarterly filing gives the context: Nvidia says that by the end of the quarter it was effectively shut out of China's data-centre compute market, pointing to US export controls and to Chinese government restrictions on buying its products. The effect is clear. An entire national market is set to zero in the company's own planning, and the guidance nonetheless points to sequential growth of roughly 12%.
On products, Nvidia said its Vera Rubin platform is moving into full production, and named CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure and Nebius among the partners where its racks are already running. A less prominent line is more telling for this story. Nvidia said Groq 3 LPX, which it describes as an interactive AI inference accelerator, is now in full production. The market leader itself now has hardware in full production that is aimed specifically at inference.
Chief executive Jensen Huang framed the quarter in commercial terms, arguing that AI output has become useful, profitable work, and summed it up as 'Now, compute is revenue.' The release also announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to set up independent financing platforms that could channel more than $500 billion of outside capital into AI infrastructure over time. Those partnerships remain subject to definitive agreements.
Training, inference and why the chip mix is changing
Building an AI model and using it are different workloads. Training means running very large volumes of data through a model repeatedly to set its parameters. It rewards flexible, high-end accelerators linked tightly together, and it happens in large, relatively infrequent campaigns.
Inference is what happens every time someone sends a prompt or an agent takes a step. The finished model produces its output token by token, continuously and at scale. Once a model is in service, its operator cares less about flexibility and more about the cost of each answer. Agent systems amplify this, because one task can involve many model calls, each running on inference hardware.
That cost depends on how much useful output a chip produces for each watt of electricity, how quickly it moves data in and out of memory, and how efficiently many chips exchange data with each other. General-purpose GPUs handle all of this well. A chip designed around one provider's models and serving software, however, can leave out capabilities that workload never uses.
This is the logic behind custom silicon, often called ASICs or XPUs. The buyer, usually a cloud provider, specifies what it needs, and a chip company contributes design expertise, intellectual property and access to manufacturing. The trade-off is a different kind of lock-in: a custom part can be excellent at the jobs it was built for and less adaptable when models or methods change.
Connectivity is the other half of the equation. Large models are spread across many chips, so the links between them can become the constraint. Qualcomm's work on optical links of up to 1.6T, meaning 1.6 terabits per second, and Nvidia's Spectrum-6 switch systems supporting co-packaged optics both reflect the same pressure: moving data between processors with light to reach further on less power.
The Qualcomm and AWS deal: release, filing and reports
Qualcomm's release of 8 September is short on specifics. It describes multi-generation customised silicon for AI inference in AWS data centres, plus optical connectivity built on SerDes and optical DSP technology. It also says Qualcomm will make greater use of AWS, including Amazon Bedrock, for chip-design workloads, with the aim of shortening design cycles. It contains no dollar figures, volumes or product dates. The executive comments centre on efficiency: Qualcomm chief executive Cristiano Amon says data-centre infrastructure needs progress in both computing and connectivity, and AWS vice president Prasad Kalyanaraman points to more cost-effective infrastructure for customers.
The financial terms come from a regulatory filing Qualcomm made the same day. On 3 September, Qualcomm issued a warrant to an Amazon affiliate for up to 25 million Qualcomm shares at an exercise price of $161.26 each, expiring on 3 September 2036. Of those, 3.75 million shares, or 15%, vested on issuance based on initial purchase commitments.
The remaining shares vest in tranches as Amazon signs commercial arrangements, places binding purchase orders and makes actual purchases. The vesting runs up to a maximum of $60 billion in payments by Amazon for Qualcomm server chip products and services.
Press coverage compressed this into a $60 billion deal. The Motley Fool described Amazon agreeing to buy up to $60 billion of chips and put the warrants at about $4 billion. That figure appears to equal the share count multiplied by the exercise price, roughly $4.03 billion, rather than a market valuation of the warrant itself. Under the filing's terms, $60 billion is the most Amazon could pay before vesting stops; it is not a firm order.
Barchart reported that chief financial officer Akash Palkhiwala said the agreement gives Qualcomm very high confidence in the $5 billion of data-centre revenue it targets for fiscal 2027, alongside a goal of $15 billion by fiscal 2029. According to the same report, revenue from the partnership begins in the quarter ending in December 2026, and Palkhiwala said a comparable arrangement is in progress with another hyperscale cloud provider, which he did not name. These are executive remarks relayed by the press, not terms in the release or the filing.
The $60 billion is a ceiling on purchases that could unlock shares, not an order book.
Structurally, the warrant works like a volume incentive paid in equity. The more Amazon buys, the more Qualcomm shares it can acquire at a set exercise price. It links the size of a customer's potential stake in its supplier to that customer's future orders, and it gives the customer a reason to keep ordering.
Custom silicon moves from side bet to core business
Qualcomm is not the only supplier reporting this demand. On 2 September Broadcom reported AI semiconductor revenue of $16.7 billion for its quarter ended 2 August, up 221% year on year, and said it expects $21.7 billion in its fourth quarter. Chief executive Hock Tan pointed to strong demand for the company's custom AI accelerators and its networking products.
For Qualcomm, the move is also about reducing dependence on phones. The Motley Fool reported that handsets made up about 75% of revenue in Qualcomm's chip segment in fiscal 2025, and that the company expects its non-handset businesses to reach $40 billion by fiscal 2029. The same report noted that Qualcomm named Meta as its first data-centre customer in June 2026, under a multi-generation supply agreement for its Dragonfly C1000 server CPUs.
Nvidia, meanwhile, is working to keep the whole buildout moving. Its quarterly filing lists $279 billion of supply and capacity commitments, primarily for memory and manufacturing facilities. It also says Nvidia has arrangements to help select customers secure the land, power, building shells and data-centre capacity they need. The financing platforms serve a similar purpose: customers who can fund more infrastructure can buy more systems.
Seen side by side, the strategies differ. Nvidia is widening its range with an inference accelerator, a CPU it says is built for AI agents and optical networking, while helping to finance demand. Cloud providers are building second sources they control. Both approaches rest on the assumption that inference volumes keep climbing.
What is disputed or still unknown
Several important questions remain open, and some of the figures in circulation are softer than they look. The first is the real size of the Qualcomm and Amazon relationship. The $60 billion figure is a vesting cap over ten years. How much Amazon actually buys will depend on product performance, delivery and Amazon's other options, and none of that is disclosed.
The second is performance. Barchart reported Qualcomm executive Durga Malladi saying that a technology the company calls High-Bandwidth Compute delivers up to six times better performance per watt than the industry-standard high-bandwidth memory approach, with a first generation shipping commercially in 2027. That is a vendor claim relayed through press coverage, not an independent benchmark. It should be treated as developing until customers or third parties publish results.
The third is Nvidia's planning assumptions. The China exclusion is a forecasting choice, not a permanent state, and export policy could move it in either direction. The financing partnerships are subject to definitive agreements, so the $500 billion figure is an aim rather than committed capital.
The last is whether cheaper silicon becomes cheaper services. A lower cost per token for a cloud provider does not automatically mean lower prices for its customers. Providers may keep the margin, use it to expand capacity or pass it on selectively. None of the announcements discussed here commits to any price change.
What this means for teams running AI in production
For organisations running agents, voice systems, robotics or other automation, the compute race matters less for share prices than for availability, cost and flexibility over the next few hardware generations. Several practical steps follow.
- Keep workloads portable. Build on serving layers and model formats that run on more than one cloud and more than one accelerator family, so a price cut in one place can be used and a capacity shortfall absorbed.
- Benchmark on your own traffic. Vendor performance-per-watt claims and headline token prices are starting points; measure cost per completed task, latency and error rates on real workloads before committing.
- Separate interactive and batch work. Latency-sensitive agents and voice systems may justify premium capacity, while evaluation runs, document processing and overnight jobs can often move to whatever is cheapest.
- Plan for regional differences. Nvidia's decision to assume no China data-centre compute revenue shows how export policy can reshape where hardware is sold; check where your providers can serve from and keep a fallback region.
- Avoid long commitments priced on current hardware. With several chip suppliers promising multi-generation roadmaps, favour contract terms that allow repricing or migration as inference costs shift.
- Treat efficiency as a design goal. Smaller models, caching and fewer unnecessary agent steps reduce exposure to whichever hardware turns out to be scarce.
Sources
- NVIDIA Announces Financial Results for Second Quarter Fiscal 2027NVIDIA Newsroom · 26 August 2026
- NVIDIA Corporation Quarterly Report for the period ended July 26, 2026US Securities and Exchange Commission (NVIDIA Form 10-Q) · 26 August 2026
- Qualcomm Announces Multi-Generational Product Collaboration with Amazon to Build Next-Generation AI Data Center InfrastructureQualcomm · 8 September 2026
- Qualcomm Incorporated Current Report, Item 3.02: warrant issued to an Amazon affiliateUS Securities and Exchange Commission (Qualcomm Form 8-K) · 8 September 2026
- Qualcomm Stock Spikes as It Bags Massive Data Center Deal With AmazonBarchart (via Yahoo Finance) · 10 September 2026
- Forget Smartphones: Qualcomm Just Landed a Massive AI Deal With AmazonThe Motley Fool · 9 September 2026
- Broadcom Inc. Announces Third Quarter Fiscal Year 2026 Financial Results and Quarterly DividendBroadcom (via PR Newswire) · 2 September 2026