Blog

The AI bill is going up: servers 15% pricier in 2027

Nvidia told its biggest customers that AI servers will cost over 15% more in 2027, because of memory. What that changes for companies that bought AI on price.

August 25, 2026 · Agência Primeira Página

The AI bill is going up: servers 15% pricier in 2027

On 22 August 2026, Bloomberg reported that Nvidia had notified some of its largest customers: AI servers will cost more than 15% extra in many cases, for Grace Blackwell and Vera Rubin systems delivered in early 2027. The warning reached the big buyers through the contract manufacturers that assemble the servers, and the size of each increase depends on the chip generation and the memory configuration.

For two years, almost every piece of news about the cost of AI pointed the same way: cheaper. This one points the other way — and for a reason that has nothing to do with artificial intelligence. It is a memory shortage.

Why it is rising: memory became almost a third of the machine

A Vera Rubin VR200 rack carries a bill of materials of roughly US$ 2.1 million. Memory, once a small slice of that, now reaches about 29% of it — close to US$ 609,000 in memory alone, when Nvidia's own target is to stay around 20% of system cost.

The workaround described in the reports was to cut capacity: SOCAMM memory on the Vera CPU would drop from around 55 terabytes to 28 terabytes per rack, while the GPU's HBM4 stays at 20.7 terabytes. In other words, on top of being more expensive, the delivered system may ship with less memory than the original design called for.

The backdrop is the memory market. TrendForce projected a 58% to 63% rise in conventional DRAM contract prices in the second quarter of 2026, and forecasts another 13% to 18% for server DRAM in the third. This is not a one-quarter blip: the firm expects successive increases into the second half of 2027, even if the pace slows.

And Nvidia chose to pass it on. The company runs a non-GAAP gross margin of 75.0% — that was the figure for the first quarter of fiscal 2027, with guidance at the same level for the quarter after. This is not a squeezed company sharing the pain with its customers; it is a decision to protect the margin.

Does this contradict what we wrote a month ago?

No, and the difference is the whole point. In July we published that the price of AI fell 80% in three weeks. That still holds — but they are two different bills.

The price of the token, which is what your company pays to use the model, falls through software: more efficient models, quantization, routing, competition between vendors. The price of the metal, which is what the vendor pays to own the machine, rises through physical scarcity of a component nobody manufactures overnight.

The two curves coexisted so far because the software drop outran the hardware rise. Nvidia's warning is the first clear sign that the second curve has started to bite. It does not mean your invoice goes up 15% in January — it means the slack that was pushing prices down has shrunk.

The clause that separates who is protected from who pays

Here is the part almost nobody discusses, and the most useful one for a mid-sized company.

TrendForce itself explains that several US cloud providers signed multi-year supply agreements, and that those agreements stop memory makers from raising prices for them. The consequence is direct: the increase does not spread evenly. It lands on whoever does not have a long-term contract.

The same logic runs down the entire chain. Whoever buys racks negotiates directly with the memory supplier; whoever buys an API from a large vendor is protected by that vendor's contract until the day they are not; and whoever ordered a system with AI built in from a small integrator sits at the most fragile end of the queue, because no contract is locking anything for them.

The practical question is not “will AI get more expensive?” It is “which step of that ladder am I on, and for how long is my price locked?”

Three decisions this changes in your company

1. Read what your contract says about price revision. Not the list price — the clause. AI list prices matter least in a market that has changed three times in twelve months. What matters is whether there is a fixed term, what the adjustment index is, how much notice the vendor must give, and what happens if they swap the model underneath the contract.

2. Know how much of your cost is actually AI. In most projects we deliver, model spend is the smallest line — the money goes into integration, into organised data, and into whoever maintains the system afterwards. If AI is 6% of your cost, a 15% increase on it moves less than 1% of the total, and that arithmetic calms the decision. If you do not know the proportion, the increase looks bigger than it is.

3. Keep a second vendor configured, not merely considered. Switching models has to be cheap. That means prompts that do not depend on a single vendor, your own evaluation with real tasks, and a plan B that has run at least once. Whoever has that negotiates; whoever does not, accepts.

Want to know where AI pays for itself first in your business — and where it only adds cost?

We run that diagnosis before writing a single line of code.

See AI implementation for business

The other side of the same coin: memory decides on phones too

That same weekend, Artificial Analysis published, in partnership with Liquid AI, an independent measurement of small models running inside the device, on the iPhone 17 Pro and the Galaxy S26 Ultra.

The result matters because it shows the same constraint from the other end. LFM2.5-2.6B ties for first place at 63 points, answering a 1,024-token prompt in 8.0 seconds using 2.3 GB of memory, while 9-billion-parameter models take over 25 seconds and 6.9 GB to reach a comparable score.

And the whole ranking shifts with available memory: with context capped at 16,000 tokens, one model lands fourth; at 64,000 it climbs. Except a 64,000-token window does not fit in a phone's memory, and at the measured speed, generating that many tokens would mean a 20-minute wait and a lot of battery.

It is the same lesson at two scales: what decides what is possible is not the size of the model, it is the memory available on the machine that will run it. In the data centre that became price; on the phone, it became a ceiling.

What you cannot conclude from this

It is worth drawing the boundaries, because price headlines tend to promise more than the data supports.

  • The warning is about servers delivered in 2027, not about next month's subscription invoice.
  • A hardware increase does not become an API increase in the same proportion — in between sit contracts, competition, model efficiency and each vendor's commercial call.
  • Nvidia did not comment publicly, and the percentages come from market reporting, not from a published price list.
  • Memory shortages have a cyclical history: they tighten, attract factory investment, then ease. What nobody knows is when.

The honest summary

AI getting cheaper was never a law of nature — it was software outrunning hardware. That lead has narrowed.

For a company that uses AI, the answer is not to postpone the project waiting for a better price. It is to stop treating price as the main criterion and look at what survives an increase: a contract with a term, a total cost you actually know, and the freedom to switch vendors without rebuilding everything. Whoever organised that rides out the increase complaining; whoever tied everything to one vendor for the promotional price rides it out rebuilding.

Sources

  • Bloomberg — “Nvidia Customers Notified About AI-Related Price Hikes Above 15%” (22 August 2026).
  • TrendForce — DRAM contract price forecasts (31 March and 9 July 2026).
  • NVIDIA — first quarter fiscal 2027 results (non-GAAP gross margin).
  • Artificial Analysis and Liquid AI — small model measurement on mobile devices (24 August 2026).

Frequently asked questions

Will my AI subscription cost 15% more?

That is not what was announced. Nvidia's warning concerns the price of servers delivered in early 2027, bought by large cloud providers. Between that cost and your invoice sit contracts, competition and efficiency gains in the models. What changes is the direction: the slack that had been pushing prices down has shrunk.

Why are AI servers getting more expensive?

Because of memory. In a Vera Rubin VR200 rack, memory reaches about 29% of a roughly US$ 2.1 million bill of materials, when Nvidia's target is to stay near 20%. TrendForce projected a 58% to 63% rise in conventional DRAM contract prices in the second quarter of 2026 and forecasts another 13% to 18% for server DRAM in the third.

Does this contradict the news that AI prices fell 80%?

No, they are two different bills. Token prices fall through software — more efficient models, quantization, competition. Hardware prices rise through physical memory scarcity. The two curves coexisted because the software drop was faster; now the hardware rise has started to bite.

What should my company do now?

Three things: read the price revision clause in your contract instead of only looking at the list price; work out what percentage of the project cost is actually the AI model, which is usually the smallest line; and keep a second vendor configured and tested, so switching is cheap.

Should I postpone an AI project because of the increase?

Generally no. Model spend is usually the smallest part of a project's cost — most of it is integration, data organisation and maintenance. Postponing does not reduce those lines and delays the return on top. What is worth doing is picking a project that finishes and measuring the result.

Do small models running on the device solve the problem?

They solve part of it, and only for some tasks. The Artificial Analysis measurement with Liquid AI shows a 2.6-billion-parameter model answering in 8 seconds using 2.3 GB on an iPhone 17 Pro, scoring the same as much larger models. But the device's memory caps the context size, and long tasks still call for a server.