AI News

Nvidia Blackwell Ultra: the GPU that powers AI agents

Announced at GTC in March 2025, the Nvidia Blackwell Ultra (B300) packs 288 GB of HBM3e and doubled performance on agentic workloads. This chip redefines the balance of power between cloud and on-premise for businesses.

5 min read
NvidiaGPUInfrastructureCloudPerformance
⚡ The news in 30 seconds

Nvidia Blackwell Ultra: the GPU for AI agents takes shape

Announced at GTC in March 2025, Nvidia's Blackwell Ultra (B300) is the GPU built for large-scale AI inference. Key figures: 288 GB of HBM3e memory, doubled performance on agentic workloads compared to the B200, and a cost per token cut in half. The first deployments are planned at AWS, Azure and GCP in the second half of 2025. In parallel, Nvidia presented the Dynamo platform to orchestrate multi-GPU inference at datacenter scale.

The era of autonomous AI agents requires a radically different compute infrastructure — and Nvidia is about to deliver it.

What this changes for you

The opportunity

Until now, running an AI agent continuously was expensive. A Claude or GPT agent working 8 hours a day on analysis tasks can represent 2,000 to 5,000 euros per month in API costs. With the Blackwell Ultra, cloud providers will be able to reduce these rates by 30 to 50%, making profitable use cases that were not yesterday.

What this concretely unlocks:

🤖

Continuous AI agents

Running an AI copilot 24/7 on customer support, market analysis or regulatory monitoring becomes economically viable for a mid-market company.

🏢

Accessible on-premise AI

For sensitive sectors (defense, healthcare, finance), the Blackwell Ultra makes high-performance on-premise AI possible at a controlled cost — without depending on a US cloud.

Real-time inference

Applications requiring responses in under 100ms (trading, autonomous driving, high-frequency chatbots) gain in reliability and speed.

The risk

⚠️

Beware of over-investing in infrastructure

The classic trap: investing in expensive GPU hardware before validating the use cases. A Blackwell Ultra server costs between 60,000 and 100,000 euros. If your inference needs do not exceed 2 million tokens per day, the cloud remains far more economical. Don't confuse available power with required power.

🔒

Nvidia's monopoly on AI

Nvidia controls more than 80% of the GPU market for AI. This dependence exposes the entire ecosystem to the decisions of a single supplier in terms of pricing, allocation and geographic priorities. Keep an eye on the alternatives: AMD MI350, Intel Gaudi 3, and the clouds' custom chips (Google TPU, AWS Trainium).

Our recommendation

For 95% of SMBs and mid-market companies, the arrival of the Blackwell Ultra is indirectly good news: AI cloud prices will fall. Here is how to take advantage of it.

1

Assess your current consumption

How much do you spend on AI APIs per month? How many tokens do you process? If it is less than 5 million tokens per day, stay on the cloud and wait for the price cuts.

2

Negotiate your cloud contracts

Cloud providers will offer Blackwell Ultra instances as early as Q3 2025. This is the time to renegotiate your reserved instances or to benchmark the AWS, Azure and GCP offerings on your real workloads.

3

Consider on-premise only if necessary

On-premise investment is only justified for sovereignty reasons (sensitive data, regulation) or volume (more than 10M tokens/day). In that case, compare the 3-year TCO between a Blackwell Ultra cluster and the cloud equivalent.

In summary

Opportunity
AI inference cost halved, continuous AI agents made viable
Risk
Infrastructure over-investment, Nvidia's monopoly on the AI GPU market
Recommended action
Stay on the cloud, negotiate rates, watch for Q3 2025 price cuts
Horizon
Cloud cost reductions as early as September-October 2025

Frequently asked questions

What is the Nvidia Blackwell Ultra?

The Blackwell Ultra (B300) is Nvidia's latest GPU, designed specifically for large-scale AI inference. It packs 288 GB of HBM3e memory (versus 192 GB for the B200), delivers doubled performance on agentic workloads, and consumes 1,200 watts. It is available from the major cloud providers (AWS, Azure, GCP) and for direct purchase for private datacenters.

Why is the Blackwell Ultra important for businesses?

Because AI agents (like Claude 4 in autonomous mode or business copilots) require massive and continuous inference power. The Blackwell Ultra makes it possible to run these agents at a cost per token halved compared to the previous generation. For a business, this means that use cases that were too expensive yesterday become profitable today.

Should you invest in an on-premise GPU or stay on the cloud?

For the vast majority of SMBs and mid-market companies, the cloud remains the best option. A server equipped with Blackwell Ultra costs between 60,000 and 100,000 euros. The cloud (AWS, Azure, GCP) gives you access to the same power on demand, with no upfront investment. On-premise is only justified if you process more than 10 million tokens per day or if you have strict sovereignty constraints.

For tech profiles

Blackwell Ultra in detail

B300 (Blackwell Ultra)

Flagship inference GPU

288 GB HBM3e, 12 TB/s memory bandwidth, 1,200W TDP, native FP4/FP8 support. Optimized for long queries and multi-step AI agents with extended context windows.

DGX B300

Turnkey AI server

8x B300, 2.3 TB of total GPU memory, NVLink 6 interconnect. Designed to run 400B+ parameter models in local inference. Indicative price: $400,000-500,000.

Pricing

B300 unit ~$60,000-100,000
DGX B300 ~$400,000-500,000
Cloud (instance) ~$30-50/h (estimated)
Availability Q3 2025 (cloud) · Q4 2025 (on-prem)

Quick comparison

CriterionBlackwell Ultra (B300)Blackwell (B200)AMD MI350
HBM memory288 GB HBM3e192 GB HBM3e288 GB HBM3e
Inference perf.2x B200Baseline~1.5x B200 (estimated)
Software ecosystemCUDA (mature)CUDA (mature)ROCm (improving)
AvailabilityQ3-Q4 2025AvailableQ4 2025

Related articles