Technology

NVIDIA DSX MaxLPS: Lambda reports 23% performance-per-watt gain in fixed-power validation

NVIDIA DSX MaxLPS helped Lambda run 19 HGX B200 nodes inside a 16-node power budget, with reported gains of 24% in throughput and 23% per watt.

Original editorial illustration of NVIDIA DSX MaxLPS power coordination across liquid-cooled AI server racks in a modern data centre
Original editorial illustration of NVIDIA DSX MaxLPS power coordination across liquid-cooled AI server racks in a modern data centre. Illustration: Reddy News.
Key points
  • Lambda and NVIDIA say a five-rack, 19-node HGX B200 proof of concept ran within the power budget normally allocated to a 16-node baseline.
  • The companies reported cluster token throughput of roughly 5 million tokens per second, up from about 4.04 million, a stated 24% increase, alongside 23% better performance per watt.
  • NVIDIA DSX MaxLPS combines site and cooling design, Dynamic Power Software and workload-level performance-per-watt tuning to use a fixed AI factory power envelope more productively.
  • NVIDIA documents Dynamic Power Software as a Developer Preview that may change and should not be used in production except in an authorised collaboration.
  • NVIDIA's up-to-40% more GPU-capacity figure is a conditional AI-factory design claim, not the measured Lambda outcome; the Lambda validation is not an independent fleet-wide benchmark.

Lambda reports more token output within a fixed power budget

Lambda and NVIDIA say a five-rack proof of concept on NVIDIA HGX B200 systems ran 19 nodes in the power budget normally allocated to 16 full-power nodes. Their jointly published case study reports approximately 5 million tokens a second under an 85% NVIDIA DSX MaxLPS policy, compared with about 4.04 million tokens a second for a 16-node baseline without that policy. The companies describe the outcome as 24% more cluster-wide token throughput and 23% better performance per watt.

The validation was disclosed at the AI Infra Summit on 15 September. It is meaningful for people who plan or operate power-constrained AI infrastructure because the limit is often the electricity and cooling available to a site, not only the number of servers that can be bought. It is nevertheless a vendor-and-customer reported proof of concept, rather than a fleet-wide, third-party benchmark or an assurance that every workload will produce the same result.

What 19 nodes inside a 16-node power budget means

The comparison is a fixed-power experiment, not a claim that 19 machines consume less electricity in absolute terms than 16. Lambda says it applied configurable power controls at 80% and 85% of maximum draw, lowering per-node consumption enough to bring additional nodes online within the stated facility budget. The reported pure-inference result used a 16-node baseline with no policy and a 19-node configuration at an 85% policy.

NVIDIA and Lambda identify the test environment as a five-rack, 19-node cluster of HGX B200 systems. The case study says it used MLPerf inference and training workloads to produce consistent, peak-level draw and reproducible, industry-comparable results. Its inference chart identifies GPT-OSS-120B at 40 queries per second per node. Those configuration details help define the result, but the publication does not provide an independent audit, a distribution of results across Lambda's wider fleet, or complete operational measures such as service latency under every workload mix.

How NVIDIA DSX MaxLPS coordinates an AI factory

NVIDIA describes DSX MaxLPS as an AI-factory design and operating framework rather than a single GPU setting. It brings together facilities and site design, NVIDIA Dynamic Power Software and techniques for improving application performance per watt. Its purpose is to raise throughput per megawatt inside a fixed site-power envelope by making more of the available IT capacity usable and then making that capacity more productive.

At the software layer, NVIDIA says Dynamic Power Software monitors GPU and rack-level consumption, identifies unused allocation and reallocates available power across configured GPU and rack groups within defined budgets and policies. Static provisioning must reserve capacity for credible peaks and failure conditions. The MaxLPS premise is that workloads do not all reach those peaks at the same moment, leaving headroom that can be managed. Facility controls and electrical protections still operate and protect the physical infrastructure; NVIDIA says MaxLPS does not replace them.

Why tokens per megawatt needs a careful reading

Tokens per second measures the output rate of an inference service, while performance per watt relates that output to the power used. For a cloud operator, the question is practical: can a constrained site serve more AI work without increasing its approved power envelope? The reported Lambda comparison answers that question for its disclosed configuration by showing more aggregate token output under the same stated budget.

That does not make tokens per megawatt a universal score. Model size, request rate, context length, latency targets, training activity, network behaviour, thermal conditions and the power policy can all affect the outcome. NVIDIA itself explains that training and inference have different power profiles. Readers should therefore treat the 24% and 23% figures as measurements reported for this proof of concept, not as a conversion factor that can be applied to another data centre or service.

Dynamic Power Software remains a Developer Preview

NVIDIA's current Dynamic Power Software documentation labels the software a Developer Preview. The notice says DPS may have limitations, may change significantly and includes functionality still being refined through testing and feedback. It also says the software should not be used in production environments except as part of an authorised collaboration. That status matters when separating a technical validation from general availability at scale.

The documentation describes DPS as a data-centre power-management system that monitors consumption, enforces power policies and integrates with baseboard-management controllers and cluster-management systems. NVIDIA's MaxLPS overview lists current support for Blackwell HGX B200 and B300 platforms, GB200 and GB300 NVL72 platforms, and Vera Rubin NVL72. A deployment also requires a validated topology, policy boundaries and suitable telemetry; the software's ability to shift allocation does not remove the need for site-specific operating safeguards.

The up-to-40% design claim is separate from Lambda's test

NVIDIA says a fully planned MaxLPS AI factory can make up to 40% more GPU capacity available within the same site-power envelope. That is a company design claim with stated conditions, not the result Lambda measured. NVIDIA's documentation says reaching the full potential requires planning for the complete MaxLPS target, including a facility design supporting 45°C cooling, and depends on site and ambient conditions that determine cooling overhead.

The difference is important. Lambda's reported configuration added three nodes to a 16-node baseline and reported throughput and efficiency changes in that specific cluster. NVIDIA's up-to-40% figure concerns a wider future-facing site design that combines cooling, electrical distribution, network capacity, rack positions, dynamic allocation and workload choices. Conflating the two would turn a conditional planning scenario into a measured customer result.

What the validation establishes and what it does not

The disclosed evidence establishes that NVIDIA and Lambda completed the described HGX B200 proof of concept and publicly reported its stated performance outcomes. Wccftech independently covered the announcement, but its report relays the companies' disclosed figures rather than presenting separate test results. The most accurate label is therefore a vendor-customer validation that has been independently reported, not independently replicated.

The result points to a useful operational hypothesis: mixed or variable workloads can leave headroom that dynamic, policy-bound allocation may recover. Whether that translates into reliable capacity for a particular organisation depends on its systems, workloads, thermal design, service objectives and risk controls. NVIDIA's documents make the same broader point in technical terms by requiring a bounded pilot, validated deployment plan and operator-approved policies. The public material supports interest in further testing, not a promise of fleet-wide performance or a financial conclusion.

Reader guide

Article questions, answered

Short answers to common reader questions based on the reporting above.

What did Lambda report for NVIDIA DSX MaxLPS?

In a five-rack, 19-node NVIDIA HGX B200 proof of concept, Lambda reported about 5 million tokens per second under an 85% DSX MaxLPS policy, versus about 4.04 million tokens per second for a 16-node baseline with no policy. Lambda and NVIDIA describe that as a 24% cluster-wide token-throughput gain and a 23% improvement in performance per watt within the same fixed power budget.

Does 19 nodes within a 16-node power budget mean the cluster used less total power?

Not necessarily. The comparison is about producing more work within a fixed facility power budget, not about a claim that the 19-node configuration used less total power than the baseline. Lambda says per-node power policies reduced the draw per node enough to operate more nodes within that stated budget.

What is the difference between NVIDIA DSX MaxLPS and Dynamic Power Software?

NVIDIA describes DSX MaxLPS as the broader AI-factory framework combining site and cooling design, Dynamic Power Software and workload performance-per-watt techniques. Dynamic Power Software, or DPS, is the software component that monitors consumption and manages available power across configured GPU and rack groups within policies and budgets.

Is the up-to-40% NVIDIA DSX MaxLPS figure a Lambda result?

No. NVIDIA presents up to 40% more GPU capacity in the same site-power envelope as a conditional MaxLPS design potential for an AI factory planned for the full approach, including 45°C cooling where site conditions support it. Lambda's disclosed validation concerns the separate 19-node HGX B200 configuration and its reported 24% throughput and 23% performance-per-watt results.

Sources and further reading

These references support the factual context used in this article. Links open the original publisher.

  1. AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI FactoriesNVIDIA · accessed 17 September 2026
  2. Lambda Maximizes Performance per Watt With NVIDIA DSX MaxLPSNVIDIA · accessed 17 September 2026
  3. NVIDIA DSX MaxLPS OverviewNVIDIA Documentation · accessed 17 September 2026
  4. NVIDIA Dynamic Power SoftwareNVIDIA Documentation · accessed 17 September 2026
  5. Maximizing AI Factory Performance per Watt with NVIDIA DSX MaxLPSNVIDIA Technical Blog · accessed 17 September 2026
  6. NVIDIA’s DSX MaxLPS Drives 40% More Token Throughput Per Megawatt, As Lambda’s Cluster Jumps To 5 Million Tokens Per SecondWccftech · accessed 17 September 2026