AI Infra Summit 2026, held September 15–17 at the Santa Clara Convention Center in California, made one shift unmistakable: AI infrastructure is no longer a race measured only in accelerator speed. The conference brought memory bandwidth, power utilization, networking, chip-design cycles and inference deployment to the center of the conversation around production-scale and agentic AI.

That matters because an AI system can have enormous compute capacity and still lose time, money or efficiency while moving data, feeding models, managing power or waiting for the next generation of silicon. The summit’s presentations came from different companies and covered different stages of development, but together they pointed to the same pressure point: the infrastructure around the accelerator is becoming just as important as the accelerator itself.

The AI Infra Summit’s central message: the bottleneck is no longer just compute

The event focused on compute, data movement, physical AI, data centers, networking, storage, software and operations. In practical terms, that means the industry is trying to solve several linked problems at once: how to keep models supplied with data, how to use constrained data-center power more effectively, how to connect large numbers of chips and how to shorten the path from chip design to deployment.

The focus was especially relevant to agentic AI, in which a system can perform multiple model calls and actions to complete a task. That pattern increases pressure on CPUs, memory systems and the network around the accelerator. A faster GPU is useful, but it does not remove every bottleneck in the path between an input, a model and a result.

The conference took place in person in Santa Clara. NVIDIA’s program scheduled a keynote by Ian Buck, its vice president and general manager for Hyperscale and HPC, titled “Advancing Infrastructure for the Era of Agentic AI.” NVIDIA also said recordings of its sessions would be published within 72 hours of the live event.

NVIDIA Vera and DSX MaxLPS target agentic workloads and stranded power

NVIDIA presented Vera CPU data for agentic workloads and associated it with a 50–100% performance advantage over AMD Turin. That figure came from NVIDIA’s presentation data. AMD executive Forrest Norrod presented a competing claim of better performance than Vera, so the two companies offered different views of the comparison.

The important takeaway is narrower than a universal winner: CPUs are becoming a more visible part of the performance equation when AI agents rapidly consume model output and initiate another step. In those workloads, the system’s ability to coordinate inference can matter alongside the accelerator’s throughput.

NVIDIA also introduced DSX MaxLPS, power-management firmware for over-provisioned GPUs. NVIDIA presented the idea as a way to apply power policy more intelligently across an AI factory. The presentation associated that approach with 40% more AI Factory revenue by reducing stranded data-center power.

That is a business claim attached to a power-management strategy, not a general performance figure for every data center. Its significance is easy to understand, though: when power availability limits how much compute a facility can operate, better utilization can be as valuable as adding more silicon.

d-Matrix and Qualcomm put stacked DRAM at the center of inference

Stacked DRAM places high-capacity memory physically closer to processing logic. The goal is to move data with less delay and less energy than a design that repeatedly reaches across a more distant memory hierarchy. For AI inference, where models must constantly access weights and intermediate data, that can directly affect throughput and efficiency.

d-Matrix presented Raptor as a planned inference accelerator built around stacked DRAM. The company claimed 10× performance at one-tenth the power and presented a result of more than 3,000 tokens per second on GLM 5.2 with a 1-million-token context. The chart associated with that presentation showed approximately 3,153 tokens per second per user at a decode batch size of 8, 2,831 at batch size 16 and 2,121 at batch size 32.

Those figures describe a presentation workload, not a universal inference rate. Raptor was described as a future product direction after the 2026 summit, while d-Matrix’s longer-term Lightening roadmap was aimed at systems supporting multiple stacked DRAM dies and models in the 20-trillion-parameter range.

The company also described a disaggregated-inference design involving NVIDIA H200 GPUs for context processing and d-Matrix Corsair for decode processing. Disaggregated inference splits those stages across different hardware so each part can be tuned for its own workload. In theory, that can make existing infrastructure more useful instead of requiring every stage to run on the same accelerator.

Qualcomm presented a different stacked-memory approach called High Bandwidth Compute. Its design places stacked DRAM on a logic die. Qualcomm’s presentation cited 6× bandwidth per watt versus HBM and 200× capacity per watt versus SRAM. Those are presentation figures tied to the company’s architecture and comparisons, with no common industry test protocol attached to them.

The broader point is straightforward: memory is no longer just a specification buried in a server sheet. For inference systems handling large models and long contexts, the movement of data can determine how much of the available compute becomes useful work.

Open scale-up networking moves toward ESUN and optical interconnects

Broadcom promoted Ethernet Scale-up Networking, or ESUN, as an open approach to connecting processors inside large AI systems. Scale-up networking links the components of a single computing system or tightly coupled cluster, where latency and bandwidth can determine whether separate chips behave like a coordinated platform.

Broadcom’s Tomahawk 6 switch series was associated with 102.4 Tbps of switching capacity and described as production-volume silicon for the wider ecosystem. ESUN was also associated with the Open Compute Project, while Optical Compute Interconnect was presented as an effort to move scale-up connections from individual racks toward multi-rack and multi-row architectures.

That direction is ambitious, but networking ecosystems do not arrive with a single chip. Hardware, software, standards and interoperability all have to mature together. The summit’s networking discussions therefore pointed to a multiyear transition rather than an instant replacement for existing proprietary scale-up fabrics.

Cognichip attacks the chip-design bottleneck

Cognichip presented physics-informed AI models for chip design rather than large language models. Physics-informed models incorporate the behavior and constraints of the physical system they are working on, which makes them a better fit for tasks such as optimizing layouts, power use or signal behavior.

Cognichip claimed its tools could accelerate chip-design cycles by up to 100×. The presentation also placed that ambition against traditional design and fabrication cycles of 18–30 months and a possible AI-assisted range of 3–6 months.

The stakes are obvious. Faster model training does little to help a hardware program that takes years to reach production. If AI tools can shorten parts of the design process, chipmakers could respond more quickly to changing model architectures and data-center requirements. The claim remains tied to Cognichip’s solution and its intended use, but it addresses a bottleneck that raw accelerator throughput cannot solve.

Positron brings a startup accelerator into the cloud discussion

Positron said it had raised $875 million at a $5 billion valuation and was deploying more than 50 Atlas racks at Oracle Cloud Infrastructure. Atlas is Positron’s inference system, and the company claimed a 3× performance-per-dollar advantage over NVIDIA DGX H200, along with more than four times NVIDIA’s throughput and context windows exceeding one million tokens.

Those comparisons are company claims tied to Positron’s systems and chosen reference hardware. They show the kind of argument startup accelerator companies are making: specialized inference hardware can compete by optimizing cost, memory capacity and workload fit rather than by matching the entire general-purpose platform built around a dominant GPU supplier.

Positron also presented Titan as a planned system targeting up to 18.4 TB of memory per system, up to 32 trillion parameters per server and context windows exceeding 10 million tokens. Those figures describe a forward-looking roadmap, not a current commercial configuration.

What the summit changes for AI infrastructure builders

For teams designing or buying AI infrastructure, the practical lesson is to evaluate the whole path from model request to completed result. Accelerator throughput still matters, but it is only one part of the decision.

Memory bandwidth and capacity affect whether a large model can be served efficiently. Power-management software affects how much of a facility’s electrical budget becomes usable compute. Scale-up networking affects how well separate chips cooperate. Software integration determines whether a specialized accelerator can fit into an existing stack. Chip-design automation affects how quickly the next hardware generation can arrive.

That is why AI Infra Summit 2026 felt broader than a conventional chip launch. NVIDIA, d-Matrix, Qualcomm, Broadcom, Cognichip and Positron were addressing different layers of the same system—and the next infrastructure gains will come from making those layers work together.