From Scale-Out to Scale-Up: Rethinking AI Interconnects for the Supernode Era

AI is pushing data centers into the Scale-Up era. As GPU clusters expand, network performance has become a critical factor in unlocking compute efficiency. OCS is emerging as a key enabler for high-performance AI Supernodes.

For modern AI workloads, communication is no longer a secondary consideration—it is part of the computing process itself. Every synchronization, parameter update, and collective communication operation depends on a high-performance interconnect capable of delivering massive bandwidth with extremely low latency.

Against this backdrop, Optical Circuit Switching (OCS) is emerging as a key enabler for next-generation AI infrastructure. By providing flexible, high-capacity optical connectivity, OCS helps unlock the full performance potential of AI Supernodes while supporting future scalability.

Today, we’d like to look at an important application of OCSenabling high-performance networking inside AI Supernodes.

Scale-Out Expands Capacity. Scale-Up Unlocks Performance.

Both scaling methodologies are essential for modern AI infrastructure, yet they solve fundamentally different challenges.

Scale-Out focuses on expanding computing capacity horizontally. More servers, more GPUs, and more racks are added into a larger cluster, allowing AI systems to train larger models or process more inference requests simultaneously. This architecture has long been the dominant approach for cloud computing because it offers excellent scalability and operational flexibility.

However, larger clusters also generate exponentially increasing communication traffic. As the number of GPUs grows, the amount of synchronization required between distributed nodes rises dramatically, making network efficiency a limiting factor rather than computing power itself.

Scale-Up addresses this challenge from a different perspective.

Instead of simply increasing the number of independent computing nodes, Scale-Up enlarges the computing domain by tightly coupling multiple GPUs into one unified high-performance system. GPUs communicate through ultra-high-bandwidth, ultra-low-latency interconnects, allowing them to behave much more like components inside a single supercomputer than isolated servers connected over a conventional network.

The objective is not to increase the number of GPUs alone, but to maximize how efficiently every GPU can work together.

Rather than competing approaches, Scale-Out and Scale-Up complement each other. Large AI infrastructure typically relies on Scale-Up to optimize communication inside each computing domain, while Scale-Out connects multiple high-performance domains into massive AI clusters capable of training trillion-parameter foundation models.

Scale-Out vs. Scale-Up——why does AI need new scaling methods?

Why AI Supernodes Matter

The rapid growth of foundation models has significantly increased communication intensity during AI training. Large Language Models (LLMs), multimodal AI systems, and recommendation models continuously exchange gradients, synchronize parameters, and execute collective communication operations such as AllReduce.

In traditional distributed AI clusters, independent servers communicate through Ethernet or InfiniBand networks. Although these technologies provide high bandwidth, communication still traverses multiple network hops, switch layers, and packet forwarding processes. As GPU counts increase into the thousands, communication overhead becomes a significant portion of the total training time.

In many large-scale AI deployments, GPUs are no longer waiting for computation—they are waiting for data.

This communication bottleneck directly reduces GPU utilization, increases training cost, and limits overall cluster efficiency.

AI Supernodes were introduced to overcome this limitation.

Instead of treating every server as an independent computing unit, AI Supernodes tightly integrate multiple computing nodes into one unified high-performance computing domain. GPUs inside the Supernode communicate through extremely fast interconnects with deterministic latency and significantly higher bandwidth.

The result is:

1. Higher bandwidth

2. Lower latency

3. Better GPU collaboration

4. Higher overall computing efficiency

As AI models continue to scale, maximizing the capability of each computing domain is becoming just as important as expanding the overall cluster.ng the capability of each computing domain is becoming just as important as expanding the overall cluster.

OCS Brings More Flexibility to Scale-Up Networks

As more GPUs are integrated into each Supernode, networking efficiency becomes a critical factor.

OCS introduces an all-optical interconnect that establishes dedicated optical paths between computing resources.

Unlike conventional packet-based networking, optical circuits can be dynamically reconfigured to adapt to changing workloads, providing greater flexibility for AI training and inference.

how does OCS connect to AI Supernode?

Unlike conventional packet-based networking, optical circuits can be dynamically reconfigured to adapt to changing workloads, providing greater flexibility for AI training and inference.

Key Advantages of OCS in AI Supernodes

OCS delivers more than high-speed connectivity—it enables a more flexible and efficient AI network.

1. Dynamic reconfiguration – Optical paths can be adjusted in milliseconds to match changing AI workloads.

2. Higher availability – Rapid path switching helps isolate failures and maintain service continuity.

3. Lower power consumption – A simplified optical architecture reduces energy usage.

4. Ultra-low latency – Dedicated optical circuits eliminate packet forwarding delays, maximizing GPU communication efficiency.

5. Scalable design – High-port-density switching supports the continued growth of AI Supernodes.

Advantages of OCS Optical Switches in Scale-Up Networks

Industrial Expectation

The evolution of AI infrastructure is increasingly shifting the industry’s attention from simply deploying more computing resources to enabling those resources to work together more efficiently.

As GPU performance continues to improve, communication efficiency is becoming one of the primary determinants of overall AI system performance. Future AI data centers will not only require faster interconnects but also more intelligent, adaptive network architectures capable of responding to constantly changing workloads.

In this context, optical networking is moving from a supporting role to a strategic foundation of AI infrastructure.

Among the emerging technologies shaping this transformation, Optical Circuit Switching stands out for its ability to combine ultra-high bandwidth, low latency, flexible topology, and energy efficiency within a unified optical interconnect architecture.

As AI Supernodes become an increasingly common building block of next-generation AI data centers, OCS is expected to play a growing role in enabling scalable, resilient, and high-performance computing infrastructure capable of supporting the next wave of foundation models and AI innovation.