Engineering Papers⌕ Search

Engineering topics

Hattink, Maarten

Publications and source records attributed to Hattink, Maarten.

Scaling comb-driven resonator-based DWDM silicon photonic links to multi-Tb/s in the multi-FSR regime

The use of chip-based micro-resonator Kerr frequency combs in conjunction with dense wavelength-division multiplexing (DWDM) enables massively parallel intensity-modulated direct-detection data transmission with low energy consumption. Resonator-based modulators and filters used in such systems can limit the number of usable wavelength channels due to practical constraints on the maximum achievable free spectral range (FSR). In this work, we introduce the design of multi-Tb/s comb-driven resonator-based silicon photonic links by leveraging the multi-FSR regime. We demonstrate the viability of the link architecture with yield estimates that are supported by extensive wafer-scale measurements of 704 micro-resonators fabricated in a commercial complementary metal–oxide–semiconductor foundry. We show that a 2.80 Tb/s link is realizable with a ≥6 σ yield (∼99.999%), and that aggregate bandwidths of 3.76 Tb/s and 4.72 Tb/s are possible if yield targets are relaxed (3 σ and 1 σ , respectively). All designs represent a 1.94−3.28× boost to aggregate link bandwidth while maintaining BER≤10 −10 performance, with a theoretical bandwidth of 10.51 Tb/s being possible for sufficiently robust resonators. We use high-speed BER measurements to inform co-optimization of data rate and aggressor spacing ( λ ag ), limiting any additional loss-based power penalties to off-resonance insertion loss (IL) and routing loss. This work demonstrates that, through the multi-FSR regime, there is a clear path toward Kerr comb-driven ultra-broadband, high bandwidth silicon photonic links that can support next-generation data centers and high-performance computers.

James, Aneek (ORCID:0000000262527807)↗

Distributed deep learning training using silicon photonic switched architectures

The scaling trends of deep learning models and distributed training workloads are challenging network capacities in today’s datacenters and high-performance computing (HPC) systems. We propose a system architecture that leverages silicon photonic (SiP) switch-enabled server regrouping using bandwidth steering to tackle the challenges and accelerate distributed deep learning training. In addition, our proposed system architecture utilizes a highly integrated operating system-based SiP switch control scheme to reduce implementation complexity. To demonstrate the feasibility of our proposal, we built an experimental testbed with a SiP switch-enabled reconfigurable fat tree topology and evaluated the network performance of distributed ring all-reduce and parameter server workloads. The experimental results show up to 3.6× improvements over the static non-reconfigurable fat tree. Our large-scale simulation results show that server regrouping can deliver up to 2.3× flow throughput improvement for a 2× tapered fat tree and a further 11% improvement when higher-layer bandwidth steering is employed. The collective results show the potential of integrating SiP switches into datacenters and HPC systems to accelerate distributed deep learning training.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Optically connected memory for disaggregated data centers

Recent advances in integrated photonics enable the implementation of reconfigurable, high-bandwidth, and low energy-per-bit interconnects in next-generation data centers. We propose and evaluate an Optically Connected Memory (OCM) architecture that disaggregates the main memory from the computation nodes in data centers. OCM is based on micro-ring resonators (MRRs), and it does not require any modification to the DRAM memory modules. We calculate energy consumption from real photonic devices and integrate them into a system simulator to evaluate performance. Here, our results show that (1) OCM is capable of interconnecting four DDR4 memory channels to a computing node using two fibers with 1.02 pJ energy-per-bit consumption and (2) OCM performs up to 5.5× faster than a disaggregated memory with 40G PCIe NIC connectors to computing nodes.

97 MATHEMATICS AND COMPUTING↗