Engineering PapersSearch

SEARCH · Engineering Papers

Results for “bandwidth efficiency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Intelligent Pixel Detectors: Towards a Radiation Hard ASIC with On-Chip Machine Learning in 28 nm CMOS

Detectors at future high energy colliders will face enormous technical challenges. Disentangling the unprecedented numbers of particles expected in each event will require highly granular silicon pixel detectors with billions of readout channels. With event rates as high as 40 MHz, these detectors will generate petabytes of data per second. To enable discovery within strict bandwidth and latency constraints, future trackers must be capable of fast, power efficient, and radiation hard data-reduction at the source. We are developing a radiation hard readout integrated circuit (ROIC) in 28nm CMOS with on-chip machine learning (ML) for future intelligent pixel detectors. We will show track parameter predictions using a neural network within a single layer of silicon and hardware tests on the first tape-outs produced with TSMC. Preliminary results indicate that reading out featurized clusters from particles above a modest momentum threshold could enable using pixel information at 40 MHz.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

On-chip probabilistic inference for charged-particle tracking at the sensor edge

Modern scientific instruments operate under increasingly extreme constraints on bandwidth, latency, and power. Inference at the sensor edge determines experimental data collection efficiency by deciding which information to save for further analysis. Particle tracking detectors at the Large Hadron Collider exemplify this challenge: pixelated silicon sensors generate rich spatiotemporal ionization patterns, yet most of this information is discarded due to data-rate limitations. Concurrently, advancements in co-design tools provide rapid turn-around for incorporating machine learning into application-specific integrated circuits, motivating designs for particle detectors with new integrated technologies. We demonstrate that neural networks embedded in the front-end electronics can infer charged-particle kinematic parameters from a single silicon layer. We regress hit positions and incident angles with calibrated uncertainties, while satisfying stringent constraints on numerical precision, latency, and silicon area. Our results establish a path toward probabilistic inference directly at the edge, opening new opportunities for intelligent sensing in high-rate scientific instruments.

Das, Arghya Ranjan [Purdue U.] (ORCID:000000018451

Approaches for the Simulation of Coupled Processes in Evolving Fractured Porous Media Enabled by Exascale Computing

Models have historically represented fractured porous media with continuum descriptions that characterize the media using bulk parameters. The impact of small-scale features is not captured in these models, although they may be controlling the performance of subsurface applications. Pore-scale models can simulate processes in small-scale features by representing the pore space geometry explicitly but are computationally expensive for large domains. The alternative multiscale approach entails the combination of pore-scale and continuum-scale descriptions in a single framework. We use Chombo-Crunch, a computational capability that discretizes complex geometries with an adaptive, embedded boundary method to contrast these two approaches. Chombo-Crunch takes advantage of recent computational performance and memory bandwidth improvements resulting from the emergence of exascale computing resources. These combined improvements enable the efficient simulation of reactive transport in fractured media with a high degree of fidelity and the ability to capture the control small-scale processes exert on the overall medium evolution.

42 ENGINEERING

A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing

Energy efficiency of training and inferencing with large neural network models is a critical challenge facing the future of sustainable large-scale machine learning workloads. This paper introduces an alternative strategy, called phantom parallelism, to minimize the net energy consumption of traditional tensor (model) parallelism, the most energy-inefficient component of large neural network training. The approach is presented in the context of feed-forward network architectures as a preliminary, but comprehensive, proof-of-principle study of the proposed methodology. We derive new forward and backward propagation operators for phantom parallelism, implement them as custom autograd operations within an end-to-end phantom parallel training pipeline and compare its parallel performance and energy-efficiency against those of conventional tensor parallel training pipelines. Formal analyses that predict lower bandwidth and FLOP counts are presented with supporting empirical results on up to 256 GPUs that corroborate these gains. Experiments are shown to deliver ∼50% reduction in the energy consumed to train FFNs using the proposed phantom parallel approach when compared with conventional tensor parallel methods. Additionally, the proposed approach is shown to train smaller phantom models to the same model loss on smaller GPU counts as larger tensor parallel models on larger GPU counts offering the possibility for even greater energy savings.

Seal, Sudip [ORNL] (ORCID:0000000332330656)

Large-scale real-time signal processing in physics experiments: the ALICE TPC FPGA pipeline

For LHC Run 3, the ALICE Time Projection Chamber was upgraded to operate in continuous readout mode. Interaction rates of up to 50 kHz in Pb-Pb collisions require real-time processing of more than 3 TB s -1 of raw detector data. This requirement is met by a custom FPGA-based processing pipeline that performs the complete front-end data treatment fully in-stream, including common-mode correction, pedestal subtraction, ion-tail filtering, zero suppression, and dense data packing. A central element of the design is a highly parallel common-mode correction algorithm operating directly on the streaming data. It robustly identifies signal-free readout channels on a time-bin basis and applies pad-dependent scaling to compensate for local variations in capacitive coupling in the GEM readout. In combination with pedestal subtraction and ion-tail filtering, this enables accurate baseline restoration under extreme high-occupancy conditions, preventing signal loss while efficiently suppressing noise prior to zero suppression. The pipeline operates continuously at the full detector bandwidth and reduces the raw input rate of approximately 3 TB s -1 to about 900 GBps for Pb-Pb collisions at the target interaction rate. Overall, it represents a large-scale FPGA-based real-time signal-processing implementation for high-energy physics detector readout.

Digital signal processing (DSP)

Multiplexed color centers in a silicon photonic cavity array

Entanglement distribution is central to the modular scaling of quantum processors and establishing quantum networks. Color centers with telecom-band transitions and long spin coherence times are suitable candidates for long-distance entanglement distribution. However, high-bandwidth memory-enhanced quantum communication is limited by high-yield, scalable creation of efficient spin-photon interfaces. Here, we develop a silicon photonics platform consisting of arrays of bus-coupled cavities. The coupling to a common bus waveguide enables simultaneous access to individually addressable cavity-enhanced T center arrays. We demonstrate frequency-multiplexed operation of two T centers in separate photonic crystal cavities. In addition, we investigate the cavity enhancement of a T center through hybridized modes formed between physically distant cavities. Our results show that bus-coupled arrays of cavity-enhanced color centers could enable efficient on-chip and long-distance entanglement distribution.

Komza, Lukasz

HPC Campaign Management: Remote data access with user-defined error bound using ADIOS and ZFP

Remote access to large-scale scientific datasets, like those generated by combustion simulations or other high-performance computing (HPC) applications, presents a significant challenge. Downloading entire datasets is often impractical due to their size and the bandwidth limitations of typical networks. To address this challenge, we propose a novel approach that enables efficient remote access to large datasets distributed across multiple facilities. Our method enables technologies to download only the data values of a select variable, in a select region of interest, to a user-defined accuracy. For this purpose, we extended the ADIOS IO library to provide read functions with user-defined accuracy, a remote data server that understands multidimensional selections of specific variables, steps and accuracy from an ADIOS dataset, and which uses lossy compression on the remote site to reduce the data to be transferred back to the client. In addition, our extension of the ADIOS library collects metadata from multiple datasets in small files called Campaign Archives, which can be shared among project participants on any HPC, cloud or laptop, and which can easily facilitate the discovery of content and pointers to the data location as well as remote access to the data by local tools as if data was local. This feature called Campaign Management, enables a group of scientists to manage related datasets stored in multiple files, across multiple facilities as if it was in a single file/database. We demonstrate the effectiveness of our approach using a 1.5 TB dataset from the S3D combustion simulation on Frontier at the Oak Ridge Leadership Facility. Even a single variable from this dataset, at 64 GB, is too large to be processed on a standard laptop. We show two different reading patterns for 2D plots and 3D visualization, with careful settings that a scientist studying combustion data would do and show that running the same Python scripts on Frontier directly takes comparable time than running them on the local laptop with remote access to the data on Frontier.

Podhorszki, Norbert [ORNL] (ORCID:000000019647542X

Non-resonant Bragg scattering four-wave mixing at near-visible wavelengths in low-confinement silicon nitride waveguides

Quantum state coherent frequency conversion processes—such as Bragg-scattering four-wave mixing (BSFWM)—hold promise as a flexible technique for networking heterogeneous and distant quantum systems. In this Letter, we demonstrate BSFWM within an extended (1.2-m) low-confinement silicon nitride waveguide and show that this system has the potential for near-unity frequency conversion in visible and near-visible wavelength ranges. Using sensitive classical heterodyne laser spectroscopy at low optical powers, we characterize the Kerr coefficient (∼1.55 W −1 m −1 ) and linear propagation loss (∼0.0175 dB/cm) of this non-resonant waveguide system, revealing a record-high nonlinear figure of merit (NFM = γ / α ≈ 3.85 W −1 ) for BSFWM of near-visible light in non-resonant silicon nitride waveguides. We predict how, at high yet achievable on-chip optical powers, this NFM would yield a comparatively large frequency conversion efficiency, opening the door to near-unity flexible frequency conversion without cavity enhancement and resulting bandwidth constraints.

Jaber, Nicholas (ORCID:000900070106809X)

A compact and portable gamma-ray spectrometer (GRASP) for inertial confinement fusion and basic science experiments

A compact and portable gamma-ray spectrometer has been designed to diagnose different components of the inertial confinement fusion-relevant γ-ray spectrum with energies between ∼3.7–17.9 MeV. The system is designed to be as compact as possible for convenient transportation and fielding in diagnostic ports on the OMEGA laser, the National Ignition Facility, and other photon-source facilities. The system consists of a conversion foil for Compton scattering in front of four magnetic spectrometer “arms,” each covering a different energy range and constructed out of cylindrical permanent magnet Halbach arrays. Monte Carlo simulations have been used to optimize and assess the performance of the conversion foil, and COSY INFINITY ion-optical simulations have been used to optimize the spectrometer magnets. The performance of the design is assessed for a simulated direct-drive γ-ray spectrum. Spanning its total γ-ray energy bandwidth and using a 1.7 mm thick boron conversion foil, the system’s total energy resolution and efficiency are ∼15.8%–4.5% and 5.4 × 10−7–3.7 × 10−7e−/γ, respectively, with room for improvement. Spectral γ-ray measurements will provide guidance to the inertial confinement fusion program toward achieving high-energy gain relevant to inertial fusion energy and enable new measurement capabilities for basic discovery science.

Instruments & Instrumentation

ExtremeMETA: High-speed Lightweight Image Segmentation Model by Remodeling Multi-channel Metamaterial Imagers

Deep neural networks (DNNs) have heavily relied on traditional computational units, such as CPUs and GPUs. However, this conventional approach brings significant computational burden, latency issues, and high power consumption, limiting their effectiveness. This has sparked the need for lightweight networks such as ExtremeC3Net. Meanwhile, there have been notable advancements in optical computational units, particularly with metamaterials, offering the exciting prospect of energy-efficient neural networks operating at the speed of light. Yet, the digital design of metamaterial neural networks (MNNs) faces precision, noise, and bandwidth challenges, limiting their application to intuitive tasks and low-resolution images. In this study, we proposed a large kernel lightweight segmentation model, ExtremeMETA. Based on ExtremeC3Net, our proposed model, ExtremeMETA maximized the ability of the first convolution layer by exploring a larger convolution kernel and multiple processing paths. With the large kernel convolution model, we extended the optic neural network application boundary to the segmentation task. To further lighten the computation burden of the digital processing part, a set of model compression methods was applied to improve model efficiency in the inference stage. The experimental results on three publicly available datasets demonstrated that the optimized efficient design improved segmentation performance from 92.45 to 95.97 on mIoU while reducing computational FLOPs from 461.07 MMacs to 166.03 MMacs. The large kernel lightweight model ExtremeMETA showcased the hybrid design’s ability on complex tasks.

large convolution kernel

Optimal Control Strategy With Efficiency and Reliability Improvement for Offshore DC Microgrids

Offshore microgrids, due to their remote location and lack of external energy support, face significant challenges in wide-range load operation and maintenance. Consequently, efficiency and reliability are critical concerns for converters in offshore dc microgrids. This article presents an optimal control strategy aimed at enhancing both efficiency and reliability. A normalized nonlinear relationship between power loss and thermal stress of a paralleled converter is first established. Based on this, a dual-objective optimization function with an active weight function as well as a system overall performance index is established. The active weight function dynamically adjusts the control priority based on converter efficiency and switching device thermal stress. Then, the optimal power-sharing strategy is derived by the Lagrange multiplier method with the proposed optimal function. Additionally, to accommodate a wide load range, an optimal selection strategy for operating converter combinations is proposed, requiring only low-bandwidth communication. Experiment verification is given to validate the effectiveness of the proposed control strategy. The experiment results demonstrate that the proposed control strategy can improve the overall performance of offshore microgrids by optimizing efficiency and reliability.

24 POWER TRANSMISSION AND DISTRIBUTION

Impact of quantum well thickness on efficiency loss in InGaN/GaN LEDs: Challenges for thin-well designs

We investigate the impact of quantum well (QW) thickness on efficiency loss in c-plane InGaN/GaN LEDs using a small-signal electroluminescence technique. Multiple mechanisms related to efficiency loss are independently examined, including injection efficiency, carrier density vs current density relationship, phase space filling, quantum-confined Stark effect, and Coulomb enhancement. An optimal QW thickness of around 2.7 nm in these InGaN/GaN LEDs was determined for QWs having constant In composition. Despite improved control of deep-level defects and lower carrier density at a given current density, LEDs with thin QWs still suffer from an imbalance in enhancement effects on the radiative and intrinsic Auger–Meitner recombination coefficients. The imbalance in enhancement effects results in a decline in internal quantum efficiency and radiative efficiency with decreasing QW thickness at low current density in LEDs with QW thicknesses below 2.7 nm. Here, we also investigate how LED modulation bandwidth varies with QW thickness, identifying the key trends and their implications for device performance.

36 MATERIALS SCIENCE

In Situ MOF Pyrolysis Construction of Hierarchical Porous Co‐Nanoparticles/Carbon Cloth Composites for Enhanced Electromagnetic Wave Shielding and Absorption

The development of high-performance electromagnetic protection materials integrating broadband absorption and effective shielding capabilities is hindered by challenges in simultaneously optimizing multiple electromagnetic properties through conventional material designs. This study pioneers a hierarchical porous Co nanoparticle/carbon cloth (Co/CC) composite via controlled annealing of a Co-MOF precursor on carbon cloth. The Co-MOF served a dual role as both magnetic source and pore-forming agent, enabling in situ generation of uniformly dispersed Co nanoparticles and creation of abundant pores/interfaces on the CC fibers during pyrolysis. This unique architecture synergistically enhanced dielectric loss (via interfacial/dipolar polarization) and magnetic loss (via natural resonance, exchange interactions, and eddy currents), significantly improving impedance matching. The hierarchical pores further functioned as integrated “absorption–reflection” units for efficient electromagnetic energy attenuation. Consequently, the Co/CC composite annealed at 800°C (Co/CC-800) achieves minimum reflection loss (−40.69 dB) and 120% effective absorption bandwidth extension (6.16 GHz) as a filler, and exhibits superior electromagnetic interference shielding effectiveness (46.66 dB) as an integrated component. Significantly, Co/CC-800 demonstrated robust photothermal and electrothermal conversion capabilities, ensuring operational stability in ice-covered and humid harsh environments. This work pioneers a pore-structure-mediated strategy to harmonize dielectric–magnetic synergy, providing a new paradigm for designing advanced multifunctional electromagnetic protection materials.

dielectric‐magnetic synergy

Fast and Scalable FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices with Application to Linear Inverse Problems Governed by Autonomous Dynamical Systems

In this work, we present an efficient and scalable algorithm for performing matrix-vector multiplications (matvecs) for block Toeplitz matrices. Such matrices, which are shift-invariant with respect to their blocks, arise in the context of solving inverse problems governed by autonomous systems, and time-invariant systems in particular. In this article, we consider inverse problems that infer unknown parameters from observational data of a linear time-invariant dynamical system given in the form of partial differential equations (PDEs). Matrix-free Newton-conjugate-gradient methods are often the gold standard for solving these inverse problems, but they require numerous actions of the Hessian on a vector. Matrix-free adjoint-based Hessian matvecs require solution of a pair of linearized forward/adjoint PDE solves per Hessian action, which may be prohibitive for large-scale inverse problems. Time invariance of the forward PDE problem leads to a block Toeplitz structure of the discretized parameter-to-observable (p2o) map defining the mapping from inputs (parameters) to outputs (observables) of the PDEs. This block Toeplitz structure enables us to exploit two key properties: (1) compact storage of the p2o map and its adjoint, and (2) efficient fast Fourier transform–based Hessian matvecs. The proposed algorithm is mapped onto large multi-GPU clusters and achieves more than 80% of peak bandwidth on NVIDIA A100 GPUs. Excellent weak scaling is shown for up to 48 A100 GPUs. For the targeted problems, the implementation executes Hessian matvecs within fractions of a second, which is orders of magnitude faster than can be achieved by conventional matrix-free Hessian matvecs via forward/adjoint PDE solves.

97 MATHEMATICS AND COMPUTING

FedCSpc: A Cross-Silo Federated Learning System With Error-Bounded Lossy Parameter Compression

Cross-Silo federated learning is widely used for scaling deep neural network (DNN) training over data silos from different locations worldwide while guaranteeing data privacy. Communication has been identified as the main bottleneck when training large-scale models due to large-volume model parameters and gradient transmission across public networks with limited bandwidth. Most previous works focus on gradient compression, while limited work tries to compress parameters that can not be ignored and extremely affect communication performance during the training. Here, to bridge this gap, we propose FedCSpc: an efficient cross-silo federated learning system with an XAI-driven adaptive parameter compression strategy for large-scale model training. Our work substantially differs from existing gradient compression techniques due to the distinct data features of gradient and parameter. The key contributions of this paper are fourfold. (1) Our designed FedCSpc proposes to compress the parameter during the training using the state-of-the-art error-bounded lossy compressor – SZ3. (2) We develop an adaptive compression error bound adjustment algorithm to guarantee the model accuracy effectively. (3) We exploit an efficient approach to utilize the idle CPU resources of clients to compress the parameters. (4) We perform a comprehensive evaluation with a wide range of models and benchmarks on a GPU cluster with 65 GPUs. Results show that FedCSpc can achieve the same model accuracy as FedAvg while reducing the data volume of parameters and gradients in communication by up to 7.39× and 288×, respectively. With 32 clients on a 4 Gb size model, FedCSpc significantly outperforms FedAvg in wall-clock time in the emulated WAN environment (at the bandwidth of 1 Gbps or lower without loss of generality).

SZ3

Femtosecond wavelength-tunable laser system using gain managed nonlinear amplifier

Optical parametric amplifiers require a seed energy at the spectral range of interest for amplification. In the case of pulsed laser systems, seed pulse characteristics such as spectral energy density and spectral uniformity influence laser system design and performance. In this letter, we present the design and modeling of a few micro-joules, wavelength-tunable, femtosecond parametric amplifier system. We employ a gain-managed nonlinear fiber amplifier output as the seed pulse. This seed pulse is relatively uniform across its bandwidth from 1010 to 1180 nm. Its spectral energy density averages 600 pJ/nm, 30 times higher than conventional counterparts. The overall amplification gain is 175, and the conversion efficiency is 17.5%. The inherent quasi-linear chirp of the seed pulse is utilized to alter the output wavelength. Our tunable femtosecond laser is an enabling technology for high-demand applications such as time-resolved spectroscopy, nonlinear and advanced material research, and biomedical imaging and microscopy modalities.

47 OTHER INSTRUMENTATION

Pushing the limits of NAND technology scaling with ferroelectrics

Artificial intelligence (AI) continues to drive transformative advancements across various industries. The data-intensive nature of AI training (and inferencing) has resulted in the generation of unprecedented volumes of data with machine-generated content surpassing human-generated data by more than 100-fold in 2025. Efficiently managing this data influx necessitates advanced digital storage technologies. However, traditional NAND flash memory, which is critical for supporting data flows in AI systems—alongside high-bandwidth memory, for AI training—faces fundamental scaling limitations as it approaches the 1000-layer milestone, encompassing more than 40 trillion transistors. This article delves into the potential of hafnia-based ferroelectric materials as a breakthrough solution to these challenges. Recent advancements indicate that the intrinsic limitations of ferroelectric field-effect transistors (FEFETs) can be mitigated through material and device-level engineering. These advancements enable FEFETs to meet the stringent density, reliability, and scalability requirements of future three-dimensional NAND technology. The role of ferroelectrics in addressing NAND scaling challenges and expanding storage capabilities presents a promising avenue for meeting the storage demands of the AI-driven era.

3D NAND