Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Adaptive Array”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning

Introduction: Neuromorphic Materials

The explosive growth in data collection and the need to process it efficiently, as well as the desire to automate increasingly complex tasks in transportation, medical care, manufacturing, security and many other fields have motivated a growing interest in neuromorphic computing. Unlike the binary, transistorbased ON/OFF logic gates and separate logic and memory functionalities employed in digital computing, neuromorphic computing is inspired by animal brains that use interconnected synapses and neurons to perform processing, storage and transmission of information at the same location, while only consuming ~20 W or less of power. Motivated by the brain’s efficiency, adaptability, self-learning and resiliency qualities, neuromorphic computing can be broadly defined as an approach to processing and storing information using hardware and algorithms inspired by models of biological neural systems. Present research in neuromorphic computing encompasses approaches that vary significantly in their degree of neuro-inspiration, from systems that only incorporate features such as asynchronous, event-driven operation or use crossbar arrays of non-volatile memory (NVM) elements to accelerate deep neural networks (DNNs), to designs that embrace the extreme parallelism, sparsity, reconfigurability, adaptability, complexity and stochasticity observed in nervous systems. The term ‘neuromorphic’ computing is often credited to Carver Mead, who in the 1980s investigated Si-based analog electronics to replicate functions of the animal retina. Earlier important advances in this field include the work of Frank Rosenblatt, who proposed the concept of the perceptron, Bernard Widrow, who used this concept to build one of the first analog neural networks, the Adaline and many other researchers (see ref. 6 for an historical perspective on neuromorphic computing). With the recent increase in the use of artificial intelligence and large language models, and rising concerns over the associated energy costs, interest in neuromorphic hardware has expanded rapidly. According to some estimates, driven largely by the drastic growth in the training use of artificial intelligence (AI) models using the current computing architectures, the energy cost of computing is projected to reach the energy supply worldwide by 2045. Furthermore, while this is not a realistic outcome, it means that, if more efficient computing technologies are not developed -- soon -- the world will soon become one where demand for energy and market constraints limit the continued increase of societal access to AI and cloud services from data centers. Data centers used for training and use of these models consume hundreds of terawatt hours of electricity, already past 4% of the US electricity demand.

Circuits

Climate adaptation in Populus trichocarpa : key adaptive loci identified for stomata and leaf traits

We investigated adaptive genetic variation in Populus trichocarpa, a potential biofuel feedstock crop, to better understand how physiological traits may influence tolerance to water limitation. Our study focused on leaf and stomatal traits, given their roles in plant–water relations and adaptation. Using a diversity panel of over 1300 genotypes, we measured 14 leaf and stomatal traits under control (well-watered) and drought (water-limited) conditions. We conducted genome-wide association studies (GWAS), climate association analyses, and transcriptome (RNA-seq) profiling to identify genetic loci associated with phenotypic variation and adaptation. Stomatal traits, including size and density, were correlated with the climate of origin, with genotypes from more arid regions tending to have smaller but denser stomata. GWAS identified multiple loci associated with trait variation, including a major-effect region on chromosome 10 linked to stomatal size and abaxial contact angle. This locus overlapped with a tandem array of 3-ketoacyl-CoA synthase (KCS) genes and showed strong allele–climate and gene expression associations. Our findings reveal genetic and phenotypic variation consistent with local adaptation and suggest that future climates may favor alleles associated with smaller stomata, particularly under increasing aridity. This work provides insights into climate adaptation and breeding strategies for resilience in perennial crops.

Populus trichocarpa

Rate expressions and kinetic parameters for metal ferrites in relation to applications of fossil fuel conversion to hydrogen: Part 1 of 2

Here, the goal of the present work was to provide the necessary reaction emulation information to enable detailed process simulation of a chemical looping H 2 production system from fossil fuels using CaFe 2 O 4 . This specifically pertained to the necessary kinetic data, reaction model development, and model rate parameters required for reaction emulation in both reducing and oxidizing environments. A logical methodology was defined, which included discretization of the reaction network, establishing a core model for reaction emulation that could be adapted based on the system phenomena, and development of a rate parameter regression tool designed around the core model. An extensive array of data sets was acquired by which parametric regressions were performed. The work presented and tabulated a comprehensive set of rate parameters for the reduction and oxidation reactions of CaFe 2 O 4 and descendent phases of Ca 2 Fe 2 O 5 , FeO, Fe 3 O 4 , Fe, and CaO to emulate reaction behavior in a looping-based process environment. This included direct reduction using CH 4 , H 2 , and CO, and direct oxidation reactions with steam, CO 2 and O 2 . Dynamic equilibrium was quantified for reactions that could utilize H 2 O and CO 2 as soft oxidants to re-saturate lattice oxygen in the depleted structure/phases. The kinetics associated with the oxidative mechanisms with the soft oxidants were quantified and compared to those of the reducing counterparts. The analysis provided critical insight to emulate reactions for a process that seeks to use natural gas (NG) or other fossil fuels as a direct reductant for the end goal of H 2 production.

calcium ferrite oxygen carriers

Radiation Effects on Network on Chips (NoC) Laboratory Directed Research and Development (LDRD) project

This project was motivated by State-of-the-Art (SOTA) technology that incorporates Network on Chips (NOC) for efficient data communication across the various computer kernels. For example, on the AMD Versal Field Programmable Gate Arrays (FPGA), an NoC has been incorporated for fast data communication from the programmable logic and other computer kernels (processing system, adaptable intelligence engines, etc.). The radiation effects on the legacy technology of this FPGA, such as the programmable logic, are well understood, and established methods exist to measure cross-sections when new families/generations are released; however, newly incorporated technologies, such as the NoC, are not fully understood and could introduce new failure points into the mission space.

36 MATERIALS SCIENCE

Controlling cantilevered adaptive X-ray mirrors

Modeling the behavior of a prototype cantilevered X-ray adaptive mirror (held from one end) demonstrates its potential for use on high-performance X-ray beamlines. Similar adaptive mirrors are used on X-ray beamlines to compensate optical aberrations, control wavefronts and tune mirror focal distances at will. Controlled by 1D arrays of piezoceramic actuators, these glancing-incidence mirrors can provide nanometre-scale surface shape adjustment capabilities. However, significant engineering challenges remain for mounting them with low distortion and low environmental sensitivity. Finite-element analysis is used to predict the micron-scale full actuation surface shape from each channel and then linear modeling is applied to investigate the mirrors' ability to reach target profiles. Using either uniform or arbitrary spatial weighting, actuator voltages are optimized using a Moore–Penrose matrix inverse, or pseudoinverse, revealing a spatial dependence on the shape fitting with increasing fidelity farther from the mount.

47 OTHER INSTRUMENTATION

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]

SymbolNet: neural symbolic regression with adaptive dynamic pruning for compression

Abstract Compact symbolic expressions have been shown to be more efficient than neural network (NN) models in terms of resource consumption and inference speed when implemented on custom hardware such as field-programmable gate arrays (FPGAs), while maintaining comparable accuracy (Tsoi et al 2024 EPJ Web Conf. 295 09036). These capabilities are highly valuable in environments with stringent computational resource constraints, such as high-energy physics experiments at the CERN Large Hadron Collider. However, finding compact expressions for high-dimensional datasets remains challenging due to the inherent limitations of genetic programming (GP), the search algorithm of most symbolic regression (SR) methods. Contrary to GP, the NN approach to SR offers scalability to high-dimensional inputs and leverages gradient methods for faster equation searching. Common ways of constraining expression complexity often involve multistage pruning with fine-tuning, which can result in significant performance loss. In this work, we propose S y m b o l N e t , a NN approach to SR specifically designed as a model compression technique, aimed at enabling low-latency inference for high-dimensional inputs on custom hardware such as FPGAs. This framework allows dynamic pruning of model weights, input features, and mathematical operators in a single training process, where both training loss and expression complexity are optimized simultaneously. We introduce a sparsity regularization term for each pruning type, which can adaptively adjust its strength, leading to convergence at a target sparsity ratio. Unlike most existing SR methods that struggle with datasets containing more than O ( 10 ) inputs, we demonstrate the effectiveness of our model on the LHC jet tagging task (16 inputs), MNIST (784 inputs), and SVHN (3072 inputs).

Tsoi, Ho Fung (ORCID:0000000225502184)

FY 2026 Midyear Report: Seismic Monitoring of Underground Vibration Sources Using Distributed Acoustic Sensing and Seismometers

Safeguards-relevant temporal changes in underground facilities can be observed using geophysical monitoring techniques. Seismic waves, in particular, provide valuable insights into subsurface activities and can serve as an important tool for detecting anomalous events that may indicate containment breaches at geological repositories. This midyear report summarizes ongoing efforts to automatically and rapidly detect and locate anomalous vibration signals that could be indicative of potential containment breaches. Previous work during FY25 focused on compiling continuous seismic datasets from two underground sites and developing a database of continuous waveforms and ground-truth event data derived from multiple sensing modalities. Building on this foundation, we are adapting anomaly detection and geolocation algorithms to explore methods for monitoring underground activities using two relatively low-maintenance sensing technologies: a dense surface geophone array deployed at the Pleasant Gap mine in Pennsylvania, and a three-dimensional fiber-optic cable array for distributed acoustic sensing (DAS) installed in the subsurface at the Sanford Underground Research Facility (SURF) in South Dakota. This report summarizes work conducted during the first two quarters of FY26, during which we refined a dynamic power spectral density (PSD)-based detector, applied it independently to each geophone station, and then combined the per‑station detections with density-based spatial clustering of applications with noise (DBSCAN) to cluster events and produce spatial maps over a nine‑day interval. In addition, we outline plans for a field trial at the Waste Isolation Pilot Plant (WIPP) in New Mexico to compare traditional seismic monitoring approaches with DAS techniques and to evaluate the benefits of combined data analysis. Activities during the past two quarters have included the preparation and submission of a Field Test Plan to WIPP for approval, as well as submission to headquarters for review and feedback.

58 GEOSCIENCES

ICED: An Integrated CGRA Framework Enabling DFVS-Aware Acceleration

oarse-grained reconfigurable arrays (CGRAs) are a promising solution to enable energy-efficient acceleration of applications from different domains. By leveraging reconfiguration at the functional level, they can adapt to significantly different computational patterns. Existing CGRA mapping approaches extract instruction-level parallelism, exploit loop-pipelining opportunities, guarantee the data dependency, and target high throughput of a given loop. However, the recurrence data-dependency in the DFG and the mismatch between required and available computing/communication resources complicate the mapping, and might lead to significant unbalances in the utilization of the CGRA's tiles. This results in wasted power for tiles with low utilization. Applying dynamic voltage and frequency scaling (DVFS) can potentially solve this challenge and improve energy efficiency by adjusting voltage and frequency of different tiles independently. CGRAs have also been successful in accelerating data-dependent streaming applications. However, in these applications, the execution time of each kernel in the pipeline might dynamically vary depending on the characteristics of the input. This also leads to under-utilization of resources for the dynamically changing kernels that do not limit the application throughput. DVFS can also improve energy efficiency for these applications by dynamically changing the voltage and frequency levels of tiles that host non performance-constraining kernels. This paper proposes ICEDTEA -- an integrated DVFS-aware framework to map applications on CGRAs that support power islands. ICEDTEA proposes a CGRA architecture supporting DVFS islands at varying granularity (from a single tile to a group of tiles) and the related DVFS-aware compilation and mapping toolchain. ICEDTEA is the first work that introduces DVFS support for spatio-temporal CGRAs at power-island levels. The experimental evaluation shows that ICEDTEA improves average utilization by 2.3$\times$ and energy-efficiency by 1.32$\times$ over a conventional CGRA. With streaming applications, ICEDTEA improves energy efficiency by 1.12$\times$ over a state-of-the-art CGRA that introduces partial dynamic reconfiguration to adapt to variations in kernels' throughput.

Tan, Cheng

Development of a Superconducting Adaptive Gap Undulator Prototype for NSLS-II

The photon flux and brightness of synchrotron radiation, crucial parameters for any light source, vary significantly depending on the type of source employed. Among the 23 Insertion Device (ID) sources at the National Synchrotron Light Source II (NSLS-II) at Brookhaven National Lab (BNL), the 12-year-old 3m-long In-Vacuum Undulator (IVU20) stands out for its superior performance, although it no longer represents the cutting edge of technology. Recently, there has been a shift in focus towards developing next-generation sources, particularly Superconducting Undulators (SCUs), characterized by smaller gaps, shorter periods, and maximum lengths. However, despite ongoing research and development efforts, SCUs have yet to surpass their predecessors, the Cryogenic Permanent Magnet Undulators (CPMUs), in terms of performance. This is largely attributed to the limitations posed by traditional superconducting wire, as well as challenges in the design of the magnetic structure and vacuum chamber. In this paper, we aim to overcome such limitations through the development of a unique prototype Superconducting Adaptive Gap Undulator (SC-AGU) magnet core and vacuum chamber design. This paper will outline a novel technical approach aimed at constructing a compact prototype magnet array utilizing state-of-the-art superconducting wire technology. This approach provides a more efficient magnetic structure, allowing for enhanced magnetic field strength and stability.

36 MATERIALS SCIENCE

Teleseismic Network Association with GENIE

In this report we investigate adapting the Graph Neural Interpretation Engine (GENIE), an associator developed for three-component dense monitoring networks, to regional to teleseismic association using a sparse network of array stations. We expand GENIE’s input features to include first-P detection time, azimuth, and slowness estimates. Additionally, we include a probability of detection (PDET) term which measures a station’s likelihood of detecting an event. To assess each feature’s relative importance, we train four models, each using an increasing set of node features and find that the PDET models perform the best. We define two measures of event complexity which demonstrate that all GENIE model versions perform better than the standard backprojection stack.

47 OTHER INSTRUMENTATION

Shifting Between Compute and Memory Bounds: A Compression-Enabled Roofline Model

In the evolving landscape of high-performance computing, especially to fight the end of Moore’s Law and Dennard’s Scaling, the ability to shift between compute-bound and memory-bound states is critical for enhancing adaptability and flexibility to diverse system and domain-specific architectures. Such capability is vital for optimizing performance across distinguished hardware configurations, such as accelerators, memory hierarchies, and cache systems. Despite that ad hoc optimization techniques, such as compressed/approximate computation, have been enabled for compute-/data-intensive computing for improved performance in distinct hardware settings, there lacks an understanding of 1) the rational behind performance improvement; 2) capability of different optimizations; 3) what optimization to respond to specific computational and memory demands. This work proposes a compression-enabled roofline model to facilitate this adaptability with data compression techniques to balance and transform between computational and memory demands. This model enables applications to adjust in response to the specific strengths and limitations of the underlying hardware and system to optimize resource utilization. The effectiveness of this approach is demonstrated with matrix multiplication kernels on different input sizes, with turning on/off various compression techniques, including 1) low-precision floating point; 2) sparse matrix formulation; and 3) compressed arrays with ZFP. By reducing memory transfer volumes and cache misses and increasing data locality and computational intensity through compression, the specific roofline model can transform between compute and memory bounds to align more efficiently with system capabilities. This advancement not only improves overall performance but also maximizes adaptability in diverse computing environments.

Naraparaju, Ramasoumya [University of Washington]

Design of a robot-automated flat plate/reflection geometry x-ray diffraction setup for accelerated materials discovery and structural screening

Here, we report the design, construction, and automation of a flat plate sample loading, alignment, and data acquisition system for X-ray diffraction measurements in reflection geometry implemented at the Stanford Synchrotron Radiation Lightsource. The system is built onto a single platform, enabling facile transferability, and is compartmentalized into sample storage, sample transfer, and sample position/alignment segments. The core feature of this system is a six-axis robotic arm that offers a large range of highly reproducible and programable movements. The degrees of freedom of the robot arm enable adaptability in which movements can be modified to fit various beamline environments and sample configurations. Samples are housed on 3D printed sample mounts, which are arranged onto a 6 × 2 array of sample cassettes capable of holding 7 samples. Using sample mounts designed for solid oxide electrolysis button cells (SOECs), the maximum tray capacity is 84 samples, which can be aligned and run in ~ 24 hours with long exposure scans. The sample array is additionally capable of accommodating a range of sample sizes and geometries due to the rapid 3D printed fabrication. The components of the setup will be described in detail and performance will be demonstrated with a set of representative SOEC and XRD standard samples. Opportunities for future developments and integration with the automated setup are summarized.

08 HYDROGEN

Single nucleotide variants drive evolutionary phage-host arms race in anaerobic carbon dioxide-converting microbiome

Microbial bioconversions are shaped by environmental perturbations and the adaptation of resident microbiomes. Prokaryotes coexist with bacteriophages, yet their coevolutionary trajectories remain underexplored. Here, we investigate the effects of a cultivation vessel leak on an anaerobic consortium performing carbon dioxide reduction. Using time-series shotgun metagenomic sequencing, we reconstruct microbial and viral genomes to track community shifts. We further apply single-nucleotide variant profiling and CRISPR array analysis to monitor viral microdiversity and host defense mechanisms. After bioaugmentation restores bioconversion efficiency, the consortium undergoes pronounced restructuring, with new dominant taxa emerging from the rare biosphere. We identify patterns consistent with phage predation selectively removing certain species, while others exhibit resilience to infection. This shift aligns with a widespread viral outbreak and a transient increased frequency of single nucleotide variants in bacterial CRISPR–Cas defense genes. Expansion of CRISPR spacers further supports that CRISPR-mediated processes influence microbial resilience. Concurrently, phages infecting resilient hosts exhibited adaptive evolution, marked by high genetic heterogeneity. Selective pressure varies across their genomes, targeting infectivity genes and protospacer-adjacent motifs. These findings highlight a dynamic evolutionary arms race driven by the selection of beneficial genetic variants, providing a mechanistic framework for multi-omics investigations, and informing biotechnological applications, including phage-based microbiome manipulation.

Ghiotto, G

Field Programmable Gate Array-Based Reactor Protection Systems and Potential for Inclusion of Secure Elements to Improve Cybersecurity

For acceptable implementations of technologies like wireless communications, remote monitoring, etc., strong mitigations must be developed and evaluated to ensure that new attack pathways do not increase risk for Advanced Reactors. Secure Elements can be adopted and adapted for this purpose based on tamper resistance and cryptographic abilities, but research must be done to properly integrate into critical components such as FPGA-based Important to Safety systems in conjunction with current and future regulations on cyber security features in Advanced Reactors. Typically, the integration of a Secure Element happens during the POST and UEFI boot of a computing platform, performed by the Operating System, which is not possible with FPGAs because they do not include these firmware components. Work must be done to identify a reliable and secure method for integration in FPGA-based systems which lack Operating Systems and therefore complex boot procedures, system calls, etc.

97 MATHEMATICS AND COMPUTING

Field Programmable Gate Array-Based Reactor Protection Systems and Potential for Inclusion of Secure Elements to Improve Cybersecurity

For acceptable implementations of technologies like wireless communications, remote monitoring, etc., strong mitigations must be developed and evaluated to ensure that new attack pathways do not increase risk for Advanced Reactors. Secure Elements can be adopted and adapted for this purpose based on tamper resistance and cryptographic abilities, but research must be done to properly integrate into critical components such as FPGA-based Important to Safety systems in conjunction with current and future regulations on cyber security features in Advanced Reactors. Typically, the integration of a Secure Element happens during the POST and UEFI boot of a computing platform, performed by the Operating System, which is not possible with FPGAs because they do not include these firmware components. Work must be done to identify a reliable and secure method for integration in FPGA-based systems which lack Operating Systems and therefore complex boot procedures, system calls, etc.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Self-assembled reconfigurable pump architectures via magnetic colloidal swarms

Self-assembled swarms of interactive active units, which are adaptive and dynamically reconfigurable to accommodate different functionalities, represent a promising platform for the development of next-generation robotics. Here, we utilize the emergent collective behavior of active magnetic colloids confined in quasi-two-dimensional arrays of overlapping wells to demonstrate the self-organization of a colloidal swarm into a dynamic pump architecture capable of controlled transport of passive cargo particles. This dynamic architecture provides a global unidirectional looping flow pattern along the entire length of the system. We show that the flow direction of the dynamic swarm-based pump can be externally controlled by a phase shift of a driving magnetic field energizing the swarm. The experimental observations are supported by computational modeling based on phenomenological coarse-grained particle dynamics coupled to shallow-water Navier-Stokes hydrodynamics. In conclusion, our findings demonstrate how the emergent collective behavior of a swarm can be orchestrated into a desired functionality by exploiting the interplay between activity and confinement potentials.

36 MATERIALS SCIENCE