Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pruning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

CSB-RNN: A Faster-Than-Realtime RNN Acceleration Framework with Compressed Structured Blocks

Recurrent Neural Networks (RNN) is widely applied to temporal sequence analysis, where real-time performance is usually in demand. However, RNN suffers a heavy computational workload as the model comes with a large weight matrix. To alleviate the pain, model compression (pruning) schemes have been proposed for RNN that pruning the redundant (near-zero) weight-values. On the one hand, the non-structured pruning methods achieve a considerable pruning rate while bringing the computational irregularity, which is un-friendly to parallel-hardware. On the other hand, the existing structured pruning methods consider the hardware parallelism; However, they suffer a poor pruning rate due to the restrict constraints on pruning structure. This paper presents CSB-RNN, an optimized full-stack RNN framework with the novel compressed structured block (CSB) technique. The CSB-pruned RNN model comes with both fine-granularity that benefits the pruning rate and regular structure that facilitates the hardware-parallelism. Further, we propose a novel hardware architecture for inferencing the CSB-pruned model. Different from conventional parallel hardware, this architecture solves the block-workload imbalance issue and achieves an over 95% hardware utilization. With the experiments on 10 RNN models in 5 application domains, the CSB-RNN realizes 7×-20× lossless compression and up to 50× acceptable lossy-compression, which is 2×-7× to the prior art. With the addition of the novel hardware, the compressed-RNN inference reaches a super real-time latency of 10-400µs with FPGA implementation.

Shi, Runbin↗

Local rules for fabricating allosteric networks

Mechanical properties of disordered networks can be significantly tailored by modifying a small fraction of their bonds. This procedure has been used to design and build mechanical metamaterials with a variety of responses. A long-range “allosteric” response, where a localized input strain at one site gives rise to a localized output strain at a distant site, has been of particular interest. This work presents an approach to incorporating allosteric responses in experimental systems by pruning disordered networks in situ. Previous work has relied on computer simulations to design and predict the response of such systems using a cost function where the response of the entire network to each bond removal is used at each step to determine which bond to prune. It is not feasible to follow such a design protocol in experiments where one has access only to local response at each site. This paper presents design algorithms that allow determination of what bonds to prune based purely on the local forces in the network without employing a cost function; using only local information, allosteric networks are designed in simulations and then built out of real materials. The results show that some pruning strategies work better than others when translated into an experimental system. A method is presented to measure local stresses experimentally in disordered networks. This approach is then used to implement pruning methods to design desired responses in situ. Results from these experiments confirm that the pruning methods are robust and work in a real laboratory material.

36 MATERIALS SCIENCE↗

O3BNN-R: An Out-Of-Order Architecture for High-Performance and Regularized BNN Inference

Binarized Neural Networks (BNN) have drawn tremendous attention due to significantly reduced computational complexity and memory demand. They have especially shown great potential in cost- and power-restricted domains, such as IoT and smart edge-devices, where reaching a certain accuracy bar is often sufficient, and real-time is highly desired.In this work, we demonstrate that the highly-condensed BNN model can be shrunk significantly further by dynamically pruning irregular redundant edges. Based on two new observations on BNN-specific properties, an out-of-order (OoO) architecture – O3BNN-R, can curtail edge evaluation in cases where the binary output of a neuron can be determined early. Similar to Instruction-Level-Parallelism(ILP), these fine-grained, irregular, runtime pruning opportunities are traditionally presumed to be difficult to exploit. In order to increase the pruning opportunities, we also optimize the training process by adding 2 regularization items in the loss function (1) for pooling pruning and (2) for threshold pruning. We evaluate our design on an FPGA platform using three well-known networks, including VggNet-16, AlexNet for ImageNet, and a VGG-like network for Cifar-10.

Geng, Tong↗

Mitigating Algorithmic Bias in Cancer Site Classification Models

Purpose Integrating artificial intelligence in cancer diagnostics has improved tumor classification beyond rule-based systems. Despite these advancements, these models may still encode demographic biases. We conducted a large-scale, applied bias-probing study of a deep learning–based cancer site classifier to quantify race information encoded in document embeddings. We then evaluated how performance changes when race-correlated embedding dimensions are removed in a post-training sensitivity analysis. Methods The cancer site classifier was trained using 3.5 million electronic cancer pathology reports from six of the National Cancer Institute's SEER registries. We trained a hierarchical self-attention network to generate 400-dimensional document embeddings. These embeddings were used to train two downstream, gradient-boosted decision tree classifiers: one to classify the cancer sites and another to predict racial categories. We identified overlapping features by intersecting the top 50 feature-importance rankings from the site and race models and computed their cumulative feature importance in each model. As a post hoc sensitivity analysis, we progressively pruned these overlapping dimensions, retrained the site model, and compared overall macro-F1 and accuracy, race-stratified macro-F1, and group fairness metrics on the basis of demographic parity and equalized odds before and after pruning. Results The analysis revealed minimal feature overlap between the cancer site and race prediction models, and the cumulative importance scores indicated a negligible influence of racial information on clinical predictions. Post-training pruning of overlapping features did not compromise the models' diagnostic accuracy, with a 0.07% loss in accuracy. Conclusion Our findings demonstrate that HiSAN-generated embeddings from SEER data can be used effectively in cancer site classification without significant demographic bias influencing the outcomes. Post-training pruning therefore functions as a practical audit and sensitivity check.

Shivanna, Abhishek [ORNL] (ORCID:0009000665228593)↗

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science↗

Binary Complex Neural Network Acceleration on FPGA

Being able to learn from complex data with phase information is imperative for many signal processing applications. Today’s real-valued deep neural networks (DNNs) have shown efficiency in latent information analysis but fall short when applied to the complex domain. Deep complex networks (DCN) , in contrast, can learn from complex data, but have high computational costs; therefore, they cannot satisfy the instant decision making requirements of many deployable systems dealing with short observations or short signal bursts. Recent, Binarized Complex Neural Network (BCNN), which integrates DCNs with binarized neural networks (BNN), shows great potential in classifying complex data in real-time. In this paper, we propose a structural pruning based accelerator of BCNN, which is able to provide more than 5000 frames/s inference throughput on edge devices. The high performance comes from both the algorithm and hardware sides. On the algorithm side, we conduct structural pruning to the original BCNN models and obtain 20 × pruning rates with negligible accuracy loss; on the hardware side, we propose a novel 2D convolution operation accelerator for the binary complex neural network. Experimental results show that the proposed design works with over 90% utilization and is able to achieve the inference throughput of 5882 frames/s and 4938 frames/s for complex NIN-Net and ResNet-18 using CIFAR-10 dataset and Alveo U280 Board.

Peng, Hongwu↗

Uncontrolled Learning: Codesign of Neuromorphic Hardware Topology for Neuromorphic Algorithms

Neuromorphic computing has the potential to revolutionize future technologies and our understanding of intelligence, yet it remains challenging to realize in practice. The learning-from-mistakes algorithm, inspired by the brain's simple learning rules of inhibition and pruning, is one of the few brain-like training methods. This algorithm is implemented in neuromorphic memristive hardware through a codesign process that evaluates essential hardware trade-offs. While the algorithm effectively trains small networks as binary classifiers and perceptrons, performance declines significantly with increasing network size unless the hardware is tailored to the algorithm. This work investigates the trade-offs between depth, controllability, and capacity—the number of learnable patterns—in neuromorphic hardware. This highlights the importance of topology and governing equations, providing theoretical tools to evaluate a device's computational capacity based on its measurements and circuit structure. The findings show that breaking neural network symmetry enhances both controllability and capacity. Additionally, by pruning the circuit, neuromorphic algorithms in all-memristive circuits can utilize stochastic resources to create local contrasts in network weights. Through combined experimental and simulation efforts, the parameters are identified that enable networks to exhibit emergent intelligence from simple rules, advancing the potential of neuromorphic computing.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Polishing the Gold Standard: The Role of Orbital Choice in CCSD(T) Vibrational Frequency Prediction

While CCSD(T) with spin-restricted Hartree-Fock (RHF) orbitals has long been lauded for its ability to accurately describe closed-shell interactions, the performance of CCSD(T) on open-shell species is much more erratic, especially when using a spin-unrestricted HF (UHF) reference. Previous studies have shown improved treatment of open-shell systems when a non-HF set of molecular orbitals, like Brueckner or Kohn-Sham density functional theory (DFT) orbitals, is used as a reference. Inspired by the success of regularized orbital-optimized second-order Møller-Plesset perturbation theory (κ-OOMP2) orbitals as reference orbitals for MP3, we investigate the use of κ-OOMP2 orbitals and various DFT orbitals as reference orbitals for CCSD(T) calculations of the corrected ground-state harmonic vibrational frequencies of a set of 36 closed-shell (29 neutrals, 6 cations, 1 anion) and 59 open-shell diatomic species (38 neutrals, 15 cations, 6 anions). The aug-cc-pwCVTZ basis set is used for all calculations. The use of κ-OOMP2 orbitals in this context alleviates difficult cases observed for both UHF orbitals and OOMP2 orbitals. Removing two multireference systems and 12 systems with ambiguous experimental data leaves a pruned data set. Overall performance on the pruned data set highlights CCSD(T) with a B97 orbital reference (CCSD(T):B97), CCSD(T) with a κ-OOMP2 orbital reference (CCSD(T):κ-OOMP2), and CCSD(T) with a B97M-rV orbital reference (CCSD(T):B97M-rV) with RMSDs of 8.48 cm -1 , and 8.50 cm -1 , and 8.75 cm -1 respectively, outperforming CCSD(T):UHF by nearly a factor of 5. Moreover, the performance on the closed- and open-shell subsets shows these methods are able to treat open-shell and closed-shell systems with comparable accuracy and robustness. CCSD(T) with RHF orbitals is seen to improve upon UHF for the closed-shell species, while spatial symmetry breaking in a number of restricted open-shell HF (ROHF) references leads CCSD(T) with ROHF reference orbitals to exhibit the poorest statistical performance of all methods surveyed for open-shell species. The use of κ-OOMP2 orbitals has also proven useful in diagnosing multireference character that can hinder the reliability of CCSD(T).

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Isochronic development of cortical synapses in primates and mice

Abstract The neotenous, or delayed, development of primate neurons, particularly human ones, is thought to underlie primate-specific abilities like cognition. We tested whether synaptic development follows suit—would synapses, in absolute time, develop slower in longer-lived, highly cognitive species like non-human primates than in shorter-lived species with less human-like cognitive abilities, e.g., the mouse? Instead, we find that excitatory and inhibitory synapses in the maleMus musculus(mouse) andRhesus macaque(primate) cortex form at similar rates, at similar times after birth. Primate excitatory and inhibitory synapses and mouse excitatory synapses also prune in such an isochronic fashion. Mouse inhibitory synapses are the lone exception, which are not pruned and instead continuously added throughout life. The monotony of synaptic development clocks across species with disparate lifespans, experiences, and cognitive abilities argues that such programs are likely orchestrated by genetic events rather than experience.

Science & Technology - Other Topics↗

Waveform processing using neural network algorithms on the front-end electronics

In a multi-channel radiation detector readout system, waveform sampling, digitization, and raw data transmission to the data acquisition system constitute a conventional processing chain. The deposited energy on the sensor is estimated by extracting peak amplitudes, area under pulse envelopes from the raw data, and starting times of signals or time of arrivals. However, such quantities can be estimated using machine learning algorithms on the front-end Application-Specific Integrated Circuits (ASICs), often termed as “edge computing”. Edge computation offers enormous benefits, especially when the analytical forms are not fully known or the registered waveform suffers from noise and imperfections of practical implementations. In this work, we aim to predict peak amplitude from a single waveform snippet whose rising and falling edges containing only 3 to 4 samples. We thoroughly studied two well-accepted neural network algorithms, Multi-Layer Perceptron (MLP) and Convolutional Neural Network (CNN) by varying their model sizes. Further, to better fit front-end electronics, neural network model reduction techniques, such as network pruning methods and variable-bit quantization approaches, were also studied. By combining pruning and quantization, our best performing model has the size of 1.5 KB, reduced from 16.6 KB of its full model counterpart. It can reach mean absolute error of 0.034 comparing to that of a naive baseline of 0.135. Such parameter-efficient and predictive neural network models established feasibility and practicality of their deployment on front-end ASICs.

47 OTHER INSTRUMENTATION↗

Adapting In Situ Accelerators for Sparsity With Granular Matrix Reordering

Neural network (NN) inference is an essential part of modern systems and is found at the heart of numerous applications ranging from image recognition to natural language processing. In situ NN accelerators can efficiently perform NN inference using resistive crossbars, which makes them a promising solution to the data movement challenges faced by conventional architectures. Although such accelerators demonstrate significant potential for dense NNs, they often do not benefit from sparse NNs, which contain relatively few non-zero weights. Processing sparse NNs on in situ accelerators results in wasted energy to charge the entire crossbar where most elements are zeros. To address this limitation, this paper proposes Granular Matrix Reordering (GMR): a preprocessing technique that enables an energy-efficient computation of sparse NNs on in situ accelerators. GMR reorders the rows and columns of sparse weight matrices to maximize the crossbars' utilization and minimize the total number of crossbars needed to be charged. The reordering process does not rely on sparsity patterns and incurs no accuracy loss. Finally, GMR achieves an average of 28% and up to 34% reduction in energy consumption over seven pruned NNs across four different pruning methods and network architectures.

97 MATHEMATICS AND COMPUTING↗

DESIVAST: Catalogs of Low-redshift Voids Using Data from the DESI Data Release 1 Bright Galaxy Survey

We present three separate void catalogs created using a volume-limited sample of the DESI Data Release 1 Bright Galaxy Survey. We use the algorithms VoidFinder and V 2 to construct void catalogs out to a redshift of z = 0.24. Excluding voids affected by the boundaries of the survey, we obtain 1489 voids with VoidFinder, 389 with V 2 using REVOLVER pruning, and 297 with V 2 using VIDE pruning. Comparing our catalogs with overlapping Sloan Digital Sky Survey void catalogs, we find generally consistent void properties but significant differences in the void volume overlap, which we attribute to differences in the galaxy selection and survey masks. These catalogs are suitable for studying the variation in galaxy properties with cosmic environment and for cosmological studies.

79 ASTRONOMY AND ASTROPHYSICS↗

Impact of biochar amendments on soil water and plant uptake dynamics under different cropping systems

Abstract Application of biochar amendments in agricultural systems has received much attention in recent years. In this study, we assess the 5‐year impacts of biochar application on soil water and plant interactions for an irrigated fresh market tomato ( Solanum lycopersicum ) and a rainfed pasture ( Poaceae ) cropping system. In particular, we focus on three varieties of locally produced biochar from agricultural waste materials—almond shell, walnut shell, and almond pruning residues that are pyrolyzed using a mobile pyrolysis unit. We used the soil hydrological model HYDRUS‐1D to explicitly track seasonal and annual soil water fluxes through changes in water retention, drainage, evaporation, and plant water uptake under biochar application. Modeling results show that the application of biochar at 5% increased soil water availability within the top 20 cm for a rainfed system, irrespective of biochar amendment type. This is clearly indicative of higher plant water uptake and greater water use efficiency (WUE) under biochar application. In contrast, a similar biochar amendment for the irrigated system did not affect WUE, instead reducing seasonal soil evaporation loss and thereby reducing irrigation demand. In both cropping systems, year‐to‐year variability in precipitation significantly impacted the total amount of water saved under biochar application with certain amendments retaining more water than others. Given that biochar application increased water retention irrespective of cropping systems, we further used a simple approach to determine yield trade‐off, if any, between control and biochar treatments. Our economic balance clearly demonstrates that the water saved by amending soil with biochar does not offset the yield disparity if compensated with carbon credits and therefore, application of biochar should be actively considered for both its direct and indirect benefits to potential greenhouse gas mitigation (e.g., diverting orchard waste from open burning), water savings, and soil health.

54 ENVIRONMENTAL SCIENCES↗

A survey of techniques for optimizing transformer inference

Recent years have seen a phenomenal rise in the performance and applications of transformer neural networks. The family of transformer networks, including Bidirectional Encoder Representations from Transformer (BERT), Generative Pretrained Transformer (GPT) and Vision Transformer (ViT), have shown their effectiveness across Natural Language Processing (NLP) and Computer Vision (CV) domains. Transformer-based networks such as ChatGPT have impacted the lives of common men. However, the quest for high predictive performance has led to an exponential increase in transformers' memory and compute footprint. Researchers have proposed techniques to optimize transformer inference at all levels of abstraction. Further, this paper presents a comprehensive survey of techniques for optimizing the inference phase of transformer networks. We survey techniques such as knowledge distillation, pruning, quantization, neural architecture search and lightweight network design at the algorithmic level. We further review hardware-level optimization techniques and the design of novel hardware accelerators for transformers. We summarize the quantitative results on the number of parameters/FLOPs and the accuracy of several models/techniques to showcase the tradeoff exercised by them. We also outline future directions in this rapidly evolving field of research. We believe that this survey will educate both novice and seasoned researchers and also spark a plethora of research efforts in this field.

97 MATHEMATICS AND COMPUTING↗

Neural architecture codesign for fast physics applications

We develop a pipeline to streamline neural architecture codesign for physics applications to reduce the need for ML expertise when designing models for novel tasks. Our method employs neural architecture search and network compression in a two-stage approach to discover hardware efficient models. This approach consists of a global search stage that explores a wide range of architectures while considering hardware constraints, followed by a local search stage that fine-tunes and compresses the most promising candidates. We exceed performance on various tasks and show further speedup through model compression techniques such as quantization-aware-training and neural network pruning. We synthesize the optimal models to high level synthesis code for FPGA deployment with the hls4ml library. Additionally, our hierarchical search space provides greater flexibility in optimization, which can easily extend to other tasks and domains. We demonstrate this with two case studies: Bragg peak finding in materials science and jet classification in high energy physics, achieving models with improved accuracy, smaller latencies, or reduced resource utilization relative to the baseline models.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Flexible AI Models for Grid Resilience

The rapid growth in size and complexity of artificial intelligence (AI) and machine learning (ML) models has led to increased energy demands, posing a threat to the reliability of the existing power grid. This project addresses the challenge of highly intermittent and energy-intensive inference workloads by (1) developing fidelity-adaptive neural networks capable of dynamic response to grid conditions and (2) integrating these networks with power flow simulations to assess their impact on power grid reliability. We will explore both top-down and bottom-up approaches to create hierarchies of submodels that provide a controlled trade-off between power draw and prediction accuracy. The top-down method utilizes NN pruning to reduce a flagship model into progressively smaller, energy-efficient variants. The bottom-up approach employs geometrically principled weight setting strategies to construct depth-efficient models from the ground up. A real-time hardware-in-the-loop (HIL) platform will be developed to simulate a scaled AC power grid, integrating live AI workload power draw and enabling dynamic model switching in response to grid feedback. This work will provide a novel framework for evaluating the impact of flexible AI/ML workloads on grid performance and establish new methodologies for energy-aware computing in data centers. The outcomes will demonstrate that adaptive AI/ML can play a critical role in improving grid stability while advancing NREL's leadership in energy-efficient computing research.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Three-dimensional phenotyping of peach tree-crown architecture utilizing terrestrial laser scanning

Tree training systems for temperate fruit have been developed throughout history by pomologists to improve light interception, fruit yield, and fruit quality. These training systems direct crown and branch growth to specific configurations. Quantifying crown architecture could aid the selection of trees that require less pruning or that naturally excel in specific growing/training system conditions. Regarding peaches [Prunus persica (L.) Batsch], access tools such as branching indices have been developed to characterize tree-crown architecture. However, the required branching data (BD) to develop these indices are difficult to collect. Traditionally, BD have been collected manually, but this process is tedious, time-consuming, and prone to human error. These barriers can be circumnavigated by utilizing terrestrial laser scanning (TLS) to obtain a digital twin of the real tree. TLS generates three-dimensional (3D) point clouds of the tree crown, wherein every point contains 3D coordinates (x, y, z). To facilitate the use of these tools for peach, we selected 16 young peach trees scanned in 2021 and 2022. These 16 trees were then modeled and quantified using the open-source software TreeQSM. As a result, “in silico” branching and biometric data for the young peach trees were calculated to demonstrate the capabilities of TLS phenotyping of peach tree-crown architecture. The comparison and analysis of field measurements (in situ) and in silico BD, biometric data, and quantitative structural model branch uncertainty data were utilized to determine the reconstructive model’s reliability as a source substitute for field measurements. Mean average deviation when comparing young tree (YT) height was approx. 5.93%, with crown volume was approx. 13.26% across both 2021 and 2022. All point clouds of the YTs in 2022 showed residuals lower than 12 mm to cylinders fitted to all branches, and mean surface coverage greater than 40% for both the trunk and primary branching orders.

09 BIOMASS FUELS↗