Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model efficiency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A Linear-Complexity Tensor Butterfly Algorithm for Compressing High-Dimensional Oscillatory Integral Operators

This paper presents a multilevel tensor compression algorithm called tensor butterfly algorithm for efficiently representing large-scale and high-dimensional oscillatory integral operators, including Green's functions for wave equations and integral transforms such as Radon transforms and Fourier transforms. The proposed algorithm leverages a tensor extension of the so-called complementary low-rank property of existing matrix butterfly algorithms. The algorithm partitions the discretized integral operator tensor into subtensors of multiple levels and factorizes each subtensor at the middle level as a Tucker-type interpolative decomposition, whose factor matrices are formed in a multilevel fashion. For a d-dimensional (d > 1) integral operator discretized into a 2d-mode tensor with n2d entries, the overall CPU time and memory requirement scale as O(nd), in stark contrast to the O(nd log n) complexity of existing matrix algorithms such as matrix butterfly algorithms and fast Fourier transforms (FFTs), where n is the number of points per direction. When comparing with other tensor algorithms such as quantized tensor train (QTT), the proposed algorithm also shows superior CPU and memory performance for tensor contraction. Remarkably, the tensor butterfly algorithm can efficiently model high-frequency Green's function interactions between two unit cubes, each spanning 512 wavelengths per direction, which represents problems of scale over 512× larger than that existing butterfly algorithms can handle, with the same amount of computation resources. On the other hand, for a problem representing 64 wavelengths per direction, which is the largest size existing algebraic matrix algorithms can handle, our tensor butterfly algorithm exhibits 200x speedups and 30× memory reduction compared with existing ones. Moreover, the tensor butterfly algorithm also permits O(nd)-complexity FFTs and Radon transforms up to d = 6 dimensions.

Kielstra, P Michael↗

Integrated Simulation of PIP-II at Fermilab

We describe progress towards a community software ecosystem for efficient modeling of the Fermilab PIP-II complex for design validation and virtual test stands.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Optimization and Commercialization of the Juvenile Eel/Lamprey Acoustic Transmitter and Micro-battery (CRADA 477)

Through collaboration with partners under this CRADA, we continued to optimize the Eel/Lamprey Acoustic Transmitter (ELAT) of the Juvenile Salmon Acoustic Telemetry System (JSATS) to enhance its capability to track sensitive species and early life stages of fish with 3D and sub-meter accuracy. The transmitter firmware was improved to offer new functionalities, and the circuit design was revised to provide significantly more accurate frequency transmission. The transmitter's transducer was replaced with a shorter, more energy-efficient model. Additionally, the manufacturing processes for the transmitter and the microbattery used in it were improved. The outcomes of this work help advance understanding of migration timing and behaviors, habitat use, and survival rates of these species, supporting more informed management decisions for new and existing hydroelectric facilities and better designs for new hydropower systems that minimize or avoid environmental impacts.

13 HYDRO ENERGY↗

Electrospinner Upgrades for Nanofiber Production

Electrospinning is an inexpensive and flexible method for producing nanofibers. Nanofibers are highly adaptable with potential applications in accelerator target systems, air and water filtration, and biomedicine. This project aims to upgrade and test an existing roll-to-roll electrospinner that is more economical for industrial nanofiber production. Varying the diameter of nanofiber can change its functional properties. The current unit cannot adjust spinneret-collector separation, which determines nanofiber diameter. The electro-spinneret channel also does not have lateral adjustment capabilities for precise alignment. Lastly, the viscosities of our polymer solutions have not been quantified. Solution injection into the channel is presently inconsistent because of imperfect nozzle sizing, resulting in waste and decreased efficiency. Modeling was done with Siemens NX CAD software, and viscosity was measured using a Brookfield DVE-LV viscometer. A dual scissor lift design for channel displacement, powered by a dual-shaft DC motor coupled to precision lead screws, was approved and construction was started. Preliminary channel modifications also were initiated. Going forward, the scissor lift system will be integrated and evaluated with our electrospinner, along with further channel modifications. Viscosity measurements of polyvinyldimethylformamide (PVDF) were recorded with inconsistent results due to inadequate testing conditions. Future viscosity trials must be completed in accordance with testing requirements. The optimization and commercialization of roll-to-roll electrospinner units can make nanofiber production more feasible for many new industries and consumer products.

Black, Niko↗

An Optimization-Based Coupling of Reduced Order Models with an Efficient Reduced Adjoint Basis Generation Approach

Optimization-based coupling (OBC) is an attractive alternative to traditional Lagrange multiplier approaches in multiple modeling and simulation contexts. However, application of OBC to time-dependent problems has been hindered by the computational cost of finding the stationary points of the associated Lagrangian, which requires primal and adjoint solves. This issue can be mitigated by using OBC in conjunction with computationally efficient reduced order models (ROMs). To demonstrate the potential of this combination, in this paper, we develop an optimization-based ROM-ROM coupling for a transient advection-diffusion transmission problem. We pursue the “optimize-then-reduce” path toward solving the minimization problem at each time step and solve reduced space adjoint system of equations, where the main challenge in this formulation is the generation of adjoint snapshots and reduced bases for the adjoint systems required by the optimizer. One of the main contributions of the paper is a new technique for an efficient adjoint snapshot collection for gradient-based optimizers in the context of optimization-based ROM-ROM couplings. In conclusion, we present numerical studies demonstrating the accuracy of the approach along with comparison between various approaches for selecting a reduced order basis for the adjoint systems, including decay of snapshot energy, average iteration counts, and timings.

coupled problems↗

HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models

As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.

Herron, Emily [ORNL] (ORCID:0000000273008172)↗

Efficient First-Order Algorithms for Large-Scale, Non-Smooth Maximum Entropy Models with Application to Wildfire Science

Maximum entropy (MaxEnt) models are a class of statistical models that use the maximum entropy principle to estimate probability distributions from data. Due to the size of modern data sets, MaxEnt models need efficient optimization algorithms to scale well for big data applications. State-of-the-art algorithms for MaxEnt models, however, were not originally designed to handle big data sets; these algorithms either rely on technical devices that may yield unreliable numerical results, scale poorly, or require smoothness assumptions that many practical MaxEnt models lack. In this paper, we present novel optimization algorithms that overcome the shortcomings of state-of-the-art algorithms for training large-scale, non-smooth MaxEnt models. Our proposed first-order algorithms leverage the Kullback–Leibler divergence to train large-scale and non-smooth MaxEnt models efficiently. For MaxEnt models with discrete probability distribution of n elements built from samples, each containing m features, the stepsize parameter estimation and iterations in our algorithms scale on the order of O(mn) operations and can be trivially parallelized. Moreover, the strong ℓ1 convexity of the Kullback–Leibler divergence allows for larger stepsize parameters, thereby speeding up the convergence rate of our algorithms. To illustrate the efficiency of our novel algorithms, we consider the problem of estimating probabilities of fire occurrences as a function of ecological features in the Western US MTBS-Interagency wildfire data set. Our numerical results show that our algorithms outperform the state of the art by one order of magnitude and yield results that agree with physical models of wildfire occurrence and previous statistical analyses of wildfire drivers.

Physics↗

Computationally efficient subglacial drainage modelling using Gaussian process emulators: GlaDS-GP v1.0

Subglacial drainage models represent water flow at the ice–bed interface through coupled distributed and channelized systems to determine water pressure, discharge, and drainage system geometry. While they are used to understand processes such as the relationship between surface melt and ice flow, the number of uncertain model parameters and the computational cost of running models makes it difficult to adequately explore the high-dimensional parameter space and evaluate uncertainty in model predictions. Here, we develop Gaussian process (GP) emulators that make fast predictions with associated uncertainty of subglacial drainage model outputs. Using a truncated principal component (PC) basis representation, we construct a GP emulator for diurnally averaged subglacial water pressure. We also explore emulation of scalar variables describing drainage efficiency and configuration. We train the emulators using ensembles of up to 512 simulations varying eight parameters of the Glacier Drainage System (GlaDS) model on a synthetic domain intended to represent an ice-sheet margin. The emulators make predictions ∼ 1000 times faster than GlaDS simulations, with errors <3 % for the water pressure field and ∼ 5 %–9 % for drainage efficiency and configuration. We apply the emulators to explore the eight-dimensional parameter space by computing variance-based parameter sensitivity indices, finding that three parameters (ice flow coefficient, bed bump aspect ratio, and the subglacial cavity system conductivity) explain 90 % of the variance in modelled water pressure in response to parameter changes. The GP emulator approach described here is well suited to integrating observational data with models to make calibrated, credible predictions of subglacial drainage.

58 GEOSCIENCES↗

JuTrack: A Julia package for auto-differentiable accelerator modeling and particle tracking

Efficient accelerator modeling and particle tracking are key for the design and configuration of modern particle accelerators. In this work, we present JuTrack, a nested accelerator modeling package developed in the Julia programming language and enhanced with compiler-level automatic differentiation (AD). With the aid of AD, JuTrack enables rapid derivative calculations in accelerator modeling, facilitating sensitivity analyses and optimization tasks. Here we demonstrate the effectiveness of AD-derived derivatives through several practical applications, including sensitivity analysis of space-charge-induced emittance growth, nonlinear beam dynamics analysis for a synchrotron light source, and lattice parameter tuning of the future Electron-Ion Collider (EIC). Through the incorporation of automatic differentiation, this package opens up new possibilities for accelerator physicists in beam physics studies and accelerator design optimization.

43 PARTICLE ACCELERATORS↗

Synergizing human expertise and AI efficiency with language model for microscopy operation and automated experiment design

With the advent of large language models (LLMs), in both the open source and proprietary domains, attention is turning to how to exploit such artificial intelligence (AI) systems in assisting complex scientific tasks, such as material synthesis, characterization, analysis and discovery. Here, we explore the utility of LLMs, particularly ChatGPT4, in combination with application program interfaces (APIs) in tasks of experimental design, programming workflows, and data analysis in scanning probe microscopy, using both in-house developed APIs and APIs given by a commercial vendor for instrument control. We find that the LLM can be especially useful in converting ideations of experimental workflows to executable code on microscope APIs. Beyond code generation, we find that the GPT4 is capable of analyzing microscopy images in a generic sense. At the same time, we find that GPT4 suffers from an inability to extend beyond basic analyses for more in-depth technical experimental design. We argue that an LLM specifically fine-tuned for individual scientific domains can potentially be a better language interface for converting scientific ideations from human experts to executable workflows. Such a synergy between human expertise and LLM efficiency in experimentation can open new doors for accelerating scientific research, enabling effective experimental protocols sharing in the scientific community.

97 MATHEMATICS AND COMPUTING↗

Divergent carbon use efficiency-growth rate tradeoff in popular biological growth models

Carbon use efficiency (CUE) is an important trait emerging from processes regulating biological growth. CUE can be computed either based on the growth of structural biomass or total biomass divided by substrate uptake rate. Nonequilibrium thermodynamics and observations suggest that, for an exponentially growing population of cells, structural biomass CUE should first increase, then peak, and finally decrease with specific growth rate; meanwhile, total biomass CUE increases asymptotically with specific growth rate. We compared predictions from six popular models that are often used for plant and microbial growth in existing ecosystem models. We found that, for an exponentially growing population of biological cells, (1) the source-driven Pirt and Compromise models predict that structural biomass CUE increase asymptotically with growth rate; (2) the apparent sink-driven modified Droop model predicts that structural biomass CUE decreases with growth rate; and (3) the sink-driven variable internal storage model and two dynamic energy budget models predict that structural biomass CUE first increases, then peaks, and finally decreases with growth rate. Moreover, the modified Droop model predicts that total biomass CUE is constant with growth rate, while all other five models predict that total biomass CUE increases with growth rate asymptotically. For non-exponential biological growth, we show that there is no static relationship between total biomass CUE or structural biomass CUE with respect to either growth rate or temperature. Therefore, we contend that biological growth models should explicitly represent interactions between substrate acquisition, substate transformation, and maintenance respiration to better capture observed CUE dynamics, and the sink-driven model should be preferred for general ecosystem biogeochemistry modeling.

Tang, Jinyun [Lawrence Berkeley National Laborator↗

A dynamic 2D Borehole Thermal Energy Storage (BTES) model for enhanced computational efficiency

Progressing toward a future increasingly reliant on renewable energy sources, the development of effective, durable energy storage solutions becomes essential to balance supply and demand fluctuations. Borehole Thermal Energy Storage (BTES) is a long-duration thermal energy storage technology that captures excess heat generated from renewable energy sources and stores it underground for later use, enabling the efficient utilization of sustainable energy. This approach is particularly valuable in district energy networks when integrated with Ground Source Heat Pumps (GSHP) to provide stable heating and cooling. However, traditional three-dimensional (3D) numerical models of BTES systems demand extensive computational resources, limiting their practicality for real-time and large-scale applications. This study introduces a novel two-dimensional (2D) modeling approach that reduces computational costs while maintaining high accuracy. By employing a radial ring-based discretization method, the model simulates heat injection, retention, and retrieval dynamics over seasonal cycles. A new thermal-mass weighted-average temperature parameter is introduced to evaluate the performance of BTES systems. Model validation against FEFLOW simulations demonstrates a 17-fold improvement in computational speed compared to traditional Computational Fluid Dynamics (CFD) models while achieving a mean absolute percentage error (MAPE) of 2 % during charging and 4 % during discharging. Additionally, a trade-off analysis between computational efficiency and accuracy is conducted, ensuring the model's applicability for real-world scenarios. The findings of this research contribute to the development of computationally efficient BTES models, facilitating better optimization, control, and integration into renewable energy systems. This work provides a foundation for further studies in techno-economic analysis, multi-year performance evaluation, and real-time operational strategies for BTES applications, supporting a more sustainable energy future.

2D modeling↗

Conditional Pseudo-Reversible Normalizing Flow for Surrogate Modeling in Quantifying Uncertainty Propagation

We introduce a conditional pseudo-reversible normalizing flow (PR-NF) that directly learns conditional probability distributions from noisy physical models to efficiently quantify both forward and inverse uncertainty propagation. Traditional surrogate modeling approaches approximate only the deterministic component of physical models, requiring separate noise characterization and computationally expensive sampling methods for inverse problems. Here, in this work, we develop the conditional PR-NF model to directly learn and efficiently generate samples from the conditional probability density functions (PDFs). The training process utilizes dataset consisting of input-output pairs without requiring prior knowledge about the noise and the function. Once trained, our model efficiently generates samples from conditional PDFs for any input within the training domain. Moreover, the pseudo-reversibility feature allows for the use of fully connected neural network architectures, which simplifies the implementation and enables theoretical analysis. We provide a rigorous convergence analysis of the conditional PR-NF model, showing its ability to converge to the target conditional PDF using the Kullback−Leibler divergence. To demonstrate the effectiveness of our method, we apply it to several benchmark tests and a real-world geologic carbon storage problem.

97 MATHEMATICS AND COMPUTING↗

Scalable training of trustworthy and energy-efficient predictive graph foundation models for atomistic materials modeling: a case study with HydraGNN

We present our work on developing and training scalable, trustworthy, and energy-efficient predictive graph foundation models (GFMs) using HydraGNN, a multi-headed graph convolutional neural network architecture. HydraGNN expands the boundaries of graph neural network (GNN) computations in both training scale and data diversity. It abstracts over message passing algorithms, allowing both reproduction of and comparison across algorithmic innovations that define nearest-neighbor convolution in GNNs. This work discusses a series of optimizations that have allowed scaling up the GFMs training to tens of thousands of GPUs on datasets consisting of hundreds of millions of graphs. Our GFMs use multitask learning (MTL) to simultaneously learn graph-level and node-level properties of atomistic structures, such as energy and atomic forces. Using over 154 million atomistic structures for training, we illustrate the performance of our approach along with the lessons learned on two state-of-the-art US Department of Energy (US-DOE) supercomputers, namely the Perlmutter petascale system at the National Energy Research Scientific Computing Center and the Frontier exascale system at Oak Ridge Leadership Computing Facility. The HydraGNN architecture enables the GFM to achieve near-linear strong scaling performance using more than 2000 GPUs on Perlmutter and 16,000 GPUs on Frontier.

97 MATHEMATICS AND COMPUTING↗

Intelligent Surrogate Model Development: Boosting Computational Efficiency for Autonomous Control of Advanced Reactors

Advanced reactors promise enhanced safety, greater efficiency, and waste reductions. To fully realize these benefits, it is crucial to address the need for autonomous or semi-autonomous control systems that require fewer operators. This research primarily supports the MARVEL autonomous control system, which requires real-time operation. However, the current RELAP5 reactor thermal hydraulic transient simulation is excessively time-consuming. Therefore, this study aims to leverage deep learning techniques to develop a surrogate model, providing a more efficient and accurate alternative for real-time performance. The model was trained using a combination of one-timestep prediction and scheduled sampling. It was then used for recursive prediction of the reactor state. This developed surrogate model significantly improves computational efficiency, achieving a 12 times acceleration.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Simulated moving bed-inspired method for continuous adsorptive denitrogenation of model fuel

Efficient, continuous routes for removing nitrogen-containing compounds from hydrothermal liquefaction-derived synthetic aviation fuel are needed to enable direct blending with conventional jet fuels. Here, we report a simulated moving bed-inspired process for adsorptive denitrogenation of a model fuel. Unlike conventional simulated moving bed systems, which are designed for sharp separations between similar solutes, this approach was run deliberately outside the classical separation region so that both pyridine and indole were removed together from the hydrocarbon stream. Alcohol solvents were used to regenerate the silica adsorbent, maintaining performance over extended operation and avoiding the downtime and energy demand associated with calcination. Under these conditions, the system demonstrates removal of more than 98% of nitrogen while cutting solvent use by 28% compared to batch operation. Classical modeling tools predicted column concentration profiles even in this nontraditional regime, suggesting a straightforward path to scaling. Together, these results motivate solvent-efficient, continuous denitrogenation strategies that could be integrated with biorefinery processes.

Adsorption↗

Investigating resource-efficient neutron/gamma classification ML models targeting eFPGAs

There has been considerable interest and resulting progress in implementing machine learning (ML) models in hardware over the last several years from the particle and nuclear physics communities. A big driver has been the release of the Python package, hls4ml, which has enabled porting models specified and trained using Python ML libraries to register transfer level (RTL) code. So far, the primary end targets have been commercial field-programmable gate arrays (FPGAs) or synthesized custom blocks on application specific integrated circuits (ASICs). However, recent developments in open-source embedded FPGA (eFPGA) frameworks now provide an alternate, more flexible pathway for implementing ML models in hardware. These customized eFPGA fabrics can be integrated as part of an overall chip design. In general, the decision between a fully custom, eFPGA, or commercial FPGA ML implementation will depend on the details of the end-use application. In this work, we explored the parameter space for eFPGA implementations of fully-connected neural network (fcNN) and boosted decision tree (BDT) models using the task of neutron/gamma classification with a specific focus on resource efficiency. We used data collected using an AmBe sealed source incident on Stilbene, which was optically coupled to an OnSemi J-series silicon photomultiplier (SiPM) to generate training and test data for this study. We investigated relevant input features and the effects of bit-resolution and sampling rate as well as trade-offs in hyperparameters for both ML architectures while tracking total resource usage. The performance metric used to track model performance was the calculated neutron efficiency at a gamma leakage of 10 -3 . The results of the study will be used to aid the specification of an eFPGA fabric, which will be integrated as part of a test chip.

47 OTHER INSTRUMENTATION↗