Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scalable deep learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Fission with Exotic Nuclei (Full Technical Report)

Despite its importance for stockpile stewardship or nuclear forensics, data on fission is fragmentary: most experiments have only been performed on stable actinide nuclei and are often incomplete, and theoretical simulations often contain many parameters hard to constrain, which results in large uncertainties. Consequently, nuclear libraries have large gaps for important isotopes. The goal of this project was to prepare to take advantage of the unprecedented yields of radioactive isotopes at the upcoming DOE Facility for Rare Ion Beams and to leverage recent progress in the field of machine learning to develop a comprehensive program of fission studies that could address some of the most pressing problems in nuclear fission over the next several decades. Our project leveraged synergies between experimental nuclear physics, nuclear theory, and data science expertise at LLNL. On the experimental front, our main objective was to develop in-house expertise for inverse kinematics reactions with relativistic beams as well as to field and test new detectors to perform correlated measurements of fission properties. On the theoretical side, the objective was to develop a novel, high-fidelity and scalable approach of fissionfragment calculations based on emulating computationally expensive, quantum-mechanical calculations of nuclear properties with deep neural networks.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Improving scalability of parallel CNN training by adaptively adjusting parameter update frequency

Synchronous SGD with data parallelism, the most popular parallelization strategy for CNN training, suffers from the expensive communication cost of averaging gradients among all workers. The iterative parameter updates of SGD cause frequent communications and it becomes the performance bottleneck. In this paper, we propose a lazy parameter update algorithm that adaptively adjusts the parameter update frequency to address the expensive communication cost issue. Our algorithm accumulates the gradients if the difference of the accumulated gradients and the latest gradients is sufficiently small. Here, the less frequent parameter updates reduce the per-iteration communication cost while maintaining the model accuracy. Our experimental results demonstrate that the lazy update method remarkably improves the scalability while maintaining the model accuracy. For ResNet50 training on ImageNet, the proposed algorithm achieves a significantly higher speedup (739.6 on 2048 Cori KNL nodes) as compared to the vanilla synchronous SGD (276.6) while the model accuracy is almost not affected (<0.2% difference).

97 MATHEMATICS AND COMPUTING↗

FrESCO: Framework for Exploring Scalable Computational Oncology

The National Cancer Institute (NCI) monitors population level cancer trends as part of its Surveillance, Epidemiology, and End Results (SEER) program. This program consists of state or regional level cancer registries which collect, analyze, and annotate cancer pathology reports. From these annotated pathology reports, each individual registry aggregates cancer phenotype information from electronic health records. This data is then used to create summary statistics about cancer incidence and mortality to facilitate population health monitoring. Extracting phenotypic information from these reports is a labor intensive task, requiring specialized knowledge about the reports and cancer. Automating the information extraction process from cancer pathology reports has the potential to improve data quality by extracting information in a consistent manner across registries. It can also improve patient outcomes by reducing the time from diagnosis, enabling rapid case ascertainment for clinical trials. Here we present FrESCO, a modular deep-learning natural language processing (NLP) library initially designed for extracting pathology information from clinical text documents. This repository is not solely limited to clinical medical text, but may also be used by researchers just getting started with NLP methods and those looking for a robust solution for their classification problems.

60 APPLIED LIFE SCIENCES↗

Vehicular Re-Identification from Uncontrolled Multiple Views

Vehicle re-identification (re-ID) across disparate sensing modalities remains a fundamental challenge for transportation research. In this work, we introduce a deep multi-view vehicle re-ID framework that leverages Siamese networks to compare pairs of vehicle images and produce matching scores, enabling robust association across drastically different viewpoints such as those from UAVs, surveillance cameras, and ground sensors. The model exploits convolutional neural networks to learn features that remain discriminative under changes in angle, distance, and illumination, supporting more generalizable re-ID performance. As part of this effort, we also developed an automated pipeline to synchronize roadside and UAV video streams, producing a multi-perspective dataset that complements preexisting real collections and a synthetic dataset generated in this study. Together, these contributions advance the capability to re-identify vehicles across wide viewing baselines; establish a foundation for scalable, reproducible research in vehicle re-ID; and open pathways for future applications, such as inferring routine behaviors, movement patterns, and daily habits of the individual associated with the vehicle.

convolutional neural networks↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

Transforming Agricultural Productivity with AI-Driven Forecasting: Innovations in Food Security and Supply Chain Optimization

Global food security is under significant threat from climate change, population growth, and resource scarcity. This review examines how advanced AI-driven forecasting models, including machine learning (ML), deep learning (DL), and time-series forecasting models like SARIMA/ARIMA, are transforming regional agricultural practices and food supply chains. Through the integration of Internet of Things (IoT), remote sensing, and blockchain technologies, these models facilitate the real-time monitoring of crop growth, resource allocation, and market dynamics, enhancing decision making and sustainability. The study adopts a mixed-methods approach, including systematic literature analysis and regional case studies. Highlights include AI-driven yield forecasting in European hydroponic systems and resource optimization in southeast Asian aquaponics, showcasing localized efficiency gains. Furthermore, AI applications in food processing, such as plasma, ozone and Pulsed Electric Field (PEF) treatments, are shown to improve food preservation and reduce spoilage. Key challenges—such as data quality, model scalability, and prediction accuracy—are discussed, particularly in the context of data-poor environments, limiting broader model applicability. The paper concludes by outlining future directions, emphasizing context-specific AI implementations, the need for public–private collaboration, and policy interventions to enhance scalability and adoption in food security contexts.

99 GENERAL AND MISCELLANEOUS↗

Robust High-Throughput Phenotyping with Deep Segmentation Enabled by a Web-Based Annotator

The abilities of plant biologists and breeders to characterize the genetic basis of physiological traits are limited by their abilities to obtain quantitative data representing precise details of trait variation, and particularly to collect this data at a high-throughput scale with low cost. Although deep learning methods have demonstrated unprecedented potential to automate plant phenotyping, these methods commonly rely on large training sets that can be time-consuming to generate. Intelligent algorithms have therefore been proposed to enhance the productivity of these annotations and reduce human efforts. We propose a high-throughput phenotyping system which features a Graphical User Interface (GUI) and a novel interactive segmentation algorithm: Semantic-Guided Interactive Object Segmentation (SGIOS). By providing a user-friendly interface and intelligent assistance with annotation, this system offers potential to streamline and accelerate the generation of training sets, reducing the effort required by the user. Our evaluation shows that our proposed SGIOS model requires fewer user inputs compared to the state-of-art models for interactive segmentation. As a case study of the use of the GUI applied for genetic discovery in plants, we present an example of results from a preliminary genome-wide association study (GWAS) of in planta regeneration in Populus trichocarpa (poplar). We further demonstrate that the inclusion of a semantic prior map with SGIOS can accelerate the training process for future GWAS, using a sample of a dataset extracted from a poplar GWAS of in vitro regeneration. The capabilities of our phenotyping system surpass those of unassisted humans to rapidly and precisely phenotype our traits of interest. The scalability of this system enables large-scale phenomic screens that would otherwise be time-prohibitive, thereby providing increased power for GWAS, mutant screens, and other studies relying on large sample sizes to characterize the genetic basis of trait variation. Our user-friendly system can be used by researchers lacking a computational background, thus helping to democratize the use of deep segmentation as a tool for plant phenotyping.

54 ENVIRONMENTAL SCIENCES↗

Mid-IR UAV-based sensing platform with deep learning to Identify and Quantify Gaseous Emission in Gas Flares

This report details the development and evaluation of a Mid-Infrared (Mid-IR) Unmanned Aerial Vehicle (UAV)-based sensing platform integrated with deep learning algorithms for the identification and quantification of gaseous emissions in gas flares. The project, spearheaded by Omega Optics, Inc., aimed to address environmental monitoring challenges by leveraging advanced photonic technologies and autonomous UAV operations. The research focused on designing, optimizing, and fabricating photonic crystal waveguides and grating couplers to enhance the sensitivity and accuracy of gas detection. A comprehensive drone-based system was developed, featuring a miniaturized sensor, GPS module, and microcontroller communication network for real-time gas concentration monitoring. The system's adaptive sampling algorithm, implemented using the Robot Operating System (ROS), enables autonomous detection and localization of gas emission sources. Preliminary results demonstrate the platform's capability to detect and monitor gas emissions with high precision, cost-effectiveness, and scalability. Future work will expand upon this foundation by introducing 3D wind model-based learning for dynamic environmental conditions and further enhancing the user interface and data processing algorithms to support broader environmental monitoring applications. Overall, this project represents a significant step forward in UAV-based environmental sensing technologies, offering robust solutions for detecting and mitigating the impacts of gaseous emissions on public health and safety.

47 OTHER INSTRUMENTATION↗

Scaling Ensembles of Data-Intensive Quantum Chemical Calculations for Millions of Molecules

Deep learning models are efficient computational tools that can accelerate the inverse design of molecules with desired functional properties by generating predictions at a fraction of the time required by traditional quantum chemical approaches. To ensure that a model maintains accuracy and transferability across broad regions of the chemical space explored during the inverse design, it must be trained on massively large volumes of simulation data. This requires running large-scale ensemble quantum chemical calculations on high-performance computing (HPC) systems for data collection. However, the efficient execution of such large ensemble calculations and the management of large volumes of output data require tools that can judiciously utilize computational resources and manage metadata overhead on the file system. Therefore, we present a high-performance, scalable, ensemble management framework for performing data-intensive quantum chemical electronic structure calculations for organic molecules. This framework provides abstractions to plug different ab initio, first principles, and first principles-based semi-empirical methods and executes them efficiently at large scale on HPC systems. It dynamically distributes tasks to resources and uses tiered storage for managing large collections of files. We employed this framework to process over ten million organic molecules and generate open-source datasets that provide UV-vis absorption spectra by running time-dependent density-functional tight-binding calculations. It is the largest database containing molecular optical spectra that were simulated with quantum chemical methods in a consistent manner.

Mehta, Kshitij↗

Quantum Reinforcement Learning for Volt-VAR Control in Power Distribution Systems

Volt-VAR control (VVC) is crucial in active distribution networks for optimizing voltage profiles and minimizing network losses. While traditional deep reinforcement learning (DRL) algorithms exhibit promise for VVC, they often require extensive computational resources to handle such a high-dimensional problem. As a potential solution, quantum reinforcement learning (QRL) algorithms integrate the computational capabilities of quantum computing into the DRL framework. However, existing QRL algorithms struggle with complex VVC problems due to the limitations of current quantum hardware. To bridge this gap, this paper proposes an innovative QRL algorithm featuring an end-to-end architecture that integrates a classical autoencoder, variational quantum circuits (VQCs), and classical post-processing layers. This design efficiently compresses high-dimensional grid states, enabling VQCs to leverage quantum advantages while producing multiple control device outputs tailored for VVC tasks. Numerical studies on three representative distribution systems verify the effectiveness and scalability of the proposed QRL algorithm, and demonstrate its enhanced performance over classical approaches with only approximately 1% of the parameters. Additionally, the robustness of our developed algorithm is validated through noisy quantum environments.

97 MATHEMATICS AND COMPUTING↗

Exploration with Scalable Gaussian Process Reinforcement Learning

Exploration is a challenging problem in reinforcement learning (RL), especially in environments with sparse rewards. Quantifying and utilizing the parametric uncertainty has been shown to be paramount for successful exploration [Osband et al., 2018]. Bayesian, or approximately Bayesian, methods present a principled means of estimating the parametric uncertainty in RL problems. Gaussian processes, nonparametric Bayesian models, are often impractical due to poor scalability and computational bottlenecks. We introduce a scalable Gaussian process RL (GPRL) method which directly induces sparsity in the covariance matrix to facilitate faster computation. This is a departure from previous GPRL methods which instead rely on data reduction and subsampling. We compare various covariance-based exploration techniques (Thompson sampling, upper confidence bound, and probabilistic maximum variance) which leverage our scalable GP framework in sparse reward environments. Finally, we show favorable comparison against the bootstrapped deep Q-Network.

97 MATHEMATICS AND COMPUTING↗

Can artificial intelligence and data-driven machine learning models match or even replace process-driven hydrologic models for streamflow simulation?: A case study of four watersheds with different hydro-climatic regions across the CONUS

With recent developments in computational techniques, Data-driven Machine Learning Models (DMLs) have shown great potential in simulating streamflow and capturing the rainfall-runoff relationship in given watersheds, which are traditionally fulfilled by Process-based Hydrologic Models (PHMs). There are debates on whether the DMLs can outperform and possibly replace the classical PHMs for streamflow simulation and river forecasting, but no clear conclusions have been made. This study aims to investigate whether the newer DMLs have any potential in further improving the simulation accuracy of classical PHMs, and vice versa. To do this, we compared a few popular PHMs and DMLs over four watersheds across the Continental US (CONUS) that are associated with different input, climate, and regional conditions. A total of five hydrologic models were chosen, including (1) two classical lumped models, i.e., the Sacramento Soil Moisture Accounting (SAC-SMA) and Xinanjiang (XAJ); (2) one modern distributed model, termed Coupled Routing and Excess Storage (CREST); (3) and two DMLs including an Artificial Neural Networks (ANN) and a deep learning model, termed Long Short Term Memory (LSTM). Our results demonstrated that the DMLs still significantly biased when using the baseline input scenario with the PHMs. However, the DMLs fed with delayed input scenarios had great potential and can reach high simulation accuracy. The DMLs, especially the ANN, outperformed other employed models under the rainfall-runoff relationship in which rainfall dominantly drives. Furthermore, the DMLs also showed better performance in the high-flow regime, while the PHMs had a better performance for the low-flow regime, implying both PHMs and DMLs have their own merits and are worthy of joint development. In general, our study indicated a great potential of using DMLs to simulate streamflow, but further studies are still needed to verify the transferability and scalability of DMLs in large-scale experiments, such as the Distributed Model Intercomparison Projects 1&2 conducted by National Weather Services but to compare modern DMLs and PHMs.

58 GEOSCIENCES↗

A fault-tolerant neutral-atom architecture for universal quantum computation

Quantum error correction (QEC) is essential for the realization of large-scale quantum computers. However, owing to the complexity of operating on the encoded ‘logical’ qubits, understanding the physical principles for building fault-tolerant quantum devices and combining them into efficient architectures is an outstanding scientific challenge. Here we use reconfigurable arrays of up to 448 neutral atoms to implement the key elements of a universal, fault-tolerant quantum processing architecture and experimentally explore their underlying working mechanisms. We first use surface codes to study how repeated QEC suppresses errors, demonstrating 2.14(13)x below-threshold performance in a four-round characterization circuit by leveraging atom loss detection and machine learning decoding. We then investigate logical entanglement using transversal gates and lattice surgery and extend it to universal logic through transversal teleportation with three-dimensional [[15,1,3]] codes, enabling arbitrary-angle synthesis with polylogarithmic overhead. Finally, we develop mid-circuit qubit reuse16, increasing experimental cycle rates by two orders of magnitude and enabling deep-circuit protocols with dozens of logical qubits and hundreds of logical teleportations with [[7,1,3]] and high-rate [[16,6,4]] codes while maintaining constant internal entropy. Our experiments show key principles for efficient architecture design, involving the interplay between quantum logic and entropy removal, judiciously using physical entanglement in logic gates and magic state generation, and leveraging teleportations for universality and physical qubit reset. These results establish foundations for scalable, universal error-corrected processing and its practical implementation in neutral atom systems.

atomic and molecular physics↗

Learning the structure of wind: A data-driven nonlocal turbulence model for the atmospheric boundary layer

In this work, we develop a novel data-driven approach to modeling the atmospheric boundary layer. This approach leads to a nonlocal, anisotropic synthetic turbulence model which we refer to as the deep rapid distortion (DRD) model. Our approach relies on an operator regression problem that characterizes the best fitting candidate in a general family of nonlocal covariance kernels parameterized in part by a neural network. This family of covariance kernels is expressed in Fourier space and is obtained from approximate solutions to the Navier–Stokes equations at very high Reynolds numbers. Each member of the family incorporates important physical properties such as mass conservation and a realistic energy cascade. The DRD model can be calibrated with noisy data from field experiments. After calibration, the model can be used to generate synthetic turbulent velocity fields. To this end, we provide a new numerical method based on domain decomposition which delivers scalable, memory-efficient turbulence generation with the DRD model as well as others. We demonstrate the robustness of our approach with both filtered and noisy data coming from the 1968 Air Force Cambridge Research Laboratory Kansas experiments. Using these data, we witness exceptional accuracy with the DRD model, especially when compared to the International Electrotechnical Commission standard.

17 WIND ENERGY↗

Sparse Convolutional Neural Networks for particle classification in ProtoDUNE-SP events

Deep Learning (DL) methods and Computer Vision are becoming important tools for event reconstruction in particle physics detectors. In this work, we report on the use of submanifold sparse convolutional neural networks (SparseNets) for the classification of track and shower hits from a DUNE prototype liquid-argon detector at CERN (ProtoDUNE-SP). By taking advantage of the three-dimensional nature of the problem we use a set of nine input features to classify sparse and locally dense hits associated to track or shower particles. The SparseNet has been trained on a test sample and shows promising results: efficiencies and purities greater than 90%. This has also been achieved with a considerable speedup and substantially less resource utilization with respect to other DL networks such as graph neural networks. This method offers great scalability advantages for future large neutrino detectors such as the planned DUNE experiment.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Automated Segmentation of Twin Boundaries in TRISO Silicon Carbide Using Deep Neural Networks

Coated particle fuels, such as the tristructural isotropic (TRISO) fuel particle, are essential for high-temperature gas reactor (HTGR) applications due to their efficiency and stability under normal and off-normal conditions. However, widespread commercialization and deployment of this technology for next-generation nuclear applications require robust quality assurance and quality control (QA/QC) methods linking fabrication, properties, and performance. Of the many important metrics for TRISO QA/QC, quantification of the silicon carbide (SiC) microstructure is critical because it correlates with fission product retention during irradiation. Previous work has shown extensive twinning of the SiC microstructure, which strongly affects microstructural metrics; however, twin grain boundaries are not expected play a significant role in fission product diffusion. This report summarizes the initial development, training, and testing of a machine learning image processing algorithm to detect twin grain boundaries in a backscattered electron image, which can be removed so that microstructural metrics can be recalculated for legacy data. Further development and deployment of this model will provide automated, scalable improvement of potential QA/QC methods for the SiC layer of TRISO particles.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Learning in continuous action space for developing high dimensional potential energy models

Reinforcement learning (RL) approaches that combine a tree search with deep learning have found remarkable success in searching exorbitantly large, albeit discrete action spaces, as in chess, Shogi and Go. Many real-world materials discovery and design applications, however, involve multi-dimensional search problems and learning domains that have continuous action spaces. Exploring high-dimensional potential energy models of materials is an example. Traditionally, these searches are time consuming (often several years for a single bulk system) and driven by human intuition and/or expertise and more recently by global/local optimization searches that have issues with convergence and/or do not scale well with the search dimensionality. Here, in a departure from discrete action and other gradient-based approaches, we introduce a RL strategy based on decision trees that incorporates modified rewards for improved exploration, efficient sampling during playouts and a “window scaling scheme" for enhanced exploitation, to enable efficient and scalable search for continuous action space problems. Using high-dimensional artificial landscapes and control RL problems, we successfully benchmark our approach against popular global optimization schemes and state of the art policy gradient methods, respectively. We demonstrate its efficacy to parameterize potential models (physics based and high-dimensional neural networks) for 54 different elemental systems across the periodic table as well as alloys. We analyze error trends across different elements in the latent space and trace their origin to elemental structural diversity and the smoothness of the element energy surface. Broadly, our RL strategy will be applicable to many other physical science problems involving search over continuous action spaces.

36 MATERIALS SCIENCE↗

GenomeFace v1.0

GenomeFace is meta-genome binning software. Metagenomic binning, the process of grouping DNA sequences into taxonomic units, is critical for understanding the functions, interactions, and evolutionary dynamics of microbial communities. We propose a deep learning approach to binning using two neural networks, one based on composition and another on environmental abundance, dynamically weighting the contribution of each based on characteristics of the input data. Trained on over 43,000 prokaryotic genomes, our network for composition-based binning is inspired by metric learning techniques used for facial recognition. Using a task-specific, multi-GPU accelerated algorithm to cluster the embeddings produced by our network, our binner leverages marker genes observed to be universally present in nearly all taxa to grade and select optimal clusters of sequences from a hierarchy of candidates. We evaluate our approach on four simulated datasets with known ground truth. Our linear time integration of marker genes recovers more near complete genomes than state of the art but computationally infeasible solutions using them, while being over an order of magnitude faster. Finally, we demonstrate the scalability and acuity of our approach by testing it on three of the largest metagenome assemblies ever performed. Compared to other binners, we produced 47%-183% more near complete genomes. From these datasets, we find over the genomes of over 3000 new candidate species which have never been previously cataloged, representing a potential 4% expansion of the known bacterial tree of life.

Lettich, Richard [Lawrence Berkeley National Labor↗