Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep transfer learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Deep learning interfacial momentum closures in coarse-mesh CFD two-phase flow simulation using validation data

Multiphase flow phenomena have been widely observed in the industrial applications while it remains a challenging yet unsolved problems. Three-dimensional computational fluid dynamics (CFD) approaches resolve the flow fields on a finer special and temporal scales which can complement the dedicated experimental study. However, closures have to be introduced to reflect the underlying physics in multiphase flow. Among them, the interfacial forces, including drag, lift, turbulent dispersion and wall lubrication forces, play in important role on the bubble’s distribution and migration in liquid-vapor two-phase flow. Development of those closures traditionally rely on the experimental data and analytical derivation with simplified assumptions which usually cannot deliver a universal solution across wide range of flow conditions. In this paper, a data-driven approach, named as Feature Similarity Measurement (FSM), is developed and applied to improve the simulation capability of two-phase flow with coarse-mesh CFD approach. Interfacial momentum transfer in adiabatic bubbly flow serves as the focus of the present study. Both a mature and a simplified set of interfacial closures are taken as the low fidelity data. Experimental data and fine mesh CFD simulations results are adopted as high-fidelity data. Qualitative and quantitative analysis are performed in this paper which reveals that FSM can substantially improve the prediction of coarse mesh CFD model regardless of the choice of interfacial closures and it provides scalability and consistency across discontinuous flow regimes. Furthermore, it demonstrates that data-driven method can aid the multiphase flow modeling by exploring the connections between local physical features and simulation errors.

97 MATHEMATICS AND COMPUTING↗

Language Model For Earth Science: Exploring Potential Downstream Applications As Well As Current Challenges

The use of deep learning techniques to build transformer language models such as SciBERT and GPT3 have transformed the natural language technology (NLT) landscape. These new NLTs are being used in speech to text and vice versa, auto-mated text classification, sentiment analysis, topic modeling, text summarization, and cognitive assistants. While Earth science has no shortage of unstructured data such as journal and conference papers, little efforts have focused on harnessing NLTs for knowledge extraction and supporting the scientific process. This paper surveys the use of language models in different science. BERT-E, a new Earth science-specific language model, is presented. BERT-E is generated using a transfer learning solution. A language model that has already been trained for general Science (SciBERT) is fine-tuned using abstracts and full text extracted from various Earth science-related articles. A downstream keywords classification application is used for evaluation, and the use of BERT-E shows improved performance. The need to develop a robust set of benchmarks in evaluating the language model such as BERT-E is discussed. Finally, example applications are presented to inspire additional ideas for applications using domain-specific language models.

R Ramachandran↗

The Artificial Scientist: in-Transit Machine Learning of Plasma Simulations

Large-scale simulations or scientific experiments produce petabytes of data per run. This poses massive challenges for I/O and storage when scientific analysis workflows are run manually offline. Unsupervised deep learning-based techniques to extract patterns and non-linear relations from these large amounts of data provide a way to build scientific understanding from raw data, reducing the need for manual pre-selection of analysis steps, but require exascale compute and memory to process the full dataset available. In this paper, we demonstrate a heterogeneous streaming workflow in which plasma simulation data is streamed directly to a Machine Learning (ML) application training a model on the simulation data in-transit, completely circumventing the capacity-constrained filesystem bottleneck. This workflow employs openPMD to provide a high level interface to describe scientific data and also uses ADIOS2, to transfer volumes of data that exceed the capabilities of the filesystem. We employ experience replay to avoid catastrophic forgetting in learning from this non-steady state process in a continual manner and adapt it to improve model convergence while learning in-transit. As a proof-of-concept, we approach the ill-posed inverse problem of predicting particle dynamics from radiation in a particle-incell (PIConGPU) simulation of the Kelvin-Helmholtz instability (KHI). We detail hardware-software co-design challenges as we scale PIConGPU to full Frontier, the Top-1 system as of June 2024 Top500 list.

Kelling, Jeffrey [Helmholtz-Zentrum Dresden Rossen↗

Deep learning of interface structures from simulated 4D STEM data: cation intermixing vs. roughening ∗

Abstract Interface structures in complex oxides remain an active area of condensed matter physics research, largely enabled by recent advances in scanning transmission electron microscopy (STEM). Yet the nature of the STEM contrast in which the structure is projected along the given direction precludes separation of possible structural models. Here, we utilize deep convolutional neural networks (DCNN) trained on simulated 4D STEM datasets to predict structural descriptors of interfaces. We focus on the widely studied interface between LaAlO 3 and SrTiO 3 , using dynamical diffraction theory and leveraging high performance computing to simulate thousands of possible 4D STEM datasets to train the DCNN to learn properties of the underlying structures on which the simulations are based. We test the DCNN on simulated data and show that it is possible (with >95% accuracy) to identify a physically rough from a chemically diffuse interface and create a DCNN regression model to predict step positions. We quantify the applicability of the model to different thicknesses and the transferability of the approach. The method shown here is general and can be applied for any inverse imaging problem where forward models are present.

42 ENGINEERING↗

Determination of Ceres Physical Parameters Using Radiometric and Optical Data

The Dawn spacecraft was launched on September 27th, 2007. Its mission is to rendezvous with and observe the two largest bodies in the main asteroid belt, Vesta and Ceres. It has completed over a year’s worth of direct observations of Vesta from early 2011 through late 2012. In the spring of 2015, the Dawn spacecraft entered orbit around the asteroid Ceres for the start of what is expected to be more than a year of science operations. The science data collected from this encounter consist of infrared (IR) images and spectra, visible images through a number of color filters, gamma ray detections and measurements of the Ceres gravity field. These data will be collected during several science phases: an Approach phase (1500000-4860 km from Ceres), a Survey orbit (4860 km radius), a High Altitude Mapping Orbit (HAMO) (1940 km radius) and a Low Altitude Mapping Orbit (LAMO) (855 km radius). The Approach phase included three Rotational Characterization (RC) opportunities. Designing each science orbit and successfully transferring into that orbit requires a sufficiently accurate estimate of Ceres physical parameters (body fixed frame, GM and harmonics). This paper focuses on work performed to estimate Ceres physical parameters using Deep Space Network (DSN) radiometric tracking data and optical measurements derived from science camera imagery. This paper describes planning for the data acquisition, as well as processing techniques and methodology. The trajectories predicted by the gravity field estimations are also compared with the actual as-flown trajectories. Observations of the gravity at high altitudes are found to be sufficient to design precision orbits at lower altitudes. Follow-up analysis after successfully reaching LAMO is included, as is a discussion of lessons learned.

Kennedy, Brian M.↗

Realistic galaxy image simulation via score-based generative models

ABSTRACT We show that a denoising diffusion probabilistic model (DDPM), a class of score-based generative model, can be used to produce realistic mock images that mimic observations of galaxies. Our method is tested with Dark Energy Spectroscopic Instrument (DESI) grz imaging of galaxies from the Photometry and Rotation curve OBservations from Extragalactic Surveys (PROBES) sample and galaxies selected from the Sloan Digital Sky Survey. Subjectively, the generated galaxies are highly realistic when compared with samples from the real data set. We quantify the similarity by borrowing from the deep generative learning literature, using the ‘Fréchet inception distance’ to test for subjective and morphological similarity. We also introduce the ‘synthetic galaxy distance’ metric to compare the emergent physical properties (such as total magnitude, colour, and half-light radius) of a ground truth parent and synthesized child data set. We argue that the DDPM approach produces sharper and more realistic images than other generative methods such as adversarial networks (with the downside of more costly inference), and could be used to produce large samples of synthetic observations tailored to a specific imaging survey. We demonstrate two potential uses of the DDPM: (1) accurate inpainting of occluded data, such as satellite trails, and (2) domain transfer, where new input images can be processed to mimic the properties of the DDPM training set. Here we ‘DESI-fy’ cartoon images as a proof of concept for domain transfer. Finally, we suggest potential applications for score-based approaches that could motivate further research on this topic within the astronomical community.

79 ASTRONOMY AND ASTROPHYSICS↗

Understanding and Leveraging the I/O Patterns of Emerging Machine Learning Analytics

The scientific community is currently experiencing unprecedented amounts of data generated by cutting-edge science facilities. Soon facilities will be producing up to 1 PB/s which will force scientist to use more autonomous techniques to learn from the data. The adoption of machine learning methods, like deep learning techniques, in large-scale workflows comes with a shift in the workflow’s computational and I/O patterns. These changes often include iterative processes and model architecture searches, in which datasets are analyzed multiple times in different formats with different model configurations in order to find accurate, reliable and efficient learning models. This shift in behavior brings changes in I/O patterns at the application level as well at the system level. These changes also bring new challenges for the HPC I/O teams, since these patterns contain more complex I/O workloads. In this paper we discuss the I/O patterns experienced by emerging analytical codes that rely on machine learning algorithms and highlight the challenges in designing efficient I/O transfers for such workflows. We comment on how to leverage the data access patterns in order to fetch in a more efficient way the required input data in the format and order given by the needs of the application and how to optimize the data path between collaborative processes. We will motivate our work and show performance gains with a study case of medical applications.

Gainaru, Ana↗

Freely scalable and reconfigurable optical hardware for deep learning

Abstract As deep neural network (DNN) models grow ever-larger, they can achieve higher accuracy and solve more complex problems. This trend has been enabled by an increase in available compute power; however, efforts to continue to scale electronic processors are impeded by the costs of communication, thermal management, power delivery and clocking. To improve scalability, we propose a digital optical neural network (DONN) with intralayer optical interconnects and reconfigurable input values. The path-length-independence of optical energy consumption enables information locality between a transmitter and a large number of arbitrarily arranged receivers, which allows greater flexibility in architecture design to circumvent scaling limitations. In a proof-of-concept experiment, we demonstrate optical multicast in the classification of 500 MNIST images with a 3-layer, fully-connected network. We also analyze the energy consumption of the DONN and find that digital optical data transfer is beneficial over electronics when the spacing of computational units is on the order of $$>10\,\upmu $$ > 10 μ m.

42 ENGINEERING↗

Concurrent Relaxation through Accelerated Deep Learning

CRADL captures performance metrics of machine learning algorithms operating on mesh data from multiphysics codes This proxy application is a tool to explore scalability of inference on HPC platforms, and also gather performance metrics for inference on new machine learning specific hardware. CRADL is designed to give users as fine a control as possible over an inference simulation. Users may select the number of cycles, amount of data, and batch size to pass to the accelerator of choice. Additionally the user may select a number of performance optimization libraries and flags. CRADL comes packaged with a repository of anonymized multi-physics simulation data, as well as a pretrained model for inference. The code allows a user to load their own pre-trained model and data if they wish. The code can operate in multiple parallelization schemes, with performance enhancing options such as half-precision libraries, PyTorch benchmarking, and pinned memory with non-blocking data transfers.

Zieb, KristoferJ.↗

1-D Convolutional Graph Convolutional Networks for Fault Detection in Distributed Energy Systems

This paper presents a 1-D convolutional and graph convolutional networks for fault detection in microgrids. The combination of 1-D convolutional neural networks (1D-CNN) and graph convolutional networks (GCN) helps extract both spatial-temporal correlations from the voltage measurements in microgrids. The fault detection scheme includes fault event detection, fault type and phase classification, and fault location. There are five neural network model training to handle these tasks. Transfer learning and fine-tuning are applied to reduce training efforts. The combined 1-D convolutional and graph convolutional networks (1D-CGCN) is compared with the traditional ANN structure on the Potsdam 13-bus microgrid dataset. The accuracy of 99.5%, 98.4%, 99.2%, and 95.5% are achieved in fault event detection, fault type classification, fault phase identification, and fault location respectively. The detailed confusion matrices of fault type and fault phase classification are provided for validation.

deep neural network↗

A domain wall-magnetic tunnel junction artificial synapse with notched geometry for accurate and efficient training of deep neural networks

Inspired by the parallelism and efficiency of the brain, several candidates for artificial synapse devices have been developed for neuromorphic computing, yet a nonlinear and asymmetric synaptic response curve precludes their use for backpropagation, the foundation of modern supervised learning. Spintronic devices—which benefit from high endurance, low power consumption, low latency, and CMOS compatibility—are a promising technology for memory, and domain-wall magnetic tunnel junction (DW-MTJ) devices have been shown to implement synaptic functions such as long-term potentiation and spike-timing dependent plasticity. In this work, we propose a notched DW-MTJ synapse as a candidate for supervised learning. Using micromagnetic simulations at room temperature, we show that notched synapses ensure the non-volatility of the synaptic weight and allow for highly linear, symmetric, and reproducible weight updates using either spin transfer torque (STT) or spin–orbit torque (SOT) mechanisms of DW propagation. We use lookup tables constructed from micromagnetics simulations to model the training of neural networks built with DW-MTJ synapses on both the MNIST and Fashion-MNIST image classification tasks. Accounting for thermal noise and realistic process variations, the DW-MTJ devices achieve classification accuracy close to ideal floating-point updates using both STT and SOT devices at room temperature and at 400 K. Our work establishes the basis for a magnetic artificial synapse that can eventually lead to hardware neural networks with fully spintronic matrix operations implementing machine learning.

42 ENGINEERING↗

Rays for Roots - Integrating Backscatter X-Ray Phenotyping, Modeling and Genetics to Increase Carbon Sequestration and Switchgrass Resource Use (Final Report)

To increase carbon (C) deposition in the soil and enhance crop resource use efficiency, characterizing root form and function is essential. Several root and soil traits have been linked to increased root-to-soil C transfer. Technology that could provide high-resolution characterization of many of these traits in field conditions would revolutionize our ability to study and understand how to increase C sequestration. In this effort, we developed an initial early prototype backscatter X-ray system for non-destructive imaging of root traits. We collected initial backscatter X-ray data in field and lab settings and carried out early analysis of these data. Along with this prototype, we also developed a suite of root phenotyping approaches including advanced minirhizotron image analysis, soil core imaging, and mesocosm imaging. Minirhizotron (MR) tubes are clear tubes inserted into the soil in the field and used to image roots and the surrounding soil. Our team has developed deep learning-based methods that can segment roots from soil that can learn from imprecise image-level labels. The ability to learn or fine-tune our deep learning algorithms from image-level labels allows easier and faster application of these approaches to new locations and new plant species. We have successfully implemented and applied our MR analysis approaches to thousands of switchgrass MR images collected across geographical regions. An advantage of MR imaging is the ability to collect root and soil images over time. Our soil core analysis included collecting hundreds of soil core samples from harvested switchgrass fields and imaging these cores with both X-ray CT and backscatter X-ray imaging. Initial segmentation approaches for the X-ray CT images of these cores have been developed and applied. An advantage of soil core analysis is that it preserves the three-dimensional structures of the roots and soil in the core collected. Our group also developed photogrammetry-based mesocosm root imaging and phenotyping approaches. In this approach, a plant was grown in a large mesocosm with a three-dimensional grid of thin supporting lines inserted throughout the mesocosm. After the plant (and, correspondingly, the root architecture is grown and established) the soil media was removed and the supporting lines approximately preserved the three-dimensional root architecture. Then, we applied photogrammetry techniques to create a three-dimensional digital representation of the root architecture for which we developed analysis algorithms including skeletonization. We carried out our phenotyping development with powerful switchgrass resources and physiological and agroecosystem modeling to deliver novel technology. This project contributes to multiple ARPA-E missions including reduction of foreign imports of energy, reduction of energy-related emissions including greenhouse gases, and ensuring that the United States maintains a technological lead in developing and deploying advanced energy technology. Furthermore, the developed tools could transform public and private plant breeding and could be broadly applicable to other crops and, potentially, other application areas. Our team of engineers, plant and soil scientists, and modelers i) developed an early prototype backscatter X-ray platform that can operate in field conditions; ii) developed a suite of root phenotyping and characterization approaches as described above; iii) developed and carried out plant biology and physiology roots studies and; iv) developed and implemented mechanistic physiological modeling.

42 ENGINEERING↗

Language model-accelerated deep symbolic optimization

Symbolic optimization methods have been used to solve varied challenging and relevant problems such as symbolic regression and neural architecture search. However, the current state of the art typically learns each problem from scratch and is unable to leverage pre-existing knowledge and datasets that are available for many applications. Here, inspired by the similarity between sequence representations learned in natural language processing and the formulation of symbolic optimization as a discrete sequence optimization problem, we propose language model-accelerated deep symbolic optimization (LA-DSO), a method that leverages language models to learn symbolic optimization solutions more efficiently. We demonstrate LA-DSO in two tasks: symbolic regression, which allows us to perform extensive experimentation due to its low computation requirements, and computational antibody optimization, which shows that our proposal accelerates learning in challenging real-world problems.

97 MATHEMATICS AND COMPUTING↗

Flow over an espresso cup: inferring 3-D velocity and pressure fields from tomographic background oriented Schlieren via physics-informed neural networks

Tomographic background oriented Schlieren (Tomo-BOS) imaging measures density or temperature fields in three dimensions using multiple camera BOS projections, and is particularly useful for instantaneous flow visualizations of complex fluid dynamics problems. We propose a new method based on physics-informed neural networks (PINNs) to infer the full continuous three-dimensional (3-D) velocity and pressure fields from snapshots of 3-D temperature fields obtained by Tomo-BOS imaging. The PINNs seamlessly integrate the underlying physics of the observed fluid flow and the visualization data, hence enabling the inference of latent quantities using limited experimental data. In this hidden fluid mechanics paradigm, we train the neural network by minimizing a loss function composed of a data mismatch term and residual terms associated with the coupled Navier–Stokes and heat transfer equations. We first quantify the accuracy of the proposed method based on a two-dimensional synthetic data set for buoyancy-driven flow, and subsequently apply it to the Tomo-BOS data set, where we are able to infer the instantaneous velocity and pressure fields of the flow over an espresso cup based only on the temperature field provided by the Tomo-BOS imaging. Moreover, we conduct an independent PIV experiment to validate the PINN inference for the unsteady velocity field at a centre plane. To explain the observed flow physics, we also perform systematic PINN simulations at different Reynolds and Richardson numbers and quantify the variations in velocity and pressure fields. Furthermore, the results in this paper indicate that the proposed deep learning technique can become a promising direction in experimental fluid mechanics.

97 MATHEMATICS AND COMPUTING↗

MnEdgeNet for accurate decomposition of mixed oxidation states for Mn XAS and EELS L2,3 edges without reference and calibration

Accurate decomposition of the mixed Mn oxidation states is highly important for characterizing the electronic structures, charge transfer and redox centers for electronic, and electrocatalytic and energy storage materials that contain Mn. Electron energy loss spectroscopy (EELS) and soft X-ray absorption spectroscopy (XAS) measurements of the Mn L2,3 edges are widely used for this purpose. To date, although the measurements of the Mn L2,3 edges are straightforward given the sample is prepared properly, an accurate decomposition of the mix valence states of Mn remains non-trivial. For both EELS and XAS, 2+, 3+, and 4+ reference spectra need to be taken on the same instrument/beamline and preferably in the same experimental session because the instrumental resolution and the energy axis offset could vary from one session to another. To circumvent this hurdle, in this study, we adopted a deep learning approach and developed a calibration-free and reference-free method to decompose the oxidation state of Mn L2,3 edges for both EELS and XAS. A deep learning regression model is trained to accurately predict the composition of the mix valence state of Mn. To synthesize physics-informed and ground-truth labeled training datasets, we created a forward model that takes into account plural scattering, instrumentation broadening, noise, and energy axis offset. With that, we created a 1.2 million-spectrum database with 1-by-3 oxidation state composition ground truth vectors. The library includes a sufficient variety of data including both EELS and XAS spectra. By training on this large database, our convolutional neural network achieves 85% accuracy on the validation dataset. We tested the model and found it is robust against noise (down to PSNR of 10) and plural scattering (up to t/λ = 1). We further validated the model against spectral data that were not used in training. In particular, the model shows high accuracy and high sensitivity for the decomposition of Mn 3 O 4 , MnO, Mn 2 O 3 , and MnO 2 . The accurate decomposition of Mn 3 O 4 experimental data shows the model is quantitatively correct and can be deployed for real experimental data. Our model will not only be a valuable tool to researchers and material scientists but also can assist experienced electron microscopists and synchrotron scientists in the automated analysis of Mn L edge data.

25 ENERGY STORAGE↗

Deep learning-enabled probing of irradiation-induced defects in time-series micrographs

Modeling time-series data with convolutional neural networks (CNNs) requires building a model to learn in batches as opposed to training sequentially. Coupling CNNs with in situ or operando techniques opens the possibility of accurately segmenting dynamic reactions and mass transport phenomena to understand how materials behave under the conditions in which they are used. In this article, in situ ion irradiation transmission electron microscopy (TEM) images are used as inputs into the CNN to assess the defect generation rate, defect cluster density, and saturation of defects. We then use the output segmentation maps to correlate with conventional TEM micrographs to assess the model’s ability to detail nanoscale interactions. Next, we discuss the implications of preprocessing and hyperparameters on model variability, accuracy when expanded to other datasets, and the role of regularization when controlling model variance. Ultimately, we eliminate human bias when extrapolating physical metrics, speed up analysis time, decouple reactions that happen at 100 ms intervals, and deploy models that are both accurate and transferable to similar experiments.

Burns, Kory (ORCID:0000000158019909)↗

Can artificial intelligence and data-driven machine learning models match or even replace process-driven hydrologic models for streamflow simulation?: A case study of four watersheds with different hydro-climatic regions across the CONUS

With recent developments in computational techniques, Data-driven Machine Learning Models (DMLs) have shown great potential in simulating streamflow and capturing the rainfall-runoff relationship in given watersheds, which are traditionally fulfilled by Process-based Hydrologic Models (PHMs). There are debates on whether the DMLs can outperform and possibly replace the classical PHMs for streamflow simulation and river forecasting, but no clear conclusions have been made. This study aims to investigate whether the newer DMLs have any potential in further improving the simulation accuracy of classical PHMs, and vice versa. To do this, we compared a few popular PHMs and DMLs over four watersheds across the Continental US (CONUS) that are associated with different input, climate, and regional conditions. A total of five hydrologic models were chosen, including (1) two classical lumped models, i.e., the Sacramento Soil Moisture Accounting (SAC-SMA) and Xinanjiang (XAJ); (2) one modern distributed model, termed Coupled Routing and Excess Storage (CREST); (3) and two DMLs including an Artificial Neural Networks (ANN) and a deep learning model, termed Long Short Term Memory (LSTM). Our results demonstrated that the DMLs still significantly biased when using the baseline input scenario with the PHMs. However, the DMLs fed with delayed input scenarios had great potential and can reach high simulation accuracy. The DMLs, especially the ANN, outperformed other employed models under the rainfall-runoff relationship in which rainfall dominantly drives. Furthermore, the DMLs also showed better performance in the high-flow regime, while the PHMs had a better performance for the low-flow regime, implying both PHMs and DMLs have their own merits and are worthy of joint development. In general, our study indicated a great potential of using DMLs to simulate streamflow, but further studies are still needed to verify the transferability and scalability of DMLs in large-scale experiments, such as the Distributed Model Intercomparison Projects 1&2 conducted by National Weather Services but to compare modern DMLs and PHMs.

58 GEOSCIENCES↗

Testing Paired Neural Network Models for Aftershock Identification

Aftershock sequences are a burden to real-time seismic monitoring. Cross-correlation can be used because aftershocks exhibit similar waveforms, but the method is computationally expensive. Deep learning may be an alternative, as it is computationally efficient, but great attention to training and testing is required in order to trust that the model can generalize to new aftershock sequences. This is problematic for aftershock sequences, because large-magnitude earthquakes are unpredictable and are globally widespread. Here, we test several paired neural network (PNN) models trained on a augmented (noise-added) earthquake dataset, to determine whether they can be generalized to process real aftershock sequences. Two aftershock datasets that were originally detected by cross-correlation and subsequently validated by an expert analyst were used. We found that current PNN models struggle to generalize to aftershock sequences. However, we identify approaches to improve training future PNN models and believe that improvements may be achieved by transfer learning.

58 GEOSCIENCES↗