Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Microstructure Segmentation With Deep Learning Encoders Pre-Trained on a Large Microscopy Dataset

This study examined the improvement of microscopy segmentation intersection over union accuracy by transfer learning from a large dataset of microscopy images called MicroNet. Many neural network encoder architectures were trained on over 100,000 labeled microscopy images from 54 material classes. These pre-trained encoders were then embedded into multiple segmentation architectures including UNet and DeepLabV3+ to evaluate segmentation performance on created benchmark microscopy datasets. Compared to ImageNet pre-training, models pre-trained on MicroNet generalized better to out-of-distribution micrographs taken under different imaging and sample conditions and were more accurate with less training data. When training with only a single Ni-superalloy image, pre-training on MicroNet produced a 72.2% reduction in relative intersection over union error. These results suggest that transfer learning from large in-domain datasets generate models with learned feature representations that are more useful for downstream tasks and will likely improve any microscopy image analysis technique that can leverage pre-trained encoders.

machine learning↗

Using Machine Learning to Estimate Surface-Level SO2 Concentrations from Satellite-Based Measurements

Sulfur dioxide (SO2) is a criteria air pollutant due to its contributions to aerosol formation, rainfall acidification, and harm to human health. The placement of air quality monitoring sites is typically biased towards urban areas, leaving large areas with very limited monitoring data. The Ozone Monitoring Instrument (OMI) has been used to provide estimates of SO2 vertical column densities (VCDs) globally at spatial resolution of 10s of kms once per day. OMI SO2 VCDs have been previously used to estimate surface SO2 concentrations using chemical transport model (CTM) simulations. The CTMs use estimated emissions and assimilated meteorological data, and simulate the chemical and physical processes that determine the vertical profile of SO2, which can be used to derive a ratio between the surface concentrations and VCDs. These models are complex, computationally expensive, and have large uncertainties in the simulated surface-to-VCD ratio due to biases in emissions and relatively coarse resolution. Machine learning techniques are comparatively easier to use, much less computationally expensive to use after training, and can produce more accurate estimations of surface concentrations than the CTM-based method. The interpretation of machine learning models often poses challenges, and in some cases, non-physical variables unrelated to SO2 are used as predictors. In this work, we create an artificial neural network (ANN) to relate OMI retrievals and archived GEOS-FP boundary layer heights to surface SO2 concentrations from the ChinaHighAirPollutants ChinaHighSO2 dataset (CHAP; Wei et al., 2023) on a seasonal average timescale from 2013-2018. Our model only utilizes five variables that are directly relevant to the satellite retrieval, lifetime, and spatial distribution of SO2. The model was trained on 16 seasons (four of each) with independent validation (one of each season) and testing datasets (one of each season) to avoid overfitting. Our ANN generates surface SO2 concentrations that are sensitive (slope = 0.51) and consistent (r = 0.74) with the CHAP data, but are underpredicted by an average of 1.2 ppbv with a mean absolute error of 2.2 ppbv. These results are better than recent studies utilizing the CTM method. To our knowledge, this is the best performing machine learning model that only uses physical variables to predict surface SO2. Our work demonstrates that a carefully constructed, simple ML model can accurately estimate surface-based SO2 concentrations from satellite VCD measurements, and this technique has future promise to expend to newer, higher resolution satellites and other air pollutants.

SO2, air quality, OMI, machine learning↗

Elastomeric load sharing device

An elastomeric load sharing device, interposed in combination between a driven gear and a central drive shaft to facilitate balanced torque distribution in split power transmission systems, includes a cylindrical elastomeric bearing and a plurality of elastomeric bearing pads. The elastomeric bearing and bearing pads comprise one or more layers, each layer including an elastomer having a metal backing strip secured thereto. The elastomeric bearing is configured to have a high radial stiffness and a low torsional stiffness and is operative to radially center the driven gear and to minimize torque transfer through the elastomeric bearing. The bearing pads are configured to have a low radial and torsional stiffness and a high axial stiffness and are operative to compressively transmit torque from the driven gear to the drive shaft. The elastomeric load sharing device has spring rates that compensate for mechanical deviations in the gear train assembly to provide balanced torque distribution between complementary load paths of split power transmission systems.

Isabelle, Charles J.↗

An Efficient GPU-Accelerated Multi-Source Global Fit Pipeline for LISA Data Analysis

The large-scale analysis task of deciphering gravitational wave signals in the LISA data stream will be difficult, requiring a large amount of computational resources and extensive development of computational methods. Its high dimensionality, multiple model types, and complicated noise profile require a global fit to all parameters and input models simultaneously. In this work, we detail our global fit algorithm, called “Erebor,” designed to accomplish this challenging task. It is capable of analysing current state-of-the-art datasets and then growing into the future as more pieces of the pipeline are completed and added. We describe our pipeline strategy, the algorithmic setup, and the results from our analysis of the LDC2A Sangria dataset, which contains Massive Black Hole Binaries, compact Galactic Binaries, and a parameterized noise spectrum whose parameters are unknown to the user. The Erebor algorithm includes three unique and very useful contributions: GPU acceleration for enhanced computational efficiency; ensemble MCMC sampling with multiple MCMC walkers per temperature for better mixing and parallelized sample creation; and special online updates to reversible-jump (or trans-dimensional) sampling distributions to ensure sampler mixing and accurate initial estimates for detectable sources in the data. We recover posterior distributions for all 15 (6) of the injected MBHBs in the LDC2A training (hidden) dataset. We catalog ∼12000 Galactic Binaries (∼8000 as high confidence detections) for both the training and hidden datasets. All of the sources and their posterior distributions are provided in publicly available catalogs.

LISA global fit↗

A new way to unravel the $^{12}$C($\alpha $,$\gamma $)$^{16}$O cross section components using neural networks

The $^{12}$C($\alpha $,$\gamma $)$^{16}$O reaction rate is crucial in determining the carbon-to-oxygen abundance ratio in stellar nucleosynthesis. Measuring this reaction’s cross section at stellar energies is challenging due to its extremely small value, approximately 10 -17 barn at E c.m. = 300 keV. To address this, R-matrix calculations are employed to extrapolate data to lower energies, requiring a comprehensive understanding of each contribution to the cross section. The dominant contributions to the cross section at stellar energies arise from electric dipole (E1) and electric quadrupole (E2) transitions to the ground state of 16 O, along with a significant cascade contribution. Traditionally, these contributions have been separated using the γ-ray angular distribution. In this work, we propose a novel technique using the energy distribution of the 16 O recoils at the focal plane. This method involves a neural network trained on detailed Monte Carlo simulations of the energy distribution of recoils transported through the recoil mass separator ERNA. This approach enables the simultaneous determination of all three contributions with errors around 10% in the energy range E c.m. = 1.0–2.2 MeV. In conclusion, by employing this new technique, we aim to significantly improve the accuracy of determining the cross section of the $^{12}$C($\alpha $,$\gamma $)$^{16}$O reaction at astrophysical energies.

Astrochemistry↗

High Performance EVA Glove Collaboration: Glove Injury Data Mining Effort

Human hands play a significant role during Extravehicular Activity (EVA) missions and Neutral Buoyancy Lab (NBL) training events, as they are needed for translating and performing tasks in the weightless environment. Because of this high frequency usage, hand and arm related injuries are known to occur during EVA and EVA training in the NBL. The primary objectives of this investigation were to: 1) document all known EVA glove related injuries and circumstances of these incidents, 2) determine likely risk factors, and 3) recommend interventions where possible that could be implemented in the current and future glove designs. METHODS: The investigation focused on the discomforts and injuries of U.S. crewmembers who had worn the pressurized Extravehicular Mobility Unit (EMU) spacesuit and experienced 4000 Series or Phase VI glove related incidents during 1981 to 2010 for either EVA ground training or in-orbit flight. We conducted an observational retrospective case-control investigation using 1) a literature review of known injuries, 2) data mining of crew injury, glove sizing, and hand anthropometry databases, 3) descriptive statistical analyses, and finally 4) statistical risk correlation and predictor analyses to better understand injury prevalence and potential causation. Specific predictor statistical analyses included use of principal component analyses (PCA), multiple logistic regression, and survival analyses (Cox proportional hazards regression). Results of these analyses were computed risk variables in the forms of odds ratios (likelihood of an injury occurring given the magnitude of a risk variable) and hazard ratios (likelihood of time to injury occurrence). Due to the exploratory nature of this investigation, we selected predictor variables significant at p≤0.15. RESULTS: Through 2010, there have been a total of 330 NASA crewmembers, from which 96 crewmembers performed 322 EVAs during 1981-2010, resulting in 50 crewmembers being injured inflight and 44 injured during 11,704 ground EVA training events. Of the 196 glove related injury incidents, 106 related to EVA and 90 to EVA training. Over these 196 incidents, 277 total injuries (126 flight; 151 training) were reported and were then grouped into 23 types of injuries. Of EVA flight injuries, 65% were commonly reported to the hand (in general), metacarpophalangeal (MCP) joint, and finger (not including thumb) with fatigue, abrasion, and paresthesia being the most common injury types (44% of total flight injuries). Training injuries totaled to more than 70% being distributed to the fingernail, MCP joint, and finger crotch with 88% of the specific injuries listed as pain, erythema, and onycholysis. Of these training injuries, when reporting pain or erythema, the most common location was the index finger, but when reporting onycholysis, it was the middle finger. Predictor variables specific to increased risk of onycholysis included: female sex (OR=2.622), older age (OR=1.065), increased duration in hours of the flight or training event (OR=1.570), middle finger length differences in inches between the finger and the EVA glove (OR=7.709), and use of the Phase VI glove (OR=8.535). Differentiation between training and flight and injury reporting during 2002-2004 were significant control variables. For likelihood of time to first onycholysis injury, there was a 24% reduction in rate of reporting for each year increase in age. Also, more experienced crewmembers, based on number of EVA flight or training events completed, were less likely to report an onycholysis injury (3% less for every event). Longer duration events also found reporting rates to occur 2.37 times faster for every hour of length. Crewmembers with larger hand size reported onycholysis 23% faster than those with smaller hand size. Finally, for every 1/10th of an inch increase in difference between the middle finger length and the glove, the rate of reporting increased by 60%. DISCUSSION: One key finding was that the Series 4000 glove had a lower injury risk than the Phase VI, which provides a platform for further evaluation. General interventions that reduce hand overexertion and repetitive use exposure through tool development, procedural changes and shorter exposures may be one mitigation path, but due to the way the training event times were reported, we cannot provide a guideline for a specific event duration change. When the finger length was different from the glove length, the risk of injury increased indicating that the use of larger finger take-ups could be contributing to injury and therefore may not be recommended. Prior to this investigation, there was one previous investigation indicating hand anthropometry may be related to onycholysis. We found different hand anthropometry variables indicated by this investigation as compared to the prior, specifically differences in middle finger length compared to glove finger length, which point more towards a sizing issue than a specific anthropometry issue. Additionally, although this investigation has identified sizing as an issue, the force and environmental-related variables of the EVA glove that could also cause injury were not accounted for.

Reid, C. R.↗

Generative models on phase space

Deep generative models such as diffusion and flow matching are powerful machine learning tools capable of learning and sampling from high-dimensional distributions. They are particularly useful when the training data appears to be concentrated on a submanifold of the data embedding space. For high-energy physics data, consisting of collections of relativistic energy-momentum 4-vectors, this submanifold can enforce extremely strong physically-motivated priors, such as energy and momentum conservation. If these constraints are learned only approximately, rather than exactly, this can inhibit the interpretability and reliability of such generative models. To remedy this deficiency, we introduce generative models which are, by construction, confined at every step of their sampling trajectory to the manifold of massless N-particle Lorentz-invariant phase space in the center-of-momentum frame. In the case of diffusion models, the "pure noise" forward process endpoint corresponds to the uniform distribution on phase space, which provides a clear starting point from which to identify how correlations among the particles emerge during the reverse (de-noising) process. We demonstrate that our models are able to learn both few-particle and many-particle distributions with various singularity structures, paving the way for future interpretability studies using generative models trained on simulated jet data.

Bogorad, Zachary [Fermilab]↗

A multi-dimensional parametric study of variability in multi-phase flow dynamics during geologic CO 2 sequestration accelerated with machine learning

Successful geologic CO 2 storage projects depend on numerical simulations to predict reservoir performance during site selection, injection verification, and post-injection monitoring phases of the project. These numerical simulations solve non-linear sets of coupled partial differential equations, while accounting for multi-phase fluid dynamics on the basis of constitutive equations that are embedded into the solution scheme. As a consequence, individual simulations often require tens to hundreds of hours to complete on high-performance computing clusters. Moreover, laboratory experiments reveal that parametric functions for capillary pressure and relative permeability exhibit substantial variability, even within the same rock type. This combination of computational expense and wide-ranging parametric variability means that there remains substantial uncertainty in the behavior of multi-phase CO 2 -water systems, particularly in the context of feedbacks between relative permeability and capillary pressure. To bridge this knowledge gap, here we develop a novel workflow that utilizes physics-based numerical simulation to train an artificial neural network (ANN) emulator for interrogating the multivariate parameter space that governs both capillary pressure and relative permeability. With this approach, the ANN is trained to emulate both fluid pressure distribution and CO 2 saturation, which are then interrogated quantitatively to generate parametric response surface mappings with high-fidelity resolution. Results from this study initially show that capillary entry pressure is the dominant control on both CO 2 plume geometry and fluid pressure propagation when considering the combined effects of capillary pressure and relative permeability, particularly when phase interference is low and residual CO 2 saturation is high. Moreover, the ANN emulator provides tremendous computational speed-up by computing 2691 individual simulations in several minutes; whereas, the same simulation ensemble would have required ~3 years of simulation time using only physics-based simulation methods (25,000 times speed up).

58 GEOSCIENCES↗

Physics-based hybrid machine learning for critical heat flux prediction with uncertainty quantification

Critical heat flux (CHF) is a key quantity in nuclear system modeling due to its impact on heat transfer, safety margins, and reactor performance. This study develops and validates an uncertainty-aware hybrid modeling approach that combines machine learning with physics-based models to predict CHF in cases of dryout. The Biasi and Bowring empirical correlations were paired with three ML uncertainty quantification (UQ) techniques: deep neural network (DNN) ensembles, Bayesian neural networks (BNNs), and deep Gaussian processes (DGPs). A pure ML model without a base model was evaluated for comparison. Model performance was assessed under plentiful (7,350 points) and limited (9 points) training data scenarios using parity, uncertainty distributions, and calibration curves. Results show that the Biasi hybrid DNN ensemble achieved the best overall performance, with a mean absolute relative error of 1.846%, and well-calibrated uncertainty estimates. The BNN-based hybrids showed slightly higher error (2.14%) but superior uncertainty calibration. DGP models underperformed, with over 6% error and poor uncertainty calibration. All hybrid models outperformed pure machine learning configurations, demonstrating resistance against data scarcity. These findings indicate that hybrid modeling significantly improves predictive accuracy, interpretability, and resilience to data scarcity. The integration of uncertainty awareness provides actionable confidence in CHF predictions, which is vital for safety-critical decisions in nuclear applications. This hybrid approach offers a viable pathway for deploying ML models in reactor analysis tools while preserving domain knowledge and physical consistency.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Generative learning of densities on manifolds

A generative modeling framework is proposed that combines diffusion models and manifold learning to efficiently sample data densities on manifolds. The approach utilizes Diffusion Maps to uncover possible low-dimensional underlying (latent) spaces in the high-dimensional data (ambient) space. Two approaches for sampling from the latent data density are described. The first is a score-based diffusion model, which is trained to map a standard normal distribution to the latent data distribution using a neural network. The second one involves solving an Itô stochastic differential equation in the latent space. Additional realizations of the data are generated by lifting the samples back to the ambient space using Double Diffusion Maps , a recently introduced technique typically employed in studying dynamical system reduction; here the focus lies in sampling densities rather than system dynamics. The proposed approaches enable sampling high dimensional data densities restricted to low-dimensional, a priori unknown manifolds. The efficacy of the proposed framework is demonstrated through a benchmark problem and a material with multiscale structure.

Double diffusion maps↗

MatPhase: Material phase prediction for Li-ion Battery Reconstruction using Hierarchical Curriculum Learning

Li-ion Batteries (LIB), one of the most efficient energy storage devices, are used extensively in many industrial applications. These batteries consist of electrodes that are put together with heterogeneous material compositions. Imaging data of these battery electrodes obtained from X-ray tomography can explain the distribution of material constituents and allow reconstructions to study electron transport pathways. Such reconstructions of material constituents help quantify various associated properties of electrodes (e.g., volume-specific surface area, porosity) which determine the performance of batteries. These images often suffer from low image contrast between multiple material constituents, hence making it difficult for humans to distinguish and characterize these constituents through visual inspection. A minor error in detecting distributions of the material constituents can lead to magnified errors in the calculated parameters of material properties (e.g., porosity). We present MatPhase, a novel hierarchical curriculum learning technique to address the complex task of estimating material constituent distribution in battery electrodes. MatPhase comprises three modules: (i) an uncertainty-aware global model trained to yield inferences conditioned upon global knowledge of material distribution, (ii) a local model to capture relatively more fine-grained (local) distributional signals, (iii) an aggregator model to appropriately fuse the local and global effects towards obtaining the final distribution. On average, MatPhase improves prediction up to 8.5% relative to other sophisticated modeling pipelines and state-of-the-arts (SOTA) object detection models employed in the performance comparison.

Tabassum, Anika↗

3D Gaussian Splatting for Volume Compression

This codebase uses machine learning to train a collection of 3D Gaussian distributions to approximate scientific volume data. Because this collection uses less memory than the original dataset, it can be used as a compressed model of the original data for applications such as visualization.

Dyken, Landon↗

How to understand limitations of generative networks

Well-trained classifiers and their complete weight distributions provide us with a well-motivated and practicable method to test generative networks in particle physics. We illustrate their benefits for distribution-shifted jets, calorimeter showers, and reconstruction-level events. In all cases, the classifier weights make for a powerful test of the generative network, identify potential problems in the density estimation, relate them to the underlying physics, and tie in with a comprehensive precision and uncertainty treatment for generative networks.

Physics↗

Integrated Design of Ultradurable, Low CO 2 Alternative Binder Systems via Machine Learning

This ARPA-E project developed a machine learning tool to use in formulation design of cementitious binders for concrete having 50% less embodied CO 2 and possessing twice the durability compared to concrete based on ordinary portland cement (OPC) binders. The technical focus was on limestone/calcined clay cement (LC3), the leading replacement for OPC. Here, hierarchical machine learning (HML) was used to model the flowability, set time, strength, and durability of LC3 concrete. This methodology identifies latent variables derived from domain knowledge and empirical models that develop an accurate model for a response surface from small datasets. For the flowability metric, particle packing was a dominant factor, while strength and durability were both strongly determined by the fraction of metakaolin and the water:solids ratio. Under constraints of water:binder ratio, material performance metrics, embodied CO 2 , and cost per tonne of OPC, multi-objective optimization was used to design binders parameterized by the mineral composition replacing OPC, particle size distributions, and water:solids ratio. The trained algorithm was able to predict multiple mixes met these performance criteria, and experimental testing validated the predictions. The model demonstrated here is relevant for North America, where pure kaolin deposits are found broadly. The approach is being taken forward into commercial application by Ansatz AI, a materials informatics company founded by PI Washburn and co-PI Poczos. Through collaborations with the cement and concrete industry, and funding from SBIR programs, a commercial software will be developed in future research.

36 MATERIALS SCIENCE↗

RICIS Symposium 1988

Integrated Environments for Large, Complex Systems is the theme for the RICIS symposium of 1988. Distinguished professionals from industry, government, and academia have been invited to participate and present their views and experiences regarding research, education, and future directions related to this topic. Within RICIS, more than half of the research being conducted is in the area of Computer Systems and Software Engineering. The focus of this research is on the software development life-cycle for large, complex, distributed systems. Within the education and training component of RICIS, the primary emphasis has been to provide education and training for software professionals.

Source record↗

Environmental awareness program development

Work this summer in the Office of Safety, Environment, and Mission Assurance began with a review of current initiatives and environmental projects at the Langley Research Center (LaRC). This involved researching many of the documents on file which detail problems which have occurred as well as various approaches which have been used to address these problems. A large portion of the time was spent interviewing and working with each of the engineers, industrial hygienists and other professionals connected with the Office of Environmental Engineering. A few of the projects I worked on include: Researching environmental compliance, and pollution prevention efforts; touring many of the facilities at LaRC to observe the environmental efforts in the work place; researching equipment needs for the recycling/reclamation center; writing scripts for in-house training videos; working with the video production department to produce a training video; developing e-mail distribution list; developing environmental coordinator's database; and working with others to research logistics of recycling and waste minimization efforts.

Steinhauer, David A.↗

A Global Model for Estimating Atmospheric Phase Scintillation Statistics

Since 2007, the National Aeronautics and Space Administration (NASA) has been collecting atmospheric phase turbulence data from various NASA ground stations throughout the world. The goal of these measurement campaigns has been to generate statistics to characterize the local site turbulence conditions and their impact on widely distributed ground based antenna arrays. This is of critical importance for the situation of uplink arraying, in which a priori knowledge of the fast varying turbulent conditions of water vapor in the troposphere may not be known, and will impact the power combining efficiency of ground based transmitting arrays. Therefore, the design of these type of systems will be dependent on the local climatology of the particular ground station site. Based on the 30+ station years of data collected characterizing atmospheric phase scintillation statistics at various sites, a global model is presented which attempts to predict the average phase statistics of a generic site based on local surface weather data, such as surface pressure, temperature, relative humidity, wind speed, and median wind direction. A model is proposed and based on a standard log power distribution similar to amplitude scintillation models trained on the existing data sets and shows reasonable accuracy against existing data sets.

propagation↗