Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Harnessing on-machine metrology data for prints with a surrogate model for laser powder directed energy deposition

In this study, we leverage the massive amount of multi-modal on-machine metrology data generated from Laser Powder Directed Energy Deposition (LP-DED) to construct a comprehensive surrogate model of the 3D printing process. By employing Dynamic Mode Decomposition with Control (DMDc), a data-driven technique, we capture the complex physics inherent in this extensive dataset. This physics-based surrogate model emphasizes thermodynamically significant quantities, enabling us to accurately predict key process outcomes. The model ingests 21 process parameters, including laser power, scan rate, and position, while providing outputs such as melt pool temperature, melt pool size, and other essential observables. Furthermore, it incorporates uncertainty quantification to provide bounds on these predictions, enhancing reliability and confidence in the results. We then deploy the surrogate model on a new, unseen part and monitor the printing process as validation of the method. Our experimental results demonstrate that the predictions align with actual measurements with high accuracy, confirming the effectiveness of our approach. Furthermore, this methodology not only facilitates real-time predictions but also operates at process-relevant speeds, establishing a basis for implementing feedback control in LP-DED.

Digital twins↗

Deconvoluting thermomechanical effects in X-ray diffraction data using machine learning

X-ray diffraction is ideal for probing the sub-surface state during complex or rapid thermomechanical loading of crystalline materials. However, challenges arise as the size of diffraction volumes increases due to spatial broadening and because of the inability to deconvolute the effects of different lattice deformation mechanisms. Here, we present a novel approach that uses combinations of physics-based modeling and machine learning to deconvolve thermal and mechanical elastic strains for diffraction data analysis. The method builds on a previous effort to extract thermal strain distribution information from diffraction data. The new approach is applied to extract the evolution of the thermomechanical state during laser melting of an Inconel 625 wall specimen which produces significant residual stress upon cooling. A combination of heat transfer and fluid flow, elasto-plasticity and X-ray diffraction simulations is used to generate training data for machine-learning (Gaussian process regression, GPR) models that map diffracted intensity distributions to underlying thermomechanical strain fields. First-principles density functional theory is used to determine accurate temperature-dependent thermal expansion and elastic stiffness used for elasto-plasticity modeling. The trained GPR models are found to be capable of deconvoluting the effects of thermal and mechanical strains, in addition to providing information about underlying strain distributions, even from complex diffraction patterns with irregularly shaped peaks.

36 MATERIALS SCIENCE↗

Improving vertical detail in simulated temperature and humidity data using machine learning

Atmospheric models used for weather forecasting and climate predictions discretise the atmosphere onto a vertical grid. There are however atmospheric phenomena that occur on scales smaller than the thickness of those model layers. The formation of low-level clouds due to temperature inversions is an example. This leads to atmospheric models underestimating, or even missing, these clouds and their radiative effects. Using radiosonde observations as training data, a machine learning model is used to improve the vertical detail of modelled profiles of temperature and specific humidity. In addition, a physics-informed machine learning model is developed and compared to the traditional approach; showing improvements in the cloud fraction profiles calculated from its predictions. The vertically enhanced profiles also improve the representation of layers of convective inhibition and anomalous refractivity gradients. This work facilitates targeted improvements to the representation of certain atmospheric processes without the burden of increased memory and computational cost from increasing vertical resolution throughout the whole model.

54 ENVIRONMENTAL SCIENCES↗

Decoding diffraction and spectroscopy data with machine learning: A tutorial

This Tutorial provides a step-by-step guide on how to apply supervised machine-learning techniques to analyze diffraction and spectroscopy data. This Tutorial details four models—a reconstruction-focused model, a regression-focused model, a hybrid reconstruction/regression model, and a multimodal model—that use x-ray diffraction profiles and vibrational density of states spectra to predict various microstructural descriptors. In this Tutorial, we cover data pre-processing steps, constructions of the models via dimensionality reduction and regression, training, and analysis of these models. Comparisons of the model’s performance are provided, highlighting the strength and weakness of the various approaches utilized.

36 MATERIALS SCIENCE↗

Macroscopic Traffic Modeling Using Probe Vehicle Data: A Machine Learning Approach

Abstract The macroscopic fundamental diagram (MFD) captures an orderly relationship among traffic flow, density, and speed at the network level. It is a simple yet powerful tool for modeling traffic dynamics in large urban networks with broad application in traffic control and management. However, empirically derived MFDs in urban regions require high-resolution traffic data from the network. Having the network flow and vehicular density estimated at the (granular) census tract level using vehicle probe data, we apply machine learning methods to predict the MFDs across U.S. urban areas and capture the impacts of location-specific input features on the network flow–density relationships at a large scale. The results show that, among the four tested machine learning approaches (Random Forest, XGBoost, Support Vector Machine, and Neural Network), XGBoost delivers the best performance in predicting network traffic flow based on vehicular density and location attributes. Using interaction Shapley Additive explanation (SHAP) values and partial correlation analysis, we examine the factors influencing MFD shapes across different locations. Our empirical findings reveal that across U.S. urban areas, network topology, transportation infrastructure, and land use are primary factors shaping MFD curves, while demand and trip-related factors play a lesser role. Specifically, higher ranking roads, centrality, and development levels correlate positively with network capacity and critical density, whereas negative associations are observed for network connectivity, mixed-use development, and road roughness levels.

Jin, Ling↗

ImageLabler: Labeling and Managing Image Data for Machine Learning in the Earth Sciences

While machine learning techniques for image classification have been around for a long time, storing and managing the vast number of images required as training data is still a problem for scientists. This is especially true for the field of Earth science, where only recently have experts begun using machine learning techniques for image-based phenomena classification. Image Labeler, a fast and scalable cloud-based tagging platform for Earth science images, seeks to improve upon existing methods of managing images and associated metadata, such as maintaining categorized folders of images on a local machine, a process that can be cumbersome and difficult to scale. The platform facilitates rapid development of image-based Earth science phenomena training datasets by allowing scientists to upload their existing imagery as well as extract new samples from open satellite imagery services made available through NASA’s Global Imagery Browse Service (GIBS). Image Labeler also supports GeoTIFF data, with capabilities such as displaying GeoTIFFs on an interactive map, drawing shapefiles over them, and tagging them with additional metadata. This allows scientists to perform spatiotemporal subsetting with geographic information and develop training data more quickly. Built using modern web technologies, Image Labeler includes additional capabilities such as team collaboration for large-scale image tagging projects. Users can download their data in a machine-learning-ready format, allowing scientists to spend time on experimentation rather than on the collection of training data. In this presentation, we demonstrate how Image Labeler seeks to become a one-stop image data management solution for machine learning applications in Earth science.

Ashish Acharya↗

High speed machining of space shuttle external tank liquid hydrogen barrel panel

Actual and projected optimum High Speed Machining data for producing shuttle external tank liquid hydrogen barrel panels of aluminum alloy 2219-T87 are reported. The data included various machining parameters; e.g., spindle speeds, cutting speed, table feed, chip load, metal removal rate, horsepower, cutting efficiency, cutter wear (lack of) and chip removal methods.

Hankins, J. D.↗

Sub-Continental-Scale Carbon Stocks of Individual Trees in African Drylands

The distribution of dryland trees and their density, cover, size, mass and carbon content are not well known at sub-continental to continental scales. This information is important for ecological protection, carbon accounting, climate mitigation and restoration efforts of dryland ecosystems. We assessed more than 9.9 billion trees derived from more than 300,000 satellite images, covering semi-arid sub-Saharan Africa north of the Equator. We attributed wood, foliage and root carbon to every tree in the 0–1,000 mm year −1 rainfall zone by coupling field data, machine learning, satellite data and high-performance computing. Average carbon stocks of individual trees ranged from 0.54 Mg C ha −1 and 63 kg C tree −1 in the arid zone to 3.7 Mg C ha −1 and 98 kg tree −1 in the sub-humid zone. Overall, we estimated the total carbon for our study area to be 0.84 (±19.8%) Pg C. Comparisons with 14 previous TRENDY numerical simulation studies23 for our area found that the density and carbon stocks of scattered trees have been underestimated by three models and overestimated by 11 models, respectively. This benchmarking can help understand the carbon cycle and address concerns about land degradation. We make available a linked database of wood mass, foliage mass, root mass and carbon stock of each tree for scientists, policymakers, dryland-restoration practitioners and farmers, who can use it to estimate farmland tree carbon stocks from tablets or laptops.

Compton Tucker↗

Construction of a Fluid Flowfield from Discrete Point Data using Machine Learning

Many verification and validation procedures in aerospace engineering involve the comparison of computational fluid dynamics (CFD) data to experimental results from sources like wind tunnel tests. However, an incongruity exists between the data available from these sources: flow visualization is available by default in computational data, whereas in most experimental setups the available data is far more discrete and far more limited: integrated forces and moments, discrete pressure and temperature probes, etc. When differences exist between quantities of interest like lift and drag coefficients, the lack of full-field flow data from the experiments complicates most attempts to reconcile why the different data sources disagree. To this end, a shallow neural network, constrained by certain fluid flow properties, was trained to approximate flow field snapshots given only discrete data like that available in a wind tunnel test. The constructed snapshots, even for complex incompressible fluid flows, were found to agree at the large scales with the true flow fields. With this tool, researchers can more readily and easily understand why quantities of interest differ between their experimental and computational datasets. This in turn improves the resulting data's uncertainty measures.

Yury Lebedev↗

Construction of a Fluid Flowfield from Discrete Point Data using Machine Learning

Many verification and validation procedures in aerospace engineering involve the comparison of computational fluid dynamics (CFD) data to experimental results from sources like wind tunnel tests. However, an incongruity exists between the data available from these sources: flow visualization is available by default in computational data, whereas in most experimental setups the available data is far more discrete and far more limited: integrated forces and moments, discrete pressure and temperature probes, etc. When differences exist between quantities of interest like lift and drag coefficients, the lack of full-field flow data from the experiments complicates most attempts to reconcile why the different data sources disagree. To this end, a shallow neural network, constrained by certain fluid flow properties, was trained to approximate flow field snapshots given only discrete data like that available in a wind tunnel test. The constructed snapshots, even for complex incompressible fluid flows, were found to agree at the large scales with the true flow fields. With this tool, researchers can more readily and easily understand why quantities of interest differ between their experimental and computational datasets. This in turn improves the resulting data's uncertainty measures.

Yury Lebedev↗

Prediction of electric and magnetic fields from spectral data using machine learning algorithms for Doppler-free saturation spectroscopy diagnostics

The prediction of electric and magnetic field amplitudes from atomic spectral data is critical for plasma control in fusion devices such as tokamaks. Conventional approaches that rely on physics-based models are computationally expensive and unsuitable for real-time applications. In this work, we develop and benchmark three machine learning algorithms—simulation-based inference (SBI), fully connected neural networks (FCNN), and histogram-based gradient boosting regression (GBR-Hist)—to infer field intensities directly from Doppler-free saturation spectroscopy (DFSS) spectra. Synthetic datasets of spectra were generated using the EZSSS code and evaluated both with and without added Poisson noise to mimic experimental conditions. We find that SBI achieves the highest accuracy and robustness, FCNN provides a strong balance of accuracy and computational efficiency for real-time applications, and GBR-Hist offers the fastest inference but is more sensitive to noise. Furthermore, these results demonstrate the potential of machine learning to accelerate DFSS analysis and enhance its utility for plasma diagnostics and control.

Doppler-free saturation spectroscopy↗

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES↗

Scalability analysis of heavy-duty gas turbines using data-driven machine learning

With the increasing integration of variable renewable energy sources into power systems, the role of flexible power generation technologies like gas turbines (GT) in rapid grid balancing remains crucial. This sustained importance underscores the need for scaled and precise modeling of GT to ensure effective integration within evolving energy frameworks. While physics-driven GT models integrate thermodynamics, fluid dynamics, and combustion principles, they often rely on approximate mathematical representations to accommodate scaling that may not capture the actual complex dynamics for GTs and inertial effects associated to GTs with different ratings. In this study, a data-driven model is proposed using machine learning (ML) techniques to conduct GT scalability analysis and performance evaluation with high accuracy. The ML model, trained on data from various operating conditions and performance parameters, aims to uncover intricate relationships and patterns, resembling GT characteristics at different scales (ratings). The model is developed to capture complex system interaction and to adapt to changing operational scenarios at different capacities, providing valuable insights of power system dynamics. In this study, the real-time digital simulator platform was employed to generate training data for the ML model and assess its dynamic characteristics. The ultimate objective was to develop a detailed modeling framework based on governing equations and data-driven ML capable of predicting key performance indicators, in thermal systems such as GTs, including power output, speed, fuel consumption, and exhaust temperature under diverse operating conditions at different scales. The developed ML framework demonstrated high accuracy, with mean relative errors for GT power prediction, reference speed, exhaust temperature, and compressor pressure ratio (CPR) parameters consistently below 0.1% across typical load fluctuation scenarios. Maximum deviations were limited to approximately 0.5 K for exhaust temperature and 0.009 for CPR, underscoring the model’s ability to replicating dynamic GT behavior with high precision. The adaptability of the ML model enables its application across diverse operational conditions and its extension to other thermal systems. By leveraging advanced ML techniques, this study presents a robust and scalable modeling framework that enhances GT simulation precision, facilitating improved integration into evolving power systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Space Flown Rodent Liver RNA Sequencing Data for Machine Learning in Space Biology Research

High-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. Data analysis has been accelerated in recent years by the adoption of artificial intelligence (AI) and machine learning (ML) techniques by biomedical researchers. In space biology research, RNAseq datasets from space-flown experimental samples are critical for characterizing the gene expression aberrations associated with exposure to spaceflight stressors. However, space biological experiments tend to be very low sample size, so identifying proper AI/ML algorithms for sequencing data analysis is an ongoing challenge since these algorithms typically require large sample size. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML”, focused on creating datasets meant for three main applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. These scientific benchmarks consist of an AI-ready dataset and a reference implementation on a specific scientific question. In this work, we focused on generating standardized datasets to allow the scientific community to benchmark AI/ML algorithms in the domain of space biology. We present here a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data as a collaboration between the NASA AI4LS (Artificial Intelligence for Life Sciences) working group. and NASA’s SMD. This dataset consists of space-flown and ground control mouse liver found in the NASA GeneLab omics database. However, to amplify the small sample number (n=112 samples) for ML purposes, we employ Gaussian noise and a generative adversarial network to extend this dataset to 6,000 synthetic samples, matching the original gene expression characteristics.

James Casaletto↗

Detecting Process Equipment Failures Using Acoustic Data and Machine Learning

Nuclear power plant (NPP) process equipment such as fans, motors, valves, and pumps generate frequent or continuous noise, and deviations from the normal operational sounds made by this equipment can indicate potential issues. These deviations can be identified via automated acoustic anomaly detection, which involves using acoustic sensors (i.e., microphones) alongside detection algorithms to continuously monitor for changes in acoustic signatures. This task is made challenging by the substantial background noise that exists, such as operators opening and closing doors, manipulating valves, and conversing—in addition to typical plant noises. In collaboration with a nuclear power utility partner, this effort assessed the efficacy of acoustic anomaly detection when using a specific acoustic sensor that compresses data into a fixed set of features that are transferable over a standard Internet of Things communication protocol, thereby improving usability but potentially degrading detection performance. Two methods of performing automated acoustic anomaly detection were evaluated: one-class support vector machine (OC-SVM) and isolation forest (iForest). To enable the use of high-quality acoustic data encompassing both normal and anomalous conditions, the study utilized the publicly available Malfunctioning Industrial Machine Investigation and Inspection dataset, which includes real measured acoustic sensor data for a range of equipment types, model numbers, and signal-to-noise ratios (SNRs), along with a benchmark set of detection results. Using this dataset, the methods were tested and then compared against the benchmark results. The results indicated that although the specific acoustic sensor did not enable as rich a feature set extraction, the proposed methods with the limited feature set performed just as well. This provides solid justification for both the methods and the use of the proposed acoustic sensor.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Breaking the barrier of human-annotated training data for machine learning-aided plant research using aerial imagery

Machine learning (ML) can accelerate biological research. However, the adoption of such tools to facilitate phenotyping based on sensor data has been limited by (i) the need for a large amount of human-annotated training data for each context in which the tool is used and (ii) phenotypes varying across contexts defined in terms of genetics and environment. This is a major bottleneck because acquiring training data is generally costly and time-consuming. This study demonstrates how a ML approach can address these challenges by minimizing the amount of human supervision needed for tool building. A case study was performed to compare ML approaches that examine images collected by an uncrewed aerial vehicle to determine the presence/absence of panicles (i.e. “heading”) across thousands of field plots containing genetically diverse breeding populations of 2 Miscanthus species. Automated analysis of aerial imagery enabled the identification of heading approximately 9 times faster than in-field visual inspection by humans. Leveraging an Efficiently Supervised Generative Adversarial Network (ESGAN) learning strategy reduced the requirement for human-annotated data by 1 to 2 orders of magnitude compared to traditional, fully supervised learning approaches. The ESGAN model learned the salient features of the data set by using thousands of unlabeled images to inform the discriminative ability of a classifier so that it required minimal human-labeled training data. This method can accelerate the phenotyping of heading date as a measure of flowering time in Miscanthus across diverse contexts (e.g. in multistate trials) and opens avenues to promote the broad adoption of ML tools.

59 BASIC BIOLOGICAL SCIENCES↗

Quantifying dispersity in size and shape of nanoparticles from small-angle scattering data using machine learning based CREASE

Here, we use machine learning (ML) enhanced computational reverse engineering analysis of scattering experiments (CREASE) to interpret small-angle X-ray scattering (SAXS) data obtained from a system of nanoparticles without a priori knowledge of their exact shapes (e.g. spheres or ellipsoids), sizes (0.5–50 nm) and distributions. The SAXS measurements yielded three categories of scattering profiles exhibiting 'strong', 'weak' and 'no' features. Diminishing features (e.g. broadening or disappearing peaks) in scattering profiles have always been attributed to the presence of significant dispersity in the system. Such featureless SAXS data are not suitable for traditional analysis using analytical models. If one were to fit a relevant analytical model (e.g. the lmfit analytical model for polydisperse spheres) to these 'weak' and 'no' SAXS profiles from our nanoparticle systems, one would obtain non-unique interpretations of the data. Relying on electron microscopy to identify the distributions of nanoparticle shapes and sizes is also unfeasible, especially in high-throughput synthesis and characterization loops. In such situations, to identify the distributions of particle sizes and shapes that could be present in the sample, one must rely on methods like ML-CREASE to interpret the data quickly and output all relevant interpretations about the structure present in the system. The ML-CREASE optimization loop takes the experimental scattering profile as input and outputs multiple candidate solutions whose computed scattering profiles match the SAXS profile input. The ML-CREASE method outputs distributions of relevant structural features, such as the volume fraction of the nanoparticles in the system and the mean and standard deviation of the particle size and aspect ratio, assuming a type of distribution (e.g. normal, log-normal) for size and aspect ratio. We find that, for the SAXS profiles analyzed here, accounting for the shape dispersity along with size dispersity of the nanoparticles using ML-CREASE improved the match between the computed scattering profiles and input experimental profiles.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗