Engineering PapersSearch

SEARCH · Engineering Papers

Results for “training data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Machine Learning‐Assisted Microearthquake Location Workflow for Monitoring the Newberry Enhanced Geothermal System

Abstract Enhanced geothermal systems (EGS) offer a sustainable energy source but face challenges in accurately locating microearthquakes induced during reservoir stimulation. Locating these microearthquakes provides reliable feedback on the stimulation progress. Current deep learning methods for locating earthquakes require extensive data sets for training, which is problematic as detected microearthquakes are often limited. To address the scarcity of training data, we propose a practical workflow using probabilistic multilayer perceptron (PMLP) which predicts microearthquake locations from cross‐correlation time lags in waveforms. Utilizing a 3D velocity model of Newberry site derived from ambient noise interferometry, we generate numerous synthetic microearthquakes and 3D acoustic waveforms for PMLP training. Accurate synthetic tests prompt us to apply the trained network to the 2012 and 2014 stimulation field waveforms. To enhance the accuracy of source localization, we carefully handpick the P‐arrival times. Predictions on the 2012 stimulation data set show major microseismic activity at depths of 0.5–1.2 km, correlating with a known casing leakage scenario. In the 2014 data set, the majority of predictions concentrate at 2.0–2.9 km depths, consistent with results obtained from conventional physics‐based inversion, and align with the presence of natural fractures from 2.0 to 2.7 km. We validate our findings by comparing the synthetic and field picks, demonstrating a satisfactory match for the first arrivals. By combining the benefits of quick inference speeds and accurate location predictions, we demonstrate the feasibility of using realistic synthetic data set to locate microseismicity for EGS monitoring.

15 GEOTHERMAL ENERGY

One-shot gas detection with transformer paired neural networks in Mako collected longwave infrared hyperspectral imagery

To date, careful data treatment workflows and statistical detectors are used to perform hyperspectral image (HSI) detection of any gas contained in a spectral library, which is often expanded with physics models to incorporate different spectral characteristics. In general, surrounding evidence or known gas-release parameters are used to provide confidence in or confirm detection capability, respectively. This makes quantifying detection performance difficult as it is nearly impossible to develop an absolute ground truth for gas target pixel presence in collected HSI. Consequently, developing and comparing new detection methods, especially machine learning (ML)-based methods, is susceptible to subjectivity in derived detection map quality. Here, in this work, we demonstrate the first use of transformer-based paired neural networks (PNNs) for one-shot gas target detection for multiple gases while providing quantitative classification and detection metrics for their use on labeled data. Terabytes of training data are generated from a database of long-wave infrared HSI obtained from historical Mako sensor campaigns over Los Angeles. By incorporating labels, singular signature representations, and a model development pipeline, we can tune and select PNNs to detect multiple gas targets that are not seen in training on a quantitative basis. We additionally assess our test set detections using interpretability techniques widely employed with ML-based predictors, but less common with detection methods relying on learned latent spaces.

Hyperspectral imaging

A Computational Framework to design 3D stiffness gradient acoustic metamaterials for impedance matching

Acoustic waves play a crucial role in various applications, including medical imaging, non-destructive testing, and sonar systems. One of the significant challenges in these applications is impedance matching, which is essential for minimizing reflections and maximizing the transfer of acoustic energy between different media. Acoustic metamaterials offer a promising solution to this challenge. In addition to impedance control, gradient stiffness can enhance structural efficiency and enable spatial control of wave propagation, making it a valuable feature in acoustic metamaterial design. In this pa- per, we present our developed computational method to design 3D stiffness gradient acoustic metamaterials for impedance matching. The key steps in our approach include generating initial designs using a periodic covariance function to provide unit cells that are both periodic on the boundaries and randomly formed inside the unit cell. Furthermore, we integrated manufacturing constraints into the design process, ensuring that the structures are interconnected for fabrication. We propose two computational optimization algorithms: GenUnit, based on a non-dominated sorting genetic algorithm (NSGA-II), and MLMatch, which leverages differentiable machine learning. The two approaches are not separate contributions but complementary com- ponents of a unified framework. GenUnit requires no training data and directly interfaces with physics-based simulations, making it highly accurate but slower for large-scale exploration. In contrast, MLMatch is data-hungry during training but, once trained, enables near-instantaneous inference and broad design-space coverage. Together, they form a hybrid strategy: ML- Match rapidly explores the global design space, and GenUnit provides local refinement with high-fidelity accuracy. This balance between training cost, inference time, and precision is the motivation for including both methods in the same study. We applied this dual-algorithm framework to generate two metallic-based metamaterial designs that match the acoustic impedance of water while exhibiting a controlled gradient in stiffness (from stiff to soft). The stiffness gradient is particularly advantageous in applications where one side of the structure must interface with soft or sensitive surfaces, such as human tissue or delicate components. Here, this work paves the way for improved materials in various acoustic applications, particularly in ultrasound devices, by providing better impedance.

Metamaterial

Multivariate Testing of Sampling Techniques to Address Class Imbalance in Building Use Type Classification

This study addresses the challenges inherent in building use type classification, particularly focusing on the issue of class imbalance in the training datasets for machine learning classifiers. We comprehensively analyze the efficacy of various class-balancing sampling techniques. Employing Monte Carlo simulations and Bayesian optimization, we evaluated the performance of multiple sampling methods, including Random Oversampling, Random Undersampling, SMOTE, Borderline-SMOTE, and ADASYN, across a dataset encompassing nine southeastern coastal states of the United States. Our findings reveal that simple random over- and undersampling techniques outperform more sophisticated methods. Additionally, we show inherent value in creating an imbalance in training data to effectively train a machine learning classifier for distinguishing between residential and nonresidential buildings. This study provides valuable guidance for future research on building use type classification research and lays essential groundwork for developing attribute-rich building stock datasets.

Adams, Daniel

Feature extraction of multispectral data

A method is presented for feature extraction of multispectral scanner data. Non-training data is used to demonstrate the reduction in processing time that can be obtained by using feature extraction rather than feature selection.

Crane, R. B.

Improved Algorithms for Accurate Retrieval of UV - Visible Diffuse Attenuation Coefficients in Optically Complex, Inshore Waters

Photochemical processes driven by high-energy ultraviolet radiation (UVR) in inshore, estuarine, and coastal waters play an important role in global bio geochemical cycles and biological systems. A key to modeling photochemical processes in these optically complex waters is an accurate description of the vertical distribution of UVR in the water column which can be obtained using the diffuse attenuation coefficients of down welling irradiance (Kd()). The Sea UV Sea UVc algorithms (Fichot et al., 2008) can accurately retrieve Kd ( 320, 340, 380,412, 443 and 490 nm) in oceanic and coastal waters using multispectral remote sensing reflectances (Rrs(), Sea WiFS bands). However, SeaUVSeaUVc algorithms are currently not optimized for use in optically complex, inshore waters, where they tend to severely underestimate Kd(). Here, a new training data set of optical properties collected in optically complex, inshore waters was used to re-parameterize the published SeaUVSeaUVc algorithms, resulting in improved Kd() retrievals for turbid, estuarine waters. Although the updated SeaUVSeaUVc algorithms perform best in optically complex waters, the published SeaUVSeaUVc models still perform well in most coastal and oceanic waters. Therefore, we propose a composite set of SeaUVSeaUVc algorithms, optimized for Kd() retrieval in almost all marine systems, ranging from oceanic to inshore waters. The composite algorithm set can retrieve Kd from ocean color with good accuracy across this wide range of water types (e.g., within 13 mean relative error for Kd(340)). A validation step using three independent, in situ data sets indicates that the composite SeaUVSeaUVc can generate accurate Kd values from 320 490 nm using satellite imagery on a global scale. Taking advantage of the inherent benefits of our statistical methods, we pooled the validation data with the training set, obtaining an optimized composite model for estimating Kd() in UV wavelengths for almost all marine waters. This optimized composite set of SeaUVSeaUVc algorithms will provide the optical community with improved ability to quantify the role of solar UV radiation in photochemical and photobiological processes in the ocean.

UV

Multi Modality Brain Mapping System (MBMS) Using Artificial Intelligence and Pattern Recognition

A Multimodality Brain Mapping System (MBMS), comprising one or more scopes (e.g., microscopes or endoscopes) coupled to one or more processors, wherein the one or more processors obtain training data from one or more first images and/or first data, wherein one or more abnormal regions and one or more normal regions are identified; receive a second image captured by one or more of the scopes at a later time than the one or more first images and/or first data and/or captured using a different imaging technique; and generate, using machine learning trained using the training data, one or more viewable indicators identifying one or abnormalities in the second image, wherein the one or more viewable indicators are generated in real time as the second image is formed. One or more of the scopes display the one or more viewable indicators on the second image.

Kateb, Babak

Tropical Cyclone Intensity Estimation Using Deep Convolutional Neural Networks

Estimating tropical cyclone intensity by just using satellite image is a challenging problem. With successful application of the Dvorak technique for more than 30 years along with some modifications and improvements, it is still used worldwide for tropical cyclone intensity estimation. A number of semi-automated techniques have been derived using the original Dvorak technique. However, these techniques suffer from subjective bias as evident from the most recent estimations on October 10, 2017 at 1500 UTC for Tropical Storm Ophelia: The Dvorak intensity estimates ranged from T2.3/33 kt (Tropical Cyclone Number 2.3/33 knots) from UW-CIMSS (University of Wisconsin-Madison - Cooperative Institute for Meteorological Satellite Studies) to T3.0/45 kt from TAFB (the National Hurricane Center's Tropical Analysis and Forecast Branch) to T4.0/65 kt from SAB (NOAA/NESDIS Satellite Analysis Branch). In this particular case, two human experts at TAFB and SAB differed by 20 knots in their Dvorak analyses, and the automated version at the University of Wisconsin was 12 knots lower than either of them. The National Hurricane Center (NHC) estimates about 10-20 percent uncertainty in its post analysis when only satellite based estimates are available. The success of the Dvorak technique proves that spatial patterns in infrared (IR) imagery strongly relate to tropical cyclone intensity. This study aims to utilize deep learning, the current state of the art in pattern recognition and image recognition, to address the need for an automated and objective tropical cyclone intensity estimation. Deep learning is a multi-layer neural network consisting of several layers of simple computational units. It learns discriminative features without relying on a human expert to identify which features are important. Our study mainly focuses on convolutional neural network (CNN), a deep learning algorithm, to develop an objective tropical cyclone intensity estimation. CNN is a supervised learning algorithm requiring a large number of training data. Since the archives of intensity data and tropical cyclone centric satellite images is openly available for use, the training data is easily created by combining the two. Results, case studies, prototypes, and advantages of this approach will be discussed.

tropical cyclone intensity

Inception of a Spaceflight-specific Mouse to Human Expression Profiling Translation Model

Rodents are foundational model organisms often utilized due to their seemingly analogous morphologies and biological responses to humans. However, recent studies have demonstrated that murine model data are limited in their applicability, particularly in inflammatory disease. In space studies, accurately predicting human response from mouse data is critical due to extreme limiting factors in both rodent and human spaceflight research. With successful prediction, spaceflight ailments can be predicted and prevented while respecting the constraints of the spaceflight industry and minimizing danger to humans. To do so, novel methodologies must be developed that predict human response from murine data after considering biological differences between rodents and humans in spaceflight. After considering terrestrial models, we determined that a spaceflight-based expression profiting translation tool should be created to accurately capture predictions of human gene expression in spaceflight from mouse data. To prepare to build this model, we organized known human spaceflight risks, chose analog human diseases as training data categories, then identified existing RNASeq disease datasets from GEO as potential training data. In addition, we classified existing Genelab mouse differential gene expression datasets for use as experimental data.

Translation

Self-supervised physics-informed generative networks for phase retrieval from a single X-ray hologram

X-ray phase contrast imaging significantly improves the visualization of structures with weak or uniform absorption, broadening its applications across a wide range of scientific disciplines. Propagation-based phase contrast is particularly suitable for time- or dose-critical in vivo/in situ/operando (tomography) experiments because it requires only a single intensity measurement. However, the phase information of the wave field is lost during the measurement and must be recovered. Conventional algebraic and iterative methods often rely on specific approximations or boundary conditions that may not be met by many samples or experimental setups. In addition, they require manual tuning of reconstruction parameters by experts, making them less adaptable for complex or variable conditions. Here we present a self-learning approach for solving the inverse problem of phase retrieval in the near-field regime of Fresnel theory using a single intensity measurement (hologram). A physics-informed generative adversarial network is employed to reconstruct both the phase and absorbance of the unpropagated wave field in the sample plane from a single hologram. Unlike most state-of-the-art deep learning approaches for phase retrieval, our approach does not require paired, unpaired, or simulated training data. This significantly broadens the applicability of our approach, as acquiring or generating suitable training data remains a major challenge due to the wide variability in sample types and experimental configurations. The algorithm demonstrates robust and consistent performance across diverse imaging conditions and sample types, delivering quantitative, high-quality reconstructions for both simulated data and experimental datasets acquired at beamline P05 at PETRA III (DESY, Hamburg), operated by Helmholtz-Zentrum Hereon. Furthermore, it enables the simultaneous retrieval of both phase and absorption information.

36 MATERIALS SCIENCE

Inductive monitoring system constructed from nominal system data and its use in real-time system monitoring

The present invention relates to an Inductive Monitoring System (IMS), its software implementations, hardware embodiments and applications. Training data is received, typically nominal system data acquired from sensors in normally operating systems or from detailed system simulations. The training data is formed into vectors that are used to generate a knowledge database having clusters of nominal operating regions therein. IMS monitors a system's performance or health by comparing cluster parameters in the knowledge database with incoming sensor data from a monitored-system formed into vectors. Nominal performance is concluded when a monitored-system vector is determined to lie within a nominal operating region cluster or lies sufficiently close to a such a cluster as determined by a threshold value and a distance metric. Some embodiments of IMS include cluster indexing and retrieval methods that increase the execution speed of IMS.

Iverson, David L.

Extracting Material Property Measurements from Scientific Literature with Limited Annotations

Extracting material property data from scientific text is pivotal for advancing data-driven research in chemistry and materials science; however, the extensive annotation effort required to produce training data for named entity recognition (NER) models for this task often makes it a barrier to extracting specialized data sets. Here, in this work, we present a comparative study of the conventional, supervised NER methodology to alternative few-shot learning architectures and large language model (LLM)-based approaches that mitigate the need to label large training data sets. We find that the best-performing LLM (GPT-4o) not only excels in directly extracting relevant material properties based on limited examples but also enhances supervised learning through data augmentation. We supplement our findings with error and data quality assessments to provide a nuanced understanding of factors that impact property measurement extraction.

36 MATERIALS SCIENCE

Machine Learning Based Crater Detection for Terrain Relative Navigation

As Lunar exploration continues to become more commonplace, reliable methods of precise Terrain Relative Navigation (TRN) are needed. While there are many TRN techniques available, one that has received increased interest in the past few years is that of crater based navigation. Crater based navigation has numerous benefits, including being a human recognizable feature (important for crewed missions), as well as the fact that craters are often possible hazards that need to be detected and avoided. The use of crater based navigation has been limited however. This has been due to the difficulty of running such algorithms on board a spacecraft, as well as the difficulty in procuring large amounts of the required training data. This paper presents a new rendering tool for generating large amounts of high quality training data. It then looks at two recently developed machine learning techniques for crater detection and crater identification in real-time on near-future space hardware.

computer vision

System and Method for Outlier Detection via Estimating Clusters

An efficient method and system for real-time or offline analysis of multivariate sensor data for use in anomaly detection, fault detection, and system health monitoring is provided. Models automatically derived from training data, typically nominal system data acquired from sensors in normally operating conditions or from detailed simulations, are used to identify unusual, out of family data samples (outliers) that indicate possible system failure or degradation. Outliers are determined through analyzing a degree of deviation of current system behavior from the models formed from the nominal system data. The deviation of current system behavior is presented as an easy to interpret numerical score along with a measure of the relative contribution of each system parameter to any off-nominal deviation. The techniques described herein may also be used to "clean" the training data.

Iverson, David J.

Unveiling the Transferability of PLSR Models for Leaf Trait Estimation: Lessons from a Comprehensive Analysis with a Novel Global Dataset

Leaf traits are essential for understanding many physiological and ecological processes. Partial least-squares regression (PLSR) models with leaf spectroscopy are widely applied for trait estimation, but their transferability across space, time and plant functional types (PFTs) remains unclear. We compiled a novel dataset of paired leaf traits and spectra, with 47,393 records for >700 species and eight PFTs at 101 globally-distributed locations across multiple seasons. Using this dataset, we conducted an unprecedented comprehensive analysis to assess the transferability of PLSR models in estimating leaf traits. While PLSR models demonstrate commendable performance in predicting chlorophyll content, carotenoid, leaf water and leaf mass per area prediction within their training data space, their efficacy diminishes when extrapolating to new contexts. Specifically, extrapolating to locations, seasons, and PFTs beyond the training data leads to reduced R 2 (0.12-0.49, 0.15-0.42, and 0.25-0.56) and increased NRMSE (3.58-18.24%, 6.27-11.55% and 7.0-33.12%) compared to nonspatial random cross-validation (NRCV). The results underscore the importance of incorporating greater spectral diversity in model training to boost its transferability. These findings highlight potential errors in estimating leaf traits across large spatial domains, diverse PFTs and time due to biased validation schemes and provide guidance for future field sampling strategies and remote sensing applications.

Leaf traits

Chemical signature characterization with hyperspectral imagery: novel deep learning model architectures and physically-motivated data augmentation techniques

The high spectral resolution afforded by Hyperspectral Imaging (HSI) sensors is poised to bring unprecedented advancements to signature characterization applications. Thus far, much of the research in the machine learning field devoted to HSI applications has focused on a few specific tasks like land-use land-cover classification. In land classification tasks, spatial information is very important, and model architectures are often designed to leverage spatial contexts. However, it is unclear how well these spatially-tuned models will translate to tasks where spectral information is critical, like the detection and characterization of chemicals. In this work, we compare spectral models (inputs are 1D spectra) and spatial-spectral models (inputs are 3D cubes) in the context of predicting chemical concentration maps. We find that spatial-spectral models perform the best, though we find a wide range in performance across the different architectures tested. Additionally, we find that model performance is impacted by the availability of training data, particularly in scenarios where the training data doesn't fully capture the true variance of real-world conditions. We find that data augmentation can help mitigate sparse coverage of observed parameter space (e.g., seasonal or geographic variability in ground cover), and present augmentation strategies that are tailored to hyperspectral data.

• Artificial intelligence (AI) / machine learning

Grid-Ready Energy Analytics Training with Data (“GREAT with Data”)

GridEd is a collaborative educational initiative consisting of the Electric Power Research Institute (EPRI), 5 Partner Universities (Stony Brook University, The University of Texas at Austin, University of California – Riverside, Virginia Tech, Washington State University), and participating industry sponsors. This educational initiative focuses on developing and training the next generation of power engineers so they can help shape the electric grid of the future by anticipating and fulfilling the needs of changing electric industry requirements. GridEd is leveraging electric industry research to educate a future electric grid workforce by empowering new and continuing education students, not only to become competent and well-informed engineers, but also to participate and influence major technological, social, and policy decisions that address critical global challenges. GridEd’s activities are centered around four core pillars: Enhancement of university power systems engineering curricula; Professional development and training for a diverse electric industry workforce; Stimulating students to join the movement for the next generation of power engineers, and; Improve workforce development efforts in the electric utility industry. Major accomplishments over the course of the project were: Over 50 unique professional short courses were delivered by more than 30 instructors across the GridEd network; Over 3,600 unique learners, many who took multiple courses, received more than 27,000 professional development hours (PDH) and over 1,000 certificates of completion; Approximately 2,900 unique learners undertook a course that was offered LIVE online or in-person; Approximately 700 unique learners undertook a course that was offered as computer-based training (CBT); Over 35 unique university courses were delivered by more than 30 instructors to 1,500 university students across the GridEd network; One-hundred-and-fifty-seven (157) students were funded to completed 43 student projects in topics of power systems and data science, and; Six (6) Historically Black Colleges and Universities (HBCUs) recruited as Affiliate Universities via Utility Partners. The professional training initiative and workforce development activities launched by this project will be sustained through EPRI’s collaborative business model with industry. Stimulating students to join the power engineering workforce of the future and the enhancement of tertiary training may continue to need government support.

14 SOLAR ENERGY

Divide and conquer: Learning chaotic dynamical systems with multistep penalty neural ordinary differential equations

Forecasting high-dimensional dynamical systems is a fundamental challenge in various fields, such as geosciences and engineering. Neural Ordinary Differential Equations (NODEs), which combine the power of neural networks and numerical solvers, have emerged as a promising algorithm for forecasting complex nonlinear dynamical systems. However, classical techniques used for NODE training are ineffective for learning chaotic dynamical systems. In this work, we propose a novel NODE-training approach that allows for robust learning of chaotic dynamical systems. Here, our method addresses the challenges of non-convexity and exploding gradients associated with underlying chaotic dynamics. Training data trajectories from such systems are split into multiple, non-overlapping time windows. In addition to the deviation from the training data, the optimization loss term further penalizes the discontinuities of the predicted trajectory between the time windows. The window size is selected based on the fastest Lyapunov time scale of the system. Multi-step penalty(MP) method is first demonstrated on Lorenz equation, to illustrate how it improves the loss landscape and thereby accelerates the optimization convergence. MP method can optimize chaotic systems in a manner similar to least-squares shadowing with significantly lower computational costs. Our proposed algorithm, denoted the Multistep Penalty NODE, is applied to chaotic systems such as the Kuramoto-Sivashinsky equation, the two-dimensional Kolmogorov flow, and ERA5 reanalysis data for the atmosphere. It is observed that MP-NODE provide viable performance for such chaotic systems, not only for short-term trajectory predictions but also for invariant statistics that are hallmarks of the chaotic nature of these dynamics.

Chaotic dynamical systems