Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “supervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Seascape Interface Control Document (V.1)

This paper serves as the Interface Control Document (ICD) for the Seascape automated test harness developed at Sandia National Laboratories. The primary purposes of the Seascape system are: (1) provide a place for accruing large, curated, labeled data sets useful for developing and evaluating detection and classification algorithms (including, but not limited to, supervised machine learning applications) (2) provide an automated structure for specifying, running and generating reports on algorithm performance. Seascape uses GitLab, Nexus, Solr, and Banana, open source codes, together with code written in the Python language, to automatically provision and configure computational nodes, queue up jobs to accomplish algorithms test runs against the stored data sets, gather the results and generate reports which are then stored in the Nexus artifact server.

97 MATHEMATICS AND COMPUTING↗

Seascape Interface Control Document (V. 2)

This paper serves as the Interface Control Document (ICD) for the Seascape automated test harness developed at Sandia National Laboratories. The primary purposes of the Seascape system are: (1) provide a place for accruing large, curated, labeled data sets useful for developing and evaluating detection and classification algorithms (including, but not limited to, supervised machine learning applications) (2) provide an automated structure for specifying, running and generating reports on algorithm performance. Seascape uses GitLab, Nexus, Solr, and Banana, open source codes, together with code written in the Python language, to automatically provision and configure computational nodes, queue up jobs to accomplish algorithms test runs against the stored data sets, gather the results and generate reports which are then stored in the Nexus artifact server.

97 MATHEMATICS AND COMPUTING↗

Development of Hopfield Artificial Neural Network for Anomaly Detection in Environmental Gamma Radiation Background: Consortium on Nuclear Security Technologies (CONNECT) (Q2 Report)

Environmental screening of gamma radiation consists of detecting weak nuisance and anomaly signal in the presence of strong and highly varying background. In a typical scenario, a mobile detector-spectrometer continuously measures gamma radiation spectra in short, e.g., one-second, signal acquisition intervals. The measurement data is a 2D matrix, where one dimension is gamma ray energy, and the other dimension is the number of measurements or total time. In principle, gamma radiation sources can be detected and identified from the measured data by their unique spectral lines. Detecting sources from data measured in a search scenario is difficult due to the highly varying background because of naturally occurring radioactive material (NORM), and low signal-to-noise ratio (S/N) of spectral signal measured during one-second acquisition intervals. The objective of this work is to explore supervised machine learning (ML) algorithms for development of a Hopfield Neural Network (HNN) in conjunction with an image processing algorithm for detection and identification of weak nuisances and anomalies events in the presence of a highly fluctuating background.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Seascape Interface Control Document

This paper serves as the Interface Control Document (ICD) for the Seascape automated test harness developed at Sandia National Laboratories. The primary purposes of the Seascape system are: (1) provide a place for accruing large, curated, labeled data sets useful for developing and evaluating detection and classification algorithms (including, but not limited to, supervised machine learning applications) (2) provide an automated structure for specifying, running and generating reports on algorithm performance. Seascape uses GitLab, Nexus, Solr, and Banana, open source software, together with code written in the Python language, to automatically provision and configure computational nodes, queue up jobs to accomplish algorithms test runs against the stored data sets, gather the results and generate reports which are then stored in the Nexus artifact server.

97 MATHEMATICS AND COMPUTING↗

Applying novel analytical tools for analyzing multidimensional secondary organic aerosol measurements

In the atmosphere, secondary organic aerosols (SOA) are often the major components of fine particulate matter and interact with clouds and radiation. SOA comprises a mixture of thousands of organic compounds. There is tremendous complexity and uncertainty in understanding SOA formation, since it is formed by oxidation and gas to particle conversion of a variety of sources: natural biogenic, anthropogenic (vehicles, cooking coal combustion) and biomass burning. The Aerosol Mass Spectrometer (AMS) produces multidimensional chemical information about SOA but analyzing this data to understand SOA sources relies on time consuming analyses (~months to years) such as the positive matrix factorization (PMF). PMF also becomes difficult for aircraft data where signal to noise ratio is weaker. There is a critical need to develop fast machine learning techniques that can analytically provide information about SOA sources using AMS data on the same timescales as the data is being collected (~minutes). We apply a machine learning supervised classification approach: the multinomial logistic regression to rapidly classify AMS data obtained from aircraft measurements.

47 OTHER INSTRUMENTATION↗

Real-time neutron multiplicity and source localization for criticality safety during fuel debris removal

Advancing neutron detection and analysis techniques for complex radiation environments is an ongoing focus in nuclear instrumentation and monitoring. This proposal presents research and development of a generalized real-time neutron monitoring and analysis system, applicable to any detector capable of producing time-tagged neutron count data. While the work is demonstrated using the Neutron Multiplication Analysis Detector (NoMAD), a modular 15-tube helium-3 (He-3) array, due to its availability, spatial resolution, and flexible deployment, the methods developed are extensible to other systems, including organic scintillators and fast digital detectors. This research investigates two complementary analytical techniques for real-time characterization of neutron emitting sources: neutron multiplicity estimation based on the Hage-Cifarelli formalism and spatial localization using supervised machine learning applied to spatial count rate patterns. These methods are designed to operate under dynamic, evolving conditions such as fuel debris retrieval or reactor startup, where neutron-emitting material geometries may be partially unknown or changing over time. By integrating statistical neutron emission data with spatial localization, this research aims to develop and evaluate methods for real time neutron monitoring, source characterization, and material verification. Key contributions include implementation of a low-latency data pipeline for continuous neutron multiplicity analysis, development and validation of machine learning models for spatial inference, and experimental evaluation of system performance under variable measurement conditions. The outcomes are intended to support applications in nuclear safeguards, verification, emergency response, and reactor startup.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Training Restricted Boltzmann Machines With a D-Wave Quantum Annealer

Restricted Boltzmann Machine (RBM) is an energy-based, undirected graphical model. It is commonly used for unsupervised and supervised machine learning. Typically, RBM is trained using contrastive divergence (CD). However, training with CD is slow and does not estimate the exact gradient of the log-likelihood cost function. In this work, the model expectation of gradient learning for RBM has been calculated using a quantum annealer (D-Wave 2000Q), where obtaining samples is faster than Markov chain Monte Carlo (MCMC) used in CD. Training and classification results of RBM trained using quantum annealing are compared with the CD-based method. The performance of the two approaches is compared with respect to the classification accuracies, image reconstruction, and log-likelihood results. The classification accuracy results indicate comparable performances of the two methods. Image reconstruction and log-likelihood results show improved performance of the CD-based method. It is shown that the samples obtained from quantum annealer can be used to train an RBM on a 64-bit “bars and stripes” dataset with classification performance similar to an RBM trained with CD. Though training based on CD showed improved learning performance, training using a quantum annealer could be useful as it eliminates computationally expensive MCMC steps of CD.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Predicting SARS-CoV-2 Variant Using Non-Invasive Hand Odor Analysis: A Pilot Study

The adaptable nature of the SARS-CoV-2 virus has led to the emergence of multiple viral variants of concern. This research builds upon a previous demonstration of sampling human hand odor to distinguish SARS-CoV-2 infection status in order to incorporate considerations of the disease variants. This study demonstrates the ability of human odor expression to be implemented as a non-invasive medium for the differentiation of SARS-CoV-2 variants. Volatile organic compounds (VOCs) were extracted from SARS-CoV-2-positive samples using solid phase microextraction (SPME) coupled with gas chromatography–mass spectrometry (GC–MS). Sparse partial least squares discriminant analysis (sPLS-DA) modeling revealed that supervised machine learning could be used to predict the variant identity of a sample using VOC expression alone. The class discrimination of Delta and Omicron BA.5 variant samples was performed with 95.2% (±0.4) accuracy. Omicron BA.2 and Omicron BA.5 variants were correctly classified with 78.5% (±0.8) accuracy. Lastly, Delta and Omicron BA.2 samples were assigned with 71.2% (±1.0) accuracy. This work builds upon the framework of non-invasive techniques producing diagnostics through the analysis of human odor expression, all in support of public health monitoring.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Subcellular Feature-Based Classification of α and β Cells Using Soft X-ray Tomography

The dysfunction of α and β cells in pancreatic islets can lead to diabetes. Many questions remain on the subcellular organization of islet cells during the progression of disease. Existing three-dimensional cellular mapping approaches face challenges such as time-intensive sample sectioning and subjective cellular identification. To address these challenges, we have developed a subcellular feature-based classification approach, which allows us to identify α and β cells and quantify their subcellular structural characteristics using soft X-ray tomography (SXT). We observed significant differences in whole-cell morphological and organelle statistics between the two cell types. Additionally, we characterize subtle biophysical differences between individual insulin and glucagon vesicles by analyzing vesicle size and molecular density distributions, which were not previously possible using other methods. These sub-vesicular parameters enable us to predict cell types systematically using supervised machine learning. We also visualize distinct vesicle and cell subtypes using Uniform Manifold Approximation and Projection (UMAP) embeddings, which provides us with an innovative approach to explore structural heterogeneity in islet cells. This methodology presents an innovative approach for tracking biologically meaningful heterogeneity in cells that can be applied to any cellular system.

3D cell mapping↗

Analyzing and Predicting Effort Associated with Finding and Fixing Software Faults

Context: Software developers spend a significant amount of time fixing faults. However, not many papers have addressed the actual effort needed to fix software faults. Objective: The objective of this paper is twofold: (1) analysis of the effort needed to fix software faults and how it was affected by several factors and (2) prediction of the level of fix implementation effort based on the information provided in software change requests. Method: The work is based on data related to 1200 failures, extracted from the change tracking system of a large NASA mission. The analysis includes descriptive and inferential statistics. Predictions are made using three supervised machine learning algorithms and three sampling techniques aimed at addressing the imbalanced data problem. Results: Our results show that (1) 83% of the total fix implementation effort was associated with only 20% of failures. (2) Both safety critical failures and post-release failures required three times more effort to fix compared to non-critical and pre-release counterparts, respectively. (3) Failures with fixes spread across multiple components or across multiple types of software artifacts required more effort. The spread across artifacts was more costly than spread across components. (4) Surprisingly, some types of faults associated with later life-cycle activities did not require significant effort. (5) The level of fix implementation effort was predicted with 73% overall accuracy using the original, imbalanced data. Using oversampling techniques improved the overall accuracy up to 77%. More importantly, oversampling significantly improved the prediction of the high level effort, from 31% to around 85%. Conclusions: This paper shows the importance of tying software failures to changes made to fix all associated faults, in one or more software components and/or in one or more software artifacts, and the benefit of studying how the spread of faults and other factors affect the fix implementation effort.

software fix implementation effort↗

Pixel-Based Model For High Latitude Dust Detection

Dust has implications on the energy budget, ocean biodiversity, and economy at regional and global scales. Dust detection relies on spectral sensitivity at visible (RGB) and infrared wavelengths. Radiative properties of high latitude dust and the background surface albedo in these regions (>40°N, >40°S) complicate current dust detection methods. Leveraging supervised machine learning (ML) methods, we propose a new method accounting for regional differences of dust occurrence.

High latitude dust↗

Pixel Based Model For High Latitude Dust Detection

Current methods of dust detection rely on spectral sensitivity at visible (RGB) and infrared wavelengths. However, their application on different regions needs to be tuned to mitigate errors associated with background properties. High latitude dust (HLD) regions are characterized by surface with variable albedos and land cover, thus further complicating the dust detection. Leveraging supervised machine learning (ML) methods, we propose a new method accounting for regional differences of dust occurrence.

High latitude dust↗

Improving Sim-to-Real Transfer in Vision-Based Robot Navigation Via Instance-Level GAN-Based Data Augmentation

Achieving robust vision-based robotic tasks requires large amounts of data, which are often difficult to obtain in real-world scenarios. Simulators and synthetic data offer a cost-effective alternative, but the visual gap between simulation and reality hinders the performance of models when deployed in real-world environments. In this paper, we present a data augmentation pipeline that integrates a foundation model (Segment Anything Model) with an unsupervised image-to-image translation model (CycleGAN) for instance-level domain transfer from simulation to reality. This pipeline enables the generation of realistic labeled data from synthetic images for training supervised machine learning models in vision-based navigation tasks. We evaluate our approach on real-world data for ego-vehicle pose estimation, a critical autonomous navigation task involving the prediction of cross-track position and heading angle relative to road center line markings. The results of our tests show that our GAN-based data augmentation pipeline significantly outperforms models trained solely on simulation data or on data processed with standard image augmentation methods for sim-to-real transfer, enhancing model robustness and generalizability in real-world scenarios. Our method provides a scalable and flexible data augmentation tool for leveraging large synthetic datasets to enhance vision-based robotic navigation tasks.

artificial intelligence↗

A Model Based Approach to Extract Health Information from Textual Data

In current nuclear power plants (NPPs) a large amount of condition-based data is being generated and stored to assess and monitor component health and performance. The format of this data can be either numeric (e.g., pump vibration data) or textual (e.g., condition report which assess component health). While assessing component health from numeric data can be performed with a large variety of methods, the extraction of information from textual data still remains a challenge. Natural language processing (NLP) methods are starting to be deployed in current NPPs mainly to filter out incident reports (IRs) that are not safety related by employing supervised machine learning methods. However, these methods do not really provide the quantitative information that might be contained in IRs. This paper presents an approach to extract information from textual data (e.g., from IRs, maintenance reports) that is based on NLP data analytics methods coupled with model-based system engineer (MBSE) models. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence; such analysis includes: part of speech (POS) tagging (i.e., identification of grammatic elements of each string - e.g., nouns, verbs), named entity recognition (i.e., identification of text entities - e.g., names, dates, events), and relation extraction (e.g., coreference resolution). On the other hand, semantic analysis is designed to analyze the logic structure of a sentence. Through a specific set of rules, our methods can identify whether a sentence contains health information of a component (e.g., degraded performance, anomaly behavior) or the causal relationship between two events (i.e., a cause-effect pair). An innovative element of our approach is that semantic analysis relies on MBSE models to identify links between textual elements. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. This paper presents in detail how the integration of NLP methods and MBSE models is performed. Few analysis examples focusing on centrifugal pumps are presented.

97 - MATHEMATICS AND COMPUTING↗

Latent Representation Learning for Structural Characterization of Catalysts

Supervised machine learning-enabled mapping of the X-ray absorption near edge structure (XANES) spectra to local structural descriptors offers new methods for understanding the structure and function of working nanocatalysts. We briefly summarize a status of XANES analysis approaches by supervised machine learning methods. We present an example of an autoencoder-based, unsupervised machine learning approach for latent representation learning of XANES spectra. This new approach produces a lower-dimensional latent representation, which retains a spectrum–structure relationship that can be eventually mapped to physicochemical properties. Furthermore, the latent space of the autoencoder also provides a pathway to interpret the information content “hidden” in the X-ray absorption coefficient. Our approach (that we named latent space analysis of spectra, or LSAS) is demonstrated for the supported Pd nanoparticle catalyst studied during the formation of Pd hydride. By employing the low-dimensional representation of Pd K-edge XANES, the LSAS method was able to isolate the key factors responsible for the observed spectral changes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predicting the oxidation states of Mn ions in the oxygen-evolving complex of photosystem II using supervised and unsupervised machine learning

Abstract Serial Femtosecond Crystallography at the X-ray Free Electron Laser (XFEL) sources enabled the imaging of the catalytic intermediates of the oxygen evolution reaction of Photosystem II (PSII). However, due to the incoherent transition of the S-states, the resolved structures are a convolution from different catalytic states. Here, we train Decision Tree Classifier and K-means clustering models on Mn compounds obtained from the Cambridge Crystallographic Database to predict the S-state of the X-ray, XFEL, and CryoEM structures by predicting the Mn’s oxidation states in the oxygen-evolving complex. The model agrees mostly with the XFEL structures in the dark S 1 state. However, significant discrepancies are observed for the excited XFEL states (S 2 , S 3, and S 0 ) and the dark states of the X-ray and CryoEM structures. Furthermore, there is a mismatch between the predicted S-states within the two monomers of the same dimer, mainly in the excited states. We validated our model against other metalloenzymes, the valence bond model and the Mn spin densities calculated using density functional theory for two of the mismatched predictions of PSII. The model suggests designing a more optimized sample delivery and illumiation systems are crucial to precisely resolve the geometry of the advanced S-states to overcome the noncoherent S-state transition. In addition, significant radiation damage is observed in X-ray and CryoEM structures, particularly at the dangler Mn center (Mn4). Our model represents a valuable tool for investigating the electronic structure of the catalytic metal cluster of PSII to understand the water splitting mechanism.

Plant Sciences↗

Supervised and unsupervised machine learning of structural phases of polymers adsorbed to nanowires

Here, we identify configurational phases and structural transitions in a polymer nanotube composite by means of machine learning. We employ various unsupervised dimensionality reduction methods, conventional neural networks, as well as the confusion method, an unsupervised neural-network-based approach. We find neural networks are able to reliably recognize all configurational phases that have been found previously in experiment and simulation. Furthermore, we locate the boundaries between configurational phases in a way that removes human intuition or bias. This could be done before only by relying on preconceived, ad hoc order parameters.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗