Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “supervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Structure of chalcogen overlayers on Au(111): Density functional theory and lattice-gas modeling

Ordering of different chalcogens, S, Se, and Te, on Au(111) exhibit broad similarities but also some distinct features, which must reflect subtle differences in relative values of the long-range pair and many-body lateral interactions between adatoms. We develop lattice-gas (LG) models within a cluster expansion framework, which includes about 50 interaction parameters. These LG models are developed based on density functional theory (DFT) analysis of the energetics of key adlayer configurations in combination with the Monte Carlo (MC) simulation of the LG models to identify statistically relevant adlayer motifs, i.e., model development is based entirely on theoretical considerations. The MC simulation guides additional DFT analysis and iterative model refinement. Given their complexity, development of optimal models is also aided by strategies from supervised machine learning. The model for S successfully captures ordering motifs over a broader range of coverage than achieved by previous models, and models for Se and Te capture the features of ordering, which are distinct from those for S. More specifically, the modeling for all three chalcogens successfully explains the linear adatom rows (also subtle differences between them) observed at low coverages of ~0.1 monolayer. The model for S also leads to a new possible explanation for the experimentally observed phase with a (5 × 5)-type low energy electron diffraction (LEED) pattern at 0.28 ML and to predictions for LEED patterns that would be observed with Se and Te at this coverage.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Extended Fayans energy density functional: optimization and analysis

The Fayans energy density functional (EDF) has been very successful in describing global nuclear properties (binding energies, charge radii, and especially differences of radii) within nuclear density functional theory. In a recent study, supervised machine learning methods were used to calibrate the Fayans EDF. Building on this experience, in this work we explore the effect of adding isovector pairing terms, which are responsible for different proton and neutron pairing fields, by comparing a 13D model without the isovector pairing term against the extended 14D model. At the heart of the calibration is a carefully selected heterogeneous dataset of experimental observables representing ground-state properties of spherical even–even nuclei. To quantify the impact of the calibration dataset on model parameters and the importance of the new terms, we carry out advanced sensitivity and correlation analysis on both models. The extension to 14D improves the overall quality of the model by about 30%. The enhanced degrees of freedom of the 14D model reduce correlations between model parameters and enhance sensitivity.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A Morphological Model to Separate Resolved–Unresolved Sources in the DESI Legacy Surveys: Application in the LS4 Alert Stream

Separating resolved and unresolved sources in large imaging surveys is a fundamental step to enable downstream science, such as searching for extragalactic transients in wide-field time-domain surveys. Here we present our method to effectively separate point sources from the resolved, extended sources in the Dark Energy Spectroscopic Instrument (DESI) Legacy Surveys (LS). We develop a supervised machine learning model based on the Gradient Boosting algorithm XGBoost. The features input to the model are purely morphological and are derived from the tabulated LS data products. We train the model using ∼2 × 10 5 LS sources in the COSMOS field with HST morphological labels and evaluate the model performance on LS sources with spectroscopic classification from the DESI Data Release 1 (∼2 × 10 7 objects) and the Sloan Digital Sky Survey Data Release 17 (∼3 × 10 6 objects), as well as on ∼2 × 10 8 Gaia stars. A significant fraction of LS sources are not observed in every LS filter, and we therefore build a “Hybrid” model as a linear combination of two XGBoost models, each containing features combining aperture flux measurements from the “blue” (gr) and “red” (iz) filters. The Hybrid model shows a reasonable balance between sensitivity and robustness, and achieves higher accuracy and flexibility compared to the LS morphological typing. With the Hybrid model, we provide classification scores for ∼3 × 10 9 LS sources, making this the largest ever machine learning catalog separating resolved and unresolved sources. The catalog has been incorporated into the real-time pipeline of the La Silla Schmidt Southern Survey (LS4), enabling the identification of extragalactic transients within the LS4 alert stream.

astrostatistics↗

Discovery of two bright high-redshift gravitationally lensed quasars revealed by Gaia

We present the discovery and preliminary characterisation of two high-redshift gravitationally lensed quasar systems in Gaia Data Release 2 (DR2). Candidates with multiple close-separation Gaia detections and quasar-like colours in WISE, Pan-STARRS, and DES are selected for follow-up spectroscopy with the New Technology Telescope. We confirm DES J215028.71-465251.3 as a $z$ = 4.130 ± 0.006 asymmetric, doubly imaged lensed quasar system and model the lensing mass distribution as a singular isothermal sphere. The system has an Einstein radius of 1.202 ± 0.005 arcsec and a predicted time delay of ~122.0 d between the quasar images, assuming a lensing galaxy redshift of $z$ = 0.5, making this a priority system for future optical monitoring. We confirm PS J042913.17+142840.9 as a $z$ = 3.866 ± 0.003 four-image quasar system in a cusp configuration, lensed by two foreground galaxies. The system is well modelled using a singular isothermal ellipsoid for the primary lens and a singular isothermal sphere for the secondary lens with Einstein radii 0.704 ± 0.006 and 0.241 ± 0.030 arcsec, respectively. A maximum predicted time delay of 9.6 d is calculated, assuming lensing galaxy redshifts of $z$ = 1.0. Furthermore, PS J042913.17+142840.9 exhibits a large flux ratio anomaly, up to a factor of 2.66 ± 0.37 in i band, that varies across optical and near-infrared wavelengths. We discuss LSST and its implications for future high-redshift lens searches and outline an extension to the search using supervised machine learning techniques.

Astronomy & Astrophysics↗

Autonomous Tuning and Charge-State Detection of Gate-Defined Quantum Dots

Defining quantum dots in semiconductor-based heterostructures is an essential step in initializing solid-state qubits. With growing device complexity and increasing number of functional devices required for measurements, a manual approach to finding suitable gate voltages to confine electrons electrostatically is impractical. Here, we implement a two-stage device characterization and dot-tuning process, which first determines whether devices are functional and then attempts to tune the functional devices to the single or double quantum-dot regime. We show that automating well-established manual-tuning procedures and replacing the experimenter’s decisions by supervised machine learning is sufficient to tune double quantum dots in multiple devices without premeasured input or manual intervention. The quality of measurement results and charge states are assessed by four binary classifiers trained with experimental data, reflecting real device behavior. We compare and optimize eight models and different data preprocessing techniques for each of the classifiers to achieve reliable autonomous tuning, an essential step towards scalable quantum systems in quantum-dot-based qubit architectures.

Darulova, J↗

VoroClust: Scalable Clustering for Remote Sensing

Although supervised machine learning provides a powerful framework for image classification and segmentation, it requires comprehensive consistent datasets, which are not available for many remote-sensing applications. Remote-sensing datasets are expensive to collect, and each is acquired under different environmental conditions or with significant variations in system operating parameters. Unsupervised clustering algorithms analyze the structure of each dataset independently, rather than drawing on similarities with existing “training” examples, and are thus well suited for practical remote-sensing applications. We introduce VoroClust, a fast density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. VoroClust runs as fast as distance-based clustering methods, while capturing complex regional geometries at least as well as current-density-based methods. It uses a data-centered sphere cover to reduce computational demands, while still capturing data topology. It then propagates clusters outward from local peaks in density. We show that VoroClust provides fast state-of-the-art clustering for both high-resolution polarimetric synthetic aperture radar and high-dimensional hyperspectral imaging datasets.

42 ENGINEERING↗

Community‐Level Metabolic Shifts Following Land Use Change in the Amazon Rainforest Identified by a Supervised Machine Leaning Approach

ABSTRACT The Amazon rainforest has been subjected to high rates of deforestation, mostly for pasturelands, over the last few decades. This change in plant cover is known to alter the soil microbiome and the functions it mediates, but the genomic changes underlying this response are still unresolved. In this study, we used a combination of deep shotgun metagenomics complemented by a supervised machine learning approach to compare the metabolic strategies of tropical soil microbial communities in pristine forests and long‐term established pastures in the Amazon. Machine learning‐derived metagenome analysis indicated that microbial community structures (bacteria, archaea and viruses) and the composition of protein‐coding genes were distinct in each plant cover type environment. Forest and pasture soils had different genomic diversities for the above three taxonomic groups, characterised by their protein‐coding genes. These differences in metagenome profiles in soils under forests and pastures suggest that metabolic strategies related to carbohydrate and energy metabolisms were altered at community level. Changes were also consistent with known modifications to the C and N cycles caused by long‐term shifts in aboveground vegetation and were also associated with several soil physicochemical properties known to change with land use, such as the C/N ratio, soil temperature and exchangeable acidity. In addition, our analysis reveals that these alterations in land use can also result in changes to the composition and diversity of the soil DNA virome. Collectively, our study indicates that soil microbial communities shift their overall metabolic strategies, driven by genomic alterations observed in pristine forests and long‐term established pastures with implications for the C and N cycles.

carbon and nitrogen cycles↗

Predictive Modeling of Local Film-Cooling Flow on a Turbine Rotor Blade

Abstract In the turbine section of a modern gas turbine engine, components exposed to the main gas path flow rely on cooling air to maintain hardware durability targets. Therefore, monitoring turbine cooling flow is essential to the diagnostic and prognostic efficacy of a condition-based operation and maintenance (CBOM) approach. This study supports CBOM goals by leveraging supervised machine learning to estimate relative changes to local film-cooling flowrate using surface temperature measured on the pressure side of a rotating turbine blade operating at engine-relevant aerothermal conditions. Throughout the lifetime of a film-cooled turbine component, characteristics of the film-cooling flow—such as film trajectory and cooling effectiveness—vary as degradation-driven geometry distortions occur, which ultimately affects the relationship between the model input and the model output—film-cooling flowrate predictions. The present study addresses this complication by testing a data-driven model on multiple turbine blades of the same nominal design, but with each blade exhibiting different localized film-cooling flow characteristics. By testing the model in this manner, strategies for mitigating the detrimental effects of film-cooling flow characteristic variations on model performance were investigated, and the corresponding flowrate prediction accuracy was quantified.

Engineering↗

Characterizing the Spread of COVID-19 from Human Mobility Patterns and SocioDemographic Indicators

Mobility is an indicator of human movement through space and time. With the increasing availability of geolocated data (from GPS, accelerometers, etc.), it is now possible to examine individual as well as group human mobility patterns. Human mobility is influenced by both intrinsic (i.e. personal motivations) and extrinsic (i.e., events like natural hazards or a pandemic like the COVID-19) factors. However, the intricate relationships between human mobility patterns and sociodemographic characteristics in the context of a pandemic are yet to be fully explored. Our goal is to overcome this gap by using human mobility data at the census block group level from mobile phones and combining those with social vulnerability indicators to examine the overall spread of COVID-19 at local spatial scales. We used 585,878 weekly visits to 37,871 points of interests (POIs) from Safegraph to quantify mobility indices and social distancing metrics in 2,820 census block groups in the city of Los Angeles (LA) - before and during lockdown as well as during the phase1 and phase 2 reopening. Finally, using supervised machine learning algorithms, we classified the census block groups in LA into High, Medium and Low categories that represented the vulnerability of these block groups based on the cumulative number of occurrences of COVID-19 cases till July 24, 2020. Our results indicate that the tree-based classifiers performed well in comparison to the Support Vector Machines and Multinomial Logit models. Gradient Boosting had the highest classification accuracy of 97.4% COVID-19 with an AUC score of 0.987. The block groups with high COVID-19 cases also had a high concentration of socially vulnerable populations, high human mobility index and a low social distancing index.

Roy, Avipsa↗

WiiBin

WiiBin is a framework to determine architecture of an unknown binary and locate opcode sections within the same binary via supervised machine learning.

Beckman, BryanR↗

Avatar Tools

Supervised machine learning is the process of using past experience to predict the future. "Ensembles" are a machine-learning meta-method that can be applied to most machine learning algorithms. Ensembles generally greatly improve accuracy, reduce or remove most of the design issues presented by machine learning, and are admirably suited to parallel and distributed computation. The Avatar Tools codes are an implementation of ensembles specifically for decision trees. Some features that distinguish Avatar Tools from other "ensembles for decision trees" codes are: (1) Does the bookkeeping necessary for out of bag (OOB) validation. (2) Can use OOB validation to automatically determine optimal ensemble size. (3) Provides an MPI-based parallel implementation, for distributed operation. (4) Provides convenient tools for cross-validation, to assess the accuracy provided by a training set. SAND2020-3858 M Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Siefert, Christopher↗

DNA Sequencing Optimizer v0.1

This is a software utility used to maximize next-generation sequencing data output without compromising quality values. The DNA Sequencing Optimizer is used to determine the optimal amount of library to load. To do this a supervised machine learning model predicts cluster density, a measure of how much DNA is bound to the sequence, from the concentration of DNA loaded. The output is a plot that is used by experimentalists to determine the best DNA library concentration to load onto the sequencer.

Costello, Zachary↗

Annotated Translated Disassembled Code

Procedure for generating function/library embedding based on disassembled binary data from a graph database and processing those embedding through supervised machine learning techniques.

Beckman, BryanR.↗

RGM: Random Geological Model Generation Package

This Fortran code is to accompany a manuscript to be submitted to Computers & Geosciences, a high-impact, peer-reviewed journal in computer methods for geosciences research. This Fortran code focuses on generation of synthetic geological models using a multi-randomization strategy. Generating high-fidelity synthetic geological models, including realistic seismic reflector migration images, faults, salt bodies, and relative geological time images, is the key for many supervised machine learning methods that aim to delineate faults and other geological properties of interest from seismic migration images. Our package contains two major functionalities: generating 2D synthetic random geological models and generating 3D synthetic random geological models. In each step of the generation process, we set random values for key properties of a geological model to improve the fidelity of the resulting geological model. The package also includes example codes on how to use the random geological model generation subroutines. We name this package RGM – Random Geological Model generation package.

Gao, Kai↗

ldrd_virus_work

This is a Python code base that takes openly-available genetic information on known viruses and performs supervised machine learning and feature importance analysis on the relationship of the viral genomes to the competence to infect humans or bind to a specific host cell receptor.

Reddy, Tyler [LANL]↗

Spectral Data Fusion From Handheld Laser-Induced Breakdown Spectroscopy (LIBS) and X-ray Fluorescence (XRF) Analyzers for Improved Detection of Cerium in a Simulated Dispersal Accident

Here, this work implements a mid-level data fusion methodology on spectral data from handheld X-ray fluorescence and laser-induced breakdown spectroscopy analyzers to quantify plutonium surrogate (CeO 2 ) contamination in soil samples for the first time. Spectral data from each analyzer were used independently to train supervised machine learning regressions to predict Ce concentration. Fused features from both data sets were then used to train the same models, comparing prediction performance by evaluating model precision and sensitivity. Fusing principal component scores from the two sensors yielded an order of magnitude improvement in precision and sensitivity of predictions made with an artificial neural network, compared to predictions made by models trained on independent sensor data. As a result, a boosted ensemble trained on the fused spectral features yielded an ideal predictor with root-mean-squared error on the order of 10 –6 and calculated limit of detection order 10 –5 wt %.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A Mixed-Method Design Approach for Empirically Based Selection of Unbiased Data Annotators

Implicit bias embedded in the annotated data is by far the greatest impediment in the effectual use of supervised machine learning models in tasks involving race, ethics, and geopolitical polarization. For societal good and demonstrable positive impact on wider society, it is paramount to carefully select data annotators and rigorously validate the annotation process. Current approaches to selecting annotators are not sufficiently grounded in scientific principles and are limited at the policy-guidance level, thereby rendering them unusable for machine learning practitioners. This work proposes a new approach based on the mixed-methods design that is functional, adaptable, and simpler to implement in selecting unbiased annotators for any machine learning problem. By demonstrating it on a real-world geopolitical problem, we also identified and ranked key inane profile characteristics towards an empirically-based selection of unbiased data annotators.

Thakur, Gautam Malviya↗

Automated Threat Recognition For Aviation Security Applications

We have developed a framework for automated threat recognition (ATR) of explosive threat materials for both single-energy and dual-energy X-ray CT systems. Under this framework, two different types of ATR have been developed. The first type of ATR employs supervised machine learning with statistical characterization of target materials for threat identification training. The reliance only on statistical characterization information for threat training uniquely enables this style of ATR to adapt quickly to evolving threats. The second type of ATR employs deep learning through convolutional neural networks. Convolutional neural networks are attractive due to their human-like capacity for learning and strong ability to identify trends and patterns. Although this method is more powerful than the first type, it requires large amounts of training data and therefore is less agile. Each ATR performs threat characterization at the voxel neighborhood level. This approach avoids the use of threat shape as a detection criterion, which is prohibited by DHS and TSA guidelines.

97 MATHEMATICS AND COMPUTING↗