Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “supervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Fault Detection on Seismic Structural Images Using a Nested Residual U-Net

Automatic identification of faults on seismic structural images is a challenging yet crucial task in quantitative seismic interpretation. Human picking or attribute-based fault detection methods may misidentify faults on noisy, complex seismic images. In this work, we develop a new automatic fault detection method using a nested residual U-shaped convolutional neural network. Each of the encoders and decoders in this neural network is a residual U-Net, leading to a nested architecture. The final fault map results from the fusion of three fault maps with low, medium, and high fault resolutions. We demonstrate the excellent fault-detection capability of our nested neural network using a series of synthetic and field seismic images. We find that our approach produces clearer and more interpretable fault maps than the current state-of-the-art U-Net fault detection method, particularly on noisy seismic images. Our new automatic fault detection method can facilitate reliable quantitative seismic interpretation on field seismic images.

58 GEOSCIENCES↗

Selecting XFEL single-particle snapshots by geometric machine learning

A promising new route for structural biology is single-particle imaging with an X-ray Free-Electron Laser (XFEL). This method has the advantage that the samples do not require crystallization and can be examined at room temperature. However, high-resolution structures can only be obtained from a sufficiently large number of diffraction patterns of individual molecules, so-called single particles. Here, we present a method that allows for efficient identification of single particles in very large XFEL datasets, operates at low signal levels, and is tolerant to background. This method uses supervised Geometric Machine Learning (GML) to extract low-dimensional feature vectors from a training dataset, fuse test datasets into the feature space of training datasets, and separate the data into binary distributions of “single particles” and “non-single particles.” As a proof of principle, we tested simulated and experimental datasets of the Coliphage PR772 virus. We created a training dataset and classified three types of test datasets: First, a noise-free simulated test dataset, which gave near perfect separation. Second, simulated test datasets that were modified to reflect different levels of photon counts and background noise. These modified datasets were used to quantify the predictive limits of our approach. Third, an experimental dataset collected at the Stanford Linear Accelerator Center. The single-particle identification for this experimental dataset was compared with previously published results and it was found that GML covers a wide photon-count range, outperforming other single-particle identification methods. Moreover, a major advantage of GML is its ability to retrieve single particles in the presence of structural variability.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Predicting industrial building energy consumption with statistical and machine-learning models informed by physical system parameters

The industrial sector consumes about one-third of global energy, making them a frequent target for energy use reduction. Variation in energy usage is observed with weather conditions, as space conditioning needs to change seasonally, and with production, energy-using equipment is directly tied to production rate. Previous models were based on engineering analyses of equipment and relied on site-specific details. Others consisted of single-variable regressors that did not capture all contributions to energy consumption. Further, new modeling techniques could be applied to rectify these weaknesses. Applying data from 45 different manufacturing plants obtained from industrial energy audits, a supervised machine-learning model is developed to create a general predictor for industrial building energy consumption. The model uses features of air enthalpy, solar radiation, and wind speed to predict weather-dependency; motor, steam, and compressed air system parameters to capture support equipment contributions; and operating schedule, production rate, number of employees, and floor area to determine production-dependency. Results showed that a model that used a linear regressor over a transformed feature space could outperform a support vector machine and utilize features more representative of physical systems. Using informed parameters to build a reliable predictor will more accurately characterize a manufacturing facility's energy savings opportunities.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Combining synchrotron X-ray diffraction, mechanistic modeling and machine learning for in situ subsurface temperature quantification during laser melting

Laser melting, such as that encountered during additive manufacturing, produces extreme gradients of temperature in both space and time, which in turn influence microstructural development in the material. Qualification and model validation of the process itself and the resulting material necessitate the ability to characterize these temperature fields. However, well established means to directly probe the material temperature below the surface of an alloy while it is being processed are limited. To address this gap in characterization capabilities, a novel means is presented to extract subsurface temperature-distribution metrics, with uncertainty, from in situ synchrotron X-ray diffraction measurements to provide quantitative temperature evolution data during laser melting. Temperature-distribution metrics are determined using Gaussian process regression supervised machine-learning surrogate models trained with a combination of mechanistic modeling (heat transfer and fluid flow) and X-ray diffraction simulation. The trained surrogate model uncertainties are found to range from 5 to 15% depending on the metric and current temperature. The surrogate models are then applied to experimental data to extract temperature metrics from an Inconel 625 nickel superalloy wall specimen during laser melting. The maximum temperatures of the solid phase in the diffraction volume through melting and cooling are found to reach the solidus temperature as expected, with the mean and minimum temperatures found to be several hundred degrees less. The extracted temperature metrics near melting are determined to be more accurate because of the lower relative levels of mechanical elastic strains. However, uncertainties for temperature metrics during cooling are increased due to the effects of thermomechanical stress.

36 MATERIALS SCIENCE↗

Performance Comparison of Clipping Detection Techniques in AC Power Time Series

In this research, a variety of methods were developed to detect clipping periods in AC power time series. AC power data streams associated with 36 unique systems across the United States were collected, and data points representing clipping periods were manually labeled by experts. Using this data set for training and validation, novel logic-based and machine learning (ML) approaches were developed to classify time series values as clipping or non-clipping. These approaches were compared to the RdTools method for detecting clipping periods. The logic-based and ML XGBoost approaches achieved F-scores of 85.0 and 77.6, respectively, when cross-validated against the manually labeled data, as compared to the current RdTools approach (F-score of 56.4), indicating a significant improvement at detecting clipping periods. Additionally, the effects of each clipping filter when evaluating system degradation rates were assessed, using 31 unique systems across the United States. Results indicate that estimated system degradation rate can vary based on the type of clipping filter used, by up to 0.6% degradation rate for some cases.

clipping↗

Machine learning and deep learning for mineralogy interpretation and CO 2 saturation estimation in geological carbon Storage: A case study in the Illinois Basin

Carbon capture and storage (CCS) is a promising approach to simultaneously maintaining energy security and reducing carbon dioxide (CO 2 ) emissions under the current energy portfolio that is dominated by fossil fuel energy. Pre-injection formation characterization and post-injection CO 2 monitoring are two critical tasks to guarantee storage efficiency in CCS. The CCS projects in the Illinois Basin, the first large-scale CO 2 injection into saline aquifers in the United States, employed conventional and the latest pulsed neutron logging (PNL) tools for mineralogy interpretation and CO 2 saturation estimation, which provide valuable references for future CCS projects. Because of the inherent fuzziness of petrophysical measurements and complex subsurface heterogeneity, interpreting well-logging data is time-consuming, and its accuracy can be user-biased. In recent years, data-driven methods have been widely used to capture the non-linear patterns between input features and interpretation results. This work applied and evaluated four commonly used machine learning (ML) models, including ridge regression (RR), random forest (RF), gradient boosting regression (GBR), support vector regression (SVR), and one deep learning (DL) model, the artificial neural network (ANN). We optimized the hyperparameters of the four ML models and the DL model using the simulated annealing algorithm and the grid search strategy, respectively. The input features of the mineralogy interpretation models were eleven conventional well-logging parameters, and the label data (i.e., ground truth) were the porosity and volumetric fractions of six minerals, including quartz, feldspar, dolomite, calcite, clay, and iron minerals. The results demonstrated that the GBR and RF models were superior in predicting volumetric fractions of minerals and porosity; label data with low coefficient of variation (CV) values tended to yield better performance. For CO 2 saturation estimation, the RF was the best-performing model, followed by SVR, ANN, GBR, and RR. Furthermore, we conducted feature importance ranking using the permutation importance algorithm and found that the formation sigma and well pressure were the most important features in this study. In conclusion, the study of CCS projects in the Illinois Basin bridges the gap between the limited knowledge and understanding of geological carbon storage and the increasing demand for reliable, cost-effective, and sustainable energy solutions.

58 GEOSCIENCES↗

State Predictor of Classification Cognitive Engine Applied to Channel Fading

This study presents the application of machine learning (ML) to a space-to-ground communication link, showing how ML can be used to detect the presence of detrimental channel fading. Using this channel state information, the communication link can be used more efficiently by reducing the amount of lost data during fading. The motivation for this work is based on channel fading observed during on-orbit operations with NASA's Space Communication and Navigation (SCaN) testbed on the International Space Station (ISS). This paper presents the process to extract a target concept (fading and not-fading) from the raw data. The pre-processing and data exploration effort is explained in detail, with a list of assumptions made for parsing and labelling the dataset. The model selection process is explained, specifically emphasizing the benefits of using an ensemble of algorithms with majority voting for binary classification of the channel state. Experimental results are shown, highlighting how an end-to-end communication system can utilize knowledge of the channel fading status to identity fading and take appropriate action. With a laboratory testbed to emulate channel fading, the overall performance is compared to standard adaptive methods without fading knowledge, such as adaptive coding and modulation.

Fading↗

Physics-guided dual implicit neural representations for source separation

Significant challenges exist in efficient data analysis of most advanced experimental and observational techniques because the collected signals often include unwanted contributions, such as background and signal distortions, that can obscure the physically relevant information of interest. To address this, we have developed a self-supervised machine-learning approach for source separation using a dual implicit neural representation framework that jointly trains two neural networks: one for approximating distortions of the physical signal of interest and the other for learning the effective background contribution. Our method learns directly from the raw data by minimizing a reconstruction-based loss function without requiring labeled data or pre-defined dictionaries. We demonstrate the effectiveness of our framework by considering a challenging case study involving large-scale simulated, as well as experimental, momentum-energy-dependent inelastic neutron scattering data in a four-dimensional parameter space, characterized by heterogeneous background contributions and unknown distortions to the target signal. The method is found to successfully separate physically meaningful signals from a complex or structured background even when the signal characteristics vary across all four dimensions of the parameter space. An analytical approach that informs the choice of the regularization parameter is presented. Our method offers a versatile framework for addressing source separation problems across diverse domains, ranging from superimposed signals in astronomical measurements to structural features in biomedical image reconstructions.

47 OTHER INSTRUMENTATION↗

ProvSec: Open Cybersecurity System Provenance Analysis Benchmark Dataset with Labels

System provenance forensic analysis has been studied by a large body of research work. This area needs fine granularity data such as system calls along with event fields to track the dependencies of events. While prior work on security datasets has been proposed, we found a useful dataset of realistic attacks and details that are needed for high-quality provenance tracking is lacking. We created a new dataset of eleven vulnerable cases for system forensic analysis. It includes the full details of system calls including syscall parameters. Realistic attack scenarios with real software vulnerabilities and exploits are used. For each case, we created two sets of benign and adversary scenarios which are manually labeled for supervised machine-learning analysis. In addition, we present an algorithm to improve the data quality in the system provenance forensic analysis. We demonstrate the details of the dataset events and dependency analysis of our dataset cases.

97 MATHEMATICS AND COMPUTING↗

A long short-term memory embedding for hybrid uplifted reduced order models

In this paper, we introduce an uplifted reduced order modeling (UROM) approach through the integration of standard projection based methods with long short-term memory (LSTM) embedding. Our approach has three modeling layers or components. In the first layer, we utilize an intrusive projection approach to model dynamics represented by the largest modes. The second layer consists of an LSTM model to account for residuals beyond this truncation. This closure layer refers to the process of including the residual effect of the discarded modes into the dynamics of the largest scales. However, the feasibility of generating a low rank approximation tails off for higher Kolmogorov n -width systems due to the underlying nonlinear processes. The third uplifting layer, called super-resolution, addresses this limited representation issue by expanding the span into a larger number of modes utilizing the versatility of LSTM. Therefore, our model integrates a physics-based projection model with a memory embedded LSTM closure and an LSTM based super-resolution model. In several applications, we exploit the use of Grassmann manifold to construct UROM for unseen conditions. We performed numerical experiments by using the Burgers and Navier-Stokes equations with quadratic nonlinearity. Finally, our results show robustness of the proposed approach in building reduced order models for parameterized systems and confirm the improved trade-off between accuracy and efficiency.

42 ENGINEERING↗

Single-cell chromatin accessibility and cis -regulatory element analyses in plants using the scPlantReg platform

Understanding gene regulation is fundamental to plant improvement, but the lack of plant-specific single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) frameworks and cross-species databases has limited insights into cell-type-specific cellular regulation. Here we present ‘scPlantReg’, an integrated framework and database for plant scATAC-seq data. scPlantReg supports end-to-end analyses from raw data processing to biological interpretation and features ‘scATACtor’, a supervised machine-learning approach that outperforms existing tools for cell-type annotation. We applied scPlantReg to pearl millet to characterize cell-type-specific chromatin accessibility and identify validated activating and repressing accessible chromatin regions (ACRs), revealing WRKY transcription factors as potential regulators of xylem development. Furthermore, we reanalysed scATAC-seq datasets from 8 plant species, spanning 11 tissues and multiple developmental stages, enabling cross-species comparisons. Furthermore, these analyses uncovered conserved regulatory programmes, including AP2/EREBP-associated ACRs linked to cell wall development and cell-type-conserved TFs across grasses. Collectively, scPlantReg provides a general framework and resource for comparative regulatory analysis in plants.

Epigenomics↗

Supervised enhancer prediction with epigenetic pattern recognition and targeted validation

Enhancers are important non-coding elements, but they have traditionally been hard to characterize experimentally. The development of massively parallel assays allows the characterization of large numbers of enhancers for the first time. Here, we developed a framework using Drosophila STARR-seq to create shape-matching filters based on meta-profiles of epigenetic features. We integrated these features with supervised machine-learning algorithms to predict enhancers. We further demonstrated that our model could be transferred to predict enhancers in mammals. We comprehensively validated the predictions using a combination of in vivo and in vitro approaches, involving transgenic assays in mice and transduction-based reporter assays in human cell lines (153 enhancers in total). The results confirmed that our model can accurately predict enhancers in different species without re-parameterization. Finally, we examined the transcription factor binding patterns at predicted enhancers versus promoters. Here, we demonstrated that these patterns enable the construction of a secondary model that effectively distinguishes enhancers and promoters.

59 BASIC BIOLOGICAL SCIENCES↗

Improving neutrino oscillation measurements through event classification

Precise neutrino energy reconstruction is essential for next-generation long-baseline oscillation experiments, yet current methods remain limited by large uncertainties in neutrino-nucleus interaction modeling. Even so, it is well established that different interaction channels produce systematically varying amounts of missing energy and therefore yield different reconstruction performance–information that standard calorimetric approaches do not exploit. We introduce a strategy that incorporates this structure by classifying events according to their underlying interaction type prior to energy reconstruction. Using supervised machine-learning techniques trained on labeled generator events, we leverage intrinsic kinematic differences among quasielastic scattering, meson-exchange current, resonance production, and deep-inelastic scattering processes. A cross-generator testing framework demonstrates that this classification approach is robust to microphysics mismodeling and, when applied to a simulated DUNE 𝜈 𝜇 disappearance analysis, yields improved accuracy and sensitivity at the 10%–20% level. These results highlight a practical path toward reducing reconstruction-driven systematics in future oscillation measurements.

Ellis, Sebastian A. R. [King's College, London (Un↗

Fundamental Parameters of ∼30,000 M dwarfs in LAMOST DR1 Using Data-driven Spectral Modeling

M dwarfs are the most common type of star in the Galaxy, and because of their small size are favored targets for searches of Earth-sized transiting exoplanets. Current and upcoming all-sky spectroscopic surveys, such as the Large Sky Area Multi Fiber Spectroscopic Telescope (LAMOST), offer an opportunity to systematically determine physical properties of many more M dwarfs than has been previously possible. Here, we present new effective temperatures, radii, masses, and luminosities for 29,678 M dwarfs with spectral types M0–M6 in the first data release (DR1) of LAMOST. We derived these parameters from the supervised machine-learning code, The Cannon, trained with 1388 M dwarfs in the Transiting Exoplanet Survey Satellite Cool Dwarf Catalog that were also present in LAMOST with high signal-to-noise ratio (>250) spectra. Our validation tests show that the output parameter uncertainties are strongly correlated with the signal-to-noise of the LAMOST spectra, and we achieve typical uncertainties of 110 K in T{sub eff} (∼3%), 0.065 R{sub ⊙} (∼14%) in radius, 0.054 M{sub ⊙} (∼12%) in mass, and 0.012 L{sub ⊙} (∼20%) in luminosity. The model presented here can be rapidly applied to future LAMOST data releases, significantly extending the samples of well-characterized M dwarfs across the sky using new and exclusively data-based modeling methods.

79 ASTRONOMY AND ASTROPHYSICS↗

Information Content of JWST NIRSpec Transmission Spectra of Warm Neptunes

Warm Neptunes offer a rich opportunity for understanding exo-atmospheric chemistry. With the upcoming James Webb Space Telescope (JWST), there is a need to elucidate the balance between investments in telescope time versus scientific yield. We use the supervised machine-learning method of the random forest to perform an information content (IC) analysis on a 11-parameter model of transmission spectra from the various NIRSpec modes. The three bluest medium-resolution NIRSpec modes (0.7–1.27 μm, 0.97–1.84 μm, 1.66–3.07 μm) are insensitive to the presence of CO. The reddest medium-resolution mode (2.87–5.10 μm) is sensitive to all of the molecules assumed in our model: CO, CO{sub 2}, CH{sub 4}, C{sub 2}H{sub 2}, H{sub 2}O, HCN, and NH{sub 3}. It competes effectively with the three bluest modes on the information encoded on cloud abundance and particle size. It is also competitive with the low-resolution prism mode (0.6–5.3 μm) on the inference of every parameter except for the temperature and ammonia abundance. We recommend astronomers to use the reddest medium-resolution NIRSpec mode for studying the atmospheric chemistry of 800–1200 K warm Neptunes; its corresponding high-resolution counterpart offers diminishing returns. We compare our findings to previous JWST IC analyses that favor the blue orders and suggest that the reliance on chemical equilibrium could lead to biased outcomes if this assumption does not apply. A simple, pressure-independent diagnostic for identifying chemical disequilibrium is proposed based on measuring the abundances of H{sub 2}O, CO, and CO{sub 2}.

79 ASTRONOMY AND ASTROPHYSICS↗

The Circular Velocity Curve of the Milky Way from 5–25 kpc Using Luminous Red Giant Branch Stars

We present a sample of 254,882 luminous red giant branch (LRGB) stars selected from the APOGEE and LAMOST surveys. By combining photometric and astrometric information from the Two Micron All Sky Survey and Gaia survey, the precise distances of the sample stars are determined by a supervised machine-learning algorithm: the gradient-boosted decision trees. To test the accuracy of the derived distances, member stars of globular clusters (GCs) and open clusters are used. The tests by cluster member stars show a precision of about 10% with negligible zero-point offsets, for the derived distances of our sample stars. The final sample covers a large volume of the Galactic disk(s) and halo of 0 < R < 30 kpc and |Z| ≤ 15 kpc. The rotation curve (RC) of the Milky Way across the radius of 5 ≲ R ≲ 25 kpc has been accurately measured with ~54,000 stars of the thin disk population selected from the LRGB sample. The derived RC shows a weak decline along R with a gradient of -1.83 ± 0.02 (stat.) ± 0.07 (sys.) km s -1 kpc -1 , in excellent agreement with the results measured by previous studies. The circular velocity at the solar position, yielded by our RC is 234.04 ± 0.08 (stat.) ± 1.36 (sys.) km s -1 , again in great consistency with other independent determinations. From the newly constructed RC, as well as constraints from other data, we have constructed a mass model for our Galaxy, yielding a mass of the dark matter halo of M 200 = (8.05 ± 1.15) × 10 11 M ⊙ with a corresponding radius of R 200 = 192.37 ± 9.24 kpc and a local dark matter density of 0.39 ± 0.03 GeV cm -3 .

79 ASTRONOMY AND ASTROPHYSICS↗