Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Unsupervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Machine learning uncovers independently regulated modules in the Bacillus subtilis transcriptome

The transcriptional regulatory network (TRN) of Bacillus subtilis coordinates cellular functions of fundamental interest, including metabolism, biofilm formation, and sporulation. Here, we use unsupervised machine learning to modularize the transcriptome and quantitatively describe regulatory activity under diverse conditions, creating an unbiased summary of gene expression. We obtain 83 independently modulated gene sets that explain most of the variance in expression and demonstrate that 76% of them represent the effects of known regulators. The TRN structure and its condition-dependent activity uncover putative or recently discovered roles for at least five regulons, such as a relationship between histidine utilization and quorum sensing. The TRN also facilitates quantification of population-level sporulation states. As this TRN covers the majority of the transcriptome and concisely characterizes the global expression state, it could inform research on nearly every aspect of transcriptional regulation in B. subtilis.

59 BASIC BIOLOGICAL SCIENCES↗

Objective Phenotyping of Root System Architecture Using Image Augmentation and Machine Learning in Alfalfa (Medicago sativa L.)

Active breeding programs specifically for root system architecture (RSA) phenotypes remain rare; however, breeding for branch and taproot types in the perennial crop alfalfa is ongoing. Phenotyping in this and other crops for active RSA breeding has mostly used visual scoring of specific traits or subjective classification into different root types. While image-based methods have been developed, translation to applied breeding is limited. This research is aimed at developing and comparing image-based RSA phenotyping methods using machine and deep learning algorithms for objective classification of 617 root images from mature alfalfa plants collected from the field to support the ongoing breeding efforts. Our results show that unsupervised machine learning tends to incorrectly classify roots into a normal distribution with most lines predicted as the intermediate root type. Encouragingly, random forest and TensorFlow-based neural networks can classify the root types into branch-type, taproot-type, and an intermediate taproot-branch type with 86% accuracy. With image augmentation, the prediction accuracy was improved to 97%. Coupling the predicted root type with its prediction probability will give breeders a confidence level for better decisions to advance the best and exclude the worst lines from their breeding program. This machine and deep learning approach enables accurate classification of the RSA phenotypes for genomic breeding of climate-resilient alfalfa.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning technique to identify grains in polycrystalline materials samples

A method of identifying grains in polycrystalline materials, the method including (a) identifying local crystal structure of the polycrystalline material based on neighbor coordination or pattern recognition machine learning, the local crystal structure including grains and grain boundaries, (b) pre-processing the grains and the grain boundaries using image processing techniques, (c) conducting grain identification using unsupervised machine learning; and (d) refining a resolution of the grain boundaries.

Sankaranarayanan, Subramanian↗

General-Purpose Unsupervised Cyber Anomaly Detection via Non-Negative Tensor Factorization

Distinguishing malicious anomalous activities from unusual but benign activities is a fundamental challenge for cyber defenders. Prior studies have shown that statistical user behavior analysis yields accurate detections by learning behavior profiles from observed user activity. These unsupervised models are able to generalize to unseen types of attacks by detecting deviations from normal behavior, without knowledge of specific attack signatures. However, approaches proposed to date based on probabilistic matrix factorization are limited by the information conveyed in a two-dimensional space. Non-negative tensor factorization, on the other hand, is a powerful unsupervised machine learning method that naturally models multi-dimensional data, capturing complex and multi-faceted details of behavior profiles. Herein, our new unsupervised statistical anomaly detection methodology matches or surpasses state-of-the-art supervised learning baselines across several challenging and diverse cyber application areas, including detection of compromised user credentials, botnets, spam e-mails, and fraudulent credit card transactions.

97 MATHEMATICS AND COMPUTING↗

Defect detection in atomic-resolution images via unsupervised learning with translational invariance

Abstract Crystallographic defects can now be routinely imaged at atomic resolution with aberration-corrected scanning transmission electron microscopy (STEM) at high speed, with the potential for vast volumes of data to be acquired in relatively short times or through autonomous experiments that can continue over very long periods. Automatic detection and classification of defects in the STEM images are needed in order to handle the data in an efficient way. However, like many other tasks related to object detection and identification in artificial intelligence, it is challenging to detect and identify defects from STEM images. Furthermore, it is difficult to deal with crystal structures that have many atoms and low symmetries. Previous methods used for defect detection and classification were based on supervised learning, which requires human-labeled data. In this work, we develop an approach for defect detection with unsupervised machine learning based on a one-class support vector machine (OCSVM). We introduce two schemes of image segmentation and data preprocessing, both of which involve taking the Patterson function of each segment as inputs. We demonstrate that this method can be applied to various defects, such as point and line defects in 2D materials and twin boundaries in 3D nanocrystals.

36 MATERIALS SCIENCE↗

Identifying Different Classes of Seismic Noise Signals Using Unsupervised Learning

Abstract Proper classification of nontectonic seismic signals is critical for detecting microearthquakes and developing an improved understanding of ongoing weak ground motions. We use unsupervised machine learning to label five classes of nonstationary seismic noise common in continuous waveforms. Temporal and spectral features describing the data are clustered to identify separable types of emergent and impulsive waveforms. The trained clustering model is used to classify every 1 s of continuous seismic records from a dense seismic array with 10–30 m station spacing. We show that dominate noise signals can be highly localized and vary on length scales of hundreds of meters. The methodology demonstrates the complexity of weak ground motions and improves the standard of analyzing seismic waveforms with a low signal‐to‐noise ratio. Application of this technique will improve the ability to detect genuine microseismic events in noisy environments where seismic sensors record earthquake‐like signals originating from nontectonic sources.

Johnson, Christopher W.↗

Unraveling Hydrogen Induced Geochemical Reaction Mechanisms through Coupled Geochemical Modeling and Machine Learning

Underground hydrogen storage (UHS) provides a promising large-scale, long-term energy storage solution. A reasonable recovery of stored hydrogen is critical for a successful storage scheme. However, in subsurface reservoirs hydrogen is subject to active geochemical reactions that might result in hydrogen loss. In this study, we implemented a geochemical modeling approach coupled with an unsupervised machine learning technique called non-negative matrix factorization (NMF) to unravel the complex brine-rock-H 2 geochemical processes responsible for hydrogen losses, with particular focus on sulfate reduction reactions. NMF is applied to modeled mineral evolution and fluid component profiles to retrieve profiles that can be interpreted to more easily assess competing processes. NMF decouples simulated competing equilibrium reactions. This facilitates separation of overlapping reaction profiles from redox processes, dissolution fronts, and secondary precipitation while considering the effects of simulation parameters such as salinity, temperature, and total H 2 pressure. NMF successfully discriminates these competing effects in nonlinear ways, allowing robust interpretation. In addition, NMF reveals subtle coupled mineral associations and reaction fronts that are invisible to conventional model analysis. This integrated approach strengthens the conceptual understanding of complex nonlinear hydrogen-brine-rock interactions and advances geochemical research on UHS systems to resolve complexities in modeled geochemical systems without the need for direct experiments or prior knowledge. Furthermore, this study highlights the efficacy of combining geochemical modeling with machine learning techniques to enhance the interpretability of the intricate geochemical simulation output through deciphering the overlapping reaction path that cannot be achieved only using conventional analysis of geochemical models alone.

08 HYDROGEN↗

Spread spectrum time domain reflectometry (SSTDR) and frequency domain reflectometry (FDR) cable inspection using machine learning

Cables are initially qualified for nuclear power plant use for 40 years. As plants extend their operating license to 60 and 80 years, justification for continued cable use must shift to a condition-based approach since it is cost prohibitive to completely replace cables that are likely still capable of performing their design function. The Pacific Northwest National Laboratory (PNNL) Accelerated and Real Time Experimental Nodal Analysis (ARENA) cable motor test bed was used to test the response of a commercial spread spectrum time domain reflectometry (SSTDR) system, a laboratory instrument software-controlled SSTDR, and a vector network analyzer-based frequency domain reflectometry (FDR) system to various cable anomalies. The three instrument systems were able to interrogate cables over a range of frequency bandwidths that can be helpful for human data analysis. Data were subjected to supervised and unsupervised machine learning (ML) analyses to distinguish normal undamaged cable responses from anomalous cable responses. Both supervised and unsupervised ML approaches produced encouraging results with an undamaged/anomalous prediction accuracy from 0.69% to 0.87%. Recommendations for further development and field implementation include increased and more balanced sample sets particularly including more training data.

SSTDR, FDR, Reflectometry, Machine Learning, ARENA↗

ACAT

A Physics-Informed Machine Learning (PIML) framework for better system vulnerability assessment and faster corrective action recommendation. This framework consists of deriving physics-informed priors and smart sampling algorithms to help reduce the data samples of grid models and the dimension of simulation outputs, yielding small yet representative subset of the complex system, and both supervised and unsupervised machine learning (ML) algorithms for designing corrective actions

Chen, Yousu↗

Comparison of Supervised and Un-Supervised Machine Learning Algorithms for Threat Detection and Scintillator Performance for Radiation Portal Monitoring

Following the events of September 11, 2001, international border crossing have been equipped with radiation portal monitors (RPMs) to identify illicit radioactive material. Polyvinyl toluene (PVT) scintillators are commonly used due to their low cost and reasonable maintainability, however they offer low spectral resolution. Despite the fact that over twenty years has transpired since this event, radioisotopes are still typically identified by hand-crafted classification algorithms, e.g., total counts or energy windowing, and exhibit relatively poor performance in detecting threats at the low false alarm rates required to support the stream of commerce. While some improvement to performance has been realized via the use of supervised machine learning, these classification algorithms typically utilize simulations in lieu of real data due to the sparsity of data for one or more classes. Accordingly, the performance of these algorithms is somewhat less than optimal when examining experiments or simulations with model mismatch. Consequently, in this work, we examine the application of a number of unsupervised machine learning, anomaly detection based algorithms, to circumvent the inverse crime when analyzing spectroscopy data for RPMs. We also compare anomaly detection results with those obtained via the use of supervised classification detection ML algorithms when model mismatch is introduced between the simulated threat items utilized for training/testing. Finally, we compared the performance of the PVT scintillators to those obtained with higher resolution detectors using both anomaly detection and supervised classification algorithms.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]↗

The Deeper, Wider, Faster programme: exploring stellar flare activity with deep, fast cadenced DECam imaging via machine learning

ABSTRACT We present our 500 pc distance-limited study of stellar flares using the Dark Energy Camera as part of the Deeper, Wider, Faster programme. The data were collected via continuous 20-s cadence g-band imaging and we identify 19 914 sources with precise distances from Gaia DR2 within 12, ∼3 deg2, fields over a range of Galactic latitudes. An average of ∼74 min is spent on each field per visit. All light curves were accessed through a novel unsupervised machine learning techniques designed for anomaly detection. We identify 96 flare events occurring across 80 stars, the majority of which are M dwarfs. Integrated flare energies range from ∼1031–1037 erg, with a proportional relationship existing between increased flare energy with increased distance from the Galactic plane, representative of stellar age leading to declining yet more energetic flare events. In agreement with previous studies we observe an increase in flaring fraction from M0 to M6 spectral types. Furthermore, we find a decrease in the flaring fraction of stars as vertical distance from the galactic plane is increased, with a steep decline present around ∼100 pc. We find that $\sim 70{{\ \rm per\ cent}}$ of identified flares occur on short time-scales of <8 min. Finally, we present our associated flare rates, finding a volumetric rate of 2.9 ± 0.3 × 10−6 flares pc−3 h−1.

Webb, S.↗

Anomaly detection search for new resonances decaying into a Higgs boson and a generic new particle $X$ in hadronic final states using $\sqrt{s}$ = 13 TeV $pp$ collisions with the ATLAS detector

A search is presented for a heavy resonance $Y$ decaying into a Standard Model Higgs boson $H$ and a new particle $X$ in a fully hadronic final state. The full Large Hadron Collider run 2 dataset of proton-proton collisions at $\sqrt{s}$ = 13 TeV collected by the ATLAS detector from 2015 to 2018 is used and corresponds to an integrated luminosity of 139 fb –1 . The search targets the high $Y$-mass region, where the $H$ and $X$ have a significant Lorentz boost in the laboratory frame. A novel application of anomaly detection is used to define a general signal region, where events are selected solely because of their incompatibility with a learned background-only model. It is constructed using a jet-level tagger for signal-model-independent selection of the boosted $X$ particle, representing the first application of fully unsupervised machine learning to an ATLAS analysis. Two additional signal regions are implemented to target a benchmark $X$ decay into two quarks, covering topologies where the $X$ is reconstructed as either a single large-radius jet or two smallradius jets. The analysis selects Higgs boson decays into $b$$\overline{b}$, and a dedicated neural-network-based tagger provides sensitivity to the boosted heavy-flavor topology. No significant excess of data over the expected background is observed, and the results are presented as upper limits on the production cross section $σ$ ($pp$ → $Y$ → $XH$ → $q$$\overline{q}$$b$$\overline{b}$) for signals with $m_Y$ between 1.5 and 6 TeV and $m_X$ between 65 and 3000 GeV

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Nowcasting Earthquakes: Imaging the Earthquake Cycle in California With Machine Learning

We propose a new machine learning-based method for nowcasting earthquakes to image the time-dependent earthquake cycle. The result is a timeseries that may correspond to the process of stress accumulation and release. The timeseries are constructed by using principal component analysis of regional seismicity. The patterns are found as eigenvectors of the cross-correlation matrix of a collection of seismicity timeseries in a coarse grained regional spatial grid (pattern recognition via unsupervised machine learning). The eigenvalues of this matrix represent the relative importance of the various eigenpatterns. Using the eigenvectors and eigenvalues, we compute the weighted correlation timeseries of the regional seismicity. This timeseries has the property that the weighted correlation generally decreases prior to major earthquakes in the region, and increases suddenly just after a major earthquake occurs. As in a previous paper, we find that this method produces a nowcasting timeseries that resembles the hypothesized regional stress accumulation and release process characterizing the earthquake cycle. We then address the problem of whether the timeseries contain information regarding future large earthquakes. For this, we compute a receiver operating characteristic and determine the decision thresholds for several future time periods of interest (optimization via supervised machine learning). We find that signals can be detected that can be used to characterize the information content of the timeseries. These signals may be useful in assessing present and near-future seismic hazards.

58 GEOSCIENCES↗

Automatic point Cloud Building Envelope Segmentation (Auto-CuBES) using Machine Learning

Modern retrofit construction practices use 3D point cloud data of the building envelope to obtain the as-built dimensions. However, manual segmentation by a trained professional is required to identify and measure window openings, door openings, and other architectural features, making the use of 3D point clouds labor-intensive. In this study, the Automatic point Cloud Building Envelope Segmentation (Auto-CuBES) algorithm is described, which can significantly reduce the time spent during point cloud segmentation. The Auto-CuBES algorithm inputs a 3D point cloud generated by commonly available surveying equipment and outputs a wire-frame model of the building envelope. Unsupervised machine learning methods were used to identify facades, windows, and doors while minimizing the number of calibration parameters. Additionally, Auto-CuBES generates a heat map of each facade indicating non-planar characteristics that are crucial for the optimization of connections used in overclad envelope retrofits. With a scan resolution of 3 mm, the resulting window dimensions showed a mean absolute error of 4.2 mm compared to manual laser measurements.

Maldonado Puente, Bryan↗

Genarris 2.0: A Random Structure Generator for Molecular Crystals

Genarris is an open source Python package for generating random molecular crystal structures with physical constraints for seeding crystal structure prediction algorithms and training machine learning models. Here we present a new version of the code, containing several major improvements. A MPI-based parallelization scheme has been implemented, which facilitates the seamless sequential execution of user-defined workflows. A new method for estimating the unit cell volume based on the single molecule structure has been developed using a machine-learned model trained on experimental structures. A new algorithm has been implemented for generating crystal structures with molecules occupying special Wyckoff positions. A new hierarchical structure check procedure has been developed to detect unphysical close contacts efficiently and accurately. New intermolecular distance settings have been implemented for strong hydrogen bonds. To demonstrate these new features, we study two specific cases: benzene and glycine. Genarris finds the experimental structures of the two polymorphs of benzene and the three polymorphs of glycine. Program summary Program Title: Genarris 2.0 Program Files doi: http://dx.doi.org/10.17632/grx6mz4pjn.1 Licensing provisions: BSD-3 Clause Programming language: Python, C External routines/libraries: Spglib, ASE, pymatgen, SciPy, mpi4py, scikit-learn, PyTorch, FHI-aims. Nature of problem: Molecular crystal structure prediction. Solution method: Genarris 2.0 generates molecular crystal structures over the 230 space groups, on general and special Wyckoff positions, using physical constraints. Down-sampling of the generated structures may be performed subsequently, based on molecular crystal packing descriptors and an unsupervised machine learning algorithm. Lastly, ab initio structure relaxation may be performed for the final pool. Depending on the user-defined workflow implemented, Genarris may be used to generate diverse molecular crystal datasets to seed evolutionary algorithms or to train machine learning algorithms or as a standalone crystal structure prediction method. Restrictions: For crystal structure generation, the molecule of interest must be semi-rigid with no bond rotational degrees of freedom. Unusual features: Genarris 2.0 is a highly distributed program, making use of MPI for Python parallelization. The user has the ability to design and implement workflows by executing a user-defined list of procedures. Genarris 2.0 offers new features including a machine learning model for estimating the molecular volume in the solid state from the single molecule structure, structure generation in special Wyckoff positions of space groups, hierarchical structure checks including rigorous treatment of non-orthogonal structures, and clustering and down-selection workflows combining first principles simulations with machine learning. (C) 2020 Elsevier B.V. All rights reserved.

Crystal structure prediction↗

Confidentiality-preserving machine learning algorithms for soft-failure detection in optical communication networks

Automated fault management is at the forefront of next-generation optical communication networks. The increase in complexity of modern networks has triggered the need for programmable and software-driven architectures to support the operation of agile and self-managed systems. In these scenarios, the European Telecommunications Standards Institute zero-touch network and service management approach is imperative. The need for machine learning algorithms to process the large volume of telemetry data brings safety concerns as distributed cloud-computing solutions become the preferred approach for deploying reliable communication network automation. This paper’s contribution is twofold. First, we propose a simple yet effective method to guarantee the confidentiality of the telemetry data based on feature scrambling. The method allows the operation of third-party computational services without direct access to the full content of the collected data. Additionally, the effectiveness of four unsupervised machine learning algorithms for soft-failure detection is evaluated when applied to the scrambled telemetry data. The methods are based on factor analysis, principal component analysis, nonlinear principal component analysis, and singular value decomposition. Most dimensionality reduction algorithms have the common property that they can maintain similar levels of fault classification performance while hiding the data structure from unauthorized access. Evaluations of the proposed algorithms demonstrate this capability.

97 MATHEMATICS AND COMPUTING↗