Automatic detection of crystallographic defects in STEM images by unsupervised learning with translational invariance
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
We present a novel approach for resolving modes of rupture directivity in large populations of earthquakes. A seismic spectral decomposition technique is used to first produce relative measurements of radiated energy for earthquakes in a spatially compact cluster. The azimuthal distribution of energy for each earthquake is then assumed to result from one of several distinct modes of rupture propagation. Rather than fitting a kinematic rupture model to determine the most likely mode of rupture propagation, we instead treat the modes as latent variables and learn them with a Gaussian mixture model. The mixture model simultaneously determines the number of events that best identify with each mode. The technique is demonstrated on four datasets in California, each with compact clusters of several thousand earthquakes with comparable slip mechanisms. We show that the datasets naturally decompose into distinct rupture propagation modes that correspond to different rupture directions, and the fault plane is unambiguously identified for all cases. We find that these small earthquakes exhibit unilateral ruptures 63–73% of the time on average. Here, the results provide important observational constraints on the physics of earthquakes and faults.
Not Available
One of the outstanding analytical problems in X-ray single-particle imaging (SPI) is the classification of structural heterogeneity, which is especially difficult given the low signal-to-noise ratios of individual patterns and the fact that even identical objects can yield patterns that vary greatly when orientation is taken into consideration. Proposed here are two methods which explicitly account for this orientation-induced variation and can robustly determine the structural landscape of a sample ensemble. The first, termed common-line principal component analysis (PCA), provides a rough classification which is essentially parameter free and can be run automatically on any SPI dataset. The second method, utilizing variation auto-encoders (VAEs), can generate 3D structures of the objects at any point in the structural landscape. Both these methods are implemented in combination with the noise-tolerant expand–maximize–compress (EMC) algorithm and its utility is demonstrated by applying it to an experimental dataset from gold nanoparticles with only a few thousand photons per pattern. Both discrete structural classes and continuous deformations are recovered. These developments diverge from previous approaches of extracting reproducible subsets of patterns from a dataset and open up the possibility of moving beyond the study of homogeneous sample sets to addressing open questions on topics such as nanocrystal growth and dynamics, as well as phase transitions which have not been externally triggered.
Explore the source record for details and available documents.
As compute clusters continue to grow in scale and complexity, the frequency of detected anomalies in their operation significantly increases. Timely detection of anomalous events is vital to maintain system efficiency and availability. This study presents an attentionbased graph neural network (GNN) for detecting anomalies in clusters at the compute node level and for providing detailed root cause analysis. We show the effectiveness of attention-based GNNs to accurately detect and localize anomalies on real-world datasets.
Predictability of wind resource conditions is critical for offshore wind design and operations. While many studies of extreme wind conditions focus on specific events such as low-level jets or ramps, these rely on threshold definitions that limit generality. Here we present a data-driven framework that combines principal component analysis (PCA), self-organizing maps (SOM), and k-means clustering to classify wind resource conditions as typical and anomalous from climatological data. Anomalies are defined not by fixed thresholds but by flagging samples located far from SOM node centers inside the baseline SOM structure. This reframes extremes as rare ebents and hence, likely difficult to anticipate by numerical weather prediction models. We applied this approach to 23 years (2000–2022) of hourly profiles from the NOW-23 hindcast model at the Humboldt Wind Energy Area. Classification is conducted on a feature space consisting of 10 m wind speed and direction, bulk shear and veer across 30–270 m, and a low-level jet index. Dimensionality reduction is achieved through PC. A 2 × 3 OM lattice trained on the PCA vectors identified six baseline regimes spanning weak to strong flow states. High quantization-error profiles are identified and re-clustered into four anomalous regimes. The baseline regimes exhibited clear seasonal and diurnal cycles. Meanwhile, the anomalous regimes represented <10 % of all hours but showed distinct combinations of speed, shear, and veer, when compared to the baseline regimes. Anomalous regimes are typically short-lived (~few hours), yet their transitions can lead to hub-height wind changes of −18 to +9 m s -1 . For a representative 15 MW turbine, these shifts imply rapid swings in capacity factor from near-full output to negligible generation. Validation with lidar buoy data showed 51% agreement in SOM labels across ~6,000 overlapping hours, with most mismatches confined to adjacent speed classes. HRRR comparisons further revealed that anomalous regimes were disproportionately associated with forecast biases exceeding 5 m s -1 . Together, these results reframe extremes in offshore wind from absolute maxima or minima to weather states that are difficult to anticipate from models.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
This paper describes a novel approach to the development of a learning control system for autonomous space robot (ASR) which presents the ASR as a 'baby' -- that is, a system with no a priori knowledge of the world in which it operates, but with behavior acquisition techniques that allows it to build this knowledge from the experiences of actions within a particular environment (we will call it an Astro-baby). The learning techniques are rooted in the recursive algorithm for inductive generation of nested schemata molded from processes of early cognitive development in humans. The algorithm extracts data from the environment and by means of correlation and abduction, it creates schemata that are used for control. This system is robust enough to deal with a constantly changing environment because such changes provoke the creation of new schemata by generalizing from experiences, while still maintaining minimal computational complexity, thanks to the system's multiresolutional nature.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Abstract not provided.
Environmental screening of gamma radiation consists of detecting weak nuisance and anomaly signal in the presence of strong and highly varying background. In a typical scenario, a mobile detector-spectrometer continuously measures gamma radiation spectra in short, e.g., one-second, signal acquisition intervals. The measurement data is a 2D matrix, where one dimension is gamma ray energy, and the other dimension is the number of measurements or total time. In principle, gamma radiation sources can be detected and identified from the measured data by their unique spectral lines. Detecting sources from data measured in a search scenario is difficult due to the highly varying background because of naturally occurring radioactive material (NORM), and low signal-to-noise ratio (S/N) of spectral signal measured during one-second acquisition intervals. The objective of this work is to explore unsupervised machine learning (ML) algorithms for development of a digital twin of gamma radiation background, and for detection and identification of weak nuisances and anomalies events in the presence of highly fluctuating background. In one segment of work, we developed a gamma background estimation model using a Longshort term memory (LSTM) network for one-step CPS time series prediction. The LSTM model was validated with two data sets of measurements from two independent NaI detectors positioned on a mobile platform. The data sets contained background radiation only and no orphan isotope sources. The LSTM model was constructed and tested using data from one of the detectors. Performance of the LSTM model was validate through one-step prediction of CPS time series of another NaI detector without re-training. This approach allows to create a digital twin for nuclear background estimation. Using LSTM, it could be possible to detect a source through subtraction of the estimated counts from the measured background. In another segment of work, we investigated detection of gamma emitting sources in the presence of complex background using unsupervised machine learning. Spectral lines of isotopes are difficult to observe in one-second measurements. Averaging over the entire measurement campaign data set reveals spectral lines of most common background isotopes. Spectral lines of orphan sources, which might appear only in a few measurements during the campaign, will be washed out if averaging is performed over the entire measurement data set. The approach we have explored consists of extracting one-second measurements containing weak spectral features through data clustering. Averaging one-second spectra in a cluster should reveal the presence of anomaly sources. We created two ML models using K-means clustering and Neural Network Self-organizing Map (SOM). Performance of these ML models was benchmarked using search data. One data set contained 137 Cs source, and another dataset contained 131 I source.
We report faults in Heating, Ventilation, and Air Conditioning (HVAC) systems of buildings result in significant energy waste in building operation. With fast-growing sensing data availability and advancement in computing, computational modeling has demonstrated strong capability to detect and diagnose HVAC system faults, hence, ensuring efficient building operation. This paper comprehensively reviews the state-of-the-art computing-based fault detection and diagnosis (FDD) for HVAC systems. Overall, the reviewed computing-based FDD methods are classified as two major approaches: knowledge-based and data-driven approaches. We then identify multiple important topics, including data availability, training data size, data quality, approach generality, capability, interpretability, and required modeling efforts, along with corresponding metrics to summarize the most updated FDD development. Generally, the knowledge-based approaches are further divided as physics-based modeling, Diagnostic Bayesian Network, and performance indicator-based methods while data-driven approaches include supervised learning, unsupervised learning, and regression and statistics-based methods. State-of-the-art FDD development, remaining challenges, and future research directions are further discussed to push forward FDD in practice. Availability of fault data, capability of existing methods to deal with complex fault situations (such as simultaneous faults), modeling interpretability for data-driven methods, and required engineering efforts for physics-based methods are identified as remaining challenges in FDD development. Improving modeling fidelity and reducing modeling efforts are essential for applying physics-based methods in real buildings. Meanwhile, addressing fault data availability, increasing algorithm adaptability, and handling multiple faults are essential to further enhance the applicability of data-driven FDD approaches.
Focus areas: (1) Data assimilation enabled by machine learning, unsupervised learning (including deep learning), and (2) Predictive modeling through the use of AI techniques and AI-derived model components, comprising a hierarchy of models.