Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Unsupervised learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Distributed fiber sensor and machine learning data analytics for pipeline protection against extrinsic intrusions and intrinsic corrosions

This paper presents an integrated technical framework to protect pipelines against both malicious intrusions and piping degradation using a distributed fiber sensing technology and artificial intelligence. A distributed acoustic sensing (DAS) system based on phase-sensitive optical time-domain reflectometry (φ-OTDR) was used to detect acoustic wave propagation and scattering along pipeline structures consisting of straight piping and sharp bend elbow. Signal to noise ratio of the DAS system was enhanced by femtosecond induced artificial Rayleigh scattering centers. Data harnessed by the DAS system were analyzed by neural network-based machine learning algorithms. The system identified with over 85% accuracy in various external impact events, and over 94% accuracy for defect identification through supervised learning and 71% accuracy through unsupervised learning.

Peng, Zhaoqiang↗

AI to Automate ModEx for Optimal Predictive Improvement and Scientific Discovery

Focal Areas: Data acquisition and assimilation enabled by machine learning, AI, advanced experimental optimization, unsupervised learning, and hardware-related AI efforts; Predictive modeling through AI techniques and AI-derived model components; Using AI to design a hierarchical model prediction system consisting and model selection; and, Interrogating complex data (observed and simulated) using AI, big data analytics, and other advanced methods such as explainable AI and physics- or knowledge-guided AI.

54 ENVIRONMENTAL SCIENCES↗

Solar forecasting using machine learned cloudiness classification

Methods and systems for predicting irradiance include learning a classification model using unsupervised learning based on historical irradiance data. The classification model is updated using supervised learning based on an association between known cloudiness states and historical weather data. A cloudiness state is predicted based on forecasted weather data. An irradiance is predicted using a regression model associated with the cloudiness state.

Hamann, Hendrik F.↗

The LHC Olympics 2020 a community challenge for anomaly detection in high energy physics

A new paradigm for data-driven, model-agnostic new physics searches at colliders is emerging, and aims to leverage recent breakthroughs in anomaly detection and machine learning. In order to develop and benchmark new anomaly detection methods within this framework, it is essential to have standard datasets. To this end, we have created the LHC Olympics 2020, a community challenge accompanied by a set of simulated collider events. Participants in these Olympics have developed their methods using an R&D dataset and then tested them on black boxes: datasets with an unknown anomaly (or not). Furthermore, methods made use of modern machine learning tools and were based on unsupervised learning (autoencoders, generative adversarial networks, normalizing flows), weakly supervised learning, and semi-supervised learning. This paper will review the LHC Olympics 2020 challenge, including an overview of the competition, a description of methods deployed in the competition, lessons learned from the experience, and implications for data analyses with future datasets as well as future colliders.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Unsupervised quantum circuit learning in high energy physics

Unsupervised training of generative models is a machine learning task that has many applications in scientific computing. Here, in this work, we evaluate the efficacy of using quantum circuit-based generative models to generate synthetic data of high energy physics processes. We use nonadversarial, gradient-based training of quantum circuit Born machines to generate joint distributions over two and three variables.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Sub-pilot-scale Production of High-Value Products from U.S. Coals

Investigators from the University of Utah, University of Wyoming and Marshall University pursued a program to study the conversion of raw coal to high-value products of carbon fiber and silicon carbide. Team members also developed an initial framework for a data portal that can incorporate laboratory data on coal processing and product quality, and also work with tools for machine learning for data analysis, data visualization and economic assessment. Experimental R&D efforts focused on the conversion of raw coal to coal tar and other byproducts, and the resulting tar intermediates were upgraded to form anisotropic and isotropic pitch materials. These pitch materials were produced from coal using both thermal (pyrolysis) and chemical (mild solvolysis liquefaction) decomposition of raw coal. Four different coals were studied: Utah bituminous coal (Sufco), Wyoming PRB coal (Black Thunder), Illinois bituminous coal (Illinois #6), and West Virginia bituminous coal (Flying Eagle). Both metallurgical-grade coking coals and lower-grade steam coals were investigated, and controlled secondary gas-phase reactions were used during a two-stage pyrolysis process to induce cracking and condensation reactions among the pyrolytic tar species. This approach successfully improved the performance of the lower grade coals for yielding pitch materials, with properties more consistent with a commercial-grade pitch that had previously demonstrated success for quality carbon fiber production. The use of waste plastic materials was also studied, to help improve physical and chemical characteristics of the intermediate tars and final pitch product; in particular, for lowering the pitch softening point to an acceptable level for melt spinning carbon fiber. Mild solvolysis liquefaction was also used as a method for producing pitch for carbon fiber production. As expected, significantly higher pitch yields were obtained using this approach, and waste plastic materials were also successfully used to reduce pitch softening point to an acceptable level. The plastic materials were also utilized to create a solvent for the mild solvolysis process, and this plastic-derived solvent was shown to provide results consistent with more expensive commercial chemical solvents, and could thus avoid the need for costly recovery and recycle of a liquefaction solvent. Additional experimental R&D focused on the production of silicon carbide (β-SiC) from the residual char byproduct from pitch production, and also on the production of carbon fiber from the anisotropic pitch. SiC was successfully synthesized using a mixture of residual char and sandstone at a ratio of 1:1. Reaction temperature and residence time were optimized and yielded a product purity of 81%. For carbon fiber production, the most successful pitch samples were obtained from the mild solvolysis liquefaction approach, combined with the use of a plastic (HDPE)-derived solvent. Fiber properties improved over time as laboratory fiber production methodologies improved, and final yields of carbon fiber were obtained with a diameter of 12.14 ± 1.10 um, Modulus of 173.73 ± 15.25 GPa, and Tensile Strength of 1.04 ± 0.10 GPa. A proof-of-concept Modern Community Research Data Portal (MCRDP) was developed and deployed for coal and coal-derived pitch characterization, with the full support of (i) remote web-based access, (ii) distributed analysis, (iii) interactive visualization and exploration, (iv) shared and long-term data access, (v) advanced query capabilities and (vi) real-time collaboration. The Coal to Products Data Portal “coaltoproducts.org” provides researchers with space to store and share data within a project, tools for analyzing and understanding data for scientific investigation, and the ability to publish data to the broader community for reproducibility. The portal leverages the Material Commons 2.0 (MC) platform developed by the Center for PRedictive Integrated Structural Materials Science (PRISMS) of the University of Michigan, to achieve long-term longevity of data collections and, more importantly, collaborative science. A number of data visualization tools were also assessed and implemented for interrogating the experimental and modeling data. The machine learning portion of this project analyzed datasets from two different coal conversion processes performed on a diverse set of coal samples from both the coal pyrolysis experiments and the solvent liquefaction experiments. The work was initiated by exploring standard regression models on the pyrolysis data, aiming to understand the impact of sample characteristics and processing conditions on key product metrics. Over the course of the project, the focus expanded to include a variety of machine learning tools, delving into both supervised and unsupervised learning methods. Models tested on the pyrolysis data included linear, ridge, lasso, elastic-net, Gaussian process, random forest regression, and AutoSklearn, and the approach was continually refined to enhance predictive accuracy and model interpretability. Similar techniques were applied to the liquefaction data with an additional focus on feature engineering. Along with mesophase content, additional outputs of interest were the pitch yield, softening point, and QI content. Insights derived from these analyses are crucial in determining the factors influencing the quality and yield of coal-derived products. As the work progressed, the research evolved from foundational model comparisons to analyses of random forests, decision paths, and feature importance scores. A thorough market analysis was performed to examine the prospects of coal-based carbon fibers. The best opportunities for coal come from its lower and more stable price relative to petroleum, particularly for subbituminous coals, which is the primary advantage that a coal refinery may have over a petroleum refinery. Before a commercial CTP production facility can be modeled, however, several things need to be understood regarding the nature of the would-be coal refinery. These include the technology to be deployed, the size of facility, the volume(s) of co-product(s), and the waste and emissions profile of the plant. The volume of co-products and waste may be substantial and will require separate market analysis to ensure viability. In the near-term, the importance of coal tar pitch, in the form of carbon pitch, to the aluminum and steel industries is likely to overshadow the alternative use of this material as an input for carbon fiber. The importance of steel and aluminum in building materials, and the need for carbon materials in their manufacturing, will ensure that demand for these products remains for the long run. In addition, carbon fiber may also be the best substitute for steel and aluminum well into the future. While society will eventually be able to shift production of much of its electricity needs to renewables, it will not be able to shift away from fossil fuels for production of high-strength construction and vehicular materials. Demand for carbon fiber is expected to increase quickly, but the volume of carbon fiber and the amount of coal that would be needed to produce even a sizeable share of this market may still be relatively small compared to current coal production. Thus, other coal-based products like graphene, graphite, carbon foams, resins, and carbon-based building products will play important roles in sustaining coal production as coal-fired power generation continues to decline.

01 COAL, LIGNITE, AND PEAT↗

A low-latency graph computer to identify metastable particles at the Large Hadron Collider for real-time analysis of potential dark matter signatures

Abstract Image recognition is a pervasive task in many information-processing environments. We present a solution to a difficult pattern recognition problem that lies at the heart of experimental particle physics. Future experiments with very high-intensity beams will produce a spray of thousands of particles in each beam-target or beam-beam collision. Recognizing the trajectories of these particles as they traverse layers of electronic sensors is a massive image recognition task that has never been accomplished in real time. We present a real-time processing solution that is implemented in a commercial field-programmable gate array using high-level synthesis. It is an unsupervised learning algorithm that uses techniques of graph computing. A prime application is the low-latency analysis of dark-matter signatures involving metastable charged particles that manifest as disappearing tracks.

47 OTHER INSTRUMENTATION↗

Visualization of hydraulic fracture using physics-informed clustering to process ultrasonic shear waves

Ultrasonic transmission is sensitive to the spatial variation in mechanical properties of materials due to the presence of cracks/fractures. Wave propagation through fractured media introduces changes in the frequency content, travel time and transmission coefficient of the wave. A workflow based on physics-informed unsupervised learning is developed to process the transmitted ultrasonic-shear waveforms to non-invasively visualize the geomechanical alterations due to hydraulic fracturing. Novelty of the work involves the assignment of both statistically consistent and physically consistent clusters to the measurements of shear waveforms acquired across the one axial and two frontal planes. Physically consistent/relevant information is incorporated by considering the travel time of the peak of spectral energy and transmission coefficient of the transmitted waveform. The proposed workflow generates maps of geomechanical alterations across the frontal and axial planes of the sample. The outputs of the workflow are in good agreement with independent techniques viz. acoustic emission and X-ray computed tomography. Finally, the proposed workflow can be adapted for improved fracture characterization in the subsurface when processing sonic-logging, cross-wellbore seismic or surface seismic waveform data.

42 ENGINEERING↗

Statistical Complexity of Quantum Learning

Abstract Learning problems involve settings in which an algorithm has to make decisions based on data, and possibly side information such as expert knowledge. This study has two main goals. First, it reviews and generalizes different results on the data and model complexity of quantum learning, where the data and/or the algorithm can be quantum, focusing on information‐theoretic techniques. Second, it introduces the notion of copy complexity, which quantifies the number of copies of a quantum state required to achieve a target accuracy level. Copy complexity arises from the destructive nature of quantum measurements, which irreversibly alter the state to be processed, limiting the information that can be extracted about quantum data. As a result, empirical risk minimization is generally inapplicable. The paper presents novel results on the copy complexity for both training and testing. To make the paper self‐contained and approachable by different research communities, an extensive background material is provided on classical results from statistical learning theory, as well as on the distinguishability of quantum states. Throughout, the differences between quantum and classical learning are highlighted by addressing both supervised and unsupervised learning, and extensive pointers are provided to the literature.

97 MATHEMATICS AND COMPUTING↗

InClass nets: independent classifier networks for nonparametric estimation of conditional independence mixture models and unsupervised classification

Abstract Conditional independence mixture models (CIMMs) are an important class of statistical models used in many fields of science. We introduce a novel unsupervised machine learning technique called the independent classifier networks (InClass nets) technique for the nonparameteric estimation of CIMMs. InClass nets consist of multiple independent classifier neural networks (NNs), which are trained simultaneously using suitable cost functions. Leveraging the ability of NNs to handle high-dimensional data, the conditionally independent variates of the model are allowed to be individually high-dimensional, which is the main advantage of the proposed technique over existing non-machine-learning-based approaches. Two new theorems on the nonparametric identifiability of bivariate CIMMs are derived in the form of a necessary and a (different) sufficient condition for a bivariate CIMM to be identifiable. We use the InClass nets technique to perform CIMM estimation successfully for several examples. We provide a public implementation as a Python package called RainDancesVI.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

AI-Based Integrated Modeling and Observational Framework for Improving Seasonal to Decadal Prediction of Terrestrial Ecohydrological Extremes

Focal Areas: (1) Insight gleaned from complex data (both observed and simulated) using artificial intelligence(AI), big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI (2) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing).

54 ENVIRONMENTAL SCIENCES↗

AI-Based Upgrades to Observational Data Centers to Facilitate Data Interoperability

Focal Areas: (1) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Focal areas 2 and 3 have critical dependencies to the modernization described. Key benefits to the focal areas: (1) Modernized observatory framework capable of agile adaptive observation, (2) Advanced instrument and data tagging supporting AI data acquisition for assimilation or validation, and (3) Widespread data interoperability bridging Earth system prediction scales

54 ENVIRONMENTAL SCIENCES↗

Elucidating and predicting the dynamic evolution of water and land systems due to natural and energy-related forcings

Focal Area(s): 3. Insight gleaned from complex data (both observed and simulated) using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI; & 1. Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Science Challenge: Interactions between water, land, and energy systems are complex and occur on a variety of scales, ranging from local to basinal to regional. Accurately predicting the behavior of ground water and surface water systems for 5-10 years and beyond requires an understanding of the current system and the ability to model both the natural system at scale and human-induced forcings related to energy and other activities. Artificial intelligence and machine learning (AI/ML) combined with modern compilation and integration efforts for U.S. groundwater and surface water systems present potential solutions to bolstering detailed physics-based models of these systems. Big data tied with ML and physics-based modeling can drive breakthroughs in understanding the earth system, but research is often impeded by data access (e.g., privacy issues), quality, formats, gaps, multi-source, multi-scale, integration, and spatiotemporal challenges. Effective integration of real data and simulated (synthetic) data that fill gaps is critical. Overcoming these complex data and model integration challenges will enable a transformational approach to acquiring enhanced understanding of environmental systems.

54 ENVIRONMENTAL SCIENCES↗

Integrating Applied Energy and BER Smart Data Capabilities to Develop a DOE Data Fabric for Energy-Water R&D

Focal Area(s): 1) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Science Challenge: DOE R&D, including DOE’s Basic Energy Research (BER)’s Environmental Systems Science Division (EESSD) program and DOE’s applied energy research (AER) programs (EERE, FE, and NE) are producers and consumers of Earth systems datasets. This white paper focuses on the first topic area from the call in relation to how crosscutting resources and innovations from DOE’s EESSD and AER can be brought to bear to mutual benefit and more efficient energy-water, Earth system data resources through improved. The overarching challenge posed by this call focuses on how DOE can directly leverage artificial intelligence (AI) to engineer a substantial (paradigm-changing) improvement in Earth System Predictability? While stemming from DOE BER’s EESSD program, this is a challenge that is faced and also being addressed by DOE’s AER programs. Over the past decade plus, FE, EERE, and NE programs have made important strides towards addressing this need. These strides are in many ways highly complementary to EESSD’s MODEX efforts. Energy water systems spanning metocean to groundwater to surface water systems all are data driven whether for basic energy or applied energy. These are remote, multi-variate, complex natural, and in many cases engineered, systems. Key needs and challenges of both EESSD and AER include developing data-focused tools to enhance data search and discovery to fill in knowledge gaps (address sparse data challenge), and rapidly transform datasets, including disparate and multi-source data. Leveraging DOE on-premise computing (HPC, exascale) infrastructure supports the computing-intensive algorithms required to execute these data acquisition and transformation processes to derive enriched knowledge and data, driving AI/ML and big data analytics for these systems. The opportunity lies in combining BER and AER efforts to provide a more robust, advanced, efficient and complete computing data fabric to address energy-water data acquisition and assimilation needs which currently pose significant impediments to AI/ML predictions and research.

54 ENVIRONMENTAL SCIENCES↗

Integrating Models with Real-time Field Data for Extreme Events: From Field Sensors to Models and Back with AI in the Loop

Focal Area(s): This whitepaper is responsive to focal area (1) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). We discuss Artificial Intelligence and Machine Learning (AI/ML) enabled integration of real-time data into the extreme event modeling workflow to improve the predictive capabilities of these models, and deliver real-time feedback to remote sensors, including software and data engineering challenges.

54 ENVIRONMENTAL SCIENCES↗

Characterization of Extreme Hydroclimate Events in Earth System Models using ML/AI

Focal Area(s): (1) We put forward the concepts of data assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization and unsupervised learning applied to downscale information within Earth System Models (ESMs). (2) We discuss predictive modeling through the use of AI techniques and other tools to design a prediction system comprising of a hierarchy of models (e.g., AI-driven model/component/parameterization selection) to improve the characterization of extreme hydroclimate events in ESMs. The AI-based models will run five-six order of magnitude times faster, yet will provide similar accuracy, allowing us to provide range bounds on uncertainty faster and thus enabling faster extreme event identification. (3) Further, we consider the insight gleaned from complex data (observed/simulated) using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI to improve the characterization of extreme hydroclimate events in ESMs.

54 ENVIRONMENTAL SCIENCES↗

Neural Density Estimation and Uncertainty Quantification for ChemCam Spectra [Slides]

The ChemCam instrument of Curiosity uses laser-induced breakdown spectroscopy (LIBS). It fires a laser at target and vaporizes rock surfaces, creating a plasma. Three spectrographs divide the plasma light into wavelengths for chemical analysis: ultraviolet, violet, and visible near-infrared. Regression methods (SVR, PCR, CNN) have been employed for calibration (prediction of the elemental composition of samples); however, labeled ChemCam samples are limited. Here, we focus on unsupervised learning and employ generative models from ChemCam analysis. Further, we use labels (supervised) in combination to the generative model to compute uncertainties related to predictions. We report generative modeling can be successfully applied to model real-world data. Normalizing flow models can be efficiently constructed on latent spaces for fast downstream inference. Unsupervised and supervised learning can be combined to form an uncertainty quantification framework.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗