Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data- limited”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Multitask Machine Learning of Collective Variables for Enhanced Sampling of Rare Events

Computing accurate reaction rates is a central challenge in computational chemistry and biology because of the high cost of free energy estimation with unbiased molecular dynamics. In this work, a data-driven machine learning algorithm is devised to learn collective variables with a multitask neural network, where a common upstream part reduces the high dimensionality of atomic configurations to a low dimensional latent space and separate downstream parts map the latent space to predictions of basin class labels and potential energies. Here, the resulting latent space is shown to be an effective low-dimensional representation, capturing the reaction progress and guiding effective umbrella sampling to obtain accurate free energy landscapes. This approach is successfully applied to model systems including a 5D Müller Brown model, a 5D three-well model, the alanine dipeptide in vacuum, and an Au(110) surface reconstruction unit reaction. It enables automated dimensionality reduction for energy controlled reactions in complex systems, offers a unified and data-efficient framework that can be trained with limited data, and outperforms single-task learning approaches, including autoencoders.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

State Estimation for Distribution Networks with Asynchronous Sensors Using Stochastic Descent: Preprint

This paper investigates the problem of state estimation for distribution networks with asynchronous sensors comprising of a mix of smart meters and phasor measurement units (PMUs) with multiple sampling and reporting rates. We consider two independent scenarios of state estimation and tracking, with either voltages or currents as states. With these two sets, we investigate estimation under (a) full data, assuming all measurements are available and (b) limited data, where an online algorithmic approach is adopted to estimate the possibly time-varying states by processing measurements as and when available. The proposed algorithm, inspired by the classical Stochastic Gradient Descent (SGD) approach updates the states based on the previous estimate and the newly available measurements. Finally, we demonstrate the estimation and tracking efficacy through numerical simulations on the IEEE-37 test network, while also highlighting how estimation with currents as states leads to faster convergence.

asynchronous sensors↗

Deriving Stable Peak Models to Fit Complex XPS Data From Cu Contaminated Pt Electrocatalysts

X-ray Photoelectron Spectroscopy spectra peak models, designed to partition photoemission signals emanating from different elements or chemical states within an atom, are fitted to data limited to an energy interval over which inelastically scattered photoemission signal can be estimated. While the choice of background approximation and line shapes of components to the peak model requires careful consideration, the energy interval used to define the data to which the peak model is optimized has a significant impact on the final peak model. The relationship between the background intensity and data intensity at the start and end of the energy interval dictates the line shapes used in the peak model. In this work, we devise a method to peak fit a complex overlapping Cu 3p and Pt 4f XPS peak structure to perform the elemental quantification. We first use an Al 2s peak to illustrate how background curves approach data at the limits of the energy interval over which the background is defined, influencing the analysis of XPS spectra. Next, we demonstrate the nature of interactions between specific line shapes (Voigt and pseudo-Voigt profiles) suitable for photoemission peaks and a specific background curve (Shirley) and a peak model is presented that includes components to the peak model that accommodates background intensity during fitting of the peak model to data. The peak model allowed for quantification of the contributions of Pt 4f peaks emanating from the substrate that exhibits strong asymmetry in the presence of the inhomogeneously distributed Cu species, mostly of Lorentzian character.

XPS↗

Graphene SHDMC Data

The data used to produce all of the figures and tables in the manuscript titled "Highly Accurate Many-Body Theory Reaches 2D Materials" can be found here. This data set includes: -SHDMC results for graphene -selected CI with and without re-normalized second-order perturbation (rPT2) theory corrections for graphene -data demonstrating that SHDMC displays an exponential rate of convergence -data used for sCI + rPT2 complete basis set extrapolation -data used to extrapolate SHDMC energies to the infinite basis limit -data used to demonstrate compactness of SHDMC wavefunction

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Natural Language Processing-Enhanced Nuclear Industry Operating Experience Data Analysis: Aggregation and Interpretation of Multi-Report Analysis Results

Industry-wide operating experience is a critical source of raw data for reliability and risk model parameter estimations for nuclear power plants. A large portion of operating experience data are failure events stored as reports that contain unstructured data, such as narratives. In current practice, a failure report is usually reviewed and manually coded by analysts. The coding is based on extracting several event characteristics such as system name, component type, sub-part type, failure mode, and failure cause. Event narratives are mostly used to help understand events and extract their characteristics. In this line of research, we aim to maximize the usage of event narratives by leveraging natural language processing (NLP) methods to automatically convert an event narrative to a causal graph. This research has promise to improve physical understanding of failure initiation and propagation and to facilitate use of non-failure data (e.g., near-misses and degradations) to complement the limited data pool of failures. In our previous work, we developed an NLP tool and applied it to analyze a number of licensee event reports submitted by U.S. nuclear power plants to the Nuclear Regulatory Commission. In this paper, we will report our recent research progress in aggregating the results of multiple reports, developing network model(s), and drawing statistical insights.

99 GENERAL AND MISCELLANEOUS↗

High-throughput spin-bath characterization of spin defects in semiconductors

Detailed knowledge of the local environments of spin defects in semiconductors, such as nitrogenvacancy (NV) centers in diamond or divacancies in silicon carbide, is crucial for optimizing control and entanglement protocols in quantum sensing and information applications. However, at present a direct experimental characterization of individual defect environments is not scalable, as conventional spin-bath measurements are time consuming and difficult to automate. Achieving high-throughput characterization requires short experiments to probe the spin bath. However, with fewer and noisier measurements, the inverse problem of recovering spin-bath properties from measured data becomes ill posed, with multiple spin baths having a high likelihood of yielding the same data. In this work, we present a set of computational tools to resolve the ill-posed inverse problem of recovering the atomic positions and hyperfine couplings of random nuclei surrounding spin defects from sparse, noisy experimental coherence data, which can be obtained in hours. Here, we use a trans-dimensional Bayesian approach that incorporates ab initio data to yield full posterior distributions over nuclear spin environments, enabling robust recovery from limited data. We also provide practical tools and guidelines to determine the limits of detectability for hyperfine couplings under specific dynamical decoupling sequences and sampling conditions. In addition, we demonstrate how the tools developed here, in combination with ab initio simulations of spin baths, can guide the design of efficient experimental protocols for application-specific high-throughput screening. To showcase the utility of our approach, we apply it to design fast dynamical decoupling experiments to characterize the spin baths often individual NV centers in diamond. While the primary focus is on accelerating spin-bath characterization of spin defects, this Bayesian approach also lays the foundation for digital-twin studies of spin defects, where a virtual model of the spin-defect system evolves in real time with ongoing experimental measurements. Together, the set of tools we designed and applied paves the way for scalable deployment of spin defects in semiconductors for quantum sensing and information applications.

Bayesian methods↗

To Fail or not to Fail: An Exploration of Machine Learning Techniques for Predictive Maintenance

Predictive maintenance refers to the ability to predict when machinery or systems need to be maintained. Making an accurate prediction is quite challenging given the costs for both over-estimating (unnecessary maintenance and reduction in availability of assets) and under-estimating (untimely breakdowns and possible loss of equipment or lives). To address these challenges researchers were able to develop new approaches for analyzing oil samples taken extracting samples from oil-wetted machinery that may provide information critical to developing predictive capabilities. We consider the problem from both supervised (though data limited) and unsupervised approaches and provide a first look into a data driven approach for identification of condition indicators. Through this work we identify a collection of candidate features that can form the basis of condition indicators for both a high level discrimination of failure vs. normal operation as well as a set for potential failure mode identification. Finally, we present an anomaly detection framework for detecting failures which can be a viable solution for an onboard analysis tool in deployed systems.

predictive maintenance, anomaly detection, Laserne↗

DOME: Directional medical embedding vectors from Electronic Health Records

Motivation: The increasing availability of Electronic Health Record (EHR) systems has created enormous potential for translational research. Recent developments in representation learning techniques have led to effective large-scale representations of EHR concepts along with knowledge graphs that empower downstream EHR studies. However, most existing methods require training with patient-level data, limiting their abilities to expand the training with multi-institutional EHR data. On the other hand, scalable approaches that only require summary-level data do not incorporate temporal dependencies between concepts. Methods: We introduce a DirectiOnal Medical Embedding (DOME) algorithm to encode temporally directional relationships between medical concepts, using summary-level EHR data. Specifically, DOME first aggregates patient-level EHR data into an asymmetric co-occurrence matrix. Then it computes two Positive Pointwise Mutual Information (PPMI) matrices to correspondingly encode the pairwise prior and posterior dependencies between medical concepts. Following that, a joint matrix factorization is performed on the two PPMI matrices, which results in three vectors for each concept: a semantic embedding and two directional context embeddings. They collectively provide a comprehensive depiction of the temporal relationship between EHR concepts. Results: We highlight the advantages and translational potential of DOME through three sets of validation studies. First, DOME consistently improves existing direction-agnostic embedding vectors for disease risk prediction in several diseases, for example achieving a relative gain of 5.5% in the area under the receiver operating characteristic (AUROC) for lung cancer. Second, DOME excels in directional drug-disease relationship inference by successfully differentiating between drug side effects and indications, correspondingly achieving relative AUROC gain over the state-of-the-art methods by 10.8% and 6.6%. Finally, DOME effectively constructs directional knowledge graphs, which distinguish disease risk factors from comorbidities, thereby revealing disease progression trajectories. The source codes are provided at https://github.com/celehs/Directional-EHRembedding.

60 APPLIED LIFE SCIENCES↗

Database of low‐temperature absorption and fluorescence spectra of native photosynthetic tetrapyrrole macrocycles

Low-temperature (77 K) absorption and fluorescence spectra of 12 naturally occurring photosynthetic tetrapyrrole macrocycles have been recorded in a frozen glass (2-methyltetrahydrofuran). The compounds encompass distinct chromophore classes: porphyrin, chlorophyll c 2 ; chlorin, chlorophylls a, b, d, f and bacteriochlorophylls c, d, e, f; and bacteriochlorin, bacteriochlorophylls a, b, g. The spectra are compared with those of the same pigment in liquid solution (predominantly 2-methyltetrahydrofuran) at room temperature (293 K). The measured Stokes shifts at 77 K across the 12 macrocycles range from ~30 to 300 cm −1 . The spectral data in digital form are made available as part of the PhotochemCAD databases. Literature searches have revealed extensive published data for Chl a (often in biological matrices) but at best rather limited data for less common macrocycles. The availability of a systematic collection of curated spectral data collected at low temperature should be useful for a variety of assessments, including reconstruction of absorption spectra of (bacterio)chlorophyll-containing protein complexes, vibrational analysis of absorption and fluorescence spectra, and calculations where knowledge of energy levels is important.

Niedzwiedzki, Dariusz M. [Washington University in↗

An automated integrated web-based smart tool for open stope design

The Stability Graph is a widely used tool for the design of open stopes in underground mining. Many users of the Stability Graph still apply this design method manually. Although the manual approach has benefits, using multiple graphs and stability number computation charts for each stope surface is time-consuming, even for the experienced mining engineer. Current practice in the use of the method also limits data sharing. This paper presents a StopeSoft web-based tool for open stope stability prediction that is developed on the basis of the Stability Graph method and is available at openstope.com. StopeSoft incorporates flexibility in terms of Stability Graph options and incorporates additional critical factors often overlooked. As a web-based tool, StopeSoft encourages and makes data sharing possible globally, focused on expanding the database and improving the current limitations of the Stability Graph to provide practical, reliable solutions for mining engineers, consultants, and academics. The StopeSoft automated process facilitates the process of open stope stability prediction, saving time and minimizing potential human errors. Statistical treatment of the data accounts for the variability of input parameters to emphasize the probabilistic nature of the Stability Graph method. The probabilistic interpretation of the stability states of stope surfaces eliminates the false feeling of absolute stope performance based on its location on the Stability Graph , as implied by the deterministic approach.

58 GEOSCIENCES↗

Sharing is caring: An extensive analysis of parameter-based transfer learning for the prediction of building thermal dynamics

In recent years deep neural networks have been proposed as a lightweight data-driven model to capture high-dimensional, nonlinear physical processes to predict building thermal responses. However, the need of a large amount of data for the training process of deep neural networks clashes with the potential limited data availability in most existing or new buildings. Transfer learning aims to enhance the performance of a target learner exploiting knowledge from related and similar environments. This study conducted a suite of experiments that leveraged 250 data-driven models based on a synthetic dataset of a building archetype to study the influence of data availability, energy efficiency level, occupancy and climate for the transfer process of thermal dynamics. The performance of the transfer learning process was compared against a classical machine learning approach. Here, the results suggest that building thermal dynamics can be effectively transferred under the same climatic conditions, increasing performance when dealing with different occupancy schedules, efficiency levels and low data availability. Furthermore, the paper compares the performance of both transfer learning and machine learning approaches in an online fashion, to support the implementation in real-world deployment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Tracking Soiling Losses: Assessment, Uncertainty, and Challenges in Mapping

Several models have been presented in the recent years to estimate the magnitude of soiling from environmental parameters. However, these models are often based on data from a single site, or at most a few sites, and only limited data are, as of yet, available on their uncertainty. The present work aims to present a first comparative analysis of soiling estimation models, using measured soiling data from various locations in the USA. The study also investigates the impact that the source of the input data can have on the estimation. The results show that the model selection is only one of the factors that can affect the evaluation. Indeed, the use of satellite-derived or ground-mounted particulate matter data can lead to the generation of different soiling maps, with factors greater than 2x between the modeled losses. The current challenges and the unanswered questions that can bias soiling estimation are discussed. Additionally, potential research directions to improve the quality of soiling modeling are identified.

14 SOLAR ENERGY↗

Identifying Adversarial Cyber-Activity in Operational Technology Environments Using Bayesian Networks

Critical infrastructure and other operational technology (OT) environments face increasing cybersecurity risks from adversarial behavior. This paper describes the development of a risk model using a Bayesian network to enhance the comprehension of observable cyber events caused by malicious activity in OT environments. The core of the Bayesian network is a process model that describes the stages of adversary behavior. The remainder of the model is based on the MITRE ATT&CK® for Industrial Control Systems (ICS) taxonomy, which includes tactics and techniques that may be used by the adversary. The observables provide evidence for adversary behavior through the intermediary technique and tactic nodes. One challenge in constructing this model is a lack of open-source data from cyber-attacks on OT systems. This paper discusses learning from limited data, the elicitation of expert opinion to construct the conditional probability tables when data is scarce, and the refinement of the most difficult conditional probabilities tables using several forms of sensitivity analyses. Finally, the Bayesian network is demonstrated using two historical case studies: the DarkSide ransomware attack on the Colonial Pipeline and the destructive cyberattack targeting the ThyssenKrupp blast furnace. Index Terms—Cybersecurity, industrial control systems, operational technology

97 - MATHEMATICS AND COMPUTING↗

Analysis of Oil and Gas Ethane and Methane Emissions in the Southcentral and Eastern United States Using Four Seasons of Continuous Aircraft Ethane Measurements

In the last decade, much work has been done to better understand methane (CH 4 ) emissions from the oil and gas (O&G) industry in the United States. Ethane (C 2 H 6 ), a gas that is co-emitted with thermogenic sources of CH 4 , is emitted in the US predominantly by the O&G sector. Here, in this study, we perform an inverse analysis on 200 h of atmospheric boundary layer C 2 H 6 measurements to estimate C 2 H 6 emissions from the US O&G sector. Measurements were collected from 2017 to 2019 as part of the Atmospheric Carbon and Transport (ACT) America aircraft campaign and encompass much of the central and eastern United States. We find that for the fall, winter, and spring campaigns, C 2 H 6 data consistently exceeds values that would be expected based on EPA O&G leak rate estimates by more than 50%. C 2 H 6 observations from the summer 2019 data set show significantly lower C 2 H 6 enhancements in the southcentral region that cannot be reconciled with data from the other three seasons, either due to complex meteorological conditions or a temporal shift in the emissions. Combining the fall, winter, and spring C 2 H 6 posterior emissions estimate to an inventory of O&G CH 4 emissions, we estimate that O&G CH 4 emissions are larger than EPA inventory values by 48%–76%. Uncertainties in the gas composition data limit the accuracy of using C 2 H 6 as a proxy for O&G CH 4 emissions. These limits could be resolved retroactively by increasing the availability of industry-collected gas composition data.

54 ENVIRONMENTAL SCIENCES↗

Residuals-based distributionally robust optimization with covariate information

We consider data-driven approaches that integrate a machine learning prediction model within distributionally robust optimization (DRO) given limited joint observations of uncertain parameters and covariates. Our framework is flexible in the sense that it can accommodate a variety of regression setups and DRO ambiguity sets. We investigate asymptotic and finite sample properties of solutions obtained using Wasserstein, sample robust optimization, and phi-divergence-based ambiguity sets within our DRO formulations, and explore cross-validation approaches for sizing these ambiguity sets. Through numerical experiments, we validate our theoretical results, study the effectiveness of our approaches for sizing ambiguity sets, and illustrate the benefits of our DRO formulations in the limited data regime even when the prediction model is misspecified.

97 MATHEMATICS AND COMPUTING↗

Characterizing manufacturing wastewater in the United States for the purpose of analyzing energy requirements for reuse

This paper seeks to inform an improved understanding of the energy tradeoff associated with on-site manufacturing water reuse in the United States from a lifecycle perspective, in part by developing an analytical framework for understanding when this tradeoff for reuse is beneficial. We survey the literature to assess the current state of reuse and its motives and barriers in the United States, before synthesizing information from publicly available EPA data on contaminants in US manufacturing wastewaters and technologies for treating them. Using the available data, we derive a set of “ubiquitous contaminants” among the top ten in terms of mass discharged in more than half of US manufacturing subsectors (NAICS 31–33) according to EPA permit data. We also present information on proven treatment trains and their energy requirements. We then compare water quality requirements for specific contaminants in reclaimed water to those characteristic of wastewater streams currently being discharged from manufacturing plants into surface waters to highlight sectors with reuse opportunities that could require little cost to realize, such as primary metals and, to a lesser extent, petroleum and coal products. We conclude by highlighting data limitations that need to be rectified before applying the framework more broadly and discussing how these data gaps could be filled. Better understanding the relationship between energy and water in the context of on-site manufacturing water reuse would allow manufacturers to improve resiliency by reducing regulatory, physical, and reputational risks while lessening their footprint on local watersheds.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

System Configuration Evaluation for Process Settling of Hanford Waste Solid Particles

Direct Feed High-Level Waste (DFHLW) is a potential flowsheet operations approach to initiating high-level waste (HLW) vitrification prior to completion of the Hanford Waste Treatment and Immobilization Plant (WTP) Pretreatment Facility. A settle/decant process has been proposed to concentrate solids prior to delivery to the WTP HLW Facility during DFHLW operations, wherein the solids in a settled layer would be remixed with the supernatant liquid remaining after decanting operations to provide the feed at required solids concentrations. Settling would be used in lieu of purpose-built filtration or other solids separation equipment. Pacific Northwest National Laboratory (PNNL) is providing baseline technical support to the Washington River Protection Solutions (WRPS) Flowsheet Integration group. To support planning for DFHLW, WRPS previously requested that PNNL evaluate the current data set available to predict the time needed for HLW solids to settle and the solids concentration and strength of that settled layer, to identify gaps in the understanding and predictive capability of HLW solids waste settling times, and to provide scoping estimates of the potential settling times. Eight technical gaps were identified for predicting settling times and characteristics of the formed sediment layers. In addition to the data gaps, an overarching observation was made that there is significant variation in behavior of settling rate and settled layer data. The settling time required to concentrate solids via a settle/decant process was determined from the limited data to have a difference of potentially more than a factor of 5,000 in the estimated settling times, varying from 0.2 to 1,060 days for example depending on process vessel depth and final sediment solids concentration. In contrast, successful processes of liquid-forward output streams resulting from in-tank settling and decanting forward liquid have been reported for operations conducted at the Hanford Site. The purpose of this current report is to further support DFHLW planning by evaluating double-shell tank (DST) and alternate vessel equipment and operational configurations to enable optimization of the settle/decant process to concentrate solids. Hanford waste processing behavior specific to liquid feed availability following a slurry transfer in a DST is summarized, including process stream characteristics and process equipment configurations. The performance of DST process equipment configurations is evaluated for possible improvements using computational fluid dynamics (CFD) and simple analytical models. Potential new vessel design(s) specific to enabling effective settle/decant processes, and cursory summary of other separate and inline solids separations processes, are also provided. The CFD results indicated that improvement in outflow solids concentration was promoted by a reduction in the slurry flow rate, angling the distributor nozzles downward, and lifting the transfer pump. The solid-liquid analysis evaluating particle trajectory confirmed that the potential for particle ingestion (in the transfer pump) was decreased with increased radial separation between the inlet and outlet (transfer pump inlet), decreased inlet flow, and decreased liquid density and viscosity for a neutrally buoyant inlet flow. An assessment was also made of the potential for inflow configuration changes to result in the discrete mounding or piling of solids within the tank. Based on the characterization of the settled waste to date, HLW sediments will be unlikely to sustain a substantial angle of repose to facilitate significant variations in the elevation of the settled solids.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Residuals-based distributionally robust optimization with covariate information

We consider data-driven approaches that integrate a machine learning prediction model within distributionally robust optimization (DRO) given limited joint observations of uncertain parameters and covariates. Our framework is flexible in the sense that it can accommodate a variety of regression setups and DRO ambiguity sets. We investigate asymptotic and finite sample properties of solutions obtained using Wasserstein, sample robust optimization, and phi-divergence-based ambiguity sets within our DRO formulations, and explore cross-validation approaches for sizing these ambiguity sets. Through numerical experiments, we validate our theoretical results, study the effectiveness of our approaches for sizing ambiguity sets, and illustrate the benefits of our DRO formulations in the limited data regime even when the prediction model is misspecified.

97 MATHEMATICS AND COMPUTING↗