Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Distribution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Bayesian Inference for the Seismic Moment Tensor Using Regional Waveforms and Teleseismic- P Polarities with a Data-Derived Distribution of Velocity Models and Source Locations

The largest source of uncertainty in any source inversion is the velocity model used in the transfer function that relates observed ground motion to the seismic moment tensor. However, standard inverse procedure often does not quantify uncertainty in the seismic moment tensor due to error in the Green’s functions from uncertain event location and Earth structure. Here, we incorporate this uncertainty into an estimation of the seismic moment tensor using a data-derived distribution of velocity models based on complementary geophysical data sets, including thickness constraints, velocity profiles, gravity data, surface-wave group velocities, and regional body-wave travel times. The data-derived distribution of velocity models is then used as a prior distribution of Green’s functions for use in Bayesian inference of an unknown seismic moment tensor using regional and teleseismic-P waveforms. The use of multiple data sets is important for gaining resolution to different components of the moment tensor. The combined likelihood is estimated using data-specific error models and the posterior of the seismic moment tensor is estimated and interpreted in terms of the most probable source type.

58 GEOSCIENCES↗

Solar Forecasting, Net Load Forecasting, and Data-Driven Distributed Solar Visibility Prizes (Final Technical Report)

The American-Made Solar Forecasting Prize, Net Load Forecasting Prize, and Data-Driven Distribution (3D) Solar Visibility Prize is a multimillion-dollar prize competition designed to energize U.S. solar innovation through a series of contests that accelerate the entrepreneurial process from years to months. The activities incentivized by these three prizes will support the governmentwide approach to increase American energy dominance by promoting innovation and early deployment of energy technologies, resulting in wider adoption, which is critical for secure, affordable, and reliable solar energy.

14 SOLAR ENERGY↗

Effectiveness of denoising diffusion probabilistic models for fast and high-fidelity whole-event simulation in high-energy heavy-ion experiments

Artificial intelligence (AI) generative models, such as generative adversarial networks (GANs), variational autoencoders, and normalizing flows, have been widely used and studied as efficient alternatives for traditional scientific simulations. However, they have several drawbacks, including training instability and inability to cover the entire data distribution, especially for regions where data are rare. This is particularly challenging for whole-event, full-detector simulations in high-energy heavy-ion experiments, such as sPHENIX at the Relativistic Heavy Ion Collider and Large Hadron Collider experiments, where thousands of particles are produced per event and interact with the detector. This work investigates the effectiveness of denoising diffusion probabilistic models (DDPMs) as an AI-based generative surrogate model for the sPHENIX experiment that includes the heavy-ion event generation and response of the entire calorimeter stack. DDPM performance in sPHENIX simulation data is compared with a popular rival, GANs. Results show that both DDPMs and GANs can reproduce the data distribution where the examples are abundant (low-to-medium calorimeter energies). Nonetheless, DDPMs significantly outperform GANs, especially in high-energy regions where data are rare. Additionally, DDPMs exhibit superior stability compared to GANs. The results are consistent between both central and peripheral centrality heavy-ion collision events. Moreover, DDPMs offer a substantial speedup of approximately a factor of 100 compared to the traditional Geant4 simulation method.

42 ENGINEERING↗

Long–short-term memory encoder–decoder with regularized hidden dynamics for fault detection in industrial processes

The ability of recurrent neural networks (RNN) to model nonlinear dynamics of high dimensional process data has enabled data-driven RNN-based fault detection algorithms. Previous studies have focused on detecting faults by identifying the discrepancies in data distribution between the faulty and normal data, as reflected in prediction errors generated by RNN models. However, in industrial processes, variations in data distribution can also result from changes in normal control setpoints and compensatory control adjustments in response to disturbances, making it hard to differentiate between normal and faulty conditions. This paper proposes a fault detection method utilizing a long short-term memory (LSTM) encoder–decoder structure with regularized hidden dynamics and reversible instance normalization (RevIN) to compactly represent high-dimensional measurements for effective monitoring. During training, the hidden states of the model are regularized to form a low-dimensional latent space representation of the original multivariate time series data. As a result, the prediction errors of the latent states can be used to monitor the abnormal dynamic variations, while the reconstruction errors of the measured variables are used to monitor the abnormal static variations. Furthermore, the proposed indices can reflect operating conditions, even when the distribution of test data changes, which helps distinguish faults from normal adjustments and disturbances that controllers can settle. Here, data from numerical simulation and the Tennessee Eastman process are used to illustrate the effectiveness of the proposed fault detection method.

42 ENGINEERING↗

Magnetic pair distribution function data using polarized neutrons and ad hoc corrections

Here, we report the first example of magnetic pair distribution function (mPDF) data obtained through the use of neutron polarization analysis. Using the antiferromagnetic semiconductor MnTe as a test case, we present high-quality mPDF data collected on the HYSPEC instrument at the Spallation Neutron Source using longitudinal polarization analysis to isolate the magnetic scattering cross section. Clean mPDF patterns are obtained for MnTe in both the magnetically ordered state and the correlated paramagnet state, where only short-range magnetic order is present. We also demonstrate significant improvement in the quality of high-resolution mPDF data through the application of ad hoc corrections that require only minimal human input, minimizing potential sources of error in the data processing procedure. We briefly discuss the current limitations and future outlook of mPDF analysis using polarized neutrons. Overall, this work provides a useful benchmark for mPDF analysis using polarized neutrons and provides an encouraging picture of the potential for routine collection of high-quality mPDF data.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Bayesian inference for the seismic moment tensor using regional waveforms and a data-derived distribution of velocity models

The largest source of uncertainty in any source inversion is the velocity model used to construct the transfer function employed in the forward model that relates observed ground motion to the seismic moment tensor. However, standard inverse procedures often does not quantify uncertainty in the seismic moment tensor due to error in the Green’s functions from uncertain event location and Earth structure. We attempt to incorporate this uncertainty into an estimation of the seismic moment tensor using a distribution of velocity models calculated in a prior effort based on different and complementary data sets. The posterior distribution of velocity models is then used to construct Green’s functions for use in Bayesian inference of an unknown seismic moment tensor using regional waveform data. The combined likelihood is estimated using data-specific error models and the posterior of the seismic moment tensor is estimated and can be interpreted in terms of most-probable source-type.

58 GEOSCIENCES↗

Reliable Measures of Spread in High Dimensional Latent Spaces

Understanding geometric properties of the latent spaces of natural language processing models allows the manipulation of these properties for improved performance on downstream tasks. One such property is the amount of data spread in a model’s latent space, or how fully the available latent space is being used. We demonstrate that the commonly used measures of data spread, average cosine similarity and a partition function min/max ratio I (V), do not provide reliable metrics to compare the use of latent space across data distributions. We propose and examine six alternative measures of data spread, all of which improve over these current metrics when applied to seven synthetic data distributions. Of our proposed measures, we recommend one principal component-based measure and one entropy-based measure that provide reliable, relative measures of spread and can be used to compare models of different sizes and dimensionalities.

97 MATHEMATICS AND COMPUTING↗

Methods and systems for evaluating data transportability in distribution grids

This document describes systems and techniques for evaluating and improving distribution-grid data transportability. These systems and techniques allow engineers to quantify the data transportability of a communication system within or connected to a distribution grid, which represents an ability to transport in real-time telemetry from source locations (e.g., sensors in the distribution grid) to control mechanisms. Distribution engineers can use the sensor readings to perform grid analytics, control operating parameters, and operate protection systems. Distribution engineers can also use the transportability of the communication system to evaluate the observability of the distribution grid, which represents an ability to combine actual measurements and various types of computations (e.g., analytics, estimators, forecasters) from a system model. Distribution engineers can then generate a sensor allocation plan that indicates the number and location of sensors to maximize observability for a fixed sensor cost and/or minimize sensor cost for predetermined observability.

Taft, Jeffrey D.↗

Refining perovskite structures to pair distribution function data using collective Glazer modes as a basis

Structural modelling of octahedral tilts in perovskites is typically carried out using the symmetry constraints of the resulting space group. In most cases, this introduces more degrees of freedom than those strictly necessary to describe only the octahedral tilts. It can therefore be a challenge to disentangle the octahedral tilts from other structural distortions such as cation displacements and octahedral distortions. This paper reports the development of constraints for modelling pure octahedral tilts and implementation of the constraints in diffpy-CMI, a powerful package to analyse pair distribution function (PDF) data. The model in the program allows features in the PDF that come from rigid tilts to be separated from non-rigid relaxations, providing an intuitive picture of the tilting. The model has many fewer refinable variables than the unconstrained space group fits and provides robust and stable refinements of the tilt components. It further demonstrates the use of the model on the canonical tilted perovskite CaTiO 3 which has the known Glazer tilt system α + β – β – . The Glazer model fits comparably to the corresponding space-group model Pnma below r = 14 Å and becomes progressively worse than the space-group model at higher r due to non-rigid distortions in the real material.

36 MATERIALS SCIENCE↗

EGS Collab Experiment 2: Distributed Fiber Optic Temperature Data (DTS)

Distributed fiber optic sensing was an important part of the monitoring system for EGS Collab Experiment #2. A single loop of custom fiber package was grouted into the four monitoring boreholes that bracketed the experiment volume. This fiber package contained two multi-mode fibers and four single-mode fibers. These fibers were connected to an array of fiber optic interrogator units, each targeting a different measurement. The distributed temperature system (DTS) consisted of a Silixa XT-DTS unit, connected to both ends of one of the two multi-mode fibers. This system measured absolute temperature along the entire length of fiber for the duration of the experiment at a sampling rate of approximately 10 minutes. This dataset includes both raw data in XML format from the XT-DTS, as well as a processed dataset with the sections of data pertaining only to the boreholes are extracted. We have also included a report that provides all of the relevant details necessary for users to process and interpret the data for themselves. Please read this accompanying report. If, after reading it, there are still outstanding questions, please do not hesitate to contact us. Happy processing.

15 GEOTHERMAL ENERGY↗

Considerations for Distributed Edge Data Centers and Use of Building Loads to Support Large Interconnections

The rapid expansion of artificial intelligence (AI) and machine learning is driving unprecedented electricity demand from data centers. It is predicted that by 2030, 90% of AI workloads will be inference-based, requiring interconnection of multiple low-latency edge data centers (<20 MW) sited closer to end users - often on already constrained distribution feeders. Although individually small, these loads can aggregate to large loads per feeder, straining infrastructure, creating multi-year interconnection delays, and driving up customer costs. This paper proposes a data center-focused grid-integration framework that combines feeder hosting capacity analysis with building energy efficiency, building load flexibility, and waste heat reuse to expand effective feeder and substation headroom. Such approaches can reduce interconnection delays, lower costs for ratepayers, and accelerate AI-ready infrastructure deployment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Merged aerosol size distribution from SMPS and OPC for SAIL

This dataset contains merged aerosol number size distribution data for the Surface Atmosphere Integrated Field Laboratory (SAIL) campaign. The merged size distribution data were constructed by combining measurements from a scanning-mobility particle sizer (SMPS) and an optical particle counter (OPC), covering a size range of 0.01–35 µm. The merging methodology follows the approach described by Hand and Kreidenweis (2002) and Marinescu et al. (2019). All aerosol data from the ARM archive were corrected to standard temperature (273.15 K) and pressure (101.3 kPa).

merged size distribution↗

Real-time and post-hoc compression for data from Distributed Acoustic Sensing

Distributed Acoustic Sensing (DAS) is an emerging sensing technology that records the strain-rate along fiber optic cables at high spatial and temporal resolution. This technique is becoming a popular tool in seismology, hydrology, and other subsurface monitoring applications. However, due to the large coverage (10’s of km) and high density of measurements (1m spacing at 100’s of Hz), a DAS installation could produce terabytes of data records per day. Because many DAS instruments are deployed in remote locations, this large data size poses significant challenges to its transfer and storage. In this paper, we explore lossless compression methods to reduce the storage requirement in both real-time and post-hoc scenarios. Here we propose a two-stage compression method to improve the compression ratio and compression speed. This two-stage compression method could reduce the storage requirement by 40%, which is 20% more than other lossless methods, such as ZSTD. We demonstrate that the compression method could complete its operation well before the DAS instrument needs to output the next file, making it suitable for real-time DAS acquisition. We also implement a parallel compression method for a post-hoc scenario and demonstrate that our method could effectively utilize a parallel computer. With 256 CPU cores, our parallel compression method achieves the speed of 26GB/second.

58 GEOSCIENCES↗

Operational Analytics Studies for ATLAS Distributed Computing: Data Popularity Forecast and Utilization of the WLCG Centers

Operational analytics is the direction of research related to the analysis of the current state of computing processes and the prediction of future states in order to anticipate imbalances and take timely measures to stabilize a complex system. There are two relevant areas in ATLAS Distributed Computing that are currently the focus of studies: user physics analysis including the forecast of popularity of data samples among users, and evaluating WLCG centers for their readiness to process user analysis payloads. Studying these areas is challenging due to the complexity involved, as it requires a comprehensive understanding of numerous boundary conditions typically found in large-scale distributed computing infrastructures. Forecasts of data popularity are problematic without the categorization of user tasks by their types (data transformation or physics analysis), which do not always appear on the surface but may induce noise, which introduces significant distortions for predictive analysis. Evaluating the WLCG resources by their analysis workloads is also a challenging task as it is necessary to find a balance between the workload of the resource, its performance, the waiting time for jobs on it, as well as the volume of jobs that it processes. This is especially difficult in a heterogeneous computing environment, where legacy resources are used along with modern high-performance machines. We will look at these areas of research in detail and discuss what tools and methods are used in our work, demonstrating results already obtained.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Extracting structural motifs from pair distribution function data of nanostructures using explainable machine learning

Characterization of material structure with X-ray or neutron scattering using e.g. Pair Distribution Function (PDF) analysis most often rely on refining a structure model against an experimental dataset. However, identifying a suitable model is often a bottleneck. Recently, automated approaches have made it possible to test thousands of models for each dataset, but these methods are computationally expensive and analysing the output, i.e. extracting structural information from the resulting fits in a meaningful way, is challenging. Our Machine Learning based Motif Extractor (ML-MotEx) trains an ML algorithm on thousands of fits, and uses SHAP (SHapley Additive exPlanation) values to identify which model features are important for the fit quality. We use the method for 4 different chemical systems, including disordered nanomaterials and clusters. ML-MotEx opens for a type of modelling where each feature in a model is assigned an importance value for the fit quality based on explainable ML.

36 MATERIALS SCIENCE↗

Outlook towards deployable continual learning for particle accelerators

Particle accelerators are high power complex machines. To ensure uninterrupted operation of these machines, thousands of pieces of equipment need to be synchronized, which requires addressing many challenges including design, optimization and control, anomaly detection and machine protection. With recent advancements, machine learning (ML) holds promise to assist in more advance prognostics, optimization, and control. While ML based solutions have been developed for several applications in particle accelerators, only few have reached deployment and even fewer to long term usage, due to particle accelerator data distribution drifts caused by changes in both measurable and non-measurable parameters. In this paper, we identify some of the key areas within particle accelerators where continual learning can allow maintenance of ML model performance with distribution drifts. Particularly, we first discuss existing applications of ML in particle accelerators, and their limitations due to distribution drift. Next, we review existing continual learning techniques and investigate their potential applications to address data distribution drifts in accelerators. By identifying the opportunities and challenges in applying continual learning, this paper seeks to open up the new field and inspire more research efforts towards deployable continual learning for particle accelerators.

43 PARTICLE ACCELERATORS↗

Data-driven Distributed Learning of Multi-agent Systems: A Koopman Operator Approach

Koopman operator theory provides a model-free technique for studying nonlinear dynamical systems purely from data. Since the Koopman operator is infinite-dimensional, researchers have developed several methods that provide a finite-dimensional approximation of the Koopman operator so that it can be applied for practical use cases. One common thing with most of the methods is that their solutions are obtained by solving a centralized minimization problem. In this work, we treat the dynamical system to be a multi-agent system and propose an algorithm to compute the finite-dimensional approximation of the Koopman operator in a distributed manner using the knowledge of the topology of the underlying multi-agent system. The proposed distributed approach is shown to be equivalent to the centralized learning problem and results in a sparse Koopman whose block structure mimics the Laplacian of the multi-agent system. Extensive simulation studies illustrate the proposed framework on the network of oscillators and the IEEE 68 bus system.

Nandanoori, Sai Pushpak↗