Optimizing Point-in-Space Continuous Monitoring System Sensor Placement on Oil and Gas Sites
Not Available
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
Abstract We show analytically that training a neural network by conditioned stochastic mutation or neuroevolution of its weights is equivalent, in the limit of small mutations, to gradient descent on the loss function in the presence of Gaussian white noise. Averaged over independent realizations of the learning process, neuroevolution is equivalent to gradient descent on the loss function. We use numerical simulation to show that this correspondence can be observed for finite mutations, for shallow and deep neural networks. Our results provide a connection between two families of neural-network training methods that are usually considered to be fundamentally different.
Single-molecule stretching experiments are widely utilized within the fields of physics and chemistry to characterize the mechanics of individual bonds or molecules, as well as chemical reactions. Analytic relations describing these experiments are valuable, and these relations can be obtained through the statistical thermodynamics of idealized model systems representing the experiments. Since the specific thermodynamic ensembles manifested by the experiments affect the outcome, primarily for small molecules, the stretching device must be included in the idealized model system. Though the model for the stretched molecule might be exactly solvable, including the device in the model often prevents analytic solutions. In the limit of large or small device stiffness, the isometric or isotensional ensembles can provide effective approximations, but the device effects are missing. Here a dual set of asymptotically correct statistical thermodynamic theories are applied to develop accurate approximations for the full model system that includes both the molecule and the device. In conclusion, the asymptotic theories are first demonstrated to be accurate using the freely jointed chain model and then using molecular dynamics calculations of a single polyethylene chain.
Reliable real-time monitoring is valuable for maintaining the operational integrity of modern electrical smart grids. Deployment of heterogeneous sensing technologies in substations has enabled high-resolution, multichannel waveform monitoring, but also introduces challenges for anomaly detection due to noise, baseline drift, and modality-dependent signal characteristics. In this work, we present a computationally efficient unsupervised method for multimodal event detection based on Rolling Root Mean Square based Event Detection (RRMSED). The method is developed using in-house, field deployed sensors collecting data at a utility substation. The sensing system comprises voltage and current sensors, triaxial accelerometers, and magnetometers, collectively capturing electrical, vibrational, and magnetic waveform measurements at high temporal resolution. RRMSED operates by extracting rolling RMS energy features and their first-order temporal differences from consecutive waveform segments for each channel and then applying channel-specific statistical thresholds learned from historical data. A persistence-based exceedance logic is employed to robustly identify transient events while suppressing impulsive noise, and to provide precise temporal localization with high resolution. The framework is designed for continuous server-side operation and can be deployed in real time without requiring complex models. Experiments on simulated waveform data with known ground truth demonstrate low false positive (FP) and false negative (FN) rates. Application to real substation data shows RRMSED to identify events that are not captured by conventional monitoring indicators including fast transient detection algorithm currently deployed in the system. These results indicate that rolling RMS based features provide an effective and practical basis for real-time multimodal event detection in smart-grid substations.
Aerosols serve as cloud condensation nuclei, shaping the microphysical properties of cloud droplets. Aerosol effects on convective clouds are complex and remain controversial. The debate centers around the process of aerosol-induced invigoration of deep convection, a phenomenon that could significantly affect convective cloud properties but lacks robust evidence due to methodological limitations in observational approaches and questions about the robustness of modeling studies. Resolving these discrepancies is crucial for understanding how aerosols affect the atmosphere. Here, this study examines the effects of meteorological and aerosol parameters in a weakly synoptic-driven convective environment, where the influence of aerosols may be more pronounced and observable. Daily atmospheric soundings and aerosol concentrations from several ground instruments collected during the summer of 2022 in Houston, Texas, as part of the Tracking Aerosol Convection interactions Experiment (TRACER) and Experiment of Sea Breeze Convection, Aerosols, Precipitation, and Environment (ESCAPE) field campaigns are analyzed. Statistical learning methods are applied to uncover the complex relationships between aerosols, meteorology, and convective cloud characteristics, such as cell area and echo-top height. The findings reveal that higher aerosol concentrations are associated with narrower convective cells, which we argue contradicts the idea of stronger convection with increased aerosol loading. However, once the data are clustered by the synoptic environment, the relationship between aerosol loading and convective cell area diminishes, indicating that the covariablity between synoptic-scale weather patterns, local thermodynamics, and aerosol loading makes it challenging to draw definitive conclusions about the specific impacts of aerosols on convective cloud properties.
Independent component analysis is an unsupervised machine learning algorithm that separates a set of mixed signals into a set of statistically independent source signals. Applied to high-quality gene expression datasets, independent component analysis effectively reveals both the source signals of the transcriptome as co-regulated gene sets, and the activity levels of the underlying regulators across diverse experimental conditions. Two major variables that affect the final gene sets are the diversity of the expression profiles contained in the underlying data, and the user-defined number of independent components, or dimensionality, to compute. Availability of high-quality transcriptomic datasets has grown exponentially as high-throughput technologies have advanced; however, optimal dimensionality selection remains an open question. We computed independent components across a range of dimensionalities for four gene expression datasets with varying dimensions (both in terms of number of genes and number of samples). We computed the correlation between independent components across different dimensionalities to understand how the overall structure evolves as the number of user-defined components increases. We then measured how well the resulting gene clusters reflected known regulatory mechanisms, and developed a set of metrics to assess the accuracy of the decomposition at a given dimension. We found that over-decomposition results in many independent components dominated by a single gene, whereas under-decomposition results in independent components that poorly capture the known regulatory structure. From these results, we developed a new method, called OptICA, for finding the optimal dimensionality that controls for both over- and under-decomposition. Specifically, OptICA selects the highest dimension that produces a low number of components that are dominated by a single gene. We show that OptICA outperforms two previously proposed methods for selecting the number of independent components across four transcriptomic databases of varying sizes. OptICA avoids both over-decomposition and under-decomposition of transcriptomic datasets resulting in the best representation of the organism’s underlying transcriptional regulatory network.
Oklahoma and Kansas experienced unprecedented seismic activity over the past decade due to earthquakes associated with unconventional hydrocarbon development. The modest natural seismicity and incomplete knowledge of the fault network in the region made it difficult to anticipate the locations of earthquakes with larger magnitudes (M w ≥ 4). Here, we show that monitoring of microearthquakes at regional scale using a pretrained neural phase picker and an earthquake relocation algorithm can illuminate unknown fault structures, and deliver information that can be synthesized for earthquake forecasting. We found that 80% of the larger earthquakes that occurred in the past decade could have been anticipated based on the spatial extent of the seismicity clusters that were formed before these earthquakes occurred. We also found that once a seismicity cluster with a length scale enough to host a larger earthquake was formed, there was a ~5% chance that it would host one or more larger earthquakes within a year. This probability is nearly an order of magnitude higher than one based on Gutenberg–Richter statistics and preceding seismicity. Applying our approach in practice can provide critical information on seismic hazards for risk management and regulatory decision making.
Abstract After decades, the theoretical study of core-collapse supernova explosions is moving from parameterized, spherically symmetric models to increasingly realistic multidimensional simulations. However, obtaining nucleosynthesis yields based on such multidimensional core-collapse supernova simulations is not straightforward. Frequently, tracer particles are employed. Tracer particles may be tracked in situ during the simulation, but often they are reconstructed in a post-processing step based on the information saved during the hydrodynamic simulation. Reconstruction can be done in a number of ways, and here we compare the approaches of backward and forward integration of the equations of motion to the results based on inline particle trajectories. We find that both methods agree reasonably well with the inline results for isotopes for which a large number of particles contribute. However, for rarer isotopes that are produced only by a small number of particle trajectories, deviations can be large. For our setup, we find that backward integration leads to better agreement with the inline particles by more accurately reproducing the conditions following freeze-out from nuclear statistical equilibrium, because the establishment of nuclear statistical equilibrium erases the need for detailed trajectories at earlier times. Based on our results, if inline tracers are unavailable, we recommend backward reconstruction to the point when nuclear statistical equilibrium was last applied, with an interval between simulation snapshots of at most 1 ms for nucleosynthesis post-processing.
Hydrologic variability is a serious threat to the poverty-stricken regions of North Africa. Meanwhile, the scientific community struggles to attribute these extreme climatic episodes to specific oceanic and land drivers. Prior modeling studies have not assessed simulated feedbacks over North Africa against an established observational benchmark, so their results are considered as untested, model-specific findings. The Coupled Model Intercomparison Project Phase Five (CMIP5) archive represents the state-of-the-art in climate projections, with most CMIP5 models now containing interactive vegetation phenology; however, there have been no studies to date of their simulated vegetation feedbacks. The capability of CMIP5 models at accurately simulating the forcing of both sea-surface temperature (SST) and regional leaf area index (LAI) anomalies on North African climate needs to be a key consideration in assessing the models’ overall credibility and determining appropriate model weighting for developing climate projections. The Generalized Equilibrium Feedback Assessment (GEFA) is a promising statistical method, which can be applied to either model output or observations, for isolating the local and remote impacts of individual oceanic or terrestrial forcings on regional climate. We performed a combined observational and modeling assessment of land-ocean-atmosphere interactions across the distinct ecological and moisture gradients of North Africa. We evaluated and demonstrated the reliability of the GEFA statistical method over North Africa using the Community Earth System Model (CESM), applied GEFA to observational data to quantify the observed forcing of ocean basin SST anomalies and regional LAI anomalies on North African climate, evaluated the CMIP5 models’ performance in terms of representing these key observed feedbacks, and formulated CMIP5 feedback performance metrics for weighting North African climate projections. Our study represented the first attempt to separate the observed roles of oceanic and vegetation feedbacks across North Africa, the first systematic assessment and intercomparison of land-ocean-atmosphere feedbacks in CMIP5, and the first exploration of vegetation feedbacks among CMIP5 models. The following work was accomplished. (1) The team generated 14 scientific publications and 30 presentations based on the DOE-funded research. (2) A stepwise version of the GEFA statistical method was developed in which unimportant forcings were dropped to increase the reliability of the results. (3) SGEFA was successfully validated through experiments with CESM in terms of its ability to isolate the atmospheric responses to individual oceanic or land forcings. (4) Observational evidence was revealed for the Sahel’s positive vegetation-rainfall feedback on the seasonal to interannual time scale, and it was attributed to a moisture recycling mechanism rather than an albedo mechanism. (5) An approach with SGEFA was developed in which the individual contributions of soil moisture versus vegetation forcings could be separated, revealing that the former forcing outweighed the latter for sub-Saharan Africa. (6) Tropical ocean temperatures were found to be key regulators of pan-tropical vegetation variability, especially for arid and semi-arid regions, including sub-Saharan Africa. (7) The CMIP5 models largely underestimated the importance of land feedbacks across the Sahel. (8) The general consensus among CMIP5 models indicates, for the late 21st century, a diminished seasonal predictability of sub-Saharan African regional climate and an elevated role of the land surface compared to oceanic drivers in regulating regional climate variability. (9) Land-ocean-atmosphere interactions were demonstrated to be key contributors to the seasonal predictability of African wildfire activity. (10) The El Djouf was determined to be of greater importance than the Bodélé depression in terms of providing trans-Atlantic dust transport to the Americas.
A method of applying Principal Component Analysis, Soft Independent Modeling of Class Analysis, and statistical analysis is described that can be applied to many types of testers to ascertain how well matched the performance of the testers in the analysis are to one another or how well matched a tester is to itself at a later time. This method is most useful for situations for which the same units have not been run across the testers being analyzed for matched performance.
Within the Partnership Center for High-Fidelity Boundary Plasma Simulation (HBPS), work at UT-Austin was aimed at improved verification, validation, and uncertainty quantification (VVUQ) for edge plasma simulations and on performing gyrokinetics simulations of pedestal instabilities and turbulence in order to expand foundational understanding of pedestal transport. Regarding VVUQ, the accomplishments can be summarized as follows. First, it was shown that the Moment Preserving Constrained Resampling technique, when applied periodically in particle-in-cell simulations in the XGC code, can dramatically improve the accuracy of the simulation at essentially equivalent computational cost. Second, a technique for estimating model correlations, which are required to solve the model selection and sample allocation problem in multifidelity UQ techniques, without sampling the highest fidelity, most computationally expensive model, was developed and demonstrated. Third, previously developed methods for estimating statistical and discretization errors were applied to numerical methods relevant to edge plasma simulations, namely in particle-in-cell-based approaches, and shown to work. Finally, benchmark studies for comparing gyrokinetic codes were developed and performed, leading to reasonable agreement between four commonly used codes. Regarding physics studies, gyrokinetic simulations to investigate microtearing modes in the DIII-D pedestal were performed using the GENE code.
A combinatorial approach has been applied to the allowable permutations of quantum electronic configurations under the constraints of Hund's rule for established ground state configurations toward an under-approximation of electronic structure entropy. Combined with a previously reported over-approximation, the approximations are used in conjunction in an attempt to bracket the upper and lower entropy limits for multiconfigurational ground state electronic structure entropy and compared to known standard molar entropies for the elements. This formality has been used for the application of a classical statistical mechanics methodology to be applied to the discrete sets of quantum mechanical states of Pu in order to calculate orbital occupancies in Pu's multiconfigurational ground state. Without consideration of the relative energies of various possible electronic configurations contributing to the multiconfigurational ground state, the calculations are performed under a general energy degeneracy assumption weighted to the number of permutations for specific configurations. The number of configurations assumed to significantly contribute is gradually constrained in order to approach a low-order approximation of orbital occupancies in Pu that are then compared to experimental and other calculated results from the literature.
Abstract If quantum information processors are to fulfill their potential, the diverse errors that affect them must be understood and suppressed. But errors typically fluctuate over time, and the most widely used tools for characterizing them assume static error modes and rates. This mismatch can cause unheralded failures, misidentified error modes, and wasted experimental effort. Here, we demonstrate a spectral analysis technique for resolving time dependence in quantum processors. Our method is fast, simple, and statistically sound. It can be applied to time-series data from any quantum processor experiment. We use data from simulations and trapped-ion qubit experiments to show how our method can resolve time dependence when applied to popular characterization protocols, including randomized benchmarking, gate set tomography, and Ramsey spectroscopy. In the experiments, we detect instability and localize its source, implement drift control techniques to compensate for this instability, and then demonstrate that the instability has been suppressed.
Abstract With the advent of the big data era, the need to combine multiple individual data sets to draw causal effects arises naturally in many medical and biological applications. Especially each data set cannot measure enough confounders to infer the causal effect of an exposure on an outcome. In this article, we extend the method proposed by a previous study to causal data fusion of more than two data sets without external validation and to a more general (continuous or discrete) exposure and outcome. Theoretically, we obtain the condition for identifiability of exposure effects using multiple individual data sources for the continuous or discrete exposure and outcome. The simulation results show that our proposed causal data fusion method has unbiased causal effect estimate and higher precision than traditional regression, meta‐analysis and statistical matching methods. We further apply our method to study the causal effect of BMI on glucose level in individuals with diabetes by combining two data sets. Our method is essential for causal data fusion and provides important insights into the ongoing discourse on the empirical analysis of merging multiple individual data sources.
Quantum chromodynamics is the theory of the strong interaction between quarks and gluons; the coupling strength of the interaction, α S , is the least precisely-known of all interactions in nature. An extraction of the strong coupling from the radiation pattern within jets would provide a complementary approach to conventional extractions from jet production rates and hadronic event shapes, and would be a key achievement of jet substructure at the Large Hadron Collider (LHC). Presently, the relative fraction of quark and gluon jets in a sample is the limiting factor in such extractions, as this fraction is degenerate with the value of αS for the most well-understood observables. To overcome this limitation, we apply recently proposed techniques to statistically demix multiple mixtures of jets and obtain purified quark and gluon distributions based on an operational definiton. We illustrate that studying quark and gluon jet substructure separately can significantly improve the sensitivity of such extractions of the strong coupling. We also discuss how using machine learning techniques or infrared- and collinear-unsafe information can improve the demixing performance without the loss of theoretical control. While theoretical research is required to connect the extract topics with the quark and gluon objects in cross section calculations, our study illustrates the potential of demixing to reduce the dominant uncertainty for the α S extraction from jet substructure at the LHC.
Natural refrigerants are increasingly adopted in next-generation heat pump systems, among which CO₂ heat pumps have attracted significant attention. However, due to their high operating pressures, the leakage risk is higher, resulting in undercharge conditions and degraded heat pump performance. Thus, developing an accurate refrigerant charge level detection technique is necessary to guarantee safe and efficient operation. Although virtual refrigerant charge (VRC) level calculation algorithms for CO₂ heat pumps exist, they typically rely on empirically selected features without a systematic selection framework, leading to multicollinearity and potential overfitting, which limit their prediction accuracy and generalizability. To address these issues, this study proposes a VRC algorithm framework with a systematic feature selection method that identifies physically meaningful and statistically significant features, and is applied using a residential CO₂ heat pump as a case study. The method is extended from previous work on conventional refrigerants to account for charge behavior in CO₂ gas coolers. The selected features include gas cooler outlet density, evaporator pressure, and superheat temperature. The results demonstrate that the proposed feature selection method significantly improves prediction accuracy compared to existing VRC approaches. A relatively small training dataset (∼30 samples) is sufficient for feature identification and model development. The developed algorithm achieves less than 3% prediction error under both undercharge and overcharge conditions, representing reductions of 46.7% and 35.3% compared to two recent reference VRC algorithms for transcritical CO₂ heat pumps reported in the literature. The proposed algorithm and feature selection method enhance leakage detection capability, facilitate the deployment of CO₂ heat pump systems, and contribute to reduced energy waste and maintenance costs.
The complex spatial and temporal structure of cumulus clouds complicates their representation in weather and climate models. Classic meteorological instrumentation struggles to fully capture these features. Networks of multiple high-resolution hemispheric cameras are increasingly used to fill this data gap, and provide information on this missing multi-dimensional spatial information. In this study, a path-tracing algorithm is used to generate virtual camera images of resolved clouds in large-eddy simulations (LES). These images are then used as a camera network simulator, allowing reconstructions of three-dimensional cloud edges from the model output. Because the actual LES cloud field is fully known, the combined path-tracing and reconstruction method can be statistically analyzed. The method is applied to LES realizations of summertime shallow cumulus at the Jülich Observatory for Cloud Evolution (JOYCE), Germany, which also routinely operates a camera network. We find that the path-tracing method allows accurate reconstruction of up to 70% of the visible cloud edges. Additional sensitivity tests show that the method is robust for changes in its hyperparameters. The sensitivity to cloud optical thickness is also investigated, finding a cloud boundary placement error of approximately 182 m. This error can be considered typical for cloud boundary reconstruction using real stereo camera imagery. The results provide proof of principle for future use of the method for evaluating LES clouds against camera network imagery, and for further optimizing the configuration of such camera networks.
Accurate 3D representations of lithium-ion battery electrodes, in which the active particles, binder and pore phases are distinguished and labeled, can assist in understanding and ultimately improving battery performance. Here, we demonstrate a methodology for using deep-learning tools to achieve reliable segmentations of volumetric images of electrodes on which standard segmentation approaches fail due to insufficient contrast. We implement the 3D U-Net architecture for segmentation, and, to overcome the limitations of training data obtained experimentally through imaging, we show how synthetic learning data, consisting of realistic artificial electrode structures and their tomographic reconstructions, can be generated and used to enhance network performance. We apply our method to segment x-ray tomographic microscopy images of graphite-silicon composite electrodes and show it is accurate across standard metrics. We then apply it to obtain a statistically meaningful analysis of the microstructural evolution of the carbon-black and binder domain during battery operation.