Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Activated relaxation in supercooled monodisperse atomic and polymeric WCA fluids: Simulation and ECNLE theory

Here, we combine simulation and Elastically Collective Nonlinear Langevin Equation (ECNLE) theory to study the activated relaxation in monodisperse atomic and polymeric Weeks–Chandler–Andersen (WCA) liquids over a wide range of temperatures and densities in the supercooled regime under isochoric conditions. By employing novel crystal-avoiding simulations, metastable equilibrium dynamics is probed in the absence of complications associated with size polydispersity. Based on a highly accurate structural input from integral equation theory, ECNLE theory is found to describe well the simulated density and temperature dependences of the alpha relaxation time of atomic fluids using a single system-specific parameter, a c , that reflects the nonuniversal relative importance of local cage and collective elastic barriers. For polymer fluids, the explicit dynamical effect of local chain connectivity is modeled at the fundamental dynamic free energy trajectory level based on a different parameter, N c , that quantifies the degree of intramolecular correlation of bonded segment activated barrier hopping. For the flexible chain model studied, a physically intuitive value of N c ≈ 2 results in good agreement between simulation and theory. A direct comparison between atomic and polymeric systems reveals that chain connectivity can speed up activated segmental relaxation due to weakening of equilibrium packing correlations but can slow down relaxation due to local bonding constraints. The empirical thermodynamic scaling idea for the alpha time is found to work well at high densities or temperatures but fails when both density and temperature are low. The rich and subtle behaviors revealed from simulation for atomic and polymeric WCA fluids are all well captured by ECNLE theory.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Propagation of partially spatially coherent laser beams in instantaneous Kerr media

The propagation of intense, partially spatially coherent laser beams in a medium with instantaneous third-order susceptibility is studied analytically and numerically. For sufficiently high power relative to that required for nonlinear self-focusing, the propagation initially proceeds in two stages. In the first stage, spatial coherence builds up, and in the second stage, the number of speckles reduces. Once the degree of coherence is sufficiently high, whole-beam self-focusing occurs. The beam power is mostly confined within the initial spot radius. Two analytical approaches for describing the evolution of the beam are presented. The method of moments leads to an analytical solution for the rms spot radius that is in excellent agreement with simulations. This method does not require any knowledge of the field statistics beyond the initial conditions and provides no information about the evolution of the individual speckles. The other approach employs a self-similar solution for the second-order coherence function of the field and assumes that the fourth-order coherence function is factorizable and obeys complex circular Gaussian random statistics. The latter method also leads to an analytical expression for the spot radius, but its predictions for the qualitative evolution of the speckles disagree with wave-optics simulations.

lasers↗

Random insights into the complexity of two-dimensional tensor network calculations

Projected entangled pair states (PEPS) offer memory-efficient representations of some quantum many-body states that obey an entanglement area law and are the basis for classical simulations of ground states in two-dimensional (2d) condensed matter systems. However, rigorous results show that exactly computing observables from a 2d PEPS state is generically a computationally hard problem. Yet approximation schemes for computing properties of 2d PEPS are regularly used, and empirically seen to succeed, for a large subclass of (“not too entangled”) condensed matter ground states. Adopting the philosophy of random matrix theory, in this work, we analyze the complexity of approximately contracting a 2d random PEPS by exploiting an analytic mapping to an effective replicated statistical mechanics model that permits a controlled analysis at a large bond dimension. Through this statistical-mechanics lens, we argue that (i) although approximately sampling wave-function amplitudes of random PEPS faces a computational-complexity phase transition above a critical bond dimension, and (ii) one can generically efficiently estimate the norm and correlation functions for any finite bond dimension. Furthermore, these results are supported numerically for various bond-dimension regimes. It is an important open question whether the above results for random PEPS apply more generally also to PEPS representing physically relevant ground states.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Automated pipeline framework for processing of large-scale building energy time series data

Commercial buildings account for one third of the total electricity consumption in the United States and a significant amount of this energy is wasted. Therefore, there is a need for “virtual” energy audits, to identify energy inefficiencies and their associated savings opportunities using methods that can be non-intrusive and automated for application to large populations of buildings. Here we demonstrate virtual energy audits applied to large populations of buildings’ time-series smart-meter data using a systematic approach and a fully automated Building Energy Analytics (BEA) Pipeline that unifies, cleans, stores and analyzes building energy datasets in a non-relational data warehouse for efficient insights and results. This BEA pipeline is based on a custom compute job scheduler for a high performance computing cluster to enable parallel processing of Slurm jobs. Within the analytics pipeline, we introduced a data qualification tool that enhances data quality by fixing common errors, while also detecting abnormalities in a building’s daily operation using hierarchical clustering. We analyze the HVAC scheduling of a population of 816 buildings, using this analytics pipeline, as part of a cross-sectional study. With our approach, this sample of 816 buildings is improved in data quality and is efficiently analyzed in 34 minutes, which is 85 times faster than the time taken by a sequential processing. The analytical results for the HVAC operational hours of these buildings show that among 10 building use types, food sales buildings with 17.75 hours of daily HVAC cooling operation are decent targets for HVAC savings. Overall, this analytics pipeline enables the identification of statistically significant results from population based studies of large numbers of building energy time-series datasets with robust results. These types of BEA studies can explore numerous factors impacting building energy efficiency and virtual building energy audits. This approach enables a new generation of data-driven buildings energy analysis at scale.

36 MATERIALS SCIENCE↗

COVID-19 Data Curation Effort: An Initial Analysis of the Data

During the COVID-19 pandemic of 2020, major case reporting outlets quickly coalesced around two or three primary vendors. Johns Hopkins University and The New York Times were among the more prominent, and all were of great value to the nation, particularly during the uncertain early stages of the pandemic. They primarily focused on three major attributes: number of new cases, deaths, and recovery. Recognizing that many states were reporting very detailed data sets (e.g., hospital beds) at a county level or finer, the ORNL Pandemic Modeling team embarked on a major data curation effort from March to June 2020 for the purpose of capturing this wealth of detailed data. The challenge of curating this data was daunting. The number of attributes reported by the states grew on almost on a weekly basis. States were routinely shifting their web tool strategies away from easily parsable HTML-based formatting to new Tableau and ArcGIS content. This growth in the sheer number of attributes, combined with the unpredictable shifts in data format, meant an aggressive and agile combination of automated scripting and manual scraping was required to capture new daily streams. Further, the team had to scale up staff and widen its approach for capture and storage. As a result, the team collected more than 11 million data points. Following the close of this data collection effort on June 30 th , 2020, the team embarked on a major effort to appraise what had been collected, including an inventory list, spatial completeness, temporal completeness, scale and geographic characteristics, and a determination. A report on this matter was submitted on September 15 th , 2020, titled “DOE COVID-19 Data Curation Effort: Overview of Data Collection Coverage”. Over 2000 unique attributes had been netted over a wide range of spatial scales, including state, county, zip codes, health regions, and census blocks. Over 11 million individual data points were collected across these attributes, and spatial coverage (in total) included all 50 states and multiple territories. What became apparent in the process is that in the absence of any data standards, many states reported a wide variety of unique attributes that were not always compatible with attributes reported in other states. As time continued, states began adding new attributes and offering finer grain detail in some older attributes. This meant that not all data streams existed for the entire time period; in fact, the number tended to increase dramatically towards the end. Often, states would begin an attribute series and then stop altogether. These highly variable and uncertain conditions illuminated the need for harmonization approaches that would reconcile and conflate changing attribute names and detail over time. For example, grouping racial data reported as either Black or African American, depending on the state, into a single harmonized attribute. These choices would make a within-state analysis possible during the time period and lead to potential between-state analytics later on. This was almost entirely a manual decision process, requiring some subjective decision-making at times, to prevent a fragmented, short-lived collection of time series fragments that would offer few insights into trends, patterns, and correlates. This report imports harmonized data for state and county into the World Spatio-Temporal Analytics and Mapping Project (WSTAMP). WSTAMP is a major space-time analysis and visualization tool developed at ORNL for the National Geospatial-Intelligence Agency specifically for this kind of exploratory analysis. WSTAMP offers a rich analytical and graphical environment consisting of a wide range of analytics. These include time series plots, statistical summaries, data mining techniques, trend and pattern detection, and hypothesis generation.

59 BASIC BIOLOGICAL SCIENCES↗

Planetary Spin and Obliquity from Mergers

In planetary systems with sufficiently small inter-planet spacing, close encounters can lead to planetary collisions/mergers or ejections. Here, we study the spin property of the merger products of two giant planets in a statistical manner using numerical simulations and analytical modeling. Planetary collisions lead to rapidly rotating objects and a broad range of obliquities. We find that, under typical conditions for two-planet scatterings, the distributions of spin magnitude S and obliquity ${\theta }_{\mathrm{SL}}$ of the merger products have simple analytical forms: f S ∝ S and ${f}_{\cos {\theta }_{\mathrm{SL}}}\propto {(1-{\cos }^{2}{\theta }_{\mathrm{SL}})}^{-1/2}$. Through parameter studies, we determine the regime of validity for the analytical distributions of spin and obliquity. Since planetary mergers are a major outcome of planet–planet scatterings, observational search for the spin/obliquity signatures of exoplanets would provide important constraints on the dynamical history of planetary systems.

79 ASTRONOMY AND ASTROPHYSICS↗

Particle-number fluctuations near the critical point of nuclear matter

Equation of state with the quantum statistics corrections is used for particle-number fluctuations ω of the isotopically symmetric nuclear matter with interparticle van der Waals and Skyrme local density interactions. Here, the fluctuations, ω ∝ 1/K, are analytically derived through the isothermal incompressibility K at first order over a small quantum-statistics parameter. Our approximate analytical results appear to be in good agreement with the results of accurate numerical calculations. These results are also close to those obtained by using more accurate Tolman and Rowlinson expansions of the incompressibility K near the critical point. More general formula for fluctuations ω, improved at the critical point, was obtained for a finite particle-number average $\langle$N$\rangle$ by neglecting, for simplicity, small quantum statistics effects. It is shown that for a large dimensionless parameter, α ∝ K 2 $\langle$N$\rangle$/K", where K" is the second derivative of the incompressibility K as function of the average particle density n, far from the critical point (α >> 1), one finds the traditional asymptote, ω ∝ 1/K, for the fluctuations ω. For a small parameter, α << 1, near the critical point, where K = 0 and α = 0, one obtains another asymptote of ω. These fluctuations, having a maximum near the critical point as function of the average density n, for finite values of $\langle$N$\rangle$ are finite and relatively small, in contrast to the results of the traditional calculations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Generalized master equation for particle transport in binary random media with renewal statistics

Particle transport in binary stochastic mixtures is classically modeled assuming Markovian or exponential mixing statistics but in many applications material memory invalidates the Markov assumption. For non-Markovian mixing characterized by alternating renewal processes, a transport-theoretic framework is presented that provides an exact description of transport in nonscattering random binary media with general non-exponential statistics. Our approach is to Markovianize the problem by augmenting the {material type, particle flux} state space with the age or distance from the last interface. A Chapman-Kolmogorov equation is formulated for the joint probability density of the material type, particle flux, and age, and subsequently reduced to a generalized Master equation (GME) in differential form. This constitutes the primary result of this work. A state-updating Monte Carlo algorithm consistent with the GME is developed and benchmarked against analytical solutions for multiple chord-length laws. For purely absorbing renewal statistical media, the GME reproduces analytical benchmarks for the equilibrium age distribution, interior mean/variance of material-conditioned fluxes, and boundary transmittance. Simulations further demonstrate that a Markov (exponential) approximation of non-exponential statistics can introduce large errors in transmittance and interior flux profiles. Lastly, the reintroduction of memory due to scattering is briefly addressed through heuristic considerations.

Fluctuations & noise↗

Systematic quark/gluon identification with ratios of likelihoods

Discriminating between quark- and gluon-initiated jets has long been a central focus of jet substructure, leading to the introduction of numerous observables and calculations to high perturbative accuracy. At the same time, there have been many attempts to fully exploit the jet radiation pattern using tools from statistics and machine learning. We propose a new approach that combines a deep analytic understanding of jet substructure with the optimality promised by machine learning and statistics. After specifying an approximation to the full emission phase space, we show how to construct the optimal observable for a given classification task. This procedure is demonstrated for the case of quark and gluons jets, where we show how to systematically capture sub-eikonal corrections in the splitting functions, and prove that linear combinations of weighted multiplicity is the optimal observable. In addition to providing a new and powerful framework for systematically improving jet substructure observables, we demonstrate the performance of several quark versus gluon jet tagging observables in parton-level Monte Carlo simulations, and find that they perform at or near the level of a deep neural network classifier. Combined with the rapid recent progress in the development of higher order parton showers, we believe that our approach provides a basis for systematically exploiting subleading effects in jet substructure analyses at the Large Hadron Collider (LHC) and beyond.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Large-Scale Trajectory Analysis via Feature Vectors

The explosion of both sensors and GPS-enabled devices has resulted in position/time data being the next big frontier for data analytics. However, many of the problems associated with large numbers of trajectories do not necessarily have an analog with many of the historic big-data applications such as text and image analysis. Modern trajectory analytics exploits much of the cutting-edge research in machine-learning, statistics, computational geometry and other disciplines. We will show that for doing trajectory analytics at scale, it is necessary to fundamentally change the way the information is represented through a feature-vector approach. We then demonstrate the ability to solve large trajectory analytics problems using this representation.

58 GEOSCIENCES↗

Genetic programming for interpretable, data-driven continuum damage models.

The damage mechanisms that lead to failure in engineering alloys have been studied extensively, but converting this knowledge into constitutive models that are suitable for engineering-scale analysis remains a challenge. Evolution laws for continuum damage have been developed in the past and have proven effective but suffer from many non-physical assumptions that inhibit the overall accuracy of the model. Further, the assumptions inherent in these existing models prevent them from being applicable to a broad class of materials. At the same time, computational models of fine-scale damage mechanisms continue to advance making it tractable to generate large training data sets through computer simulation. Data-driven machine learning approaches can leverage these data sets to avoid making limiting assumptions, and instead produce models directly from the results of microstructural simulations and/or experiments. Many of these machine learning approaches are rapid and accurate, but they offer little to no insight into the underlying relationships among state variables being discovered. Conversely, genetic programming symbolic regression (GPSR) is a machine learning method that produces analytic expressions relating the state variables, allowing maximal insight and interpretability. To that end, we propose using GPSR as a data-driven method of obtaining microstructurally informed continuum damage models. Data is generated using microstructural simulations of damage evolution, parameterized over microstructural statistics (i.e., pore shape) and nominally applied deformations. Analytic expressions for damage evolution are obtained from the data using GPSR, and these expressions are then utilized within a continuum constitutive model. Overall, this approach is a promising method of automatically obtaining analytic relations describing constitutive phenomena in a material.

Buche, Michael Robert↗

Many-Body Level Statistics of Single-Particle Quantum Chaos

We consider a noninteracting many-fermion system populating levels of a unitary random matrix ensemble (equivalent to the q = 2 complex Sachdev-Ye-Kitaev model)—a generic model of single-particle quantum chaos. We study the corresponding many-particle level statistics by calculating the spectral form factor analytically using algebraic methods of random matrix theory, and match it with an exact numerical simulation. Despite the integrability of the theory, the many-body spectral rigidity is found to have a surprisingly rich landscape. In particular, we find a residual repulsion of distant many-body levels stemming from single-particle chaos, together with islands of level attraction. These results are encoded in an exponential ramp in the spectral form factor, which we show to be a universal feature of nonergodic many-fermion systems embedded in a chaotic medium.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Trust Your Gut: Comparing Human and Machine Inference from Noisy Visualizations

People commonly utilize visualizations not only to examine a given dataset, but also to draw generalizable conclusions about the underlying models or phenomena. Prior research has compared human visual inference to that of an optimal Bayesian agent, with deviations from rational analysis viewed as problematic. However, human reliance on non-normative heuristics may prove advantageous in certain circumstances. We investigate scenarios where human intuition might surpass idealized statistical rationality. In two experiments, we examine individuals’ accuracy in characterizing the parameters of known data-generating models from bivariate visualizations. Our findings indicate that, although participants generally exhibited lower accuracy compared to statistical models, they frequently outperformed Bayesian agents, particularly when faced with extreme samples. Participants appeared to rely on their internal models to filter out noisy visualizations, thus improving their resilience against spurious data. However, participants displayed overconfidence and struggled with uncertainty estimation. They also exhibited higher variance than statistical machines. Our findings suggest that analyst gut reactions to visualizations may provide an advantage, even when departing from rationality. These results carry implications for designing visual analytics tools, offering new perspectives on how to integrate statistical models and analyst intuition for improved inference and decision-making. The data and materials for this paper are available at https://osf.io/qmfv6

human-machine collaboration↗

A hybrid calibration approach to Hertz-type contact parameters for discrete element models

This study aims at providing a hybrid calibration framework to estimate Hertz-type contact parameters (particle-scale shear modulus and Poisson ratio) for both two-dimensional and three-dimensional discrete element modelling (DEM). On the basis of statistically isotropic granular packings, a set of analytical formulae between macroscopic material parameters (Young modulus and Poisson ratio) and particle-scale Hertz-type contact parameters for granular systems are derived under small-strain isotropic stress conditions. However, the derived analytical solutions are only estimated values for general models. By viewing each DEM modelling as an implicit mathematical function taking the particle-level parameters as independent variables and employing the derived analytical solutions as the initial input parameters, an automatic iterative scheme is proposed to obtain the calibrated parameters with higher accuracies. Considering highly nonlinear features and discontinuities of the macro-micro relationship in Hertz-based discrete element models, the adaptive moment estimation algorithm is adopted in this study because of its capacity of dealing with noise gradients of cost functions. Here, the proposed method is validated with several numerical cases including randomly distributed monodisperse and polydisperse packings. Noticeable improvements in terms of calibration efficiency and accuracy have been made.

Constitutive law↗

Detecting Low-level Radiation Sources Using Border Monitoring Gamma Sensors

We consider a problem of detecting a low-level radiation source using a network of Gamma spectral sensors placed on the periphery of a monitored region. We propose a computationally light-weight, correlation-based method which is primarily intended for systems with limited computing capacity. Sensor measurements are combined at the fusion by first generating decisions at each time step and then taking their majority vote within a time widow. At each time step, decisions are generated using two strategies: (i) SUM method based on a threshold decision on a correlation statistic derived from measurements from all sensors, and (ii) OR method based on logical-OR of threshold decisions based on correlations statistics of individual sensor measurements. We derive analytical performance bounds for false alarm rates of SUM and OR methods, and show that their performance is enhanced by the temporal smoothing of majority vote within a time window. Using measurements from a test campaign, we generate a border monitoring scenario with twelve 2"x2" NaI Gamma sensors deployed on the periphery of 42m x 42m outdoor region. A Cs-137 source is moved in a straight-line across this region, starting several meters outside and finally moving away from it. We illustrate the performance of both correlation-based detection methods, and compare their performances with each other and with a particle filter method. Overall, under small false-alarm conditions, the OR fusion is found to produce better detection performance.

Sen, Satyabrata↗

Statistically-informed deep learning for gravitational wave parameter estimation

We introduce deep learning models to estimate the masses of the binary components of black hole mergers, $(m_1,m_2)$, and three astrophysical properties of the post-merger compact remnant, namely, the final spin, $a_\mathrm f$, and the frequency and damping time of the ringdown oscillations of the fundamental $\ell = m = 2$ bar mode, $(\omega_\mathrm R, \omega_\mathrm I)$. Our neural networks combine a modified WaveNet architecture with contrastive learning and normalizing flow. We validate these models against a Gaussian conjugate prior family whose posterior distribution is described by a closed analytical expression. Upon confirming that our models produce statistically consistent results, we used them to estimate the astrophysical parameters $(m_1,m_2, a_\mathrm f, \omega_\mathrm R, \omega_\mathrm I)$ of five binary black holes: GW150914, GW170104, GW170814, GW190521 and GW190630. We use PyCBC Inference to directly compare traditional Bayesian methodologies for parameter estimation with our deep learning based posterior distributions. Our results show that our neural network models predict posterior distributions that encode physical correlations, and that our data-driven median results and 90% confidence intervals are similar to those produced with gravitational wave Bayesian analyses. This methodology requires a single V100 NVIDIA GPU to produce median values and posterior distributions within two milliseconds for each event. Furthermore, this neural network, and a tutorial for its use, are available at the Data and Learning Hub for Science.

79 ASTRONOMY AND ASTROPHYSICS↗

DEEPEN 3D PFA Weights for Exploration Datasets in Magmatic Environments

DEEPEN stands for DE-risking Exploration of geothermal Plays in magmatic ENvironments. As part of the development of the DEEPEN 3D play fairway analysis (PFA) methodology for magmatic plays (conventional hydrothermal, superhot EGS, and supercritical), weights needed to be developed for use in the weighted sum of the different favorability index models produced from geoscientific exploration datasets. This GDR submission includes those weights. The weighting was done using two different approaches: one based on expert opinions, and one based on statistical learning. The weights are intended to describe how useful a particular exploration method is for imaging each component of each play type. They may be adjusted based on the characteristics of the resource under investigation, knowledge of the quality of the dataset, or simply to reduce the impact a single dataset has on the resulting outputs. Within the DEEPEN PFA, separate sets of weights are produced for each component of each play type, since exploration methods hold different levels of importance for detecting each play component, within each play type. The weights for conventional hydrothermal systems were based on the average of the normalized weights used in the DOE-funded PFA projects that were focused on magmatic plays. This decision was made because conventional hydrothermal plays are already well-studied and understood, and therefore it is logical to use existing weights where possible. In contrast, a true PFA has never been applied to superhot EGS or supercritical plays, meaning that exploration methods have never been weighted in terms of their utility in imaging the components of these plays. To produce weights for superhot EGS and supercritical plays, two different approaches were used: one based on expert opinion and the analytical hierarchy process (AHP), and another using a statistical approach based on principal component analysis (PCA). The weights are intended to provide standardized sets of weights for each play type in all magmatic geothermal systems. Two different approaches were used to investigate whether a more data-centric approach might allow new insights into the datasets, and also to analyze how different weighting approaches impact the outcomes. The expert/AHP approach involved using an online tool (https://bpmsg.com/ahp/) with built-in forms to make pairwise comparisons which are used to rank exploration methods against one-another. The inputs are then combined in a quantitative way, ultimately producing a set of consensus-based weights. To minimize the burden on each individual participant, the forms were completed in group discussions. While the group setting means that there is potential for some opinions to outweigh others, it also provides a venue for conversation to take place, in theory leading the group to a more robust consensus then what can be achieved on an individual basis. This exercise was done with two separate groups: one consisting of U.S.-based experts, and one consisting of Iceland-based experts in magmatic geothermal systems. The two sets of weights were then averaged to produce what we will from here on refer to as the "expert opinion-based weights," or "expert weights" for short. While expert opinions allow us to include more nuanced information in the weights, expert opinions are subject to human bias. Data-centric or statistical approaches help to overcome these potential human biases by focusing on and drawing conclusions from the data alone. More information on this approach along with the dataset used to produce the statistical weights may be found in the linked dataset below.

15 GEOTHERMAL ENERGY↗

Freely jointed chain models with extensible links

We report analytical relations for the mechanical response of single polymer chains are valuable for modeling purposes, on both the molecular and the continuum scale. These relations can be obtained using statistical thermodynamics and an idealized single-chain model, such as the freely jointed chain model. To include bond stretching, the rigid links in the freely jointed chain model can be made extensible, but this almost always renders the model analytically intractable. Here, an asymptotically correct statistical thermodynamic theory is used to develop analytic approximations for the single-chain mechanical response of this model. The accuracy of these approximations is demonstrated using several link potential energy functions. This approach can be applied to other single-chain models, and to molecular stretching in general.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗