Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Automated pipeline framework for processing of large-scale building energy time series data

Commercial buildings account for one third of the total electricity consumption in the United States and a significant amount of this energy is wasted. Therefore, there is a need for “virtual” energy audits, to identify energy inefficiencies and their associated savings opportunities using methods that can be non-intrusive and automated for application to large populations of buildings. Here we demonstrate virtual energy audits applied to large populations of buildings’ time-series smart-meter data using a systematic approach and a fully automated Building Energy Analytics (BEA) Pipeline that unifies, cleans, stores and analyzes building energy datasets in a non-relational data warehouse for efficient insights and results. This BEA pipeline is based on a custom compute job scheduler for a high performance computing cluster to enable parallel processing of Slurm jobs. Within the analytics pipeline, we introduced a data qualification tool that enhances data quality by fixing common errors, while also detecting abnormalities in a building’s daily operation using hierarchical clustering. We analyze the HVAC scheduling of a population of 816 buildings, using this analytics pipeline, as part of a cross-sectional study. With our approach, this sample of 816 buildings is improved in data quality and is efficiently analyzed in 34 minutes, which is 85 times faster than the time taken by a sequential processing. The analytical results for the HVAC operational hours of these buildings show that among 10 building use types, food sales buildings with 17.75 hours of daily HVAC cooling operation are decent targets for HVAC savings. Overall, this analytics pipeline enables the identification of statistically significant results from population based studies of large numbers of building energy time-series datasets with robust results. These types of BEA studies can explore numerous factors impacting building energy efficiency and virtual building energy audits. This approach enables a new generation of data-driven buildings energy analysis at scale.

36 MATERIALS SCIENCE↗

COVID-19 Data Curation Effort: An Initial Analysis of the Data

During the COVID-19 pandemic of 2020, major case reporting outlets quickly coalesced around two or three primary vendors. Johns Hopkins University and The New York Times were among the more prominent, and all were of great value to the nation, particularly during the uncertain early stages of the pandemic. They primarily focused on three major attributes: number of new cases, deaths, and recovery. Recognizing that many states were reporting very detailed data sets (e.g., hospital beds) at a county level or finer, the ORNL Pandemic Modeling team embarked on a major data curation effort from March to June 2020 for the purpose of capturing this wealth of detailed data. The challenge of curating this data was daunting. The number of attributes reported by the states grew on almost on a weekly basis. States were routinely shifting their web tool strategies away from easily parsable HTML-based formatting to new Tableau and ArcGIS content. This growth in the sheer number of attributes, combined with the unpredictable shifts in data format, meant an aggressive and agile combination of automated scripting and manual scraping was required to capture new daily streams. Further, the team had to scale up staff and widen its approach for capture and storage. As a result, the team collected more than 11 million data points. Following the close of this data collection effort on June 30 th , 2020, the team embarked on a major effort to appraise what had been collected, including an inventory list, spatial completeness, temporal completeness, scale and geographic characteristics, and a determination. A report on this matter was submitted on September 15 th , 2020, titled “DOE COVID-19 Data Curation Effort: Overview of Data Collection Coverage”. Over 2000 unique attributes had been netted over a wide range of spatial scales, including state, county, zip codes, health regions, and census blocks. Over 11 million individual data points were collected across these attributes, and spatial coverage (in total) included all 50 states and multiple territories. What became apparent in the process is that in the absence of any data standards, many states reported a wide variety of unique attributes that were not always compatible with attributes reported in other states. As time continued, states began adding new attributes and offering finer grain detail in some older attributes. This meant that not all data streams existed for the entire time period; in fact, the number tended to increase dramatically towards the end. Often, states would begin an attribute series and then stop altogether. These highly variable and uncertain conditions illuminated the need for harmonization approaches that would reconcile and conflate changing attribute names and detail over time. For example, grouping racial data reported as either Black or African American, depending on the state, into a single harmonized attribute. These choices would make a within-state analysis possible during the time period and lead to potential between-state analytics later on. This was almost entirely a manual decision process, requiring some subjective decision-making at times, to prevent a fragmented, short-lived collection of time series fragments that would offer few insights into trends, patterns, and correlates. This report imports harmonized data for state and county into the World Spatio-Temporal Analytics and Mapping Project (WSTAMP). WSTAMP is a major space-time analysis and visualization tool developed at ORNL for the National Geospatial-Intelligence Agency specifically for this kind of exploratory analysis. WSTAMP offers a rich analytical and graphical environment consisting of a wide range of analytics. These include time series plots, statistical summaries, data mining techniques, trend and pattern detection, and hypothesis generation.

59 BASIC BIOLOGICAL SCIENCES↗

Planetary Spin and Obliquity from Mergers

In planetary systems with sufficiently small inter-planet spacing, close encounters can lead to planetary collisions/mergers or ejections. Here, we study the spin property of the merger products of two giant planets in a statistical manner using numerical simulations and analytical modeling. Planetary collisions lead to rapidly rotating objects and a broad range of obliquities. We find that, under typical conditions for two-planet scatterings, the distributions of spin magnitude S and obliquity ${\theta }_{\mathrm{SL}}$ of the merger products have simple analytical forms: f S ∝ S and ${f}_{\cos {\theta }_{\mathrm{SL}}}\propto {(1-{\cos }^{2}{\theta }_{\mathrm{SL}})}^{-1/2}$. Through parameter studies, we determine the regime of validity for the analytical distributions of spin and obliquity. Since planetary mergers are a major outcome of planet–planet scatterings, observational search for the spin/obliquity signatures of exoplanets would provide important constraints on the dynamical history of planetary systems.

79 ASTRONOMY AND ASTROPHYSICS↗

Determination of the ground albedo and the index of absorption of atmospheric particulates by remote sensing. I - Theory

A statistical technique is developed for inferring the optimum values of the ground albedo and the effective imaginary term of the complex refractive index of atmospheric particulates. The procedure compares measurements of the ratio of the hemispheric diffuse to directly transmitted solar flux density at the earth's surface with radiative transfer computations of the same as suggested by Herman et al. (1975). A detailed study is presented which shows the extent to which the ratio of diffuse to direct solar radiation is sensitive to many of the radiative transfer parameters. Results indicate that the optical depth and size distribution of atmospheric aerosol particles are the two parameters which uniquely specify the radiation field to the point where ground albedo and index of absorption can be inferred. Varying the real part of the complex refractive index of atmospheric particulates as well as their vertical distribution is found to have a negligible effect on the diffuse-direct ratio. The statistical procedure utilizes a semi-analytic gradient search method from least-squares theory and includes a detailed error analysis.

King, M. D.↗

Search strategy effects on PN acquisition performance

The present paper focusses on 'random' and 'expanding window' PN acquisition search strategies and analytically develops the PN acquisition time statistics as functions of salient system parameters such as prediction SNR, detection and false alarm probabilities and a priori information on epoch location. The significance of this analysis is its general applicability to arbitrary postdetection processing schemes. Computed performance results account for the above salient parameters, wherein sequential detection is employed in conjunction with random and selected expanding window search strategies.

Weinberg, A.↗

Possibility of measuring gravity-wave momentum flux by single beam observation of MST radar

Vincent and Reid (1983) proposed a technique to measure gravity-wave momentum fluxes in the atmosphere by mesosphere-stratosphere-troposphere (MST) radars using two or more radar beams. Since the vertical momentum fluxes are assumed to be due to gravity waves, it appears possible to make use of the dispersion and polarization relations for gravity waves in extracting useful information from the radar data. In particular, for an oblique radar beam, information about both the vertical and the horizontal velocities associated with the waves are contained in the measured Doppler data. Therefore, it should be possible to extract both V sub Z and V sub h from a single beam observational configuration. A procedure is proposed to perform such an analysis. The basic assumptions are: the measured velocity fluctuations are due to gravity waves and a separable model gravity-wave spectrum of the Garrett-Munk type that is statistically homogeneous in the horizontal plane. Analytical expressions can be derived that relate the observed velocity fluctuations to the wave momentum flux at each range gate. In practice, the uncertainties related to the model parameters and measurement accuracy will affect the results. A MST radar configuration is considered.

Liu, C. H.↗

Effects of intensity modulations on the power spectra of random processes

Intensity-modulated random processes (IMRPs), defined as the products of (1) deterministic modulating functions or processes and (2) stationary modulated processes statistically independent of (1), are investigated analytically. Instantaneous power spectra are derived for IMRPs with different classes of (1), and a spectrum series expansion with a known locally stationary approximation as its first term is obtained. Numerical results for sample IMRPs with known (deterministic), stationary, nonstationary, ergodic, onset, and bell-shaped types of (1) are presented graphically and briefly characterized.

Mark, W. D.↗

Speckle in the diffraction patterns of Hendricks-Teller and icosahedral glass models

It is shown that the X-ray diffraction patterns from the Hendricks-Teller model for layered systems and the icosahedral glass models for the icosahedral phases show large fluctuations between nearby scattering wave vectors and from sample to sample, that are quite analogous to laser speckle. The statistics of these fluctuations are studied analytically for the first model and via computer simulations for the second. The observability of these effects is discussed briefly.

Garg, Anupam↗

Proceedings of the Workshop on Change of Representation and Problem Reformulation

The proceedings of the third Workshop on Change of representation and Problem Reformulation is presented. In contrast to the first two workshops, this workshop was focused on analytic or knowledge-based approaches, as opposed to statistical or empirical approaches called 'constructive induction'. The organizing committee believes that there is a potential for combining analytic and inductive approaches at a future date. However, it became apparent at the previous two workshops that the communities pursuing these different approaches are currently interested in largely non-overlapping issues. The constructive induction community has been holding its own workshops, principally in conjunction with the machine learning conference. While this workshop is more focused on analytic approaches, the organizing committee has made an effort to include more application domains. We have greatly expanded from the origins in the machine learning community. Participants in this workshop come from the full spectrum of AI application domains including planning, qualitative physics, software engineering, knowledge representation, and machine learning.

Lowry, Michael R.↗

Particle-number fluctuations near the critical point of nuclear matter

Equation of state with the quantum statistics corrections is used for particle-number fluctuations ω of the isotopically symmetric nuclear matter with interparticle van der Waals and Skyrme local density interactions. Here, the fluctuations, ω ∝ 1/K, are analytically derived through the isothermal incompressibility K at first order over a small quantum-statistics parameter. Our approximate analytical results appear to be in good agreement with the results of accurate numerical calculations. These results are also close to those obtained by using more accurate Tolman and Rowlinson expansions of the incompressibility K near the critical point. More general formula for fluctuations ω, improved at the critical point, was obtained for a finite particle-number average $\langle$N$\rangle$ by neglecting, for simplicity, small quantum statistics effects. It is shown that for a large dimensionless parameter, α ∝ K 2 $\langle$N$\rangle$/K", where K" is the second derivative of the incompressibility K as function of the average particle density n, far from the critical point (α >> 1), one finds the traditional asymptote, ω ∝ 1/K, for the fluctuations ω. For a small parameter, α << 1, near the critical point, where K = 0 and α = 0, one obtains another asymptote of ω. These fluctuations, having a maximum near the critical point as function of the average density n, for finite values of $\langle$N$\rangle$ are finite and relatively small, in contrast to the results of the traditional calculations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Generalized master equation for particle transport in binary random media with renewal statistics

Particle transport in binary stochastic mixtures is classically modeled assuming Markovian or exponential mixing statistics but in many applications material memory invalidates the Markov assumption. For non-Markovian mixing characterized by alternating renewal processes, a transport-theoretic framework is presented that provides an exact description of transport in nonscattering random binary media with general non-exponential statistics. Our approach is to Markovianize the problem by augmenting the {material type, particle flux} state space with the age or distance from the last interface. A Chapman-Kolmogorov equation is formulated for the joint probability density of the material type, particle flux, and age, and subsequently reduced to a generalized Master equation (GME) in differential form. This constitutes the primary result of this work. A state-updating Monte Carlo algorithm consistent with the GME is developed and benchmarked against analytical solutions for multiple chord-length laws. For purely absorbing renewal statistical media, the GME reproduces analytical benchmarks for the equilibrium age distribution, interior mean/variance of material-conditioned fluxes, and boundary transmittance. Simulations further demonstrate that a Markov (exponential) approximation of non-exponential statistics can introduce large errors in transmittance and interior flux profiles. Lastly, the reintroduction of memory due to scattering is briefly addressed through heuristic considerations.

Fluctuations & noise↗

Systematic quark/gluon identification with ratios of likelihoods

Discriminating between quark- and gluon-initiated jets has long been a central focus of jet substructure, leading to the introduction of numerous observables and calculations to high perturbative accuracy. At the same time, there have been many attempts to fully exploit the jet radiation pattern using tools from statistics and machine learning. We propose a new approach that combines a deep analytic understanding of jet substructure with the optimality promised by machine learning and statistics. After specifying an approximation to the full emission phase space, we show how to construct the optimal observable for a given classification task. This procedure is demonstrated for the case of quark and gluons jets, where we show how to systematically capture sub-eikonal corrections in the splitting functions, and prove that linear combinations of weighted multiplicity is the optimal observable. In addition to providing a new and powerful framework for systematically improving jet substructure observables, we demonstrate the performance of several quark versus gluon jet tagging observables in parton-level Monte Carlo simulations, and find that they perform at or near the level of a deep neural network classifier. Combined with the rapid recent progress in the development of higher order parton showers, we believe that our approach provides a basis for systematically exploiting subleading effects in jet substructure analyses at the Large Hadron Collider (LHC) and beyond.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Large-Scale Trajectory Analysis via Feature Vectors

The explosion of both sensors and GPS-enabled devices has resulted in position/time data being the next big frontier for data analytics. However, many of the problems associated with large numbers of trajectories do not necessarily have an analog with many of the historic big-data applications such as text and image analysis. Modern trajectory analytics exploits much of the cutting-edge research in machine-learning, statistics, computational geometry and other disciplines. We will show that for doing trajectory analytics at scale, it is necessary to fundamentally change the way the information is represented through a feature-vector approach. We then demonstrate the ability to solve large trajectory analytics problems using this representation.

58 GEOSCIENCES↗

Genetic programming for interpretable, data-driven continuum damage models.

The damage mechanisms that lead to failure in engineering alloys have been studied extensively, but converting this knowledge into constitutive models that are suitable for engineering-scale analysis remains a challenge. Evolution laws for continuum damage have been developed in the past and have proven effective but suffer from many non-physical assumptions that inhibit the overall accuracy of the model. Further, the assumptions inherent in these existing models prevent them from being applicable to a broad class of materials. At the same time, computational models of fine-scale damage mechanisms continue to advance making it tractable to generate large training data sets through computer simulation. Data-driven machine learning approaches can leverage these data sets to avoid making limiting assumptions, and instead produce models directly from the results of microstructural simulations and/or experiments. Many of these machine learning approaches are rapid and accurate, but they offer little to no insight into the underlying relationships among state variables being discovered. Conversely, genetic programming symbolic regression (GPSR) is a machine learning method that produces analytic expressions relating the state variables, allowing maximal insight and interpretability. To that end, we propose using GPSR as a data-driven method of obtaining microstructurally informed continuum damage models. Data is generated using microstructural simulations of damage evolution, parameterized over microstructural statistics (i.e., pore shape) and nominally applied deformations. Analytic expressions for damage evolution are obtained from the data using GPSR, and these expressions are then utilized within a continuum constitutive model. Overall, this approach is a promising method of automatically obtaining analytic relations describing constitutive phenomena in a material.

Buche, Michael Robert↗

Many-Body Level Statistics of Single-Particle Quantum Chaos

We consider a noninteracting many-fermion system populating levels of a unitary random matrix ensemble (equivalent to the q = 2 complex Sachdev-Ye-Kitaev model)—a generic model of single-particle quantum chaos. We study the corresponding many-particle level statistics by calculating the spectral form factor analytically using algebraic methods of random matrix theory, and match it with an exact numerical simulation. Despite the integrability of the theory, the many-body spectral rigidity is found to have a surprisingly rich landscape. In particular, we find a residual repulsion of distant many-body levels stemming from single-particle chaos, together with islands of level attraction. These results are encoded in an exponential ramp in the spectral form factor, which we show to be a universal feature of nonergodic many-fermion systems embedded in a chaotic medium.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Trust Your Gut: Comparing Human and Machine Inference from Noisy Visualizations

People commonly utilize visualizations not only to examine a given dataset, but also to draw generalizable conclusions about the underlying models or phenomena. Prior research has compared human visual inference to that of an optimal Bayesian agent, with deviations from rational analysis viewed as problematic. However, human reliance on non-normative heuristics may prove advantageous in certain circumstances. We investigate scenarios where human intuition might surpass idealized statistical rationality. In two experiments, we examine individuals’ accuracy in characterizing the parameters of known data-generating models from bivariate visualizations. Our findings indicate that, although participants generally exhibited lower accuracy compared to statistical models, they frequently outperformed Bayesian agents, particularly when faced with extreme samples. Participants appeared to rely on their internal models to filter out noisy visualizations, thus improving their resilience against spurious data. However, participants displayed overconfidence and struggled with uncertainty estimation. They also exhibited higher variance than statistical machines. Our findings suggest that analyst gut reactions to visualizations may provide an advantage, even when departing from rationality. These results carry implications for designing visual analytics tools, offering new perspectives on how to integrate statistical models and analyst intuition for improved inference and decision-making. The data and materials for this paper are available at https://osf.io/qmfv6

human-machine collaboration↗

Accuracy of trace element determinations in alternate fuels

A review of the techniques used at Lewis Research Center (LeRC) in trace metals analysis is presented, including the results of Atomic Absorption Spectrometry and DC Arc Emission Spectrometry of blank levels and recovery experiments for several metals. The design of an Interlaboratory Study conducted by LeRC is presented. Several factors were investigated, including: laboratory, analytical technique, fuel type, concentration, and ashing additive. Conclusions drawn from the statistical analysis will help direct research efforts toward those areas most responsible for the poor interlaboratory analytical results.

Greenbauer-Seng, L. A.↗

Estimation of the accuracy of dynamic flight-determined coefficients

This paper discusses means of assessing the accuracy of maximum likelihood parameter estimates obtained from dynamic flight data. The commonly used analytical predictors of accuracy are compared from both statistical and simplified geometric standpoints. Emphasizing practical considerations, such as modeling error, the accuracy predictions are evaluated with real and simulated data. Improved computations of the Cramer-Rao bound to correct large discrepancies caused by colored noise and modeling error are presented. This corrected Cramer-Rao bound is the best available analytical predictor of accuracy. Engineering judgement, aided by such analytical tools, is the final arbiter of accuracy estimation.

Maine, R. E.↗