Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical graphics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

HL-LHC Analysis With ROOT: ROOT Project Input to the HL-LHC Computing Review (Stage 2)

ROOT is high energy physics' software for storing and mining data in a statistically sound way, to publish results with scientific graphics. It is evolving since 25 years, now providing the storage format for more than one exabyte of data; virtually all high energy physics experiments use ROOT. With another significant increase in the amount of data to be handled scheduled to arrive in 2027, ROOT is preparing for a massive upgrade of its core ingredients. As part of a review of crucial software for high energy physics, the ROOT team has documented its R&D plans for the coming years.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Classification of Cloud Particle Imagery from Aircraft Platforms Using Convolutional Neural Networks

Abstract A vast amount of ice crystal imagery exists from a variety of field campaign initiatives that can be utilized for cloud microphysical research. Here, nine convolutional neural networks are used to classify particles into nine regimes on over 10 million images from the Cloud Particle Imager probe, including liquid and frozen states and particles with evidence of riming. A transfer learning approach proves that the Visual Geometry Group (VGG-16) network best classifies imagery with respect to multiple performance metrics. Classification accuracies on a validation dataset reach 97% and surpass traditional automated classification. Furthermore, after initial model training and preprocessing, 10 000 images can be classified in approximately 35 s using 20 central processing unit cores and two graphics processing units, which reaches real-time classification capabilities. Statistical analysis of the classified images indicates that a large portion (57%) of the dataset is unusable, meaning the images are too blurry or represent indistinguishable small fragments. In addition, 19% of the dataset is classified as liquid drops. After removal of fragments, blurry images, and cloud drops, 38% of the remaining ice particles are largely intersecting the image border (≥10% cutoff) and therefore are considered unusable because of the inability to properly classify and dimensionalize. After this filtering, an unprecedented database of 1 560 364 images across all campaigns is available for parameter extraction and bulk statistics on specific particle types in a wide variety of storm systems, which can act to improve the current state of microphysical parameterizations.

54 ENVIRONMENTAL SCIENCES↗

Active learning of chemical reaction networks via probabilistic graphical models and Boolean reaction circuits

Discerning networks of many reactions among multiple interconverting species is challenging. Here, we present a reaction network identification methodology. Our methodology enumerates all stoichiometrically and chemically feasible reactions and requires statistical evidence from effluent concentrations for the inclusion or exclusion of each from the reaction network, contrasting with the commonly seen incremental approach and other work of relying heavily upon chemical intuition and assuming the reactions occurring. Using graph theory alongside an active learning design of experiments that propose maximally informative feeds, we identify the underlying reaction network with minimal laboratory runs. Here, we introduce chemistry-probabilistic graphical modeling and Boolean reaction circuits to statistically quantify which reactions occur from effluent concentrations. Our methodology accurately discerns active reactions, as showcased upon a laboratory network of cross-ketonization of furoic and lauric acid and validated upon simulated networks of thermal and CO 2 -assisted ethane dehydrogenation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Edge counts for the auxiliary pair graph within the graphical unitary group approach

Closed-form expressions are presented for the numbers of edges in the auxiliary pair graphs (APGs) associated with non spin-orbit and spin-orbit Shavitt graphs for full configuration interaction expansions. A Shavitt graph is a visual representation of a configuration state function expansion space constructed via the graphical unitary group approach (GUGA). An APG is an organisational aid and a programmatic tool generated from a Shavitt graph. The number of edges in an APG determines bounds on the computational scaling as a function of the total numbers of electrons, orbitals, and spin multiplicities. The edge counts extend a suite of Shavitt graph statistics based on these functional parameters. The derivation and the presentation of the formulas for the edge counts has been assisted by the bra-ket interchange symmetry and the particle-hole interchange symmetry in the GUGA formalism. Furthermore, these symmetry operators produce one-to-one correspondences between various sets of edges, and this yields identities among some edge count formulas. There are 208 possible edge types. Of these, some do not contribute to two-electron operators, some are related by bra-ket interchange symmetry, and some are related by particle-hole interchange symmetry. For the remaining unique edge types, explicit expressions are derived for the numbers of edges.

74 ATOMIC AND MOLECULAR PHYSICS↗

Computationally efficient Bayesian estimation of graphical networks for omics data

Graphical networks are useful, widely-used modeling approaches to represent complex biological processes with biological measurements generated by platforms such as mass spectrometry. Bayesian analyses of graphical networks for omics data have several advantages over their frequentist counterparts, such as the inclusion of prior knowledge in the estimation of models. However, Bayesian approaches to date have only been feasible for data with a couple hundred biomolecules due to prohibitive computational time, but omics data often contains tens of thousands of biomolecules. Here, we present and illustrate a more computationally efficient approach named BPlane (Bayesian PseudoLikelihood-based Algorithm for Network Estimation) to extend Bayesian modeling capabilities for larger-sized datasets, such as most untargeted proteomics data. Via simulation, we demonstrate that BPlane produces substantial computational savings over a current state-of-the-art Bayesian algorithm while maintaining competitive edge detection accuracy. On a SARS-CoV2 proteomics data with 7000 proteins, the competing algorithm takes three times as long to complete the first iteration as BPlane takes to converge after over 100 iterations.

EM algorithm↗

Empirically-calibrated H100 node power models for accurate AI training energy estimation

Accurately quantifying the energy use of artificial intelligence (AI) training is critical for infrastructure planning, carbon accounting, and sustainable data center operation, but few studies have directly measured the power consumption of production workloads on contemporary hardware. By combining empirical measurements from Brookhaven National Laboratory during AI training on 8-graphics-processing-unit H100 systems with open-source benchmarking data, we develop statistical models relating computational intensity to node-level power consumption. We measure the gap between manufacturer-rated thermal design power (TDP) and actual power demand during AI training. Our analysis reveals that even computationally intensive workloads operate at only 76% of the 10.2 kW TDP rating. Our architecture-specific model, calibrated to floating-point operations, predicts energy consumption with 11.4% mean absolute percentage error, significantly outperforming TDP-based approaches (27%–37% error). We identified distinct power signatures between transformer and convolutional neural network architectures, with transformers showing characteristic fluctuations that may impact grid stability. These results provide a measurement-grounded basis for improving AI training energy estimates, enabling more reliable infrastructure sizing, cost projections, and environmental impact assessments.

Newkirk, Alex C↗

A probabilistic graphical model foundation for enabling predictive digital twins at scale

A unifying mathematical formulation is needed to move from one-off digital twins built through custom implementations to robust digital twin implementations at scale. This work proposes a probabilistic graphical model as a formal mathematical representation of a digital twin and its associated physical asset. We create an abstraction of the asset–twin system as a set of coupled dynamical systems, evolving over time through their respective state spaces and interacting via observed data and control inputs. The formal definition of this coupled system as a probabilistic graphical model enables us to draw upon well-established theory and methods from Bayesian statistics, dynamical systems and control theory. The declarative and general nature of the proposed digital twin model make it rigorous yet flexible, enabling its application at scale in a diverse range of application areas. Here, we demonstrate how the model is instantiated to enable a structural digital twin of an unmanned aerial vehicle (UAV). The digital twin is calibrated using experimental data from a physical UAV asset. Its use in dynamic decision-making is then illustrated in a synthetic example where the UAV undergoes an in-flight damage event and the digital twin is dynamically updated using sensor data. The graphical model foundation ensures that the digital twin calibration and updating process is principled, unified and able to scale to an entire fleet of digital twins.

42 ENGINEERING↗

COVID-19 Data Curation Effort: An Initial Analysis of the Data

During the COVID-19 pandemic of 2020, major case reporting outlets quickly coalesced around two or three primary vendors. Johns Hopkins University and The New York Times were among the more prominent, and all were of great value to the nation, particularly during the uncertain early stages of the pandemic. They primarily focused on three major attributes: number of new cases, deaths, and recovery. Recognizing that many states were reporting very detailed data sets (e.g., hospital beds) at a county level or finer, the ORNL Pandemic Modeling team embarked on a major data curation effort from March to June 2020 for the purpose of capturing this wealth of detailed data. The challenge of curating this data was daunting. The number of attributes reported by the states grew on almost on a weekly basis. States were routinely shifting their web tool strategies away from easily parsable HTML-based formatting to new Tableau and ArcGIS content. This growth in the sheer number of attributes, combined with the unpredictable shifts in data format, meant an aggressive and agile combination of automated scripting and manual scraping was required to capture new daily streams. Further, the team had to scale up staff and widen its approach for capture and storage. As a result, the team collected more than 11 million data points. Following the close of this data collection effort on June 30 th , 2020, the team embarked on a major effort to appraise what had been collected, including an inventory list, spatial completeness, temporal completeness, scale and geographic characteristics, and a determination. A report on this matter was submitted on September 15 th , 2020, titled “DOE COVID-19 Data Curation Effort: Overview of Data Collection Coverage”. Over 2000 unique attributes had been netted over a wide range of spatial scales, including state, county, zip codes, health regions, and census blocks. Over 11 million individual data points were collected across these attributes, and spatial coverage (in total) included all 50 states and multiple territories. What became apparent in the process is that in the absence of any data standards, many states reported a wide variety of unique attributes that were not always compatible with attributes reported in other states. As time continued, states began adding new attributes and offering finer grain detail in some older attributes. This meant that not all data streams existed for the entire time period; in fact, the number tended to increase dramatically towards the end. Often, states would begin an attribute series and then stop altogether. These highly variable and uncertain conditions illuminated the need for harmonization approaches that would reconcile and conflate changing attribute names and detail over time. For example, grouping racial data reported as either Black or African American, depending on the state, into a single harmonized attribute. These choices would make a within-state analysis possible during the time period and lead to potential between-state analytics later on. This was almost entirely a manual decision process, requiring some subjective decision-making at times, to prevent a fragmented, short-lived collection of time series fragments that would offer few insights into trends, patterns, and correlates. This report imports harmonized data for state and county into the World Spatio-Temporal Analytics and Mapping Project (WSTAMP). WSTAMP is a major space-time analysis and visualization tool developed at ORNL for the National Geospatial-Intelligence Agency specifically for this kind of exploratory analysis. WSTAMP offers a rich analytical and graphical environment consisting of a wide range of analytics. These include time series plots, statistical summaries, data mining techniques, trend and pattern detection, and hypothesis generation.

59 BASIC BIOLOGICAL SCIENCES↗

Graph-based featurization methods for classifying small molecule compounds

For over a decade, drug-induced liver injury (DILI) has posed significant drawbacks in the synthesis and development of drugs and remains a consequential concern. With finite success within the existing preclinical models, DILI is one of the main causes of drug withdrawal or termination from the market. Particularly, this withdrawal occurs during the late stages of drug development (Kullak-Ublick, 2017). Since DILI is difficult to diagnose and treat, it has become an obstacle in the drug production market that in turn affects clinicians, pharmaceutical companies, and consumers. We propose a method for learning features of DILI-positive drugs based on the graphical relationships and patterns they possess within a network of biological databases. We also train various statistical and machine learning models on these learned features in order to classify the drugs as DILI-positive or negative. Our methods include Random Forest, Neural networks, and logistic regression classification. We utilize labeled DILI-positive and DILI-negative datasets, which were developed by the FDA and the National center for toxicological research, as well as additional literature datasets (Thakkar, 2020) in order to validate our results and assess our featurization and model accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Optimization Studies of Radiation Shielding for the PIP-II Project at Fermilab

The PIP-II project at Fermilab, which includes an 800-MeV superconducting LINAC, demands rigorous radiation shielding optimization to meet safety requirements. We updated the MARS geometry model to reflect new magnet and collimator designs and introduced high-resolution detector planes to better capture radiation field distributions. To overcome the significant computational demands, we implemented a well-known branching technique that drastically reduced simulation runtimes while maintaining statistical integrity. This was achieved through particle splitting and the application of Russian Roulette techniques. Additionally, new graphical tools were created to streamline data visualization and MARS code usability.

Makovec, Alajos [Fermilab] (ORCID:0000000286157492↗

Big Data Analytics for Long-Term Meteorological Observations at Hanford Site

A growing number of physical objects with embedded sensors with typically high volume and frequently updated data sets has accentuated the need to develop methodologies to extract useful information from big data for supporting decision making. This study applies a suite of data analytics and core principles of data science to characterize near real-time meteorological data with a focus on extreme weather events. To highlight the applicability of this work and make it more accessible from a risk management perspective, a foundation for a software platform with an intuitive Graphical User Interface (GUI) was developed to access and analyze data from a decommissioned nuclear production complex operated by the U.S. Department of Energy (DOE, Richland, USA). Exploratory data analysis (EDA), involving classical non-parametric statistics, and machine learning (ML) techniques, were used to develop statistical summaries and learn characteristic features of key weather patterns and signatures. The new approach and GUI provide key insights into using big data and ML to assist site operation related to safety management strategies for extreme weather events. Specifically, this work offers a practical guide to analyzing long-term meteorological data and highlights the integration of ML and classical statistics to applied risk and decision science.

54 ENVIRONMENTAL SCIENCES↗

Gauges, loops, and polynomials for partition functions of graphical models

Graphical models represent multivariate and generally not normalized probability distributions. Computing the normalization factor, called the partition function, is the main inference challenge relevant to multiple statistical and optimization applications. The problem is #P-hard that is of an exponential complexity with respect to the number of variables. Here, aimed at approximating the partition function, we consider multi-graph models where binary variables and multivariable factors are associated with edges and nodes, respectively, of an undirected multi-graph. We suggest a new methodology for analysis and computations that combines the Gauge function technique from Chertkov and Chernyak with the technique developed in Anari and Oveis Gharan 2017 arXiv:1702.02937; Gurvits 2011 arXiv:1106.2844; Straszak and Vishnoi 2017 55th Annual Allerton Conf. on Communication, Control, and Computing, based on the recent progress in the field of real stable polynomials. We show that the Gauge function, representing a single-out term in a finite sum expression for the partition function which achieves extremum at the so-called belief-propagation gauge, has a natural polynomial representation in terms of gauges/variables associated with edges of the multi-graph. Moreover, Gauge function can be used to recover the partition function through a sequence of transformations allowing appealing algebraic and graphical interpretations. Algebraically, one step in the sequence consists of the application of a differential operator over gauges associated with an edge. Graphically, the sequence is interpreted as a repetitive elimination/contraction of edges resulting in multi-graph models on decreasing in size (number of edges) graphs with the same partition function as in the original multi-graph model. Even though the complexity of computing factors in the sequence of the derived multi-graph models and respective Gauge functions grow exponentially with the number of eliminated edges, polynomials associated with the new factors remain bi-stable if the original factors have this property. Moreover, we show that BP estimations in the sequence do not decrease, each low-bounding the partition function.

97 MATHEMATICS AND COMPUTING↗

Search for new physics in high-mass diphoton events from proton-proton collisions at $ \sqrt{\textrm{s}} $ = 13 TeV

Results are presented from a search for new physics in high-mass diphoton events from proton-proton collisions at $ \sqrt{s} $ = 13 TeV. The data set was collected in 2016–2018 with the CMS detector at the LHC and corresponds to an integrated luminosity of 138 fb$^{−1}$. Events with a diphoton invariant mass greater than 500 GeV are considered. Two different techniques are used to predict the standard model backgrounds: parametric fits to the smoothly-falling background and a first-principles calculation of the standard model diphoton spectrum at next-to-next-to-leading order in perturbative quantum chromodynamics calculations. The first technique is sensitive to resonant excesses while the second technique can identify broad differences in the invariant mass shape. The data are used to constrain the production of heavy Higgs bosons, Randall-Sundrum gravitons, the large extra dimensions model of Arkani-Hamed, Dimopoulos, and Dvali (ADD), and the continuum clockwork mechanism. No statistically significant excess is observed. The present results are the strongest limits to date on ADD extra dimensions and RS gravitons with a coupling parameter greater than 0.1.[graphic not available: see fulltext]

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Sparsity-Independent Lyapunov Exponent in the Sachdev-Ye-Kitaev Model

The saturation of a recently proposed universal bound on the Lyapunov exponent has been conjectured to signal the existence of a gravity dual. This saturation occurs in the low-temperature limit of the dense Sachdev-Ye-Kitaev (SYK) model, N Majorana fermions with q body ( q > 2 ) infinite-range interactions. We calculate certain out-of-time-order correlators (OTOCs) for N ≤ 64 fermions for a highly sparse SYK model and find no significant dependence of the Lyapunov exponent on sparsity up to near the percolation limit where the Hamiltonian breaks up into blocks. This provides strong support to the saturation of the Lyapunov exponent in the low-temperature limit of the sparse SYK. A key ingredient to reaching N = 64 is the development of a novel quantum spin model simulation library that implements highly optimized matrix-free Krylov subspace methods on graphical processing units. This leads to a significantly lower simulation time as well as vastly reduced memory usage over previous approaches, while using modest computational resources. Strong sparsity-driven statistical fluctuations require both the use of a much larger number of disorder realizations with respect to the dense limit and a careful finite size scaling analysis. The saturation of the bound in the sparse SYK points to the existence of a gravity analog that would enlarge substantially the number of field theories with this feature. Published by the American Physical Society 2024

Physics↗

BioXTAS RAW 2 : new developments for a free open-source program for small-angle scattering data reduction and analysis

BioXTAS RAW is a free open-source program for reduction, analysis and modelling of biological small-angle scattering data. Here, the new developments in RAW version 2 are described. These include improved data reduction using pyFAI ; updated automated Guinier fitting and D max finding algorithms; automated series ( e.g. size-exclusion chromatography coupled small-angle X-ray scattering or SEC-SAXS) buffer- and sample-region finding algorithms; linear and integral baseline correction for series; deconvolution of series data using regularized alternating least squares ( REGALS ); creation of electron-density reconstructions using electron density via solution scattering ( DENSS ); a comparison window showing residuals, ratios and statistical comparisons between profiles; and generation of PDF reports with summary plots and tables for all analysis. Furthermore, there is now a RAW API, which can be used without the graphical user interface (GUI), providing full access to all of the functionality found in the GUI. In addition to these new capabilities, RAW has undergone significant technical updates, such as adding Python 3 compatibility, and has entirely new documentation available both online and in the program.

97 MATHEMATICS AND COMPUTING↗

Bayesian chain graph models to characterize microbe-environment dynamics

Microbiome data require statistical models that can simultaneously decode microbes' reaction to the environment and interactions among microbes. While a multiresponse linear regression model seems like a straight-forward solution, we argue that treating it as a graphical model is problematic given that the regression coefficient matrix does not encode the conditional dependence structure between response and predictor nodes. This observation is especially important in biological settings when we have prior knowledge on the edges from specific experimental interventions that can only be properly encoded under a conditional dependence model. Here, we propose a chain graph model with two sets of nodes (predictors and responses) whose solution yields a graph with edges that indeed represent conditional dependence, thus agreeing with the experimenter's intuition on the average behavior of nodes under treatment. The solution to our model is sparse via the Bayesian linear regression (LASSO). In addition, we propose an adaptive extension so that different shrinkages can be applied to different edges to incorporate edge-specific prior knowledge. Our model is computationally inexpensive through an efficient Gibbs sampling algorithm and can account for binary, counting, and compositional responses via an appropriate hierarchical structure. We test the performance of our model in a variety of simulated datasets, thereby showing superior performance to state-of-the-art approaches. We further apply our model to human gut and soil microbial compositional datasets, and we highlight that CG-LASSO can estimate biologically meaningful network structures in the data.

compositional data↗

High-Performance Semiempirical Excited-State Molecular Dynamics Powered by Graphics Processing Units

Here, this Letter introduces excited-state molecular dynamics in PYSEQM, a GPU-accelerated semiempirical quantum chemistry engine implemented in PyTorch. The new module enables Born–Oppenheimer molecular dynamics (BOMD) using configuration-interaction singles and random phase approximation for excited states, allowing long trajectories and large statistical ensembles to be simulated efficiently on a single GPU. We also implement an extended Lagrangian excited-state BOMD (XL-ESMD) scheme that propagates auxiliary electronic variables, enabling relaxed ground and excited-state convergence thresholds without compromising energy conservation. The excited-state BOMD implementation scales smoothly from small chromophores to a nearly 900-atom dendrimer (taking 6.5 s per MD step). PYSEQM also supports batched execution, allowing many geometries or trajectories to be evaluated in a single GPU launch, substantially increasing throughput and making ensemble-based protocols routine. As a demonstration, we compute absorption, emission, and infrared spectra from trajectories propagated on the ground and first excited states. The XL-ESMD scheme yields identical spectra at significantly lower computational cost, establishing the role of extended Lagrangian based dynamics for efficient excited-state BOMD simulations. Beyond raw performance, PYSEQM’s PyTorch foundation provides automatic differentiation for forces, efficient GPU batching, and seamless interfacing with machine learning models. These capabilities position PYSEQM as a practical platform for machine learning-augmented excited-state dynamics and lay the foundation for future data-driven nonadiabatic excited-state dynamics modeling of ultrafast spectroscopic probes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Imaging Bragg Edge Analysis TooLs for Engineering Structures (iBeatles)

The Spallation Neutron Source (SNS) at Oak Ridge National Laboratory (ORNL) provides pulsed neutrons with energies varying from epithermal to cold. In preparation for VENUS, the neutron imaging beamline to be located at beam port 10, we have performed a series of experiments focused on wavelength-dependent radiography and computed tomography for a broad range of applications, from materials science to biological tissues.One of the time-of-flight (TOF) techniques that is of interest to the scientific community is the 2-dimensional mapping of phases and average crystalline plane orientation in samples both ex-situ and during applied stresses such as tensile loading and heating. This technique is known as Bragg edgeimaging and relies on the identification of changes of transmission values, fitting of the edge to measure its displacement, and thus identify the shift in lattice parameter due to stresses. One of the challenges of TOF imaging measurements is the amount of data and the inability to observe Bragg edge shifts in real time during an experiment. Thus, we have been focusing on creating a Python-based interface that allows fast data processing and instantaneous mapping and fitting of the Bragg edges, and their evolution through time. Python libraries and Jupyter notebooks have been implemented to facilitate decision making during an experiment. The advantage of the notebooks is the possibility to guide an experiment as they can quickly process and display Bragg edge data. These notebooks can be used independently, or can be combined in a Python Graphical User Interface (GUI) tool called iBeatles. This interface permits visualization and fitting of the Bragg edges, and ultimately back-projects the fitting results onto the radiographs to display a strain map. Assuming data collection has sufficient statistics, the strain mapping analysis can be performed on a pixel-by-pixel basis. This development is a step forward toward a better user experience at the future VENUS beamline in terms of live feedback and productivity. Analysis that used to take days of switching between different applications can now be done in minutes within the

Bilheux, JeanChristophe [Oak Ridge National Labora↗