Engineering PapersSearch

SEARCH · Engineering Papers

Results for “functional data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Window Observables for Benchmarking Parton Distribution Functions

Global analysis of collider and fixed-target experimental data and calculations from lattice quantum chromodynamics (QCD) are used to gain complementary information on the structure of hadrons. We propose novel “window observables” that allow for higher precision cross-validation between the different approaches, a critical step for studies that wish to combine the datasets. Global analyses are limited by the kinematic regions accessible to experiment, particularly in a range of Bjorken-𝑥, and lattice QCD calculations also have limitations requiring extrapolations to obtain the parton distributions. We provide two different window observables that can be defined within a region of 𝑥 where extrapolations and interpolations in global analyses remain reliable and where lattice QCD results retain sensitivity and precision.

lattice QCD

Visualization of Noisy and Less Noisy Computational Basis States in Quantum Computing

Quantum computing technology holds substantial promise as a reliable computational paradigm. However, current noisy intermediate scale quantum (NISQ) systems, are significantly impacted by noise originating from hardware inconsistencies. This noise causes errors and lowers output fidelity. So we must find which basis states cause errors. However, there are two main challenges in analyzing noise corresponding to basis states. First, the noise distribution data is high dimensional in nature, thereby making its analysis challenging. Second, although functional box plots have been used in the state of the art research to understand such a high dimensional data, they suffer from clutter and occlusion issues because of overplotting. In this study, we introduce an innovative visualization pipeline to address the aforementioned challenges to provide a clear depiction of noisy and less-noisy basis states. Specifically, our proposed visualization pipeline comprises three stages namely, low dimensional embedding, clustering, and violin plot visualization, to reduce visual clutter and effectively analyze high-dimensional noise distribution data. Our analysis uses quantum machine learning (QML) circuits as case study for drawing a distinction between noisy and less noisy basis states.

Senapati, Priyabrata [Kent State University]

Window observables for benchmarking parton distribution functions

Global analysis of collider and fixed-target experimental data and calculations from lattice quantum chromodynamics (QCD) are used to gain complementary information on the structure of hadrons. We propose novel ``window observables'' that allow for higher precision cross-validation between the different approaches, a critical step for studies that wish to combine the datasets. Global analyses are limited by the kinematic regions accessible to experiment, particularly in a range of Bjorken-x, and lattice QCD calculations also have limitations requiring extrapolations to obtain the parton distributions. We provide two different ``window observables'' that can be defined within a region of x where extrapolations and interpolations in global analyses remain reliable and where lattice QCD results retain sensitivity and precision.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Analysis of fusion alphas interaction with RF waves in D-T plasma at JET

This work studies the influence of RF waves in ICRH range of frequency on fusion alphas during the recent JET D-T campaign. Fusion alphas from D-T reactions are born with energies of about 3.5MeV and therefore have significant Doppler shift enabling synergistic interaction between them and RF waves at broad range of frequencies including the ones foreseen for future fusion machines ITER and SPARC. Resonant interaction between RF waves and alphas, also called synergistic effects, will modify the alpha distribution and ultimately will have an impact on alpha orbit losses and heating. Data from JET 3.43T/2.3MA pulses based on the hybrid scenario during the DTE2 campaign were used for the analysis in this study. The impact of synergistic effects on alpha orbit losses and alpha heating is assessed. Conclusions are based on analysis of experimental data for fast alphas losses, i.e. measurements from neutral particle analyser, fast ion losses scintillator detector, Faraday cups, and TRANSP simulations. Experimental data and TRANSP analysis indicate that there are indeed changes in the alphas' distribution function due to interaction with RF waves. Data from the scintillator detector and the Faraday cups were compared for pulses with and without ICRH power and versus cases with enhanced alpha losses due to MHD activities. The trends from these diagnostics consistently show no additional alpha losses due to interaction with RF waves. TRANSP predictions for the impact of the synergistic effects on alpha heating show up to 42% increase in alpha electron heating and up to 25% increase in alpha ion heating. These effects however become negligibly small, less than 1%, when alpha heating is compared to the total auxiliary hearting power in the investigated JET pulses.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Operando pair distribution function analysis of nanocrystalline functional materials: the case of TiO 2 -bronze nanocrystals in Li-ion battery electrodes

Structural modelling of operando pair distribution function (PDF) data of complex functional materials can be highly challenging. To aid the understanding of complex operando PDF data, this article demonstrates a toolbox for PDF analysis. The tools include denoising using principal component analysis together with the structureMining , similarityMapping and nmfMapping apps available through the online service `PDF in the cloud' ( PDFitc , https://pdfitc.org/). The toolbox is used for both ex situ and operando PDF data for 3 nm TiO 2 -bronze nanocrystals, which function as the active electrode material in a Li-ion battery. The tools enable structural modelling of the ex situ and operando PDF data, revealing two pristine TiO 2 phases (bronze and anatase) and two lithiated Li x TiO 2 phases (lithiated versions of bronze and anatase), and the phase evolution during galvanostatic cycling is characterized.

Chemistry

Uncertainty Visualization of Critical Points of 2D Scalar Fields for Parametric and Nonparametric Probabilistic Models

This paper presents a novel end-to-end framework for closed-form computation and visualization of critical point uncertainty in 2D uncertain scalar fields. Critical points are fundamental topological descriptors used in the visualization and analysis of scalar fields. The uncertainty inherent in data (e.g., observational and experimental data, approximations in simulations, and compression), however, creates uncertainty regarding critical point positions. Uncertainty in critical point positions, therefore, cannot be ignored, given their impact on downstream data analysis tasks. Here, in this work, we study uncertainty in critical points as a function of uncertainty in data modeled with probability distributions. Although Monte Carlo (MC) sampling techniques have been used in prior studies to quantify critical point uncertainty, they are often expensive and are infrequently used in production-quality visualization software. We, therefore, propose a new end-to-end framework to address these challenges that comprises a threefold contribution. First, we derive the critical point uncertainty in closed form, which is more accurate and efficient than the conventional MC sampling methods. Specifically, we provide the closed-form and semianalytical (a mix of closed-form and MC methods) solutions for parametric (e.g., uniform, Epanechnikov) and nonparametric models (e.g., histograms) with finite support. Second, we accelerate critical point probability computations using a parallel implementation with the VTK-m library, which is platform portable. Finally, we demonstrate the integration of our implementation with the ParaView software system to demonstrate near-real-time results for real datasets.

97 MATHEMATICS AND COMPUTING

DESI DR1 Ly$α$ forest: 3D full-shape analysis and cosmological constraints

We perform an analysis of the full shapes of Lyman-$α$ (Ly$α$) forest correlation functions measured from the first data release (DR1) of the Dark Energy Spectroscopic Instrument (DESI). Our analysis focuses on measuring the Alcock-Paczynski (AP) effect and the cosmic growth rate times the amplitude of matter fluctuations in spheres of $8$$h^{-1}\text{Mpc}$, $fσ_8$. We validate our measurements using two different sets of mocks, a series of data splits, and a large set of analysis variations, which were first performed blinded. Our analysis constrains the ratio $D_M/D_H(z_\mathrm{eff})=4.525\pm0.071$, where $D_H=c/H(z)$ is the Hubble distance, $D_M$ is the transverse comoving distance, and the effective redshift is $z_\mathrm{eff}=2.33$. This is a factor of $2.4$ tighter than the Baryon Acoustic Oscillation (BAO) constraint from the same data. When combining with Ly$α$ BAO constraints from DESI DR2, we obtain the ratios $D_H(z_\mathrm{eff})/r_d=8.646\pm0.077$ and $D_M(z_\mathrm{eff})/r_d=38.90\pm0.38$, where $r_d$ is the sound horizon at the drag epoch. We also measure $fσ_8(z_\mathrm{eff}) = 0.37\; ^{+0.055}_{-0.065} \,(\mathrm{stat})\, \pm 0.033 \,(\mathrm{sys})$, but we do not use it for cosmological inference due to difficulties in its validation with mocks. In $Λ$CDM, our measurements are consistent with both cosmic microwave background (CMB) and galaxy clustering constraints. Using a nucleosynthesis prior but no CMB anisotropy information, we measure the Hubble constant to be $H_0 = 68.3\pm 1.6\;\,{\rm km\,s^{-1}\,Mpc^{-1}}$ within $Λ$CDM. Finally, we show that Ly$α$ forest AP measurements can help improve constraints on the dark energy equation of state, and are expected to play an important role in upcoming DESI analyses.

Cuceu, Andrei [LBL, Berkeley; Chicago U., KICP] (O

Extending quantum-mechanical benchmark accuracy to biological ligand-pocket interactions

Predicting the binding affinity of ligands to protein pockets is key in the drug design pipeline. The flexibility of ligand-pocket motifs arises from a range of attractive and repulsive electronic interactions during binding. Accurately accounting for all interactions requires robust quantum-mechanical (QM) benchmarks, which are scarce for ligand-pocket systems. Additionally, disagreement between “gold standard” Coupled Cluster (CC) and Quantum Monte Carlo (QMC) methods casts doubt on many benchmarks for larger non-covalent systems. We introduce the “QUantum Interacting Dimer” (QUID) benchmark framework containing 170 non-covalent (non-)equilibrium systems modeling chemically and structurally diverse ligand-pocket motifs. Symmetry-adapted perturbation theory shows that QUID broadly covers non-covalent binding motifs and energetic contributions. Robust binding energies are obtained using complementary CC and QMC methods, achieving agreement of 0.5 kcal/mol. The benchmark data analysis reveals that several dispersion-inclusive density functional approximations provide accurate energy predictions, though their atomic van der Waals forces differ in magnitude and orientation. Contrarily, semiempirical methods and empirical force fields require improvements in capturing non-covalent interactions (NCIs) for out-of-equilibrium geometries. The wide span of NCIs, highly accurate interaction energies, and analysis of molecular properties take QUID beyond the “gold standard” for QM benchmarks of ligand-protein systems.

Puleva, Mirela [University of Luxembourg, Luxembou

Deviations from the Porter-Thomas Distribution due to Nonstatistical 𝛾 Decay below the 150 Nd Neutron Separation Threshold

We introduce a new method for the study of fluctuations of partial transition widths based on nuclear resonance fluorescence experiments with quasimonochromatic linearly polarized photon beams below particle separation thresholds. It is based on the average branching of decays of 𝐽=1 states of an even-even nucleus to the 2$^{+}_{1}$ state in comparison to the ground state. Between 5 and 7 MeV, a constant average branching ratio for 𝛾 decays from 1 − states of 0.490(16) is observed for the nuclide 150 Nd. Assuming 𝜒 2 -distributed partial transition widths, this average branching ratio is related to a degree of freedom of 𝜈 = 1.93⁢(12), rejecting the validity of the Porter-Thomas distribution, requiring 𝜈 = 1. The observed deviation can be explained by nonstatistical effects in the 𝛾-decay behavior with contributions in the range of 9.4(10)% up to 94(10)%.

150 ≤ A ≤ 189

Strong coupling from hadronic τ -decay data including τ → π − π 0 ν τ from Belle

In previous work we have combined the π − π 0 , 2 π − π + π 0 , and π − 3 π 0 spectral data obtained from hadronic τ decays measured by the ALEPH and OPAL experiments, together with electroproduction data for several of the subleading hadronic modes and data for the K K ¯ mode to construct an inclusive nonstrange vector spectral function entirely based on experimental data, with no Monte-Carlo generated input. In this paper, we include, for the first time, the Belle τ → π − π 0 ν τ high-statistics decay data to construct a new inclusive nonstrange vector spectral function that combines more of the world’s available data. As no Belle data are at present available for the two 4 π modes, this requires a revised data analysis in comparison with our previous work. From the resulting new spectral function, we obtain a new determination of the strong coupling, α s , using our previously developed strategy based on finite-energy sum rules. We find, at the Z mass scale, α s ( m Z 2 ) = 0.1159 ( 14 ) . We discuss the smaller central value and larger error of our new result compared to our previous result, showing the shifts to be due mainly to significant changes in updated HFLAV results for the π − 3 π 0 decay mode. Published by the American Physical Society 2025

Boito, Diogo (ORCID:0000000244267984)

Data for Spatial Analysis of Cell Patterning to Aid Genetic and Phenotypic Understanding of Grass Stomatal Density: A Case Study in Maize

Biological processes involve complex hierarchies where composite traits result from multiple component traits. However, holistically understanding of how sets of component traits interact to underpin genotype-to-phenotype relationships is generally lacking. Stomatal density (SD) is a tractable model system for exploring how high-throughput phenotyping (HTP) data could be exploited by a new spatial analysis approach to better understand a developmentally and functionally important trait. SD is a composite trait, resulting from various components related to cell identity and size, which are themselves governed by a series of spatio-developmental processes. Data from 192 recombinant inbred lines of maize [Zea mays (L.)] were analyzed by a new stomatal patterning phenotype (SPP) to (1) describe the average spatial probability distribution of the nearest neighboring stomata; (2) derive a core set of component traits related to cell size, cell packing, and positional probabilities; (3) build a structural equation model of component traits underlying SD; and (4) identify stomatal patterning quantitative trait loci (QTL). The core set of SPP-derived traits explained 74% of the variation in SD. Analyzing SPP component traits allowed some loci previously identified as generic SD QTL to be recognized as specific to lateral versus longitudinal elements of stomatal patterning. Therefore, this study highlights how novel insights can be gained by decomposing a composite trait (e.g., SD) into a set of component traits that were present in HTP data but not previously exploited.

AI/ML

Short and medium range structure in elastic deformation of metallic and covalent glasses

Here, we present a concise methodology to analyze structural response to the applied stress in amorphous solids, including metallic glasses (MG), glassy selenium, silica and polycarbonate, using high energy x-ray diffraction and atomic pair distribution function (PDF) analysis. To assess the structural anisotropy induced by applied axial stress, diffraction data were expanded into spherical harmonics. Using Bessel transformation, components of the structure function were converted into isotropic and anisotropic PDFs. The PDFs were compared to the expected model behavior for ideal elastic deformation to separate homogeneous affine strain from local non-affine strains. In metallic glass the range of non-affine deformation is limited to the nearest neighbor shell, suggesting local strain relaxation under stress that occurs even in the elastic regime. Beyond the second atomic shell strain is uniform. However, in glassy silica, polycarbonate and selenium strong local bonding inhibits local displacements and strain in short range order is accommodated by rotation of local units. Interestingly, beyond a molecular unit, deformation in covalent systems is similar to MG, and response of the medium range order scales with the macroscopic stress.

glassy structure

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES

A Life Cycle Analysis Framework for Point Source Capture Systems

NETL studies the costs and benefits of PSC for electricity, industry, and mobile applications. Mobile point source capture (MPSC) and storage applied to freight modes captures emissions directly from exhaust. This poster presents a framework for conducting LCA of PSC systems applied to heavy-duty trucks, freight trains, and marine vessels. The framework defines a wheels-to-storage (gate-to-grave) boundary, including energy demands (electricity, heat, and cooling requirements), solvent use and cycling, onboard system components, carbon storage in a saline aquifer, and upstream manufacturing impacts for equipment, with a suggested functional unit of 1 tonne-km. Potential data sources for analysis include material, energy, and operational data from Oak Ridge National Laboratory, GREET (Greenhouse gases, Regulated Emissions, and Energy use in Technologies) model, and scientific literature. The suggested analytical approach includes comparison to publicly available business-as-usual systems without capture across all modes of transportation, sensitivity to composition of the capture solvent, and sensitivity to capture rate variation, all of which would support a wholistic PSC business case analysis. For future consideration, analysis can be augmented with consideration of different sources of electricity (e.g., nuclear), fuel substitution, deploying supportive infrastructure such as pipeline offloading points, and downstream applications like enhanced oil recovery (EOR).

life cycle analysis (LCA)

Spin structure of the proton from global QCD analysis

In this talk we review recent results for spin-dependent parton distribution functions extracted in global QCD analysis of high energy scattering data by the JAM collaboration, including inclusive and semi-inclusive deep-inelastic scattering, jet and weak boson production in polarised hadron-hadron collisions. In particular, we focus on the determination of the gluon polarisation in the proton, whose sign and magnitude have been the subject of debate recently.

Melnitchouk, Wally [Thomas Jefferson National Acce

A road map to cosmological parameter analysis with third-order shear statistics: III. Efficient estimation of third-order shear correlation functions and an application to the KiDS-1000 data

Context. Third-order lensing statistics contain a wealth of cosmological information that is not captured by second-order statistics. However, the computational effort it takes to estimate such statistics in forthcoming stage IV surveys is prohibitively expensive. Aims. We derive and validate an efficient estimation procedure for the three-point correlation function (3PCF) of polar fields such as weak lensing shear. We then use our approach to measure the shear 3PCF and the third-order aperture mass statistics on the KiDS-1000 survey. Methods We constructed an efficient estimator for third-order shear statistics that builds on the multipole decomposition of the 3PCF. We then validated our estimator on mock ellipticity catalogs obtained from N -body simulations. Finally, we applied our estimator to the KiDS-1000 data and presented a measurement of the third-order aperture statistics in a tomographic setup. Results. Our estimator provides a speedup of a factor of ∼100–1000 compared to the state-of-the-art estimation procedures. It is also able to provide accurate measurements for squeezed and folded triangle configurations without additional computational effort. We report a significant detection of tomographic third-order aperture mass statistics in the KiDS-1000 data (S/N = 6.69). Conclusions. Our estimator will make it computationally feasible to measure third-order shear statistics in forthcoming stage IV surveys. Furthermore, it can be used to construct empirical covariance matrices for such statistics.

Astronomy & Astrophysics

Isospin dependence of the nuclear EMC effect from a global QCD analysis

We perform a new global QCD analysis of unpolarized parton distribution functions (PDFs) in the nucleon from proton, deuteron, and A = 3 data, including recent measurements of He 3 / D and H 3 / D cross section ratios from the MARATHON experiment at Jefferson Lab. Simultaneously inferring the PDFs and nucleon off-shell corrections allows both to be determined consistently, without theoretical assumptions about the isospin dependence of nuclear effects. The analysis provides strong evidence for the need of nucleon off-shell corrections to describe the A = 3 data, with large isoscalar and a suggestion of nonzero isovector contributions in A ≤ 3 nuclei. We find that the extracted EMC ratios of nuclear to nucleon structure functions for A = 2 and 3 differ from those naively extrapolated from heavy nuclei down to low A .

Cocuzza, C. [William & Mary] (ORCID:00000003492292

JGI-Trichoderma v1.0

There is a series of Python and bash scripts to parse genomics datasets used to evaluate the coevolution of gene families and the feature importance of gene families using an SVM classifier. - Cover analysis: takes a list of single-copy genes in a set of genomes, aligns and builds the gene trees to determine if two gene families have a signature of covariation with one another. It parses the files to run phykit cover script described here: https://jlsteenwyk.com/PhyKIT/usage/index.html - SVM-classifier: This Python script is an SVM-based genomic classifier designed for biological data analysis. It combines machine learning with feature selection to identify important genomic markers and classify biological samples. Core Functionality: The script uses Support Vector Machines from scikit-learn to classify genomic data, incorporating SelectKBest for automated feature selection and leave-one-out cross-validation for performance assessment. It operates in multiple modes: feature ranking, optimal combination discovery, and sample prediction. Primary Applications: Genomic sample classification and biomarker discovery Feature importance analysis in high-dimensional biological datasets Prediction of sample categories based on genomic profiles Research applications requiring robust classification of biological data Key Advantages: High-dimensional handling: SVMs excel with genomic data's typical high feature-to-sample ratios Integrated feature selection: Reduces noise and computational overhead while identifying key markers Probability estimation: Provides confidence scores essential for biological interpretation Validation robustness: Leave-one-out cross-validation ensures reliable performance metrics Operational flexibility: Multiple analysis modes support different research phases from exploration to prediction

Stecca Steindorff, Andrei [Lawrence Berkeley Natio