Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Functional data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Data-driven analysis of dipole strength functions using artificial neural networks

Here, we present a data-driven analysis of dipole strength functions across the nuclear chart, employing an artificial neural network to model nuclear dipole responses. We train the network on a dataset of experimentally measured dipole strength functions for 216 different nuclei. To assess its predictive capability, we test the trained model on an additional set of 10 new nuclei, where experimental data exist. We demonstrate that the artificial neural network not only accurately reproduces known data but also identifies potential inconsistencies in experimental datasets, indicating which results may warrant further review or possible rejection. For nuclei where experimental data are sparse or unavailable, the network confirms theoretical calculations, reinforcing its utility as a predictive tool in nuclear physics. Finally, utilizing the predicted electric dipole polarizability, we extract the value of the symmetry energy at saturation density and find it consistent with results from the literature.

artificial neural networks↗

VEESA R package

SAND2024-04584O R package for applying the VEESA pipeline method is a technique used for explainable machine learning with functional data. The VEESA pipeline makes use of the elastic-shape analysis framework for functional data. It also implements functional principal component analysis and permutation feature importance. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Tucker, James↗

Application of an energy-dependent instrument response function to analysis of nTOF data from cryogenic DT experiments

Neutron time-of-flight (nTOF) detectors are used to diagnose the conditions present in inertial confinement fusion (ICF) experiments and basic laboratory physics experiments performed on an ICF platform. The instrument response function (IRF) of these detectors is constructed by convolution of two components: an x-ray IRF and a neutron interaction response. The shape of the neutron interaction response varies with incident neutron energy, changing the shape of the total IRF. Analyses of nTOF data that span a broad range of energies must account for this energy-dependence in order to accurately infer plasma parameters and nuclear properties in ICF experiments. This work briefly reviews a matrix multiplication approach to convolution which allows for an energy-dependent change in the shape of the IRF. This method is applied to synthetic data resembling symmetric cryogenic DT implosions to examine the effect of the energy-dependent IRF on the inferred areal density. Here, results of forward fits that infer ion temperatures and areal densities from nTOF data collected during cryogenic DT experiments on OMEGA are also discussed.

47 OTHER INSTRUMENTATION↗

Machine learning analysis of RB-TnSeq fitness data predicts functional gene modules in Pseudomonas putida KT2440

ABSTRACT There is growing interest in engineering Pseudomonas putida KT2440 as a microbial chassis for the conversion of renewable and waste-based feedstocks, and metabolic engineering of P. putida relies on the understanding of the functional relationships between genes. In this work, independent component analysis (ICA) was applied to a compendium of existing fitness data from randomly barcoded transposon insertion sequencing (RB-TnSeq) of P. putida KT2440 grown in 179 unique experimental conditions. ICA identified 84 independent groups of genes, which we call fModules (“functional modules”), where gene members displayed shared functional influence in a specific cellular process. This machine learning-based approach both successfully recapitulated previously characterized functional relationships and established hitherto unknown associations between genes. Selected gene members from fModules for hydroxycinnamate metabolism and stress resistance, acetyl coenzyme A assimilation, and nitrogen metabolism were validated with engineered mutants of P. putida . Additionally, functional gene clusters from ICA of RB-TnSeq data sets were compared with regulatory gene clusters from prior ICA of RNAseq data sets to draw connections between gene regulation and function. Because ICA profiles the functional role of several distinct gene networks simultaneously, it can reduce the time required to annotate gene function relative to manual curation of RB-TnSeq data sets. IMPORTANCE This study demonstrates a rapid, automated approach for elucidating functional modules within complex genetic networks. While Pseudomonas putida randomly barcoded transposon insertion sequencing data were used as a proof of concept, this approach is applicable to any organism with existing functional genomics data sets and may serve as a useful tool for many valuable applications, such as guiding metabolic engineering efforts in other microbes or understanding functional relationships between virulence-associated genes in pathogenic microbes. Furthermore, this work demonstrates that comparison of data obtained from independent component analysis of transcriptomics and gene fitness datasets can elucidate regulatory-functional relationships between genes, which may have utility in a variety of applications, such as metabolic modeling, strain engineering, or identification of antimicrobial drug targets.

09 BIOMASS FUELS↗

Elastic Bayesian Model Calibration

Functional data are ubiquitous in scientific modeling. For instance, quantities of interest are modeled as functions of time, space, energy, density, etc. Uncertainty quantification methods for computer models with functional response have resulted in tools for emulation, sensitivity analysis, and calibration that are widely used. However, many of these tools do not perform well when the computer model’s parameters control both the amplitude variation of the functional output and its alignment (or phase variation). This paper introduces a framework for Bayesian model calibration when the model responses are misaligned functional data. The approach generates two types of data out of the misaligned functional responses: (1) aligned functions so that the amplitude variation is isolated and (2) warping functions that isolate the phase variation. These two types of data are created for the computer simulation data (both of which may be emulated) and the experimental data. The calibration approach uses both types so that it seeks to match both the amplitude and phase of the experimental data. The framework is careful to respect constraints that arise, especially when modeling phase variation, and is framed in a way that it can be done with readily available calibration software. In conclusion, we demonstrate the techniques on two simulated data examples and on two dynamic material science problems: a strength model calibration using flyer plate experiments and an equation of state model calibration using experiments performed on the Sandia National Laboratories’ Z-machine.

97 MATHEMATICS AND COMPUTING↗

Uncertainty Quantification for Smooth Functional Data with Application to Material Properties

This document outlines a method for processing functional output (i.e., curves) for the ultimate purpose of sampling curves under specified input conditions for use in modeling and simulation uncertainty quantification (UQ) studies. A set of benchmark curves sufficiently representative of the relevant scenario(s) being simulated are provided to the process and formatted as described in Section 1. Principal Component Analysis (PCA) is utilized to discover the components of uncertainty in the benchmark curves and is outlined in Section 2. Section 3 describes the application of uncertainty quantification to the PCA results for the purpose of sampling curves to be used in UQ analysis. Section 4 applies these techniques to an example benchmark dataset. Concluding remarks are provided in the final section.

36 MATERIALS SCIENCE↗

Multimodal Bayesian registration of noisy functions using Hamiltonian Monte Carlo

Functional data registration is a necessary processing step for many applications. The observed data can be inherently noisy, often due to measurement error or natural process uncertainty; which most functional alignment methods cannot handle. A pair of functions can also have multiple optimal alignment solutions, which is not addressed in current literature. In this paper, a flexible Bayesian approach to functional alignment is presented, which appropriately accounts for noise in the data without any pre-smoothing required. Additionally, by running parallel MCMC chains, the method can account for multiple optimal alignments via the multi-modal posterior distribution of the warping functions. To most efficiently sample the warping functions, the approach relies on a modification of the standard Hamiltonian Monte Carlo to be well-defined on the infinite-dimensional Hilbert space. In this work, this flexible Bayesian alignment method is applied to both simulated data and real data sets to show its efficiency in handling noisy functions and successfully accounting for multiple optimal alignments in the posterior; characterizing the uncertainty surrounding the warping functions.

97 MATHEMATICS AND COMPUTING↗

Data processing pipeline for Tianlai experiment

The Tianlai project is a 21cm intensity mapping experiment for detecting dark energy by measuring the baryon acoustic oscillation (BAO) features in the large scale structure power spectrum. This experiment provides an opportunity to test the data processing methods for cosmological 21cm signal extraction, which is still a great challenge in current radio astronomy research. The 21cm signal is much weaker than the foregrounds and easily aected by the imperfections in the instrumental responses. Furthermore, processing the large volumes of interferometer data poses a practical challenge. We have developed a data processing pipeline called tlpipe to process the drift scan survey data from the Tianlai experiment. It performs oine data processing tasks such as radio frequency interference (RFI) agging, array calibration, binning, and map-making, etc. It also includes utility functions needed for the data analysis, such as data selection, transformation, visualization and others. A number of new algorithms are implemented, for example the eigenvector decomposition method for array calibration and the Tikhnov regularization for m-mode analysis. In this paper we describe the design and implementation of the pipeline and illustrate its functions with some analysis of real data. Finally, we outline directions for future development of this publicly code.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Collaboration 51 Correlation Function Analysis Suite (c51_corr_analysis) v0.1.0

This is a data analysis package designed for analyzing correlation functions generated with lattice QCD calculations. The purpose is to make a centralized software suite for use by all members of my collaboration, so that various members can spend less time developing their own analysis codes, and more time extracting the interesting physics from our calculations. There are also 3 independent collaborations that have expressed interest in using this code, so there is some expectation it will be used in the broader international lattice QCD community.

Walker-Loud, Andre↗

Array-Based Machine Learning for Functional Group Detection in Electron Ionization Mass Spectrometry

Mass spectrometry is a ubiquitous technique capable of complex chemical analysis. The fragmentation patterns that appear in mass spectrometry are an excellent target for artificial intelligence methods to automate and expedite the analysis of data to identify targets such as functional groups. To develop this approach, we trained models on electron ionization (a reproducible hard fragmentation technique) mass spectra so that not only the final model accuracies but also the reasoning behind model assignments could be evaluated. The convolutional neural network (CNN) models were trained on 2D images of the spectra using transfer learning of Inception V3, and the logistic regression models were trained using array-based data and Scikit Learn implementation in Python. Our training dataset consisted of 21,166 mass spectra from the United States’ National Institute of Standards and Technology (NIST) Webbook. The data was used to train models to identify functional groups, both specific (e.g., amines, esters) and generalized classifications (aromatics, oxygen-containing functional groups, and nitrogen-containing functional groups). We found that the highest final accuracies on identifying new data were observed using logistic regression rather than transfer learning on CNN models. It was also determined that the mass range most beneficial for functional group analysis is 0–100 m/z. We also found success in correctly identifying functional groups of example molecules selected from both the NIST database and experimental data. Beyond functional group analysis, we also have developed a methodology to identify impactful fragments for the accurate detection of the models’ targets. The results demonstrate a potential pathway for analyzing and screening substantial amounts of mass spectral data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis↗

Non-universal stellar initial mass functions: large uncertainties in star formation rates at z ≈ 2–4 and other astrophysical probes

ABSTRACT We explore the assumption, widely used in many astrophysical calculations, that the stellar initial mass function (IMF) is universal across all galaxies. By considering both a canonical broken-power-law IMF and a non-universal IMF, we are able to compare the effect of different IMFs on multiple observables and derived quantities in astrophysics. Specifically, we consider a non-universal IMF that varies as a function of the local star formation rate, and explore the effects on the star formation rate density (SFRD), the extragalactic background light, the supernova (both core-collapse and thermonuclear) rates, and the diffuse supernova neutrino background. Our most interesting result is that our adopted varying IMF leads to much greater uncertainty on the SFRD at $z \approx 2-4$ than is usually assumed. Indeed, we find an SFRD (inferred using observed galaxy luminosity distributions) that is a factor of $\gtrsim 3$ lower than canonical results obtained using a universal IMF. Secondly, the non-universal IMF we explore implies a reduction in the supernova core-collapse rate of a factor of $\sim 2$, compared against a universal IMF. The other potential tracers are only slightly affected by changes to the properties of the IMF. We find that currently available data do not provide a clear preference for universal or non-universal IMF. However, improvements to measurements of the star formation rate and core-collapse supernova rate at redshifts $z \gtrsim 2$ may offer the best prospects for discernment.

79 ASTRONOMY AND ASTROPHYSICS↗

On single-crystal total scattering data reduction and correction protocols for analysis in direct space

Data reduction and correction steps and processed data reproducibility in the emerging single-crystal total-scattering-based technique of three-dimensional differential atomic pair distribution function (3D-ΔPDF) analysis are explored. All steps from sample measurement to data processing are outlined using a crystal of CuIr 2 S 4 as an example, studied in a setup equipped with a high-energy X-ray beam and a flat-panel area detector. Computational overhead as pertains to data sampling and the associated data-processing steps is also discussed. Various aspects of the final 3D-ΔPDF reproducibility are explicitly tested by varying the data-processing order and included steps, and by carrying out a crystal-to-crystal data comparison. Situations in which the 3D-ΔPDF is robust are identified, and caution against a few particular cases which can lead to inconsistent 3D-ΔPDFs is noted. Although not all the approaches applied herein will be valid across all systems, and a more in-depth analysis of some of the effects of the data-processing steps may still needed, the methods collected herein represent the start of a more systematic discussion about data processing and corrections in this field.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Real-space visualization of short-range antiferromagnetic correlations in a magnetically enhanced thermoelectric

Short-range magnetic correlations can significantly increase the thermopower of magnetic semiconductors, representing a noteworthy development in the decades-long effort to develop high-performance thermoelectric materials. Here, we reveal the nature of the thermopower-enhancing magnetic correlations in the antiferromagnetic semiconductor MnTe. Using magnetic pair distribution function analysis of neutron scattering data, we obtain a detailed, real-space view of robust, nanometer-scale, antiferromagnetic correlations that persist into the paramagnetic phase above the Neel temperature $T_N$ = 307 K. In this work, the magnetic correlation length in the paramagnetic state is significantly longer along the crystallographic c axis than within the ab plane, pointing to anisotropic magnetic interactions. Ab initio calculations of the spin-spin correlations using density functional theory in the disordered local moment approach reproduce this result with quantitative accuracy. These findings constitute the first real-space picture of short-range spin correlations in a magnetically enhanced thermoelectric and inform future efforts to optimize thermoelectric performance by magnetic means.

36 MATERIALS SCIENCE↗

Low-energy interband transition in the infrared response of the correlated metal SrVO 3 in the ultraclean limit

We studied the low-energy electronic response of the prototypical correlated metal SrVO 3 in the ultraclean and disordered limit using infrared spectroscopy and density functional theory plus dynamical mean field theory calculations (DFT+DMFT). A strong optical excitation at 70 meV is observed in the optical response of the ultraclean samples but is hidden by the low-energy Drude-like response from intraband excitations in the more disordered samples. DFT+DMFT calculations reveal that this optical excitation originates from interband transitions between the bands split by orbital off-diagonal hopping, which has often been ignored in cubic systems, such as SrVO 3 . A memory function analysis of the optical data shows that this interband transition can lead to deviations of optical self-energy from the expected Fermi-liquid behavior. Our findings demonstrate that analysis schemes employed to extract many-body effects from optical spectra may be oversimplified to study the true electronic ground state and that improvements in material quality can guide efforts to refine theoretical approaches.

36 MATERIALS SCIENCE↗

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology ↗

Trimming and Decontamination of Metagenomic Data can Significantly Impact Assembly and Binning Metrics, Phylogenomic and Functional Analysis

Background: Investigators using metagenomic sequencing to study microbiomes often trim and decontaminate reads without knowing their effect on downstream analyses. Objective: This study was designed to evaluate the impacts JGI trimming and decontamination procedures have on assembly and binning metrics, placement of MAGs into species trees, and functional profiles of MAGs extracted from complex rhizosphere metagenomes, as well as how more aggressive trimming impacts these binning metrics. Methods: Twenty-three Miscanthus x giganteus rhizosphere metagenomes were subjected to different combinations and thresholds of force, kmer, and quality trimming and decontamination using BBDuk. Reads were assembled and binned in KBase. Phylogenomic and statistical analyses were applied to evaluate the effects of trimming and decontamination on downstream analyses. Results: We found that JGI trimmed and decontaminated reads had significant impacts on assembly and binning metrics compared to raw reads, including significantly higher total contig counts, more contigs greater than 10k bp in length, and larger total lengths of raw assemblies compared to QC assemblies, and 2.0% lower average contamination of QC MAGs compared to raw MAGs. We also found that differences in the placement of MAGs in species trees increased with decreasing completeness and contamination thresholds. Furthermore, aggressive trimming (Q20) was found to significantly reduce MAG counts. Conclusion: Trimming and decontamination of metagenomics reads prior to assembly can change an investigator’s answer to the questions, “Who is there and what are they doing?” However, mild trimming and decontamination of metagenomic reads with high-quality scores are recommended for removing sample processing and sequencing artifacts.

Whitham, Jason M.↗