Generation and Application of NCF Data Network Layers for Risk Analysis via Functional Decomposition .
Abstract not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Abstract not provided.
SASPDF, a method for characterizing the structure of nanoparticle assemblies (NPAs), is presented. The method is an extension of the atomic pair distribution function (PDF) analysis to the small-angle scattering (SAS) regime. The PDFgetS3 software package for computing the PDF from SAS data is also presented. An application of the SASPDF method to characterize structures of representative NPA samples with different levels of structural order is then demonstrated. The SASPDF method quantitatively yields information such as structure, disorder and crystallite sizes of ordered NPA samples. The method was also used to successfully model the data from a disordered NPA sample. The SASPDF method offers the possibility of more quantitative characterizations of NPA structures for a wide class of samples.
This is a data analysis package designed for analyzing correlation functions generated with lattice QCD calculations. The purpose is to make a centralized software suite for use by all members of my collaboration, so that various members can spend less time developing their own analysis codes, and more time extracting the interesting physics from our calculations. There are also 3 independent collaborations that have expressed interest in using this code, so there is some expectation it will be used in the broader international lattice QCD community.
Mass spectrometry is a ubiquitous technique capable of complex chemical analysis. The fragmentation patterns that appear in mass spectrometry are an excellent target for artificial intelligence methods to automate and expedite the analysis of data to identify targets such as functional groups. To develop this approach, we trained models on electron ionization (a reproducible hard fragmentation technique) mass spectra so that not only the final model accuracies but also the reasoning behind model assignments could be evaluated. The convolutional neural network (CNN) models were trained on 2D images of the spectra using transfer learning of Inception V3, and the logistic regression models were trained using array-based data and Scikit Learn implementation in Python. Our training dataset consisted of 21,166 mass spectra from the United States’ National Institute of Standards and Technology (NIST) Webbook. The data was used to train models to identify functional groups, both specific (e.g., amines, esters) and generalized classifications (aromatics, oxygen-containing functional groups, and nitrogen-containing functional groups). We found that the highest final accuracies on identifying new data were observed using logistic regression rather than transfer learning on CNN models. It was also determined that the mass range most beneficial for functional group analysis is 0–100 m/z. We also found success in correctly identifying functional groups of example molecules selected from both the NIST database and experimental data. Beyond functional group analysis, we also have developed a methodology to identify impactful fragments for the accurate detection of the models’ targets. The results demonstrate a potential pathway for analyzing and screening substantial amounts of mass spectral data.
A cloud web platform for analysis and interpretation of atomic pair distribution function (PDF) data ( PDFitc ) is described. The platform is able to host applications for PDF analysis to help researchers study the local and nanoscale structure of nanostructured materials. The applications are designed to be powerful and easy to use and can, and will, be extended over time through community adoption and development. The currently available PDF analysis applications, structureMining, spacegroupMining and similarityMapping , are described. In the first and second the user uploads a single PDF and the application returns a list of best-fit candidate structures, and the most likely space group of the underlying structure, respectively. In the third, the user can upload a set of measured or calculated PDFs and the application returns a matrix of Pearson correlations, allowing assessment of the similarity between different data sets. structureMining is presented here as an example to show the easy-to-use workflow on PDFitc . In the future, as well as using the PDFitc applications for data analysis, it is hoped that the community will contribute their own codes and software to the platform.
Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).
This paper describes the Jas4pp framework for exploring physics cases and for detector-performance studies of future particle collision experiments. Jas4pp is a multi-platform Java program for numeric calculations, scientific visualization in 2D and 3D, storing data in various file formats and displaying collision events and detector geometries. It also includes complex data-analysis algorithms for function minimization, regression analysis, event reconstruction (such as jet reconstruction), limit settings and other libraries widely used in particle physics. The framework can be used with several scripting languages, such as Python/Jython, Groovy and JShell. Several benchmark tests discussed in the paper illustrate significant improvements in the performance of the Groovy and JShell scripting languages compared to the standard Python implementation in C. Furthermore, the improvements for numeric computations in Java are attributed to recent enhancements in the Java Virtual Machine.
ABSTRACT We explore the assumption, widely used in many astrophysical calculations, that the stellar initial mass function (IMF) is universal across all galaxies. By considering both a canonical broken-power-law IMF and a non-universal IMF, we are able to compare the effect of different IMFs on multiple observables and derived quantities in astrophysics. Specifically, we consider a non-universal IMF that varies as a function of the local star formation rate, and explore the effects on the star formation rate density (SFRD), the extragalactic background light, the supernova (both core-collapse and thermonuclear) rates, and the diffuse supernova neutrino background. Our most interesting result is that our adopted varying IMF leads to much greater uncertainty on the SFRD at $z \approx 2-4$ than is usually assumed. Indeed, we find an SFRD (inferred using observed galaxy luminosity distributions) that is a factor of $\gtrsim 3$ lower than canonical results obtained using a universal IMF. Secondly, the non-universal IMF we explore implies a reduction in the supernova core-collapse rate of a factor of $\sim 2$, compared against a universal IMF. The other potential tracers are only slightly affected by changes to the properties of the IMF. We find that currently available data do not provide a clear preference for universal or non-universal IMF. However, improvements to measurements of the star formation rate and core-collapse supernova rate at redshifts $z \gtrsim 2$ may offer the best prospects for discernment.
Data reduction and correction steps and processed data reproducibility in the emerging single-crystal total-scattering-based technique of three-dimensional differential atomic pair distribution function (3D-ΔPDF) analysis are explored. All steps from sample measurement to data processing are outlined using a crystal of CuIr 2 S 4 as an example, studied in a setup equipped with a high-energy X-ray beam and a flat-panel area detector. Computational overhead as pertains to data sampling and the associated data-processing steps is also discussed. Various aspects of the final 3D-ΔPDF reproducibility are explicitly tested by varying the data-processing order and included steps, and by carrying out a crystal-to-crystal data comparison. Situations in which the 3D-ΔPDF is robust are identified, and caution against a few particular cases which can lead to inconsistent 3D-ΔPDFs is noted. Although not all the approaches applied herein will be valid across all systems, and a more in-depth analysis of some of the effects of the data-processing steps may still needed, the methods collected herein represent the start of a more systematic discussion about data processing and corrections in this field.
We make use of individual (epoch) detection data from the Pan-STARRS “3π” survey for 2863 optical ICRF3 counterparts in the five wavelength bands g, r, i, z, and y, published as part of the Data Release 2. A dedicated method based on the Functional Principal Component Analysis is developed for these sparse and irregularly sampled data. With certain regularization and normalization constraints, it allows us to obtain uniform and compatible estimates of the variability amplitudes and average magnitudes between the passbands and objects. We find that the starting assumption of affinity of the light curves for a given object at different wavelengths is violated for several percent of the sample. The distributions of rms variability amplitudes are strongly skewed toward small values, peaking at ∼0.1 mag with tails stretching to 2 mag. Statistically, the lowest variability is found for the r band and the largest for the reddest y band. A small “brighter-redder” effect is present, with amplitudes in y greater than amplitudes in g in 57% of the sample. The variability versus redshift dependence shows a strong decline with z toward redshift 3, which we interpret as the time dilation of the dominant time frequencies. The colors of radio-loud ICRF3 quasars are correlated with redshift in a complicated, wavy pattern governed by the emergence of brightest emission lines within the five passbands.
Understanding protein structure-function relationships is a key challenge in computational biology, with applications across the biotechnology and pharmaceutical industries. While it is known that protein structure directly impacts protein function, many functional prediction tasks use only protein sequence. In this work, we isolate protein structure to make functional annotations for proteins in the Protein Data Bank in order to study the expressiveness of different structure-based prediction schemes. We present PersGNN - an end-to-end trainable deep learning model that combines graph representation learning with topological data analysis to capture a complex set of both local and global structural features. While variations of these techniques have been successfully applied to proteins before, we demonstrate that our hybridized approach, PersGNN, outperforms either method on its own as well as a baseline neural network that learns from the same information. PersGNN achieves a 9.3% boost in area under the precision recall curve (AUPR) compared to the best individual model, as well as high F1 scores across different gene ontology categories, indicating the transferability of this approach.
Short-range magnetic correlations can significantly increase the thermopower of magnetic semiconductors, representing a noteworthy development in the decades-long effort to develop high-performance thermoelectric materials. Here, we reveal the nature of the thermopower-enhancing magnetic correlations in the antiferromagnetic semiconductor MnTe. Using magnetic pair distribution function analysis of neutron scattering data, we obtain a detailed, real-space view of robust, nanometer-scale, antiferromagnetic correlations that persist into the paramagnetic phase above the Neel temperature $T_N$ = 307 K. In this work, the magnetic correlation length in the paramagnetic state is significantly longer along the crystallographic c axis than within the ab plane, pointing to anisotropic magnetic interactions. Ab initio calculations of the spin-spin correlations using density functional theory in the disordered local moment approach reproduce this result with quantitative accuracy. These findings constitute the first real-space picture of short-range spin correlations in a magnetically enhanced thermoelectric and inform future efforts to optimize thermoelectric performance by magnetic means.
We studied the low-energy electronic response of the prototypical correlated metal SrVO 3 in the ultraclean and disordered limit using infrared spectroscopy and density functional theory plus dynamical mean field theory calculations (DFT+DMFT). A strong optical excitation at 70 meV is observed in the optical response of the ultraclean samples but is hidden by the low-energy Drude-like response from intraband excitations in the more disordered samples. DFT+DMFT calculations reveal that this optical excitation originates from interband transitions between the bands split by orbital off-diagonal hopping, which has often been ignored in cubic systems, such as SrVO 3 . A memory function analysis of the optical data shows that this interband transition can lead to deviations of optical self-energy from the expected Fermi-liquid behavior. Our findings demonstrate that analysis schemes employed to extract many-body effects from optical spectra may be oversimplified to study the true electronic ground state and that improvements in material quality can guide efforts to refine theoretical approaches.
The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.
This this paper, we propose a new family of depth measures called the elastic depths that can be used to greatly improve shape anomaly detection in functional data. Shape anomalies are functions that have considerably different geometric forms or features from the rest of the data. Identifying them is generally more difficult than identifying magnitude anomalies because shape anomalies are often not distinguishable from the bulk of the data with visualization methods. The proposed elastic depths use the recently developed elastic distances to directly measure the centrality of functions in the amplitude and phase spaces. Measuring shape outlyingness in these spaces provides a rigorous quantification of shape which, in turn, gives the elastic depths a strong theoretical and practical advantage over other methods in detecting shape anomalies. A simple boxplot and thresholding method are introduced to identify shape anomalies using the elastic depths. We assess the elastic depth's detection skill on simulated shape outlier scenarios and compare them against popular shape anomaly detectors. Finally, bond yields, image outlines, and hurricane trajectories are used to demonstrate our method's applicability to functional data observed on three different manifolds.
Background: Investigators using metagenomic sequencing to study microbiomes often trim and decontaminate reads without knowing their effect on downstream analyses. Objective: This study was designed to evaluate the impacts JGI trimming and decontamination procedures have on assembly and binning metrics, placement of MAGs into species trees, and functional profiles of MAGs extracted from complex rhizosphere metagenomes, as well as how more aggressive trimming impacts these binning metrics. Methods: Twenty-three Miscanthus x giganteus rhizosphere metagenomes were subjected to different combinations and thresholds of force, kmer, and quality trimming and decontamination using BBDuk. Reads were assembled and binned in KBase. Phylogenomic and statistical analyses were applied to evaluate the effects of trimming and decontamination on downstream analyses. Results: We found that JGI trimmed and decontaminated reads had significant impacts on assembly and binning metrics compared to raw reads, including significantly higher total contig counts, more contigs greater than 10k bp in length, and larger total lengths of raw assemblies compared to QC assemblies, and 2.0% lower average contamination of QC MAGs compared to raw MAGs. We also found that differences in the placement of MAGs in species trees increased with decreasing completeness and contamination thresholds. Furthermore, aggressive trimming (Q20) was found to significantly reduce MAG counts. Conclusion: Trimming and decontamination of metagenomics reads prior to assembly can change an investigator’s answer to the questions, “Who is there and what are they doing?” However, mild trimming and decontamination of metagenomic reads with high-quality scores are recommended for removing sample processing and sequencing artifacts.
The objective of the project is to develop a robotics enabled eddy current testing system (REECTS) in automatic probe deployment, inspection, and data acquisition and analysis. The main functions of the REECTS are to: 1) identify geometry and locations of heat exchange tubes with assistance of an imaging recognition system; 2) precisely control the position and motion speed of ECT probes by an adaptive control system; 3) facilitate data analysis and real-time decision making for autonomous inspection assisted by machine learning algorithms.