Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data analysis methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Domain Adaptation for Measurements of Strong Gravitational Lenses

Upcoming surveys are predicted to discover galaxy-scale strong lenses on the order of $10^5$, making deep learning methods necessary in lensing data analysis. Currently, there is insufficient real lensing data to train deep learning algorithms, but the alternative of training only on simulated data results in poor performance on real data. Domain Adaptation may be able to bridge the gap between simulated and real datasets. We utilize domain adaptation for the estimation of Einstein radius ($\Theta_E$) in simulated galaxy-scale gravitational lensing images with different levels of observational realism. We evaluate two domain adaptation techniques - Domain Adversarial Neural Networks (DANN) and Maximum Mean Discrepancy (MMD). We train on a source domain of simulated lenses and apply it to a target domain of lenses simulated to emulate noise conditions in the Dark Energy Survey (DES). We show that both domain adaptation techniques can significantly improve the model performance on the more complex target domain dataset. This work is the first application of domain adaptation for a regression task in strong lensing imaging analysis. Our results show the potential of using domain adaptation to perform analysis of future survey data with a deep neural network trained on simulated data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Domain Adaptation for Measurements of Strong Gravitational Lenses

Upcoming surveys are predicted to discover galaxy-scale strong lenses on the magnitude of 105, making deep learning methods necessary in lensing data analysis. Currently, there is insufficient real lensing data to train deep learning algorithms, but training only on simulated data results in poor performance on real data. Domain adaptation can bridge the gap between simulated and real datasets. We adopt domain adaptation on the estimation of Einstein radius in simulated galaxy-scale gravitational lensing images. We evaluate two domain adaptation techniques - domain adversarial neural networks (DANN) and maximum mean discrepancy (MMD). We train on a source domain of simulated lenses and apply it to a target domain with emulation of DES survey conditions. We show that both domain adaptation techniques can significantly improve the model performance on the more complex target domain datasets. Our results show the potential of using domain adaptation to perform analysis on future survey data with a deep neural network trained on simulated data.

79 ASTRONOMY AND ASTROPHYSICS↗

A High-Quality Genome-Scale Model for Rhodococcus opacus Metabolism

Rhodococcus opacus is a bacterium that has a high tolerance to aromatic compounds and can produce significant amounts of triacylglycerol (TAG). Here, we present iGR1773, the first genome-scale model (GSM) of R. opacus PD630 metabolism based on its genomic sequence and associated data. The model includes 1773 genes, 3025 reactions, and 1956 metabolites, was developed in a reproducible manner using CarveMe, and was evaluated through Metabolic Model tests (MEMOTE). We combine the model with two Constraint-Based Reconstruction and Analysis (COBRA) methods that use transcriptomics data to predict growth rates and fluxes: E-Flux2 and SPOT (Simplified Pearson Correlation with Transcriptomic data). Growth rates are best predicted by E-Flux2. Flux profiles are more accurately predicted by E-Flux2 than flux balance analysis (FBA) and parsimonious FBA (pFBA), when compared to 44 central carbon fluxes measured by 13C-Metabolic Flux Analysis (13C-MFA). Under glucose-fed conditions, E-Flux2 presents an R2 value of 0.54, while predictions based on pFBA had an inferior R2 of 0.28. We attribute this improved performance to the extra activity information provided by the transcriptomics data. For phenol-fed metabolism, in which the substrate first enters the TCA cycle, E-Flux2’s flux predictions display a high R2 of 0.96 while pFBA showed an R2 of 0.93. We also show that glucose metabolism and phenol metabolism function with similar relative ATP maintenance costs. These findings demonstrate that iGR1773 can help the metabolic engineering community predict aromatic substrate utilization patterns and perform computational strain design.

Roell, Garrett W.↗

Dynamics retrieval from stochastically weighted incomplete data by low-pass spectral analysis

Time-resolved serial femtosecond crystallography (TR-SFX) provides access to protein dynamics on sub-picosecond timescales, and with atomic resolution. Due to the nature of the experiment, these datasets are often highly incomplete and the measured diffracted intensities are affected by partiality. To tackle these issues, one established procedure is that of splitting the data into time bins, and averaging the multiple measurements of equivalent reflections within each bin. This binning and averaging often involve a loss of information. Here, we propose an alternative approach, which we call low-pass spectral analysis (LPSA). In this method, the data are projected onto the subspace defined by a set of trigonometric functions, with frequencies up to a certain cutoff. This approach attenuates undesirable high-frequency features and facilitates retrieving the underlying dynamics. A time-lagged embedding step can be included prior to subspace projection to improve the stability of the results with respect to the parameters involved. Subsequent modal decomposition allows to produce a low-rank description of the system's evolution. Using a synthetic time-evolving model with incomplete and partial observations, we analyze the LPSA results in terms of quality of the retrieved signal, as a function of the parameters involved. We compare the performance of LPSA to that of a range of other sophisticated data analysis techniques. We show that LPSA allows to achieve excellent dynamics reconstruction at modest computational cost. Finally, we demonstrate the superiority of dynamics retrieval by LPSA compared to time binning and merging, which is, to date, the most commonly used method to extract dynamical information from TR-SFX data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Determining biomolecular structures near room temperature using X-ray crystallography: concepts, methods and future optimization

For roughly two decades, cryocrystallography has been the overwhelmingly dominant method for determining high-resolution biomolecular structures. Competition from single-particle cryo-electron microscopy and micro-electron diffraction, increased interest in functionally relevant information that may be missing or corrupted in structures determined at cryogenic temperature, and interest in time-resolved studies of the biomolecular response to chemical and optical stimuli have driven renewed interest in data collection at room temperature and, more generally, at temperatures from the protein–solvent glass transition near 200 K to ∼350 K. Fischer has recently reviewed practical methods for room-temperature data collection and analysis [Fischer (2021), Q. Rev. Biophys. 54 , e1]. Here, the key advantages and physical principles of, and methods for, crystallographic data collection at noncryogenic temperatures and some factors relevant to interpreting the resulting data are discussed. For room-temperature data collection to realize its potential within the structural biology toolkit, streamlined and standardized methods for delivering crystals prepared in the home laboratory to the synchrotron and for automated handling and data collection, similar to those for cryocrystallography, should be implemented.

59 BASIC BIOLOGICAL SCIENCES↗

RESULTS OF A VIRTUAL ROUND ROBIN STUDY TO ESTIMATE PROBABILITY OF DETECTION FOR DISSIMILAR METAL WELDS

This paper presents efforts to overcome challenges with empirical probability of detection (POD) estimations in the nuclear power industry through the utilization of a novel virtual flaw method. A virtual round robin (VRR) study was conducted under the Program for Investigation Of NDE by International Collaboration (PIONIC), organized by the United States Nuclear Regulatory Commission (NRC) utilizing data generated by the virtual flaw method. Analysis of results from the VRR was performed by teams from Pacific Northwest National Laboratory (PNNL), Electric Power Research Institute (EPRI), and Aalto University. Empirically derived POD estimations are presented, and challenges associated with obtaining these estimations are discussed. The virtual flaw method is introduced and some details of its implementation for the VRR activity are described. Results from POD analysis of the VRR data by PNNL, EPRI, and Aalto University are presented and a discussion regarding differences in analysis results is provided. Finally, potential future efforts to improve the application of the virtual flaw method and its estimation of POD are discussed.

Probability of detection, Dissimilar Metal Weld, P↗

Structural basis of differential gene expression at eQTLs loci from high-resolution ensemble models of 3D single-cell chromatin conformations

Abstract Motivation Techniques such as high-throughput chromosome conformation capture (Hi-C) have provided a wealth of information on nucleus organization and genome important for understanding gene expression regulation. Genome-Wide Association Studies have identified numerous loci associated with complex traits. Expression quantitative trait loci (eQTL) studies have further linked the genetic variants to alteration in expression levels of associated target genes across individuals. However, the functional roles of many eQTLs in noncoding regions remain unclear. Current joint analyses of Hi-C and eQTLs data lack advanced computational tools, limiting what can be learned from these data. Results We developed a computational method for simultaneous analysis of Hi-C and eQTL data, capable of identifying a small set of nonrandom interactions from all Hi-C interactions. Using these nonrandom interactions, we reconstructed large ensembles (×105) of high-resolution single-cell 3D chromatin conformations with thorough sampling, accurately replicating Hi-C measurements. Our results revealed many-body interactions in chromatin conformation at the single-cell level within eQTL loci, providing a detailed view of how 3D chromatin structures form the physical foundation for gene regulation, including how genetic variants of eQTLs affect the expression of associated eGenes. Furthermore, our method can deconvolve chromatin heterogeneity and investigate the spatial associations of eQTLs and eGenes at subpopulation level, revealing their regulatory impacts on gene expression. Together, ensemble modeling of thoroughly sampled single-cell chromatin conformations combined with eQTL data, helps decipher how 3D chromatin structures provide the physical basis for gene regulation, expression control, and aid in understanding the overall structure-function relationships of genome organization. Availability and implementation It is available at https://github.com/uic-liang-lab/3DChromFolding-eQTL-Loci.

Du, Lin (ORCID:0009000289869812)↗

SAS PDF: pair distribution function analysis of nanoparticle assemblies from small-angle scattering data

SASPDF, a method for characterizing the structure of nanoparticle assemblies (NPAs), is presented. The method is an extension of the atomic pair distribution function (PDF) analysis to the small-angle scattering (SAS) regime. The PDFgetS3 software package for computing the PDF from SAS data is also presented. An application of the SASPDF method to characterize structures of representative NPA samples with different levels of structural order is then demonstrated. The SASPDF method quantitatively yields information such as structure, disorder and crystallite sizes of ordered NPA samples. The method was also used to successfully model the data from a disordered NPA sample. The SASPDF method offers the possibility of more quantitative characterizations of NPA structures for a wide class of samples.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Fission Product Yields: SCALE/ORNL Perspective [Slides]

This presentation covers a few key points. First and foremost, that ORIGEN is a consumer of FPY data. Secondly, TRITON and POLARIS are assets which can be used to validate new FPY data, e.g., from existing radiochemical benchmarks and new irradiation experiments. Lastly, this presentation finds that new capabilities with covariance data should be developed, and V and V methods and analysis are also needed for the covariance data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Benchmarking blockchain-based gene-drug interaction data sharing methods: A case study from the iDASH 2019 secure genome analysis competition blockchain track

Blockchain distributed ledger technology is just starting to be adopted in genomics and healthcare applications. Despite its increased prevalence in biomedical research applications, skepticism regarding the practicality of blockchain technology for real-world problems is still strong and there are few implementations beyond proof-of-concept. We focus on benchmarking blockchain strategies applied to distributed methods for sharing records of gene-drug interactions. We expect this type of sharing will expedite personalized medicine. We generated gene-drug interaction test datasets using the Clinical Pharmacogenetics Implementation Consortium (CPIC) resource. We developed three blockchain-based methods to share patient records on gene-drug interactions: Query Index, Index Everything, and Dual-Scenario Indexing. We achieved a runtime of about 60 s for importing 4,000 gene-drug interaction records from four sites, and about 0.5 s for a data retrieval query. Our results demonstrated that it is feasible to leverage blockchain as a new platform to share data among institutions.

60 APPLIED LIFE SCIENCES↗

Harmonizing tau positron emission tomography in Alzheimer's disease: The CenTauR scale and the joint propagation model

Abstract INTRODUCTION Tau‐positron emission tomography (PET) outcome data of patients with Alzheimer's disease (AD) cannot currently be meaningfully compared or combined when different tracers are used due to differences in tracer properties, instrumentation, and methods of analysis. METHODS Using head‐to‐head data from five cohorts with tau PET radiotracers designed to target tau deposition in AD, we tested a joint propagation model (JPM) to harmonize quantification (units termed “CenTauR” [CTR]). JPM is a statistical model that simultaneously models the relationships between head‐to‐head and anchor point data. JPM was compared to a linear regression approach analogous to the one used in the amyloid PET Centiloid scale. RESULTS A strong linear relationship was observed between CTR values across brain regions. Using the JPM approach, CTR estimates were similar to, but more accurate than, those derived using the linear regression approach. DISCUSSION Preliminary findings using the JPM support the development and adoption of a universal scale for tau‐PET quantification. Highlights Tested a novel joint propagation model (JPM) to harmonize quantification of tau PET. Units of common scale are termed “CenTauRs”. Tested a Centiloid‐like linear regression approach. Using five cohorts with head‐to‐head tau PET, JPM outperformed linearregressionbased approach. Strong linear relationship was observed between CenTauRs values across brain regions.

Neurosciences & Neurology↗

An iterative method to deblend AGN-Host contributions for Integral Field spectroscopic observations

ABSTRACT We present a new iterative deblending method to separate the host galaxy (HG) and their Active Galactic Nuclei (AGNs) emission with the use of Integral Field spectroscopic (IFS) data. The method decomposes the resolved HG emission from the unresolved AGN emission by modelling the two-dimensional surface brightness (SB) profile of the point-spread function (PSF) and the two-dimensional SB HG continuum simultaneously per each monochromatic slide. Our method does not require any prior information about the observed SB profile or a detailed fitting of the PSF, making it ideal for the automatic analysis of large galaxy samples. In this work, we test the quality of our method, its advantages, and its disadvantages. We test our method by using a set of IFS mock data cubes to quantify the reliability of our deblending process and further compare our method with the qdblend3d analysis tool. Furthermore, we applied our method to three data cubes selected from the MaNGA survey according to the dominance of either its HG or its AGN. We show that our deblending method is capable of disengaging the bright, non-resolved AGN emission from the HG continuum and its narrow emission lines. However, the decoupling depends on how well the IFS spatially resolves the PSF, and on the relative flux intensity of the HG-AGN. Therefore, the method is ideal for disentangling the bright-flux contribution from AGN-dominated spectra.

Ibarra-Medel, H. (ORCID:0000000297906313)↗

DeepCare: Improving Patient Care using Deep Learning on Electronic Health Records

Coordinating patient care using electronic health records (EHR) data presents an exciting but formidable opportunity in data extraction, analysis and modeling. Traditional methods use a manual feature driven approach to model patients with age, family history and symptoms to predict disease outcomes. We propose a novel approach to model patients based on their streaming electronic health records data combined with information from medical knowledge bases, which has been gained over years of medical research. Using a combination of representation learning and long short term memory (LSTM) networks we plan to model patient evolution over time, leading to more accurate and individualized predictive models for patient’s diseases. Our approach will be transformative in providing critical decision support for patient care, enabling accurate understanding and evolution of diseases in patients.

60 APPLIED LIFE SCIENCES↗

Analysis of human performance differences between students and operators when using the Rancor Microworld simulator

Here, from within the umbrella of the Simplified Human Error Experimental Program (SHEEP) framework, this paper analyzes human performance differences between professional and student operators when using a simplified simulator (i.e., Rancor Microworld). This paper represents a crucial step in understanding the fidelity of the simplified simulators and student operators within the SHEEP study. This paper explores a randomized factorial experimental design that features two independent variables: participant type and event class. Six human performance measurements are considered in the experiment. The experiment is conducted using 20 professional reactor operators employed at actual nuclear power plants (NPPs), along with 20 trained students. The experimental data are analyzed via statistical analysis methods. Finally, this paper examines the differences in human performance between actual operators and students when using Rancor Microworld.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Considerations for AMI-Based Operations for Distribution Feeders

More than $5 billion in investments in advanced metering infrastructure (AMI) technologies, AMI deployments, as pervasive secondary network voltage monitoring systems, provide opportunities for utility operations and controls. This paper focuses on the considerations for AMI-based tools and techniques as the industry moves toward operationalizing such large data sets. Phase identification is a first such tool. Numerous distribution network analysis, monitoring, and control applications - including volt/volt-ampere reactive control, state estimation, and distribution automation - require accurate phase connectivity information in the system models. The phase connectivity database maintained by utilities is inaccurate because of a significant amount of missing data, restoration activities, and network reconfiguration. Existing phase identification techniques that estimate phase connectivity work well in distribution feeders that have low or no photovoltaic (PV) generation; however, they fail to identify the phases accurately when considerable PV generation is present. This work addresses the phase identification problem in the presence of high PV generation using statistical analysis methods. Further, insights into the AMI data requirements for this application in terms of data window length and resolution are provided using sensitivity analysis performed on an actual distribution feeder model of San Diego Gas & Electric Company. The results of this study show that the phase connectivity, even in the presence of high PV generation, can be accurately identified using statistical analysis of AMI data of 1 day.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Common risk segment mapping: Streamlining exploration for carbon storage sites, with application to coastal Texas and Louisiana

Large-scale deployment of Carbon Capture and Storage (CCS) will require a commensurately large number of sites. Efficient screening methods are needed to create investment assurance and focus efforts on the most promising sites. The problem is similar to petroleum exploration, for which there are well-developed (though seldom published) workflows, including Common Risk Segment (CRS) mapping. In brief, the process requires 1) defining the key play elements; 2) identifying candidate geologic intervals for each; 3) creating fact-based maps for those intervals; 4) determining minimum criteria for the success of each element; 5) reinterpreting the fact-based maps in terms of chance of success; and 6) combining the individual maps to form a composite, basin-scale view of prospectivity. We adapt the CRS process to screening for CO 2 storage sites. Critically, we redefine the process in terms of cost of characterization and development, rather than chance of success. For illustration, we apply the process to the example of the Lower Miocene on the Texas and Louisiana Gulf Coast. We show that the predictions are consistent with historic hydrocarbon production volumes and rates. The power of the CRS method is that it creates a systematic approach to geologic evaluation and translates complex, multidimensional analysis into clear, graphical and easily comprehended business inputs. The results highlight sweet spots and identifies critical risks, suggesting a focus for further data collection and analysis. Furthermore, the method developed here can be applied to both surface and subsurface factors anywhere that there is interest in geologic storage of CO 2 .

54 ENVIRONMENTAL SCIENCES↗

Using the Monte Carlo Method to Evaluate the Reliability of Screening Multifamily Housing for Radon

When screening for radon in a multifamily housing complex using a fixed sample density (e.g., testing 1 in 10 (10%) or 1 in 4 units (25%)), the statistical confidence is dependent upon the assumed elevated radon frequency. For example, testing 25% of units in a complex estimated to have eight units with elevated radon levels will provide 90% confidence. However, if it is assumed that, in the same complex, there are only three units with elevated radon levels, the confidence drops to around 58% for the same sample density. Furthermore, in the previous example, if elevated radon levels are not found during the screening, all that can be stated is that the screening provides 90% confidence that there are no more than seven units in the complex with elevated radon levels. To more fully illustrate this uncertainty, ten separate multifamily housing radon data sets with 1 to 10 units with radon levels ≥4 pCi/L (Table 1) were selected for analysis using the Monte Carlo statistical method. Unlike other mathematically based statistical approaches, the Monte Carlo statistical method relies on repeated analysis of randomly selected data. from a 100% sampled complex at various sample densities. Success for each simulation is defined as finding at least one unit with elevated radon levels. In this statistical method, confidence in the overarching conclusion can be greatly enhanced by repeating the simulated screening hundreds or even thousands of times. For this study, each of the ten data sets was simulated 1,000 times. For each simulation, the data set was randomized three times before the fixed percentage of data was selected.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Unsupervised Machine Learning for Exploratory Data Analysis of Exoplanet Transmission Spectra

Abstract Transit spectroscopy is a powerful tool for decoding the chemical compositions of the atmospheres of extrasolar planets. In this paper, we focus on unsupervised techniques for analyzing spectral data from transiting exoplanets. After cleaning and validating the data, we demonstrate methods for: (i) initial exploratory data analysis, based on summary statistics (estimates of location and variability); (ii) exploring and quantifying the existing correlations in the data; (iii) preprocessing and linearly transforming the data to its principal components; (iv) dimensionality reduction and manifold learning; (v) clustering and anomaly detection; and (vi) visualization and interpretation of the data. To illustrate the proposed unsupervised methodology, we use a well-known public benchmark data set of synthetic transit spectra. We show that there is a high degree of correlation in the spectral data, which calls for appropriate low-dimensional representations. We explore a number of different techniques for such dimensionality reduction and identify several suitable options in terms of summary statistics, principal components, etc. We uncover interesting structures in the principal component basis, namely well-defined branches corresponding to different chemical regimes of the underlying atmospheres. We demonstrate that those branches can be successfully recovered with a K-means clustering algorithm in a fully unsupervised fashion. We advocate for lower-dimensional representations of the spectroscopic data in terms of the main principal components, in order to reveal the existing structure in the data and quickly characterize the chemical class of a planet.

Matchev, Konstantin T. (ORCID:0000000341829096)↗