Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data analysis methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Apparatus and method for safety analysis evaluation with data-driven workflow

An apparatus and method for system safety analysis evaluation is provided, the apparatus including processing circuitry configured for generating a calculation matrix for a system, generating a plurality of models based on the calculation matrix, performing a benchmarking or convolution analysis of the plurality of models, identifying a design envelope based on the benchmarking or convolution analysis, deriving uncertainty models from the benchmarking or convolution analysis, deriving an assessment judgment based on the uncertainty models and acceptance criteria, defining one or more limiting scenarios based on the design envelope, and determining a safety margin in at least one figure-of-merit for the system based on the design envelope and the acceptance criteria.

Martin, Robert P.↗

Apparatus and method for safety analysis evaluation with data-driven workflow

An apparatus and method for system safety analysis evaluation is provided, the apparatus including processing circuitry configured for generating a calculation matrix for a system, generating a plurality of models based on the calculation matrix, performing a benchmarking or convolution analysis of the plurality of models, identifying a design envelope based on the benchmarking or convolution analysis, deriving uncertainty models from the benchmarking or convolution analysis, deriving an assessment judgment based on the uncertainty models and acceptance criteria, defining one or more limiting scenarios based on the design envelope, and determining a safety margin in at least one figure-of-merit for the system based on the design envelope and the acceptance criteria.

Martin, Robert P.↗

Covariate Dependent Sparse Functional Data Analysis

This study proposes a method to incorporate covariate information into sparse functional data analysis. The method aims at cases where each subject has a limited number of longitudinal measurements and is associated with static covariates. This research is motivated by several use cases in practice. One representative example is void swelling, a nuclear-specific material degradation mechanism. Void swelling is affected by many covariates, including alloy composition and irradiation type. How to accurately model the complicated joint effects of such covariates on the swelling process is the key to mitigating the effect of swelling and ensuring safe operation. Unlike most of the existing methods, the proposed method can handle high-dimensional covariates with the informative covariate identification procedure and sparse and irregularly spaced measurements, that is, does not require complete or dense observations. The main innovation of the proposed method is that we model the variation coming from covariates and the variation left conditioned on covariates, such that the functional principal component analysis and Gaussian process can be conducted in a unified manner. Further, we also propose a systematic approach to identify important covariates in the hypothesis testing context. The methodology is demonstrated on applications in nuclear engineering and healthcare and simulation studies.

42 ENGINEERING↗

DeepMerge – II. Building robust deep learning algorithms for merging galaxy identification across domains

In astronomy, neural networks are often trained on simulation data with the prospect of being used on telescope observations. Unfortunately, training a model on simulation data and then applying it to instrument data leads to a substantial and potentially even detrimental decrease in model accuracy on the new target dataset. Simulated and instrument data represent different data domains, and for an algorithm to work in both, domain-invariant learning is necessary. Here we employ domain adaptation techniques— Maximum Mean Discrepancy (MMD) as an additional transfer loss and Domain Adversarial Neural Networks (DANNs)— and demonstrate their viability to extract domain-invariant features within the astronomical context of classifying merging and non-merging galaxies. Additionally, we explore the use of Fisher loss and entropy minimization to enforce better in-domain class discriminability. We show that the addition of each domain adaptation technique improves the performance of a classifier when compared to conventional deep learning algorithms. We demonstrate this on two examples: between two Illustris-1 simulated datasets of distant merging galaxies, and between Illustris-1 simulated data of nearby merging galaxies and observed data from the Sloan Digital Sky Survey. The use of domain adaptation techniques in our experiments leads to an increase of target domain classification accuracy of up to ~20%. With further development, these techniques will allow astronomers to successfully implement neural network models trained on simulation data to efficiently detect and study astrophysical objects in current and future large-scale astronomical surveys.

galaxies: interactions↗

Merger identification through photometric bands, colours, and their errors

Aims. We present the application of a fully connected neural network (NN) for galaxy merger identification using exclusively photometric information. Our purpose is not only to test the method’s efficiency, but also to understand what merger properties the NN can learn and what their physical interpretation is. Methods. We created a class-balanced training dataset of 5860 galaxies split into mergers and non-mergers. The galaxy observations came from SDSS DR6 and were visually identified in Galaxy Zoo. The 2930 mergers were selected from known SDSS mergers and the respective non-mergers were the closest match in both redshift and r magnitude. The NN architecture was built by testing a different number of layers with different sizes and variations of the dropout rate. We compared input spaces constructed using: the five SDSS filters: u, g, r, i, and z; combinations of bands, colours, and their errors; six magnitude types; and variations of input normalization. Results. We find that the fibre magnitude errors contribute the most to the training accuracy. Studying the parameters from which they are calculated, we show that the input space built from the sky error background in the five SDSS bands alone leads to 92.64 ± 0.15% training accuracy. We also find that the input normalization, that is to say, how the data are presented to the NN, has a significant effect on the training performance. Conclusions. We conclude that, from all the SDSS photometric information, the sky error background is the most sensitive to merging processes. This finding is supported by an analysis of its five-band feature space by means of data visualization. Moreover, studying the plane of the g and r sky error bands shows that a decision boundary line is enough to achieve an accuracy of 91.59%.

79 ASTRONOMY AND ASTROPHYSICS↗

Better together: Elements of successful scientific software development in a distributed collaborative community

Many scientific disciplines rely on computational methods for data analysis, model generation, and prediction. Implementing these methods is often accomplished by researchers with domain expertise but without formal training in software engineering or computer science. This arrangement has led to underappreciation of sustainability and maintainability of scientific software tools developed in academic environments. Some software tools have avoided this fate, including the scientific library Rosetta. We use this software and its community as a case study to show how modern software development can be accomplished successfully, irrespective of subject area. Rosetta is one of the largest software suites for macromolecular modeling, with 3.1 million lines of code and many state-of-the-art applications. Since the mid 1990s, the software has been developed collaboratively by the RosettaCommons, a community of academics from over 60 institutions worldwide with diverse backgrounds including chemistry, biology, physiology, physics, engineering, mathematics, and computer science. Developing this software suite has provided us with more than two decades of experience in how to effectively develop advanced scientific software in a global community with hundreds of contributors. Here we illustrate the functioning of this development community by addressing technical aspects (like version control, testing, and maintenance), community-building strategies, diversity efforts, software dissemination, and user support. We demonstrate how modern computational research can thrive in a distributed collaborative community. The practices described here are independent of subject area and can be readily adopted by other software development communities

97 MATHEMATICS AND COMPUTING↗

BEYONDPLANCK II. CMB mapmaking through Gibbs sampling

We present a Gibbs sampling solution to the mapmaking problem for cosmic microwave background (CMB) measurements that builds on existing destriping methodology. Gibbs sampling breaks the computationally heavy destriping problem into two separate steps: noise filtering and map binning. Considered as two separate steps, both are computationally much cheaper than solving the combined problem. This provides a huge performance benefit as compared to traditional methods and it allows us, for the first time, to bring the destriping baseline length to a single sample. Here, we applied the Gibbs procedure to simulated Planck 30 GHz data. We find that gaps in the time-ordered data are handled efficiently by filling them in with simulated noise as part of the Gibbs process. The Gibbs procedure yields a chain of map samples, from which we are able to compute the posterior mean as a best-estimate map. The variation in the chain provides information on the correlated residual noise, without the need to construct a full noise covariance matrix. However, if only a single maximum-likelihood frequency map estimate is required, we find that traditional conjugate gradient solvers converge much faster than a Gibbs sampler in terms of the total number of iterations. The conceptual advantages of the Gibbs sampling approach lies in statistically well-defined error propagation and systematic error correction. This methodology thus forms the conceptual basis for the mapmaking algorithm employed in the BEYONDPLANCK framework, which implements the first end-to-end Bayesian analysis pipeline for CMB observations.

79 ASTRONOMY AND ASTROPHYSICS↗

Anomaly detection in Hyper Suprime-Cam galaxy images with generative adversarial networks

ABSTRACT The problem of anomaly detection in astronomical surveys is becoming increasingly important as data sets grow in size. We present the results of an unsupervised anomaly detection method using a Wasserstein generative adversarial network (WGAN) on nearly one million optical galaxy images in the Hyper Suprime-Cam (HSC) survey. The WGAN learns to generate realistic HSC-like galaxies that follow the distribution of the data set; anomalous images are defined based on a poor reconstruction by the generator and outlying features learned by the discriminator. We find that the discriminator is more attuned to potentially interesting anomalies compared to the generator, and compared to a simpler autoencoder-based anomaly detection approach, so we use the discriminator-selected images to construct a high-anomaly sample of ∼13 000 objects. We propose a new approach to further characterize these anomalous images: we use a convolutional autoencoder to reduce the dimensionality of the residual differences between the real and WGAN-reconstructed images and perform UMAP clustering on these. We report detected anomalies of interest including galaxy mergers, tidal features, and extreme star-forming galaxies. A follow-up spectroscopic analysis of one of these anomalies is detailed in the Appendix; we find that it is an unusual system most likely to be a metal-poor dwarf galaxy with an extremely blue, higher-metallicity H ii region. We have released a catalogue with the WGAN anomaly scores; the code and catalogue are available at https://github.com/kstoreyf/anomalies-GAN-HSC; and our interactive visualization tool for exploring the clustered data is at https://weirdgalaxi.es.

79 ASTRONOMY AND ASTROPHYSICS↗

Assessing Low-Temperature Geothermal Play Types: Relevant Data and Play Fairway Analysis Methods

The U.S. Department of Energy (DOE) Geothermal Technologies Office (GTO) is supporting the Geothermal Heating and Cooling Geospatial Datasets and Analysis project conducted by the National Renewable Energy Laboratory (NREL) as part of a broader effort to demonstrate the multi-faceted value of integrating geothermal power and geothermal heating and cooling (GHC) technologies into national decarbonization plans and community energy plans. Currently, there is a need to establish baseline low-temperature geothermal resource datasets and evaluate methods for deploying these technologies to provide the basis for supporting private sector investment. This project is focused on collecting baseline datasets, updating conceptual models, and creating Play Fairway Analysis (PFA) workflows for low-temperature (<150 degrees Celsius) geothermal resources of different geothermal play types (i.e., sedimentary basin, orogenic belts, and radiogenic geothermal play types) that could be used for geothermal heating and cooling (GHC), combined heat and power (CHP), and other geothermal direct uses (GDU) applications. Low-temperature geothermal resources are defined as reservoirs - natural or engineered - with temperatures <150 degrees Celsius. While the focus in the NREL effort is on GHC, resources at the upper end of this temperature range can also be used for small-scale power generation. This project does not include Ground Source Heat Pumps (GSHPs) technologies because they can be effectively developed almost anywhere. Low-temperature geothermal resources have not been studied as extensively as higher- to medium-temperature geothermal resources, but there is recent interest in improving understanding of these types of resources with an uptick of interest in geothermal technologies for decarbonizing heating and cooling systems. In addition, Enhanced Geothermal Systems (EGS) and other emerging technologies for exploiting petrothermal resources have opened the possibility of utilizing deep sedimentary basin systems, where porous media provide permeability and high temperatures can be reached at great depths. This project takes the approach of classifying low- temperature geothermal resources by geothermal play type (GPT). We defined and characterized three major classes of low-temperature GPT: sedimentary basins, orogenic systems, and radiogenic systems. We develop methodologies for evaluating and analyzing the potential for these resources building off the PFA approach to de-risking geothermal exploration and characterization. The proposed PFA approach for low-temperature geothermal resources includes: 1) identifying relevant data (e.g., datasets such bottom-hole temperatures from oil and gas wells, heat flow data, Quaternary faults and stress field data, geophysical data, etc.); 2) grouping and weighting of relevant datasets into PFA criteria (e.g., geological, risk, and economic criteria); 3) uncertainty quantification; 4) developing favorability or common risk maps for low-temperature geothermal resources to identify potential locations for more focused data collection; and 5) estimating electric power generation and heating potential at those locations using the GeoRePORT Resource Size Assessment Tool (RSAT). This project should facilitate future deployment of GHC, CHP, and GDU by providing data, tools, and a workflow applicable to low-temperature geothermal resources. Increased deployment of GHC and GDU will help achieve national and local decarbonization goals.

15 GEOTHERMAL ENERGY↗

Advances in geophysical forensic event monitoring

Forensic analysis of man-made, non-nuclear events (such as industrial accidents, explosion experiments and mine collapses) has become more frequent and detailed owing to advancements in geophysical monitoring. Here, in this Technical Review, we demonstrate how geophysical forensic monitoring using seismic, infrasound and hydroacoustic recordings provides insights on events in the solid earth, atmosphere and underwater. Advanced techniques, including machine-learning-based models, have been developed to detect, identify and investigate these events, providing information on location, subevents, sources and explosive yield. The increase in data availability, application of advanced methods and computation and the growth of multitechnology approaches have increased the accuracy of forensic event analysis and enabled more realistic characterization of uncertainties. For example, the 2020 Beirut explosion in Lebanon demonstrated that various seismic, acoustic and other methods could be used to estimate explosive yield (and yield uncertainties) of about 1 ktonne, providing confidence in the application of these methods to smaller events where data are available. However, forensic investigations remain largely limited to known events with identified sources. Increased access to data, sophisticated analysis methods and high-resolution earth models will improve forensic event analysis further, enabling civil and scientific applications, such as localization in the search for the lost ARA San Juan submarine.

geophysics↗

Summary of NDC Capacity Building Workshop and Regional Seismic Travel Time (RSTT) in combination with Data Sharing and Integration Training

The NNSA Seismic Cooperation Program (SCP) sponsored Stephen Myers (LLNL), Michael Begnaud (LANL), Brian Young (SNL) and Istvan Bondar (Research Center for Astronomy and Earth Sciences, Hungary) to serve as a presenters/trainers at the “NDC Capacity Building Workshop and Regional Seismic Travel Time (RSTT) in combination with Data Sharing and Integration Training” September 4-8 2022 in Muscat, Oman (See Appendix A for the agenda). The workshop and training (workshop from here forward) was organized by the Comprehensive Nuclear-Test-Ban Treaty Organization (CTBTO) Provisional Technical Secretariat (PTS). The first half of the week was devoted to NDC workshop activities, and the second half was devoted to RSTT training. Fifty-five participants from 27 countries and the CTBTO-PTS attended the 5-day workshop (See Appendix B for list of participants and countries of origin). Presentations from the PTS described the International Monitoring System (IMS), International Data Centre (IDC) products, and metrics of regional data utilization. Contributed presentations from each country’s scientists included descriptions of regional and national networks, methods of data analysis, and needs for material and technical assistance. Training included an overview of the RSTT method and instruction on how to locate seismic events with the iLoc program, which utilizes RSTT travel times to reduce bias in event location estimates. Methods of seismic tomography and the need for a high-quality tomographic set, including seismological “ground truth”, were emphasized. Seismological “ground truth” or “GT” is a term that has come to mean both events with known location and events with well-characterized locations that are estimated using seismological data, typically with epicenter accuracy of 5 km or better. Notably, the instructional platform has migrated from UNIX shell scripts to Jupyter Notebooks. Jupyter Notebooks have the advantage being more visually intuitive, including display of graphics within the notebook. Each notebook includes every processing step that participants need to reproduce the entire exercise.

58 GEOSCIENCES↗

Seven open problems in applied combinatorics

We present and discuss seven different open problems in applied combinatorics. Additionally, the application areas relevant to this compilation include quantum computing, algorithmic differentiation, topological data analysis, iterative methods, hypergraph cut algorithms, and power systems.

97 MATHEMATICS AND COMPUTING↗

Active anomaly detection for time-domain discoveries

Aims. We present the first evidence that adaptive learning techniques can boost the discovery of unusual objects within astronomical light curve data sets. Methods. Our method follows an active learning strategy where the learning algorithm chooses objects which can potentially improve the learner if additional information about them is provided. This new information is subsequently used to update the machine learning model, allowing its accuracy to evolve with each new information. For the case of anomaly detection, the algorithm aims to maximize the number of scientifically interesting anomalies presented to the expert by slightly modifying the weights of a traditional Isolation Forest (IF) at each iteration. In order to demonstrate the potential of such techniques, we apply the Active Anomaly Discovery (AAD) algorithm to 2 data sets: simulated light curves from the Photometric LSST Astronomical Time-Series Classification Challenge (PLAsTiCC) and real light curves from the Open Supernova Catalog. We compare the AAD results to those of a static IF. For both methods, we performed a detailed analysis for all objects with the ~2% highest anomaly scores. Results. We show that, in the real data scenario , AAD was able to identify ~80% more true anomalies than the IF. This result is the first evidence that AAD algorithms can play a central role in the search for new physics in the era of large scale sky surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Estimating the Diffusion Coefficient of Lithium in Graphite: Extremely Fast Charging and a Comparison of Data Analysis Techniques

Galvanostatic intermittent titration experiments were performed in three-electrode cells to characterize the effect of C/2, 2-C and 4-C charge rates on the observed lithium diffusion coefficient. As part of the data analysis process, we compared the classic Weppner-Huggins analysis of polarization data with a newer (Wang et al.) analysis method for depolarization data. At low values of x in Li x C 6 , both analysis methods showed the same general trend in the apparent lithium diffusion coefficient, 4-C > 2-C > C/2. The two techniques differed in the magnitude of the estimated diffusion coefficient by about a factor of 100. The observed increase in diffusion coefficient does not last over a large compositional range. Since the estimates from the method of Weppner and Huggins may contain artifacts due to the use of particulate electrodes and high charge rates, the method of Wang et al. may produce better values.

25 ENERGY STORAGE↗

Deconvoluting thermomechanical effects in X-ray diffraction data using machine learning

X-ray diffraction is ideal for probing the sub-surface state during complex or rapid thermomechanical loading of crystalline materials. However, challenges arise as the size of diffraction volumes increases due to spatial broadening and because of the inability to deconvolute the effects of different lattice deformation mechanisms. Here, we present a novel approach that uses combinations of physics-based modeling and machine learning to deconvolve thermal and mechanical elastic strains for diffraction data analysis. The method builds on a previous effort to extract thermal strain distribution information from diffraction data. The new approach is applied to extract the evolution of the thermomechanical state during laser melting of an Inconel 625 wall specimen which produces significant residual stress upon cooling. A combination of heat transfer and fluid flow, elasto-plasticity and X-ray diffraction simulations is used to generate training data for machine-learning (Gaussian process regression, GPR) models that map diffracted intensity distributions to underlying thermomechanical strain fields. First-principles density functional theory is used to determine accurate temperature-dependent thermal expansion and elastic stiffness used for elasto-plasticity modeling. The trained GPR models are found to be capable of deconvoluting the effects of thermal and mechanical strains, in addition to providing information about underlying strain distributions, even from complex diffraction patterns with irregularly shaped peaks.

36 MATERIALS SCIENCE↗

Analysis of Needlet Internal Linear Combination performance on B -mode data from sub-orbital experiments

The observation of primordial B modes in cosmic microwave background (CMB) polarisation data represents the main scientific goal of most of the future CMB experiments. This signal is predicted to be much lower than polarised Galactic emission (foregrounds) in any region of the sky, pointing to the need for effective component separation methods. Aims. Among all the techniques, the blind Needlet Internal Linear Combination (NILC) is of great relevance given our current limited knowledge of the B-mode foregrounds. In this work, we explore the possibility of employing NILC for the analysis of B modes reconstructed from partial-sky data, specifically addressing the complications that such an application yields such as E–B leakage, needlet filtering, and beam convolution. We consider two complementary simulated datasets of future experiments: the balloon-borne Short Wavelength Instrument for the Polarisation Explorer (SWIPE) of the Large Scale Polarisation Explorer, which targets the observation of both reionisation and recombination peaks of the primordial CMB B-mode angular power spectrum, and the ground-based Small Aperture Telescope of Simons Observatory, which, instead, is designed to observe only the recombination bump at ℓ ~ 80. We assessed the performance of the following two alternative techniques to correct for the CMB E–B leakage: the recycling technique and the Zhao-Baskaran method. We find that both techniques reduce the E–B leakage residuals at a negligible level given the sensitivity of the considered experiments, except for the recycling method in the SWIPE footprint at ℓ < 20. Thus, we implemented two extensions of the pipeline, the iterative B decomposition and the diffusive inpainting, which enabled us to recover the input CMB B-mode power for ℓ ≥ 5. For the considered experiments, we demonstrate that needlet filtering and beam convolution do not affect the CMB B-mode reconstruction. Finally, with an appropriate masking strategy, we find that NILC foregrounds subtraction allows one to achieve sensitivities on the tensor-to-scalar ratio in agreement with the targets of the considered CMB experiments.

79 ASTRONOMY AND ASTROPHYSICS↗

A review on recent machine learning applications for imaging mass spectrometry studies

Imaging mass spectrometry (IMS) is a powerful analytical technique widely used in biology, chemistry, and materials science fields that continue to expand. IMS provides a qualitative compositional analysis and spatial mapping with high chemical specificity. The spatial mapping information can be 2D or 3D depending on the analysis technique employed. Due to the combination of complex mass spectra coupled with spatial information, large high-dimensional datasets (hyperspectral) are often produced. Therefore, the use of automated computational methods for an exploratory analysis is highly beneficial. The fast-paced development of artificial intelligence (AI) and machine learning (ML) tools has received significant attention in recent years. These tools, in principle, can enable the unification of data collection and analysis into a single pipeline to make sampling and analysis decisions on the go. There are various ML approaches that have been applied to IMS data over the last decade. Here, in this review, we discuss recent examples of the common unsupervised (principal component analysis, non-negative matrix factorization, k-means clustering, uniform manifold approximation and projection), supervised (random forest, logistic regression, XGboost, support vector machine), and other methods applied to various IMS datasets in the past five years. The information from this review will be useful for specialists from both IMS and ML fields since it summarizes current and representative studies of computational ML-based exploratory methods for IMS.

47 OTHER INSTRUMENTATION↗