Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data analysis methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Unsupervised Machine Learning for Exploratory Data Analysis of Exoplanet Transmission Spectra

Abstract Transit spectroscopy is a powerful tool for decoding the chemical compositions of the atmospheres of extrasolar planets. In this paper, we focus on unsupervised techniques for analyzing spectral data from transiting exoplanets. After cleaning and validating the data, we demonstrate methods for: (i) initial exploratory data analysis, based on summary statistics (estimates of location and variability); (ii) exploring and quantifying the existing correlations in the data; (iii) preprocessing and linearly transforming the data to its principal components; (iv) dimensionality reduction and manifold learning; (v) clustering and anomaly detection; and (vi) visualization and interpretation of the data. To illustrate the proposed unsupervised methodology, we use a well-known public benchmark data set of synthetic transit spectra. We show that there is a high degree of correlation in the spectral data, which calls for appropriate low-dimensional representations. We explore a number of different techniques for such dimensionality reduction and identify several suitable options in terms of summary statistics, principal components, etc. We uncover interesting structures in the principal component basis, namely well-defined branches corresponding to different chemical regimes of the underlying atmospheres. We demonstrate that those branches can be successfully recovered with a K-means clustering algorithm in a fully unsupervised fashion. We advocate for lower-dimensional representations of the spectroscopic data in terms of the main principal components, in order to reveal the existing structure in the data and quickly characterize the chemical class of a planet.

Matchev, Konstantin T. (ORCID:0000000341829096)↗

A Combined Computer Vision and Deep Learning Approach for Rapid Drone-Based Optical Characterization of Parabolic Troughs

Optical accuracy is a primary driver of parabolic trough concentrating solar power (CSP) plant performance, but can be damaged by wind loads, gravity, error during installation, and regular plant operation. Collecting and analyzing optical measurements over an entire operating parabolic trough plant is difficult, given the large scale of typical installations. Distant Observer, a software tool developed at the National Renewable Energy Laboratory, uses images of the absorber tube reflected in the collector mirror to measure both surface slope in the parabolic mirror and offset of the absorber tube from the ideal focal point. This technology has been adapted for fast data collection using low-cost commercial drones, but until recently still required substantial human labor to process large amounts of data. A new method leveraging advanced deep learning and computer vision tools can drastically reduce the time required to process images. This new method addresses the primary analysis bottleneck, identifying featureless, reflective mirror corner points to a high degree of accuracy. Recent work has shown promising results using computer vision methods. The combined deep learning and computer vision approach presented here proved highly effective and has the potential to further automate data collection and analysis, making the tool more robust. The method presented in this paper automatically identified 74.3% of mirror corners within 2 pixels of their manually marked counterparts and 91.9% within 3 pixels. This level of accuracy is sufficient for practical Distant Observer analysis within a target uncertainty. A commercial drone collected video of over 100 parabolic trough modules at an operating CSP plant to demonstrate the deep learning and computer vision method's usefulness in processing large amounts of data. These troughs were successfully analyzed using Distant Observer, paired with the new deep learning and computer vision algorithm, and can provide plant operators and trough designers with valuable insight about plant performance, operating strategies, and plant-wide optical error trends.

computer vision↗

Accelerating Multivariate Functional Approximation Computation with Domain Decomposition Techniques⋆

Modeling large datasets through Multivariate Functional Approximations (MFA) provide an elegant way to handle many visualization and scientific analysis workflows. The process necessitates scalable data partitioning methods to compute MFA representations efficiently without compromising the accuracy or continuity of the reconstructed solution. We propose a domain -decomposed method for computing the MFA with B -spline bases, which reduces the total work per task and uses a restricted Additive Schwarz (RAS) method to converge the control point data degrees -of -freedom along subdomain boundaries. We provide an in-depth analysis of the parallel approach with domain decomposition solvers, aiming to minimize local subdomain error residuals and recover high -order continuity at subdomain interfaces with appropriate choices of knot overlaps. The communication cost, determined by the overlap regions in the RAS implementation, is optimized to recover the numerical error profile of the single subdomain case. Our proposed method stands in contrast to previous methods, which typically only recover either C 0 or at best C 1 continuity for arbitrary B -spline degree expansions, or those that require post -processing to blend discontinuities in the reconstructed data. We demonstrate the effectiveness of our approach using analytical and real -world datasets in 1D, 2D, and 3D through both strong and weak scaling studies. The performance results indicate that the overall cost of computing the approximation is directly proportional to the underlying nearest -neighbor communication implementation, and is only weakly dependent on the overlap region size that determines the size of the messages. This finding underscores the efficiency and scalability of our proposed method, making it a promising solution for handling large datasets in scientific workflows.

additive Schwarz solvers↗

Mitigate: An Adaptive Network Data Anonymization Tool Using Condensation-Based Differential Privacy

Modern network devices collect a large amount of data that can be analyzed to identify bottlenecks, anomalies, cyber-attacks, etc. Therefore, there is often a need to analyze such collections of network data quite often by an external expert or by the research community. However, these collections of data contain sensitive, proprietary information. In order for the network data to be shared, it must first be anonymized. The overall objective of this project is to develop an innovative privacy management tool to anonymize network data and achieve sufficient privacy, acceptable data utility, and efficient data analysis at the same time. No existing anonymization methods can achieve all of these at the same time. The core of this technology is a differential private clustering algorithm that provides strong privacy protection, preserves data properties important for subsequent analysis, and allows the party receiving the anonymized data to conduct analysis directly on anonymized data without the need of decryption or any extra processing. The research carried out was to design, implement and verify a solution to this problem by completing the following tasks: 1) developing the core technology; 2) developing a context based method that automatically recommends fields that must be anonymized; 3) conducted experiments showing superior results using our approach compared to existing tools, and 4) developed an intuitive but basic user interface. The research that was conducted generated novel algorithmic techniques that utilize state-of-the-art methods such as condensation, differential privacy preservation, clustering, automated tuning based on contextual awareness, and recommendation techniques to specify columns to users for anonymization leading to optimal privacy that allows research analysis on the dataset. Experiments were conducted to evaluate the efficacy of these novel algorithmic techniques by performing analysis on original non-anonymized datasets, then conducting analysis on the same yet anonymized datasets and comparing the results of the analyses. Overall, the anonymized analysis results were within 1% of the original results, verifying that the generated technology not only guarantees a high level of privacy but also enables research analysis as if it were conducted on the original dataset. Potential applications of this technology include anonymization of any type of structured network datasets that contain sensitive identifiers, such as IP addresses, that can be used in multiple applications. For example, to create an AI or machine learning model for cyber security, e.g., to detect attacks, or for performance analysis, e.g., identify bottlenecks or predict performance. In addition, a market analysis that was conducted for potential applications of this technology identified a broader range of applications of our anonymization technology beyond the network sector that includes healthcare, banking, insurance, securities, finance (FISB), data brokering, cloud services, ad sales, and government.

97 MATHEMATICS AND COMPUTING↗

Commercial, industrial, and institutional discount rate estimation for efficiency standards analysis: Sector-level data 1998–2021

Underlying each of the Department of Energy’s (DOE’s) federal appliance and equipment energy conservation standards are a set of complex analyses of the projected costs and benefits of regulation. Any new or amended standard must be designed to achieve significant additional energy conservation, provided that it is technologically feasible and economically justified (42 U.S.C. 6295(o)(2)(A)). DOE determines economic justification based on whether the benefits exceed the burdens, considering a variety of factors, including the economic impact of the standard on consumers of the product and the savings in lifetime operating cost compared to any increase in price or maintenance expenses (42 U.S.C. 6295(o)(2)(B)). As part of this determination, DOE conducts a Life-Cycle Cost (LCC) analysis, which models the combined impact of appliance first cost and operating cost changes on a representative commercial building sample in order to identify the fraction of customers achieving LCC savings or incurring net cost at the considered efficiency levels. Thus, the commercial discount rate value(s) used to calculate the present value of energy cost savings within the LCC model implicitly plays a role in estimating the economic impact of potential standard levels. This report provides an in-depth discussion of the commercial discount rate estimation process. It is an update to previous reports on estimating commercial discount rates from firm-level financial data (Fujita, 2016). Major topics covered in this report include: Discount rate estimation methods and rationale; -Data sources used and data limitations; -Discount rate distributions for use in standards analysis; -Discount rate estimation methods and distributions specific to the small business subgroup analysis. Going forward, this report will be updated as data allow and analyses necessitate.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Utilization of an Advanced Sensor network to determine fuel heating value and Real-Time net unit heat rate during transient operation

Coal-fired utility boilers are being increasingly used as variable electricity generation to resolve the imbalance in the energy market from the expansion of intermittent renewable energy. The frequent transient operation required to meet residual energy demand has created a challenge for coal-fired units to operate efficiently. This work utilizes an advanced sensor network (ASN) to calculate net unit heat rate (NUHR) of a coal-fired boiler in real time through combustion calculations and statistical correlations to provide the tools for optimizing dynamic operation. Real-time heating values that were necessary to determine fuel input energy to calculate accurate NUHR were found using both fundamental and data-driven methods. Real-time NUHR shows distinct shifts that reflect changes in process conditions that will improve the ability to optimize transient operation. Data-driven heating value correlations had 24% lower root mean square error (RMSE) than the fundamental combustion calculation approach when compared to daily retrospective proximate analysis. Furthermore, the data-driven method RMSE improved by 7% with the inclusion of ASN data. Future work is to validate by comparing unit performance with and without the inclusion of NUHR as a control parameter for the dynamic neural network.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The CosmoVerse White Paper: Addressing observational tensions in cosmology with systematics and fundamental physics.

The standard model of cosmology has provided a good phenomenological description of a wide range of observations both at astrophysical and cosmological scales for several decades. This concordance model is constructed by a universal cosmological constant and supported by a matter sector described by the standard model of particle physics and a cold dark matter contribution, as well as very early-time inflationary physics, and underpinned by gravitation through general relativity. There have always been open questions about the soundness of the foundations of the standard model. However, recent years have shown that there may also be questions from the observational sector with the emergence of differences between certain cosmological probes. In this White Paper, we identify the key objectives that need to be addressed over the coming decade together with the core science projects that aim to meet these challenges. These discordances primarily rest on the divergence in the measurement of core cosmological parameters with varying levels of statistical confidence. These possible statistical tensions may be partially accounted for by systematics in various measurements or cosmological probes but there is also a growing indication of potential new physics beyond the standard model. After reviewing the principal probes used in the measurement of cosmological parameters, as well as potential systematics, we discuss the most promising array of potential new physics that may be observable in upcoming surveys. We also discuss the growing set of novel data analysis approaches that go beyond traditional methods to test physical models. These new methods will become increasingly important in the coming years as the volume of survey data continues to increase, and as the degeneracy between predictions of different physical models grows. There are several perspectives on the divergences between the values of cosmological parameters, such as the model-independent probes in the late Universe and model-dependent measurements in the early Universe, which we cover at length. The White Paper closes with a number of recommendations for the community to focus on for the upcoming decade of observational cosmology, statistical data analysis, and fundamental physics developments.

Dienes, Keith [Univ. of Arizona, Tucson, AZ (Unite↗

Iterative Bragg peak removal on X-ray absorption spectra with automatic intensity correction

This study introduces a novel iterative Bragg peak removal with automatic intensity correction (IBR-AIC) methodology for X-ray absorption spectroscopy (XAS), specifically addressing the challenge of Bragg peak interference in the analysis of crystalline materials. The approach integrates experimental adjustments and sophisticated post-processing, including an iterative algorithm for robust calculation of the scaling factor of the absorption coefficients and efficient elimination of the Bragg peaks, a common obstacle in accurately interpreting XAS data, particularly in crystalline samples. The method was thoroughly evaluated on dilute catalysts and thin films, with fluorescence mode and large-angle rotation. The results underscore the technique's effectiveness, adaptability and substantial potential in improving the precision of XAS data analysis. While demonstrating significant promise, the method does have limitations related to signal-to-noise ratio sensitivity and the necessity for meticulous angle selection during experimentation. Overall, IBR-AIC represents a significant advancement in XAS, offering a pragmatic solution to Bragg peak contamination challenges, thereby expanding the applications of XAS in understanding complex materials under diverse experimental conditions.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Developing High-Resolution Constrained Variational Analysis of Vertical Velocity and Advective Tendencies within the Range of ARM Scanning Radars at the SGP

Research progress has been made in two areas. One is about the incorporation of the ARM variationally constrained objective analysis method into the WRF GSI data assimilation system. The other is the development of high resolution ARM data and its applications. Specially, we developed a new data assimilation algorithm by adding dynamical constraints to the WRF GSI data assimilation system using hybrid ensemble variational system to derive 3-D fields of atmospheric dynamics and thermodynamics over the ARM SGP sites. We also developed 4x4 km high-resolution constrained variational analysis data over the SGP during the PECAN and made them available to the community. Details are in the attached report.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning in the Context of Laser-Induced Breakdown Spectroscopy

The integration of machine learning (ML) with Laser-Induced Breakdown Spectroscopy (LIBS) has revolutionized the analytical capabilities of LIBS. The combi-nation of both methods enables more accurate and efficient data analysis. While LIBS itself is a powerful technique for elemental analysis, the vast amount of spectral data it generates can be hard to interpret. Machine learning addresses these challenges by leveraging algorithms that can learn from data, identify patterns, and make predictions without explicit programming for the interpretation of each specific task. In LIBS application, ML techniques are used to enhance various analytical processes. For example, ML algorithms can classify materials based on their spectral fingerprints, predict the concentration of elements in a sample, and identify underlying patterns within complex datasets. Here, this application improves the precision of LIBS analyses while significantly reducing the time required for data processing and interpretation. In this chapter, the fundamental concepts of ML will be discussed first. Following this, the process of data splitting and the importance of feature selection will be examined. Several machine learning methods will then be closely examined, exploring how each can benefit LIBS analysis and highlighting their respective advantages and shortcomings. This structured approach will provide a comprehensive understanding of the integration of ML in the context of LIBS analysis.

47 OTHER INSTRUMENTATION↗

Digital Tools for the Preventive Conservation of Built Heritage: The Church of Santa Ana in Seville

Historic Building Information Modelling (HBIM) plays a pivotal role in heritage conservation endeavours, offering a robust framework for digitally documenting existing structures and supporting conservation practices. However, HBIM’s efficacy hinges upon the implementation of case-specific approaches to address the requirements and resources of each individual asset and context. This paper defines a flexible and generalisable workflow that encompasses various aspects (i.e., documentation, surveying, vulnerability assessment) to support risk-informed decision making in heritage management tailored to the peculiar conservation needs of the structure. This methodology includes an initial investigation covering historical data collection, metric and condition surveys and non-destructive testing. The second stage includes Finite Element Method (FEM) modelling and structural analysis. All data generated and processed are managed in a multi-purpose HBIM model. The methodology is tested on a relevant case study, namely, the church of Santa Ana in Seville, chosen for its historical significance, intricacy and susceptibility to seismic action. The defined level of detail of the HBIM model is sufficient to inform the structural analysis, being balanced by a more accurate representation of the alterations, through linked orthophotos and a comprehensive list of alphanumerical parameters. This ensures an adequate level of information, optimising the trade-off between model complexity, investigation time requirements, computational burden and reliability in the decision-making process. Field testing and FEM analysis provide valuable insight into the main sources of vulnerability in the building, including the connection between the tower and nave and the slenderness of the columns.

Chaves, Estefanía↗

TopoSZ: Preserving Topology in Error-Bounded Lossy Compression

Existing error-bounded lossy compression techniques control the pointwise error during compression to guarantee the integrity of the decompressed data. However, they typically do not explicitly preserve the topological features in data. When performing post hoc analysis with decompressed data using topological methods, preserving topology in the compression process to obtain topologically consistent and correct scientific insights is desirable. In this paper, we introduce TopoSZ, an error-bounded lossy compression method that preserves the topological features in 2D and 3D scalar fields. Specifically, we aim to preserve the types and locations of local extrema as well as the level set relations among critical points captured by contour trees in the decompressed data. The main idea is to derive topological constraints from contour-tree-induced segmentation from the data domain, and incorporate such constraints with a customized error-controlled quantization strategy from the SZ compressor (version 1.4). In conclusion, our method allows users to control the pointwise error and the loss of topological features during the compression process with a global error bound and a persistence threshold.

97 MATHEMATICS AND COMPUTING↗

PV-Finder: ML Based Algorithm for Primary Vertex Identification

he CMS detector at the High-Luminosity Large Hadron Collider (HL-LHC) will operate in challenging conditions with expected pile-up of up to 200 collisions per bunch crossing, necessitating the development of a more resilient primary vertex (PV) reconstruction method to ensure the integrity of data analysis and the efficiency of the CMS triggering system. This contribution describes preliminary studies on a new ML based PV-Finder method for PV identification. The method is based on a model trained using Kernel Density Estimations (KDEs) derived from the positions of reconstructed tracks at the beamline, incorporating uncertainties from track parameters. It also utilizes target histograms, modeled as Gaussian distributions centered on the actual ground truth values of specific primary vertices.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

DeepAdversaries: examining the robustness of deep learning models for galaxy morphology classification

With increased adoption of supervised deep learning methods for work with cosmological survey data, the assessment of data perturbation effects (that can naturally occur in the data processing and analysis pipelines) and the development of methods that increase model robustness are increasingly important. In the context of morphological classification of galaxies, we study the effects of perturbations in imaging data. In particular, we examine the consequences of using neural networks when training on baseline data and testing on perturbed data. We consider perturbations associated with two primary sources: (a) increased observational noise as represented by higher levels of Poisson noise and (b) data processing noise incurred by steps such as image compression or telescope errors as represented by one-pixel adversarial attacks. We also test the efficacy of domain adaptation techniques in mitigating the perturbation-driven errors. We use classification accuracy, latent space visualizations, and latent space distance to assess model robustness in the face of these perturbations. For deep learning models without domain adaptation, we find that processing pixel-level errors easily flip the classification into an incorrect class and that higher observational noise makes the model trained on low-noise data unable to classify galaxy morphologies. On the other hand, we show that training with domain adaptation improves model robustness and mitigates the effects of these perturbations, improving the classification accuracy up to 23% on data with higher observational noise. Domain adaptation also increases up to a factor of ${\approx}2.3$ the latent space distance between the baseline and the incorrectly classified one-pixel perturbed image, making the model more robust to inadvertent perturbations. Successful development and implementation of methods that increase model robustness in astronomical survey pipelines will help pave the way for many more uses of deep learning for astronomy.

79 ASTRONOMY AND ASTROPHYSICS↗

Illustrated formalisms for total scattering data: a guide for new practitioners

The total scattering method is the simultaneous study of both the real- and reciprocal-space representations of diffraction data. While conventional Bragg-scattering analysis (employing methods such as Rietveld refinement) provides insight into the average structure of the material, pair distribution function (PDF) analysis allows for a more focused study of the local atomic arrangement of a material. Generically speaking, a PDF is generated by Fourier transforming the total measured reciprocal-space diffraction data (Bragg and diffuse) into a real-space representation. However, the details of the transformation employed and, by consequence, the resultant appearance and weighting of the real-space representation of the system can vary between different research communities. As the worldwide total scattering community continues to grow, these subtle differences in nomenclature and data representation have led to conflicting and confusing descriptions of how the PDF is defined and calculated. This paper provides a consistent derivation of many of these different forms of the PDF and the transformations required to bridge between them. Some general considerations and advice for total scattering practitioners in selecting and defining the appropriate choice of PDF in their own research are presented. This contribution aims to benefit people starting in the field or trying to compare their results with those of other researchers.

36 MATERIALS SCIENCE↗

Micro-architected material design for mechanical response

Rapid advances in additive manufacturing (AM) have enabled the creation of micro-architected materials—also known as mechanical metamaterials—with unprecedented control over fine-scale geometries and arrangements of multiple material constituents. These “materials” can achieve unique and extraordinary effective mechanical properties through their complex architectures rather than composition alone. A key challenge is to design for these bespoke effective mechanical responses within the constraints of available AM techniques (i.e., given a set of desired effective properties), identify a (often nonunique) micro-architecture and selection of material constituents that achieves them. Two main strategies have emerged. Gradient-based methods use sensitivity analysis to iteratively refine candidate designs, while data-driven methods learn micro-architecture-constituent relationships from existing examples to propose new designs. This article reviews these design approaches for micro-architected materials with tailored mechanical responses that can be fabricated by AM as well as their applications.

Spadaccini, Christopher M [Lawrence Livermore Nati↗

Comparison and validation of the QuEChERSER mega-method for determination of per- and polyfluoroalkyl substances in foods by liquid chromatography with high-resolution and triple quadrupole mass spectrometry

Instances of food contamination with per- and polyfluoroalkyl substances (PFAS) continue to occur globally, but sample preparation and analytical methods are quite limited and often monitor for a small percentage of known PFAS. This study aimed to evaluate, validate, and compare performance of two instruments with the recently developed “quick, easy, cheap, effective, rugged, safe, efficient, and robust” (QuEChERSER) sample preparation mega-method – a method developed to monitor chemicals over a broad range of physicochemical properties. Initial evaluation of the QuEChERSER mega-method for determination of PFAS in food demonstrated recoveries, matrix interferences, and co-extractive removal comparable to (or better than) US Food and Drug Administration (FDA) and USDA Food Safety and Inspection Service (FSIS) methods. Subsequent validation of QuEChERSER in beef, catfish, chicken, pork, liquid eggs, and powdered eggs on a high-resolution mass spectrometer achieved acceptable recoveries (70–120%) and precision (RSDs ≤20%) for all 33 target analytes at the 1 and 5 ng g –1 levels and 67–88% of analytes at the 0.1 ng g –1 level, depending on the matrix. Additional validation was performed by tandem mass spectrometry on a triple quadrupole instrument. This approach provided no non-detects and better recoveries at the 0.1 ng g –1 level than the HRMS method but exhibited more variability at 1 and 5 ng g –1 spiking levels. Analysis of NIST SRMs 1946 and 1947 gave accuracies of 70–117%. Furthermore, these results demonstrate the capability of combining PFAS analysis with a mega-method previously validated for 350 analytes, while collecting non-target data for future retrospective analysis of emerging alternatives with a high-resolution mass spectrometry method.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Classifying and analyzing small-angle scattering data using weighted k nearest neighbors machine learning techniques

A consistent challenge for both new and expert practitioners of small-angle scattering (SAS) lies in determining how to analyze the data, given the limited information content of said data and the large number of models that can be employed. Machine learning (ML) methods are powerful tools for classifying data that have found diverse applications in many fields of science. Here, ML methods are applied to the problem of classifying SAS data for the most appropriate model to use for data analysis. The approach employed is built around the method of weighted k nearest neighbors (wKNN), and utilizes a subset of the models implemented in the SasView package (https://www.sasview.org/) for generating a well defined set of training and testing data. The prediction rate of the wKNN method implemented here using a subset of SasView models is reasonably good for many of the models, but has difficulty with others, notably those based on spherical structures. A novel expansion of the wKNN method was also developed, which uses Gaussian processes to produce local surrogate models for the classification, and this significantly improves the classification accuracy. Further, by integrating a stochastic gradient descent method during post-processing, it is possible to leverage the local surrogate model both to classify the SAS data with high accuracy and to predict the structural parameters that best describe the data. The linking of data classification and model fitting has the potential to facilitate the translation of measured data into results for both novice and expert practitioners of SAS.

97 MATHEMATICS AND COMPUTING↗