Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A Co-design Framework for Online Data Analysis and Reduction

Science applications preparing for the exascale era are increasingly exploring in situ computations comprising of simulation-analysis-reduction pipelines coupled in-memory. Efficient composition and execution of such complex pipelines for a target platform is a codesign process that evaluates the impact and tradeoffs of various application- and system-specific parameters. In this article, we describe a toolset for automating performance studies of composed HPC applications that perform online data reduction and analysis. We describe Cheetah, a new framework for composing parametric studies on coupled applications, and Savanna, a runtime engine for orchestrating and executing campaigns of codesign experiments. Furthermore, this toolset facilitates understanding the impact of various factors such as process placement, synchronicity of algorithms, and storage versus compute requirements for online analysis of large data. Ultimately, we aim to create a catalog of performance results that can help scientists understand tradeoffs when designing next-generation simulations that make use of online processing techniques. We illustrate the design of Cheetah and Savanna, and present application examples that use this framework to conduct codesign studies on small clusters as well as leadership class supercomputers.

97 MATHEMATICS AND COMPUTING↗

Multiresolution classification of turbulence features in image data through machine learning

During large-scale simulations, intermediate data products such as image databases have become popular due to their low relative storage cost and fast in-situ analysis. Serving as a form of data reduction, these image databases have become more acceptable to perform data analysis on. In this work, we present an image-space detection and classification system for extracting vortices at multiple scales through wavelet-based filtering. A custom image-space descriptor is used to encode a large variety of vortex-types and a machine learning system is trained for fast classification of vortex regions. By combining a radial-based histogram descriptor, a bag of visual words feature descriptor, and a support vector machine, our results show that we are able to detect and classify vortex features at various sizes at multiple scales. Once trained, our framework enables the fast extraction of vortices on new, unknown image datasets for flow analysis.

97 MATHEMATICS AND COMPUTING↗

Workflows Community Summit 2022: A Roadmap Revolution

Scientific workflows have become integral tools in broad scientific computing use cases. Science discovery is increasingly dependent on workflows to orchestrate large and complex scientific experiments that range from the execution of a cloud-based data preprocessing pipeline to multi-facility instrument-to-edge-to-HPC computational workflows. Given the changing landscape of scientific computing (often referred to as a computing continuum) and the evolving needs of emerging scientific applications, it is paramount that the development of novel scientific workflows and system functionalities seek to increase the efficiency, resilience, and pervasiveness of existing systems and applications. Specifically, the proliferation of machine learning/artificial intelligence (ML/AI) workflows, need for processing large-scale datasets produced by instruments at the edge, intensification of near real-time data processing, support for long-term experiment campaigns, and emergence of quantum computing as an adjunct to HPC, have significantly changed the functional and operational requirements of workflow systems. Workflow systems now need to, for example, support data streams from the edge-to-cloud-to-HPC, enable the management of many small-sized files, allow data reduction while ensuring high accuracy, orchestrate distributed services (workflows, instruments, data movement, provenance, publication, etc.) across computing and user facilities, among others. Further, to accelerate science, it is also necessary that these systems implement specifications/standards and APIs for seamless (horizontal and vertical) integration between systems and applications, as well as enable the publication of workflows and their associated products according to the FAIR principles.

97 MATHEMATICS AND COMPUTING↗

A Digital Twin Framework Utilizing Machine Learning for Robust Predictive Maintenance: Enhancing Tire Health Monitoring

We introduce a novel digital twin (DT) framework for the predictive maintenance of long-term physical systems. Using monitoring tire health as an application, we show how the DT framework can be used to enhance automotive safety and efficiency, and how the technical challenges can be overcome using a three-step approach. First, to manage the data complexity over a long operation span, we employ data reduction techniques to concisely represent physical tires using historical performance and usage data. Relying on these data, for fast real-time prediction, we train a transformer-based model offline on our concise dataset to predict future tire health over time, represented as remaining casing potential (RCP). Based on our architecture, our model quantifies both epistemic and aleatoric uncertainties, providing reliable confidence intervals around predicted RCP. Second, to incorporate real-time data, we update the predictive model in the DT framework, ensuring its accuracy throughout its lifespan with the aid of hybrid modeling and the use of the discrepancy function. Third, to assist decision-making in predictive maintenance, we implement a tire state decision algorithm, which strategically determines the optimal timing for tire replacement based on RCP forecasted by our transformer model. This approach ensures that our DT accurately predicts system health, continually refines its digital representation, and supports predictive maintenance decisions. Furthermore, our framework effectively embodies a physical system, leveraging big data and machine learning (ML) for predictive maintenance, model updates, and decision-making.

advanced computing infrastructure↗

Advanced Image Reconstruction for MCP Detector in Event Mode

A two-step data reduction framework is proposed in this study to reconstruct a radiograph from the data collected with a micro-channel plate (MCP) detector operating under event mode. One clustering algorithm and three neutron event back-tracing models are proposed and evaluated using both example data and a full scan data. The reconstructed radiographs are analyzed, the results of which are used to suggest future development.

Zhang, Chen↗

Bringing Different Views Together: A Hybrid Cooperative Perception Framework for Connected Autonomous Vehicles

Cooperative perception will be essential for connected autonomous vehicles to enhance object recognition and optimize path planning by sending data information about the surrounding environment. However, an inherent challenge in existing systems is the high bandwidth cost of transmitting information in real-time, which restricts cooperative perception’s practicality. Here, this work presents a hybrid cooperative perception fusion framework aimed at mitigating this issue by optimizing data transmission according to available bandwidth or through data reduction techniques. Our methods ensure that vehicles can rapidly transmit high-confidence data without overwhelming the network. Experimental results indicate that our methodology substantially diminishes data transmission sizes while maintaining object detection accuracy. For cooperative perception in autonomous vehicle systems, our approach provides a scalable and effective way to get past the bandwidth barrier.

Carrillo, Dominic [Univ. of North Texas, Denton, T↗

Progressive Tree-Based Compression of Large-Scale Particle Data

Scientific simulations and observations using particles have been creating large datasets that require effective and efficient data reduction to store, transfer, and analyze. However, current approaches either compress only small data well while being inefficient for large data, or handle large data but with insufficient compression. Toward effective and scalable compression/decompression of particle positions, we introduce new kinds of particle hierarchies and corresponding traversal orders that quickly reduce reconstruction error while being fast and low in memory footprint. Our solution to compression of large-scale particle data is a flexible block-based hierarchy that supports progressive, random-access, and error-driven decoding, where error estimation heuristics can be supplied by the user. For low-level node encoding, we introduce new schemes that effectively compress both uniform and densely structured particle distributions. Our proposed methods thus target all three phases of a tree-based particle compression pipeline, namely tree construction, tree traversal, and node encoding. In conclusion, the improved efficacy and flexibility of these methods over existing compressors are demonstrated through extensive experimentation, using a wide range of scientific particle datasets.

97 MATHEMATICS AND COMPUTING↗

Design and Testing of the Endcap Concentrator ASICs for the CMS High-Granularity Calorimeter Upgrade

A major upgrade of the High-Granularity Calorimeter (HGCAL) in the CMS detector is planned for Long-Shutdown 3 (LS3), currently expected to start in 2026. This upgrade will contain over 6 million channels and the electronics will be required to be low power and to withstand a radiation environment with a High-Energy Hadron flux of 3x10**6 particles per square centimeter per second. The solution to this significant challenge is the two Endcap Concentrator (ECON) ASICs, ECON-T and ECON-D, working in tandem with the HGCROC front-end ASIC. The two ECON ASICs provide critical on-detector data reduction for both the 40 MHz trigger path (ECON-T) and 750 kHz data acquisition path (ECON-D) of the HGCAL. The ASICs are fabricated in 65nm CMOS. They are rad-tolerant to 600 Mrad with low power consumption (<2.5 mW/channel). This presentation will be a comprehensive description of each ECON design, including the infrastructure that they share as well as the elements that make each unique. The presentation will also include functionality and radiation tests for both ASICs, and the first high statistics characterization results from the full production of 75k ECON-D and ECON-T ASICs.

Hoff, James R. [Fermilab] (ORCID:0000000163514592)↗

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)↗

Fast Semi-automated Filtration Method for Non-targeted LC-QTOF Data of Aged Nitroplasticizer Samples

A full dataset of aged nitroplasticizer (NP) is composed of more than 2000 unique mass-to-charges (m/z) when combining the non-targeted data obtained from both positive and negative electrospray ionization modes in time-of-flight mass spectrometry. Therefore, manual processing of these data often takes days, weeks, or even months to scrutinize for mechanistic insights. To effectively extract meaningful signals that represent vital degradation intermediates in the early NP degradation mechanism, a semi-automated postprocessing workflow for data filtering, tailored to the aging experiment of NP, has been developed. The automated portion of this workflow is written in a Python code (using pandas, numpy, and matplotlib libraries), which removes more than 65% of potential false signals within seconds via four threshold-based adjustable filters: signal sensitivity, coefficient of variation, number of measurements, and retention time variability. As for the manual portion, a pattern-based inspection method is employed to reduce another 23% or more false positives, which greatly simplifies data visualization and results in less than 3% of potential candidate m/z needing in-depth data interpretation. As a positive control, known compounds are verified. Using this semi-automated data reduction method, the amount of time required is reduced to a matter of hours for data filtering in the non-targeted datasets of aged NP, which saves more time and effort for compound identification.

36 MATERIALS SCIENCE↗

Efficient Data Query for Gaussian Process Compressed Data through Value Range Estimation [Slides]

When the resolution of the data increases, data reduction methods are applied to simulation output, including Gaussian process, neural representation and compression algorithms. Lots of data analysis/visualization techniques requires data query, but data query from reduced representation is still challenging. This report will provide examples and provide possible answers to why data query from reduced representation is still challenging.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Archetype-based Redshift Estimation for the Dark Energy Spectroscopic Instrument Survey

We present a computationally efficient galaxy archetype-based redshift estimation and spectral classification method for the Dark Energy Survey Instrument (DESI) survey. The DESI survey currently relies on a redshift fitter and spectral classifier using a linear combination of principal component analysis–derived templates, which is very efficient in processing large volumes of DESI spectra within a short time frame. However, this method occasionally yields unphysical model fits for galaxies and fails to adequately absorb calibration errors that may still be occasionally visible in the reduced spectra. Our proposed approach improves upon this existing method by refitting the spectra with carefully generated physical galaxy archetypes combined with additional terms designed to absorb data reduction defects and provide more physical models to the DESI spectra. We test our method on an extensive data set derived from the survey validation (SV) and Year 1 (Y1) data of DESI. Our findings indicate that the new method delivers marginally better redshift success for SV tiles while reducing catastrophic redshift failure by 10%–30%. At the same time, results from millions of targets from the main survey show that our model has relatively higher redshift success and purity rates (0.5%–0.8% higher) for galaxy targets while having similar success for QSOs. These improvements also demonstrate that the main DESI redshift pipeline is generally robust. Additionally, it reduces the false-positive redshift estimation by 5%–40% for sky fibers. We also discuss the generic nature of our method and how it can be extended to other large spectroscopic surveys, along with possible future improvements.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Film condensation with high heat fluxes and scaled experiments using pure steam for reactor containment cooling

Condensation tests were performed using a newly developed test facility for scaling the passive containment cooling system (PCCS) to a small modular reactor (SMR). The PCCS of the SMR plays a pivotal role in ensuring greater safety, reliability, and compactness than what is afforded by traditional reactors. Therefore, a well-designed PCCS is essential to SMRs. However, previous studies and test data were unsuitable for scaling, due to high variation in the test geometry and operating conditions. This study intends to close this research gap by using a novel designed scaled test facility consisting of vertical condensing test sections featuring 1-, 2-, and 4-inch-diameter condensing tubes with annular water cooling, and by applying superheated and saturated steam with different steam mass flow ranges of 5–25 g/s. Further, the primary test data, including axial temperatures, mass flow rates, and pressures, were used in conjunction with a standard data reduction method to estimate critical parameters such as heat fluxes, heat transfer coefficients, and condensation rates. These scaled test data would support improving empirical correlations and validating condensation models to identify scaling distortion for SMR PCCSs.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Big Data Meets Geothermal Exploration (CRADA Final Report)

As part of the Cyclotron Road program, Zanskar Geothermal & Minerals, Inc. investigated the application of micro-earthquake and ambient noise seismology methods to imaging and characterizing the structural characteristics and hydrothermal flux of subsurface faults. Significant advances in what could be resolved were enabled by two major developments in seismology: 1) the availability of large-n arrays of low-cost seismometers, and 2) the availability of increased computational power and semi-automated data reduction algorithms. In tandem, these advances may improve the signal-to-noise ratio and spatial precision of the data collected and enable higher-resolution characterization of subsurface fracture systems and their spatio-temporal evolution. These tools supported efforts to reduce dry-hole risk and to improve wellfield productivity for geothermal resource development. In particular, two applications of these advances were evaluated: 1) fracture-seismic imaging, which was used to detect ambient emissions from fluid-filled fractures, and 2) reservoir tomography, which used information about travel paths, source locations, and source parameters of micro-earthquakes to identify areas of enhanced permeability. Integration of these methods provided guidance for siting wells and served as prior constraints for reservoir models, informing forecasts of power potential and production and injection strategies aimed at minimizing temperature decline and improving overall resource productivity.

15 GEOTHERMAL ENERGY↗

TensorID v1.0

This Python software package includes new and efficient algorithms for satellite and core interpolative decomposition of tensor data. In general, these algorithms target high-dimensional data reduction and compression. The software is purely numerical and can be applied by others to many important sources of tensor data generated by computation or experiment.

Zhang, Yifan [Lawrence Berkeley National Laborator↗

Towards Autonomous Experiments by Connecting High Performance Microscopy with High Performance Computing

The digitization of controls, data, and analysis in microscopy is bringing the idea of autonomous microscopes closer to reality than ever before. Automated transmission electron microscopy (TEM) is already fairly routine for some experiments the only require simple repetitive tasks such as imaging biological macromolecules for single particle cryoEM [1], tilt series for electron tomography [2], and movies for crystallography [3]. The vast majority of TEM experiments are conducted completely by human operators who choose the regions of interest, optimize experimental parameters, and make decisions about data quality visually during an experiment. The field is still a long way from having completely autonomous TEMs that can adapt to sample difficulties and tune experimental parameters based on data quality and desired experimental outcomes. Part of the issue is the lack of capability for feeding information learned from on-line, live data analysis back into the on-going experiment [4]. Furthermore, this presentation will discuss current capabilities for large scale data reduction and analysis using high performance computing (i.e. supercomputing) and progress towards developing a true feed-back loop that places data analysis and theory in the experimental loop.

97 MATHEMATICS AND COMPUTING↗

Methods for Incorporating Model Uncertainty into Exoplanet Atmospheric Analysis

A key goal of exoplanet spectroscopy is to measure atmospheric properties, such as abundances of chemical species, in order to connect them to our understanding of atmospheric physics and planet formation. In this new era of high-quality JWST data, it is paramount that these measurement methods are robust. When comparing atmospheric models to observations, multiple candidate models may produce reasonable fits to the data. Typically, conclusions are reached by selecting the best-performing model according to some metric. This ignores model uncertainty in favor of specific model assumptions, potentially leading to measured atmospheric properties that are overconfident and/or incorrect. In this paper, we compare three ensemble methods for addressing model uncertainty by combining posterior distributions from multiple analyses: Bayesian model averaging, a variant of Bayesian model averaging using leave-one-out predictive densities, and stacking of predictive distributions. We demonstrate these methods by fitting the Hubble Space Telescope (HST) + Spitzer transmission spectrum of the hot Jupiter HD 209458b using models with different cloud and haze prescriptions. All of our ensemble methods lead to uncertainties on retrieved parameters that are larger but more realistic and consistent with physical and chemical expectations. Since they have not typically accounted for model uncertainty, uncertainties of retrieved parameters from HST spectra have likely been underreported. We recommend stacking as the most robust model combination method. Our methods can be used to combine results from independent retrieval codes and from different models within one code. They are also widely applicable to other exoplanet analysis processes, such as combining results from different data reductions.

79 ASTRONOMY AND ASTROPHYSICS↗